TUA-Bench tests terminal agents on 120 real-world tasks
New benchmark reveals terminal agents still fall short outside coding tasks
A team of researchers led by Shoufa Chen has released TUA-Bench, a new benchmark designed to evaluate general-purpose terminal-use agents (TUAs). Unlike prior terminal benchmarks that focus narrowly on programming and system administration, TUA-Bench includes 120 manually crafted real-world tasks across five families: routine digital activities (document editing, email management, live-web information seeking) and scientific/engineering workflows co-designed with PhD experts (e.g., using specialized software). Each task runs in a real terminal with deterministic setup and execution-based scoring.
The benchmark's results reveal that even the strongest frontier agent — Claude Code with Claude Opus 4.8 at max reasoning effort — achieves only 65.8% overall performance, leaving substantial room for improvement across both routine and specialized tracks. TUA-Bench exposes the gap between task-specific coding assistants and truly general-purpose agents that can reliably handle diverse digital environments. By providing a broad, realistic evaluation, the benchmark aims to push the field toward agents that can operate beyond the narrow scope of existing shell-oriented or GUI-focused benchmarks.
- 120 real-world tasks across 5 families including email, document editing, and scientific workflows
- Claude Code with Claude Opus 4.8 max reasoning effort achieves 65.8% overall performance
- Prior benchmarks focused on programming or GUI, leaving a gap for general terminal use
Why It Matters
As AI agents expand beyond coding, TUA-Bench provides a realistic test for terminal automation in daily workflows.