New benchmark reveals AI's 'augment vs automate' dilemma
7 real-world tasks prove automation wins are misleading for human-AI teams
A team of researchers from UC Berkeley's Haas School of Business and Department of Economics has challenged the AI benchmarking status quo with CentaurBench, a new framework that evaluates large language models (LLMs) not just on their ability to automate tasks, but on their capability to augment human or weaker agent performance.
The researchers tested seven real-world work scenarios ranging from customer support to legal document analysis, using a standardized lower-capacity worker model as the baseline. In automation mode, the assistant LLM produced outputs directly, while in augmentation mode it provided guidance to the worker model. Surprisingly, rankings between the two regimes showed only modest correlation - the top automation model lost in augmentation scenarios on five of seven tasks. Even more striking, the unaided worker model outperformed every AI-assisted condition in three tasks, and only one model's guidance consistently beat no guidance across all scenarios. The findings suggest that current benchmark practices may be optimizing for the wrong metric.
The paper proposes that future benchmarks must evaluate models based on their specific roles in human-AI and multi-agent systems, rather than solely on autonomous performance metrics.
- CentaurBench evaluates both automation and augmentation capabilities across 7 real-world tasks using standardized worker models
- Automation leaders lost augmentation scenarios 5 out of 7 times; unaided workers outperformed AI assistance in 3 tasks
- LLM judge panels with task-specific rubrics scored outputs through blind pairwise comparisons across 10 experimental runs
Why It Matters
Benchmarking must shift from pure automation to human-AI collaboration effectiveness for real-world productivity gains