ICML paper: AI evaluation should score human-AI teams, not lone autonomy
A new ICML 2026 position paper claims autonomy-first benchmarks are steering AI development wrong.
A new position paper accepted to the ICML 2026 Position Paper Track argues that the AI industry's obsession with autonomous, superhuman benchmarks is actively steering development in the wrong direction. In 'AI Evaluation Should Work With Humans,' authors Jan Kulveit, Gavin Leech, Tomáš Gavenčiak, and Raymond Douglas challenge the dominant evaluation paradigm that implicitly targets replacing humans. Instead, they contend, the field should be scoring the performance of human-AI teams. Traditional metrics that highlight standalone model capabilities—whether coding, reasoning, or general knowledge—encourage systems optimized for solo operation. The researchers believe this misaligns incentives: developers chase independent mastery rather than complementing human judgment and context. The paper was posted to arXiv with identifier 2608.13577 and published July 6, 2026.
The proposed pivot to team-based evaluation would treat AI as a teammate, not a substitute. That means measuring joint outcomes: human experts working with AI assistance, faster and safer task completion, and better handling of ambiguous, high-stakes decisions. According to the authors, this collaborative framing rewards systems that are transparent, controllable, and responsive to human feedback—properties that standalone superhuman performance often ignores. For enterprises and policymakers, the implications are significant: procurement and red-teaming would shift from 'can it do the job alone?' to 'does it make humans better at the job?' The paper stops short of prescribing specific benchmarks, but it urges the AI community to reorient evaluation around human-AI complementarity as a route to far better societal outcomes. As AI adoption accelerates, this position paper provides a timely counterweight to the prevailing 'replace the human' narrative.
- Accepted to the ICML 2026 Position Paper Track and published as arXiv:2608.13577
- Argues that autonomous superhuman benchmarks misguide AI development toward replacing humans
- Proposes evaluating human-AI team performance to build complementary, controllable AI systems
Why It Matters
Could change how enterprises benchmark and buy AI—favoring tools that amplify human teams over those that replace workers.