AI crosses expert baselines every 7 months, but frontier remains jagged: study
Between 2023-2026, AI surpassed humans on grad-level tasks but still lags in reliability.
A comprehensive study from researchers at the University of Central Florida and other institutions, titled 'Faster AI, Uneven Frontier' and published on arXiv in July 2026, documents that between 2023 and 2026, frontier AI systems crossed documented human expert baselines on a growing set of bounded cognitive tasks. These include graduate-level science questions (e.g., GPQA), competition-level mathematics (e.g., MATH), software engineering benchmarks (e.g., SWE-bench), and structured diagnostic reasoning. The paper notes a striking trend: the length of tasks that these systems can complete at 50% reliability has doubled roughly every seven months, indicating rapid scaling of autonomous capability. Yet the authors caution that benchmark scores often overstate deployed performance due to contamination, construct validity issues, vendor self-evaluation, and the gap between 50% reliability and the near-perfect reliability required for economic work.
The 'jagged frontier' concept captures the unevenness: humans still decisively outperform AI on long-horizon reliability (sustaining performance over hours or days), genuinely novel problems (out-of-distribution reasoning), calibrated self-knowledge (knowing what they don't know), sample-efficient learning (learning from few examples), and embodied action (physical tasks). The paper also reviews evidence on cognitive offloading—the use of AI as a mental extension—noting that early field data suggests costs to unaided skill, though meta-analytic evidence on prior technologies points the other way, leaving the question open for generative AI. Crucially, experiments on human-AI collaboration show that naive combination often underperforms the stronger partner, implying that human judgment must be repositioned toward specification (framing problems correctly), verification (checking outputs for correctness), and oversight (monitoring for alignment and safety). The paper's central thesis: humans must redesign their role rather than defend it against automation.
- AI crossed expert baselines on graduate science, competition math, and software engineering benchmarks between 2023-2026.
- The length of tasks AI can complete at 50% reliability doubles every ~7 months, but economic reliability requires far higher consistency.
- Human role must shift from execution to specification, verification, and oversight—experiments show naive human-AI collaboration underperforms the better partner.
Why It Matters
Professionals must pivot from doing the work to specifying, verifying, and overseeing AI—a fundamental shift in human judgment's role.