Stanford's EconEvals beats OpenAI's GDPval at 500x lower cost
New benchmark covers 47% of US occupations, but Claude usage lags far behind potential.
Researchers from Stanford, led by Alexander Wan, Percy Liang, and Rishi Bommasani, released EconEvals, an open-source evaluation suite designed to assess language models' performance on economically valuable tasks across the US labor market. Grounded in real user queries and synthetic data, EconEvals offers dramatically better coverage than OpenAI's GDPval—which only covers 5% of US occupations—at a 500x lower cost. The suite also introduces a simulation-based exposure measure that estimates potential time savings per occupation.
According to the paper, current language models could save workers substantial time on at least half of their tasks in 47% of all US occupations. However, actual usage of Claude (Anthropic's model) for those tasks is far lower, indicating a gap between potential and adoption. The researchers identify privacy concerns and proprietary systems as the primary bottlenecks preventing further time savings, rather than inherent model limitations. This infrastructure can be continuously updated as capabilities evolve, providing a more grounded way to predict AI's labor-market impact.
- EconEvals covers all US occupations, improving on OpenAI's GDPval which only covered 5% at 500x lower cost.
- Current models could save substantial time on half of tasks in 47% of occupations, per simulation-based exposure estimates.
- Observed Claude usage is low for high-potential tasks; privacy and proprietary systems are key bottlenecks.
Why It Matters
Accurate benchmarks reveal where AI truly helps workers, highlighting adoption gaps and privacy hurdles.