GAUGE benchmark reveals AI agents flunk financial valuation tests
Best AI agent scores 53.4 vs senior analysts' 88.3 on 196-task financial modeling benchmark
A team led by Jiacheng Lu has released GAUGE, a benchmark designed to evaluate how well AI agents build financial valuation models compared to human analysts. Traditional benchmarks grade outputs against a single expert answer, but financial models often have multiple reasonable answers — analysts disagree on implied price within 10% even for the same company. GAUGE solves this by using 1,001 vendor-classified analyst workbooks and a 196-task evaluation set, with a three-layer observed-practice envelope covering 56 auditable facets, eight validity gates, and deterministic structural checks. This allows it to grade agents on both mechanical model construction and subjective valuation judgment.
The benchmark was validated with a 55-participant known-groups study: senior analysts average 88.3 on the failure-aware score φ₀, junior analysts 66.0, and finance students 43.2. Across 24 agents and 1,011 scored generations, the best agent scored 53.4 — slightly above student average but below every senior and most juniors. Agents pass 93% of mechanical facets (e.g., using correct formulas) but only 78% of judgment facets (e.g., discount rate selection), with a fleet-median gap of 26 points. The researchers conclude that current agents are substantially stronger at model construction than valuation judgment. GAUGE’s methodology, a gated de-identified data tier, controlled training split, versioned 48-task evaluation core, and withheld refresh pool are being publicly released.
- GAUGE uses 1,001 analyst workbooks and 196 tasks, with 56 auditable facets and eight validity gates to avoid single-answer bias
- Senior analysts average 88.3 on the failure-aware score, juniors 66.0, students 43.2 — best agent scores just 53.4
- Agents pass 93% of mechanical facets but only 78% of judgment facets, highlighting a key limitation in current AI financial modeling
Why It Matters
As AI takes on financial analysis, GAUGE provides the first rigorous benchmark to measure agent judgment vs. human expertise.