FinProBench grades financial AI agents with rubrics from real pro deliverables
New benchmark uses 1,723 real financial deliverables to evaluate AI agents—humans still edge out AI.
Evaluating financial AI agents has been hampered by rubrics that ignore the tacit standards visible in real professional deliverables. To fix this, researchers from academia and industry introduce FinProBench, a benchmark paired with Role-Grounded Rubric Construction (RGRC)—a four-stage pipeline that collects deliverables, extracts competencies, synthesizes rubrics, and validates them. The benchmark spans 1,723 curated deliverables across 57 occupations, 8 financial sub-industries, and 161 deliverable types, with an initial evaluation set of 20 complete tasks covering 20 roles in 7 sub-industries.
Results show a clear split: for conventional roles where model priors are strong, prompt-only rubrics nearly match RGRC (89.2% vs 90.7%), but for role-specialized roles RGRC dominates (99.1% vs 78.0%). When tested with heterogeneous LLM judges, human deliverables rank first on average (73.7 out of 100) versus four AI systems (70.3, 70.2, 69.6), though confidence intervals overlap, indicating complementary strengths. Reusing role-level rubrics reduces per-task construction effort 6.7x compared to building each from scratch.
- FinProBench includes 1,723 deliverables from 57 occupations, 8 sub-industries, and 161 deliverable types
- RGRC rubrics outperform prompt-only rubrics on role-specialized tasks: 99.1% vs 78.0%
- Human experts still beat all AI systems on average (73.7 vs ~70.0), with overlapping confidence intervals
Why It Matters
For finance teams deploying AI, this benchmark reveals where generic prompts fail and why role-specific evaluation is critical.