Research & Papers

GPT-OSS, DeepSeek-V4 Flash, Gemma-4 match frontier judges at 100x lower cost

Cheap open-weight models now grade math proofs as reliably as Claude Opus 4.7 for pennies.

Deep Dive

A new paper by Benjamin Grayzel (arXiv:2608.00004) tackles the sky-high cost of grading natural-language mathematical proofs in AI evaluation. Frontier models like Claude Opus 4.7 and Gemini 3.1 Pro are reliable judges, but their inference costs quickly balloon when scoring large benchmark suites. Grayzel tested whether cheap open-weight LLMs—GPT-OSS 120B, DeepSeek-V4 Flash, and Gemma-4 31B—could serve as reliable substitutes when given a candidate proof, a ground-truth proof, and a human-grading rubric.

On a 200-instance validation sample of IMO-GradingBench, all three cheap models agreed with human pass/fail decisions at rates statistically indistinguishable from the frontier models, while costing up to 100x less. Extending to the full 1,000-instance benchmark, Grayzel found that a unanimous all-three-pass rule—where all three models must pass a proof—yielded the highest pass-agreement and precision, with the smallest run-to-run variance across four replicates. The paper recommends this rule as a deployable default, though it was identified post-hoc and needs independent replication.

Key Points
  • GPT-OSS 120B, DeepSeek-V4 Flash, and Gemma-4 31B match Claude Opus 4.7 and Gemini 3.1 Pro on IMO-GradingBench pass/fail agreement.
  • Cheap judges deliver up to 100x lower cost than frontier models without statistical loss in accuracy.
  • All-three-pass unanimous consensus maximizes precision and stability, making it the recommended default for deployment.

Why It Matters

Frontier-grade math proof grading is now accessible to any lab, slashing evaluation costs and democratizing AI reasoning benchmarks.

📬 Get the top 10 AI stories daily