The AI That Grades Other AI Got It Wrong 77% of the Time
A cheaper replacement was 300x less expensive — and far more accurate.
Many companies now build software that turns plain-English questions into database queries — you type "how many customers cancelled last month?" and the system writes the code. But someone has to check whether the answer is right. Increasingly, that someone is another AI, called a "judge" model. The authors of this paper checked whether their judge was actually any good. It was not.
The judge in question was gpt-4o-mini, a small, cheap OpenAI model. Measured against two human experts, it scored a kappa of 0.04 on a hard test set — kappa is a agreement score where 0 means random guessing and 1 means perfect agreement. In plain terms, it was nearly useless. It flagged 77.1% of the answers humans judged perfectly correct as problems, usually because it imagined errors that weren't there — a failure mode the authors named "grade-hallucination."
The fix was surprisingly mundane. A self-hosted open model, Qwen3.6-27B, scored 0.72 — roughly the same as Anthropic's top-tier Claude Opus 4.7 at 0.71 — while costing about 1/300th as much per check. That's the difference between paying a dollar and a third of a cent for the same verdict. The authors also found that mixing a weak judge with a strong one makes things worse, not better, like averaging a good student's answers with a bad one's. Three strong judges that must agree unanimously reached 0.79 agreement while automatically handling 89.7% of cases without human review.
When they ran the same audit on a different, expert-built dataset called BIRD-financial, it flagged 25.5% of the human-written "correct" answers as questionable. Either those datasets have more errors than anyone admitted, or the auditing method is too trigger-happy. Both possibilities matter, because these datasets are how the industry decides which AI systems are good.
- A cheap AI 'referee' approved the wrong answers and rejected correct ones — flagging 77.1% of good answers as broken.
- A self-hosted model (Qwen3.6-27B) matched top-tier Claude Opus accuracy at roughly 1/300th the cost per check.
- Combining weak and strong AI judges made results worse, but three strong judges agreeing unanimously hit 89.7% automatic coverage.
Why It Matters
If companies trust unchecked AI referees, bad data and wrong answers slip through quietly — and you pay for it.