The AI Grading Other AI Is Often Wrong — Here's the Fix
Bad AI report cards mean worse apps, wasted money, and shaky decisions you rely on.
More and more, companies don't hire people to grade their AI. They hire another AI. This "AI judge" reads thousands of chatbot answers and scores them quickly and cheaply. It's how many teams decide whether a new version of their AI is better than the old one, and whether it's safe enough to release.
The problem is checking the grader. To trust an AI judge, you need to know how often it agrees with real humans. But budgets are tight, so most answers get graded by only one person — or none. This paper shows that thin human coverage is the single biggest reason companies make the wrong call. At just 5% overlap (one in twenty answers independently checked by a human), wrong decisions hit 25%, and the odds of crowning the wrong winner among ten candidate judges are 65%.
The researchers offer two fixes. First, quantity: roughly one in four answers should be double-checked by a human — enough for clear-cut judges, though genuinely borderline cases stay hard. Second, a free tweak to which answers get checked. Instead of picking at random, group similar items together and sample smartly, which halves false rejections without adding a cent to the budget. They confirmed this across ten AI judges and four evaluation setups, covering image assessment, reasoning, and summarization tasks.
The takeaway for you: AI judges sit behind hiring screens, content moderation, customer service bots, and the "which model should we use" decisions inside tools you already use. If those judgments are unreliable, the products get worse — and nobody notices until users complain. If a company can't tell you how its AI was validated, treat its claims of quality with healthy skepticism.
- Companies increasingly let one AI grade another AI's work — a cheap shortcut that can quietly produce bad decisions.
- When humans independently check only 5% of answers, there's a 65% chance of picking the wrong 'best' AI judge from a group of ten.
- Checking about one in four answers, and choosing which ones smartly rather than randomly, roughly halves the errors at no extra cost.
Why It Matters
Unreliable AI graders can pick worse tools behind your apps — and you get the sloppy results.