Research & Papers

AI Code Reviewers Are Bluffing — Now They Can Admit It

⚡When one AI grades another AI's code, it often just guesses.

Deep Dive

Companies increasingly use one AI to check work done by another — including computer code. The researchers tested a published system called MARCH that breaks a judgment into smaller, checkable pieces, with several AI helpers each checking one claim (multi-agent verification). Across 80 measurements on two code-judging tests, it declared both programs equally good 78 to 95% of the time. Its accuracy was 4.4%. The exact same AI, simply asked 'which program is better?', scored 43.7%.

Why does the fancy version fail? The method works when the evidence is a stack of retrieved documents — those exist separately from the answer and differ between the two candidates. With code, that second condition collapses: the code is the thing being judged, so the helpers have nothing independent to compare. Yet the system still returns a confident verdict with reasoning attached, impossible to tell apart from a verdict it actually had grounds for. Making the problems easier, or swapping in a bigger AI as the judge, did not help.

The team's real contribution isn't a smarter judge — it's a way to catch a judge that has no basis for its answer. Two measurements pulled from the pipeline's own logs explain the failure without needing any correct answers to compare against (a big deal, since labels are expensive). Gating on one of those measurements lets the system decline comparisons it cannot ground. Accuracy jumps from 20.7% to 36.9%, while it still answers about half of all comparisons.

So what does this mean for you? If you rely on AI to check AI's output — reviewing code, drafts, contracts, spreadsheets — a confident answer is not proof of a correct one. The most valuable upgrade is an AI that says 'I'm not sure.' Expect tools that flag their own uncertainty, which could keep you from shipping or signing off on work that was never actually verified.

Key Points
  • One AI judging another AI's code called it a tie 78 to 95% of the time and got only 4.4% of comparisons right.
  • The same AI asked directly scored 43.7%, so the extra multi-step setup actually made it worse, not better.
  • Two signals from the AI's own logs let it skip comparisons it can't judge, lifting accuracy from 20.7% to 36.9% while still answering half.

Why It Matters

If AI checks AI's work, it must say 'I'm not sure' rather than confidently guess.

📬 Get the top 10 AI stories daily