Research & Papers

Can AI Really Judge Your Work? New Study Reveals Surprising Flaws

AI reviewers missed 78% of errors and gave overly generous scores compared to humans

Deep Dive

AI is increasingly being used to review academic papers, grant proposals, and even job applications. But can these AI tools actually scrutinize work critically? A new study tested two advanced AI models, Qwen2.5 and Pixtral, by having them review 165 academic submissions for a major 2026 conference. The AI reviewers were given papers with hidden errors—some as obvious as missing data or incorrect calculations—and asked to evaluate them.

The results were eye-opening. On average, AI reviewers gave scores between 7.0 and 8.1 out of 10, far higher than human reviewers, who scored papers between 3.4 and 6.8. Worse, the AI models only spotted 12% of the inserted errors under normal conditions. Even when researchers added a simple verification instruction, the detection rate only rose to 22%. That means 78% of the errors went unnoticed—a worrying sign for anyone relying on AI for fair or accurate reviews.

Surprisingly, adding figures to papers made the AI reviewers *less* likely to catch errors, while also inflating their scores. The AI models also failed to reliably spot visual errors when comparing text to figures. Even stranger, half the AI-generated reviews described figures that weren’t even in the papers. Author identities—like prestigious universities or unknown names—had no effect on the AI’s scoring or error detection.

Perhaps most concerning, the AI’s editorial decisions (accept/reject) matched what you’d get by simply averaging scores, with no sign of deeper analysis. In other words, the AI didn’t *understand* the papers—it just followed a basic formula.

Key Points
  • AI reviewers scored papers 20–30% higher than humans and missed 78% of inserted errors
  • Adding figures or verification prompts only slightly improved error detection (from 12% to 22%)
  • AI decisions matched simple score averages—suggesting no advanced critical thinking

Why It Matters

AI may be speeding up reviews, but this study shows it can’t be trusted alone for fair or accurate evaluations

📬 Get the top 10 AI stories daily