Research & Papers

The Shortcut Companies Use to Grade AI Is Quietly Wrong

⚡The way AI graders are checked for fairness may be misleading you.

Deep Dive

A growing number of companies use one AI to grade another AI's answers — checking which chatbot reply is better, which summary is more accurate, which résumé screener works best. It's far cheaper and faster than paying humans. To cut costs further, some evaluation tools skip the AI's written reasoning and just peek at the very first word it would say: 'A' or 'B'. A team of researchers tested that shortcut, and found it frequently isn't reading a judgment at all.

For three of the AI judges tested, between 12% and 49% of the time the model didn't start with a verdict. When the researchers forced a read anyway, the tool reported whichever answer happened to be shown first. The proof: when they simply swapped the order of the two answers, the 'verdict' flipped 89.7% of the time. A real judgment wouldn't care about order — this one clearly did.

The distortion is sneaky because it hides where it hurts most. It inflated 'position bias' — an AI's tendency to favor whichever answer comes first — by about 42 points. Yet accuracy scores barely budged. That means a company testing whether its AI grader is fair could see a wrong answer, while ordinary users notice nothing at all. The people misled aren't the ones using the AI; they're the ones auditing it.

The researchers' fix is refreshingly cheap: report how often a judge actually leads with a verdict. It takes one quick calculation, no human labels required, and should sit next to any fairness number. The takeaway for you: when a company claims its AI grader 'matches human judgment' or 'shows no bias,' it's worth asking how they measured that — and whether they looked past the first word.

Key Points
  • Many companies use AI to grade other AI, and some do it by reading only the first word the grader would say.
  • The shortcut overstated bias in every test: swapping the two answers flipped the 'winner' 89.7% of the time.
  • Accuracy scores looked fine while fairness numbers were wrong — so quality checks can pass while audits mislead.

Why It Matters

If you rely on AI to score résumés, essays, or support chats, its fairness checks may be unreliable.

📬 Get the top 10 AI stories daily