Research & Papers

AI Leaderboards Are Rigged by Test Setup — One Model Scored 31% to 89%

The 'best' AI might just be the one that got the easiest test.

Deep Dive

Think of AI leaderboards like race results. You assume every runner ran the same track. But a new study shows that's not true. Researchers tested 12 AI models on the same 3,679 questions across 26 slightly different test setups. The only thing that changed was the 'harness' — the way the test is given: the order of multiple-choice options, the wording of instructions, and how a model's answer is counted. The result shocked them: one model scored between 31% and 89% depending purely on those setup choices.

That's like a sprinter finishing anywhere from last place to first depending on which lane they ran in. Even worse, four out of the 12 models could reach the #1 rank somewhere in those 26 setups. So the 'winner' of a leaderboard is often chosen by the test's own invisible preferences, not by the model's actual ability. The study found that the biggest factor wasn't option order — it was how the answers were scored. That's a detail most test designers consider minor, but it turns out to be the one that decides the winner.

Why should you care? Because people use these rankings to pick which AI to trust with homework, emails, medical questions, or even code. Companies use them to decide which model to build into apps. If a leaderboard can flip winners by random test tweaks, then 'this AI is #1' is more like a coin flip than a fact. The researchers released all their data and scripts so anyone can re-run the test and see for themselves.

The catch: the study only looked at open-weight AI models, not all the big commercial ones like ChatGPT or Gemini. And it doesn't mean all AI tests are useless — it means we need a standardized, honest testing method. Until then, when you see an AI claiming to be 'best in class,' ask: best at what, and under whose rules?

Key Points
  • One AI model scored between 31% and 89% on the same questions — the only change was test setup.
  • Four of the 12 models tested could be ranked #1 just by tweaking the test format.
  • How answers are scored matters more than answer order, the study found.
  • The researchers released all data so anyone can verify the results.

Why It Matters

Don't trust 'best AI' claims blindly — rankings can be manipulated by hidden test-settings choices.

📬 Get the top 10 AI stories daily