Research & Papers

AI Test Scores Can Be Off by 2,600x, Study Finds

The AI scores you see on leaderboards may be guesswork dressed up as math.

Deep Dive

AI companies love a certain kind of number: give the model enough tries, and it solves almost any problem. This is measured by something called pass@k — the chance an AI gets a problem right if you let it attempt it k times. Testing is expensive, so labs usually collect only 10 or 16 tries per problem, then do math to project what would happen at 100 or 1,000 tries. Two statisticians, Pranav Singh and Prashant Singh, asked a simple question: does that projection actually mean anything?

Their answer is uncomfortable. If you collect n tries per problem, you can only directly measure performance up to k = n. Beyond that, many completely different 'worlds' fit your data equally well but predict wildly different outcomes. On a public dataset where every problem was tried 10,000 times, they pretended they only had 16 tries and asked what the failure rate would be at 1,000 tries. The honest answer was a range so wide — from 1.5 times to over 2,600 times different — that the number is basically unusable.

Why should you care? Because AI buying decisions, investment theses, news headlines, and even regulation increasingly lean on these numbers. If a company says its model 'solves 90% of coding problems when given enough attempts,' that figure may be a projection, not an observation. You could be paying for capability that was never actually demonstrated, or dismissing a tool that's better than its score suggests.

The authors are careful: they do not claim AI performance stops improving with more tries, or that scaling laws are wrong. They simply provide the honest baseline that any such claim has to beat. They also propose a clearer reporting standard that separates what was measured, what is a range, and what is a model-based forecast. This is a statistics paper, not a product launch, so nothing changes tomorrow — but it's a quiet warning shot at an industry built on impressive-looking scores.

Key Points
  • If an AI is tested with only a few tries per problem, any claim about how it does with hundreds of tries is extrapolation — often pure guesswork.
  • On a public dataset, the estimated failure rate at 1,000 attempts varied by factors ranging from 1.5x to over 2,600x, depending on assumptions nobody can verify.
  • The authors propose a reporting standard that clearly separates measured results, honest ranges, and model-based forecasts.

Why It Matters

AI hype numbers may overstate real ability, affecting what you buy, trust, or invest in.

📬 Get the top 10 AI stories daily