Research & Papers

AI Quiz Leaks Inflate Scores but Don't Change Winners, Study Finds

Worried AI cheats on its exam? New research shows the ranking is still fair.

Deep Dive

Think of AI benchmarks like a pop quiz for models such as ChatGPT. But what if the answer key accidentally leaked into what the AI already studied? That would be like a student who had seen the test before. In AI, this is called "contamination." For years, people worried it might make AI leaderboards unreliable. A new paper from researchers at Stanford and the City University of Macau pours cold water on that fear.

The researchers built a clever way to check for cheating. Instead of just giving the AI the original quiz, they also gave it paraphrased questions, which ask the exact same thing but use different words. If an AI memorized the original answer, it would do worse when the wording changed. If it truly understood the subject, its score would stay similar. This let them measure contamination instead of just guessing.

They tested 47 well-known AI models against this system, and the results were surprisingly reassuring. Yes, contamination inflated raw scores, pushing them up by about 0.19 points on average. But this inflation was nearly the same across every model, which means it acted like a small blanket bonus rather than an unfair advantage. The official rankings matched the "paraphrase-proof" rankings almost perfectly, with a 99.7% agreement rate. Only 3 out of 188 individual cases showed contamination strong enough to shift a model's position in the ranks.

So what should you actually take away? When people compare AIs by saying "this one scores 92%," that absolute number might be overstated. But the phrase "this one is better than that one" is generally trustworthy. The researchers suggest leaderboards show both standard scores and paraphrase-controlled rankings, plus a range of uncertainty, so users can see which results are rock solid. For most of us, this means we can keep picking the top-rated AI without panicking over hidden test leaks.

Key Points
  • AI 'contamination' means test questions leaked into the AI's training data, like a student finding the exam key.
  • Scores inflated by about 0.19 points on average, but because it happened to all models equally, rankings barely changed.
  • Leaderboards matched a paraphrased-question version 99.7% of the time, with only 3 borderline cases out of 188.
  • The researchers recommend adding 'paraphrase-controlled' rankings to leaderboards for transparency.

Why It Matters

This means AI rankings are more honest than once feared, so you can trust top picks without worrying about unfair cheating.

📬 Get the top 10 AI stories daily