Research & Papers

AI Test Scores Can Be Inflated by the Test Itself, Study Finds

⚡That glowing AI score may just mean the exam was written by AI too.

Deep Dive

Imagine two students who study from almost identical textbooks. One gets straight A's, the other scrapes by. That's what researchers at the MathNet-Retrieve benchmark (a test that checks whether AI can match a math problem to another document stating the same problem) found — and the difference wasn't intelligence. It was who wrote the practice questions.

The team trained two copies of the same AI model, matched in size and settings. The only difference was their training files. One learned from pairs of math problems written by another company's AI. The other learned from pairs verified by computer algebra, with no AI writing anything. On the easiest tier of the test, the AI-trained version won by 45 points. That's like one runner finishing a race nearly a lap ahead — for reasons that had nothing to do with running.

Digging further, the researchers found that half to two-thirds of that gap came simply from the examples being AI-written at all. When they rewrote the same problems using different instructions, the lead shrank to 30 points, then 22. And the last 15 to 25 points? They vanished entirely on genuine duplicate problems — the same question in two different languages — that no AI had generated. In other words, the model had learned the test's writing style, not the underlying math.

The most striking finding: on one benchmark, a small tweak to a single training file pushed the score up while real-world performance went down. A perfect inversion — the model got better at the test and worse at the job. The researchers released their data and three trained models so others can check the work. The takeaway for anyone buying AI: ask what a score actually measures before trusting it.

Key Points
  • Two identical AI models scored 45 points apart purely because one trained on test questions written by another AI.
  • Most of that advantage vanished on real problems no AI had written — like the same question in two languages.
  • On at least one test, the score went up while real performance went down — a warning for anyone buying AI based on rankings.

Why It Matters

AI leaderboard scores may overstate real ability, so buyers should test tools on their own work first.

📬 Get the top 10 AI stories daily