New Study: AI Grading Its Own Homework Gets Inflated Scores
If the same rules make the test and grade it, the results are worthless.
Here's the setup. Real recordings of Korean regional dialects exist, but the companies that own them won't let anyone share the files. So researchers did what AI teams increasingly do: they used computers to generate fake dialect examples instead, and trained their models on those. It's like practicing for a driving test using a simulator because you can't get access to a real car.
The promise of this approach is huge. Generating training data is cheap, fast, and legal. So how well does it actually work? That's what this paper set out to measure — and the answer turned out to depend entirely on who's grading.
When the researcher scored the AI using the exact same list of dialect markers that the computer had been told to insert, the numbers looked fantastic: 92% accuracy on "does this sound like a dialect?" and even 119% on "which region is this?" — beating real human data. That last number should have been a red flag, since you can't be better than 100% against the real thing.
So the researcher ran a cleaner test: hold back 20% of the dialect markers so the computer could never generate them, then grade against those. Same amount of training data. The dialect score crashed from 92% to 8%. The region score fell from 102% to 26%. Meanwhile, three truly independent tests didn't budge at all — the model had simply never learned what it claimed to.
The takeaway isn't that synthetic data is useless. It's that when the people building the training data also build the ruler used to measure it, the ruler bends. This is a quiet, widespread problem across AI, from chatbots to medical tools — and it means flashy accuracy numbers deserve a second question: who wrote the test?
- Synthetic data (computer-made training examples) can look nearly perfect when graded with the same rules used to create it.
- Hiding just 20% of the dialect markers crashed one score from 92% to 8% — with no real loss of learning.
- Three independent tests showed no drop at all, proving the original high scores were mostly a measuring trick.
Why It Matters
Flashy AI accuracy claims may be self-graded, so ask who wrote the test before trusting the numbers.