AI Coding Helpers Are Barely Tested on Real People, Study Finds
The AI writing your software is mostly graded by machines — not humans.
Two researchers published a wide-ranging review of how we test 'code recommender systems' — the AI assistants that suggest code as you type, like GitHub Copilot or Cursor. They combed through 92 published studies from 2017 to 2024 and analyzed what those studies actually measured. Their conclusion is blunt: the testing is lopsided, and it has been for years.
Specifically, they found that 'offline evaluations' dominate. That means a computer runs the suggested code against a fixed set of test problems and scores the result. It's cheap, fast, and repeatable — but it says nothing about what happens when a tired, distracted human being uses the tool on a messy real project. Studies that watch real developers work, or track them over weeks, are rare. Testing also clusters around one narrow moment: writing brand-new code. Less attention goes to debugging, reviewing, or maintaining code, which is how most programmers actually spend their days.
The authors also note that most studies do admit their own weaknesses — usually about the test material they used, since small or artificial examples don't reflect real-world software. But admitting a flaw isn't the same as fixing it. The bigger issue is that most tools are judged on a single quality at a time: does it produce correct code? Rarely: is it faster, less annoying, or safer in practice? Almost never do researchers combine machine testing with human testing in one study.
Why this matters to you even if you never write code: AI-generated software increasingly runs the apps, websites, and banking systems you rely on. If these tools are adopted faster than they are honestly evaluated, the risk isn't just buggy code — it's the false confidence that comes with a good test score. The catch: this paper is a review of research, not proof these tools are bad. Some do get user testing. But the authors' point stands — as AI coding assistants spread, the evidence on whether they truly help people is thinner than the marketing suggests.
- 92 studies reviewed from 2017-2024: most tested AI coding tools with automated checks, not actual people
- Testing focuses narrowly on writing new code — not debugging, fixing, or maintaining it
- Almost no studies combine machine testing with real-user testing, so we get a narrow picture
Why It Matters
AI now writes code you depend on daily — and it's mostly judged by machines, not people.