AI Now Aces College Coding Homework — and Teachers Can Hardly Tell
Nearly 30,000 student coding submissions show AI-style answers are quietly spreading.
A team of university researchers compared nearly 30,000 real student submissions from ten introductory Python labs, collected in 2021, 2023, and 2025, against 90,000 solutions generated by three top AI chatbots. Their question was simple: if a teacher suspects a student used AI, does matching the code against a bank of AI answers actually prove anything? The answer turns out to be messy — but revealing.
The AI models almost always wrote working code, and on most assignments they kept landing on the same handful of solutions. That matters because a student who genuinely solved the problem might write nearly identical code. By 2025, student submissions resembled the AI's answers far more often than in 2021 — including among submissions that passed every hidden instructor test. On narrow, tightly defined tasks, both students and AI squeezed into a few nearly identical code shapes; on open-ended tasks, both stayed varied.
There's a practical catch. Most matches were short snippets, so how long a match must be before it counts is a judgment call that changes the results. And a match at the population level tells you something about a class, not about one person. To say a specific student used AI, you'd need more: their chat prompts, draft versions, intermediate code, or their own admission.
The researchers' advice for teachers is to run AI on an assignment before giving it out, build a reference bank of those answers, and spot which tasks the AI tends to repeat. Then redesign those tasks to ask for reasoning, tests, and working steps — the stuff a chatbot can't fake as easily. The bigger takeaway for everyone: you can measure how AI is changing a whole classroom, but you can't reliably fingerprint one person's honesty from code alone.
- Students' code looked more like AI-written code by 2025 than in 2021 — even when it passed all the hidden tests.
- On simple, tightly specified problems, AI tends to produce just a few identical solutions, so a match isn't proof of cheating.
- Most overlaps were short snippets, so teachers should judge by how long a match is, not just whether one exists.
Why It Matters
Schools and hiring managers can't reliably spot AI-written code yet — raising real questions about grading and trust.