New Test Shows AI Can Almost Build Software From Scratch
It gets 99% of the code right — and that last 1% still breaks everything.
Most AI coding tools today are used to fix or add to software that already exists. A group of 25 researchers wanted to know something harder: can an AI agent (software that can take actions on its own) build an entire project from nothing? Their new test, Zero2Repo, gives an AI three things — a plain-English description of what to build, a list of required connections, and an empty workspace — and then grades the finished product with automated tests the AI never gets to see.
The tests aren't easy to cheat. The team converted real, popular open-source projects into specifications, ran a known-good version to prove the tests pass, and checked that the tests reject bad versions. The AI gets no hints and no human judge — every test must pass or it scores zero. The current version covers four major programming languages: Python, TypeScript, Go, and C++.
The headline result is both impressive and humbling. On just 11 tasks — drawn from code the AI very likely saw while training — the strongest agent solved only 10. That sounds close, but here's the twist: every single failure had already passed 90-99% of the hidden tests. For the two best agents, most failed tests came down to one missed detail or one rarely-used rule mentioned in the instructions, not a whole missing feature. In other words, the AI builds the house, wires the electricity, and then forgets to install one doorknob.
Why does that matter? Because software doesn't work in percentages. A program that's 99% correct still crashes. This research suggests the big remaining gap isn't whether AI can write code — it clearly can — but whether it can reliably follow every small instruction without drifting. The good news for the field: each failure is a specific, fixable thing, not a mystery. For everyone else, the takeaway is that an AI junior developer is arriving fast, but it still needs a human to check the fine print before shipping.
- A new test called Zero2Repo asks AI to build complete software projects from a written description and an empty folder — not just fix existing code.
- The best AI solved only 10 of 11 tasks, and every failure had already passed 90-99% of the hidden tests.
- Most mistakes came from one missed detail in the instructions, not from missing whole features — meaning AI coders get the big picture but still miss the fine print.
Why It Matters
AI coding assistants are nearly ready to build apps on their own, but small oversights still break things — so humans must review.