Developer Tools

Study: AI Still Can't Be Trusted to Review Engineering Specs

AI misses half of project requirement flaws — a costly blind spot.

Deep Dive

Engineers write requirement documents — detailed statements like "the bridge must withstand 80 mph winds" — that guide every step of a project. If a requirement is vague, impossible, or missing, it can cause expensive redesigns and delays down the line. Reviewing these documents carefully is slow, skilled work, so there was growing hope that AI chatbots could do it automatically.

This new study put that idea to the test. Researchers from the University of Arizona and their colleagues took 10 well-known AI models from OpenAI and Anthropic and asked them to spot defects in two real-world requirement sets. They compared the AI's answers against a panel of human experts, running 100 trials per model. The results were sobering: even the best-performing model caught fewer than half of the actual issues, while flagging about 1 in 10 problems that didn't actually exist. And the hardest catches — judging whether a requirement is truly necessary or correctly worded — were almost always missed.

What makes this especially interesting is that newer AI models didn't consistently outperform older ones. Sometimes the newest models did worse. That means companies can't assume "latest version" means "better at reviewing." The researchers also found that tweaking the AI's "creativity" settings barely changed the results. The flaws aren't random; they're built into how these models think.

The practical message: AI is not ready to independently approve engineering requirements — at least not yet. If you use it to "check" critical project specs, you could end up with faulty requirements that quietly turn into expensive problems. But that doesn't mean AI is useless. As a junior assistant that flags potential issues for a human expert to double-check, it could still help speed things up. For now, the human in the loop isn't a nice-to-have — it's essential.

Key Points
  • Best AI reviewed caught only 47% of real requirement flaws — worse than a coin flip at finding defects.
  • Newer AI models weren't more accurate than older ones, so 'upgrade' doesn't mean 'better.'
  • AI still makes sense as a helper for human experts, but trusting it alone risks expensive project errors.

Why It Matters

If AI reviews specs alone, half of dangerous project flaws slip through — costing time, money, and safety.

📬 Get the top 10 AI stories daily