Research & Papers

New Test Shows AI Still Can't Think Like a Real Engineer

Top AI scored just 63% on real engineering problems — and that gap matters to you.

Deep Dive

A team of researchers has released MechReason, a new way to test whether artificial intelligence can actually reason like a mechanical engineer. Today's AI tools are good at simple tasks, like reading a single chart or recognizing a machine part in a drawing. MechReason asks something much harder: look at several images, a data table, and a design drawing at once, then connect them using physics and engineering rules to answer one real question.

The test was built from actual mechanical engineering papers, so the questions come from real work rather than made-up puzzles. It contains 12,000 question-and-answer pairs and 21,000 images, covering charts, CAD models (computer-drawn 3D part designs), microscopic photos, simulation results, and factory flowcharts. Questions fall into eight categories across four skills: explaining why something works, predicting what will happen, designing a solution, and diagnosing what went wrong.

The most striking result: even the most advanced AI models answered only 62.89% of the questions correctly. In other words, the best AI on the market today would still get roughly one in three real engineering questions wrong. That is a big deal because mistakes in this field aren't just embarrassing — a wrong call on a material, a tolerance, or a stress limit can mean a failed part, a recall, or an injury.

The takeaway isn't that AI is useless for engineering. It's that AI is a fast, tireless assistant that still needs a qualified human to check its reasoning, especially when a decision depends on combining several pieces of evidence. For anyone who hires engineers, buys manufactured products, or works alongside AI tools, this study is a clear reminder: impressive demos are not the same as dependable judgment. The researchers are releasing MechReason publicly so others can measure progress — and so the next generation of AI can be graded on real problems, not easy ones.

Key Points
  • MechReason is a new public test built from real engineering papers, with 12,000 questions and 21,000 images.
  • The best AI models scored only 62.89%, meaning they'd get about one in three real engineering questions wrong.
  • The test requires combining several images, charts, and physics rules — the kind of multi-step thinking humans do naturally.

Why It Matters

AI can help engineers draft and check work, but you still need a human expert before anything gets built.

📬 Get the top 10 AI stories daily