Research & Papers

New Test Reveals AI Still Fails at Hands-On Physics Problems

The AI that aces your homework can't figure out how a ball rolls.

Deep Dive

Researchers behind a new study wanted to know how well AI chatbots actually understand physics. Instead of the usual approach — handing the AI a textbook problem with all the numbers filled in — they built a virtual playground and made the AI figure things out for itself. The test is called PhysMent, and it runs on a physics simulator (software that mimics how real objects move, like in a video game).

Here's how it worked. The AI was dropped into 105 different scenes: ramps, blocks, falling objects and hidden shapes. To answer a question, it had to take actions — push a box, check how heavy something is, fast-forward time, or even build a new object from scratch. Think of it less like a written exam and more like a driving test. You can't pass by reciting the manual; you have to actually steer.

The results were split. On simple, word-based questions, the AI did well — up to 80% correct. But when it needed precise numbers gathered through several careful steps, most models dropped below 30%. Across the seven AI systems tested, scores ranged from 25% to 67%. The researchers say the failures weren't about not knowing physics. The AI knew the concepts. It just ran experiments badly — quitting too soon, exploring inefficiently, or contradicting what the simulator had already told it.

Why should you care? Today's AI is being sold as a tutor, a research assistant and a coding partner. This study is a reminder that it can sound confident while skipping the work. That matters if you use AI for homework help, engineering estimates, or any decision where getting the number right is the whole point. The good news: the test was published publicly, so the next generation of AI can be measured on doing, not just talking.

Key Points
  • AI aced easy physics questions but stumbled badly when it had to run the experiments itself.
  • Across seven AI models, scores ranged from 25% to 67% — most fell below 30% on the hardest tasks.
  • The weak spot is process, not knowledge: the AI quits early and ignores its own test results.

Why It Matters

Don't trust AI for science homework, engineering estimates, or any answer where precision and real-world accuracy matter.

📬 Get the top 10 AI stories daily