Research & Papers

AI Just Caught a 29-Year-Old Bug in Scientists' Math Software

⚡The mistake hid in code scientists rely on daily — and AI finally found it.

Deep Dive

PETSc is a software toolbox that scientists and engineers use to run heavy simulations — airflow over wings, weather models, stress on bridges. When it gives a wrong answer, the consequences can be expensive. Yet libraries like this are almost never formally checked, because writing the rulebook by hand takes an expert a very long time. Now three researchers used AI to draft those rules from existing documentation, then a strict automated checker compared the code against them. It found a genuine bug in a function called MatAYPX.

That bug had been sitting there since 1997 — roughly 29 years. Think of it as a typo in a calculator that millions of people use and nobody noticed, because it only misbehaves in certain situations. Notably, the AI did not do the whole job. It produced only a small rule sheet, which a human then reviewed and approved. The pattern matters: AI drafts, a person approves, a machine verifies. That is likely how AI will slip into high-stakes work — quietly, in the boring middle steps.

The catch is scale. The team tested three functions. PETSc contains thousands. The researchers also warn that letting AI write code freely from a single prompt is unreliable — you just get more untrusted code to audit. And the paper does not claim the 1997 bug ever produced a wrong real-world result. It was found, not proven harmful.

So why care? Because the software underneath medicines, bridges and climate forecasts is often quietly unaudited. This shows AI can make checking far cheaper. If that spreads, the simulations your doctor, your city planner, or your weather app trust could get measurably more reliable — and you would never see the change.

Key Points
  • AI wrote the plain-English rules for what a science-math function should do, then a checker tested the code against them.
  • It found a bug that had been hiding in PETSc since 1997 — about 29 years.
  • Humans still approved the AI's rules, and only 3 of thousands of functions were tested.

Why It Matters

The math software behind medicines, bridges and weather forecasts gets cheaper to audit — meaning fewer silent wrong answers.

📬 Get the top 10 AI stories daily