New AI System Proves Math and Code Correctly — And Shows Its Work
A double-checker for AI-written code could keep buggy software out of your life.
AI assistants can now write mathematical proofs and verification code on their own. The trouble is that a proof can pass the computer's basic check without actually proving the thing you asked for — like a student submitting neat work for the wrong question. A new framework called FORALL-LEAN-AGENT wraps around these AI assistants and treats their output like a job that needs a second opinion. It isolates each task, compares the finished proof to the original request, checks whether hidden assumptions were sneaked in, and has a fresh reviewer sign off. Every approval stays attached to the exact piece of work it refers to, so nothing gets lost in translation.
The results are modest in accuracy but striking in cost. On 100 software-verification tasks, adding this framework pushed the benchmark-rule success rate from 93 to 100 for one top model, while the bill dropped from $69 to $62. On PutnamBench, a set of notoriously hard college-level math competition problems, it accepted all 672 problems at roughly $4.72 apiece. That is cheaper than a coffee per proof, and it comes with evidence instead of just a final score.
Why should you care? Because software runs your bank, your car, and your medical devices, and "it compiled" is not the same as "it is correct." As AI writes more of that code, the real question becomes who checked it and on what grounds. An auditable trail means that when something breaks, you can trace back which proof was approved, by what reasoning, and under what assumptions — the digital equivalent of a signed inspection report, not a shrug.
The catch is real. The accuracy gain was seven extra tasks out of a hundred, measured under benchmark rules that favor this kind of formal work. Lean — a language that lets computers verify proofs step by step — covers only a narrow slice of real-world software, and a human expert still needs to interpret the results. This makes AI reasoning more checkable, not automatically trustworthy.
- A new tool makes AI prove it solved the right problem, not just any problem that looks similar.
- On 100 software tasks it raised the success score from 93 to 100 and lowered the cost from $69 to $62.
- It solved all 672 hard PutnamBench math problems for about $4.72 each — with review receipts attached.
Why It Matters
Cheaper, double-checked AI could mean fewer software bugs and a clear paper trail when AI-written code causes harm.