OpenAI Dumps 400 Math Proofs, Mathematicians Overwhelmed
AI just solved hundreds of math problems—but can we trust them?
This week, OpenAI dropped a bombshell on the math world: nearly 400 AI-generated mathematical results, spread across more than 700 manuscripts. The topics range from number theory to geometry to physics. Mathematicians described the move as 'staggering' and 'pure insanity.' Many said just reading the table of contents took an hour. The sheer volume makes it hard to know what's real and what's not. Some results come with Lean code, a computer-verifiable proof, but fewer than half have that backup. OpenAI admits the results are at different stages of verification and promises to add more later.
For mathematicians, this is both exciting and terrifying. Exciting because AI might have cracked problems they've worked on for years. Terrifying because their careers could be upended overnight. As one professor put it, if the AIs disappeared now, we'd study this for a decade. But there's a big catch: many results may be 'slop'—low-quality, possibly wrong AI output. Kevin Buzzard, a professor at Imperial College London, found only about six theorems in his area that stood out, and few were formally verified. He worries he'll have to read 'possibly-not-correct slop' or wait for others to check it.
The lack of verification is a major concern. Even when Lean code is provided, researchers must check that it proves what the paper claims—a time-consuming process. Some found inconsistencies between the code and the written claims. This mirrors a broader problem: AI-generated academic papers are flooding fields, making it harder to separate real breakthroughs from noise. OpenAI's move forces mathematicians to adapt quickly, but many fear the company will move on before they can catch up.
So what does this mean for the rest of us? AI is now capable of producing massive amounts of complex work, but without proper checks, it's hard to trust. For math, this could speed up discoveries—or bury them in a pile of unverified claims. For other fields, it's a warning: AI can generate content faster than we can verify it. The next few years will be about figuring out how to separate the gold from the slop.
- OpenAI released nearly 400 AI-generated math results, but fewer than half are formally verified.
- Mathematicians are overwhelmed and worried about 'slop'—low-quality AI content that could be wrong.
- This could change how math research is done, but also flood the field with unverified claims.
Why It Matters
AI can now produce complex math faster than humans can check it, risking trust in research.