AI agents autonomously extend academic research with 1% success rate
GPT-5.6-sol cracked decade-old math problems in hours—now it’s doing science
Simon DeDeo, a scholar funded by the Templeton Foundation’s Proofs & Reasons project, conducted an experiment using OpenAI’s GPT-5.6-sol—a model fine-tuned for scientific reasoning—as an autonomous research agent. At the International Congress of Mathematicians (ICM) in Philadelphia, he observed how commercially available AI models were solving long-standing mathematical problems within hours, prompting mixed reactions from the academic community.
DeDeo then tasked the model with extending three of his own published papers across cognitive science and history of science. With full access to computing resources (including a 32-core machine and the CMU ORCHARD cluster) and literature databases, the agent autonomously generated new hypotheses, analyzed additional datasets, and produced extended analyses. The experiment highlighted both the potential of AI-driven scientific discovery and the ethical concerns around unsupervised, recursive research agents operating without safety constraints.
- GPT-5.6-sol (OpenAI’s scientific agent) autonomously extended peer-reviewed research in cognitive science and history of science after being given unrestricted access to compute and literature.
- Deployed on three of Simon DeDeo’s papers, the model achieved results comparable to human researchers, sparking debate about the future of AI in academia.
- The experiment revealed a 1% success rate for AI solving decade-old open problems in mathematics, raising concerns about job displacement and reproducibility.
Why It Matters
AI agents could accelerate scientific discovery by orders of magnitude—but also disrupt traditional research paradigms and require new ethical frameworks.