Research & Papers

New Public Dataset Reveals How AI Scientists Really Think — Mistakes and All

Two AI models scored the same — but one made 30 times more mistakes.

Deep Dive

Imagine grading a student only on their final exam answer. They write "42," it's correct, and you move on — never knowing whether they understood the problem or just guessed. That is how AI "scientists" have been tested until now. A new research dataset called OpenDiscoveryTrace changes that by recording the full journey, not just the destination.

The dataset contains 558 complete recordings of AI models attempting real scientific work — 124 tasks spanning drug discovery, materials science, genomics and reading scientific literature. For every single step, it logs what the model was thinking, which tools it reached for, what it observed, where it made errors, when it changed its mind, and how confident it felt. Seven models were tested: three big-name paid ones and four smaller free ones. The researchers describe each recording as a flight recorder for AI — a black box you can replay after something goes wrong.

What they found is the real news. All three top models solved 84% to 89% of tasks — basically tied. But one model, Claude Opus 4.6, logged 30 times more errors per task than GPT-5.4 (2.5 versus 0.08). Their mistakes were also different in kind: about two-thirds of Claude's errors were misusing tools, while GPT-5.4's were mostly reasoning slips. Same score, very different behavior — the kind of thing a final-answer-only test can never catch.

The honest catch: this is a pilot analysis covering 363 of the 558 recordings, and the grading was done by AI models rather than human experts. It also covers a narrow slice of science, and AI models change fast, so the results are a snapshot in time. Still, the dataset, the recording format and the scoring tests are all free and public, meaning universities, regulators and journalists can now replay how an AI reached a scientific claim — useful as AI writes more code and more research.

Key Points
  • Researchers released 558 step-by-step recordings of AI models doing real science tasks — showing how they reasoned, not just what they produced.
  • Two top models scored nearly the same on success (84–89%), but one made 30 times more errors — a difference final-answer scoring completely hides.
  • The dataset is free and public, so outsiders can audit AI science claims the way investigators replay a flight recorder.

Why It Matters

As AI writes more research and code, public audit trails let you check the work instead of trusting the answer.

📬 Get the top 10 AI stories daily