Agent Frameworks

This New AI System Cuts Research Costs by 90% — And Does It Better

Cheaper, smarter AI research assistants could speed up everything from medicine to space travel.

Deep Dive

Researchers built an AI teammate that doesn't just run experiments — it learns from them. Meet Praxist, a system that lets AI research agents track which ideas worked, which failed, and why, instead of starting from scratch each time. Today's autonomous agents often repeat the same mistakes across long projects because they treat every attempt as isolated. Praxist fixes this by creating an "evidence graph" — a connected map of findings and failures that future agents inherit, so lessons actually stick.

The results are striking. On MLE-bench, a standardized test of 75 research tasks, Praxist earned 60 medals (49 gold) versus 55 medals (34 gold) for a Claude Code baseline using the same underlying AI model. More impressive is the cost: Praxist spent about $3,054 on computing, while the baseline burned $38,370. That's a 12x savings — the difference between a research budget that's sustainable and one that isn't.

Praxist also worked on real-world engineering challenges, improving results in algorithmic stock trading, robotic navigation (SLAM), nuclear fusion reactor control (tokamaks), and rocket landings. In each case, it beat the task's existing baseline while leaving a clear audit trail — so you can see exactly which design choice made the difference.

For everyday people, this matters because scientific and engineering breakthroughs often stall on cost and repetition. If AI can do R&D at a fraction of today's price while keeping better records, we could see faster development of better batteries, safer self-driving cars, cleaner fusion energy, and more. The catch: this is early research, not yet available as a product, but it hints at a future where top-tier research capability is no longer a luxury.

Key Points
  • Praxist tracks which AI experiments worked and why, so agents don't repeat mistakes and build on past wins.
  • On a standardized 75-task R&D benchmark, it won more medals than the leading baseline at one-twelfth the cost (about $3,000 vs $38,000).
  • It also tackled real problems — trading, robotic vision, fusion control, and rocket landings — all with an auditable record of how each solution was found.

Why It Matters

Faster, cheaper AI research could accelerate breakthroughs in medicine, energy, and robotics — and make advanced science accessible to smaller labs.

📬 Get the top 10 AI stories daily