AI Just Got Better at Inventing Its Own Experiments
New test shows AI can design research faster and cheaper than human experts.
A new benchmark called TasteVal puts AI models in the shoes of a research scientist. Instead of just solving problems, the AI must design experiments and interpret results—a skill called "research taste." The test uses 8 real AI research tasks, like improving language models. A separate coding agent runs the experiments on a single GPU. The AI's score is based on how much experimental compute it needs to match human experts. Less compute means better taste.
The results are striking. The best AI model, Opus 5.5, outperformed a baseline of human experts from top labs like OpenAI and Google DeepMind. It needed only 2.3 times less compute to reach the same score—meaning it has 2.3 times better experimental taste. Even more impressive, it did this at roughly 1/30th the cost per attempt. The benchmark also found that AI's research taste has been doubling every 3 months since December 2025. If this trend continues, it could dramatically speed up AI progress.
But there are important caveats. The tasks are fast and cheap to verify, which may favor AI over humans. The human baseline didn't include the very best researchers, so the AI's advantage might be overstated. Also, TasteVal doesn't measure the ability to choose which problems are worth solving—only how to solve a given problem. So while the results are exciting, they don't mean AI is ready to run entire research projects on its own.
Still, the implications are huge. If AI can design experiments better and cheaper than humans, it could accelerate scientific discovery in many fields. It also raises the possibility of AI improving itself, potentially leading to superintelligence sooner than expected. For now, the benchmark is a promising step toward measuring a key ingredient of AI progress.
- AI model Opus 5.5 beat human experts at designing experiments, using 2.3x less compute.
- AI's experimental research taste is doubling every 3 months, which could speed up AI development.
- The test only measures one part of research and may overstate AI's advantage due to a weaker human baseline.
Why It Matters
Faster AI research could lead to quicker breakthroughs in medicine, tech, and more—but also raises risks of rapid, unchecked AI progress.