AI scientist benchmark: FARS papers beat Sakana, CycleResearcher by 2x
First quantitative benchmark for AI scientists finds FARS scores 2x higher on Gemini and Claude.
A new arXiv study from Vaibhava Lakshmi Ravideshik and Mayank Kejriwal tackles a growing problem: how to objectively evaluate AI systems that autonomously generate scientific research. Dubbed the first quantitative benchmark for AI Scientist systems, the study proposes an automated peer-review protocol using frontier large language models—GPT-5.4, Gemini, and Claude—to score papers on originality, scientific rigor, clarity, and significance.
They ran four leading frameworks—Sakana AI v1 & v2, CycleResearcher, and Data-to-Paper—on a consistent set of 15 research proposals from FARS, a commercial autonomous scientist. Each generated 4 papers, producing 60 total, which were benchmarked against 15 FARS-generated papers. Results show FARS papers significantly outperformed all competitors, scoring 2.14–2.47 out of 5, versus 1.00–1.87 for other systems. On Gemini and Claude evaluations, FARS scored more than 2x higher than the next-best system. Inter-rater reliability was strong: Gemini and Claude agreed closely (Spearman ρ=0.907, p<0.001) and both correlated extremely strongly with the synthesis score (ρ=0.961, p<0.001), suggesting the rubric is consistent. However, GPT-5.4 diverged sharply (ρ≈0.32), implying it applies different evaluation criteria—a caveat for future benchmark design. The authors argue this multi-model LLM review provides a scalable, reliable framework for judging autonomous research quality, potentially accelerating progress toward self-improving AI scientists.
- FARS benchmark papers scored 2.14–2.47 vs 1.00–1.87 for Sakana AI, CycleResearcher, and Data-to-Paper across 60 generated papers.
- Gemini and Claude showed strong reviewer agreement (ρ=0.907) and highly correlated synthesis scores (ρ=0.961), while GPT-5.4 diverged (ρ≈0.32).
- First quantitative benchmark for AI Scientist systems, enabling standardized comparison across 4 frameworks using 3 frontier LLM reviewers.
Why It Matters
Provides a standardized, scalable way to evaluate autonomous research systems—critical as AI scientists enter mainstream R&D workflows.