Research & Papers

New SciConBench shows AI agents score only 0.337 F1 on scientific synthesis

Best frontier agent achieves just 33.7% factual F1 on 9,110 expert-validated questions

Deep Dive

A team of researchers from Princeton, Stanford, and other institutions introduced SciConBench, a large-scale live benchmark of 9,110 questions and expert-written conclusions drawn from systematic reviews. The benchmark uses an automated evaluation pipeline that decomposes conclusions into atomic facts and measures correctness via factual precision and recall. To prevent data leakage, they also built SciConHarness, a clean-room evaluation harness that equips agents with controlled web interaction. When testing 8 frontier models and deep research agents, the best performer under clean-room conditions achieved a factual F1 of only 0.337, indicating that even the most advanced AI agents struggle to synthesize accurate and comprehensive scientific conclusions.

Notably, the clean-room setting consistently reduced performance compared to unconstrained evaluation, suggesting that many prior benchmarks overestimate true synthesis capabilities due to data leakage. The researchers also audited consumer-facing agents including Google AI Overview and OpenEvidence, finding they frequently generate incomplete or contradictory conclusions even when the ground-truth answer is available. The study, which spans 79 pages, 34 figures, and 17 tables, underscores that reliable synthesis of scientific conclusions remains an unsolved challenge, and that clean-room evaluation is essential for assessing open-domain AI agents in high-stakes domains like health.

Key Points
  • SciConBench contains 9,110 expert-validated questions from systematic reviews, with automated atomic fact decomposition
  • Best agent (frontier model) achieved only 0.337 factual F1 under clean-room conditions, with data leakage inflating scores
  • Consumer agents (Google AI Overview, OpenEvidence) generated incomplete or contradictory conclusions despite available ground truth

Why It Matters

AI agents used in health decisions cannot yet reliably synthesize scientific evidence, risking flawed conclusions.

📬 Get the top 10 AI stories daily