Research & Papers

FirstResearch framework makes LLM scientific questions auditable with certificates

New method scores 4.86/5, outperforming AI Scientist and Gemini baselines

Deep Dive

Yufeng Wang's FirstResearch framework addresses a critical gap in LLM-driven scientific discovery: the inability to audit the first research question an agent proposes. The system's core artifact is a structured Research Question Certificate that explicitly records primitive definitions, key assumptions, a mechanism model, a tension or contradiction, a falsifiable hypothesis, a minimal decisive test, and a failure update rule. This makes the reasoning behind any proposed question fully inspectable before downstream execution—whether literature synthesis, experiment planning, or report generation.

In rigorous evaluations using both a primary DeepSeek-blind-judge protocol and a Gemini-2.5-Flash independent rescore across 40 baseline packages from AI co-scientist, Agent Laboratory, and AI Scientist-v2, FirstResearch achieved the highest score of 4.86/5 versus 4.38/5 for the best baseline. The judges showed strong agreement (Pearson r=0.865). Ablation studies further confirmed the certificate is the strongest component: removing it crashed scores below 1/5 under both judges. While results are preliminary and use LLM judges rather than human domain experts, the framework demonstrates that explicit derivation constraints are a promising mechanism for making LLM-generated scientific questions more auditable.

Key Points
  • FirstResearch's Research Question Certificate includes 7 explicit components: definitions, assumptions, mechanism model, tension, falsifiable hypothesis, minimal decisive test, and failure update rule.
  • Scored 4.86/5 vs 4.38/5 for AI Scientist-v2 (strongest baseline) across 10 LLM-agent research topics using blind-judge protocols.
  • Certificate-only ablation reached 4.90/5; removing the certificate caused scores to drop below 1/5 under both DeepSeek and Gemini judges.

Why It Matters

Enables transparent and trustworthy LLM-driven scientific discovery by making first research questions auditable before execution.

📬 Get the top 10 AI stories daily