Research & Papers

AI agents ace individual tasks but fail end-to-end science pipeline

General-purpose coding agents can't yet replace scientists' intuition and visual analysis

Deep Dive

A new arXiv preprint (Horstmann et al., June 2026) presents an empirical evaluation of general-purpose coding agents on a real neuroscience pipeline for fly optogenetics. The researchers assessed agents on tasks significantly larger than existing benchmarks, using datasets orders of magnitude bigger and evaluation criteria grounded in domain expert standards. Agents could solve several individual stages of the pipeline, suggesting that stage-level automation is tractable and could save scientists weeks of work. However, when it came to connecting all stages end-to-end, agents consistently failed.

The study reveals that the core challenge is scientific judgment. Agents struggle most when they must self-assess their own solutions without a predefined criterion to iterate on. They occasionally attempt visual inspection of intermediate outputs—mirroring human scientific practice—but largely fail to interpret what they see or take appropriate corrective actions. Additional obstacles include managing computational resources effectively and generalizing to large, held-out data collections that differ from training distributions. The authors distill principles for constructing scientific tasks and rigorous evaluation criteria for open-ended problems, emphasizing that current benchmarks miss these crucial aspects of real-world research.

Key Points
  • Agents succeeded at individual pipeline stages (e.g., data processing) but couldn't chain them together end-to-end.
  • Lack of predefined optimization criteria forced agents to use scientific judgment, which they failed to apply effectively.
  • Visual inspection of intermediate outputs was attempted but agents misinterpreted results, revealing a gap in reasoning and self-correction.

Why It Matters

Highlights critical gap between current AI agents and the nuanced judgment required in real scientific discovery.

📬 Get the top 10 AI stories daily