Research & Papers

Prompt-to-Paper: AI writes bioinformatics papers with real experiments for $0.31

Multi-agent framework grounds claims in 60–100 papers and runs actual code experiments.

Deep Dive

Recent large language models can automate manuscript generation, but they often fabricate claims and results. Prompt-to-Paper, developed by Kamran et al., directly tackles these deficiencies with three integrated innovations. First, a deterministic retrieval-augmented generation pipeline uses section-aware relevance scoring and snowball citation expansion to ground every claim in a verifiable corpus of 60–100 papers — eliminating out-of-range citations. Second, an autonomous coding agent replaces synthetic outputs with genuine numerical results by executing real computational biology experiments. Third, an eight-dimensional automated quality scorer, benchmarked against published paper statistics and augmented with hallucination penalties, provides standardized, reproducible assessments.

The system includes a quality-driven improvement loop: a context-rich reviser routes each iteration to one of three researcher actions, and a deep research cycle fires every ten iterations to re-run experiments and rewrite the manuscript from stronger outputs. Validated on five bioinformatics case studies, all five compiled submission-formatted PDFs with zero out-of-range citations. The improvement loop raised manuscript quality by an average of +17.96 points on a 0–100 scale (max +26.04). As an external check, a human reviewer scored the five manuscripts at an average of 7.0 out of 10. Remarkably, complete manuscripts are produced at approximately $0.31 per paper, making high-quality automated research writing extremely cost-effective.

Key Points
  • Deterministic RAG pipeline with snowball citation expansion ensures every claim traces to 60–100 verifiable papers, eliminating fabricated references.
  • Autonomous coding agent executes real computational biology experiments, replacing synthetic results with genuine numerical outputs.
  • Quality improvement loop boosts manuscript scores by an average of +18 points on a 100-point scale, with human reviewers rating outputs 7/10.

Why It Matters

Automates rigorous bioinformatics research writing while eliminating fabrication, cutting costs to $0.31 per paper.

📬 Get the top 10 AI stories daily