New AI Reads 58,000 Peer Reviews to Judge Research Ideas
It could spot a weak research idea before you waste months on it.
AI can now produce research ideas faster than anyone can judge them. That's the new bottleneck. A team of researchers (led by Rongcan Pei and nine co-authors) built a system called DeepInstructor to fix it. Most existing AI judges just answer from memory — they've read a lot, so they guess. DeepInstructor instead digs through a library of 58,607 real peer reviews, the expert critiques scientists write about each other's papers, and pulls out specific evidence for each question it's asked: Is this idea new? Does it matter? Can it actually be done?
The system is what's called "agentic" — meaning it can take steps on its own, searching and checking before it answers, rather than blurting out a single response. The researchers also built a test set of head-to-head idea matchups, scored on novelty, significance and feasibility. When the AI was asked to pick the better idea, its top choice matched a human expert's top choice 24.4% more often than older tools — and its top-two picks matched 29.7% more often.
Why should you care? Expert review is slow, expensive and in short supply. Journals, funding agencies and universities are drowning in submissions. If a machine can give a quick, evidence-backed first read — and show its reasoning so you can check it — that could mean faster feedback for researchers, students hunting thesis topics, and startups pitching new science. It could also help decide where research money goes.
The catch: this AI learned from past reviews, so it inherits old biases and blind spots. Truly original ideas look strange to it precisely because nothing like them exists yet. It's a preprint, not a product, and the authors aren't claiming it replaces human judgment — only that it can guide it. Treat it as a fast, well-read assistant, not a final verdict.
- DeepInstructor reviews research ideas using evidence from 58,607 real peer reviews instead of guessing from memory.
- It matched human expert judgment 24% to 30% better than older AI evaluators in head-to-head tests.
- Built for a world where AI churns out ideas faster than human experts can review them — a real crunch at journals and funding agencies.
Why It Matters
Faster, cheaper expert feedback on research ideas could speed up science and shape who gets funded.