Research & Papers

AI Now Fact-Checks Peer Reviewers — And It Agrees With Humans 90% Of The Time

New tool checks whether reviewers' criticisms match the actual paper — science could move faster.

Deep Dive

Peer review is the gatekeeper of science: before a study gets published, volunteer experts read it and write up their concerns. Editors then have to figure out which of those concerns are real problems and which are misunderstandings. Today that checking is almost entirely manual, done by tired humans on tight deadlines — which means mistakes and delays. A team of researchers has now built Peerify, a system that automates that verification step.

Here's how it works, in plain terms. You feed Peerify two things: a research paper and one reviewer comment. It breaks the comment into "atomic claims" — small, single statements that can be checked one at a time, like "the sample size is too small." Then it searches the manuscript for the passages that would support or contradict each claim, and decides. To test it, the team assembled 800 claims pulled from real reviews at NeurIPS 2024 and ICLR 2024, two major AI conferences. Humans hand-labeled 300 of those claims to make sure the automated scoring was fair.

The results were promising. Peerify's automatic labels agreed with what human experts collectively decided 90.3% of the time — about as close as two humans usually get. By contrast, off-the-shelf "entailment models" (simpler AI that just guesses whether one text follows from another) performed terribly, scoring below 0.24. The winning recipe was searching the paper for evidence first, plus splitting reviews into small pieces. But interpretive, judgment-call comments — "this contribution feels incremental" — remain genuinely hard for any machine.

So what does this mean for you? Right now, Peerify is a research pipeline, not something you can sign into. But the trend matters. Faster, more consistent review means scientific findings reach doctors, engineers and the public sooner. It also hints at where this same "does this claim match the source?" technology is heading: fact-checking news articles, verifying legal contracts, or flagging unsupported health claims online. The core skill — checking a statement against the document it came from — is useful far beyond academia.

Key Points
  • Peerify splits a reviewer's comment into small checkable claims, then searches the paper for evidence that supports or contradicts each one.
  • Tested on 800 real reviews from NeurIPS and ICLR 2024, its labels matched human consensus 90.3% of the time, while simpler off-the-shelf models scored below 0.24.
  • Vague, opinion-style reviewer comments ("this feels incremental") still trip it up — it checks support, not whether the reviewer's judgment is correct.

Why It Matters

Could shrink review delays so trusted science, medicine and research reach you faster and with fewer unchecked claims.

📬 Get the top 10 AI stories daily