LLM peer review study reveals reliability flaws and risks
New arXiv study exposes critical vulnerabilities in AI-powered academic review
A new arXiv paper titled 'LLM-Based Scientific Peer Review: Methods, Benchmarks, and Reliability Challenges' (arXiv:2606.25057) by Thi Huyen Nguyen and Zahra Ahmadi delivers a sobering assessment of AI's readiness for academic peer review. The survey analyzes current LLM approaches—including prompt-based, supervised, retrieval-augmented (RAG), and alignment-optimized systems—against two core functions: critique generation and score prediction.
The authors highlight critical reliability gaps, including dataset constraints and domain concentration biases that limit benchmark validity. More concerning are identified robustness risks such as prompt injection, data poisoning, retrieval vulnerabilities, and reward hacking that could enable strategic manipulation of automated review pipelines. The paper reframes peer review as a 'high-stakes, multi-objective decision problem' and calls for robust, transparent, and trustworthy AI systems to address subjective disagreement and cross-domain generalization challenges.
- First comprehensive taxonomy of LLM approaches (prompt-based, supervised, RAG, alignment-optimized) for scientific peer review
- Identified critical vulnerabilities including prompt injection, data poisoning, and reward hacking in AI review systems
- Proposed roadmap for developing robust, transparent AI-assisted evaluation amid reliability and generalization challenges
Why It Matters
AI peer review could reshape academic publishing but requires addressing critical reliability risks before deployment.