Research & Papers

arXiv study: official conference guidelines beat AI imitations for LLM peer review

Strict rubrics hurt AI review quality—subjective scoring wins.

Deep Dive

A new study by Haowen Li, Yoichi Ishibashi, and Masafumi Oyamada (arXiv 2607.22553, ACL 2026 Findings) systematically evaluates how reviewer guideline design impacts the quality of automated peer reviews generated by large language models (LLMs). The researchers tested three types of guidelines: official conference guidelines (e.g., ACL reviewing criteria), reviewer-imitating guidelines created by prompting an LLM to generate criteria based on high-quality human reviews, and strict rubric-style scoring templates. Using LLMs to produce reviews under each condition, they compared the outputs against human judgments.

The results are striking: official conference guidelines yielded reviews that most closely matched human assessments. This suggests that evaluation criteria refined over years of conference practice serve as effective guidance not only for human reviewers but also for automated systems. In contrast, reviewer-imitating guidelines—despite being derived from actual high-quality reviews—were generally less effective. Most notably, enforcing a rigid rubric-style scoring format consistently degraded review quality. The authors argue that preserving room for subjective, holistic scoring is critical for LLM-based peer review, as overly structured rubrics strip away nuance that human reviewers naturally incorporate. The paper is 18 pages with 2 figures and is published in ACL 2026 Findings.

Key Points
  • Official conference guidelines produce automated reviews most aligned with human judgments.
  • Reviewer-imitating guidelines (LLM-generated from human examples) underperformed official ones.
  • Strict rubric-style scoring consistently degraded performance, favoring subjective/holistic scoring.
  • Study published in ACL 2026 Findings, 18 pages, 2 figures (arXiv 2607.22553).

Why It Matters

Overhauling how we design AI review prompts could make automated peer review more reliable and human-like.

📬 Get the top 10 AI stories daily