Research & Papers

LLM-as-a-Judge flips 13.6% of pairwise preferences on repeat runs

GPT-4o-mini shows 72% first-position bias; 11 trials needed for reliable verdicts.

Deep Dive

A new study from Abel Yagubyan (arXiv:2606.13685) puts the widely used "LLM-as-a-Judge" paradigm under the microscope, revealing concerning levels of unreliability and bias. Testing OpenAI's GPT-4o-mini and GPT-4.1-mini on 29 tasks across 10 categories, the researcher ran 50 pairwise and 50 pointwise trials per question, plus temperature and prompt-sensitivity ablations. The headline finding: pairwise preferences flip on average 13.6% of the time when the identical evaluation is repeated. 28% of questions exceeded a 20% flip rate, and one question reached 56% flips—essentially a coin flip. GPT-4o-mini also showed a significant first-position bias, choosing option A 72% of the time (p=0.024). Interestingly, mean pointwise score gaps were small (0.19–0.36 on a 10-point scale) and not statistically significant, creating a pairwise-pointwise gap: judges often declare a winner despite scalar scores showing no meaningful quality difference.

Beyond within-judge instability, cross-judge agreement between the two models was only 76% (Cohen's κ = 0.51), and semantically equivalent prompt templates changed majority outcomes in 25% of tested cases. Deterministic decoding reduced but didn't eliminate inconsistency. A reliability curve analysis determined that 11 repeated trials are needed for a majority vote to recover the 50-trial reference verdict with 95% probability on average, rising to 15 for high-variance questions. The author strongly recommends multi-trial aggregation, position randomization, and explicit uncertainty reporting as standard practice for high-stakes evaluation. Because both judges come from a single provider (OpenAI), cross-provider replication remains a critical next step. The paper essentially argues that a single LLM-as-a-Judge call is often too noisy for reliable ranking, leaderboard construction, or reward model training.

Key Points
  • Pairwise preference flips average 13.6% on repeated identical trials; 28% of questions exceed 20% flip rate, one hit 56%.
  • GPT-4o-mini exhibits a 72% first-position bias (p=0.024), choosing option A significantly more often.
  • Cross-judge agreement between GPT-4o-mini and GPT-4.1-mini is only 76% (κ=0.51); 11–15 repeated trials needed for reliable majority verdict.

Why It Matters

Single-trial LLM judging is unreliable; multi-trial aggregation and bias mitigation are essential for trustworthy AI evaluation.

📬 Get the top 10 AI stories daily