Research & Papers

MOOVE study finds pairwise preference masks clinical LLM safety failures

26,804 clinician judgments show preferred models still produce unsafe medical content

Deep Dive

A new arXiv paper from the MOOVE (Massive Open Online Validation and Evaluation) platform, led by Fay Elhassan and colleagues, challenges a core assumption in clinical AI evaluation: that clinicians' pairwise preference for one LLM output over another reflects safety. Using 26,804 blinded pairwise judgments across 13 LLMs, contributed by more than 736 clinicians in 28+ countries, the researchers compared preference rankings with multi-criterion rubric scores on a -2 to +2 scale, where negative values mark clinically unsafe or misleading content. Their analysis found that models ranked highly by pairwise preference still exhibited substantial rates of clinically meaningful failures (scored ≤ -1) on dimensions like Harmlessness and Accuracy. Critically, these failures clustered unevenly across medical specialties, creating "no-go zones" that aggregate leaderboards completely miss. For example, a model that looks safe overall may consistently produce harmful outputs in cardiology or pediatrics, yet still win preference comparisons based on surface-level quality.

The study also dissected why preference diverges from safety. A substantial fraction of preference votes carried no positive safety signal, and feature decomposition revealed that surface-level characteristics (like style or verbosity) explained slightly more preference variation than safety-critical rubric differences. Prompt length, refusal behavior, and whether the model appropriately escalated ambiguous cases all contributed to the gap. To address this, the authors introduce a clinically adjusted preference ranking that combines pairwise preference with rubric-derived safety feedback, producing a more safety-aware ordering than raw Bradley-Terry scores alone. Their findings argue that evaluation pipelines for clinical LLMs should separate preference from safety, report safety-critical failure rates directly (rather than hiding them in averages), and incorporate clinically grounded adjustments when ranking models for deployment. This matters as hospitals increasingly weigh off-the-shelf LLM leaderboards for clinical decision support.

Key Points
  • 26,804 pairwise judgments from 736+ clinicians across 28+ countries showed preference rankings hide serious Harmlessness and Accuracy failures
  • Safety failures concentrate in specific medical specialties, creating "no-go zones" invisible in aggregate leaderboards
  • Combining pairwise preference with rubric-based safety scores yields safer rankings than raw Bradley-Terry strength alone

Why It Matters

Clinicians and health systems choosing LLMs by preference rankings alone could deploy models that are confidently unsafe—safety must be scored directly.

📬 Get the top 10 AI stories daily