LLM chatbots flunk medical study retrieval accuracy
ChatGPT GPT-5.5 leads but only finds 63% of expert-referenced medical studies
A new study published on arXiv evaluates how well major AI chatbots retrieve clinically relevant studies compared to expert-curated Cochrane reviews. The research team tested Claude Sonnet 5, Google’s Gemini 3.1 Pro, and ChatGPT GPT-5.5 across 720 queries simulating patient, clinician, and researcher roles. Each query was repeated four times using clinical questions from the 2026 Cochrane Database.
The results reveal significant variability: ChatGPT GPT-5.5 led with 63.1% average recall of Cochrane-included studies, followed by Claude Sonnet 5 at 37.0% and Gemini 3.1 Pro at just 17.3%. Notably, the researcher role consistently outperformed others, retrieving 42.8% of relevant studies. The study also found a strong bias toward citing large clinical trials—every one-unit increase in log sample size boosted retrieval odds by 1.8x, suggesting LLMs favor high-profile, high-sample-size studies over smaller but equally valid ones.
- ChatGPT GPT-5.5 achieved 63.1% recall of Cochrane-included studies vs 37.0% for Claude Sonnet 5 and 17.3% for Gemini 3.1 Pro
- Researcher role outperformed clinician and patient roles in study retrieval (42.8% vs 38.6% vs 36.1%)
- LLMs disproportionately cite large-sample clinical trials (odds ratio 1.80 per log sample size increase)
Why It Matters
AI chatbots in healthcare may miss critical evidence, skewing clinical recommendations toward high-profile studies.