Research & Papers

Gemini 2.5 Pro, Claude Opus 4 lead AI polyp diagnosis but miss clinical bar

New study pits GPT-5, Claude Opus 4, Gemini 2.5 Pro on 132 colorectal polyp images.

Deep Dive

A new retrospective study published on arXiv (2608.07543) has evaluated how well multimodal large language models (MLLMs) can perform optical diagnosis of colorectal polyps—a task that typically requires expert endoscopists. Researchers led by Samir Grover of St. Michael's Hospital tested five frontier models: Claude Opus 4, Google Gemini 2.5 Pro, GPT-o3, GPT-4o, and GPT-5. Using the PRIME dataset, which contains white light and narrow-band imaging (NBI) images from 132 cases, the team compared each model's outputs against expert consensus across three classification schemes: Paris, NICE, and predicted histology. F1 scores and percent correct scores were calculated, with statistical testing via Cochran's Q and McNemar's test.

The results show impressive baseline performance: every MLLM achieved an F1 score above 0.9 when distinguishing neoplastic from non-neoplastic polyps, meaning they reliably identify whether a polyp is cancerous or precancerous. However, finer discrimination proved much harder. Gemini 2.5 Pro achieved the highest F1 scores for invasive vs. non-invasive polyps (0.560) and low- vs. high-grade adenoma (0.492)—still low for clinical use. Claude Opus 4 and GPT-5 scored highest on the Paris classification system, with only 41.7% percent correct, statistically beating the other models but far from expert performance. The authors conclude that while these models show promise, their sensitivity and specificity do not meet European Society of Gastrointestinal Endoscopy (ESGE) standards, and they call for prospective multicenter trials and human-in-the-loop workflows before any clinical deployment.

Key Points
  • All 5 MLLMs (Claude Opus 4, Gemini 2.5 Pro, GPT-o3, GPT-4o, GPT-5) hit F1 >0.9 for neoplastic vs. non-neoplastic polyp classification on 132 cases.
  • Gemini 2.5 Pro led invasive vs. non-invasive (F1 0.560) and low- vs. high-grade adenoma (F1 0.492) discrimination.
  • Claude Opus 4 and GPT-5 achieved best Paris classification accuracy at 41.7%, but none met ESGE clinical thresholds.

Why It Matters

AI endoscopy tools risk overpromising; these benchmarks show MLLMs still need human oversight and clinical validation before aiding polypectomy decisions.

📬 Get the top 10 AI stories daily