Research & Papers

Audit finds up to 19.8% image overlap in medical AI benchmarks, detectors unreliable

New study flags potential contamination in medical vision-language models but exposes detector flaws

Deep Dive

A new preprint from Bruce Changlong Xu, Lan Wu, and Alexander Ryu investigates pretraining contamination in public medical vision-language benchmarks. The authors audit open VLMs on datasets like SLAKE-En, PathVQA, VQA-RAD, and an OmniMedVQA mirror using four detector families: image-side nearest-neighbor overlap, canonical-order exchangeability, cohort-relative Min-K%++ tail enrichment, and cross-model top-K overlap. On SLAKE-En, SigLIP-B-16 flagged 19.8% of images as overlapping with PMC-OA-beta, but manual adjudication revealed same-modality, same-projection matches to different patients—not pixel-level duplicates. This suggests source or distributional overlap rather than memorization. Qwen2.5-VL showed a canonical-order exchangeability signal on SLAKE-En's text side, surviving ordering ablation and external baselines. On the OmniMedVQA mirror, exchangeability fired for five models while BLIP-2 remained clean. However, the cohort-relative Min-K%++ and cross-model top-K detectors collapsed under an external pre-domain baseline: BLIP-2 reproduced positive signals despite lacking plausible medical-VQA exposure, indicating these detectors are unreliable for small medical VLM cohorts. The study underscores the nuanced challenge of detecting contamination in specialized domains and calls for more robust evaluation methodologies.

Key Points
  • 19.8% of SLAKE-En images flagged by SigLIP-B-16 for overlap, but manual review confirmed only same-modality matches—no pixel-level duplicates
  • Qwen2.5-VL exhibited canonical-order exchangeability signals on SLAKE-En text side, surviving ablation tests
  • BLIP-2 produced false positives on Min-K%++ detectors, showing these methods are unreliable as standalone membership-inference signals in small medical VLM cohorts

Why It Matters

Highlights contamination risks in medical AI benchmarks and the need for more robust evaluation detectors.

📬 Get the top 10 AI stories daily