PanDent benchmark exposes MLLM failures in dental X-ray diagnosis
9,524 dental X-rays reveal GPT-class models can't localize teeth accurately
PanDent, introduced by Xiaohan Li and colleagues, is a new benchmark designed to evaluate multimodal large language models (MLLMs) on dental panoramic radiography. The dataset includes 9,524 high-quality orthopantomograms (OPGs), each annotated by experienced dentists and validated by an oral and maxillofacial radiologist. These fine-grained, tooth-level annotations provide clinically reliable supervision, with radiology reports constructed using clinician-defined logic to ensure explicit correspondence between structured findings and free-text descriptions.
Experiments on proprietary, open-source, and medical-specific MLLMs reveal a critical gap: current models generate fluent radiology reports but lack clinical consistency, particularly in fine-grained localization and tooth-level diagnosis. Fine-tuning on PanDent substantially improves visual localization and diagnostic accuracy, proving the benchmark's value as a rigorous evaluation tool and a resource for developing clinically grounded dental AI.
- PanDent contains 9,524 expert-validated panoramic dental X-rays with tooth-level annotations from experienced dentists and a radiologist
- Current MLLMs, including SOTA proprietary models, fail at fine-grained tooth localization and produce clinically inconsistent reports
- Fine-tuning on PanDent significantly enhances visual localization accuracy and diagnostic correctness in dental AI models
Why It Matters
PanDent sets the bar for clinically reliable dental AI, pushing MLLMs beyond fluent text toward accurate, actionable diagnosis.