Research & Papers

AI X-Ray Report Writers Are Being Graded on Style, Not Accuracy

If AI ever writes your dental scan report, this study is a warning worth reading.

Deep Dive

Cone beam CT is a 3D X-ray dentists and jaw surgeons use to see teeth, jawbone and sinuses. This paper is about teaching AI to write the written report that normally accompanies those scans. But the real subject is grading: how do you score a machine-written medical report? The researchers took apart the scoring system used in a 2026 challenge and found something uncomfortable.

The score had two parts. Eighty percent came from a large language model judging whether the report was factually correct. Only twenty percent came from word-matching — comparing the AI's text to a human's. Crucially, only the word-matching part was visible to teams while they were building. So teams optimized what they could see. A report chosen for word overlap scored 0.2909; one chosen for overall quality scored 0.4122. Chasing matching words pushed factual precision down from 0.522 to 0.266. Better-sounding reports, worse medicine.

Then came the bigger red flag. A small AI model trained on the actual scan images scored 0.486 — essentially a coin flip, no better than always guessing the most common answer. Meanwhile, nine numbers pulled from the scan file's header (machine settings, not pictures) predicted jaw coverage at 0.945 and joint coverage at 0.872. Which clinic the scan came from predicted sentence choices at 0.718, beating the image-based model's 0.663. In other words, the AI was learning how different hospitals dictate reports, not what anatomy looks like.

The team still delivered a working system: eight fixed statements plus five conditional ones driven by scanner metadata, scoring 0.3542 on 50 unseen scans from a hospital it had never trained on. The lesson isn't that the AI is useless. It's that measuring medical AI by how closely its words match a human's rewards imitation of house style — and quietly punishes actually being right.

Key Points
  • The grading system was 80% about medical accuracy, but only the 20% word-matching part was visible to teams while building — so that's what they chased.
  • Optimizing for matching words cut factual precision from 0.522 to 0.266, meaning better-sounding reports that were less medically correct.
  • Machine settings in the scan file predicted anatomy better than the scan image itself (0.945 vs a near coin flip), showing the AI learned hospital writing habits, not medicine.

Why It Matters

Medical AI scored on word-matching may learn hospital writing habits instead of real anatomy.

📬 Get the top 10 AI stories daily