Image & Video

AI That Reads Eye Scans Gets Different Grades From Different Doctors

Same AI, same eye scans — but the score swings 2.6 points depending on who's grading.

Deep Dive

Researchers evaluated retinal vessel segmentation using all 28 CHASE DB1 images and both human annotations, under a fixed seven-fold protocol that keeps both eyes of each of the 14 subjects together. Random forests and Extra Trees were fitted against observer 1 with three random seeds, giving 42 fits, and five threshold policies shared identical score maps. The same observer-1-tuned random-forest masks scored 73.66% against observer 1 but 71.06% against observer 2 — the number moves with the annotation it is measured against, not the model. For random forests, maximin tuning changed the threshold in 19 of 21 fits, yet worst-observer Dice went from 70.53% to 70.45%, a paired difference of -0.073 percentage points with a conditional subject-bootstrap 95% interval of [-0.384, 0.238]; Extra Trees showed the same direction. The authors say the results support explicitly reporting both the threshold-selection reference and the evaluation reference, and do not support an accuracy benefit from maximin tuning in this cohort. All splits, raw predictions, metrics and code are supplied, and AI assistance is disclosed.

Key Points
  • The very same AI scored 73.66% against one doctor's markings and 71.06% against another's — a 2.6-point swing with zero change to the AI.
  • A fairness trick called 'maximin' changed the cutoff in 19 of 21 test runs but improved the worst-case score by only 0.07 points — basically noise.
  • The study used just 28 eye scans from 14 people, so it's a warning about how medical AI gets graded, not proof any specific product is flawed.

Why It Matters

Medical AI accuracy claims can be inflated by grading choices — so ask who set the standard before trusting the number.

📬 Get the top 10 AI stories daily