Blood-Cell AI Fails When Machines Change — Study Warns
This could mean wrong blood test results at your hospital.
A team of researchers wanted to know if AI models that analyze blood cells can be trusted in real hospitals. These models, called "foundation models," are trained on huge collections of blood sample images to spot different types of white blood cells. In controlled lab tests, they achieve almost perfect accuracy — above 98% — which sounds great. But real-world hospitals use different scanners, stains, and preparation methods, and that's where things fall apart.
The study tested 15 popular AI models on blood images from four different sources. Once the models encountered data from a new machine or lab, their accuracy plunged by 34% to 72%. The model that performed best in the original lab dropped to 10th place on the most different dataset. Even worse, when these models were wrong, they were often confidently wrong — meaning they flagged incorrect results with high certainty, which is dangerous for diagnosis.
Why does this happen? The models seem to pick up on shortcuts, like the exact color or pattern produced by a specific scanner, instead of learning the actual features of the cells. When a new machine is used, those shortcuts no longer apply. The paper also points out that some "new" data wasn't truly new — it came from the same source that trained the models, making the results look more reliable than they really are.
The takeaway is clear: AI blood-cell analysis needs to be tested with real-world diversity — different hospitals, machines, and conditions — before we can trust it with patient health. The researchers propose a simple fix that helps but doesn't fully solve the problem, and they urge the medical AI field to adopt stricter evaluation standards.
- AI blood cell models drop 34–72% accuracy when tested on images from different scanners or labs.
- The best lab performer fell from 1st to 10th place on shifted data, showing results often can't be generalized.
- Models were dangerously overconfident when wrong, which could lead to misdiagnoses if used in hospitals.
- Researchers call for testing AI on diverse real-world data before approving it for clinical use.
Why It Matters
AI diagnostics are only safe if they work across hospitals — this study shows they often don't.