Medical AI Gives Different Answers Depending on Which Computer Runs It
The same scan, the same AI — but two different diagnoses. That's a real risk.
Medical AI is increasingly used to help spot tumors, read scans, and summarize patient notes. To make these models fast enough for a busy hospital, engineers split the work across multiple computer chips and machines at once — a setup called "distributed inference" (spreading one AI job across many computers). The common assumption has been that this just makes things faster, not different.
This new paper says that assumption is wrong. Researchers took the same medical AI models, gave them the exact same patient inputs, and ran them two ways: on a single standard setup, and on matched multi-computer setups. The outputs disagreed in measurable, repeatable ways. For models that handle one type of data (like images only), only 21 to 43 percent of tests passed. For models that handle multiple types at once (like images plus written notes), pass rates ranged from 32 to 98 percent — meaning in some configurations, most tests failed.
Why does this happen? When work gets split up, the math is done in a slightly different order and at slightly different precision — like rounding numbers differently at each step of a long calculation. Each tiny difference is harmless alone, but they add up. In a chatbot, a slightly different sentence is annoying. In a model flagging a suspicious lung nodule, it could mean a missed diagnosis or an unnecessary biopsy.
The team built a benchmark — a standard test kit — so hospitals and AI companies can check this before deploying. The message is simple: an AI that passed its offline exam isn't guaranteed to behave the same way in the real hospital system. It needs to be re-tested after deployment.
- Hospitals run medical AI across many computers at once to make it fast — and that changes the answers it gives.
- In tests, only 21 to 43 percent passed for image-only models; pass rates for combined image-and-text models ranged from 32 to 98 percent.
- The researchers built a standard test kit so hospitals can check an AI still works after it's installed, not just before.
Why It Matters
An AI that passes its offline exam may behave differently once installed — potentially affecting real diagnoses and treatment decisions.