Research & Papers

Google's Medical AI Reads Lung Scans as Well as Some Radiologists

It never gets tired or wavers — but it's still not better than your doctor.

Deep Dive

Lung cancer screening works like this: you get a CT scan, and a radiologist grades what they see using a standard checklist called Lung-RADS. Sounds straightforward, but even with that checklist, different doctors often grade the same scan differently. This study put that problem to the test, pitting a medical AI called MedGemma — built on Google's Gemini technology — against 12 radiologists reading the same set of scans.

On a standard accuracy scale where 1.0 is perfect and 0.5 is a coin flip, the radiologists averaged about 0.90, though individual scores ranged from 0.80 to 0.94. The untrained AI managed only 0.70, which the researchers called not good enough for real clinical use. After additional training specifically on lung cancer cases, it climbed to 0.83 — roughly matching the weaker end of the human group.

The AI's genuine advantage isn't peak accuracy, it's consistency. Give it the same scan twice and you get the same answer, every time. Radiologists vary — between people, and even for the same person on a different day. In clinics short on specialists, or in places where screening expertise is thin, that steady second opinion could be genuinely useful.

There's an important catch. The test used a hand-picked, case-enriched sample drawn from a large US screening trial, not a typical screening population, and it was never checked against outside data. That means real-world accuracy in an ordinary clinic is still unknown. The takeaway: this is a promising assistant, not a replacement — and it isn't in your doctor's office yet.

Key Points
  • Untrained, MedGemma scored 0.70 on accuracy — too weak to help. After extra training it hit 0.83, matching the least accurate of 12 radiologists.
  • The AI's biggest strength is consistency: identical scans always produce identical answers, while doctors' scores varied by 0.14 points.
  • The test used a hand-picked set of cases with no outside validation, so real-world performance in ordinary clinics remains unproven.

Why It Matters

A tireless second opinion could help clinics short on radiologists — but it isn't accurate enough to replace them yet.

📬 Get the top 10 AI stories daily