AI Safety

Hospitals Test AI to Catch Missed ER Mistakes — But ChatGPT Overreacts

AI could flag the ER visits that go wrong — if it learns when to stay quiet.

Deep Dive

When a patient is sent home from the emergency room and comes back days later, hospitals want to know why. Was it a missed diagnosis? A complication? Something the first doctor should have caught? These "revisits" get reviewed by clinicians, but only a narrow slice of them — usually returns within 48 to 72 hours — because reading medical charts by hand is slow and expensive work. That narrow window means real problems outside it can slip by unexamined.

Researchers at a multi-hospital health system tested whether AI could widen that net without burying doctors in paperwork. They took 99 pairs of diagnoses — the reason for the first visit and the reason for the return visit — and asked two to three human clinicians plus ChatGPT's GPT-4 model one simple question: does this pair deserve a closer look? The humans were selective. GPT-4 was not. It flagged 94% of pairs as needing follow-up, roughly 4 to 13 times more often than the clinicians did.

So the team built something more disciplined. Using GPT-4 to help organize medical knowledge into a structured map of connected conditions, they created an automated screening tool. That tool was right 83 to 100% of the time when at least one clinician agreed a case warranted review. In other words: a smarter setup got much closer to human judgment than raw ChatGPT did. The lesson is that how you ask the AI matters enormously.

The catch: this is one small, early study of 99 cases, done after the fact, at a single health system. No hospital is using it. And an AI that cries wolf on almost everything is worse than useless — it just moves the paperwork burden around. Still, the direction is promising: the goal isn't to replace the doctors reviewing cases, but to help them sift through far more of them with the same amount of time.

Key Points
  • Today hospitals only review emergency-room return visits within about 48-72 hours, so problems outside that window go unchecked.
  • Asked to judge 99 case pairs, ChatGPT flagged 94% as needing review — up to 13 times more often than human doctors, which would swamp staff.
  • A more structured AI tool built on a medical knowledge map matched human judgment 83-100% of the time, suggesting screening could expand without adding workload.

Why It Matters

AI could one day help hospitals catch missed ER diagnoses sooner — but only if it stops flagging almost everything.

📬 Get the top 10 AI stories daily