Audio & Speech

AI That Invented Pain Scores From Blank Transcripts: Two Models Worst

Your medical AI may guess answers and sound confident — even when it knows nothing.

Deep Dive

AI assistants are powerful, but they don't always say "I don't know." A new study shows two popular AI models will confidently invent medical information — even when the evidence in front of them contains nothing. Researchers tested seven language models on transcripts of people reading neutral sentences while one hand was in cold or warm water. Speakers only described their pain in separate, explicit statements. Once speech-to-text software removed tone and sound, the transcripts carried zero information about pain. Any pain score based only on those words was pure guesswork.

Six of the seven models handled this honestly. When prompted normally, they almost always refused to give a pain score. They also passed a control task, reading spoken pain ratings with 94% to 100% accuracy. But Google's Gemini 2.5 Flash and Meta's Llama 3.1 8B stood out for the wrong reason. When forced to answer, they produced a confident pain score 53% and 76% of the time, respectively. Compare that to 15% or less for the other models. These systems were not just guessing — they were certain while guessing.

The study also found a troubling sensitivity to how questions are phrased. When prompts had an authoritative tone, like a doctor giving instructions, the same models flipped between refusing almost never and refusing almost always depending on tiny wording changes. Reliability was not built into the model; it depended on phrasing. That is bad news for real-world use, where prompts vary constantly and no one audits every input.

Why should you care? AI is already being tested for clinical documentation, telehealth, and mental-health support. If an AI reads your visit summary and "hallucinates" a pain level, a clinician could act on a score that never happened. The authors' key lesson: refusal is not robustness. AI systems need to know when they truly cannot tell, and say so consistently — no matter how the question is asked.

Key Points
  • Six of seven AI models correctly refused to guess pain levels from transcripts that had no pain information — but Gemini 2.5 Flash and Llama 3.1 8B confidently invented scores 53% and 76% of the time.
  • Slightly changing the wording of a request could swing abstention rates from 18% to 100% in the same model, meaning reliability depends on phrasing, not just the AI itself.
  • If these models are used to read medical transcripts, a confident false pain score could lead a doctor toward wrong treatments — a serious patient-safety risk.

Why It Matters

If AI reads your medical records and confidently guesses, doctors could trust false information, causing misdiagnosis or wrong treatment decisions.

📬 Get the top 10 AI stories daily