Scientists Pinpoint Why Voice Assistants Sometimes Invent Words
Ever had Siri hear one thing and type another? Now we know why.
Sometimes speech-recognition systems generate fluent text that has nothing to do with the spoken audio. A study of two Conformer-Large recognizers found a consistent failure point: the final encoder stage. Bypassing that stage makes the transcript drift away from the audio on nearly every utterance, while removing earlier blocks has little effect. Yet the authors stress that this intervention produces garbled or repetitive output, not fluent fabrications. So it reveals a mechanical precondition for hallucination—the loss of grounding—not the complete explanation for natural, fluent hallucinations.
- Voice AI can make up fluent text that has nothing to do with what was actually said — these are called hallucinations.
- Researchers found that both major speech-recognition designs break down at the same stage: the final listening layer of the AI.
- Knowing exactly where this goes wrong could lead to fewer embarrassing voice-assistant mistakes and safer AI for medical or legal transcription.
Why It Matters
Voice AI will make fewer embarrassing mistakes and become safer for crucial tasks like medical notes and court transcripts.