AI That Reads Lips Fails Badly in Real Life, Study Finds
Your video captions may not be as reliable as the slick demos suggest.
Speech-recognition AI is getting a new superpower: watching your lips while it listens. The idea, called audio-visual speech recognition (AVSR), promises captions that keep working even when a room is loud or a microphone is bad. To test whether that promise holds up, researchers Rishabh Jain and Naomi Harte ran three leading systems through six very different situations — from polished TV news broadcasts to messy, spontaneous group conversations.
On broadcast news, the systems were almost flawless, missing fewer than one word in a hundred. Step outside that bubble and things fell apart fast. Accuracy dropped when people spoke with exaggerated mouth movements, when speakers were amateurs rather than trained professionals, and especially when faces turned sideways. At a 90-degree profile view, the AI mostly gave up reading lips and leaned on sound alone. Camera position mattered less than how clearly a person actually spoke.
The authors say this exposes a "generalization gap" — a fancy way of saying the systems learned to ace one kind of video and assumed everything else looked the same. That matters because most real footage isn't a news studio. It's your video call, a classroom, a hospital consultation, or a crowded café. If AI captions or transcription tools quietly fail in those settings, people could miss important information without realizing it.
To help fix this, the team released RoomReader-AV, a new test set covering tougher, more realistic conditions, plus free software to prepare the data. The takeaway for anyone buying or building speech tools: ask how a system performs on real recordings, not just on a demo reel. Better lip-reading AI is coming — but it isn't as close as the highlight videos suggest.
- AI that reads lips plus listens scored near-perfect on TV news clips, missing under 1 word in 100
- Accuracy crashes when faces turn sideways — at 90 degrees, systems fall back on sound alone
- The team released a free new test set (RoomReader-AV) so companies and researchers can check real-world performance
Why It Matters
Captions, video calls, and transcription tools may be less dependable than advertised outside quiet, well-lit studios.