New AI Can Transcribe Laughs and Sighs, Not Just Words
Laughter, sighs and coughs carry meaning — your voice assistant has been ignoring them.
Speech recognition — the tech behind captions, Siri and meeting notes — has always thrown away the sounds between words. Laughter, sighs, breaths and coughs get deleted, even though they carry a lot of meaning. A sigh after the word "fine" changes everything. A team of researchers from several Chinese universities built a bilingual system, for Mandarin and English, that transcribes both the words and 16 categories of these nonverbal sounds, tagging each one right where it occurs in the sentence.
Their trick was to adapt Whisper, OpenAI's free and widely used transcription model, rather than train something from scratch. Instead of teaching it a brand-new dictionary, they remapped the vocabulary it already has, so everyday words and sound-tags can come out in one stream. The harder problem was data: these sounds are rare in clean recordings. So the team cleaned up public sound collections, added acoustic variations, used an AI filter to weed out bad examples, and mined spontaneous laughter and sighs from real-world audio and video.
The payoff showed up in an official competition, the NVVSpeech Challenge at ISCSLP 2026. Their score jumped from 33.32 to 53.61 — about 60% better than where they started. For you, that could mean auto-captions and call transcripts that read "[laughs]" or "[sighs]" instead of silently dropping the human part. That matters for accessible video, customer-service recordings, telehealth notes and meeting summaries, where tone often carries more information than the words alone.
The catch: this is a research entry, not a product you can download today, and it only handles Mandarin and English. It relies on labeled public recordings and clean audio, so accents, background noise and other languages are still open questions. And there's a privacy angle — a system that captures more of what you didn't mean to say also captures more of what you'd rather it didn't.
- Most speech-to-text tools delete laughter, sighs, breaths and coughs; this one writes them into the transcript and marks where they happened.
- It's built on Whisper, OpenAI's free transcription model, and scored 53.61 versus 33.32 in an official bilingual challenge — about 60% better.
- Practical uses include captions that show [laughs], call-center and doctor-visit transcripts, and meeting notes that reflect how something was said, not just what was said.
Why It Matters
Captions and call notes may soon mark [laughs] and [sighs], so AI transcripts capture what people actually meant.