Audio & Speech

Voice AI Finally Learns Minnan, a Dialect Tens of Millions Speak

⚡If Siri ignores your grandmother's dialect, this research changes that.

Deep Dive

Researchers introduce WenetSpeech-Min, an open-source corpus of around 10,000 hours of Minnan speech collected from diverse online media, with paired Minnan and Mandarin transcripts for every utterance. The motivation: progress in dialectal speech technology is hindered by the scarcity of large-scale, real-world corpora, and few existing Minnan resources provide paired Minnan and Mandarin transcripts at scale. The team also establishes an ASR benchmark covering both transcript types and a TTS benchmark using Minnan transcripts, each with manually verified evaluation sets. Models trained on the corpus outperform the evaluated open-source systems on most metrics and achieve competitive performance against commercial systems. The corpus, benchmarks, and models will be released to facilitate reproducible research on Minnan speech technology.

Key Points
  • Researchers collected about 10,000 hours of real Minnan speech — roughly a year of nonstop listening — from online videos and audio.
  • Each clip comes with two written versions, Minnan and Mandarin, which teaches AI how the dialect relates to a language it already knows.
  • The models trained on this data beat most free alternatives and rival paid commercial systems, and everything is being released publicly for free.

Why It Matters

Voice assistants may soon understand dialect speakers, so millions aren't shut out of technology because of how they talk at home.

📬 Get the top 10 AI stories daily