Audio & Speech

Multilingual training boosts speech emphasis model generalization across 7 languages and 34 emotions

Monolingual emphasis models fail zero-shot across languages; multilingual training fixes it.

Deep Dive

Researchers Megan Wei, Deepali Aneja, Jiaqi Su, Yunyun Wang, Haonan Chen, and Zeyu Jin present MMEE (Multilingual Multi-Emotion Emphasis), a corpus of 10,000 professionally recorded expressive utterances (14.13 hours) across 7 languages and 34 emotion/style categories, each annotated with three-level perceptual labels (10 annotations per sample). The dataset addresses a critical gap: existing emphasis detection models are trained and evaluated almost exclusively on monolingual neutral read speech, limiting their real-world applicability. The team benchmarked two state-of-the-art architectures under monolingual, cross-lingual, multilingual, cross-emotion, cross-dataset, and data-scale settings.

Key findings reveal that monolingual models show limited zero-shot transfer, degrading significantly when tested on typologically distant languages. However, multilingual training substantially improves robustness, enabling models to maintain performance across languages. The models also transfer robustly between high- and low-arousal emotions, and the bidirectional transfer between synthetic and perceptual benchmarks suggests shared prosodic structure. Performance remains stable even at smaller training scales. Accepted at Interspeech 2026, this work provides a vital resource and insights for building speech AI that understands emphasis across languages and emotional contexts, with implications for multilingual voice assistants, emotion-aware dialog systems, and cross-cultural communication tools.

Key Points
  • MMEE corpus includes 10,000 utterances across 7 languages and 34 emotion/style categories with three-level perceptual labels.
  • Monolingual emphasis detection models degrade sharply in zero-shot transfer to typologically distant languages.
  • Multilingual training substantially improves cross-lingual and cross-emotion robustness, with stable performance even at smaller data scales.

Why It Matters

Enables speech AI to detect emphasis across languages and emotions, powering more natural multilingual and empathetic voice interactions.

📬 Get the top 10 AI stories daily