Researchers propose personalized emotional TTS with cultural adaptation
New TTS framework adapts to individual emotion perception across cultures using interactive AI optimization
Researchers Wangzixi Zhou, Bagus Tris Atmaja, and Sakriani Sakti have proposed a novel approach to emotional Text-to-Speech (TTS) systems that moves beyond traditional one-size-fits-all models. Their framework, detailed in a paper accepted for INTERSPEECH 2026, leverages Russell's arousal-valence (A-V) modelβa dimensional approach to emotion representation that offers finer control than discrete labels. However, the team identified a critical flaw in existing systems: emotional perception varies widely across individuals and cultures, leading to mismatches between modeled emotions and listener perception.
To solve this, they developed a personalized and culturally adaptive system using an Interactive Genetic Algorithm. This AI-driven optimization process adjusts the A-V space for each listener, tailoring emotional expression in speech to better match individual and cultural preferences. Tests with participants from Japan, China, and Indonesia demonstrated that this approach produces speech with more perceptually aligned emotional expression. The research highlights the importance of personalization in emotional TTS, offering a path beyond the limitations of averaged emotional models.
- Uses Russell's arousal-valence (A-V) model for finer emotion control than discrete labels
- Employs Interactive Genetic Algorithm to optimize emotion perception spaces individually
- Evaluated with Japanese, Chinese, and Indonesian participants showing improved alignment
Why It Matters
This breakthrough could revolutionize AI voice assistants, audiobooks, and customer service bots by delivering emotionally resonant speech tailored to each user's cultural background.