Audio & Speech

ISL framework lets expressive TTS learn from unlabeled speech

This new paper iteratively pseudo-labels emotions and prominence, cutting annotation costs dramatically.

Deep Dive

Expressive text-to-speech systems often rely on explicit conditioning labels—like word-level prominence and utterance-level emotion—to give users direct, interpretable control over how a voice sounds. But manually annotating these attributes at scale is expensive and time-consuming, and no prior semi-supervised method has directly addressed that pain point. Existing approaches focus on scarcity of paired speech-text data or transcriptions, not expressive labels.

In a new arXiv paper (2608.15910), Nicholas Sanders and colleagues from the Centre for Speech Technology Research (CSTR) propose Iterative Self-Learning (ISL), built on Invert-Classify, a classifier-free method that recovers discrete expressive labels by inverting a frozen generative model. ISL works in a loop: the current model pseudo-labels unlabeled speech, the model retrains on the combined labeled and pseudo-labeled dataset, and repeat. Validated on two expressive tasks—word-level prominence and utterance-level emotion—across multiple low-resource data splits, iterative refinement consistently improves pseudo-label accuracy over single-pass baselines. These labeling gains translate into better expressive label adherence and synthesis quality, confirmed by objective metrics and human listening tests. In the most data-scarce conditions, ISL-trained models outperform single-pass pseudo-labeling and edge closer to fully supervised performance, showing that gradient-based ISL is an effective answer to expressive label scarcity in low-resource TTS.

Key Points
  • ISL framework uses Invert-Classify to pseudo-label unlabeled speech by inverting a frozen generative model, no classifier needed
  • Tested on word-level prominence and utterance-level emotion across multiple low-resource data splits, improving pseudo-label accuracy over single-pass baselines
  • In data-scarce settings, ISL-trained models approach fully supervised performance, confirmed by objective metrics and human listening tests

Why It Matters

This could slash the cost of expressive TTS development, enabling high-quality synthetic voices in underrepresented languages and niche use cases.

📬 Get the top 10 AI stories daily