OscillaTTS: New AI Speech Model Captures Human-Like Prosody
Sharp pitch variations? This TTS model uses oscillatory bias for expressive speech.
Researchers Sandipan Dhar et al. propose OscillaTTS, a diffusion-based TTS that introduces an adaptive oscillatory nonlinearity to model sharp prosodic dynamics. Unlike the standard Snake activation function, their method enables controllable periodic modulation while maintaining signal stability through a linear bypass component. Tested on the LJSpeech and Emotional Speech Dataset, OscillaTTS shows consistent improvements across objective and subjective evaluations. The paper has been accepted at INTERSPEECH 2026.
- OscillaTTS replaces standard Snake activations with adaptive oscillatory nonlinearity for better prosody modeling.
- Maintains signal stability through a linear bypass component while enabling controllable periodic modulation.
- Outperforms baselines on LJSpeech and Emotional Speech Dataset, accepted at INTERSPEECH 2026.
Why It Matters
More expressive TTS enables lifelike virtual assistants, audiobooks, and accessibility tools with natural emotional range.