TTSYoruba brings natural speech to Yoruba with 651 tone-aware diphone units
Deployed on YorubaName.com, this rule-based system handles contour tones with 5 tonal variants per CV pair.
TTSYoruba is a rule-based concatenative diphone speech synthesizer specifically designed for Yoruba, a tonal language with complex phonology. The system takes tone-marked text as input and produces audio using an inventory of 651 diphone units that cover five tonal variants for every consonant-vowel combination. A custom phonological rule architecture handles tone file selection, three-way nasal disambiguation (oral /n/, nasalized vowel, syllabic nasal), and derivation of rising and falling contour tones from level-tone inputs. The researchers also introduced orthographic extensions—using caron and circumflex as single-vowel contour tone markers—integrated into the TTS normalization pipeline and the WriteYoruba keyboard input tool. This approach preserves tonal accuracy critical for meaning in Yoruba.
Deployed as part of the YorubaName.com open dictionary of personal names, TTSYoruba is currently the only publicly available text-to-speech system for the language. The system was evaluated via a listener study with 50 native speakers, with Mean Opinion Scores (MOS) reported in the paper. The work demonstrates that low-resource languages can achieve high-quality speech synthesis without massive datasets, using linguistic knowledge and careful rule engineering. This has immediate practical applications for language preservation, education, and accessibility for over 50 million Yoruba speakers worldwide.
- 651 diphone units across 5 tonal variants covering all CV combinations in Yoruba
- Novel phonological rules for contour tones, nasal disambiguation, and tonal file selection
- Listener study with 50 speakers; system deployed on YorubaName.com for public use
Why It Matters
Brings high-quality TTS to a low-resource tonal language, aiding preservation and accessibility for 50M+ Yoruba speakers.