Fine-tuning ZS-TTS captures Singlish accent, study finds
Chatterbox and CosyVoice 3 can now sound like Singaporeans after targeted fine-tuning
Deep Dive
Researchers fine-tuned zero-shot TTS models Chatterbox and CosyVoice 3 on 50 Singlish speakers from Singapore's IMDA corpus. Fine-tuning improved accent similarity for both in-domain and held-out speakers, moving generated speech measurably toward real Singlish. This is the first systematic study of Singlish-accented TTS.
Key Points
- Fine-tuned Chatterbox and CosyVoice 3 on 50 Singlish speakers from IMDA National Speech Corpus
- Accent similarity improved for both in-domain and held-out speakers
- First systematic study of Singapore English (Singlish) in zero-shot TTS
Why It Matters
Enables accurate regional accent synthesis, expanding TTS accessibility for Singapore's diverse linguistic landscape.