Audio & Speech

Fine-tuning ZS-TTS captures Singlish accent, study finds

Chatterbox and CosyVoice 3 can now sound like Singaporeans after targeted fine-tuning

Deep Dive

Researchers fine-tuned zero-shot TTS models Chatterbox and CosyVoice 3 on 50 Singlish speakers from Singapore's IMDA corpus. Fine-tuning improved accent similarity for both in-domain and held-out speakers, moving generated speech measurably toward real Singlish. This is the first systematic study of Singlish-accented TTS.

Key Points
  • Fine-tuned Chatterbox and CosyVoice 3 on 50 Singlish speakers from IMDA National Speech Corpus
  • Accent similarity improved for both in-domain and held-out speakers
  • First systematic study of Singapore English (Singlish) in zero-shot TTS

Why It Matters

Enables accurate regional accent synthesis, expanding TTS accessibility for Singapore's diverse linguistic landscape.

📬 Get the top 10 AI stories daily