Audio & Speech

Synthetic speech boosts ASR fine-tuning for Hindi, Kannada, Telugu

New research shows synthetic data rivals real recordings for Indic language ASR...

Deep Dive

A new paper from Sujith Pulikodan and colleagues at the Indian Institute of Science (IISc) systematically evaluates the effectiveness of synthetic speech data for fine-tuning Automatic Speech Recognition (ASR) systems in three Indic languages: Hindi, Kannada, and Telugu. The study, posted on arXiv (2606.17662), explores augmenting real-world recordings with synthetic speech generated by different text-to-speech (TTS) models, including voice-cloned samples. Results show that synthetic data can produce performance gains comparable to using only real data, especially when combined with even small amounts of authentic speech.

The researchers also analyzed how the source of the script used to generate synthetic speech impacts ASR performance, and how performance varies with the number of distinct cloned voices in the training set. They tested multiple TTS synthesis models, finding that voice cloning with a sufficient number of speakers yields the best results. This work is significant for scaling ASR to low-resource languages where collecting large volumes of transcribed speech is expensive. The team notes that synthetic data can serve as a viable alternative or supplement, potentially accelerating deployment of voice interfaces across India's diverse linguistic landscape.

Key Points
  • Augmenting real ASR data with synthetic speech yields performance gains nearly on par with real-only data for Hindi, Kannada, and Telugu.
  • Voice cloning with multiple distinct speakers outperforms single-voice synthetic data in ASR fine-tuning.
  • Different TTS models produce varying quality; the choice of synthesis model and script source significantly affects downstream ASR accuracy.

Why It Matters

Enables scalable, low-cost ASR for under-resourced Indic languages, expanding voice AI access to millions.

📬 Get the top 10 AI stories daily