SraVaani 1.0 brings open-source ASR to 65 Indian languages
Covering 65 languages and 31K hours, it beats state-of-the-art on low-resource speech...
SraVaani 1.0 is a multilingual speech recognition model that covers 65 Indian languages and dialects, many of which currently have no publicly available or competing ASR system. Built on a FastConformer architecture, it is trained from scratch through a three-stage process: self-supervised pretraining on 31,255 hours of unlabelled speech, an audio-image alignment stage using paired visuals and speech, and end-to-end fine-tuning on 31,263 hours of labelled multilingual Indian speech from 24 public datasets. Across eight benchmarks, SraVaani 1.0 achieves the lowest word error rate on a large number of language-dataset pairs while staying competitive with top systems on high-resource languages. It is also the only open-source evaluated model providing transcription for multiple low-resource and tribal Indian languages, assessed exclusively on the VAANI benchmark.
- SraVaani 1.0 covers 65 Indian languages and dialects, including low-resource and tribal languages with no existing ASR support.
- Built on FastConformer with a three-stage training pipeline: 31,255 hours of self-supervised pretraining, multimodal audio-image alignment, and TDT-CTC fine-tuning on 31,263 hours.
- Achieves the lowest word error rate (WER) across eight benchmarks, and is the only open-source model transcribing several tribal languages.
Why It Matters
Makes speech recognition accessible to hundreds of millions of Indian speakers, enabling voice tech in languages previously ignored.