Vaani Benchmark V1.0: Hindi ASR dataset from 104 Indian districts
Three transcriptions per audio enable robust multi-reference evaluation for Hindi speech recognition.
Sujith Pulikodan and seven co-authors have released Vaani Benchmark V1.0, an inclusive multimodal Hindi automatic speech recognition (ASR) dataset. Collected from 104 districts across India, the dataset captures spontaneous speech elicited using image prompts in real-world acoustic conditions. It covers a diverse demographic and geographic range, addressing the current lack of representative benchmarks for Hindi ASR. Each audio segment is annotated with three independent transcriptions, enabling multi-reference evaluation that accounts for permissible orthographic and lexical variations. This design supports more robust, inclusive, and realistic ASR evaluation compared to single-reference benchmarks.
To demonstrate its utility, the authors benchmarked several open-source and proprietary ASR models, reporting comparative performance on the dataset. The benchmark is hosted on arXiv and includes full dataset and model evaluation details. By including regional diversity and realistic recording conditions, Vaani V1.0 aims to become a standard for Hindi speech recognition research. It is particularly valuable for developing and testing systems that need to handle India's linguistic diversity and real-world noise. The dataset and associated code are available via the paper's arXiv page.
- Dataset collected from 104 districts across India with spontaneous speech from image prompts.
- Each audio clip has three independent transcriptions for multi-reference evaluation.
- Benchmarks multiple open-source and proprietary ASR models, reporting comparative performance.
Why It Matters
Sets a new, inclusive standard for Hindi speech recognition evaluation, crucial for India's linguistic diversity.