GigaSpeechBench exposes ASR blind spots with 680-hour multilingual benchmark
New benchmark covers 12 low-resource languages and 6 Chinese dialects to reveal real-world failures
GigaSpeechBench, created by a large team of researchers led by Yujie Tu, is a comprehensive multilingual and multidimensional benchmark for automatic speech recognition (ASR) and automatic speech translation (AST). It comprises 680 hours of human-annotated speech designed to test models under real-world conditions. The benchmark is organized into five modules: (1) 12 low-resource languages from the Middle East and Southeast Asia—plus challenging Japanese and Korean; (2) 6 Chinese dialects; (3) 6 English accents; (4) dense terminology from 12 vertical domains for Chinese and English; and (5) older adult and child speech. Human-annotated Chinese and English translations are also provided for 11 languages to support AST evaluation.
Extensive testing of leading foundation models and commercial APIs revealed that current systems degrade substantially on these challenging settings, particularly on low-resource languages, accented speech, and domain-specific terminology. The benchmark uncovers evaluation blind spots that standard high-resource benchmarks miss, providing a more realistic measure of ASR robustness. GigaSpeechBench aims to push the industry toward more inclusive and reliable speech technology, especially for over a billion underserved speakers across the Middle East and Southeast Asia.
- Covers 12 low-resource Middle Eastern and Southeast Asian languages plus Japanese and Korean
- Includes 6 Chinese dialects, 6 English accents, and dense terminology across 12 vertical domains
- Leading ASR models show significant performance degradation, exposing critical evaluation blind spots
Why It Matters
Helps developers identify where ASR systems fail in diverse real-world conditions across languages and domains