Iberian Language Study: Language Mismatch Trumps Speaker Variability in Voice Verification
HuBERT system tested across 5 languages reveals language, not speaker, causes performance loss.
Researchers Pol Buitrago and Javier Hernando have tackled a fundamental challenge in cross-lingual speaker verification: isolating the effect of language mismatch from natural speaker variability. Standard evaluation protocols typically compare different speakers across languages, confounding the two factors. To overcome this, the team created a bilingual same-speaker evaluation set covering five Iberian languages—Spanish, Catalan, Galician, Basque, and Portuguese—allowing them to hold speaker identity constant while varying the language of enrollment and test utterances. They applied this setup to a HuBERT-based speaker verification system previously known to exhibit strong language dependence, and analyzed pairwise cross-lingual transfer using a Cross-Lingual Transfer Matrix (CLTM).
Their results provide a critical clarification: while inter-speaker variability does account for some of the observed performance degradation when languages differ, language mismatch itself is the dominant factor driving cross-lingual verification errors. This finding underscores that even if a system can robustly recognize a speaker in one language, switching languages introduces a substantial drop in accuracy that cannot be explained away by speaker differences alone. The work, submitted to IberSPEECH 2026, offers a more precise evaluation methodology and quantitative insights that could guide the development of language-robust speaker verification systems for real-world multilingual applications.
- Introduced a bilingual same-speaker evaluation set covering 5 Iberian languages (Spanish, Catalan, Galician, Basque, Portuguese)
- Used HuBERT-based SV system with Cross-Lingual Transfer Matrix to isolate language vs. speaker effects
- Found language mismatch is the primary driver of cross-lingual performance loss, surpassing inter-speaker variability
Why It Matters
Paves the way for more accurate multilingual voice authentication by decoupling language from speaker identity in testing.