Multimodal speaker verification cracks anonymization with just 5 utterances
Combining audio and text reduces error rate by over 15%, exposing privacy flaws.
A new study from Ashi Garg and colleagues challenges the effectiveness of speaker anonymization, the technique used to hide a speaker's identity in audio recordings. Traditional automatic speaker verification (ASV) systems analyze single utterances, but the authors demonstrate that aggregating information across multiple anonymized utterances—especially using both audio and text—dramatically re-identifies speakers. The paper, submitted to arXiv in July 2026, shows that with just five anonymized utterances, a multimodal system combining audio, prosodic features, and linguistic data reduces the equal error rate (EER) by over 15% compared to audio-only aggregation.
The researchers compared aggregation strategies and found that frame-level fusion yields the lowest EER, outperforming utterance-level or score-level methods. This suggests that current anonymization techniques, which primarily mask vocal characteristics like pitch and timbre, leave rich speaker-discriminative information in prosodic patterns and word choices. The work has immediate implications for privacy tools used in whistleblower protection, medical transcription, and voice assistants—if an adversary can collect multiple anonymized recordings, identity could be exposed.
- Multimodal aggregation (audio + text) reduces EER by >15% over audio-only with just 5 anonymized utterances.
- Frame-level fusion across multiple utterances outperforms utterance-level or score-level aggregation.
- Existing speaker anonymization techniques fail to remove prosodic and linguistic cues, leaving privacy vulnerable.
Why It Matters
Speaker anonymization tools may be ineffective if attackers can aggregate multiple utterances—threatening privacy in sensitive voice applications.