Chinese researchers build AI that isolates speakers in noisy rooms
New models QIANGDA and VOXBLINK2-AVSE extract target voices with 82% accuracy from crowd noise...
A new arXiv paper introduces QIANGDA, a Mandarin audio-visual target speaker extraction benchmark with 77 scenes and 7,598 clips (11.84 hours), including 6,042 dual-annotated real two-speaker mixtures. It also curates VOXBLINK2-AVSE, a dataset of 250,828 synchronized audio–lip-ROI pairs from 28,421 identities (766.17 hours). The proposed extractor uses frozen AV-HuBERT features and target-conditioned training, and its best checkpoint achieves 0.2261 CER, 82.22% strict output correctness, and 69.53% both-output strict success.
- QIANGDA benchmark includes 7,598 clips (11.84 hours) with dual-annotated mixtures for Mandarin AV-TSE training
- VOXBLINK2-AVSE dataset contains 250,828 audio-lip pairs (766.17 hours) from 28,421 identities
- Best model achieves 0.2261 CER and 82.22% strict output correctness using AV-HuBERT features
Why It Matters
Enables precise speaker isolation in crowded meetings, surveillance, and video conferencing with 82% accuracy.