New diagnostic framework reveals deepfake audio detectors rely on silence
Non-speech intervals are a dominant shortcut for audio spoofing detection models.
The authors propose an intervention-based diagnostic framework for deepfake audio detection systems. Using controlled acoustic perturbations targeting non-speech structure, spectral content, and signal energy, they evaluate XLS-R-300M with RawGAT-ST across ASVspoof challenges datasets. Results show that non-speech interventions produce the largest performance shifts, confirming non-speech intervals as a dominant shortcut. Accepted at Odyssey 2026.
- Intervention-based diagnostic framework using directed graphical models to isolate shortcut dependencies in audio deepfake detection.
- Controlled perturbations target non-speech structure, spectral content, and signal energy on XLS-R-300M with RawGAT-ST over ASVspoof datasets.
- Non-speech intervals produce the largest performance shifts, confirming them as the dominant shortcut exploited by current models.
Why It Matters
Exposes why deepfake audio detectors fail in real-world scenarios, enabling development of more robust systems.