New acoustic simulation trains replay speech detectors without real data
Synthetic multi-channel audio defeats voice assistant spoofing attacks...
Researchers Michael Neri and Tuomas Virtanen have developed a new acoustic simulation framework that generates synthetic multi-channel replay speech data to train robust voice assistant security systems. Replay attacks—where an adversary records a user's voice and plays it back to trick a voice-controlled device—pose a growing threat in smart environments. Traditional single-channel detectors struggle to generalize across different acoustic conditions, and multi-channel datasets are scarce. The framework solves this by simulating realistic multi-channel replay scenarios using publicly available resources, enabling training without any real-world recordings.
Using this synthetic data, the authors trained the state-of-the-art multi-channel replay detector M-ALRAD and tested it on the real ReMASC corpus. To better exploit spatial cues, they extended M-ALRAD with inter-channel phase difference features computed for adjacent microphone pairs, adding directional awareness to the beamformed representation. The results show that the synthetic-only training generalizes effectively to real environments, marking a significant step toward scalable, dataset-free defenses for voice assistants. The framework and datasets have been released publicly, and the work is accepted at IEEE MMSP 2026.
- Framework generates realistic multi-channel replay speech data using only publicly available resources
- M-ALRAD detector extended with inter-channel phase difference features for improved spatial awareness
- Trained entirely on synthetic data, generalizes to real ReMASC corpus without any real training recordings
Why It Matters
Scalable synthetic training slashes data collection costs for securing smart speakers and voice-controlled devices against replay attacks.