Audio & Speech

New Benchmark ADD-C Exposes Weaknesses in Audio Deepfake Detection

Researchers reveal that today's audio deepfake detectors collapse under real-world compression and packet loss.

Deep Dive

Audio deepfake detection (ADD) systems have advanced rapidly, but their performance in real-world communication scenarios—where audio is compressed by codecs and degraded by packet loss—remains poorly understood. To bridge this gap, Haohan Shi and five co-authors developed ADD-C, a rigorous new test dataset that simulates diverse communication conditions by varying combinations of audio codecs (e.g., Opus, AAC) and packet loss rates (0% to 20%).

When they benchmarked three state-of-the-art baseline ADD models on ADD-C, all exhibited a substantial decline in detection accuracy—some dropping by over 30 percentage points compared to clean, uncompressed audio. To address this fragility, the team proposed a novel data augmentation strategy that injects realistic communication degradations during training. The augmented models showed significantly improved robustness on ADD-C, reducing the performance gap to near-negligible levels. The paper, accepted at EUSIPCO 2025, provides a much-needed standard for evaluating and building ADD systems that can work reliably in everyday phone calls, VoIP, and virtual meetings.

Key Points
  • ADD-C dataset simulates real-world codec compression (Opus, AAC) and packet loss rates (0–20%).
  • Three baseline ADD models saw detection accuracy drops of over 30 percentage points on ADD-C.
  • A new data augmentation strategy recovered most of the lost performance, enabling more practical ADD systems.

Why It Matters

As voice cloning becomes cheap, robust deepfake detection in calls and meetings is critical for enterprise security.

📬 Get the top 10 AI stories daily