Audio & Speech

AffectDF benchmark exposes speech deepfake detectors' emotional blind spot

260 hours, 21 attacks, near-random performance — current SDD systems crumble on emotional spoofs.

Deep Dive

AffectDF is a new benchmark from Aurosweta Mahapatra and nine co-authors at UT Dallas, A*STAR, and other institutions, designed to stress-test speech deepfake detection (SDD) systems against emotionally expressive spoofing attacks. It contains roughly 260 hours of speech generated using 21 distinct attacks — covering text-to-speech (TTS), voice conversion (VC), emotional VC, and large audio-language model (LALM)-based methods — across both acted and spontaneous emotional speech in five emotional states. The dataset fills a major gap: existing emotional spoofing datasets are limited in scale and attack diversity, and conventional SDD benchmarks mostly ignore emotionally expressive and LALM-generated audio.

The researchers benchmarked state-of-the-art SDD systems, including LALM-based detectors tested with both inference-only prompting and supervised fine-tuning. Results show severe robustness degradation when models trained on conventional benchmarks are evaluated on AffectDF — several systems perform near-randomly. Surprisingly, even large-scale emotional training does not consistently improve cross-domain robustness, suggesting current SDD systems fail to learn generalized spoof representations under emotional and prosodic variability. Detection performance also varies widely across emotional states, attack families, and acted vs. spontaneous speech. These findings expose fundamental limits of today's detectors and position AffectDF as a standard benchmark for building more robust speech anti-spoofing models.

Key Points
  • AffectDF includes 260 hours of speech from 21 spoofing attacks across TTS, VC, emotional VC, and LALM methods.
  • State-of-the-art SDD models drop to near-random performance on emotionally expressive deepfakes.
  • Even large-scale emotional training fails to provide consistent cross-domain robustness, revealing a need for generalized spoof representations.

Why It Matters

As voice-cloning scams rise, emotionally expressive deepfakes can bypass current detectors — AffectDF forces the field to build truly robust defenses.

📬 Get the top 10 AI stories daily