Audio & Speech

BONSAI framework spots elderly audio deepfakes at 1.66% EER

Elderly voices are vulnerable to codec deepfakes, but BONSAI beats top detectors.

Deep Dive

A new study highlights a critical blind spot in audio deepfake detection: elderly speech. Researchers led by Orchid Chetia Phukan introduce the Elderly CodecFake Detection (ECFD) task, releasing the Elderly-CodecFake (ECF) dataset with both English and Chinese samples synthesized by neural audio codecs. They demonstrate that current state-of-the-art codec fake (CF) detectors, trained on standard benchmarks, generalize extremely poorly to elderly voices—revealing a vulnerability that could be exploited for fraud targeting older adults. The team hypothesizes that multimodal foundation models (FMs), pretrained on diverse cross-modal data including elderly content, are better suited for this task.

To address the gap, they propose BONSAI, a novel detection framework that fuses two multimodal FMs—LanguageBind (LB) and ImageBind (IB)—using Jensen-Shannon Divergence as the fusion mechanism. This approach leverages complementary audio-visual representations to improve robustness. BONSAI achieves an average equal error rate (EER) of just 1.66%, significantly outperforming individual FMs and all competitive baselines. The work, accepted at INTERSPEECH 2026, establishes a new benchmark for ECFD and underscores the importance of age-specific deepfake detection as neural codec synthesis becomes more prevalent in voice assistants and telephony.

Key Points
  • Introduces the Elderly CodecFake Detection (ECFD) task with a new bilingual dataset (English/Chinese).
  • Current state-of-the-art codec fake detectors fail on elderly speech, exposing a critical vulnerability.
  • BONSAI fuses LanguageBind and ImageBind via Jensen-Shannon Divergence, achieving 1.66% EER—a new SOTA benchmark.

Why It Matters

Elderly voice fraud is rising; BONSAI provides age-aware detection that current systems miss.

📬 Get the top 10 AI stories daily