Audio & Speech

SingFox: 113K audio clips across 20 languages for singing deepfake detection

New benchmark dataset tackles AI-generated singing voices in 20 languages

Deep Dive

A team of researchers from India has released SingFox, a comprehensive dataset for detecting AI-generated singing voices (singfakes). The corpus contains 113,802 audio clips spanning 20 languages, totaling over 126 hours of singing data from 1,150 vocalists. To push detection models toward real-world robustness, the dataset is split into six tracks (T1–T6) that introduce specific novelties: language diversity (global and Indian languages), genre-specific music, and alternative deepfake generation methods. This design ensures that models are tested not just on familiar patterns but on the unpredictable variations they would face in production environments.

In cross-dataset evaluation settings—a stringent measure of generalization—the best-performing model achieved 77.84% accuracy. The dataset also supports source verification tasks, helping researchers trace which singing deepfake system generated a given clip. Accepted at INTERSPEECH 2026, SingFox aims to become a standard benchmark for the field. The authors have made all code and resource reproduction tools publicly available, lowering the barrier for labs worldwide to build and compare singing deepfake detectors.

Key Points
  • 113,802 audio clips, 126+ hours, 1,150 singers across 20 languages
  • Six evaluation tracks (T1–T6) covering language, genre, and generation novelty
  • Best model achieves 77.84% accuracy in cross-dataset generalization tests

Why It Matters

As AI-generated singing becomes indistinguishable, SingFox provides the first robust, multi-lingual benchmark to protect artists and content authenticity.

📬 Get the top 10 AI stories daily