Audio & Speech

CLeaD reveals inflated depression detection scores from speaker identity leakage

Prior Mandarin F1 scores of 0.954 were an artifact of identity leakage

Deep Dive

Speech-based depression detection has shown promise in monolingual settings, but cross-lingual generalization remains a challenge due to disparities in diagnosis and clinical presentation across languages. Prior work often used segment-level random splits without speaker grouping, inadvertently allowing the model to learn speaker identities rather than depression-related cues. This inflation artificially boosted reported metrics. To address this, researchers from the University of Southern California and other institutions introduce CLeaD (Contrastive Learning for Depression), a supervised contrastive alignment framework that maps WavLM embeddings from English and Mandarin into a shared clinical space without parallel data or target-language fine-tuning.

Evaluating 52 Mandarin speakers under rigorous leave-one-speaker-out conditions, CLeaD modestly outperforms the baseline (F1: 0.640 vs. 0.622) and improves depressed-class recall at intermediate layers (7-8). However, the small test set limits generalizability. Two robust findings emerge: model scaling degrades cross-lingual performance even as it boosts monolingual English results, and speaker identity leakage previously inflated reported Mandarin F1 scores to 0.954—an artifact the team reproduces and quantifies. The work highlights the critical need for proper evaluation protocols in multilingual mental health AI.

Key Points
  • CLeaD uses contrastive alignment on WavLM embeddings for English–Mandarin depression detection without parallel data or target-language fine-tuning.
  • Under correct leave-one-speaker-out evaluation, CLeaD achieves F1 0.640 vs 0.622 baseline; previous Mandarin F1 of 0.954 was due to identity leakage.
  • Model scaling hurt cross-lingual performance while improving monolingual English, revealing a fundamental tradeoff.

Why It Matters

Exposes evaluation pitfalls in cross-lingual mental health AI, demanding rigorously validated models for fair, non-biased deployment.

📬 Get the top 10 AI stories daily