Audio & Speech

CoSTA boosts Alzheimer's detection accuracy to 85.83% with TTS augmentation

Synthetic speech from ASR transcripts improves AD detection by 4.16%

Deep Dive

Speech-based Alzheimer's disease detection is limited by scarce pathological speech data. To address this, researchers from multiple institutions propose CoSTA (Cognitive-State-Conditioned TTS Data Augmentation). The framework adapts two state-of-the-art TTS models—CosyVoice2 and F5-TTS—to synthesize speech with distinct Alzheimer's disease (AD) and healthy control characteristics. Additionally, CoSTA constructs a transcript pool using manual transcripts and 36 different Automatic Speech Recognition (ASR) transcripts, allowing systematic evaluation of text source impact on augmentations. This enables richer, more diverse synthetic training data.

Experiments on the ADReSS benchmark show that cognitive-state-conditioned TTS significantly improves synthetic speech utility, and ASR-driven augmentation often outperforms manual transcript-based approaches. CoSTA delivers a 4.16% absolute gain over the baseline, achieving 85.83% audio-only accuracy on the ADReSS test set, outperforming prior methods. The work, accepted at Interspeech 2026, demonstrates a scalable path for non-invasive, speech-based AD screening by overcoming data scarcity through targeted synthetic data generation.

Key Points
  • Uses CosyVoice2 and F5-TTS to generate AD-specific and healthy speech via cognitive-state-conditioned TTS
  • Evaluates 36 ASR transcript variants alongside manual transcripts for augmentation optimization
  • Achieves 85.83% audio-only accuracy on ADReSS, a 4.16% improvement over baseline

Why It Matters

Enables scalable, non-invasive Alzheimer's screening via synthetic speech augmentation, overcoming data scarcity in healthcare AI.

📬 Get the top 10 AI stories daily