CoSTA boosts Alzheimer's detection accuracy to 85.83% with TTS augmentation
Synthetic speech from ASR transcripts improves AD detection by 4.16%
Speech-based Alzheimer's disease detection is limited by scarce pathological speech data. To address this, researchers from multiple institutions propose CoSTA (Cognitive-State-Conditioned TTS Data Augmentation). The framework adapts two state-of-the-art TTS models—CosyVoice2 and F5-TTS—to synthesize speech with distinct Alzheimer's disease (AD) and healthy control characteristics. Additionally, CoSTA constructs a transcript pool using manual transcripts and 36 different Automatic Speech Recognition (ASR) transcripts, allowing systematic evaluation of text source impact on augmentations. This enables richer, more diverse synthetic training data.
Experiments on the ADReSS benchmark show that cognitive-state-conditioned TTS significantly improves synthetic speech utility, and ASR-driven augmentation often outperforms manual transcript-based approaches. CoSTA delivers a 4.16% absolute gain over the baseline, achieving 85.83% audio-only accuracy on the ADReSS test set, outperforming prior methods. The work, accepted at Interspeech 2026, demonstrates a scalable path for non-invasive, speech-based AD screening by overcoming data scarcity through targeted synthetic data generation.
- Uses CosyVoice2 and F5-TTS to generate AD-specific and healthy speech via cognitive-state-conditioned TTS
- Evaluates 36 ASR transcript variants alongside manual transcripts for augmentation optimization
- Achieves 85.83% audio-only accuracy on ADReSS, a 4.16% improvement over baseline
Why It Matters
Enables scalable, non-invasive Alzheimer's screening via synthetic speech augmentation, overcoming data scarcity in healthcare AI.