F-TDNN Just Made Dysarthric Speech 4.6% More Accurate – Here's the Hidden Reason Why
Pitch features and optimal frame overlap cut speech variability errors significantly.
A new preprint on arXiv (2606.19793) presents a systematic study of dysarthric speech recognition, addressing the high acoustic variability caused by impaired articulatory precision. The authors—Paban Sapkota, Hemant Kumar Kathania, Mikko Kurimo, Sudarsana Reddy Kadiri, and Shrikanth Narayanan—explored various combinations of spectral features and acoustic models, with a focus on the Factorized Time Delay Neural Network (F-TDNN). They found that incorporating Pitch features significantly boosted performance, especially for sentence-level tasks. A critical insight was the deliberate selection of the number of overlapping frames between consecutive training example chunks, which helped compensate for speech variability.
The team tested their approach on the TORGO database, a benchmark for dysarthric speech. Their optimized F-TDNN model achieved a 4.65% relative improvement in isolated word recognition and a 4.63% improvement in sentence recognition compared to prior state-of-the-art results. The study provides clear guidance on which acoustic features work best with different models, offering practical selections for researchers and engineers working on assistive speech technologies. This work has direct implications for improving voice-controlled interfaces and communication aids for individuals with dysarthria, a common condition in cerebral palsy, ALS, and stroke survivors.
- 4.65% relative improvement in isolated word recognition using optimized F-TDNN on TORGO database
- Pitch features found particularly effective for sentence recognition in dysarthric speech
- Optimal overlapping frame selection between training chunks compensates for articulatory variability
Why It Matters
Better dysarthric speech recognition means more accessible voice assistants for millions with motor speech disorders.