SpeechDx benchmark tests clinical speech AI across 27 tasks
12 datasets, 27 tasks, 12 encoders—and no model generalizes well yet.
Clinical speech AI has advanced through isolated condition-specific studies, making cross-task comparison and generalization nearly impossible. To address this, researchers Sejal Bhalla, Larry Kieu, Aina Merchant, Eyal de Lara, and Alex Mariakakis from the University of Toronto introduce SpeechDx—a large-scale benchmark that aggregates 12 diverse datasets covering 27 tasks across a spectrum of health conditions. The benchmark uniquely structures tasks by the stage of speech production they disrupt: conceptualization (brain formulating ideas), formulation (translating ideas into language), and articulation (physical production of sound). This organizational framework allows researchers to test whether AI models capture shared clinical mechanisms rather than just dataset-specific artefacts. SpeechDx also includes tasks with limited labeled data and evaluates the same condition across multiple sources to distinguish generalizable patterns from spurious correlations.
To establish baselines, the team systematically evaluated 12 state-of-the-art audio encoders (including large-scale speech models and domain-specific fine-tuned variants) on all tasks, both in standard fine-tuning and zero-shot cross-condition transfer settings. Their results reveal a clear hierarchy: large-scale general speech models (like wav2vec 2.0 and HuBERT) achieve the strongest overall performance, but domain-specific models only improve results on tasks closely related to their training data. Critically, no single encoder generalizes reliably across the full clinical speech landscape—meaning a model that excels at detecting dementia from speech may fail on voice disorder or Parkinson’s tasks. SpeechDx thus establishes a much-needed shared evaluation framework, enabling the research community to track progress toward truly general-purpose clinical speech representations.
- SpeechDx aggregates 12 datasets and 27 tasks across diverse health conditions, organized by speech production stage (conceptualization, formulation, articulation).
- 12 state-of-the-art audio encoders were evaluated; large-scale speech models lead, but no current representation generalizes reliably across the clinical landscape.
- Domain-specific models only improve on closely matched tasks—highlighting the need for a general-purpose clinical speech AI framework.
Why It Matters
Provides a standard benchmark to measure and accelerate development of generalizable clinical speech AI for disease diagnosis.