New Framework Overhauls Voice Reconstruction Evaluation for Speech Disorders
MOS scores are unreliable; a dual-reference measure captures intelligibility vs identity trade-off.
A new evaluation framework for Text-to-Speech (TTS) voice reconstruction—aimed at people with speech disorders—combines subjective Best Worst Scaling (BWS) with situational framing to measure intelligibility and speaker identity, and introduces a novel dual-reference distributional measure for objective assessment. Tested on 17 zero-shot TTS systems and 193 speakers, the approach provides a reliable and task-aligned alternative to standard MOS-based evaluations.
- MOS-based evaluation lacks sensitivity for highly unintelligible speakers and fails to predict reconstruction success.
- Subjective component uses Best Worst Scaling (BWS) with situational framing for more reliable assessment of intelligibility and speaker identity.
- Objective component introduces a dual-reference distributional measure to capture the trade-off between intelligibility and speaker identity.
Why It Matters
Provides a reliable evaluation method to improve TTS voice reconstruction for people with speech disorders.