Audio & Speech

New Framework Overhauls Voice Reconstruction Evaluation for Speech Disorders

MOS scores are unreliable; a dual-reference measure captures intelligibility vs identity trade-off.

Deep Dive

A new evaluation framework for Text-to-Speech (TTS) voice reconstruction—aimed at people with speech disorders—combines subjective Best Worst Scaling (BWS) with situational framing to measure intelligibility and speaker identity, and introduces a novel dual-reference distributional measure for objective assessment. Tested on 17 zero-shot TTS systems and 193 speakers, the approach provides a reliable and task-aligned alternative to standard MOS-based evaluations.

Key Points
  • MOS-based evaluation lacks sensitivity for highly unintelligible speakers and fails to predict reconstruction success.
  • Subjective component uses Best Worst Scaling (BWS) with situational framing for more reliable assessment of intelligibility and speaker identity.
  • Objective component introduces a dual-reference distributional measure to capture the trade-off between intelligibility and speaker identity.

Why It Matters

Provides a reliable evaluation method to improve TTS voice reconstruction for people with speech disorders.

📬 Get the top 10 AI stories daily