Audio & Speech

ELSA: New AI metric evaluates text-to-audio with 10% better human correlation

Reference-free metric matches human ratings by checking event-level alignment in generated audio

Deep Dive

ELSA, a reference-free evaluation metric for text-to-audio (TTA) generation, is proposed. Unlike coarse CLAP-based methods, ELSA decomposes audio into acoustic events derived from the text query and assesses event-level alignment. Tested across four TTA benchmarks, it shows higher correlation with human subjective ratings than prior metrics, enabling reliable automatic TTA evaluation without costly human ratings.

Key Points
  • ELSA evaluates text-to-audio by decomposing queries into distinct acoustic events and scoring event-level alignment
  • Tested on 4 benchmarks, shows significantly higher correlation with human ratings than CLAP-based methods
  • Fully reference-free (no ground-truth audio needed), lowering evaluation costs for TTA model development

Why It Matters

Cheaper, faster, and more reliable TTA evaluation will accelerate audio AI research without expensive human raters.

📬 Get the top 10 AI stories daily