Audio & Speech

Age-Aware Adapters Cut Children's Speech Recognition Error by 3%

One-size-fits-all child ASR adapters leave performance on the table...

Deep Dive

Children's speech recognition has long posed a challenge due to developmental variation and differences from adult speech. While adapter tuning—fine-tuning small modules on large pretrained ASR models—has shown promise, a single shared child adapter fails to account for age-dependent acoustic differences. In a new paper, researcher Jialu Li presents one of the first systematic studies of age-aware adapter tuning for child ASR, targeting speech from ages 3 to 12 and older.

The study proposes two strategies: age-specialized adapters trained separately for each age group, and a unified age-conditioned FiLM adapter. Using ground-truth age routing, the age-specialized approach improved overall WER from 12.6% to 12.3% and macro WER from 18.4% to 17.6%, with consistent gains across all age groups. Remarkably, using predicted age (without ground-truth labels at inference) achieved nearly identical results: 12.3% overall WER and 17.8% macro WER. The unified FiLM adapter delivered smaller improvements, confirming that a single adapter cannot fully capture developmental variation. The work points to practical, age-aware fine-tuning as a key direction for more accurate child ASR in applications like educational tools and virtual assistants.

Key Points
  • Age-specialized adapters reduce overall WER from 12.6% to 12.3% and macro WER from 18.4% to 17.6%
  • Predicted-age routing achieves 12.3% overall WER without requiring age labels at inference
  • Unified FiLM conditioning yields smaller gains, indicating a single adapter insufficient for developmental variation

Why It Matters

Better child ASR accuracy enables more reliable voice interfaces in education and accessibility tools.

📬 Get the top 10 AI stories daily