Hybrid Continual Learning Models Identify Endangered Aboriginal Languages Without Forgetting
Two novel methods prevent catastrophic forgetting when adapting speech AI to low-resource languages...
Language identification is a critical step toward integrating endangered Australian Aboriginal languages (AALs) into speech technologies for revitalisation and digital inclusion. However, extreme data scarcity limits model performance. Transfer learning from high-resource languages often suffers from catastrophic forgetting when adapting to new low-resource languages. To address this, researchers introduce two hybrid continual learning methods: Replay Augmented Elastic Weight Consolidation (RA-EWC) and Constraint Guided Knowledge Distillation (CGKD). These approaches allow pretrained speech models to learn new AALs without losing previously acquired knowledge.
Experiments on Warlpiri, Dalabon, and Dharawal demonstrate that the proposed methods significantly outperform both standard fine-tuning and existing continual learning baselines. They improve adaptation to multiple AALs while maintaining performance on previously learned high-resource languages. The work, accepted at Interspeech 2026, provides a practical path for preserving endangered languages through speech AI.
- Two hybrid continual learning methods: RA-EWC and CGKD prevent catastrophic forgetting in speech models.
- Tested on three endangered Aboriginal languages: Warlpiri, Dalabon, and Dharawal.
- Outperforms fine-tuning and existing CL baselines, maintaining high-resource language performance.
Why It Matters
Enables scalable speech technology for language revitalisation, preserving endangered Aboriginal languages for future generations.