AI safety datasets fail non-English languages, study finds
French data passes tests while Hausa and Swahili fall short in safety training
A new arXiv paper (arXiv:2608.13695) titled 'Language-Specific Gaps in AI Safety Training Datasets' exposes critical deficiencies in how AI models are trained to handle non-English languages safely. Led by researchers Chialuka Prisca-Mary Onuoha, Bright Etornam Sunu, and Rashidat Sikiru, the study audited 21 safety datasets across 25 language slices, including low-resource Hausa, mid-resource Swahili, and high-resource French.
The audit revealed that safety claims often collapse at the language-specific level. For example, a Hausa-language slice failed its own translation-quality acceptance threshold, while Swahili passed comfortably in the same pipeline—proving these gaps are addressable but persistently overlooked. Critically, self-harm and sexual-content categories had *no native-language coverage* in either African language tier, a gap that defies resource-level explanations. The findings align with observed asymmetries in multilingual jailbreak robustness, where single-turn attacks are mitigated but multi-turn attacks remain effective, particularly in underrepresented languages.
- Audited 21 safety datasets across 25 language slices (Hausa, Swahili, French) and found systemic gaps in provenance, annotation, and harm-taxonomy coverage
- Self-harm and sexual-content categories lacked native-language coverage entirely for African languages
- Hausa data failed translation-quality thresholds while Swahili passed in the same pipeline, proving gaps are addressable but persistent
Why It Matters
Underserved languages face higher risks of unsafe AI outputs, exposing global inequities in model safety training and deployment.