New framework exposes Llama Guard's child safety blind spots
Llama Guard fails to detect unsafe prompts in educational AI interactions.
As generative AI becomes common in classrooms and children's apps, existing safety evaluations fall short by ignoring risks unique to minors. A paper accepted at the HEAL Workshop at CHI 2026 proposes a framework that integrates expert-defined hazard categories with real-world AI incident databases to create synthetic test sets for child safety. The framework applies domain-specific risk factors—rather than generic content filters—to evaluate how well models detect unsafe interactions.
Applying the framework to the education sector, the study tested three Llama Guard models on detecting unsafe user prompts. Results show that current Llama Guard models struggle to identify education-related unsafe prompts, highlighting a significant gap in current safety measures. The authors emphasize the need to extend evaluation to more risk categories and involve domain experts throughout the pipeline to ensure AI systems are safe for children.
- Framework built from expert child safety guidelines and real AI incident data
- Applied to education domain, tested Meta's Llama Guard models
- Current models fail to detect many education-related unsafe prompts
Why It Matters
Without child-specific safety evaluations, AI tools risk harming young users in educational settings.