AI Image Description for STEM: Survey Finds Progress, Persisting Hallucination Gaps
New survey of 20 studies reveals AI alt-text still struggles with STEM accuracy.
A systematic survey published on arXiv (arXiv:2607.21611) evaluated 20 peer-reviewed studies on AI-generated image descriptions for STEM domains (Science, Technology, Engineering, Mathematics). The research team, led by Marco Cardia, applied PRISMA methodology and ROBIS-based risk-of-bias assessment to analyze the types of STEM visuals targeted, AI/ML architectures used, datasets and evaluation metrics, and interaction modalities for delivering descriptions. The survey reveals a clear shift from one-shot alt text toward interactive, multimodal systems that integrate conversational interfaces, keyboard navigation, and audio or haptic feedback.
Despite this progress, critical challenges persist. The authors highlight factual inaccuracies and hallucinations in AI outputs, a scarcity of accessibility-first datasets co-designed with blind and low-vision users, and heavy reliance on automatic text-overlap metrics (like BLEU/ROUGE) that poorly capture perceived usefulness and trust. Key future directions include user-controlled verbosity, explainable and verifiable AI pipelines, and integrating accessible description tools into mainstream STEM authoring and learning environments. The work appears in the International Journal of Human-Computer Interaction.
- 20 peer-reviewed studies analyzed covering AI/ML architectures for STEM image description
- Shift from static alt-text to interactive multimodal systems with haptic and audio feedback
- Challenges: hallucinations, lack of co-designed datasets, and weak evaluation metrics (e.g., BLEU)
Why It Matters
Better AI-generated STEM image descriptions can democratize technical education for the 285M visually impaired people worldwide.