Research & Papers

AI visual QA tools fail both blind and sighted scientists on research papers

Vague AI descriptions and errors cause 100% of users to abandon tools...

Deep Dive

A new study from Arnavi Chheda-Kothary, Lucy Lu Wang, Joseph Chee Chang, and Jonathan Bragg (University of Washington / Allen Institute for AI) explores how scientists interact with AI tools to query visual content in multimodal scientific papers. Interviews with five blind/low-vision (BLV) and five sighted scientists across STEM fields revealed that both groups use tools like ChatGPT and Gemini to ask questions about figures, diagrams, and tables. However, vague or incomplete image descriptions—along with outright incorrect AI outputs—frequently led participants to abandon AI-based workflows entirely.

The paper also introduces a publicly available dataset of 115 real queries and AI responses from these interactions. The findings underscore that current AI systems are not reliable enough for professional scientific use, even for sighted users, and highlight specific accessibility gaps for BLV scientists. The authors call for future AI-powered scientific QA systems to prioritize accurate, detailed visual descriptions and robust error handling across diverse user abilities and domains.

Key Points
  • 5 BLV and 5 sighted scientists interviewed across STEM fields using ChatGPT and Gemini to query scientific figures
  • Both groups abandoned AI tools due to vague image descriptions and incorrect outputs
  • Study contributes a dataset of 115 real user queries and AI responses for future research

Why It Matters

Highlights critical reliability and accessibility gaps in AI tools for scientific research, affecting all users.

📬 Get the top 10 AI stories daily