Agent Frameworks

Study: MLLMs like Gemini and GPT often fabricate chart interpretations

1,224 AI chart descriptions analyzed—only context helped, not images.

Deep Dive

A new study by Ishrat Jahan Eliza and Md Dilshadur Rahman examines how multimodal large language models (MLLMs) generate claims about charts for accessible visualization. The researchers tested three MLLMs—including Gemini and GPT—on 102 visualizations from four sources under four input conditions: image only, source-specific accessible chart context, full context (image + context), and a withheld-context framing prompt. Across 1,224 descriptions, they labeled each model output as DIRECT (directly from data), DERIVED (inferred from data), or SPECULATIVE (outside evidence), and audited numeric consistency.

Key findings: Accessible chart context shifted Gemini and GPT toward generating more DIRECT claims and improved numeric agreement for some models. Surprisingly, adding the image to full context did not yield a consistent numeric benefit, and the “withheld-context” prompt—designed to encourage caution—did not reliably increase hedging language. The prompt-defined “Real-World Significance” section remained predominantly SPECULATIVE across all conditions. The authors argue these results motivate accessible description systems that clearly distinguish claims supported by supplied evidence from model-supplied interpretation, a critical step for trustworthy AI-assisted data accessibility.

Key Points
  • 3 MLLMs tested on 102 visualizations across 4 input conditions, producing 1,224 descriptions.
  • Accessible chart context improved DIRECT claims and numeric agreement for Gemini and GPT.
  • Adding the image to context didn't help; withheld-context prompts failed to boost cautious language.

Why It Matters

Ensures AI-generated chart descriptions clearly separate data-backed facts from speculative interpretation, critical for accessibility and trust.

📬 Get the top 10 AI stories daily