Research & Papers

Frontier VLMs fail Theory of Mind: 78% egocentric errors, ASD-like bias

Nine top VLMs score closer to autistic adults than neurotypical adults on social reasoning...

Deep Dive

A new arXiv paper by Kejia Zhang and colleagues probes whether frontier vision-language models (VLMs) possess a coherent Theory of Mind (ToM)—the ability to infer others' mental states. The study tested nine unnamed frontier VLMs on two classic psychology benchmarks: the Keysar Director Task, which measures visual perspective-taking under egocentric interference, and the Frith-Happé animated triangles test, which assesses intention attribution from motion cues, scored with the Castelli rubric.

Results reveal a stark dissociation. On the Director Task, without chain-of-thought reasoning, the model panel made egocentric errors on 78% of trials—a child-like pattern—though reasoning prompts rescued several models. On the animated triangles, models under-attributed intention, producing ToM profiles more than three times closer to the high-functioning-autistic-adult (HF-ASD) group mean than to the typical-development-adult (TD) mean, while Goal-Directed and Random motion remained near TD. Critically, no single model achieved adult-like performance on both tasks: the best Director Task model fell on the HF-ASD side for triangles, and the most TD-like triangle model remained child-like on the Director Task. The authors emphasize these are group-level descriptions, not diagnostic labels for any individual model. The findings suggest current VLMs lack a unified social-cognitive architecture, instead exhibiting task-specific, fragmented ToM capabilities that mirror distinct neuropsychological profiles.

Key Points
  • 9 frontier VLMs tested on Keysar Director Task and Frith-Happé animated triangles
  • Without chain-of-thought, models make egocentric errors on 78% of Director Task trials
  • Model ToM profile on triangles is 3x closer to HF-ASD mean than to neurotypical adult mean; no model is adult-like on both tasks

Why It Matters

Highlights that even frontier VLMs lack robust, coherent social reasoning—critical for AI agents in human-facing roles like therapy, education, and customer service.

📬 Get the top 10 AI stories daily