Research & Papers

LLMs struggle with diagrams in statics: new study reveals reasoning gap

LLMs ace text-only statics problems but accuracy plummets when diagrams enter.

Deep Dive

A new study from researchers Tanner Culleton and Hung-Fu Chang systematically investigated how large language models (LLMs) handle engineering statics problems—a topic largely overlooked in prior AI education research. Instead of feeding textbook questions directly to an LLM, the team used a model distillation process to extract 25 text-only statics questions from ChatGPT. They then created two additional datasets: one adding diagrams to each question, and another modifying the numerical values while keeping diagrams. The goal was to isolate the factors affecting LLM problem-solving ability.

Results reveal a clear performance gap: LLMs solved text-only statics questions with reasonable accuracy, but accuracy dropped sharply once diagrams were introduced, especially for problems requiring multi-step reasoning. Further analysis suggests the bottleneck is not image recognition (the models could parse visual data) but rather the ability to consistently apply extracted visual information across successive reasoning steps. The paper, presented at the Engineering and Technology Symposium 2026, highlights that current LLMs still struggle with tasks that require integrating visual and textual cues over multiple inference stages—a critical skill in fields like mechanical engineering.

Key Points
  • 25 text-only statics questions were distilled from ChatGPT, then augmented with diagrams and varied numeric values.
  • Accuracy dropped significantly when diagrams were added, especially for multi-step problems.
  • Performance decline stems from difficulty in applying visual information across consecutive reasoning steps, not image recognition limits.

Why It Matters

Highlights a key weakness in LLMs for engineering education: visual+textual multi-step reasoning still fails.

📬 Get the top 10 AI stories daily