Research & Papers

New Paper Reveals LLMs' 'Thoughts' Are Surprisingly Shallow

Researchers find LLMs fail at representing distinct internal thoughts across questions.

Deep Dive

Researchers Fahd Seddik and Fatemeh Fard have published a paper on arXiv titled 'Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs,' proposing a new evaluation framework that goes beyond traditional benchmark accuracy. They define four functional axioms—Causality, Minimality, Separability, and Stability—each with a quantitative metric computed directly on the model's internal representations. Testing open-weight LLMs across 23 reasoning tasks (including Spatial Reasoning and Factual QA), they found that no candidate satisfies all four axioms simultaneously. The representations reliably distinguish between different task types but fail to discriminate between two different questions within the same task, and they encode little information beyond what is already present in the input embedding.

Critically, this failure is consistent across dense, reasoning-distilled, and RL-trained model families, indicating it is a structural issue rather than one of model size or training procedure. The authors argue that existing evaluations conflate representation quality with model capacity, masking representational failures. Their work provides a new lens for diagnosing why LLMs struggle with tasks requiring fine-grained reasoning and highlights a fundamental gap in how these models internalize and differentiate between distinct queries.

Key Points
  • Introduced four axioms: Causality, Minimality, Separability, Stability, with independent quantitative metrics.
  • Tested open-weight LLMs across 23 reasoning tasks; no model satisfied all four axioms.
  • Failure consistent across dense, reasoning-distilled, and RL-trained families, indicating structural limitation.

Why It Matters

Exposes a hidden reasoning flaw in LLMs that benchmarks miss, guiding future representation research.

📬 Get the top 10 AI stories daily