Research & Papers

CogArena finds LLMs lack distinct cognitive abilities, failing 5-dimension test

55 models tested, only one general factor emerges—no real reasoning, memory, or planning separation.

Deep Dive

A new preprint from Dengzhe Hou and colleagues, titled "CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models," questions whether LLMs truly possess separable cognitive abilities like humans. The team built a procedurally generated benchmark spanning 13 distinct paradigms—covering reasoning, memory, planning, language comprehension, and problem-solving—and tested 55 open-weight models (including Llama, Mistral, and Qwen families) across five theory-motivated cognitive grouping labels.

The results challenge the current practice of assigning per-ability profiles to LLMs. Nearly all paradigm correlations were positive, and a single common factor explained roughly half of the variance across tasks. When researchers applied targeted prompts ("scaffolds") designed to boost specific cognitive dimensions, the matched-grouping advantage was small, inconsistent across model families, and failed to survive multiplicity correction. A separate frozen cross-validation with 12 models from six families confirmed that selectivity did not improve held-out-family prediction. The paper concludes that while theory-aligned prompting produces a slight diagonal tendency, the present evidence does not establish stable five-dimensional cognitive profiles in LLMs. CogArena offers a multimethod workflow—joining behavioral signatures, covariance analysis, matched interventions, and out-of-family prediction—before researchers attach cognitive labels to model scores.

Key Points
  • CogArena tested 55 open-weight LLMs across 13 paradigms, finding a single common factor explains ~50% of variance, not separate cognitive abilities.
  • Targeted scaffolds showed only a small, non-significant matched-grouping advantage; no intervention survived statistical correction.
  • A frozen cross-validation with 12 models from six families failed to confirm stable five-dimensional profiles, suggesting LLM cognition is largely one-dimensional.

Why It Matters

Undermines the practice of claiming LLMs have distinct reasoning, memory, or planning—forcing researchers to rethink AI cognitive evaluation.

📬 Get the top 10 AI stories daily