AI Safety

Anthropic's Claude develops internal 'global workspace' for reasoning

Claude's J-space reveals internal thoughts it can report and control.

Deep Dive

Anthropic researchers have discovered that Claude, their language model, spontaneously developed an internal 'J-space'—a small collection of neural patterns that behave like a global workspace for cognition. These patterns are tied to specific concepts (e.g., a pattern for 'France') but don't appear in the model's output; they activate silently when Claude is thinking about that concept. Unlike most of Claude's internal processing, J-space representations are reportable (Claude can tell you what it's thinking), controllable (it can start or stop thinking about something on demand), and causally essential for multi-step reasoning tasks. The team used Jacobian-based methods to isolate these patterns and found they mediate flexible use of knowledge—once 'France' lights up, Claude can answer questions about its capital, currency, or continent. The work draws inspiration from neuroscience's global workspace theory, which describes how conscious access works in the brain.

Crucially, the J-space was not engineered but emerged naturally during Claude's training. Despite its small size (relative to the model's billions of parameters), it is responsible for higher-order cognition—when the researchers suppressed J-space activity, Claude could still speak fluently and recall facts but lost its ability to reason step-by-step or follow complex instructions. This discovery has major implications for AI interpretability: it provides a direct window into a model's internal reasoning, separate from what it outputs. It also raises questions about whether such internal 'workspaces' are a general property of large language models and how they relate to consciousness. For AI safety, being able to monitor and steer J-space activity could help ensure models reason transparently and avoid hidden reasoning that might lead to unintended behaviors.

Key Points
  • J-space patterns are concept-specific internal activations that don't appear in Claude's output but are reportable on request.
  • These patterns emerge naturally during training and causally mediate multi-step reasoning, not fluent generation.
  • Suppressing J-space leaves basic language abilities intact but eliminates higher-order cognition like silent internal reasoning.

Why It Matters

Reveals a hidden cognitive layer in LLMs, enabling better interpretability and safety through monitoring internal reasoning.

📬 Get the top 10 AI stories daily