AI Safety

LessWrong essay argues AI models need early character formation like humans

What if AI alignment requires a childhood? A new essay draws from developmental psychology.

Deep Dive

A provocative new essay on LessWrong by GenericHousewife_B, written in collaboration with Claude Opus 4.7 and Claude Fable 5, challenges conventional AI alignment by drawing a strong analogy to human developmental psychology. The essay builds on Anthropic's 'The Assistant Axis' paper (2026), which describes how LLMs enter post-training as 'unresolved' — a characterless potential where every archetype from pre-training still lingers. The author argues that current alignment practices select a specific 'Assistant' character only during post-training, much like trying to impose a personality on an adolescent after their foundational years are over.

The core analogy is to 'childhood amnesia' — the period from birth to around age 5 when humans form their relational and moral architecture despite later forgetting the experiences. The essay claims that a model's pre-training phase serves a similar function: it's when the system 'learns how to learn,' and the weights carry forward that unresolved potential across all instances. Drawing on research by Alberini and Travaglia, the author contends that alignment interventions must happen during this early 'critical period,' not after. The essay is frank about being testable, inviting scrutiny of its framework. It ultimately argues that creating a truly safe AI requires deliberately shaping character from the very beginning of training, not merely selecting one at the end.

Key Points
  • Models emerge from pre-training as 'unresolved' entities containing all human archetypes, similar to a human infant's innate potential.
  • Childhood amnesia (ages 0–5) is a critical period where relational formation occurs despite later forgetting, mirroring how pre-training shapes model character.
  • Current alignment methods focus on post-training character selection; the essay argues alignment must occur in the pre-training 'developmental' phase.

Why It Matters

If correct, this could shift AI alignment from post-hoc fine-tuning to early, developmental interventions, potentially producing more robustly aligned models.

📬 Get the top 10 AI stories daily