AI Safety

Brainlike AGI Alignment: Predicting Failure Modes of LLM-Based Superhuman AI

A novel approach maps how LLMs will gain human-like cognition and become dangerous.

Deep Dive

The researcher argues that most alignment work falls into two camps: prosaic alignment (studying current systems) or agent foundations (abstract theory about idealized agents). Both share hidden assumptions—either that future AI will resemble today's models, or that we must solve the fully general alignment problem without knowing the AI's form. They advocate for a third, neglected approach: carefully predicting the properties of the first transformative AI in mechanistic detail, so we can preempt failure modes and design targeted interventions.

Drawing on their rare background in computational cognitive neuroscience, the researcher predicts the first TCAI will be an LLM augmented with human-like cognitive capacities—specifically continuous learning and executive function. They emphasize this is the most probable route to an AI capable of controlling humanity's future, not certain but more plausible than alternatives like a “brain in a box.” They call for more researchers to engage in such mechanistic speculation about the LLM-to-TCAI transition, noting that while interest has grown in three years, it remains insufficient for comfort.

Key Points
  • Proposes a third alignment approach beyond prosaic alignment and agent foundations: predicting the actual form of transformative AI.
  • Foresees first takeover-capable AI as an LLM enhanced with human-like continuous learning and executive function systems.
  • Uses computational cognitive neuroscience to bridge neuroscience and AI, offering a unique angle on failure mode prediction.

Why It Matters

Could shift alignment research toward the most likely superhuman AI architecture, increasing chances of safe deployment.

📬 Get the top 10 AI stories daily