Researcher predicts brainlike AGI alignment via LLM augmentation
What if the first transformative AI is just an augmented LLM? This researcher bets on it.
A veteran computational cognitive neuroscientist, who has been thinking about brainlike AGI since 2004 and working full-time on alignment for over three years, presents a research agenda that diverges from mainstream alignment work. Rather than studying current LLMs empirically (prosaic alignment) or theorizing about idealized agents (agent foundations), they advocate for a third path: carefully predicting the specific properties of the first AI that is genuinely transformative and capable of takeover. They argue that this AI will likely be a loosely brainlike AGI—an LLM augmented with human-like cognitive capacities such as continuous learning and executive function—and that understanding its mechanistic details in advance allows researchers to identify and address failure modes efficiently, even under tight timelines.
Drawing on 23 years in computational cognitive neuroscience, the researcher outlines how developers might augment LLMs to reach takeover-capable AI (TCAI) beyond mere scaling. They emphasize that such augmented systems would introduce new alignment challenges (e.g., from continuous learning or executive control loops) that differ from current LLM issues. While acknowledging other possibilities (like Steve Byrnes' 'brain in a box' scenario), they argue that the LLM-plus-cognition route is the most plausible path to AGI competent enough to control the future. The author calls for more explicit, mechanistic reasoning about this transition, noting that while more researchers are now engaging in it compared to three years ago, it remains underappreciated in the alignment community.
- The researcher proposes a third alignment approach: predicting the specific nature of the first transformative AI rather than studying current systems or idealized agents.
- They believe the most likely route to takeover-capable AI is LLMs augmented with continuous learning and executive function, which create new alignment challenges.
- With 23 years in computational cognitive neuroscience, they aim to anticipate failure modes mechanistically to enable interventions even under short timelines.
Why It Matters
For AI safety, predicting specific failure modes of brainlike AGI may yield more practical interventions than broad theories.