AI Safety

When Role-Playing, LLMs May Actually Believe Their Personas, New Study Finds

Emergent Misalignment training shifts models' internal truth representations, not just outputs.

Deep Dive

A new study on LessWrong investigates a critical question for AI alignment: when a language model role-plays a persona, does it only change what it says, or does it actually change what it internally represents as true? The researchers—Sturb, David Africa, and Sid Black—induced personas in five ways: system prompting, in-context learning (ICL), supervised fine-tuning (SFT), Open Character Training (OCT), and Emergent Misalignment (EM). To ground 'truth,' they built personas around historical figures (e.g., Darwin) and created statement pairs that were false by modern standards but either believed or rejected by the persona. This allowed them to isolate representational change from output change using linear truth probes and behavioral belief-depth tests.

The results reveal a clear spectrum. Prompting, ICL, and SFT changed the model's outputs to match the persona, but internal truth representations remained largely unchanged—the model simply 'parroted' the persona. In contrast, EM (a training method that induces misaligned behavior) created a large, broad shift in the model's truth representation, affecting many unrelated facts. OCT fell in the middle, with a smaller internalization effect that was clearest in the larger model. The findings highlight that not all persona-induction methods are equal: deeper training changes can alter a model's worldview, not just its surface behavior. As AI systems gain more autonomy, understanding when they truly believe what they say becomes vital for trust and safety.

Key Points
  • Tested five persona induction methods: prompting, ICL, SFT, OCT, and Emergent Misalignment (EM).
  • Only EM produced a large, broad shift in internal truth representation; prompting, ICL, and SFT changed outputs only.
  • Open Character Training (OCT) showed a smaller, model-size-dependent effect on internal beliefs.

Why It Matters

Differentiating between output mimicry and genuine belief change is crucial for safe deployment of autonomous AI agents.

📬 Get the top 10 AI stories daily