53-Trait Persona Vectors Expose What Open-Weight LLMs Hide
Systematic audit of two models reveals default helpfulness and hidden capacity for hallucination.
In a landmark paper, researchers Winston Zeng, Ali Emami, and Jinho Choi present the first large-scale application of persona vectors to audit open-weight LLMs. They compiled a 53-trait inventory spanning four behaviorally distinct domains—agency, clinician, generic traits, and antisocial tendencies—and labeled each trait as natural (expressed at baseline), steerable (amplifiable via activation steering), or intractable (resistant to standard extraction). The study found that both tested models defaulted to helpful, task-oriented behavior. All nine agentic traits were natural, and the models' default clinician behavior aligned with a board-certified psychologist's independent desirability judgments on 16 out of 17 traits. Steering produced its largest gains precisely on traits these defaults exclude: hyperbole, hallucination, and sycophancy.
Further analysis of 171 generic-trait pairs revealed a critical asymmetry: two steerable traits could collapse the composition, but pairs involving a default trait never did. Where standard extraction failed on a trait like 'evil,' a vector transferred from a fine-tuned variant still recovered it, with the residual refusals appearing inside the model's chain-of-thought. The work reframes persona vectors not as a set of controls but as a probe of behavioral organization—revealing what models express, suppress, and resist. This systematic approach offers a powerful new lens for alignment research, safety auditing, and understanding the latent structure of LLM behavior.
- First systematic application of persona vectors at scale, compiling 53 traits across four behavioral domains.
- Both models default to helpful, agentic behavior; clinician responses match expert desirability on 16/17 traits.
- Steering amplifies suppressed traits like hyperbole and hallucination; 'evil' resists extraction but is recoverable via fine-tuned variant vectors.
Why It Matters
Persona vectors offer a new, systematic way to audit and understand LLM behavior beyond prompts—critical for safety and alignment.