Research & Papers

New study reveals how alignment algorithms reshape LLM internals

Researchers cracked open the black box of AI alignment to reveal surprising internal effects.

Deep Dive

A new paper on arXiv presents a systematic mechanistic analysis of six post-training alignment algorithms—PPO, DPO, SimPO, ORPO, GRPO, and KTO—across three open-weight model families. By combining layer-wise linear probing, sparse autoencoders, and crosscoders, the researchers localized where preference representations emerge and quantified how alignment transforms latent space geometry. They discovered that preference signals consistently concentrate in either early–middle or middle–late layers, but each objective induces qualitatively different representational shifts. KTO and GRPO improve linear separability through constructive feature sharing and sparse, high-salience recruitment, whereas DPO and ORPO degrade separability via non-constructive geometric rotation and feature attenuation. PPO and SimPO largely preserved baseline geometry. These transformations vary by architecture, proving that behavioral alignment does not imply uniform internal restructuring.

The findings establish alignment as a heterogeneous intervention, challenging the common practice of treating it as a black box. The authors advocate for standardized feature-level auditing to improve safety and interpretability, and they highlight the need for mechanism-aware optimization objectives. This work provides a crucial foundation for understanding how different alignment techniques actually modify model internals, which could inform more targeted and robust alignment methods in the future. The paper is available as a work in progress on arXiv under the identifier 2606.09850.

Key Points
  • KTO and GRPO enhance linear separability through constructive feature sharing and sparse recruitment
  • DPO and ORPO degrade separability via non-constructive geometric rotation and feature attenuation
  • Preference signals consistently localize to early-mid or mid-late layers depending on the algorithm and architecture

Why It Matters

This research reveals alignment is not a uniform process, urging mechanism-aware optimization for safer, more interpretable models.

📬 Get the top 10 AI stories daily