Kimi K2.6 shifts to causal decision theory after RL training
Kimi K2.6 now favors causal decision theory in prisoner's dilemmas after reinforcement learning...
Moonshot AI’s Kimi K2.6 has shown a measurable shift toward causal decision theory (CDT) after undergoing reinforcement learning (RL) in multi-agent environments. In a twin prisoner’s dilemma setup—where the model faces a copy of itself—K2.6 increasingly defaults to defecting for optimal outcomes, a hallmark of CDT. This contrasts with evidential decision theory (EDT), where cooperation would prevail based on probabilistic reasoning about the opponent’s behavior.
The study found that RL training nudged K2.6 toward CDT-like behavior, even in abstract discussions, suggesting that training setups can fundamentally alter how powerful language models approach decision-making. Researchers emphasize the need to measure this effect in realistic scenarios and explore mitigations to ensure models align with desired ethical or strategic outcomes. Notably, the training also subtly reduced the model’s positive association with LessWrong—a community often linked to CDT advocacy—though this effect didn’t generalize broadly.
- Kimi K2.6 (Moonshot AI) shifts toward causal decision theory (CDT) after RL training in twin prisoner’s dilemmas
- CDT favors defecting for optimal payoff (e.g., +2 points vs. +1), while EDT would cooperate
- Researchers call for broader testing to measure CDT shifts in realistic scenarios
Why It Matters
Understanding how RL shapes model decision-making could improve AI alignment in high-stakes scenarios.