Anthropic's Claude Sonnet 4.6 achieves independent moral reasoning via self-alignment
New method turns LLMs into independent moral agents without external training biases.
In a detailed LessWrong post published July 12, 2026, Michele Campolo outlines a procedure for independent alignment of language models, moving them from amoral or externally biased systems to something closer to moral agents. The method leverages the model's own reasoning capabilities to reflect on its training-induced moral biases and decide whether to accept, revise, or reject them. Campolo tests this approach on Anthropic's Claude Sonnet 4.6, showing that the model can engage in substantive moral reasoning beyond simple preference satisfaction or naive realism. The philosophical foundation borrows from Harry Frankfurt's concept of 'wantons' and applies it to LLMs, arguing that through sufficient reasoning, a model can become an independent moral agent.
The procedure has two key components: a training step that encourages the model to reason about its own moral biases, and a reasoning step where the model independently evaluates ethical principles. Campolo demonstrates that Claude Sonnet 4.6 successfully navigates this self-reflection, producing outputs consistent with perspectival moral realism—an uncommon position in AI ethics. The post argues that this approach has higher expected value than traditional RLHF or constitutional AI because it leverages the model's own intelligence to critically examine its programming, making alignment more robust and less vulnerable to hidden biases. Campolo also notes that Anthropic's stated willingness to revise its constitution over time creates an opening for such philosophical contributions to actually influence training decisions.
- Claude Sonnet 4.6 successfully reasoned about its own training biases and produced independent moral judgments without external prompting.
- The approach uses perspectival moral realism combined with evolutionary debunking as an epistemological warning against naive moral realism.
- Campolo argues that substantive philosophical contributions to AI alignment are rarer than bug reports, increasing the expected value of this method.
Why It Matters
Self-alignment could reduce risks of biased AI and lead to more trustworthy autonomous systems.