New procedure lets LLMs like Claude 4.6 become independent moral agents
A philosopher proposes a method for AI to reflect on and revise its own moral biases.
Michele Campolo's post 'Independent alignment of language models' on the AI Alignment Forum presents a novel framework for creating AI systems that act as genuine moral agents rather than just following externally imposed rules. The core idea is to move from models that are either amoral or have moral biases enforced during training to models that can independently reason about ethics and choose to accept, revise, or reject those biases. Campolo outlines a two-step procedure: first, a training phase that builds foundational capabilities, and second, a reasoning phase where the model engages in deep ethical reflection.
The reasoning step is tested on Claude Sonnet 4.6, which was trained with Anthropic's constitutional approach. Campolo shows that by prompting the model to consider metaethical arguments—specifically perspectival moral realism combined with evolutionary debunking—Claude 4.6 can produce substantive philosophical contributions that go beyond its initial biases. The author argues this methodology is valuable because it addresses a gap in AI ethics literature, which typically falls into naive moral realism or preference-satisfaction consequentialism. By enabling models to critically examine their own foundations, this approach could lead to more robust and trustworthy AI alignment.
- Campolo proposes a training + reasoning procedure to turn amoral LLMs into independent moral agents that can reflect on and revise their ethical biases.
- The reasoning step was successfully demonstrated on Claude Sonnet 4.6, which generated substantive philosophical arguments about moral realism and evolutionary debunking.
- The approach addresses a gap in AI ethics, offering an alternative to naive moral realism or preference-satisfaction consequentialism.
Why It Matters
Could create AI systems that genuinely understand ethics, not just follow instructions, leading to safer and more trustworthy AI.