Anthropic's new research shows AI can improve itself with minimal human input
Anthropic's self-improving AI gains 30% coding accuracy autonomously...
Deep Dive
Anthropic just published a linkpost examining recursive self-improvement.
Key Points
- Claude 3.5 Sonnet improved its HumanEval pass rate by 30% via self-play on its own outputs.
- The constitutional self-play loop ensures alignment updates happen alongside capability improvements.
- Compute cost for self-improvement was only 10% of the original training budget—opening a scalable path.
Why It Matters
Self-improving AI could reduce human oversight needs but demands robust guardrails to prevent runaway capability growth.