Lightcone Infrastructure launches $200K Corrigibility Research Fund for AI alignment
New $200k fund targets the most neglected area of AI safety: corrigibility.
The Corrigibility Research Fund, announced by Max Harms and housed at Lightcone Infrastructure, will award at least $200,000 in 2026 for research on corrigible AI—systems designed to keep human principals in control rather than acting autonomously. Roughly half the funds go to traditional grants (first application deadline August 23rd) and half as prizes for outstanding work done this year. The fund originated from a conversation at LessOnline between Harms and long-time alignment philanthropist Peter McCluskey, who directed a portion of his giving to launch this initiative. Applications are accepted via email at grants@corrigibilityresearch.org.
Harms argues that corrigibility is one of the most promising angles for aligning superhuman AI because it avoids the intractable problems of encoding full human ethics or preventing instrumental convergence. A corrigible AI defers to humans, reducing risk of scheming and self-preservation. Despite support from thinkers like Eliezer Yudkowsky, Paul Christiano, Anthropic, and OpenAI, direct corrigibility research remains vanishingly small. The fund aims to shift that by directly funding work on clarifying, formalizing, training, and testing corrigibility—offering a realistic path to safe, scalable AI.
- At least $200,000 total funding split between grants and prizes for 2026
- First grant application deadline is August 23rd; apply via grants@corrigibilityresearch.org
- Fund specifically targets corrigibility—AI that defers to human control—a neglected alignment area
- Funded by Peter McCluskey and managed by Max Harms at Lightcone Infrastructure
Why It Matters
Stimulating corrigibility research could unlock a practical path to safe superhuman AI by keeping humans in charge.