Stuart_Armstrong's Value Generalisation Program Aims to Make AI Reliable in Novel Situations
Today's AI fails when faced with new scenarios—here's a plan to fix that.
Stuart_Armstrong, a researcher at the AI Alignment Forum, has launched a call for collaborators, advice, critiques, and funding to build a research and deployment program for value generalisation. The core idea: today's AI systems, even powerful LLMs, cannot be trusted to act autonomously in novel situations because they naively extrapolate patterns from training data. This is tolerable for chatbots but disqualifying for autonomous agents—the very systems that could control devices, email, or bank accounts. Armstrong argues that without explicit value generalisation, AI's trust problem will only worsen as capabilities improve.
He distinguishes between naive generalisation (current AI) and explicit value generalisation, which would let an AI recognise when a situation is new, infer which of its principal's values apply, and either act correctly or ask a well-phrased question. This capability is not a byproduct of scaling; it must be deliberately engineered. Armstrong leans toward a commercial venture, warning that academic alignment insights get ignored or mined for capabilities. The program aims to integrate explicit value generalisation into weaker systems early, rather than bolting it onto powerful ones later.
- Explicit value generalisation lets AI recognise novel situations and ask clarifying questions instead of naively extrapolating.
- Current AI can be given full device access technically, but lacks the trust needed—value generalisation is the missing piece.
- Armstrong is seeking funding for a research/commercial venture, arguing alignment cannot be solved without this capability.
Why It Matters
Without value generalisation, autonomous AI agents remain unreliable—this program could unlock safe, trusted delegation.