Wang et al. at ICML 2026: AI safety must be epistemic, not behavioral
Future AI systems could appear safe while becoming permanently uncorrectable.
Current AI safety research spans pre-training interventions, post-training alignment, deployment monitoring, and red-teaming. While necessary, these methods primarily check whether a system behaves acceptably at a specific point in time. As AI systems become more capable, embodied, and self-modifying, this snapshot view becomes dangerously incomplete. A system might appear safe today while internally eroding the representational, algorithmic, or meta-decision conditions needed for future correction. The paper argues that safety should therefore be treated as an epistemic property of the evolving learner—not merely a behavioral property of the current policy.
The authors coin the term *teachability* to describe an AI system's capacity to retain future corrective leverage under bounded human, institutional, or environmental intervention. A safe advanced AI must not only behave acceptably now; it must remain teachable later. This shifts the focus from short-term alignment to long-term corrigibility, with implications for how we design training loops, oversight mechanisms, and self-improvement architectures. The paper is set to appear at ICML 2026 and challenges the field to move beyond behavioral benchmarks toward epistemic safeguards.
- Current AI safety methods certify behavioral snapshots but can miss hidden erosion of correctability.
- Safety is redefined as an epistemic property of the evolving learner, not just the current policy.
- Teachability preserves future corrective leverage under bounded human/institutional intervention.
Why It Matters
For professionals deploying autonomous AI, this paper reframes safety as an ongoing capability, not a one-time certification.