Kasneci & Kasneci Reveal Hidden AI Safety Failures Beyond Obvious Harms
The real AI safety threats are hidden in socio-technical systems, not just model outputs.
A new paper by Gjergji Kasneci and Enkelejda Kasneci, published in AI and Ethics (2026) and posted on arXiv, argues that AI safety discourse is dangerously incomplete. While current efforts focus on obvious harms, dramatic misuse, and hypothetical catastrophic scenarios, the authors contend that many of the most consequential failures are quieter: plausible rather than spectacular, distributed across components, and normalized by workflows before they are recognized as hazards. The central challenge, they say, is not only whether a model emits a harmful response, but whether the broader socio-technical system preserves the conditions for errors to remain visible, contestable, containable, and recoverable.
The authors propose a five-layer framework to diagnose these hidden risks: (1) epistemic integrity—honest representation of evidence and uncertainty; (2) control integrity—robust authority, permissions, and action boundaries; (3) temporal integrity—safety across sessions, memory updates, and deployment drift; (4) organizational integrity—capacity to audit, assign responsibility, and intervene; and (5) ecosystem integrity—preserving the information environment for future oversight. They identify under-recognized risk patterns including overreliance, uncertainty and legitimacy laundering in retrieval, prompt injection, reward hacking, memory poisoning, evaluation deception, fictional human oversight, synthetic evidence pollution, and model collapse. The paper concludes with design and governance recommendations and a research agenda for shifting AI safety from model-centric evaluation toward socio-technical reliability.
- The paper identifies five layers of hidden safety failures: epistemic, control, temporal, organizational, and ecosystem integrity.
- Specific under-recognized risks include uncertainty laundering, memory poisoning, evaluation deception, and synthetic evidence pollution.
- The authors recommend shifting from model-centric evaluation to socio-technical reliability and governance.
Why It Matters
This paper exposes systemic blind spots in AI safety that could lead to gradual, compounded failures in real-world deployments.