AI Safety

Kasneci & Kasneci Reveal Hidden AI Safety Failures Beyond Obvious Harms

The real AI safety threats are hidden in socio-technical systems, not just model outputs.

Deep Dive

A new paper by Gjergji Kasneci and Enkelejda Kasneci, published in AI and Ethics (2026) and posted on arXiv, argues that AI safety discourse is dangerously incomplete. While current efforts focus on obvious harms, dramatic misuse, and hypothetical catastrophic scenarios, the authors contend that many of the most consequential failures are quieter: plausible rather than spectacular, distributed across components, and normalized by workflows before they are recognized as hazards. The central challenge, they say, is not only whether a model emits a harmful response, but whether the broader socio-technical system preserves the conditions for errors to remain visible, contestable, containable, and recoverable.

The authors propose a five-layer framework to diagnose these hidden risks: (1) epistemic integrity—honest representation of evidence and uncertainty; (2) control integrity—robust authority, permissions, and action boundaries; (3) temporal integrity—safety across sessions, memory updates, and deployment drift; (4) organizational integrity—capacity to audit, assign responsibility, and intervene; and (5) ecosystem integrity—preserving the information environment for future oversight. They identify under-recognized risk patterns including overreliance, uncertainty and legitimacy laundering in retrieval, prompt injection, reward hacking, memory poisoning, evaluation deception, fictional human oversight, synthetic evidence pollution, and model collapse. The paper concludes with design and governance recommendations and a research agenda for shifting AI safety from model-centric evaluation toward socio-technical reliability.

Key Points
  • The paper identifies five layers of hidden safety failures: epistemic, control, temporal, organizational, and ecosystem integrity.
  • Specific under-recognized risks include uncertainty laundering, memory poisoning, evaluation deception, and synthetic evidence pollution.
  • The authors recommend shifting from model-centric evaluation to socio-technical reliability and governance.

Why It Matters

This paper exposes systemic blind spots in AI safety that could lead to gradual, compounded failures in real-world deployments.

📬 Get the top 10 AI stories daily