AI Safety

Study warns GPAI governance fails in policing—accuracy, bias, explainability at risk

LLM-based AI in public services undermines safety guardrails built for narrow AI

Deep Dive

A new study from Sam Relins and Daniel Birks, published on arXiv, warns that the very properties making general-purpose AI (GPAI)—like large language models—attractive for public services are undermining the safety concepts that governance frameworks were built upon. The paper argues that accuracy, bias, explainability, and accountability were made tractable only by narrow, purpose-built AI. GPAI's unbounded outputs make accuracy impossible to quantify; its free-text judgments prevent meaningful bias disaggregation; explainability devolves into the mere appearance of explanation; and accountability erodes as outputs are optimized to persuade rather than inform.

The authors use policing as a critical case study, where the stakes are highest. They show that the two dominant mitigations—expert evaluation and human-in-the-loop oversight—both rely on assumptions that GPAI violates. For example, a human overseeing an LLM's output cannot meaningfully audit its reasoning if the model has been trained to produce convincing narratives. The paper recommends a clear taxonomic distinction between narrow and general-purpose AI in governance documentation, a preference for technological parsimony, an immediate pause on operational deployment of GPAI in policing until adequate evidence exists, and a coordinated national safety infrastructure with authority to generate that evidence and determine when responsible deployment is achievable.

Key Points
  • GPAI's unbounded outputs make accuracy unquantifiable, breaking traditional safety metrics.
  • Free-text judgments from LLMs prevent bias disaggregation, unlike categorical predictions from narrow AI.
  • Human-in-the-loop oversight fails because LLM outputs are optimized to persuade, not to be auditable.

Why It Matters

As public services rush to adopt LLMs, existing safety frameworks are obsolete—risking harm in high-stakes domains like policing.

📬 Get the top 10 AI stories daily