Research & Papers

Fine-tuned LLMs discover 1 in 5 UK police logs show mental health issues

Researchers tested open-weight LLMs on 3,000 de-identified incident narratives with surprising accuracy gaps.

Deep Dive

Researchers Sam Relins and Daniel Birks adapted a fine-tuned LLM classification pipeline—originally trained on open-source US police data—to estimate the prevalence of four vulnerability indicators (mental ill health, substance misuse, alcohol dependence, and homelessness) in nearly 3,000 de-identified UK police incident logs. The pipeline combined repeated model inference, label aggregation, structured human review, and statistical correction, running entirely on a locally hosted open-weight LLM to meet secure police environments. Results showed mental ill health indicators in roughly one in five incidents, with lower prevalence for the other categories.

However, naive deployment proved unreliable: single-pass classifications were unstable, and aggregated outputs systematically over-assigned indicators relative to human judgment. Correcting these biases required substantial human input and statistical adjustment, leaving considerable uncertainty. At the population level, defensible estimates were achievable but resource-intensive; at the individual level, errors remained frequent and unpredictable, limiting suitability for operational decisions. The study highlights both the potential and constraints of LLM-based measurement in high-stakes applied settings.

Key Points
  • Fine-tuned open-weight LLM analyzed 3,000 de-identified UK police incident logs for vulnerability indicators.
  • Mental ill health appeared in ~20% of incidents, but single-pass classifications were unstable and aggregated outputs over-assigned indicators.
  • Correcting biases required substantial human review and statistical adjustment; individual-level errors made operational use unreliable.

Why It Matters

Shows LLMs can scale vulnerability detection in policing but need careful oversight before deployment in sensitive decisions.

📬 Get the top 10 AI stories daily