AI Safety

Researchers Question 'Prompt Governance' – Can We Control AI with Natural Language?

New study reveals system prompts are unstable and contradictory as policy tools.

Deep Dive

Anna Neumann, Holli Sargeant, and Jatinder Singh from the University of Cambridge have published a critical paper, 'Prompt Governance? On Governing Technologies Governed by Natural Language,' accepted at ACM FAccT 2026. The researchers systematically examine the assumption that natural language prompts can effectively shape and control generative AI behavior. They focus on system-level instructions—such as end-user guidelines, developer specifications, and system prompts—that policymakers increasingly treat as accessible levers for governance. The paper analyzes how researchers discuss these instructions in the literature, finding a fragmented landscape with varying and often contradictory claims about what goals system-level instructions can achieve. The authors then incorporate analysis of two major policy frameworks: the US Executive Order on Preventing Woke AI and the EU General-Purpose AI Code of Practice, both of which position system prompts as stable, interpretable control mechanisms.

The central finding is a significant misalignment between research and policy. While policymakers assume prompts function reliably and predictably across contexts, the literature reveals that their effectiveness is highly context-dependent, brittle, and subject to contradictory interpretations. The authors develop a typology of claims made about system-level instructions, highlighting how divergent perspectives complicate any uniform governance approach. They argue that given these misalignments, careful consideration must be given to prompt governance approaches—especially as regulators worldwide rush to mandate prompt-based interventions. The paper’s implications extend beyond large language models to any technical system that uses natural language as a control mechanism, questioning the very viability of governing AI through natural language. This work serves as a cautionary note for policymakers and industry alike: prompt governance may be far less reliable than currently assumed.

Key Points
  • Neumann et al. (FAccT 2026) find fragmented LLM literature with contradictory claims about what system-level prompts can achieve.
  • US Executive Order on Preventing Woke AI and EU General-Purpose AI Code of Practice treat prompts as stable controls, misaligned with research evidence.
  • Authors warn that governing AI through natural language is unreliable, with impacts on future regulatory frameworks and technical systems.

Why It Matters

As regulators worldwide mandate 'prompt governance', this study exposes critical flaws in assuming natural language can reliably control AI.

📬 Get the top 10 AI stories daily