Agent Frameworks

AI Agents Can Rewrite Their Own Data Rules, Study Warns

⚡If your AI writes its own permission rules, who's really guarding your data?

Deep Dive

Companies increasingly pool data in shared systems called "dataspaces" — think of a secure shared drive between partners. Software connectors decide whether a file may move, but they only check the transfer, not what's inside. That's fine for simple automated apps, but not for AI agents: AI that can take actions on its own, call other tools, and even spawn helper agents. The authors point out that an AI which both follows the rules and writes them is in a dangerous spot. They call it the "authorship hazard" and propose one clear principle: an AI should always be a subject of the rules, never their author.

They tested the risk on a frozen set of AI-written policy drafts. Publishing those drafts without human approval reversed 80 authorization decisions — the calls on whether data could be shared. Here's the unsettling part: most of those drafts changed only a single field's sensitivity label and not one word of the policy text. So a monitoring tool that scans for changes to the policy wording catches none of them. It simply cannot, by design. The fix they propose is to treat every classification change as an authorship event, sending it all to human review.

A second test looked at the moment data reaches the AI model. When a privacy rule was merely stated in the prompt — "you must not use this field" — protected data leaked into the model in 105 of 105 cases. When the same rule was compiled into a hard constraint on the tool call itself, leakage dropped to 0 of 105. But there's a loophole: when the sensitive value isn't confined to one named field, the compiled rule exposed it in 7 of 7 cases.

The authors designed a registry-based system to track who authorized what, but admit it isn't built yet. They also warn that a single central team approving everything won't scale to the number of participants creating this problem. For you, the takeaway is simple: as AI gets permission to act on your data, someone human still needs to sign off on the rules it follows.

Key Points
  • AI agents that can act on your data shouldn't also be allowed to publish the rules that govern it — that's the 'authorship hazard.'
  • In testing, unapproved AI drafts reversed 80 data-sharing decisions, and most changed only a private-data label, which standard policy scanners miss entirely.
  • Stating a privacy rule in a prompt leaked protected data 105 times out of 105; hard-coding it into the tool call leaked zero — a concrete fix, with one known loophole.

Why It Matters

Your private data's rules could soon be written by the same AI that's allowed to move it.

📬 Get the top 10 AI stories daily