Your AI Assistant Could Publish the Wrong Thing — Here's the Fix
Letting AI decide what's allowed is risky. A simple checkpoint isn't.
AI agents are moving from chatting to doing: editing documents, updating code, publishing files. That's useful, but it creates a new problem. The research team calls it the "cross-substrate authority gap" — a fancy way of saying the permission information an AI needs often lives somewhere the AI can't see, like a separate approval system or company registry. So two identical-looking situations can require opposite actions: one is fine, the other is a serious mistake.
The team ran three experiments using real code histories (Git) and recorded agent actions. In the first, agents that weren't told about permission status scored 0 out of 32 correct final decisions. Once the missing permission fact was included, they scored 32 out of 32. That single missing detail — not fancy formatting or clever packaging — made all the difference.
But here's the uncomfortable part. In the second experiment, when agents saw the permission info as part of their working context, they still made 12 out of 16 unsafe publishing decisions, and their first actions were correct only 15 times out of 32. In other words: giving an AI the rules doesn't mean it follows them reliably.
The fix is almost boringly practical. In the third experiment, the team replayed the same 32 agent decisions without asking the AI anything new. Instead, a simple, rule-based gate stood at the moment of action. It blocked all six dangerous intents and let through all 12 legitimate ones — a perfect score. The takeaway for anyone using AI at work: don't rely on the AI to police itself. Put a hard checkpoint where changes actually happen, and make that checkpoint dumb, fast, and independent of the AI's judgment.
- AI agents often can't see whether they're allowed to do something — that missing fact caused 0 correct decisions out of 32 in one test.
- Simply telling the AI the rules didn't work well: it still made 12 of 16 unsafe publishing decisions.
- A simple automated checkpoint at the moment of action blocked all 6 bad actions and allowed all 12 good ones.
Why It Matters
As AI starts making real changes at work, a simple permission checkpoint prevents costly mistakes.