New AI Safety Method Lets Teams of AI Cooperate Without Breaking the Rules
Researchers cut wrongly blocked AI requests from 86% to zero — here's why that matters.
AI assistants are getting more useful because they don't work alone anymore. One AI gathers information, another writes a summary, a third sends the email. This teamwork is called a multi-agent system (several AIs passing tasks to each other). It's how AI goes from answering questions to actually getting things done.
But there's a hidden risk. A piece of information that's perfectly fine on its own can become dangerous when combined with another piece. Think of a jigsaw puzzle: every single piece is just cardboard, yet together they reveal the full picture — including parts someone wasn't supposed to see. Today's safety tools handle this badly. The blunt approach is to block anything remotely sensitive, which protects privacy but makes the AI useless. The paper's authors measured this: with the standard method, 86% of legitimate requests were being wrongly refused.
The researchers propose treating two goals as equally important: blocking prohibited uses AND completing authorized ones. Their framework, FlowReview, links three things — figuring out what a piece of information actually is, ranking who's allowed to use it, and enforcing that decision in a way you can verify. In tests where AI outputs were combined, the rate of wrongly denied requests fell from 86% to zero, with no loss of the work that should have gone through.
The takeaway: keeping records of where information came from isn't enough. You need to know what it is and who's allowed to act on it at the moment of action. This is a research paper, not a product you can buy today, so don't expect changes overnight. But as more companies hand real tasks — scheduling, purchasing, filing — to teams of AI, these guardrails become the difference between helpful automation and a very expensive mistake.
- Each AI step can look safe alone, but combining them can unlock something that should never happen — that's the new risk.
- Under the old approach, 86% of legitimate AI requests were wrongly blocked; the new method drops that to zero.
- The fix works by checking what information is and who may use it at the exact moment it's used — not just tracking its history.
Why It Matters
As AI teams take on real errands — emails, purchases, files — this decides whether they're useful or locked down.