AI Safety

New Research: AI That Acts Alone Needs Rules Set Before It Starts

Your AI can book, buy and email before you notice — who is accountable?

Deep Dive

WHAT HAPPENED: Four researchers published a philosophy paper arguing that our favourite safety measure for AI — having a human check and approve what it does — breaks down with "agentic" AI. That means software that plans, breaks a big goal into smaller steps, and then acts over hours or days. Think of an AI travel assistant that books your flights, rebooks them when prices change, and emails the hotel on its own.

WHY THE OLD APPROACH FAILS: Approving every single step defeats the whole point of automation — you may as well do the job yourself. But reviewing only the overall pattern is too blurry. Some harm only becomes visible after it piles up: a trading AI making thousands of tiny trades that look fine individually but wreck a portfolio together, or an email agent slowly committing your company to things it should not. By the time anyone notices, the damage is done.

THE PROPOSAL: Set the rules before the AI acts. The authors call this an "anticipatory" mode of oversight — essentially a constitution for the agent, written in advance. It spells out what is allowed, what is forbidden, and which situations must trigger a stop-and-ask-a-human moment. Those same triggers are what bring old-fashioned after-the-fact checking back into play. The authors say the rules should be refined over time: when you write them, while the AI is running, and when you review what it did.

THE BIGGER POINT — BLAME: The paper argues this creates a clear responsibility chain by design. Whoever sets the rules takes on a duty beforehand, and afterwards stays answerable for what the agent does — even if no single action was their fault — because they had the chance to put precautions in place. The authors also tackle obvious objections, such as whether this is just an illusion of control. The catch: this is an argument on paper, not tested software, so nobody yet knows how well it works in the messy real world.

Key Points
  • Agentic AI (software that plans and takes actions on its own) moves too fast and chains too many steps together for a human to approve each move
  • The suggested fix is to set boundaries and explicit "stop and ask me" triggers before the AI starts, rather than correcting it afterwards
  • It also answers the blame question: whoever deployed the agent stays answerable, even when no single step was clearly their fault

Why It Matters

As AI starts booking, buying and emailing on our behalf, someone has to own the mistakes — and it will not be the AI.

📬 Get the top 10 AI stories daily