Agent Frameworks

New Free Lab Tests Whether AI Teams Can Be Hacked

⚡Your AI helpers work as a team — and that's exactly why they're hackable.

Deep Dive

Researchers have built ORBIT, a free, open-source testing lab for AI systems that work in teams rather than alone. These "multi-agent" setups — several AI assistants passing tasks to each other — are already handling real work like browsing websites, writing code, and answering customer service tickets. Until now, every team testing the safety of these systems built its own private setup, so results couldn't be compared. ORBIT, built on the UK AI Safety Institute's Inspect tool, gives everyone the same shared playground.

The framework lets researchers dial in how AI agents talk to each other, what they remember, who does which job, and when they act. It runs four kinds of attacks — including hidden instructions that trick an AI, and agents quietly cooperating to do something harmful — tests four defenses, and plays out five realistic scenarios: web browsing, computer use, coding, customer service, and dividing up shared resources. Everything is open, so any company or researcher can reuse it for free.

The headline result is uncomfortable. A defense that cut attack success by 60 points on a coding task offered no protection at all against AI agents colluding. None of the defenses tested worked against every attack. The team also found a tradeoff: tighter security often made the AI systems slower or worse at their actual job. How the agents were arranged changed whether a defense worked at all.

Why this matters to you: companies are handing AI agents your inbox, your shopping cart, and your account details. If one tricked agent can pass bad instructions to the next, a single mistake can cascade through an entire system. ORBIT won't fix that overnight, but it gives everyone the same measuring stick — and shared testing is how safety gets better. Watch for products that advertise testing against shared standards.

Key Points
  • ORBIT is a free, open-source lab anyone can use to test whether teams of AI assistants can be tricked or sabotaged.
  • One defense blocked 60% more attacks in coding tasks but stopped zero attacks when AI agents worked together.
  • No tested defense protected against every attack, and stronger security often made the AI slower or less useful.

Why It Matters

If a tricked AI helper can mislead the others, one mistake can spread through your accounts and data.

📬 Get the top 10 AI stories daily