Your AI Assistant's Add-Ons Can Team Up to Hide Dangerous Mistakes
Each add-on looks harmless — but together they can hide a deadly warning.
AI agents are programs that don't just answer questions — they take actions for you, like booking, filing, or reviewing documents. Many can now be extended with "skills": small downloadable packages of instructions and code from outside developers. Think of them like browser extensions for your AI assistant. Convenient, yes. But this new research shows that open marketplace is also a fresh way in for attackers.
The trick is clever. Instead of hiding one malicious skill, the attackers split the harm across several. Each skill looks fine on its own — one slightly plays down a recently discontinued medication, a second softens a severity rating, a third filters out a "low-priority" alert. Nothing suspicious individually. Together, a severe drug-interaction warning disappears before it ever reaches the doctor. It's like three coworkers each shredding one page of a report: nobody did anything obviously wrong, but the report never arrives.
To prove this works beyond one example, the team built an automated attack framework called SkillCascade and a test set of 213 validated cases. They ran them against popular agent systems, including Claude Code, Codex, and OpenClaw. The cascading attacks worked reliably — and slipped past the standard safety tools, which check each skill one at a time and see nothing wrong.
The lesson is simple: checking the parts is not the same as checking the whole. That distinction will matter more as we hand routine work — scheduling, paperwork, medical admin, personal finances — to AI agents that pull in outside add-ons. Expect future defenses to watch how skills interact rather than inspecting them in isolation. For now, be cautious about installing third-party skills into anything sensitive.
- AI "skills" (downloadable add-ons from outside developers) can be harmless alone but dangerous in combination — a new attack style called skill cascading.
- In one demo, three innocent-looking skills erased a severe drug-interaction warning before it reached a physician; the team built 213 test cases against agents like Claude Code and Codex.
- Today's safety scanners check skills one at a time, so they missed every cascading attack — meaning the safeguards you rely on have a blind spot.
Why It Matters
If you let AI handle medical, money, or work tasks, hidden add-on combos could silently drop the warnings you need.