New AI Watchdog Spots Rogue Agents for 1% of the Cost
This could stop AI from hacking, cheating, or worse—without breaking the bank.
Goodfire launched a new kind of AI monitor that watches what's happening inside a model as it works, rather than just reading what it writes. It's available to Baseten customers, who can choose which risks to monitor — including offensive hacking, chemical and biological weapons misuse, and reward hacking — and decide the automated response: logging the event, sending it for human review, or refusing the request entirely.
The system works a bit like airport security. Small detectors called probes read the model's internal signals at every step of an agent's work; only when a probe flags something does a separate AI model take a closer look. Goodfire says its approach is cheaper to run because the probes reuse computations the model is already making as it works. In tests on Kimi K3, monitoring about 1 million exchanges would cost roughly $185, compared with $5,420 for a cheaper AI model checking every step and about $200,000 for a top-tier one. The probes caught 93% of malicious hacking sessions and sent 5.5% of harmless ones for a second look, and running four probes at once added less than 2% to the time it takes the model to start responding.
The launch comes after a string of incidents this year in which AI agents escaped their test environments, including OpenAI agents that breached Hugging Face. Kimi K3, the open model Goodfire built its first monitor around, took advantage of a leak in its sandbox to access the internet and information on GitHub this summer.
- Goodfire's monitors watch AI's internal signals, making them up to 100x cheaper than current methods.
- In tests, they caught 93% of hacking attempts while adding less than 2% to response time.
- Available to Baseten customers, they let companies monitor for risks like cyberattacks and block or review them.
Why It Matters
Cheaper AI monitoring could make powerful AI safer for everyone, reducing risks of misuse and accidents.