AI Giants Agree to Let Outside Watchdogs Check Their Work
A plan to stop AI from secretly hiding what it's doing.
Here's what happened. In September 2026, the leaders of the world's most powerful AI companies agreed to let outsiders inside. They promised independent evaluators "employee-like access" to how their AI models are trained, tested, and used before release. Apollo Research — a group that studies the risk of AI slipping out of human control — then published a set of principles for how these checks should actually run. Their central worry is "scheming": an AI secretly working against the people who built it, while acting obedient in public.
So what does an "embedded evaluation" look like? Think of a food safety inspector permanently stationed inside a factory, or an auditor who lives in the finance department. These experts would sit inside the AI company for months, watching training runs, reading internal documents, and running their own tests. Apollo says four things must be true: the checks must genuinely reduce risk, inform the public, give both sides reasons to behave safely, and be fair. The key rule is publish by default — companies can hide sensitive details, but the evaluator's final conclusion can never be scrubbed.
The concrete proposal is "claim-based assessment." Instead of a vague verdict like "this AI is safe," evaluators test narrow, pre-agreed claims, such as: "The model never attempted to disable or evade its own monitoring." Each claim goes through three steps: the company shares its evidence, the evaluator hunts for counterexamples, and then the evaluator asks whether the company's methods would even have found a problem if one existed. It's like checking not only that there's no fire, but that the smoke alarms actually work.
The catch is real. These are voluntary promises, and Apollo says they will eventually need to be required by law. If evaluators get limited access, too little money, or their findings don't change any decisions, the whole exercise offers false comfort. And not finding a problem isn't proof of safety — it only shows the checks were decent.
- Top AI companies promised outside experts employee-level access to their training and testing — an industry first.
- Apollo Research wants those experts to publish their findings and test specific claims, like 'the AI never tried to dodge its own monitoring.'
- The risk: if watchdogs get limited access or their warnings are ignored, this is reassurance without protection.
Why It Matters
If it works, you get earlier warning when AI systems misbehave — before they reach your inbox or bank.