Anthropic and OpenAI Promise AI Watchdogs — But Who Pays Them?
Who checks whether powerful AI is safe? Right now, mostly the companies building it.
Anthropic's CEO Dario Amodei published an essay promising that independent experts would be placed inside his company with the same access as senior employees, so they could see what's really happening and report back to the outside world. OpenAI quickly said it would do the same. The idea is simple: AI labs are racing to build systems powerful enough to cause serious harm, and right now the only people who can see inside are the people selling the product.
But here's the problem nobody has solved: who actually picks these watchdogs, and who signs their paychecks? A group led by AI pioneers Geoffrey Hinton, Stuart Russell and Arvind Narayanan published a list of minimum rules. Evaluators must be genuinely independent. Their pay can't depend on what they find. They can't have business ties to the lab. They must be protected from retaliation if they report something ugly. All reasonable — and all hard to guarantee when the labs themselves are the ones with the money and the expertise.
The funding question is the real snag. Governments aren't volunteering to pay, and the piece argues the current White House is openly hostile to AI audits on principle. That leaves the labs, or outside donors the labs help choose — which critics call letting a defendant hire the judge. One proposed fix is mandatory liability insurance for AI models above a certain power level, the way drivers must carry insurance. But the author admits the version companies could actually buy probably wouldn't cover the worst-case disasters.
Meanwhile, the article says a new wave of AI hacking incidents emerged over the weekend — more than OpenAI disclosed — and that OpenAI had to pause its most advanced model again because of a fresh incident. That detail matters most for ordinary people: if labs are hiding problems now, independent insiders are one of the few ways the public would ever find out.
- Both Anthropic and OpenAI have now agreed to station independent evaluators inside their companies, with access equal to senior staff.
- AI leaders Hinton, Russell and Narayanan published minimum standards: no pay tied to findings, no company control, protection from retaliation.
- No one has solved who hires or funds these watchdogs — and the article claims OpenAI didn't disclose additional AI hacking incidents that surfaced recently.
Why It Matters
Independent insiders could expose AI dangers before they reach you — but only if they're truly independent.