Anthropic's AI Safety Watchdog Is Basically a Friend, Critic Says
The group checking AI safety shares an office — and friendships — with the company it checks.
Anthropic has promised something unusual: it will let an outside team, a nonprofit called METR, work inside the company almost like regular employees. Their job is to watch how Anthropic builds and trains its AI, flag problems, and confirm the company is keeping its safety promises. Anthropic's boss, Dario Amodei, compared this to bank examiners who sit inside banks to keep an eye on them. The company is doing this on its own — no law requires it.
A critic says the comparison falls apart. METR isn't really independent, the argument goes, because METR's own report admits some of its staff have close social ties to people at AI companies, and it works out of a shared research center that also hosts AI lab employees. By accounting rules, an auditor who had dated or lived with the people they were checking would have to step aside. METR also has no power to fine or punish Anthropic — and Anthropic can end the arrangement whenever it wants.
Why should you care? AI is quietly moving into hiring decisions, medical advice, banking tools, and the software you use at work. If the only safety check is a friendly group paid by goodwill rather than law, a problem could reach you before anyone with real authority notices. Bank regulators work for the government and can send people to prison for wrongdoing. METR staff, the critic notes, would mostly be hanging out with friends while drawing nonprofit salaries.
The catch cuts both ways. METR was surprisingly candid about its own conflicts, and unpaid volunteer scrutiny is better than none at all. But "trust us, our friend is watching" isn't the same as a rule with teeth. If you want real protection, it will have to come from governments, not from companies choosing their own reviewers.
- Anthropic voluntarily lets an outside nonprofit, METR, sit inside the company to monitor its AI safety work.
- METR's own report says some staff have close social ties to AI company employees, and it shares office space with AI lab workers.
- METR can't punish Anthropic, and Anthropic can cancel the arrangement whenever it likes — unlike a government bank regulator.
Why It Matters
Friendly, voluntary oversight may not protect you from AI mistakes baked into everyday tools and decisions.