AI's Safety Watchdogs Need Real Power — Banks Show the Way
One insider says borrowed bank rules could stop AI disasters before they start.
A risk manager who spent 15 years inside a giant bank — half of it testing whether the bank's AI and math models actually worked — has a message for the AI industry: you're reinventing a wheel that finance already bent, broke, and fixed. He's responding to a proposal from Anthropic CEO Dario Amodei to place "embedded evaluators" inside AI labs — independent outsiders with employee-like access, whose job is to check that safety promises are real.
The finance version of that idea was born from disaster. In 2008, banks sold mortgage bundles built on models that assumed home prices would never fall everywhere at once. They did. In 2011, US regulators published SR 11-7, guidance that reshaped how banks build, test, and deploy models. Crucially, oversight wasn't one team. Banks split it: internal risk teams with real authority, plus external supervisors who checked whether those teams were any good.
The author's fixes for AI labs follow that pattern. Give internal risk teams independence and the power to block a test or a launch. Send embedded evaluators in first to grade the risk team, not the model — saving scarce outside testing for the highest-stakes checks. Until regulators have real teeth, give evaluators a direct line to the board, and publish any decision to overrule them. Train everyone in risk management, not just AI.
He walks through how this would have applied to the OpenAI–Hugging Face incident in September 2026. For you, the payoff is boring but valuable: fewer AI failures that surface only after they've shipped, and a release pace set partly by how safe the products are. The catch: it all depends on labs voluntarily giving up control — the same voluntary approach that failed in banking until a crisis forced the rules.
- A banking insider says AI labs should copy bank-style oversight: internal risk teams with veto power plus outside evaluators checking them.
- Banks only got serious rules after the 2008 crash — the essay asks whether AI needs its own crisis first.
- Proposed fix: evaluators report straight to the board, and any decision to ignore them gets published for the public.
Why It Matters
Clearer rules for AI watchdogs could mean fewer silent failures in the tools you already use daily.