New Plan Lets Outside Watchdogs Work Inside AI Companies
Independent experts would get employee-style access to the biggest AI labs.
Right now, when someone checks whether a powerful AI system is safe, they usually test the finished product from the outside — typing prompts into a chat box, like inspecting a restaurant by ordering takeout. A new paper argues that misses the dangers that actually matter, because those live inside the company: how the AI is used by staff, what it's allowed to touch, and who is watching it. So the authors propose "embedded assessments" — letting independent evaluators in as if they were employees.
Why does inside access matter? Imagine an AI assistant that can read company email, move money, or change code. Whether it's dangerous depends on things you can't see from outside: what data it can reach, which coworkers can switch it on, and what happens when it makes a mistake. Companies have already promised this kind of access, and the paper says it can work while keeping safety tight — evaluators get employee-level visibility, but under strict security controls that protect sensitive systems and company secrets.
The authors lay out seven design questions, from how long an assessment lasts to what happens if a serious problem turns up. Their main recommendations: start now, and cover at least three areas first — how companies monitor their internal AI helpers, what those helpers are permitted to do, and whether the AI reliably does what humans intend. Assessments should run continuously, evaluators should publish detailed public reports at least every three months, and there should be a clear way to raise alarms.
The catch is that none of this is required by law. Companies choose the scope, pay the bills, and could walk away. That raises real questions about how independent these watchdogs can be, and the paper calls its ideas a starting point rather than a finished fix. But it matters to you: if AI is handling your job applications, medical paperwork or money, this is about whether anyone outside the company can genuinely verify it's safe.
- Instead of testing AI from the outside, independent experts would work inside AI companies like employees.
- The paper says to check three things first: how companies monitor their internal AI helpers, what those helpers are allowed to do, and whether the AI does what humans intend.
- Evaluators should publish detailed reports at least every three months — but nothing forces companies to take part.
Why It Matters
If it works, you'd get independent proof that the AI handling your data and decisions is actually safe.