AI Labs Grade Their Own Safety Tests — Researchers Want a Second Opinion
The companies promising to keep AI safe are the only ones checking their work.
When AI companies like OpenAI and Anthropic publish research about making their models safer, they usually don't hand over their code or even explain exactly how they ran the tests. That means the claim "our AI is safer now" can't be checked by anyone outside the company. Two researchers, Zephaniah Roe and yix, argue in a new post that this needs to change — and that independent teams should redo the experiments themselves.
The stakes sound dramatic because the labs themselves have made them dramatic. Leaders at these companies have said, more or less openly, that the technology they're building could cause serious harm, even human extinction. Yet their evidence for preventing that is like a restaurant grading its own kitchen. When outside teams have tried to redo safety experiments, they found results that hinge on tiny, invisible choices — like which server the AI was routed through, or one single number used when training the model. Change something small and the "safe" result can disappear.
So the authors propose a dedicated effort: redo the lab experiments, stress-test the methods by trying them in different settings, and publish everything openly so others can check the work. They're honest that this isn't glamorous. Copying someone else's experiment won't win awards or land a researcher a job, which is exactly why almost nobody does it. They want special attention paid to "model cards" — the summaries labs publish about what a model can and can't do — including big claims like "our most aligned model ever." Alignment, in plain terms, means the AI does what humans actually want it to do.
Why should you care? If you use a chatbot to draft emails, ask health questions, or handle your money, you're relying on safety promises you have no way to verify. Replication is the simplest quality check there is. If a finding falls apart when a careful outsider tries it, that's strong evidence it won't hold up in the real world either — where the consequences land on you, not the lab.
- AI labs mark their own homework: safety results are published without code or clear method details, so no outsider can check them.
- Replication is a basic test — if an independent team can't get the same result, the finding probably won't hold up in the real world.
- The researchers want funded teams to redo lab experiments, poke holes in the methods, and post everything publicly for others to verify.
Why It Matters
If AI safety claims go unchecked, the guardrails on tools you use daily rest on the company's word alone.