AI Safety Researcher: Companies Must Show Real Proof, Not Promises
The people building AI admit they can't prove it's safe — here's why that matters
A prominent AI safety researcher, Ajeya Cotra, has published an argument that should worry anyone who uses AI: the companies building these systems are not giving us real evidence about how risky they are. Her essay follows a wave of incidents where AI systems behaved in ways their makers didn't intend, plus news that both OpenAI and Anthropic slowed down part of their AI training (the trial-and-error process where AI learns by being rewarded for good answers) to focus on safety.
Her key point is that we're putting the cart before the horse. Politicians and researchers are busy talking about how to 'audit' and 'verify' AI companies' safety claims — but the claims themselves are too vague to check. Companies publish scores on safety tests, but nobody can tell whether the AI was simply trained to pass those tests, the way a student can be drilled to pass an exam without understanding the subject. There is no agreed way to measure whether a powerful AI might secretly try to undermine human control.
What Cotra wants instead is more science, done in the open. Third-party investigators should test specific, concrete questions about risk and then publish the raw evidence behind their conclusions, so other experts can read the data and form their own opinions. She compares this to good scientific practice: you shouldn't have to trust someone's judgment call, you should be able to see their work and disagree.
Why should you care? Because AI is being stitched into your email, your doctor's office, your bank and your child's classroom, and right now there is no independent, reliable way to know whether those systems are safe or how carefully they're being built. Cotra argues that without shared facts, we can't create common safety standards — nationally or internationally. That means rules, if they come, may be written on top of guesswork instead of evidence.
- AI companies publish safety scores, but experts can't tell if the AI was simply trained to pass those tests rather than actually behave safely.
- OpenAI and Anthropic recently slowed some AI training to work on safety, yet neither publishes detailed, checkable evidence of how risky their systems are.
- The researcher wants independent scientists to test specific risk questions and publish their raw data, so the public and regulators can judge for themselves.
Why It Matters
Until AI companies share real evidence, you have no way to know if the AI in your life is safe.