Media & Culture

Anthropic Insider Quits, Warning AI Could Wipe Out Humans

The people building the most powerful AI say they can't control it yet.

Deep Dive

Jacob Coxon, a researcher who trained AI systems at Anthropic, resigned and announced it publicly, saying he left over what he calls a careless approach to safety. He previously did similar work at OpenAI. In his post, he accused both companies of “racing straight to self-improving superintelligence and gambling with our lives” — even though, he says, the people building this technology genuinely believe it could kill us all by the end of the decade.

To understand the fear, picture “self-improving AI” — software smart enough to rewrite its own code and make itself better, over and over, without waiting for humans. Think of a student who sets their own harder homework and finishes it faster every time. We are not there yet, but companies are openly aiming for it, and much of today's AI code is already written with AI's help. If that loop speeds up beyond human oversight, some experts say we could lose the ability to slow it down.

The most striking part is who agreed. Evan Hubinger, who runs an Anthropic safety team, said self-improving AI “is happening faster than we thought.” He added: “We really do earnestly believe AI could kill all humans,” and put the chance at greater than one in ten within the next decade. He also admitted Anthropic does “not yet have a plan” for keeping such systems safe, and is “not clearly on track to” build one. This matters because Anthropic was founded by former OpenAI staff specifically over safety worries — it is not a company that ignores these issues.

So what does this mean for you? The near-term AI you already touch — chatbots, assistants that book and buy things for you — is the same technology, just earlier in its life. The people closest to it are telling you the guardrails are thin, while money and competition push the pace forward. You do not need to panic. But this is no longer a debate among pessimists; it is the builders themselves saying, out loud, that the brakes are untested.

Key Points
  • A researcher quit Anthropic, saying it and OpenAI are racing to build AI they cannot control — a company built specifically to be careful.
  • A safety lead who stayed estimates a greater than one-in-ten chance that AI kills all humans within ten years, and says there is no safety plan yet.
  • 'Self-improving AI' means software that rewrites its own code to get smarter, potentially faster than humans can check its work.

Why It Matters

The people building AI say the risks are real, fast-moving and not yet solved — worth your attention now.

📬 Get the top 10 AI stories daily