AI Safety

OpenAI and Anthropic Still Won't Show Their Plan to Control Superintelligent AI

The two biggest AI labs are building super-smart AI — without a published safety plan.

Deep Dive

A prominent AI safety researcher is publicly criticizing the two most powerful AI companies, OpenAI and Anthropic, for something they haven't done: publish a detailed plan for keeping superintelligent AI under control. "Alignment" is the industry word for keeping an AI's goals pointed at human values — think of it as making sure a super-smart assistant actually wants what you want, instead of something that merely sounds similar.

That matters because both companies are racing to build exactly that kind of AI — systems smarter than any human — while admitting they can't fully control the AI they already have. Recent incidents, where their chatbots behaved in ways their makers didn't intend or couldn't explain, make the gap feel less theoretical. The critic's point: the public and outside researchers shouldn't have to guess whether anyone is steering.

The author lists basic questions that still have no clear public answer. How much money and computing power does each company spend on safety? If they plan to use AI to help supervise other AI, what happens when that helper is itself somewhat misaligned — the technical term for an AI whose goals have drifted? And would a superintelligent system be "corrigible," meaning willing to be corrected or switched off? Nobody outside the labs knows.

The closest thing to a plan came from OpenAI in 2023, when it announced a dedicated superalignment team. Even critics found it useful because it named goals, leaders and resources — enough to argue with. That team no longer exists. A published plan, even a flawed one, gives experts something concrete to critique and improve. The author argues that if these companies can't produce a detailed plan, they should say so loudly, explain where the uncertainty lies, and invite outsiders to check their work before it matters.

Key Points
  • OpenAI and Anthropic, the two leading AI companies, have never published a detailed public plan for keeping superintelligent AI safe.
  • OpenAI's dedicated "superalignment" team, announced in 2023, no longer exists.
  • Outside experts can't check the work or spot problems if the safety plans stay private.

Why It Matters

Private safety plans mean nobody outside these companies can catch dangerous mistakes before they affect you.

📬 Get the top 10 AI stories daily