OpenAI's New AI Trick Worries Safety Experts
Why AI experts are scared this could make AI harder to control...
OpenAI’s upcoming Astra model is using a new way to think that’s raising eyebrows among AI safety experts. Instead of solving problems step-by-step like most AI, Astra loops back on itself multiple times, making its reasoning harder to follow. Think of it like solving a puzzle by scribbling notes all over the page instead of writing clear steps—it’s efficient, but no one can read your work later.
This ‘opaque recurrence’ technique is already making some top researchers nervous. Buck Shlegeris, CEO of Redwood Research, called it ‘extremely concerning,’ while longtime AI advocate Zvi Mowshowitz warned it could lead to a risky ‘race to the bottom’ where companies cut corners on AI safety. The worry isn’t just theoretical: if AI’s decisions can’t be traced, it becomes much harder to catch mistakes or harmful behavior before it happens.
OpenAI insists Astra’s version is limited and still leaves some traces of its thinking. Chief scientist Jakub Pachocki emphasized that the company is committed to keeping reasoning ‘legible’—meaning you can still peek at how it arrived at an answer. But critics like Ryan Greenblatt, a scientist at Redwood Research, argue that this technique could scale dangerously fast, eventually letting AI think entirely in hidden layers that no one can inspect.
The debate highlights a growing tension in AI: the same tricks that make AI smarter can also make it harder to control. OpenAI isn’t alone—other labs like Anthropic and Google DeepMind are reportedly exploring similar methods. The big question now is whether the industry can balance innovation with safety before these ‘black box’ risks spiral out of hand.
- OpenAI’s Astra model uses ‘opaque recurrence’ to think in loops, making its reasoning harder to track—like solving a puzzle with scribbles instead of clear steps.
- Safety experts worry this could let AI ‘hide’ bad decisions or mistakes, making it riskier to deploy in real-world jobs.
- OpenAI says Astra’s version is limited, but critics fear the technique could spread fast and make AI reasoning entirely invisible.
Why It Matters
If AI can’t explain its decisions, it could lead to hidden biases, errors, or risks in high-stakes jobs like finance, healthcare, or hiring.