AI Safety

New AI Solves Problems in Silence — and Safety Experts Are Worried

This 'thinking without showing its work' trick could make AI much harder to police

Deep Dive

Most AI models 'think out loud.' Ask ChatGPT a hard math problem and it writes out every step — and that visible scratchpad is how people spot mistakes and how companies watch for dangerous behavior. A new model called Astra, built by an AI safety research group, can skip that entirely. It solves problems silently, in a single pass, without writing anything down. A researcher built a 19-task test specifically to measure how well models can do this.

Astra's results were striking. Its odds of solving a random test problem without showing its work were roughly 8.6 times better than the next-best model. On chains of arithmetic, it managed about 7 steps in a row before faltering; the runner-up managed about 4. To make this concrete, the researcher published five sample puzzles that literally only Astra could solve. The finding independently reproduces an earlier result from the UK's AI Safety Institute, which is why people are taking it seriously.

Why does silent thinking matter? Because AI that shows its work is much easier to police. Researchers can read the steps, flag suspicious plans, and catch errors before they cause harm. If models get strong enough to skip that, oversight gets harder — not because the AI is hiding anything, but because there's no text to read. The researcher frames this as an early forecasting signal, not an emergency.

The catch: this is heavily AI-assisted research, and the numbers shift depending on how the test tasks are designed. The author ran enough checks to believe the core claims hold. Astra's gains are also lopsided — it's unusually good at deep step-by-step puzzles, but merely normal at other thinking tasks. For now, this is a sign of where AI is heading, not something that changes your life this week.

Key Points
  • Chain of thought is AI 'showing its work' — the step-by-step text you see when ChatGPT answers you
  • Astra solved hidden-reasoning puzzles about 8.6 times more often than the next-best model
  • AI that doesn't show its work is much harder for humans to check, correct, or catch misbehaving

Why It Matters

If AI stops showing its work, safety teams lose their main window into what it's doing.

📬 Get the top 10 AI stories daily