AI Safety

Goodfire's Silico reproduces RLFR, challenging the 'Most Forbidden Technique' taboo

A new platform uses internal model probes as reward signals, reigniting a key AI safety debate.

Deep Dive

Goodfire has launched a private beta of Silico, an LLM training platform that reproduces RLFR (reinforcement learning from probes)—a technique that uses internal model probes as reward signals for reinforcement learning. This has rapidly drawn comparisons to the 'Most Forbidden Technique' concept from the classic LessWrong post, which warned against training directly on model internals due to risks of obfuscation and gaming of safety metrics. Goodfire's accompanying post, authored by Rauno Arike, argues that such blanket objections are overblown and that training on internals can be acceptable under specific conditions.

Arike's literature review cites The Obfuscation Atlas by Taufeeque et al., which provides nuanced guidance on when using model internals in training is warranted. The key red line, per alignment researcher Daniel Kokotajlo, is maintaining a held-out test set that is not trained on or correlated with training proxies. Arike emphasizes that probe-based signals (like RLFR) should not be categorically forbidden, but rather used with caution by ensuring validation measures remain uncorrelated with training objectives. Silico's release reignites the debate over alignment vs. safety, offering a practical tool while pushing the community to move beyond cached memes.

Key Points
  • Goodfire launches private beta of Silico, an LLM training platform reproducing RLFR (reinforcement learning from probes).
  • RLFR uses internal model probes as reward signals, which many label the 'Most Forbidden Technique' from LessWrong.
  • Post argues training on internals is acceptable under four conditions, referencing The Obfuscation Atlas by Taufeeque et al.

Why It Matters

This debate redefines how AI labs can responsibly use internals for alignment training without sacrificing safety.

📬 Get the top 10 AI stories daily