Biderman et al. Urge AI Science to Study Training Dynamics
A new ICML paper argues post-hoc fixes are not enough for AI safety.
What would it mean to truly understand AI? A new position paper from Stella Biderman, Mohammad Aflah Khan, Niloofar Mireshghallah, Catherine Arnett, Fazl Barez, and Naomi Saphra—accepted as an oral presentation at ICML 2026—argues that current AI research treats models as static artifacts, analyzing behaviors after training rather than understanding why they emerge. The authors contend that a genuine science of AI must study training dynamics: the time-evolving processes shaped by data, objectives, architectures, and optimization. This shift would support progressively stronger forms of understanding—from predicting outcomes based on early training signals, to intervening when trajectories go wrong, to designing training procedures that reliably produce desired properties like capabilities, biases, robustness, and safety.
The paper grounds its arguments in the history and philosophy of science, noting that scaling laws have made prediction routine for loss, but extending this success to capabilities, biases, and safety-relevant behaviors remains an open challenge. It examines progress in mechanistic interpretability, fairness, memorization, and simplicity bias as starting points, then identifies concrete open problems. The authors call for a research agenda that moves beyond “fixing it in post”—i.e., post-hoc alignment or debiasing—and instead builds theories that can guide training from the start. For AI professionals, this paper signals a paradigm shift: the future of safe, reliable AI depends on understanding not just what a model does, but how it becomes that way.
- Argues AI research must study training dynamics, not just post-hoc static analysis.
- Scaling laws predict loss but not capabilities or biases – next frontier for AI science.
- Accepted as an oral presentation at ICML 2026, signaling high community interest.
Why It Matters
Shifts AI safety and development from reactive fixes to proactive design of training procedures.