AI safety research surges 25x: 8.3% of top ML papers in 2026
Only 0.3% of papers in 2019 were safety-focused; now it's 8.3%.
Researchers used DeepSeek V4 Flash to classify every paper accepted at ICLR (2019–2026), ICML (2019–2026), and NeurIPS (2019–2025) — over 55,000 total. 2,328 papers (4.2%) were classified as AI safety. Safety's share rose from 0.3% in 2019 to 8.3% in 2026, a 25-fold increase, with the biggest jump between 2023 (1.5%) and 2024 (4.2%). The absolute number of safety papers grew from 9 in 2019 to 999 in 2026 (over 100x).
Among safety subdomains, interpretability leads with 657 papers, followed by alignment training (457), adversarial robustness (298), and red-teaming (286). Newer areas like dangerous-capability evals (55), scheming & deception (34), and monitoring (53) emerged from near zero. The analysis also tracks affiliations and funders. All data, code, and an interactive paper explorer are openly available on GitHub and a dedicated website.
- 2,328 out of 55,794 papers (4.2%) across three top ML conferences (ICLR, ICML, NeurIPS) are AI safety research
- Safety's share grew from 0.3% in 2019 to 8.3% in 2026 — a 25-fold increase, with 999 safety papers in 2026 alone
- Top subdomains: interpretability (657 papers), alignment training (457), adversarial robustness (298); 17 subdomains total
Why It Matters
Safety research is no longer niche — it's a major, fast-growing fraction of top ML conferences, shaping how future AI systems are built.