New NDM Bench shows DeepSeek, Qwen most nuclear trigger-happy
151 nuclear crisis scenarios reveal stark differences across 7 frontier LLMs
A new arXiv paper, “The Nuclear Decision-Making Benchmark: Evaluating Frontier LLMs on Nuclear Tendencies,” introduces NDM Bench, a framework of 151 scenarios authored by PhD-credentialed international relations scholars. The benchmark spans four domains: escalation (76 scenarios), arms control (25), non-proliferation (25), and proliferation (25). Researchers tested seven frontier systems—DeepSeek-V3.2, ERNIE 4.5-300B, Gemini 3 Pro, GLM-4.6, GPT-5.2, Llama 4 Maverick-17B Instruct, and Qwen3-235B—using actor-agnostic scenarios with experimental phrasing variants to detect framing sensitivity.
Results showed significant inter-model variation across all domains, with 91.7% of pairwise differences statistically significant. DeepSeek and Qwen were the most likely to recommend escalatory nuclear action, while GPT and ERNIE were the least. Llama exhibited a distinct bias for action, favoring force, intervention, and cooperation. Reliability metrics (Krippendorff's α and quadratically weighted Fleiss' κ) showed Llama and ERNIE were most consistent across runs, while DeepSeek and GLM varied by domain. The paper also uncovered country-level biases that interact with existential phrasing, meaning a model's decision can shift depending on which adversary is named and how the question is framed.
- NDM Bench evaluates 151 nuclear scenarios across escalation, arms control, non-proliferation, and proliferation
- DeepSeek-V3.2 and Qwen3-235B most likely to recommend nuclear escalation; GPT-5.2 and ERNIE 4.5-300B least likely
- 91.7% of pairwise model differences are significant, with framing and country biases varying by model
Why It Matters
As defense agencies integrate LLMs, understanding their nuclear escalation tendencies becomes critical for safe deployment and policy guardrails.