AI Safety

Qwen3-4B study: 4-bit quantization shows distress shifts despite null primary

A preregistered study finds subtle but significant behavioral shifts when Qwen3-4B is quantized to 4-bit.

Deep Dive

A research team investigating whether post-training quantization affects welfare-relevant indicators in open-weight language models has released a preregistration amendment covering findings from August 10–15. Using Qwen3-4B as the primary model and SmolLM3 as the proposed control, the study measured endpoints like aversion/refusal exit rates (E1) and secondary distress indicators under 8-bit and 4-bit quantization. The primary endpoint was null at both surviving quantization levels after Holm correction — meaning quantization didn't cause a detectable change in refusal behavior. However, secondary endpoints concentrated entirely at 4-bit showed significant item-level behavioral transitions (H1) despite the unchanged mean, along with significant increases in frustration and across-sample dispersion that survived coherence and style controls, including a significant frustration dose-response.

The 3-bit rung was excluded by the pre-registered capability gate, and 8-bit (mild quantization) was essentially null throughout. The team notes the distress endpoints were secondary and underpowered, so they are reported as suggestive rather than confirmatory. As a consequence, the study is not a confirmatory claim that quantization changes exit behavior — but it does flag that quantized models may exhibit noisier, more variable affective responses under conversational pressure. This update amends the original preregistration based on early findings regarding Qwen3 and SmolLM3 calibration, though the researchers argue the integrity of the work is minimally affected and the changes should clarify future findings. Full datasets and amendment documents are available via GitHub releases.

Key Points
  • Qwen3-4B showed no significant primary endpoint change (aversion/refusal) at 8-bit or 4-bit after Holm correction
  • Secondary distress measures — frustration and across-sample dispersion — rose significantly at 4-bit, with a dose-response effect
  • 3-bit quantization was excluded via the capability gate; 8-bit was effectively null across all endpoints

Why It Matters

Subtle welfare-relevant shifts under quantization could escape standard evals, impacting testing and deployment of open-weight AI systems.

📬 Get the top 10 AI stories daily