The 120B parameter model gap: Are midsize AI models going extinct?
No new 100-120B MoE models in months, releases jump from 35B to 200B+.
A clear trend has emerged in the AI model landscape: the 100–120B parameter range is being abandoned. The first notable entry was GPT-OSS-120B, released roughly 10 months ago, followed by several other ~120B MoE models — GLM-4.5-Air, Nemotron-3-Super, Qwen3.5-122B, and Mistral-Small-4-119B. However, all of these are now at least three months old, and no new models in this size range have been released since. The pattern mirrors the earlier death of the 70B/80B parameter class, which also saw a flurry of releases before going silent.
Today's frontier releases split sharply into two camps: efficiency-focused small models (25–35B parameters, e.g., Gemma4 and Qwen3.6) and ultra-large systems exceeding 200B parameters (e.g., Step 3.5/3.7 Flash, DeepSeek-V4-Flash, MiniMax-M3, and Nemotron-3-Ultra). This suggests that the mid-range 120B MoE configuration may have been a temporary sweet spot — overtaken by better training data, architecture innovations in smaller models, or the scaling benefits of massive models. If the pattern holds, future releases will likely skip 120B entirely, leaving it as a brief historical footnote in the AI model evolution.
- First 120B MoE model (GPT-OSS-120B) is 10 months old; all subsequent releases in that range are at least 3 months old.
- Recent AI model launches are either sub-35B (Gemma4, Qwen3.6) or 200B+ (Step 3.5/3.7 Flash, DeepSeek-V4-Flash).
- The 120B family appears to be following the same extinction pattern as the earlier 70B/80B parameter class.
Why It Matters
Developers and enterprises should expect a binary choice — small efficient models or massive frontier systems — with the mid-range 120B class soon obsolete.