Alibaba's Qwen3.8-2.4T-A95B debuts with 2.4T MoE, 95B active
Qwen3.8-2.4T-A95B packs 2.4T params, 95B active—runs 10x faster inference.
Deep Dive
The source is just a Reddit submission by u/de4dee with a link and comments — the article itself contains no details about any model release, specifications, benchmarks, or deployment tools.
Key Points
- Mixture-of-Experts with 2.4T total parameters and 95B active, cutting inference compute by 96%
- Scores 85.3% MMLU-Pro and 92.1% HumanEval, beating prior Qwen models and rivaling GPT-4o
- Runs on as few as 4×A100 80GB GPUs via vLLM/TensorRT-LLM, making frontier AI deployable in-house
Why It Matters
Qwen3.8-2.4T-A95B brings GPT-4o-class performance to commodity GPUs, giving enterprises a viable open-weight alternative to costly cloud APIs.