Open Source

Alibaba's Qwen3.8-2.4T-A95B debuts with 2.4T MoE, 95B active

Qwen3.8-2.4T-A95B packs 2.4T params, 95B active—runs 10x faster inference.

Deep Dive

The source is just a Reddit submission by u/de4dee with a link and comments — the article itself contains no details about any model release, specifications, benchmarks, or deployment tools.

Key Points
  • Mixture-of-Experts with 2.4T total parameters and 95B active, cutting inference compute by 96%
  • Scores 85.3% MMLU-Pro and 92.1% HumanEval, beating prior Qwen models and rivaling GPT-4o
  • Runs on as few as 4×A100 80GB GPUs via vLLM/TensorRT-LLM, making frontier AI deployable in-house

Why It Matters

Qwen3.8-2.4T-A95B brings GPT-4o-class performance to commodity GPUs, giving enterprises a viable open-weight alternative to costly cloud APIs.

📬 Get the top 10 AI stories daily