Alibaba's New AI Cuts Costs to Pennies, Matches Top Models
This could make advanced AI tools 10x cheaper for everyone.
Alibaba unveiled Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture — now open-weight. The production version will be available soon via the QwenCloud API at $0.16 per 1M input tokens and $0.47 per 1M output tokens.
It has 125B parameters plus 51B N-gram embeddings, with just 6B activated per token. It was trained at just 1/9 the cost of Qwen3.7-Plus while outperforming it across the board, with especially strong gains in coding and office tasks — scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI).
It offers 262K native context, extensible to 1M with YaRN. Alibaba is also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture being explored for Qwen4.
- Qwen3.8-Flash costs only $0.16 per million input words, making it one of the cheapest powerful AIs available.
- It can handle text and images, and performs well in coding and office tasks, beating Alibaba's older models.
- The model is open-weight, so anyone can use it, leading to more affordable AI apps for everyone.
Why It Matters
This could make AI assistants and coding helpers drastically cheaper, saving you money on everyday tools.