DeepSeek V4 runs on Huawei Ascend day one — 1.6T parameters, 1M context
Huawei's Ascend 950DT delivers full-speed performance for DeepSeek's massive MoE model from launch.
DeepSeek V4, a massive 1.6 trillion parameter Mixture-of-Experts model with support for 1 million token context, launched with seamless day-zero compatibility on Huawei's Ascend 950DT AI chips. The collaboration between DeepSeek and Huawei ensured that CANN software was pre-optimized with fused operators and intelligent dual-core scheduling, eliminating the typical early-stage performance bottlenecks. The model uses FP4 precision, which aligns perfectly with Ascend's architectural strengths, enabling smooth inference without manual tuning.
This rare instant support for a large open-source model means enterprises and cloud providers could deploy DeepSeek V4 on Ascend SuperNode clusters immediately. Huawei's hardware-software co-design, highlighted in a SemiAnalysis report, gives Ascend a practical edge over other accelerators that struggled with speed or missing optimizations. The launch signals that Huawei's domestic chip ecosystem is ready for frontier AI workloads, offering faster, more independent deployment options for companies in China and beyond.
- DeepSeek V4 has 1.6 trillion parameters and 1M token context, using Mixture-of-Experts architecture with FP4 precision.
- Huawei Ascend 950DT delivered day-zero support via CANN software with fused operators and dual-core optimizations.
- SemiAnalysis reports Huawei's hardware-software co-design outperforms other accelerators for modern MoE models.
Why It Matters
Huawei's Ascend chips now support cutting-edge open-source AI from day one, reducing dependency on Western hardware.