DeepSeek V4 Pro trained on Huawei Ascend 910C chips, boosting China’s AI self-reliance
China’s first fully domestic LLM training run: 1,000 Huawei chips, 1,500 error-free upgrades.
A research team from Huawei, Shenzhen Loop Area, Harbin Institute of Technology, and Shenzhen Institute of Big Data has successfully executed the post-training of DeepSeek V4 Pro on a cluster of roughly 1,000 Huawei Ascend 910C chips. The process involved full-parameter training—upgrading the entire model structure without pruning—and completed over 1,500 training upgrades without any errors. This demonstrates that Huawei’s Ascend processors can now handle the complex computational and communication demands of LLM post-training, which includes aligning models with human instructions, safety rules, and mathematical reasoning.
Previously, Chinese AI firms relied on Nvidia H800 or AMD chips for training (e.g., DeepSeek V3 used 2,048 H800s). With current U.S. export restrictions, domestic alternatives like Ascend are critical. The successful run shows that Chinese chips can move beyond inference (running pre-trained models) to the more demanding training phase. The team noted that the new capability adds “complex flyovers and loops” to the model’s one-way inference path, instantly multiplying computational demands. This breakthrough enhances China’s AI self-reliance and reduces dependence on foreign hardware.
- DeepSeek V4 Pro post-training completed on ~1,000 Huawei Ascend 910C chips with zero failures.
- Full-parameter training upgraded the entire model, improving math reasoning and instruction-following.
- Marks China's transition from Nvidia-reliant training to domestic chip capability, critical under export restrictions.
Why It Matters
Domestic training on Huawei chips reduces China’s reliance on Nvidia for advanced LLMs, reshaping global AI supply chains.