DeepSeek trains 1.6T-parameter V4-Pro on 1,000+ Huawei Ascend chips, bypassing Nvidia
Chinese researchers complete full-parameter post-training on domestic hardware—a milestone for AI self-sufficiency.
DeepSeek, in collaboration with Huawei and three Shenzhen-based research institutes, has completed a full-parameter post-training run of its V4-Pro model using more than 1,000 Huawei Ascend 910C chips. The model contains approximately 1.6 trillion parameters and was pre-trained on a dataset exceeding 32 trillion tokens—a process that still relied on Nvidia hardware. This post-training phase refined the model’s instruction-following, safety controls, and specialized behaviors by updating every internal weight, rather than just applying a smaller adapter. The Ascend 910C, Huawei's flagship AI accelerator, previously showed about 60% of an Nvidia H100's inference performance in DeepSeek’s tests, but training performance metrics (duration, utilization, benchmarks) were not disclosed.
Despite the milestone, DeepSeek has not fully replaced Nvidia. Earlier reports indicated that the company struggled to complete a successful training run for its R2 model on Ascend hardware due to unstable performance, slow chip-to-chip communication, and shortcomings in Huawei's CANN software ecosystem. As a result, DeepSeek returned to Nvidia processors for pre-training while continuing to use Ascend chips for inference. The Shenzhen announcement lacked technical details such as training duration, hardware utilization rates, or independent benchmarks. DeepSeek's V4-Pro, released in April, was its first model designed from the start around Ascend hardware, but the company has not publicly commented on this reported achievement. The news underscores China's progress in domestic AI hardware while highlighting persistent gaps in training capability.
- DeepSeek completed full-parameter post-training of its 1.6T-parameter V4-Pro model on over 1,000 Huawei Ascend 910C chips.
- Ascend 910C delivers ~60% of Nvidia H100 inference performance but has faced training instability and ecosystem issues.
- DeepSeek still relies on Nvidia for pre-training; this achievement is limited to post-training, not full model training from scratch.
Why It Matters
Signals China’s growing ability to train AI on domestic chips, though full independence remains elusive amid export restrictions.