DeepSeek said it open-sourced a set of infrastructure components for Huawei’s Ascend platform on GitHub on Sept. 30, covering both compute libraries and distributed communication libraries. The company said the releases correspond one by one to components it had previously open-sourced for Nvidia GPU platforms.
DeepGEMM Ascend released with support for the Ascend 950 series
Documentation for DeepGEMM Ascend says the first version was released on Sept. 30 and supports the Ascend 950 series. The matrix multiplication library is described as fully compatible with the original DeepGEMM API, with support for BF16, FP8 and FP4 formats. After installation, developers can keep using the same APIs and development workflow used on other platforms.
The documentation also says the implementation adds a lightweight abstraction layer over Ascend matrix multiply-accumulate primitives, packaging details such as fractal matrix layout, alignment constraints and address calculation.
DeepEP-Ascend exposes a buffer API aligned with the Nvidia version
Another component, DeepEP-Ascend, is a communication library for training and inference. It provides the expert-parallel communication operations required by mixture-of-experts, or MoE, models. Its documentation says the external buffer API is aligned with the Nvidia version of DeepEP.
According to QbitAI, the open-source release also includes the TileLang compiler tool, FlashMLA, TileKernel and DeepSelect.
Benchmark data shows near-hardware-limit bandwidth at smaller scale
DeepEP-Ascend documentation includes benchmark data from tests run on Ascend 950DT with CANN 9.2.0. Under a configuration of eight expert-parallel units, dispatch bandwidth came in at 373-375 GB/s, while combine bandwidth was 345-347 GB/s. When the scale increased to 128 units, dispatch fell to 313-320 GB/s and combine dropped to 272-278 GB/s.
The documentation says sustained dispatch bandwidth reaches about 90% to 95% of the physical payload bandwidth ceiling at up to 32 units. Larger-scale runs and combine are still being tuned, and combine carries additional overhead from local reduction.
The same documentation notes that the figures were obtained in a manually configured proof-of-concept firmware environment. It says support for other Ascend generations or other CANN versions was not verified by this batch of tests.
QbitAI says Huawei provided SuperPoD Flex and UBL128 networking plans
QbitAI reported that Huawei said it provided Ascend SuperPoD Flex and UBL128 networking schemes defined jointly with DeepSeek. The report said the setup can reach a network scale of 128 cards with 3.2 Tb/s single-layer switching.
QbitAI also said Huawei open-sourced the deployment work through the CANN community, including methods for low-latency inference deployment, single-card and single-node deployment, and large-scale training.
Offline inference throughput and firmware timing
Data cited by QbitAI showed that in offline inference mode with 32 expert-parallel units, the DeepSeek-V4.1-Flash base model delivered output throughput of 2,469 tokens per second per card at 5 milliseconds of latency per output token. The figure was marked as collected in offline inference mode and does not include the effects of service scheduling or framework load balancing.
Documentation for DeepEP-Ascend also says full bandwidth requires Huawei’s third-quarter commercial firmware release for Atlas 850E. The document says that version is expected to be made public in mid-October 2026, around Oct. 15, though the actual schedule remains subject to Huawei’s release plan.
Related context
Compute options for laboratories in China have remained in flux recently. ChainCatcher previously reported that The Information said China was preparing to allow ByteDance and Alibaba to buy Nvidia’s new chips.

