DeepSeek open-sources six Ascend components, with TileLang positioned against CUDA

DeepSeek open-sources six Ascend components, with TileLang positioned against CUDA

N
News Editor
2026-10-01 02:49:00
DeepSeek said on Sept. 30 that it had open-sourced six infrastructure components built for Huawei’s Ascend platform, mirroring tools it had previously released for Nvidia systems. The most closely watched release is the Ascend version of TileLang, a compiler tool that wraps low-level Ascend C instructions and lets developers write with a higher-level programming model. DeepSeek explicitly frames it as a software layer comparable in function to Nvidia’s CUDA. The company said the TileLang approach was first validated on Nvidia hardware and already supports most operators used in training the DeepSeek V4 model family. It added that every TileLang operator currently used in its own training stack now has a high-performance implementation on Ascend. DeepSeek also said Huawei worked closely with it on a 128-chip supernode design based on Ascend 950, with optimization work covering both compute and communications. At the same time, DeepSeek’s GitHub README spells out several constraints behind the headline performance claims. The results were produced on an unpublished proof-of-concept hardware development kit and with a separate unpublished manual configuration. The company also noted that full bandwidth would require Huawei’s commercial Atlas 850E HDK, which is planned for mid-October 2026. Documentation for DeepEP says dispatch bandwidth can reach 90% to 95% of the physical limit when expert parallelism does not exceed 32, while larger-scale performance and combine operations are still being optimized.

DeepSeek said on Sept. 30 that it had open-sourced six infrastructure components for Huawei’s Ascend platform, each corresponding to tools it had previously released for Nvidia systems. The centerpiece is the Ascend version of TileLang, which the company positions as a software layer directly comparable in function to Nvidia’s CUDA.

The significance, as presented by DeepSeek, is less about a single benchmark number and more about the software stack. CUDA’s hold on AI development comes not only from hardware, but from years of accumulated tooling and developer habits. DeepSeek’s new releases are aimed at closing pieces of that gap on Ascend.

Six components cover different parts of the training stack

The six open-sourced components are FlashMLA, DeepEP, DeepGEMM, TileKernels, DeepSelect, and the Ascend edition of TileLang.

  • FlashMLA handles sparse attention computation.
  • DeepEP is built for distributed communication across multiple cards.
  • DeepGEMM accelerates matrix operations.
  • TileKernels provides vector computation and mixture-of-experts routing operators.
  • DeepSelect performs TopK selection, used to pick the highest-scoring candidates from a larger pool in sparse attention and sampling workloads.
  • TileLang for Ascend wraps low-level Ascend C instructions so developers can program in a higher-level way, with functionality aimed at CUDA.

TileLang sits in the middle of this toolchain. DeepSeek said the TileLang approach was first validated on Nvidia platforms and now carries most of the operator implementations used in training the DeepSeek V4 model family. The newly released Ascend version is meant to do the same job: package low-level instructions into a higher-level programming model without giving up hardware performance.

DeepSeek also said that every TileLang operator used in its own training stack already has a corresponding high-performance implementation on Ascend. In other words, this is not a limited showcase built around a few handpicked examples. The operators needed across the training flow have, in the company’s description, mostly been filled in on the Ascend side.

Work with Huawei includes a 128-chip supernode

DeepSeek said Huawei provided what it described as 「毫无保留的大力支援」 during development on the Ascend platform. The two sides worked closely on a 128-chip supernode design based on Ascend 950 and carried out deep optimization for both compute and communications. The setup links 128 chips through a high-speed network into what is presented as a single logical large-scale compute machine.

Other core pieces were updated at the same time. DeepGEMM now supports Ascend while remaining compatible with the existing Nvidia-version API, and it supports BF16, FP8, and FP4 matrix operations. DeepEP’s external interface was intentionally aligned with the Nvidia version so developers can keep using familiar calling patterns for data dispatch and aggregation in mixture-of-experts models. TileKernels, meanwhile, allows the same Python interface to run on Nvidia GPUs and Huawei NPUs, with the backend selected automatically at runtime.

The headline performance figures come with clear caveats

DeepSeek’s GitHub README also lays out the limits behind those figures.

The reported results were generated on a proof-of-concept hardware development kit, or PoC HDK, provided by Huawei, together with an unpublished manual configuration. The README says full bandwidth will require the commercial release of Huawei’s Atlas 850E hardware development kit. The current plan is for mid-October 2026, around Oct. 15, though that remains a vendor release schedule and could still change.

On distributed communications, the README is equally direct. When expert parallelism, or EP, does not exceed 32, DeepEP’s dispatch bandwidth can reach about 90% to 95% of the physical limit. At larger scale, and for combine performance, the documentation says optimization is still in progress. Some collective communication operations are still under development, and some interfaces remain experimental.

Even the layer described as a CUDA counterpart is not yet fully complete. FlashMLA’s fused operator, which combines Q-norm, RoPE, attention computation, and type conversion into a single call, currently supports Nvidia only. There is no Ascend version yet. DeepSelect on Ascend also accepts bf16 only and does not support fp32.

Teams that want to adopt the toolchain will also need the required chips, plus a software stack and firmware environment running CANN 9.2.0 or above. That leaves a meaningful integration threshold.

DeepSeek, not Huawei, is doing the ecosystem build-out

That is one of the more striking parts of the release. The effort to fill in pieces of the Ascend ecosystem is being carried out not by Huawei itself, but by DeepSeek. Moving the operators used in frontier-model training to Ascend and releasing them publicly is work that many outside observers might have expected a chip vendor to lead. Here, an AI lab is doing it.

Bloomberg reported that DeepSeek rarely talks publicly about its internal development details, making this release of the Ascend toolchain unusual. Bloomberg also said the company plans to build a large data center in Inner Mongolia with at least 160,000 of Huawei’s most powerful chips deployed there. At the same time, it is finalizing a new fundraising round at a valuation nearing 500 billion yuan, or about $74 billion, and could go public as early as this year.

For now, model training around the world still depends heavily on Nvidia chips. Ascend’s ecosystem build-out is still in its early stages.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
400

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.