DeepSeek is working with Huawei to move a full set of AI training and inference components onto the Ascend platform, according to a BlockBeats news brief. The newly open-sourced tools cover operator development, matrix computing, inter-card communication, Attention, and data screening. BlockBeats said the stack broadly corresponds to the lower-level tools DeepSeek had previously built for Nvidia GPUs.
Reuters described the development as the latest step by Chinese technology companies looking for alternatives to Nvidia’s ecosystem. In the BlockBeats report, the most important piece is TileLang. A large number of operators used in DeepSeek V4 training have already been implemented with it. The tool tackles a low-level question with direct impact on performance: how developers can run model computation efficiently on chips.
TileLang targets a layer long tied to CUDA
According to the report, this capability had long been built heavily around CUDA. DeepSeek and Huawei are now building a corresponding version for Ascend. BlockBeats framed that effort as striking at a concern Nvidia CEO Jensen Huang raised months ago.
In April, Huang said on Dwarkesh Patel’s podcast that it would be a “bad outcome” for the US if DeepSeek were first released on Huawei chips. As cited by BlockBeats, his concern was not only Huawei hardware itself, but the possibility that models would start to be optimized around Huawei’s architecture. If developers around the world end up using the same model and it runs better by default on non-US hardware, the advantage held by US chips would be eroded over time.
From model adaptation to training chips, the software layer is now being filled in
BlockBeats said DeepSeek’s later moves have largely followed that path. V4 has already started adapting to Ascend 950, and Huawei has said its own chips took part in part of V4 training.
The Information reported last week that Liang Wenfeng has made expanded training on domestic chips a key bet for DeepSeek and expects to receive new Huawei training chips as early as the fourth quarter. What is being added now is the software layer. From model adaptation and training chips to matrix computing, MoE communication, Attention, and operator development, DeepSeek is shifting part of the lower-level capabilities that had been deeply dependent on Nvidia and CUDA onto Huawei Ascend, according to the report.
The competition is shifting toward the full technical stack
The BlockBeats brief said AI competition between China and the US had previously focused more on model strength and GPU supply. It is now moving toward a different question: who can build a full stack spanning chips, operators, communication, and training frameworks.
In that framing, the hardest part of CUDA to replace was never just a graphics card. It was the software system that grew around it. DeepSeek and Huawei are now working on that layer.

