DeepSeek and Huawei move more AI infrastructure onto Ascend, extending an alternative to Nvidia’s stack

DeepSeek and Huawei move more AI infrastructure onto Ascend, extending an alternative to Nvidia’s stack

N
News Editor
2026-09-30 04:24:54
DeepSeek is working with Huawei to shift a full set of AI training and inference components onto Ascend, according to a BlockBeats brief that cited recent developments around the company’s software stack. The newly open-sourced tools cover operator development, matrix computing, inter-card communication, Attention, and data screening, broadly matching the lower-level toolkit DeepSeek had previously built around Nvidia GPUs. Reuters described the move as the latest sign that Chinese tech companies are looking for alternatives to Nvidia’s ecosystem. BlockBeats said one of the most important pieces is TileLang, which has already been used to implement a large number of operators in DeepSeek V4 training. The tool addresses a basic but critical issue: how developers translate model computation into efficient execution on chips. The report also linked the shift to earlier remarks by Nvidia CEO Jensen Huang, who said in April on Dwarkesh Patel’s podcast that it would be a “bad outcome” for the US if DeepSeek were first released on Huawei chips. BlockBeats added that DeepSeek V4 has already begun adapting to Ascend 950, Huawei has said its chips were used in part of V4 training, and The Information reported last week that Liang Wenfeng has made expanded domestic-chip training a key bet for the company.

DeepSeek is working with Huawei to move a full set of AI training and inference components onto the Ascend platform, according to a BlockBeats news brief. The newly open-sourced tools cover operator development, matrix computing, inter-card communication, Attention, and data screening. BlockBeats said the stack broadly corresponds to the lower-level tools DeepSeek had previously built for Nvidia GPUs.

Reuters described the development as the latest step by Chinese technology companies looking for alternatives to Nvidia’s ecosystem. In the BlockBeats report, the most important piece is TileLang. A large number of operators used in DeepSeek V4 training have already been implemented with it. The tool tackles a low-level question with direct impact on performance: how developers can run model computation efficiently on chips.

TileLang targets a layer long tied to CUDA

According to the report, this capability had long been built heavily around CUDA. DeepSeek and Huawei are now building a corresponding version for Ascend. BlockBeats framed that effort as striking at a concern Nvidia CEO Jensen Huang raised months ago.

In April, Huang said on Dwarkesh Patel’s podcast that it would be a “bad outcome” for the US if DeepSeek were first released on Huawei chips. As cited by BlockBeats, his concern was not only Huawei hardware itself, but the possibility that models would start to be optimized around Huawei’s architecture. If developers around the world end up using the same model and it runs better by default on non-US hardware, the advantage held by US chips would be eroded over time.

From model adaptation to training chips, the software layer is now being filled in

BlockBeats said DeepSeek’s later moves have largely followed that path. V4 has already started adapting to Ascend 950, and Huawei has said its own chips took part in part of V4 training.

The Information reported last week that Liang Wenfeng has made expanded training on domestic chips a key bet for DeepSeek and expects to receive new Huawei training chips as early as the fourth quarter. What is being added now is the software layer. From model adaptation and training chips to matrix computing, MoE communication, Attention, and operator development, DeepSeek is shifting part of the lower-level capabilities that had been deeply dependent on Nvidia and CUDA onto Huawei Ascend, according to the report.

The competition is shifting toward the full technical stack

The BlockBeats brief said AI competition between China and the US had previously focused more on model strength and GPU supply. It is now moving toward a different question: who can build a full stack spanning chips, operators, communication, and training frameworks.

In that framing, the hardest part of CUDA to replace was never just a graphics card. It was the software system that grew around it. DeepSeek and Huawei are now working on that layer.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.