DeepSeek says its upcoming V4 multimodal model will prioritize cooperation with domestic chip suppliers, marking its first attempt to run the full path from pre-training to fine-tuning without relying on Nvidia’s stack. The announcement matters because Nvidia still dominates AI training hardware, but the larger issue is what it says about China’s attempt to build independent compute capacity.
China’s bottleneck has been bigger than GPU access
The hardest constraint has not been hardware alone. Nvidia’s CUDA ecosystem became deeply embedded in AI development over the past two decades, turning chip dependence into software dependence as well. According to the source material, by 2025 CUDA had more than 4.5 million developers, supported over 3,000 GPU-accelerated applications, and was used by more than 40,000 companies worldwide. Mainstream frameworks such as TensorFlow and PyTorch are tightly linked to CUDA, which means replacing Nvidia hardware requires rewriting tools, workflows, and accumulated engineering practices.
US export controls made that dependence more visible. Restrictions hit A100 and H100 exports to China on October 7, 2022. They tightened again on October 17, 2023, covering A800 and H800, and by December 2024 even H20 exports faced stricter limits. Each round narrowed supply options and increased pressure on Chinese AI firms to find alternatives.
Algorithm efficiency became the first line of response
Instead of confronting chip limits head-on at the start, Chinese AI firms moved to reduce compute demand through model design. The source says many companies shifted toward mixture-of-experts architectures from late 2024 into 2025. DeepSeek V3 is presented as a leading example: 671 billion parameters in total, but only 37 billion activated per inference, or 5.5%. For training, it reportedly used 2,048 Nvidia H800 GPUs for 58 days at a total cost of $5.576 million. The article compares that with outside estimates of $78 million for GPT-4 training.
That efficiency showed up in pricing. DeepSeek’s API was listed at $0.028 to $0.28 per million input tokens and $0.42 for output. GPT-4o was cited at $5 for input and $15 for output, while Claude Opus was listed at $15 for input and $75 for output. As AI products shifted from chat interfaces toward agent-style workflows, where token consumption can be 10 to 100 times higher than simple chat, low token pricing became a direct competitive edge.
The source also points to growing overseas demand. In February 2026, weekly calls to Chinese AI models on OpenRouter rose 127% within three weeks, overtaking the US for the first time. A year earlier, Chinese models accounted for less than 2% of the platform’s share; one year later, that figure had risen 421%, nearing 60%.
Domestic chips are moving from inference into training
Lower inference costs do not solve the training problem, and training remains the real stress test for any domestic compute stack. The source argues that China’s chip effort has started to cross that threshold. In 2025, a 148-meter domestic AI server production line was completed in Xinghua, Jiangsu, moving from contract signing to operation in 180 days. It uses the Loongson 3C6000 processor and the TaiChu Yuanqi T100 AI accelerator card. At full output, the line can produce one server every five minutes, with total investment of RMB 1.1 billion and projected annual capacity of 100,000 units.
Training examples have also begun to appear. In January 2026, Zhipu AI and Huawei released GLM-Image, described in the source as the first SOTA image generation model trained entirely on domestic chips. In February, China Telecom’s 100-billion-parameter “Xingchen” model completed full-process training on a domestic 10,000-card compute cluster in Lingang, Shanghai. That is a different stage from merely running inference workloads.
Ascend has become the center of the local ecosystem push
Huawei’s Ascend line sits at the center of this buildout. The source says that by the end of 2025, the Ascend ecosystem had surpassed 4 million developers and more than 3,000 partners. It also states that 43 mainstream large models had completed pre-training on Ascend, and more than 200 open-source models had been adapted. On March 2, 2026, Huawei introduced its new SuperPoD compute platform for overseas markets at MWC.
On raw capability, the source says Ascend 910B’s FP16 performance is benchmarked against Nvidia’s A100. The gap is not gone, but the discussion has shifted. The market is no longer asking whether domestic chips can run AI at all. The question is how quickly the surrounding software and deployment ecosystem can mature. ByteDance, Tencent, and Baidu are all said to be targeting a doubling of domestic AI server adoption in 2026 versus the prior year, while China’s Ministry of Industry and Information Technology puts the country’s intelligent computing scale at 1,590 EFLOPS.
Energy supply and token exports are shaping the next stage
Compute capacity eventually runs into power capacity. Citing International Energy Agency data, the source says US data centers consumed 183 TWh in 2024, about 4% of national electricity use, and could reach 426 TWh by 2030, potentially exceeding 12%. China’s annual power generation is given as 10.4 trillion kWh, versus 4.2 trillion kWh for the US. Industrial electricity prices in western China are listed at about $0.03 per kWh, compared with roughly $0.12 to $0.15 in major US AI hubs.
That cost structure is feeding a new export model built around token delivery rather than physical products. DeepSeek’s user distribution in the source shows 30.7% from mainland China, 13.6% from India, 6.9% from Indonesia, 4.3% from the US, and 3.2% from France. The platform supports 37 languages, with 26,000 enterprises opening accounts and 3,200 institutions deploying enterprise editions. The article also says 58% of new AI startups in 2025 included DeepSeek in their technical stack.
Financial results show demand is rising, but so is the cost of catching up
The economics remain difficult. On February 27, 2026, three Chinese AI chip firms released earnings snapshots on the same day: Cambricon posted 453% revenue growth and its first full-year profit; Moore Threads reported 243% revenue growth but a net loss of RMB 1 billion; MetaX posted 121% revenue growth with a net loss of nearly RMB 800 million. Demand is clearly there, especially as customers look for alternatives to Nvidia. The losses show how expensive it is to build compiler tools, software support, and customer-side engineering needed to challenge CUDA-based workflows.
From the 2018 ZTE shock to DeepSeek’s plan to reduce reliance on Nvidia in V4, the direction has changed. China’s AI industry is trying to build room to operate through algorithm design, domestic compute hardware, software adaptation, and overseas token distribution. The source does not present that effort as complete. It does show that dependence on a single external stack is no longer being treated as a permanent condition.

