DeepSeek founder Liang Wenfeng confirmed internally that the next-generation flagship model V4 will officially launch in late April. Sources cited by Sina Finance reveal that while no exact date has been announced, developer communities are already sensing momentum: the V4-Lite variant is being tested on API nodes, showing a 30% inference speed improvement over the previous generation and a 94% recall rate on 128K token contexts.
Trillion Parameters, Million-Token Window
Unofficial leaked specs show V4 retains a Mixture-of-Experts (MoE) design with roughly 1 trillion total parameters, but only about 37 billion are activated per token, maintaining DeepSeek's hallmark of computational efficiency. On context, V4 leverages a new Engram module that enables up to 1 million tokens of ultra-long context. Engram employs conditional memory queries with O(1) complexity for knowledge access, bypassing linear scaling with sequence length.
Benchmark leaks suggest HumanEval scores of 90% and SWE-bench Verified above 80%. If accurate, these put V4's coding ability on par with leading flagship models. The model natively supports text, image, and video input, with input pricing at roughly $0.30/MTok, continuing DeepSeek's low-cost strategy.
Full Deployment on Huawei Chips: The Biggest Geopolitical Signal
Beyond specs, the most striking shift is hardware: the entire model will run exclusively on Huawei's Ascend 950 PR chips, with zero reliance on NVIDIA GPUs. Alibaba, ByteDance, and Tencent have already placed large orders for Huawei's next-generation chips. If V4 successfully proves that Ascend can handle both training and inference for a top-tier model, it would be the most compelling real-world case yet for chip autonomy in China's AI supply chain.
US export restrictions on NVIDIA components, in this context, could ironically accelerate the maturation of China's native ecosystem.

