DeepSeek’s next flagship model, V4, may have had its technical specs exposed ahead of launch. On July 22, Princeton AI Lab researcher Yifan Zhang posted a detailed configuration sheet on X, and the post was widely interpreted as referring to DeepSeek’s upcoming V4 release. According to the leaked information, the main model carries 1.6T parameters, a lighter 285B-parameter version called V4-Lite is also planned, and the final system could support as much as 1 million tokens of context after reinforcement learning.
A flagship model and a new Lite variant
The source material states that Yifan Zhang is not currently employed by DeepSeek, though he previously worked on ByteDance’s Seed team. He had also posted “V4, next week.” on July 19, which added to speculation once the longer spec list appeared three days later. Based on that leak, the V4 family would include two models: the 1.6T flagship and a 285B V4-Lite.
The architecture details are unusually specific. V4 is said to use a mixture-of-experts setup with 384 experts, activating 6 experts per pass, for roughly 25B active parameters. The leak also mentions a Fused MoE Mega-Kernel for higher compute efficiency. On the attention side, it lists DSA2, a head dimension of 512, and Sparse MQA combined with sliding window attention.
Training details point to a 1M-token context window
The post also describes changes in optimization and training. It claims the model uses the matrix-level optimizer Muon and adopts Hyper-Connections for residual links. Pretraining context length is listed at 32K, while a later reinforcement learning phase using GRPO with corrected KL is said to extend the usable window to 1M tokens.
That combination stands out in the current model race. A million-token context window is already a headline feature on its own; paired with a new Lite edition, it suggests an attempt to cover both top-end performance and lighter deployment scenarios. Even so, none of these details has been confirmed by DeepSeek.
Text-only claim splits reactions
The most debated point is not the scale, but the reported modality choice. The leaked sheet labels V4 as “Text only”, implying no native multimodal support for images, voice or video. That would put it on a different track from models such as GPT-4o and Gemini, which have been pushing multimodal integration aggressively.
Responses on social media quickly split. Some users described the leaked specs as strong enough to compete at the SOTA level. Others questioned why a new flagship would stay with pure text at a time when multimodal capability has become a central benchmark. A number of developers also expressed doubts about authenticity, noting how detailed the sheet is and the lack of any official confirmation or denial from DeepSeek.
Still, the source notes that items such as the Muon optimizer and corrected KL usage fit DeepSeek’s established focus on algorithmic efficiency and cost control. Whether V4 will actually arrive next week, as the earlier teaser suggested, remains unverified.

