DeepSeek V4 Specs Leak Claims 1.6T Parameters, 1M Context and Text-Only Design

DeepSeek V4 Specs Leak Claims 1.6T Parameters, 1M Context and Text-Only Design

N
News Editor 01
2026-07-22 15:05:13
A post by Princeton AI researcher Yifan Zhang claims DeepSeek V4 will feature 1.6 trillion parameters, a 1 million-token context window and a 285B Lite version, while reportedly remaining text-only.
DeepSeeklarge-language-modelsartificial-intelligenceAI-modelsmultimodal

DeepSeek’s next flagship model, V4, may have had its technical specs exposed ahead of launch. On July 22, Princeton AI Lab researcher Yifan Zhang posted a detailed configuration sheet on X, and the post was widely interpreted as referring to DeepSeek’s upcoming V4 release. According to the leaked information, the main model carries 1.6T parameters, a lighter 285B-parameter version called V4-Lite is also planned, and the final system could support as much as 1 million tokens of context after reinforcement learning.

A flagship model and a new Lite variant

The source material states that Yifan Zhang is not currently employed by DeepSeek, though he previously worked on ByteDance’s Seed team. He had also posted “V4, next week.” on July 19, which added to speculation once the longer spec list appeared three days later. Based on that leak, the V4 family would include two models: the 1.6T flagship and a 285B V4-Lite.

The architecture details are unusually specific. V4 is said to use a mixture-of-experts setup with 384 experts, activating 6 experts per pass, for roughly 25B active parameters. The leak also mentions a Fused MoE Mega-Kernel for higher compute efficiency. On the attention side, it lists DSA2, a head dimension of 512, and Sparse MQA combined with sliding window attention.

Training details point to a 1M-token context window

The post also describes changes in optimization and training. It claims the model uses the matrix-level optimizer Muon and adopts Hyper-Connections for residual links. Pretraining context length is listed at 32K, while a later reinforcement learning phase using GRPO with corrected KL is said to extend the usable window to 1M tokens.

That combination stands out in the current model race. A million-token context window is already a headline feature on its own; paired with a new Lite edition, it suggests an attempt to cover both top-end performance and lighter deployment scenarios. Even so, none of these details has been confirmed by DeepSeek.

Text-only claim splits reactions

The most debated point is not the scale, but the reported modality choice. The leaked sheet labels V4 as “Text only”, implying no native multimodal support for images, voice or video. That would put it on a different track from models such as GPT-4o and Gemini, which have been pushing multimodal integration aggressively.

Responses on social media quickly split. Some users described the leaked specs as strong enough to compete at the SOTA level. Others questioned why a new flagship would stay with pure text at a time when multimodal capability has become a central benchmark. A number of developers also expressed doubts about authenticity, noting how detailed the sheet is and the lack of any official confirmation or denial from DeepSeek.

Still, the source notes that items such as the Muon optimizer and corrected KL usage fit DeepSeek’s established focus on algorithmic efficiency and cost control. Whether V4 will actually arrive next week, as the earlier teaser suggested, remains unverified.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.