ByteDance’s Seed team said it found a repeatable weakness in DeepSeek-V4 during long-context tasks. In the test setup, researchers asked the model to complete a piece of code. The code itself stayed the same, but when they only extended an unrelated comment placed before it, the model’s answer alternated between correct and incorrect outputs.
According to the findings, the base V4-Flash model repeated this pattern every four tokens, while V4.1-Flash showed a similar effect on a two-token cycle. ByteDance linked the behavior to DeepSeek’s chunked KV cache compression, a method used to save memory by merging information from several consecutive tokens. The team said that compression can preserve information unevenly across positions, making some parts harder to retrieve accurately later.
ByteDance described the periodic retrieval gap as “phase sensitivity.” In a retrieval test spanning about 128,000 tokens, the base V4-Flash model showed a maximum accuracy gap of 40.2 percentage points across positions. After post-training, that gap narrowed to 19.1 percentage points. V4.1-Flash reduced it further to 6.1 percentage points, though the periodic fluctuation remained. The team also trained several control models from scratch and said models without chunked compression did not show the same pattern.
ByteDance’s Seed team said it found a regular weakness in DeepSeek-V4 during long-context processing. Researchers asked the model to complete a piece of code. The code did not change, but when they lengthened an unrelated comment placed before it, the model’s output flipped back and forth between correct and incorrect answers.
A fixed cycle appeared in the errors
In the reported results, the base V4-Flash model repeated the pattern every 4 tokens. V4.1-Flash showed a similar issue, with the cycle shifting to 2 tokens.
ByteDance called this periodic retrieval difference “phase sensitivity.”
The issue was tied to chunked KV cache compression
The team said the behavior was related to DeepSeek’s “chunked KV cache compression,” which is used to cut memory use. The method merges information from several consecutive tokens so long-context processing costs less.
That compression, however, can preserve information from different positions to different degrees. Some positions become harder to retrieve accurately later, which leads to a regular fluctuation when the input position changes.
Test results and control models
In a retrieval test covering about 128,000 tokens, the base V4-Flash model showed a maximum accuracy gap of 40.2 percentage points across different positions. After post-training, the gap narrowed to 19.1 percentage points. V4.1-Flash reduced it further to 6.1 percentage points, but the periodic fluctuation still remained.
ByteDance also trained several control models from scratch. It found that the fluctuation period matched the compression stride, while control models that did not use chunked compression did not show the same pattern.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.