DeepSeek has announced a major pricing update for its model lineup, cutting the cost of input cache hits across the entire series to one-tenth of the original launch price. The move could lower operating costs for users who rely on repeated prompts, high-volume inference, and optimized API workflows.
Cache pricing reduced across the lineup
The pricing change applies to DeepSeek’s full model family and focuses on input cache hits, an important cost component for developers and businesses running recurring requests. Lower cache pricing can meaningfully improve cost efficiency for applications that reuse context or process similar prompts at scale.
V4-Pro gets an extra temporary discount
In addition to the broader pricing cut, DeepSeek said DeepSeek-V4-Pro is available with a 25% discount until May 5, 2026, 23:59 Beijing Time. Under the promotion, the model’s input price falls to 0.025 yuan per million tokens, offering further savings for customers considering deployment.
From a market perspective, pricing adjustments like this often signal stronger competition for developer adoption and usage volume. As model costs become a more important factor in procurement and product planning, lower input and cache fees can make a platform more attractive to cost-conscious teams.
So far, the disclosed information centers on the pricing update and temporary promotion, with no additional technical changes detailed in the source material. For users evaluating model integrations, the reduced pricing is likely to be the most immediate takeaway.

