DeepSeek released DeepSeek-V4.1-Flash on Sept. 10, introducing what it describes as the smallest model in its new architecture family. The company published the weights on Hugging Face under an MIT license and launched the model across its web and mobile apps.
According to SiliconANGLE, DeepSeek said a range of tests showed the open-weight model surpassed its much larger flagship V4-Pro in performance, cost, speed, and total execution time. Chain News had previously reported that the production version of V4 Pro launched on Aug. 13, meaning a smaller in-house model moved ahead of it in less than a month.
552B backbone parameters with 8B to 16B activated per token
The model card says V4.1-Flash uses a mixture-of-experts architecture with 552B backbone parameters, nearly double the 284B listed for V4-Flash. Even so, the model activates only 8B parameters per token during the prefilling stage and 16B during decoding.
SiliconANGLE said the model needs just 890 bytes of KV cache per token, compared with about 3,560 bytes for V4-Flash, a factor tied to lower inference cost. DeepSeek lists support for a 1 million-token context window, native image input, and text output.
API naming changes and retirement of the V4 Flash line
DeepSeek’s API update log says V4.1-Flash is now offered under the name deepseek-flash. The older deepseek-v4-flash endpoint and the experimental vision model name will temporarily route to V4.1-Flash, while the previous V4 Flash series has been retired.
The same log says DeepSeek V4 Pro API service will continue after Sept. 14, with pricing unchanged.
Published benchmarks put it above Claude Opus 5
In the agent-style benchmarks listed on the model card at the highest reasoning setting, V4.1-Flash scored 90.6 on Terminal-Bench 2.1, above Claude Opus 5 at 89.1. On DeepSWE v1.1, it posted 74.2%, roughly in line with Opus 5 at 74.0% and well above V4-Pro at 62.7%.
SiliconANGLE also cited DeepSeek data showing GPT-5.6 Sol at 88.8 on Terminal-Bench 2.1.
At the base-model level, V4.1-Flash-Base scored 74.1 on MMLU-Pro, 79.4 on HumanEval, and 93.0 on GSM8K, slightly ahead of V4-Pro-Base at 73.5, 76.8, and 92.6, respectively.
All of those figures came from DeepSeek’s own published evaluations. Independent third-party validation has not yet been completed.
Off-peak input pricing starts at $0.15 per million tokens
SiliconANGLE reported that API pricing for off-peak hours is set at $0.15 per million input tokens for uncached requests and $0.60 per million output tokens. Peak-hour pricing is double the off-peak rate.
For V4-Pro, output pricing was cut from $3.96 to $1.20, a drop of about 70%. The report noted that DeepSeek had raised V4-Pro pricing in mid-August and reduced it again alongside the new model launch.
Open-weight rankings continue to shift
Chain News reported in June that Zhipu GLM-5.2 had overtaken DeepSeek V4 to top the open-weight rankings. Leadership among open-weight models has changed hands repeatedly over the past few months.
Set against other launches this week, ABMedia said DeepSeek is sticking with a strategy of pushing its smallest model to flagship-level performance and then open-sourcing it at very low cost. The article contrasted that with OpenAI’s API product pricing and Sakana’s approach of lowering costs through orchestration systems, framing the three as different pricing paths in the same week.

