DeepSeek has released V4.1 Flash, a new model with 552 billion parameters built on what it calls a Causal-Encoder-Decoder architecture. The company said the model natively supports image understanding and is currently the smallest model in this new architecture family. A key design change is that input and output are handled separately: only 8 billion parameters are activated when reading the input, while 16 billion are activated when generating a response. DeepSeek said that setup helps reduce inference costs for a model at the 552B scale. The company also said that after new pretraining and larger-scale reinforcement learning, V4.1 Flash has surpassed V4 Pro on its official benchmarks. Cache usage has also been reduced sharply, with HBM memory demand cut to one-quarter of the previous generation and SSD storage demand reduced to one-eighth. DeepSeek said those changes are aimed at lowering costs for long-context workloads and repeated Agent calls. The model is now available through the API under the name deepseek-flash.
DeepSeek has officially released V4.1 Flash, a 552-billion-parameter model built on a new Causal-Encoder-Decoder architecture. The company said the model also natively supports image understanding.
A new architecture that splits input and output processing
According to DeepSeek, V4.1 Flash is currently the smallest model in this new architecture family. The system handles input and output separately: it activates 8 billion parameters when reading the input and 16 billion when generating a response. DeepSeek said this design helps bring down inference costs even for a 552B-scale model.
Training and infrastructure changes
DeepSeek said V4.1 Flash has surpassed V4 Pro on the company’s official benchmarks after new pretraining and larger-scale reinforcement learning. The cache footprint has also been reduced substantially. Compared with the previous generation, HBM memory demand has been cut to one-quarter, while SSD storage demand has been reduced to one-eighth. DeepSeek said those reductions are mainly intended to lower costs in long-context use cases and repeated Agent calls.
API access is live
V4.1 Flash is now available through the API. Users can call it by switching the model name to deepseek-flash.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.