On 10 September DeepSeek released V4.1-Flash, the smallest model of a new architecture family, with 552 billion parameters and only 8 billion active when reading input. From 14 September all requests to V4-Pro are routed to it at its prices. The weights are under MIT.
- Causal encoder-decoder: 8 billion active parameters for input, 16 for output.
- Per DeepSeek, the key-value cache needs a quarter of the HBM and an eighth of the SSD space compared with the previous generation.
- V4-Flash and V4-Flash-Vision-Exp are retired; new prices apply from 10 September, with off-peak at 50% of peak.
DeepSeek V4-Pro lasted four weeks as a generally available model before the company announced it is phasing it out.
And it isn't replacing it with something bigger. It is replacing it with Flash, the smaller and cheaper one, which by their numbers beats it.
Why input and output cost differently
Agents read far more than they write. A whole file, the conversation history, a tool result, then three lines of answer. A model that spends less precisely on reading is built for them.
DeepSeek says plainly why: cache-hit charges often make up a large share of an agent's costs, and compressing the cache cuts them. This is engineering driven by the bill, and I trust it more than a benchmark.
A note on the numbers. 552 billion is the official figure from the announcement. The model card lists a separate conditional memory of 196 billion parameters, so if you count the files on Hugging Face you will arrive at more. There is no error, it is simply counted differently.
From 14 September your API requests to deepseek-v4-pro go to a different model without you having changed a line of code. Run your own tests before that date, not after.