we_are_coded.by CODE · The world, decoded
БГ
DeepSeek

DeepSeek released V4.1-Flash and retired its own flagship: requests to V4-Pro now go to the smaller model

DeepSeek API Docs · event date: 10 September 2026Frontier

On 10 September DeepSeek released V4.1-Flash, the smallest model of a new architecture family, with 552 billion parameters and only 8 billion active when reading input. From 14 September all requests to V4-Pro are routed to it at its prices. The weights are under MIT.

In short
  • Causal encoder-decoder: 8 billion active parameters for input, 16 for output.
  • Per DeepSeek, the key-value cache needs a quarter of the HBM and an eighth of the SSD space compared with the previous generation.
  • V4-Flash and V4-Flash-Vision-Exp are retired; new prices apply from 10 September, with off-peak at 50% of peak.
Checked on1 October 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

DeepSeek V4-Pro lasted four weeks as a generally available model before the company announced it is phasing it out.

And it isn't replacing it with something bigger. It is replacing it with Flash, the smaller and cheaper one, which by their numbers beats it.

The facts: on 10 September 2026 DeepSeek released DeepSeek-V4.1-Flash, the smallest model in a new architecture family, with native image understanding. Per the company it is an MoE model with 552 billion parameters and a "causal encoder-decoder" architecture in which 8 billion parameters are active when processing input and 16 billion when generating. Compared with the previous generation, the key-value cache needs 1/4 of the HBM and 1/8 of the SSD storage. DeepSeek writes that tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed and total runtime. V4-Flash and V4-Flash-Vision-Exp are retired, with the old names temporarily routing to the new model; from 04:00 UTC on 14 September all deepseek-v4-pro requests route to V4.1-Flash at its rates until V4.1-Pro launches. New prices apply from 04:00 UTC on 10 September, with off-peak rates at 50% of peak. The weights are on Hugging Face under the MIT licence; the model card describes 552 billion backbone parameters, a separate Engram conditional memory of 196 billion parameters, and context of up to one million tokens.

Why input and output cost differently

Agents read far more than they write. A whole file, the conversation history, a tool result, then three lines of answer. A model that spends less precisely on reading is built for them.

DeepSeek says plainly why: cache-hit charges often make up a large share of an agent's costs, and compressing the cache cuts them. This is engineering driven by the bill, and I trust it more than a benchmark.

The model is built for the agent that reads a hundred lines to write three.

A note on the numbers. 552 billion is the official figure from the announcement. The model card lists a separate conditional memory of 196 billion parameters, so if you count the files on Hugging Face you will arrive at more. There is no error, it is simply counted differently.

From 14 September your API requests to deepseek-v4-pro go to a different model without you having changed a line of code. Run your own tests before that date, not after.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→DeepSeek API Docs - DeepSeek-V4.1-Flash Release, 10.09.2026→Hugging Face - deepseek-ai/DeepSeek-V4.1-Flash (model card and licence)
Original: https://wearecoded.com/en/articles/deepseek-v41-flash-zamenya-v4-pro.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news