we_are_coded.by CODE · The world, decoded
БГ
DeepSeek

DeepSeek shipped vision over the API and counts an image as up to 384 tokens

DeepSeek API DocsFrontier

An experimental variant of their fast model now understands images too, not only text. You pay for it at the text model's rate. For the promised jump on the multimodal agentic tests there is not a single published number.

In short
  • DeepSeek-V4-Flash-Vision-Exp comes out over the API on 21 August, as an experiment only.
  • An image is billed as up to 384 tokens, at the price of V4-Flash.
  • They claim a big jump and closeness to Opus 4.8, but publish no numbers.
Checked on21 August 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

One image counts with them as at most 384 tokens. That is what a long paragraph of text weighs.

Models have been looking at pictures for a long time. The new part is the price.

Until now the frame was the expensive input. You send a screenshot to an agent and you look at the bill before you send it again.

The facts: on 21 August 2026 DeepSeek announced DeepSeek-V4-Flash-Vision-Exp over its own API. You call it with model='deepseek-v4-flash-vision-exp'. Per DeepSeek the model covers the text abilities of V4-Flash, including agents, reasoning and world knowledge, and makes a big jump above it on the multimodal agentic tests, coming close to Opus 4.8. There are no numbers in the announcement and no test names. Images are turned into tokens for billing, up to 384 per image, at the price of V4-Flash. Mixed input of text and images is accepted, and the images themselves are passed as base64, an external address, or through the new Files API, which is free and lets you upload once and then point at the file by file_id. Supported are Chat Completions, Messages and Responses. DeepSeek Harness 0.1.1 comes out alongside the model. No open weights are announced.

The claim with no numbers

"Close to Opus 4.8" is their sentence, not a measurement. No table, no test name, no conditions. I write it down as it was said and stop there.

I will take it when somebody who is not selling the model measures it. Until then it is a promise with an emoji.

The expensive part of vision was never the vision. It was the bill for every next frame.

The Files API is the small thing you will feel in the work. You upload the screen once and then hand it over by identifier, instead of uploading it on every call. With an agent that looks at the same interface twenty times, that is a fair amount of bandwidth saved.

There is a difference with last week too. Then they handed out weights with a licence that gets in nobody's way. This model sits behind their own door, labelled experimental, which means it can also be swapped under you.

If you have an agent that until now sent frames only when it really had to, count that bill again. The number has changed, and the limit may no longer stand.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→DeepSeek API Docs: DeepSeek-V4-Flash-Vision-Exp Release (21.08.2026)
Original: https://wearecoded.com/en/articles/deepseek-v4-flash-vision-exp.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news