An experimental variant of their fast model now understands images too, not only text. You pay for it at the text model's rate. For the promised jump on the multimodal agentic tests there is not a single published number.
- DeepSeek-V4-Flash-Vision-Exp comes out over the API on 21 August, as an experiment only.
- An image is billed as up to 384 tokens, at the price of V4-Flash.
- They claim a big jump and closeness to Opus 4.8, but publish no numbers.
One image counts with them as at most 384 tokens. That is what a long paragraph of text weighs.
Models have been looking at pictures for a long time. The new part is the price.
Until now the frame was the expensive input. You send a screenshot to an agent and you look at the bill before you send it again.
The claim with no numbers
"Close to Opus 4.8" is their sentence, not a measurement. No table, no test name, no conditions. I write it down as it was said and stop there.
I will take it when somebody who is not selling the model measures it. Until then it is a promise with an emoji.
The Files API is the small thing you will feel in the work. You upload the screen once and then hand it over by identifier, instead of uploading it on every call. With an agent that looks at the same interface twenty times, that is a fair amount of bandwidth saved.
There is a difference with last week too. Then they handed out weights with a licence that gets in nobody's way. This model sits behind their own door, labelled experimental, which means it can also be swapped under you.
If you have an agent that until now sent frames only when it really had to, count that bill again. The number has changed, and the limit may no longer stand.