On 31 August DeepSeek uploaded DeepSeek-V4-Flash-Vision-Exp to Hugging Face under the MIT licence. Ten days earlier the model was API-only, without a single image benchmark. Now the table is out, comparing the new model with their older one and with Opus 4.8.
- The first experimental multimodal model of the V4 family, built on V4-Flash with added vision modules.
- Per DeepSeek: 36.5 on ApexBench against 26.2 for V4-Flash-0731; on text agent tasks it stays close to its predecessor.
- The licence is MIT. At about 305 billion parameters, this is not a laptop model.
On 21 August we wrote that there was not a single number behind the promised jump on images. On 31 August the numbers came out, and together with the weights.
The order is backwards, but at least it is there.
One footnote under the table takes away half the excitement. The old model does not see the images at all in those two benchmarks, so the jump from 26.2 to 36.5 is the difference between blind and sighted. The honest comparison is with Opus 4.8. There DeepSeek trails slightly on ApexBench and leads slightly on Agents' Last Exam.
The licence is the other good news. MIT lets you take the model, change it and build a product on it without asking. With Qwen last week we wrote about a community licence with conditions. Here there are none.
You will not run it at home. About 305 billion parameters is a job for a provider. But any provider can now host it legally, and that is the road by which it will reach you.