we_are_coded.by CODE · The world, decoded
БГ
DeepSeek

DeepSeek released the weights of its vision model under MIT and this time showed the numbers

Hugging Face · event date: 31 August 2026Frontier

On 31 August DeepSeek uploaded DeepSeek-V4-Flash-Vision-Exp to Hugging Face under the MIT licence. Ten days earlier the model was API-only, without a single image benchmark. Now the table is out, comparing the new model with their older one and with Opus 4.8.

In short
  • The first experimental multimodal model of the V4 family, built on V4-Flash with added vision modules.
  • Per DeepSeek: 36.5 on ApexBench against 26.2 for V4-Flash-0731; on text agent tasks it stays close to its predecessor.
  • The licence is MIT. At about 305 billion parameters, this is not a laptop model.
Checked on1 October 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

On 21 August we wrote that there was not a single number behind the promised jump on images. On 31 August the numbers came out, and together with the weights.

The order is backwards, but at least it is there.

The facts: on 31 August 2026 DeepSeek created the repository deepseek-ai/DeepSeek-V4-Flash-Vision-Exp on Hugging Face under the MIT licence. In the company's words, this is the first experimental multimodal model of the DeepSeek V4 family, built on the DeepSeek-V4-Flash architecture with added visual modules and continued training. According to the repository metadata the model has about 305 billion parameters. Per DeepSeek, on multimodal agent benchmarks it scores 36.5 on ApexBench (against 26.2 for DeepSeek-V4-Flash-0731 and 39.4 for Opus-4.8) and 27.3 on Agents' Last Exam (against 25.2 and 25.7), while on text agent tasks it stays close to its predecessor: 83.9 on Terminal Bench 2.1 against 82.7, and a slight drop on Cybergym, 75.3 against 76.7. DeepSeek notes that V4-Flash-0731 ignores the images in the input of the two multimodal benchmarks.

One footnote under the table takes away half the excitement. The old model does not see the images at all in those two benchmarks, so the jump from 26.2 to 36.5 is the difference between blind and sighted. The honest comparison is with Opus 4.8. There DeepSeek trails slightly on ApexBench and leads slightly on Agents' Last Exam.

Beating your own blind model is easy. What matters is how close it stands to Opus.

The licence is the other good news. MIT lets you take the model, change it and build a product on it without asking. With Qwen last week we wrote about a community licence with conditions. Here there are none.

You will not run it at home. About 305 billion parameters is a job for a provider. But any provider can now host it legally, and that is the road by which it will reach you.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→Hugging Face - deepseek-ai/DeepSeek-V4-Flash-Vision-Exp (model card and licence)
Original: https://wearecoded.com/en/articles/deepseek-v4-flash-vision-tegla-mit.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news