we_are_coded.by CODE · The world, decoded
БГ
Alibaba

The new Qwen runs 6 billion parameters out of 125 - and its licence is no longer Apache

Alizila (Alibaba Group) · event date: 27 August 2026Frontier

Alibaba announced Qwen3.8-Flash on 27 August as a preview of the architecture behind Qwen4. The weights really are free to download, but they are named Qwen3.8-Flash-Next and come under a community licence that requires a separate agreement if you build a coding or office assistant.

In short
  • 125 billion parameters in the main model, 51 billion additional N-gram embeddings, 6 billion active per token.
  • The press release and the repository describe two different things: Qwen3.8-Flash-Next is the weights, Qwen3.8-Flash is the paid version in their cloud.
  • The licence is Qwen Community 1.0, not Apache 2.0: model-as-a-service or a coding assistant needs separate permission.
Checked on28 August 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

I opened the repository first and the press release second. Turns out the order matters.

On 27 August Alibaba announced Qwen3.8-Flash - a multimodal mixture-of-experts model they say holds its own against far more expensive rivals, and an early look at the architecture behind Qwen4. The numbers are good: 125 billion parameters in the main model, 51 billion more for N-gram embeddings, and only 6 billion doing work on each token. Native context of 262 thousand tokens, stretchable to a million.

The weights, however, are not called that. What sits on Hugging Face is Qwen3.8-Flash-Next and its FP8 sibling, uploaded three days earlier. Their own model card explains it honestly: Qwen3.8-Flash is the official version in their paid cloud, built on Next, with a million tokens of context by default and built-in tools. So you download one thing, and the numbers in the announcement belong to another.

The facts: on 27 August 2026 Alibaba announced Qwen3.8-Flash through Alizila, its own news channel. Per the company, the model is a multimodal MoE with 125 billion parameters in the main model, 51 billion N-gram embeddings and 6 billion active parameters per token, with a native context of 262,144 tokens extendable to 1 million. Pricing through Model Studio and Qwen Cloud is 1 yuan (USD 0.16) per million input tokens and 3 yuan (USD 0.47) per million output tokens. Alibaba compares it with DeepSeek-V4-Flash and Claude-Opus-4.6 on SWE-bench Pro, CoWorkBench, Toolathlon Verified, MathVision, AndroidWorld and ERQA. The architectural changes listed are hybrid attention combining Gated DeltaNet with Qwen Sparse Attention, a Gated Residual mechanism, N-gram embeddings and the Muon optimiser; training cost roughly one ninth of the resources of Qwen3.7-Plus, a model three times its size. In QwenWork the redesigned standard mode cuts tokens consumed per task by 75 per cent and roughly doubles generation speed. The weights on Hugging Face are published as Qwen/Qwen3.8-Flash-Next and Qwen/Qwen3.8-Flash-Next-FP8, created on 24 August 2026, under the Qwen Community License 1.0. The repository Qwen/Qwen3.8-Flash is not publicly accessible. For comparison: Qwen3.8-27B from 5 August is under Apache 2.0, while the Qwen3.8-2.4T-A95B giant from 8 August is also under the community licence.

What the licence actually says

I read all of it, because for us that is a required step before any outside model goes near real work. The community licence grants the right to use, modify, distribute and sell - and then attaches two conditions.

The first is soft: cross 100 million monthly users or 20 million dollars in monthly revenue, and the model's name has to appear prominently in your product's interface. The second is not soft. If your business is giving access to models as a service, or building an assistant for coding and office work, you need separate permission from Qwen before any commercial use. Internal use stays free, as long as you do not expose the model, its outputs or its capabilities to a third party.

Open weights and an open licence stopped being the same sentence a while ago.

I am not writing this to scold them. The clause aimed at coding assistants is understandable - they sell exactly that product and would rather not fund their own competition. I am writing it because if you build software and read a headline with the words open weights, it is easy to decide you are free. A downloaded model is not the same as permitted work.

Download it, try it, measure it - nobody is asking about that. But before it goes under a paid product, open the LICENSE file in their repository and work out which of the two paragraphs you fall into.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→Alizila (Alibaba Group) - Alibaba Releases Qwen3.8-Flash, 27.08.2026→Hugging Face - Qwen/Qwen3.8-Flash-Next (model card and licence)
Original: https://wearecoded.com/en/articles/qwen38-flash-next-tegla-i-licenz.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news