we_are_coded.by CODE · The world, decoded
БГ
MiniMax

MiniMax released a model that writes a full five-minute song, and handed out the weights

MiniMax-Music3, MiniMax's official profile (Hugging Face)Audio

Not a thirty-second clip, but a song with a beginning, a middle and an end that holds its theme all the way through. Underneath sit two language models on different levels, and the big one started life as Qwen3-8B.

In short
  • MiniMax Music 3 lands on 7 August with open weights and writes songs up to five minutes.
  • The architecture has two levels: a large model for structure, a small one for the sound frame by frame.
  • The licence is their own community one, not Apache and not MIT.
Checked on8 August 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

Thirty seconds of good sound is something everyone can do by now. The problem has always been what comes after them.

A song has construction. A verse that sets up, a chorus that returns and has to be the same while sounding bigger, a bridge that breaks the expectation right before the last chorus. Models fell apart there: a minute and a half in, they would start inventing a different song.

This one holds for five minutes.

The facts, per MiniMax's official profile on Hugging Face and the model card: MiniMax Music 3 was published on 7 August 2026, with 2.43 billion parameters in total in the downloadable weights. The architecture is hierarchical and separates the global from the local: a global language model of 8 billion parameters carries the long structure of the song, while a local one of 0.6 billion restores the fine acoustic detail within each frame. The global model is initialised from Qwen3-8B. Synthesis runs through Flow Matching with 2.4 billion parameters. The model takes instructions on arrangement: primary and secondary instruments, section-level evolution, groove, bass, percussion. The licence is their own, called the MiniMax-Music3 Community License, granting use, modification and distribution under conditions listed in the document itself.

The split that does the work

A large model thinks about where the song is going. A small model looks after how every moment sounds. The large one does not deal with acoustic detail, the small one does not try to remember what happened two minutes ago.

That is exactly the division of labour in a studio. One person holds the arrangement in their head, another sits over a single bar sorting out how the bass breathes.

The difference between sound and a song is memory. One fills seconds, the other remembers where it is going.

The small detail I liked most: the global model started out as Qwen3-8B. So they took a language model trained on text and taught it to think in musical units.

The same trick that made speech models good is now heading into music. The language part is not about words, it is about order.

The licence, before you get excited

It is not Apache and it is not MIT. It is their own community licence, with the conditions in the document itself. It grants use, modification and distribution, but it gets read before you put anything from it into a product you sell.

And one more thing that keeps me up. If a whole song now comes out of an instruction about groove and instruments, the next argument will not be a technical one. It will be about whose song it is, and what it was taught on in order to make it.

If you make music: download it and find where it breaks in the fourth minute. You will learn more there than from any demo on their site.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→MiniMax-Music3, MiniMax's official profile (Hugging Face)→MiniMax-Music3 Community License (the licence text)
Original: https://wearecoded.com/en/articles/minimax-music3-cyala-pesen-pet-minuti.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news