How the biggest models can afford to be big: by never working all at once.
A hospital does not have its entire medical staff examine you. At the front desk they route you: this one is for the cardiologist, this for the eye doctor, this for the lab. The other doctors are seeing other patients meanwhile. An MoE model is a hospital for words: a small dispatcher decides which specialist takes exactly this piece of the sentence, while the rest stay silent and burn no power.
The clever part is that the specialization is not assigned by people - it self-organizes during training. Nobody says “you will handle chemistry.” Some parts simply start firing on code, others on numbers, others on rare languages. There are no labels; the division works.
Why this matters if you don't build models
Because it explains the numbers in the news. When you read that a model has an enormous parameter count yet runs fast and cheap, MoE is almost always underneath: the total size is the warehouse, the active part is the bill. And when you compare models by size, the comparison is no longer fair if one is dense and the other a mixture of experts.
Our read: MoE is one more proof that progress in AI does not come only from “more”. It comes from more, arranged more cleverly. The warehouse grows, but you only pay for what leaves it.