The Allen Institute has shown a rebuilt training system for mixture-of-experts models. Per Ai2, on 8 NVIDIA B300 chips a 47 billion parameter model trains 2.7 times faster than with their old code, and the largest test run passes 1.2 trillion parameters on 512 GPUs. They ship a report, code and a demo. No models.
- Experts grow from 8 to 128 at roughly 3.2 billion active parameters per token; total capacity jumps from 4.6 to 47 billion with under 5% loss in speed.
- The largest run: 1.2 trillion parameters (58.36 billion active) on 512 GPUs, peaking at 858 TFLOP/s per chip. With DeepEP v2 they reach 2.38 trillion.
- Released: a technical report, code on GitHub and a demo. No checkpoints in the post. The code in the allenai/Olmo-core repository is under Apache 2.0.
52,000 against 19,400. Tokens per second on a single chip, same model, same hardware, new code. The rest of Ai2's post is how they got there.
Be careful about what this is and what it is not. There is no new model. Nobody downloads a new Olmo today. This is the tool a model is made with, and it comes from a team that so far has published data, code and weights together.
This is backstage work, like at any event: power, cables and the schedule nobody in the audience sees. The network between 512 chips is exactly that kind of work. When 128 experts have to swap tokens over thousands of links, you lose time waiting, not computing. Ai2 says the jump from 8 to 128 experts costs it under 5% in speed. Their number, their tests.
Who is it for? Universities and small labs that have chips but not a hundred engineers to write this kind of code. Until now they chose between the closed code of the big players and their own, which does not scale. Now there is a third way, and it is free.
The test is simple. It will show when someone outside Ai2 trains a model with it and releases it. Until then it is good code with a good report.