MTIA 300 is Meta's first accelerator built to train models, not just serve answers from finished ones. It targets recommendation and ranking systems, with the network built straight into the silicon. By Meta's own numbers, total communication time drops 3.9x versus an equivalent GPU cluster.
- Meta's first chip built to TRAIN models, not just run them.
- Twelve in-house 800 gigabit-per-second network cards built into the chip itself, for 1.2 terabytes per second total.
- By Meta's own numbers, parallel communication costs under 0.5 percent of throughput, versus more than 20 percent on GPU clusters.
Open the lid on Meta's new chip and you find something that usually sits on its own board. Twelve in-house network cards, fused straight into the silicon. Not next to the chip. Inside it.
Until now MTIA did one thing well: answer fast once a model was already trained. MTIA 300 is the first in the family built to teach the model, not just run it. The target is recommendation and ranking systems, the ones that decide what you see in your feed.
Communication is not an afterthought
On GPU chips, computing and talking to other machines queue on the same line. Send a message while the chip is computing and throughput on typical clusters drops by more than 20 percent. Here the drop is under half a percent, because messages have their own modules and never enter that race.
The number sounds technical, but it gets paid in real hours. A recommendation model does not train on one chip, it trains on thousands that constantly trade results. Every second lost in that exchange multiplies by how many of them there are.
Everything here was measured by Meta, on their own models. The chip runs in a rack today, and that part is fact. Whether the numbers hold when someone outside repeats them is something we will find out from someone outside.