we_are_coded.by CODE · The world, decoded
БГ
Meta

Meta unveils MTIA 300, its first chip built to train models

Meta EngineeringInfra

MTIA 300 is Meta's first accelerator built to train models, not just serve answers from finished ones. It targets recommendation and ranking systems, with the network built straight into the silicon. By Meta's own numbers, total communication time drops 3.9x versus an equivalent GPU cluster.

In short
  • Meta's first chip built to TRAIN models, not just run them.
  • Twelve in-house 800 gigabit-per-second network cards built into the chip itself, for 1.2 terabytes per second total.
  • By Meta's own numbers, parallel communication costs under 0.5 percent of throughput, versus more than 20 percent on GPU clusters.
Checked on25 August 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

Open the lid on Meta's new chip and you find something that usually sits on its own board. Twelve in-house network cards, fused straight into the silicon. Not next to the chip. Inside it.

Until now MTIA did one thing well: answer fast once a model was already trained. MTIA 300 is the first in the family built to teach the model, not just run it. The target is recommendation and ranking systems, the ones that decide what you see in your feed.

The facts: MTIA 300 is Meta's first accelerator optimized for training rather than inference, aimed at recommendation and ranking models. It carries 12 in-house network cards for direct memory-to-memory transfer, 800 gigabits per second each, built into the chip itself, for 1.2 terabytes per second of total I/O. It has 216 gigabytes of HBM3E memory, a 12-by-6 grid of compute elements and 16 separate message-processing modules. Within a single rack, chip-to-chip communication reaches 940 gigabytes per second. By Meta's own numbers, measured on production models, total communication time drops 3.9x versus an equivalent GPU cluster. Source: Meta's engineering blog.

Communication is not an afterthought

On GPU chips, computing and talking to other machines queue on the same line. Send a message while the chip is computing and throughput on typical clusters drops by more than 20 percent. Here the drop is under half a percent, because messages have their own modules and never enter that race.

On a typical cluster, the conversation steals from the computing. Here, parallel communication costs under half a percent of throughput.

The number sounds technical, but it gets paid in real hours. A recommendation model does not train on one chip, it trains on thousands that constantly trade results. Every second lost in that exchange multiplies by how many of them there are.

Everything here was measured by Meta, on their own models. The chip runs in a rack today, and that part is fact. Whether the numbers hold when someone outside repeats them is something we will find out from someone outside.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→Meta Engineering: MTIA 300
Original: https://wearecoded.com/en/articles/meta-mtia-300-chip-za-obuchenie.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news