A big model isn't trained on one chip, but on thousands, constantly talking to each other, and that network between them weighs as much as the chips themselves. Worth knowing, because it explains why companies spend billions not just on chips, but on the connections between them.
Why thousands of chips talk to each other
Training a big model gets split into pieces across thousands of chips, and each one holds part of the calculation. At a given step, all of them have to gather what they've computed, combine it in one place, and then each one gets the combined result back, to keep going. That gathering and re-sending is called a collective operation. Everyone takes part at once, and no one can move on until the last one finishes its piece.
That's exactly why a lost packet of data here costs something an ordinary internet connection never does. At home, if a page loads half a second slower, you don't notice. In a cluster of thousands of chips, if one packet gets lost and has to be sent again, the whole exchange stops, and thousands of chips sit and wait on the single one that hasn't gotten its piece yet.
Who claimed the number
Meta built twelve network cards directly into its training chip, the MTIA 300, instead of keeping them separate on the board. By their numbers, the slowdown from parallel communication fell under half a percent, while the same slowdown in ordinary clusters runs twenty percent.
The same company also wrote its own protocol for connecting chips, MetaRoCE, and opened it publicly through the Open Compute Project, so others could check it themselves instead of taking their word for it. That's the second half of the story, and it repeats often: whoever builds very big ends up writing their own network too.