Google Cloud split its new TPU chips in two - 8t for training, 8i for inference. Price-performance climbs up to 2.7x, and a single training cluster reaches over a million chips.
- Google split the TPU generation in two: 8t for training and 8i for inference (April 22).
- 8t: 12.6 PFLOPs and up to 2.7x better price-performance; 8i: 10.1 PFLOPs and up to 80% better inference.
- The scale: over 1 million chips in a single training cluster; the Virgo network links 134 000+ chips.
One chip for training. Another for thinking. Google Cloud split the new TPU generation exactly that way - 8t for training, 8i for inference (execution and reasoning). The architecture splits because the work is no longer one thing.
What sticks and what's noise here? The key is the split itself - training versus thinking. For a long time one chip did everything. Now Google admits that training a model and running it are different jobs - and it builds different hardware for each, to save power and money separately.
The number 'a million chips' sounds abstract, but it means something concrete. The next giant models don't hinge on one smarter chip, but on whether you can orchestrate an entire factory of them without the network choking. From there on, the race is as much logistics as silicon.