Google announces Trillium, its sixth-generation TPU
Trillium delivers a claimed 4.7x peak compute increase per chip over the prior TPU generation and trains models including Gemini 1.5 Flash and Gemma 2.
- Compute & infrastructure
- Minor
Google announced Trillium, its sixth-generation Tensor Processing Unit, at its I/O developer conference. Compared with the prior generation, TPU v5e, Google claimed a 4.7x increase in peak compute per chip, doubled high-bandwidth memory capacity and interconnect bandwidth between chips, and 67% better energy efficiency — figures reported by Google itself and not independently benchmarked at announcement.
The chip was already in internal use: Google said Trillium had trained Gemini 1.5 Flash, Imagen 3 and Gemma 2, and would be used for future Gemini models. A single pod could scale to 256 chips, and with Google’s multislice networking, deployments could link many pods into a building-scale cluster of tens of thousands of chips. General availability through Google Cloud was set for later in 2024, with customers able to register interest at announcement rather than order immediately.
The announcement was part of a broader push by Google to position its own silicon, developed over several TPU generations, as an alternative to buying Nvidia GPUs for both training and inference at scale — a pitch aimed less at matching any single Nvidia chip than at offering customers a second source of AI compute.