Timeline

Google launches Ironwood, its seventh-generation inference-focused TPU

Google said a full 9,216-chip pod delivered 42.5 exaflops, more than 24 times the compute of the El Capitan supercomputer, with roughly four times Trillium's per-chip compute in the FP8 format.

  • Compute & infrastructure
  • Notable

Google announced Ironwood, its seventh-generation Tensor Processing Unit, at Google Cloud Next 25, describing it as the first TPU generation designed specifically for inference rather than training. Each chip offered 192GB of HBM memory — six times the previous generation, Trillium — at 7.37TB/s of memory bandwidth, and native support for the lower-precision FP8 format, which Google said gave it roughly four times Trillium’s peak compute per chip.

Ironwood shipped in two pod configurations, a 256-chip version and a 9,216-chip version; Google said the larger pod delivered 42.5 exaflops of aggregate compute, which it described as more than 24 times the compute power of El Capitan, then the world’s largest supercomputer. Google also claimed roughly double Trillium’s performance per watt, and said the design was aimed at “thinking models” — large reasoning and mixture-of-experts systems whose inference cost, run continuously in production, was becoming a larger share of total compute spend than training.

The launch reflected a broader shift in how AI infrastructure spending was being justified: as reasoning models made inference itself computationally expensive, chip designers who had until then optimised primarily for training throughput began building hardware explicitly for serving those models at scale, with Ironwood positioned as Google’s answer to Nvidia’s inference-oriented GPU lines.