Timeline

NVIDIA announces the Blackwell architecture

The GB200 NVL72 system was claimed to cut LLM inference cost and energy per token by up to 25 times against Hopper-generation hardware.

  • Compute & infrastructure
  • Major

NVIDIA unveiled the Blackwell GPU architecture, presented by chief executive Jensen Huang, at its GTC developer conference, positioning it as the successor to the Hopper generation that had powered most of the 2023–24 AI training boom. Each Blackwell GPU packs 208 billion transistors, built on a custom TSMC 4NP process, and the company said the architecture was designed specifically to make training and running models up to roughly 10 trillion parameters practical.

The headline system, the GB200 NVL72, links 72 Blackwell GPUs through a fifth-generation NVLink interconnect offering 1.8TB/s of bidirectional bandwidth per GPU, functioning as a single very large accelerator for large-model workloads. NVIDIA said the system delivered up to 30 times the inference performance of an equivalent number of H100 GPUs on large-language-model workloads, while cutting cost and energy consumption per token by up to 25 times — figures that, like most NVIDIA-supplied performance claims, described a specific workload configuration rather than a general multiplier, and were not independently benchmarked at announcement.

NVIDIA listed essentially every major cloud and AI company as a Blackwell customer or partner, including Amazon, Google, Microsoft, Meta, OpenAI, Oracle, Tesla and xAI, alongside server makers such as Dell, HP Enterprise, Lenovo and Supermicro, and cloud providers including CoreWeave.

The announcement mattered less for any individual specification than for what it confirmed about the pace of the buildout: NVIDIA was committing to an annual architecture cadence, and every major AI lab and hyperscaler was lined up to buy the next generation before the current one had finished shipping. Blackwell chips subsequently faced well-reported production and thermal-design delays before reaching customers in volume, but the announcement itself set the demand and capital-expenditure expectations that shaped compute planning across the industry for the following two years.