Timeline

OpenAI partners with Cerebras for 750MW of compute

The multi-year, reportedly $10bn-plus deal covers wafer-scale chips for fast inference — not model training — deployed in phases through 2028.

  • Compute & infrastructure
  • Notable

OpenAI and Cerebras Systems announced a multi-year agreement to deploy 750 megawatts of Cerebras’s wafer-scale AI chips for OpenAI’s inference workloads — the compute used to serve responses to users, rather than to train new models. Reporting put the value of the deal above $10 billion, with capacity coming online in phases through 2028.

Cerebras chips integrate compute, memory and interconnect on a single, dinner-plate-sized wafer rather than the smaller dies used in GPUs, which the companies said let them return chatbot and coding-agent responses up to fifteen times faster than GPU-based inference systems. OpenAI framed the deal as aimed at low-latency use cases — coding agents and voice interaction in particular — where response speed affects usability more directly than in offline or batch workloads, rather than as a substitute for the Nvidia and AMD GPU capacity underpinning its training runs.

For Cerebras, a company that had built its business selling wafer-scale hardware as a niche alternative to Nvidia, a multi-year commitment of this size from OpenAI — one of the largest buyers of AI compute in the world — represented a significant validation of the wafer-scale approach and diversified OpenAI’s compute supply chain beyond its existing large commitments to Nvidia, AMD, Microsoft, Oracle and others, part of a broader pattern in which OpenAI spent 2025 and early 2026 assembling compute capacity from as many suppliers as it could sign.

In the commentary

What people were saying around this time — external links, from the record's commentary rail.