Timeline

OpenAI adds a much faster response tier for its flagship model, built on Cerebras chips

The 'Ultrafast' preview reaches up to 750 tokens per second, as much as 14 times OpenAI's standard speed, at the same intelligence level as GPT-5.6 Sol.

  • Models & capabilities
  • Compute & infrastructure
  • Minor

OpenAI began previewing “Ultrafast,” a new API tier for GPT-5.6 Sol that runs on Cerebras inference chips rather than OpenAI’s usual infrastructure. The company said the tier reaches up to 750 output tokens per second — as much as 14 times its standard response speed — while producing answers at what it described as the same intelligence level as the standard version of Sol, rather than a faster but weaker variant.

The pitch is latency, not capability: OpenAI framed Ultrafast around applications where response time matters as much as answer quality, including financial research, incident response, customer support, voice applications and live commerce, where a multi-second wait for a reasoning model breaks the interaction. Cerebras, which builds wafer-scale chips designed for fast sequential token generation rather than the parallel throughput GPUs optimise for, said the partnership let it “power” the tier rather than merely host it, an unusual arrangement for a frontier lab that otherwise runs almost entirely on Nvidia-based infrastructure.

The launch is a limited preview available initially to select customers rather than a general release, and neither company published independent benchmarking of the speed claim beyond their own figures. It is nonetheless a notable diversification: it marks one of OpenAI’s first public uses of non-Nvidia inference hardware at the frontier-model tier, at a moment when inference cost and latency, rather than raw capability, are an increasingly visible competitive axis among leading labs.