Timeline

OpenAI releases GPT-5.3-Codex-Spark, an ultra-low-latency coding model

Served on Cerebras' Wafer Scale Engine 3 rather than OpenAI's usual infrastructure, the smaller model hit over 1,000 tokens per second — about 15 times the standard Codex model's speed.

  • Models & capabilities
  • Minor

OpenAI released GPT-5.3-Codex-Spark as a research preview, a smaller sibling of GPT-5.3-Codex built specifically for real-time interactive coding rather than raw capability. Instead of running on OpenAI’s usual serving infrastructure, Codex-Spark runs on Cerebras’ Wafer Scale Engine 3, a chip purpose-built for low-latency inference.

The headline figure was speed rather than a benchmark score: OpenAI said the model generates more than 1,000 tokens per second, roughly 15 times faster than the standard Codex model, while still handling real coding tasks. At launch the model was text-only, with a 128,000-token context window, and was made available to ChatGPT Pro subscribers through the Codex app, its command-line interface and the VS Code extension.

The release followed the standard GPT-5.3-Codex model earlier in the month and reflected a wider push among coding-model vendors toward interactive, low-latency serving — using specialised inference hardware to make edit-and-see-results loops feel closer to instantaneous, rather than pursuing further gains on capability benchmarks.