OpenAI publishes first benchmark results for its Jalapeno chip
Benchmarked on the public InferenceX suite against three open-weight models, the chip delivered 1.5–1.9 times more throughput per watt than comparison hardware.
- Compute & infrastructure
- Benchmarks & progress
- Notable
OpenAI published the first disclosed benchmark results for Jalapeño, the inference chip it has been developing with Broadcom since 2024, testing it against Nvidia’s current-generation GB200 and GB300 systems on InferenceX, a public benchmark maintained by SemiAnalysis that normalises results for power draw. Running three open-weight models chosen to span different workload sizes — GPT-OSS 120B, DeepSeek R1 and the roughly one-trillion-parameter Kimi K2.5 — OpenAI reported 1.5 to 1.9 times more tokens processed per kilowatt than the comparison systems, and 1.7 to 3.6 times lower end-to-end latency, running at 700 watts against the roughly 1.2 to 1.4 kilowatts the Nvidia parts draw.
The figures were OpenAI and Broadcom’s own, tested on workloads and comparison points they chose, rather than results from an independent lab — though InferenceX’s methodology is public and the same suite has been used to benchmark other vendors’ hardware. They also gave OpenAI something more specific to point to than the launch-day claim from Broadcom chief executive Hock Tan that early samples ran “roughly 50% cheaper” than typical GPUs, a figure disclosed in June without a stated comparison baseline.
OpenAI vice president of hardware Richard Ho told reporters on a press call that deployment would begin at the end of 2026 “in very small volumes,” with more substantial rollout following in 2027 — a more cautious framing than the “before end of 2026” timeline given at unveiling. He said a second-generation chip aimed at further improving performance per watt was already well into development, with a third generation targeting cheaper, lower-latency serving in earlier planning.
The release coincided with an essay from OpenAI chief financial officer Sarah Friar arguing that gains like Jalapeño’s compound with model efficiency and routing to lower the cost of a given amount of useful AI work — positioning the chip as evidence for that argument rather than as a standalone hardware story.
In the commentary
What people were saying around this time — external links, from the record's commentary rail.