Timeline

OpenAI brings its full-duplex voice model to the API

GPT-Live-1, which listens and speaks at once and delegates reasoning to a back-end model, opens to developers at $0.05/minute; OpenAI says it lifts Full Duplex Bench interactivity 30 points over its previous real-time model.

  • Models & capabilities
  • Minor

OpenAI released GPT‑Live‑1 in its API, a full-duplex voice model — one that listens and speaks at the same time — that had previously powered ChatGPT’s voice mode. Rather than the usual chain of separate speech-to-text, reasoning and text-to-speech systems, GPT‑Live‑1 handles the audio layer in a single model and delegates deeper reasoning and tool calls to a back-end model such as GPT‑6 Astra or a third-party model, which OpenAI argues removes the latency and brittle handoffs of stitched-together voice stacks.

OpenAI reported that GPT‑Live‑1 improves interactivity on the Full Duplex Bench by 30 percentage points over its previous GPT‑Realtime‑2.1 model, with turn-taking latency roughly halved (0.80 vs 1.41 seconds), and that paired with Astra it ranked first on Tau3, a benchmark of end-to-end voice-agent tasks. It cited early customers including Yelp, the language-learning app Speak — which said the model cut interruptions of learners by almost 80% — Cognition and Fin.

The model is available at $0.05 per minute for the front-end voice layer, with the back-end reasoning model and agent harness billed separately. The pricing and API access position it for high-volume telephony uses — reservations, order-taking, customer support — extending the agent-tooling OpenAI has been shipping through 2026 into real-time voice.

In the commentary

What people were saying around this time — external links, from the record's commentary rail.