OpenAI launches GPT-4o with real-time voice
A single model handling text, vision and audio end to end, with conversational latency — and a voice that led to a public dispute with Scarlett Johansson.
- Models & capabilities
- Culture & impact
- Era-defining
OpenAI released GPT-4o — “o” for omni — a model trained natively across text, vision and audio rather than chaining separate systems together. The company reported average audio response latency of 320 milliseconds, comparable to human conversational turn-taking, where the previous voice mode had averaged several seconds because it pipelined transcription, generation and speech synthesis.
The demonstration emphasised interaction rather than benchmarks: interrupting the model mid-sentence, asking it to adjust its tone, pointing a camera at a maths problem, having two instances converse. It matched GPT-4 Turbo on English text and code while being faster and half the price through the API, and it was made available to free ChatGPT users — the first time a frontier-class model was not paywalled.
Within a week the launch was overtaken by a dispute about one of its voices. Scarlett Johansson said OpenAI had approached her twice about lending her voice, that she declined, and that the released voice named Sky was similar enough that friends could not tell them apart; Altman’s one-word post on the day, “her,” referencing the film in which Johansson voices an AI companion, was widely read as confirmation. OpenAI paused the voice and said it belonged to a different professional actress cast before any contact with Johansson.
The episode is a compact example of a recurring pattern in the period: a genuine technical advance whose public meaning was set by a question of consent that the developers had not treated as central. It also foreshadowed the fights over voice and likeness that followed with video generation.