Timeline

Google ships Gemini 3.8 Live voice models

The Extended Thinking variant reasons mid-conversation and makes background tool calls without breaking the dialogue; both support automatic switching across 97+ languages.

  • Models & capabilities
  • Minor

Google shipped Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, a pair of real-time speech-to-speech models, alongside a separate transcription model, Gemini 3.5 Transcribe.

The base model is built for low-latency, efficient voice interaction; the Extended Thinking variant is designed to reason through harder requests mid-conversation — Google described it as able to “reason and execute tasks while maintaining the flow of conversations” — including making background tool or API calls without pausing the spoken exchange. Both models support automatic language switching across more than 97 languages and can incorporate visual context. Google said the Extended Thinking model topped Artificial Analysis’ Speech-to-Speech Quality Index, while the base model placed second on a separate speech-agent leaderboard. Gemini 3.5 Transcribe, a dedicated speech-to-text model, was quoted at a 4.0% word error rate for streaming transcription and 2.6% for non-streaming, across more than 85 languages.

The models are available through the Gemini API, Google AI Studio and Gemini Enterprise, and are being rolled into consumer products including Search Live, the Gemini app and Google Workspace for paid subscribers. The release continues a pattern of frontier labs treating low-latency voice as a distinct capability track from text generation, competing on measures — turn-taking, mid-conversation reasoning, tool use without dead air — that standard text or image benchmarks do not capture.

In the commentary

What people were saying around this time — external links, from the record's commentary rail.