Timeline

Google releases Gemini 3.5 Transcribe, a speech-to-text model

Google reported a 5.04% word-error rate on the FLEURS benchmark and a 70% latency cut from its previous Chirp 3 model, across more than 85 languages.

  • Models & capabilities
  • Benchmarks & progress
  • Minor

Google released Gemini 3.5 Transcribe, a speech-to-text model offered through both a real-time streaming API and a pre-recorded-audio API. The model automatically detects and transcribes more than 85 languages, and adds speaker attribution with timestamps for up to three speakers in pre-recorded audio, with support for larger groups described as experimental.

On Google’s own measurements against the FLEURS benchmark, the model scored a 5.04% word-error rate in non-streaming mode and 5.50% in streaming mode, and the company said time to final transcription improved by 70% compared with its previous Chirp 3 model — enough, it said, for sub-second latency in interactive voice applications.

The model is rolling out in public preview across developer and enterprise channels alike: through the Live and Interactions APIs, inside products including the Gboard keyboard’s Rambler dictation feature on Android, Google Antigravity, the Gemini app on macOS and Google AI Studio, with Chrome integration described as forthcoming, and through the Gemini Enterprise Agent Platform and Gemini Enterprise for Customer Experience. The breadth of the rollout — spanning consumer keyboard input, developer tooling and enterprise customer-service products in a single release — situated Transcribe as infrastructure for other Google products rather than a standalone consumer launch.

In the commentary

What people were saying around this time — external links, from the record's commentary rail.