Timeline

Google introduces Gemini 1.5 Flash

A lighter, cheaper sibling of Gemini 1.5 Pro, produced by distilling the larger model and matching its one-million-token context window.

  • Models & capabilities
  • Minor

At Google I/O 2024, Google announced Gemini 1.5 Flash, a smaller, faster sibling of the Gemini 1.5 Pro model it had shipped three months earlier. Google said it built Flash through “distillation” — transferring what it called the most essential knowledge and skills from the larger 1.5 Pro model into a lighter-weight one — and positioned it for high-volume, high-frequency tasks such as chat, summarisation and data extraction from large documents, where latency and serving cost mattered more than peak capability.

Despite the smaller footprint, Flash launched in public preview with the same one-million-token context window as 1.5 Pro, extending the long-context capability introduced with the Pro model to a cheaper tier rather than confining it to Google’s flagship. The company described Flash as the fastest model then served through the Gemini API. Both models became available in over 200 countries as a preview, with general availability expected the following month.

The release established a pattern Google repeated through later Gemini generations: a “Pro” model for maximum capability alongside a “Flash” variant trading some quality for speed and lower cost, aimed at developers running the model at scale rather than end users seeking the best possible answer. That two-tier structure — later joined by an even lighter “Flash-Lite” tier — became a fixture of Google’s model releases and was mirrored, in various forms, by competitors offering their own fast, cheap alternatives to flagship models.

Referenced by