Timeline

Google launches Gemini 2.5 Flash in preview

A cheaper, faster sibling of Gemini 2.5 Pro with a configurable 'thinking budget' letting developers trade reasoning depth against cost and speed.

  • Models & capabilities
  • Minor

Google released Gemini 2.5 Flash in preview, calling it its first fully hybrid reasoning model: developers could turn the model’s step-by-step “thinking” on or off, or set a configurable thinking budget — up to 24,576 tokens — to trade reasoning depth against cost and response time. With the budget set to zero, Google said the model matched the speed of the prior Gemini 2.0 Flash generation while still improving on its output quality.

The model arrived three weeks after Gemini 2.5 Pro and was pitched as the cost-efficient counterpart to it — Google described 2.5 Flash as sitting between 2.0 Flash’s speed and 2.5 Pro’s reasoning strength, and said it led on price-to-performance among reasoning models it compared itself against. On the LMArena leaderboard’s “Hard Prompts” category, Google said the model ranked second only to 2.5 Pro itself. It was made available through Google AI Studio, Vertex AI and the Gemini API under the model ID gemini-2.5-flash-preview-04-17.

The release extended the hybrid-reasoning approach Google had introduced with 2.5 Pro down to a cheaper, faster tier, following the same pattern set by Gemini 1.5 Flash the previous year: a flagship model establishes a new capability, and a Flash variant then makes an economical version of it available for high-volume production use. The explicit thinking-budget control — rather than a fixed amount of reasoning baked into the model — became a recurring feature Google carried into subsequent Gemini releases.