Gemini 1.5 Pro ships a million-token context window
A mixture-of-experts model that matched Gemini 1.0 Ultra on many tasks at lower compute, offered in limited preview with up to a million tokens of context.
- Models & capabilities
- Major
Google introduced Gemini 1.5 Pro, describing it as a mid-sized model built on a mixture-of-experts architecture that reached performance comparable to the larger Gemini 1.0 Ultra on many benchmarks while using substantially less compute. The general-availability context window was 128,000 tokens, but Google offered a limited preview extending that to one million tokens — roughly an order of magnitude beyond what production models had previously offered — through Vertex AI and AI Studio.
Google’s own demonstrations leaned on that expanded window rather than on conventional benchmark scores: the model correctly answering questions about a 402-page transcript of the Apollo 11 mission, retrieving plot details after being shown a 44-minute silent film, reasoning over more than 100,000 lines of code, and, in an oft-repeated example, learning to translate into Kalamang — a language with fewer than 200 speakers, with no meaningful presence in web text — after being given only a grammar manual, a dictionary and some parallel sentences in its context.
The release landed the same day as OpenAI’s Sora video-generation preview, which drew more immediate public attention, but Gemini 1.5 Pro’s context-length jump reset expectations about what “in-context learning” could substitute for. Competitors’ context windows had mostly sat in the 8,000–200,000 token range; within the year, long-context support at or near a million tokens became a standard entry in frontier-model system cards rather than a distinguishing feature, and “needle in a haystack” retrieval tests over very long inputs became a routine part of model evaluation.