Timeline

Google ships Gemini 3

Gemini 3 Pro reported a 1501 Elo score on LMArena and 91.9% on GPQA Diamond, prompting OpenAI to reportedly declare an internal 'code red' days later.

  • Models & capabilities
  • Benchmarks & progress
  • Major

Google released Gemini 3, its next flagship model, in two configurations: Gemini 3 Pro, available immediately across the Gemini app, Search’s AI Mode, Google AI Studio, Vertex AI and Google’s Antigravity coding environment, and Gemini 3 Deep Think, a higher-compute reasoning mode held back for further safety evaluation before release to Google AI Ultra subscribers.

Google reported state-of-the-art results across a wide set of benchmarks: a 1501 Elo score that topped the LMArena leaderboard, 91.9% on GPQA Diamond (93.8% for Deep Think), 37.5% on Humanity’s Last Exam without external tools (41.0% for Deep Think), 76.2% on SWE-bench Verified, and a top score on WebDev Arena. The model carries a one-million-token context window and was pitched as needing less explicit prompting to infer what a user wants, extending a general industry push toward agentic, tool-using systems rather than pure chat.

The release was read in the industry primarily through a competitive lens. Gemini 3’s benchmark lead over OpenAI’s then-current GPT-5.1 and Anthropic’s Claude models was large enough, and consistent enough across independent evaluations, that it shifted market and press narrative toward Google having retaken technical leadership in large language models — a position it had not clearly held since the early GPT-4 era. Reporting later in December said the release prompted an internal “code red” at OpenAI, with Sam Altman reportedly telling staff to prioritise ChatGPT’s core product over other initiatives; OpenAI shipped GPT-5.2 three weeks later citing gains on the same class of professional and coding benchmarks.

As with prior releases, the reported scores were self-published by the developer rather than independently audited, and benchmark performance did not resolve the separate, recurring argument about how much such gains translate into reliability on real-world tasks — a gap Gemini 3’s own release notes did not address.

Referenced by