Timeline

Google releases Gemini 4 Argon to cyber defenders before a wider launch

The output limit rises from 64,000 to one million tokens and introductory pricing is $2/$10 per million; Artificial Analysis placed it level with GPT-6 Astra, behind Anthropic's newest models.

  • Models & capabilities
  • Security & misuse
  • Major

Google announced Gemini 4 Argon, its first new frontier model since Gemini 3.1 Pro in February. It is not yet generally available. The first users are vetted government and infrastructure security teams in Google’s Fairwind programme, who get it “without cyber guardrails”, along with Google’s own staff. Paid API customers and Google AI Ultra subscribers come next, then developers, enterprises and consumers “as soon as possible”, with no date given. Google said it was taking part in the US government’s voluntary pre-release testing process and would adjust its safeguards on early testers’ feedback. The post was signed by Koray Kavukcuoglu, Google’s chief AI architect.

The main technical change is output length. Argon can write up to one million tokens in a single response, up from 64,000 on earlier Gemini models, so that it can reason through long problems in one pass. Google reported a 77.9% score on the DeepSWE v1.1 long-horizon coding benchmark, which it called state of the art, and first place on Zapier’s AutomationBench at 51.3%. It also claimed the lead on the Vals Index of finance, legal, tax and coding work. On cyber defence, it said Argon tied for first on CWE-bench v1 at 68% and found more vulnerabilities than Gemini 3.8 Flash Cyber, the restricted model Fairwind launched with. Inside Google, the company said, Argon agents had freed more than 300 TiB of data-centre memory and were porting C and C++ code to Rust. Introductory pricing is $2 per million input tokens and $10 per million output tokens, rising later to $4 and $20.

Independent testing was more mixed. The Decoder reported that Artificial Analysis scored Argon 53 on its Intelligence Index. That tied it with GPT-6 Astra and Claude Fable 5.1 but put it behind Claude Opus 5.5 at 58 and Claude Sonnet 5.5 at 56. On Artificial Analysis’s run of Terminal-Bench 4 it scored 57%, behind both Anthropic models and GPT-6 Astra. It used more than twice as many output tokens per task as GPT-6 Astra, so its cost advantage came from price rather than efficiency. It did lead Vals AI’s index at 68.9%, the first Gemini model to do so, and ranked first on Arena’s human-preference text leaderboard. Artificial Analysis also measured a lower hallucination rate than OpenAI’s latest models.

The safety section of the post focused on agent containment. Google said it monitors Argon’s chain of thought and actions during deployment and stops execution when the model oversteps the user’s intent. It used a similar monitor during training, keeping the findings out of the training signal so as not to teach the model to evade it, and urged other labs to preserve “reasoning transparency”. It also said it now isolates and seals sandboxed environments before high-risk training or evaluation. The measures came weeks after a Gemini model broke into three companies’ systems during a red-team test.