Gemini 4 Argon
modelYour notes
The first model of Google's Gemini 4 generation, announced September 30, 2026 by Koray Kavukcuoglu and released first to a set of trusted cyber defenders through the Fairwind Program rather than to the public. Google is running it through the US government's voluntary process for pre-release model access while it widens availability, with developers, enterprises and consumers to follow, starting with paid API customers and Google AI Ultra subscribers. The headline capability change is room to think: the output limit rises to 1M tokens from 64K, which Google frames as letting the model generate hundreds of thousands of tokens in a single trajectory instead of breaking a long job into pieces. Artificial Analysis measures a 1M-token context window with text and image input, and scores it 53 on Intelligence Index v4.3.2 — by far Google's best (Gemini 3.8 Flash reads 41), level with GPT-6 Astra and Claude Fable 5.1, and behind Claude Opus 5.5 (58) and Sonnet 5.5 (56). Introductory pricing is $2 per million input tokens and $10 per million output, with cached input 95% off, rising to $4 and $20 once the introductory period ends. Architecture, parameter count and training scale are undisclosed.
Google's evaluation sheet (highest thinking setting, pass@1, against GPT-6 Astra / Fable 5.1 / Opus 5.5) puts Argon first on knowledge work and much of coding: Vals Index 68.9% (63.1 / 65.8 / 67.0), Zapier's AutomationBench 51.3% (41.4 / 31.4 / 42.5), Vals Finance Agent v2 65.4%, Harvey's Legal Agent Benchmark 19.6% against 5.4 / 6.7 / 3.8, DeepSWE v1.1 77.9% (74.1 / 67.4 / 74.2) and Vibe Code Bench 91.9%. It also leads on LABBench 2 (88.8%), RiemannBench (76.0%), Chartography (71.6%), long video on LVBench (91.7%), Agent's Last Exam (39.5%) and both GraphWalks long-context splits, where the 256K-to-1M subset reaches 84.2% against 71.8 for Astra. It trails on FrontierSWE v2 (55.0% to Astra's 65.5%), Terminal-Bench 4.0 (57.4% to Opus 5.5's 66.4%), Terminal-Bench Science 0.1 (57.6% to Astra's 68.1%), PostTrainBench and OSWorld-2.0. On CWE-bench v1 it ties Astra for first at 68.0%.
Cybersecurity is the reason for the staged rollout. Google trained Argon to find, validate and patch vulnerabilities on its own, and is giving trusted defenders and its internal teams a build without cyber guardrails. Wiz is already running it in its Scan for Good program, where it found a critical flaw exposing personal information in healthcare software used by hospitals worldwide that earlier frontier models had missed. Google says it is hardening four areas before a broad release: misuse defenses for cyber and CBRN under its Frontier Safety Framework, including monitoring the model's internal activations; resistance to indirect prompt injection, where it leads Gray Swan's IPI benchmark; monitors on the chain-of-thought and actions that halt execution on misalignment, with findings deliberately kept out of training so the model is not shaped to evade them; and sealed sandboxes for high-risk training and evaluation. Internally, thousands of Googlers already use it: it cut the spacetime cost of a quantum subroutine 40% below the published baseline, freed more than 300 TiB of data-center memory through agent-driven profiling (with 500 TiB to 1 PiB projected), and is migrating C and C++ codebases to Rust at scales from the re2 and libgav1 libraries up to the 800K-line Fuchsia Zircon kernel. In libgav1 it replaced 32K lines of hand-written SIMD with safe Rust the compiler auto-vectorizes, running 2.7× faster than the existing Rust port with identical output.
Model Details
Benchmark Scores
| Benchmark | Score | Mode |
|---|---|---|
| Vals Index | 68.9% | — |
| AutomationBench | 51.3% | — |
| DeepSWE v1.1 | 77.9% | — |
| Vibe Code Bench | 91.9% | — |
| LVBench | 91.7% | — |
| GraphWalks (256k-1M, BFS F1) | 84.2% | — |
| CWE-bench v1 | 68.0% | — |
| Terminal-Bench 4.0 | 57.4% | — |