On September 30, 2026, Google DeepMind published a blog post signed by Koray Kavukcuoglu, the company's Senior Vice President and Chief AI Architect, introducing Gemini 4 Argon. It is the first model in the Gemini 4 family, and Google is positioning it for a very specific kind of work: long-horizon, complex workflows such as real-world software engineering, enterprise knowledge work in legal and finance, and cybersecurity defense. The headline numbers are aggressive. The model can produce up to one million tokens in a single reply, roughly sixteen times the previous generation's cap. On DeepSWE v1.1, a long-horizon software engineering benchmark, Google reports a score of 77.9 percent, which it claims is the best result published so far. And inside Google, the model has already been used to migrate more than 800,000 lines of the Fuchsia Zircon kernel from C and C++ to Rust.
Those are striking claims. They are also, it must be said from the start, Google-reported figures. Every benchmark number in this article comes from DeepMind's own announcement, and no independent party has reproduced them yet. Just as important: most readers cannot buy Argon today. Google is running a controlled rollout through a partner program called Fairwind, and the general public and the paid API are still locked out, with no release date given. That combination, big claims plus limited access, is exactly why this release deserves a careful, honest read rather than a victory lap.
The one-million-token reply
Start with the output window, because it changes what "one reply" can mean. The previous generation of Gemini models capped output at 64,000 tokens, which was already generous by industry standards. Argon raises that ceiling to one million tokens, about sixteen times more. Google describes it as industry-leading, and on raw capacity it is hard to argue.
Why does output length matter for agentic work? When a model is refactoring a large codebase, the expensive failure mode is rarely a wrong answer on a single function. It is losing the thread halfway through a multi-thousand-file change, or forcing the orchestrating system to stitch together dozens of short replies, each one a chance to drift. A million-token reply lets the model hold an entire migration plan, the code it touches, and the test results in one continuous generation. We looked at the economics of very long single outputs in our million-token output comparison, and the conclusion there holds here: capacity alone does not guarantee quality, but insufficient capacity guarantees failure on tasks of this size.
The Google-reported benchmark table
DeepMind's announcement includes a benchmark table, and it is worth reading in full, including the rows where Argon loses. All figures below are Google-reported and have not been independently verified.
| Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|
| DeepSWE v1.1 (long-horizon software engineering) | 77.9 | 74.1 | 67.4 | 74.2 |
| AutomationBench (Zapier) | 51.3 | 41.4 | 31.4 | 42.5 |
| Vals Index | 68.9 | not listed | not listed | not listed |
| Vals Finance Agent v2 | 65.4 | not listed | not listed | not listed |
| Harvey legal agent | 19.6 | 5.4 | not listed | not listed |
| GraphWalks 256k to 1M, BFS F1 | 84.2 | 71.8 | not listed | not listed |
| Agent's Last Exam | 39.5 | not listed | not listed | not listed |
| CWE-bench v1 | 68 | tied first | not listed | not listed |
| LVBench | 91.7 | not listed | not listed | not listed |
Three rows stand out. DeepSWE v1.1 at 77.9 percent is the flagship claim, and the margin over GPT-6 Astra at 74.1 and Claude Opus 5.5 at 74.2 is real but modest, less than four points. AutomationBench, which measures whether an agent can operate Zapier-style automation tools, shows a much wider gap: Argon at 51.3 versus 42.5 for Opus 5.5 and 31.4 for Fable 5.1. And the Harvey legal agent result is fascinating precisely because both scores look low. Argon manages 19.6 against Astra's 5.4. Astra's number is single digits, which tells you how far the whole field still is from reliable legal work, even as Google advertises a relative win.
Note also what these numbers do and do not allow. Each row compares models on the same benchmark version, which is the only kind of comparison that means anything. A score on DeepSWE v1.1 cannot be set against a score on a different benchmark, and we have kept every comparison here within the same row for that reason.
Where Argon loses
The most credible thing in DeepMind's announcement is that the weakness rows are published at all. Google's own table shows Argon trailing on three benchmarks:
- FrontierSWE v2: Argon 55.0 versus GPT-6 Astra 65.5
- Terminal-bench 4.0: Argon 57.4 versus Claude Opus 5.5 66.4
- Terminal-Bench Science 0.1: Argon 57.6 versus GPT-6 Astra 68.1
An honesty line is warranted here, and Google has effectively written it for us: on benchmarks that probe harder, more recent software engineering tasks and terminal-driven scientific work, Argon is not the leader, and the gaps are double digits, not rounding errors. The pattern suggests a model tuned for very long, very structured agentic runs rather than for the sharpest single-shot performance on adversarial or fast-moving task suites. Terminal-driven suites tend to reward tight, short interaction loops with tools: many small commands, fast feedback, and little room to recover from an early mistake. Long-horizon agentic suites reward planning, consistency, and the ability to hold a coherent strategy across hours of work. A model optimized for the latter can look merely average on the former, and the Google-reported spread here is consistent with a deliberate design choice rather than a uniform capability deficit. If your workload lives in FrontierSWE v2 territory, the Google-reported numbers say Astra is currently the stronger pick. Our ongoing flagship coding model comparison keeps track of how these trade-offs shake out across the current generation.
Inside Google: the Zircon kernel and the 2.7x decoder
The benchmark table is abstract. The internal case studies are not, and they are the most vivid part of the announcement.
The flagship story is a migration of the Fuchsia Zircon kernel from C and C++ to Rust, with the largest single conversion covering more than 800,000 lines. To put that in perspective, 800,000 lines is a mid-sized operating system subsystem, the kind of code that engineering organizations staff with dedicated teams for years. Google says the converted code was deployed only after passing automated tests, simulation, and human review, which is the right reading of how this should work: the model does the grinding, humans and test suites keep the gate.
Why does a kernel migration make sense as a showcase? Zircon is the heart of Fuchsia, Google's independent operating system, and its C and C++ code carries decades of memory-management assumptions. Rust rewrites of C code usually fail for boring reasons: a pointer cast that was legal for twenty years, a lock held across a call the compiler never saw, a buffer whose size is documented only in a comment. Doing this at 800,000-line scale in machine-generated form is a claim about process as much as about the model. It also matches the million-token output window, because a conversion of this size cannot be chopped into a hundred independent prompts without losing global consistency, so the reply capacity and the migration story reinforce each other.
The second story is more technically interesting. Argon replaced about 32,000 lines of hand-tuned SIMD code in libgav1, Google's open-source AV1 video decoder, with a Rust implementation that runs 2.7 times faster than the original Rust port, with frame-identical output. The frame-identical part matters most. Video decoding is a domain where an off-by-one error produces subtle corruption rather than a crash, so bit-exact equivalence across frames is a strong correctness signal, not a marketing phrase.
There is more in the same register. Google credits Argon-driven work with freeing more than 300 TiB of memory in its datacenters, with an estimated total of 500 TiB to one PiB once fully rolled out. A quantum computing subroutine written with the model's help beat a published baseline by roughly 40 percent in minutes of work. And Google says Argon identified a sensitive information exposure vulnerability in medical software used by hospitals worldwide, one that previous frontier models had failed to find.
Treat these as Google-reported anecdotes rather than reproducible results, but the shape of the pattern is consistent: very long runs, verifiable outputs, and human or test-suite gates before deployment.
Pricing: an intro window worth watching
Google has announced pricing even though general access is not open yet. During the introductory period, Argon will cost 2 dollars per million input tokens and 10 dollars per million output tokens, with cached input priced at 95 percent off the input rate. After the introductory period ends, pricing moves to 4 dollars per million input and 20 dollars per million output.
Two observations. First, the introductory rate is aggressive for a flagship agentic model, and the cached-input discount is unusually deep; teams running repetitive agent loops against large context could see meaningful savings. Second, the post-intro price of 4 and 20 is where the real economics of the model will settle, and it roughly doubles the entry price. Anyone planning production workloads should model costs at the post-intro rate and treat the intro window as a discount on experimentation, not a baseline.
Fairwind: you cannot buy Argon yet
Here is the part that tempers everything above. Argon is not generally available. Google is running a controlled, staged rollout through the Fairwind Program, which counts more than 650 partners. Priority access goes to governments, critical infrastructure operators, and core platform teams. A notable subgroup, described by Google as trusted defenders in cybersecurity, receives a version of the model without cyber guardrails, presumably so that defensive research is not hobbled by safety filters. Google also says it is participating in a voluntary pre-release channel with the United States government.
For everyone else, including paid API customers, the door is closed, and Google has not given a date for opening it. That is an unusual posture for a consumer-facing company, and it reads as a deliberate strategy: gather real-world feedback from a curated set of sophisticated users before exposing the model broadly. The mechanics of a staged program like this matter more than they look. Partners in the priority tiers are the users most likely to wire Argon into pipelines with strong test coverage and audit trails, which means early failure modes surface in environments equipped to catch and report them. The unfiltered build for trusted defenders is a double-edged signal: it acknowledges that cyber guardrails would blunt defensive research, and it concentrates significant capability in a small group Google vets directly. The practical consequence for buyers is simple. Everything in this article, from the 77.9 to the 800,000 lines, is currently a promise backed by Google's own measurements, and the market cannot vote with its wallets until general availability. For context on how the competition is shipping, Claude's Sonnet 5.5 release took a more conventional public rollout path.
Safety claims, briefly
DeepMind's safety framing centers on agentic risk. Google claims Argon is its most resilient model against indirect prompt injection, the attack class where malicious instructions hide inside documents, web pages, or tool outputs that an agent consumes mid-task. The announced stack includes monitoring of the model's chain of thought and actions, with the ability to halt execution when monitoring flags a problem, plus hardened sandbox isolation around tool use. These are claims, not audit results, but prompt injection resilience is the right thing to be competing on for an agent model that reads untrusted content all day.
What this means for the rest of 2026
Three takeaways. First, the output window arms race is over for now: one million tokens out is the new flagship ceiling, and the interesting question shifts from capacity to what agents do with it. Second, the benchmark landscape is fragmenting into long-horizon agentic suites where Argon leads, Google-reportedly, and adversarial terminal suites where it trails, and buyers should match benchmarks to their actual workloads instead of chasing a single number. Third, the controlled rollout model may become the template for the most capable models, which means the public's first hands-on impression will arrive months after the announcement does.
One more competitive note: Argon's lead is not uncontested across the board, and open-weight models keep closing the gap on the same benchmarks. DeepSeek's V4.1 Flash open-source release, which we covered separately, is a reminder that the closed frontier no longer has the stage to itself.
Search Keywords
- when will gemini 4 argon be publicly available
- google argon api pricing
- gemini 4 argon deepswe 77.9 benchmark
- gemini 4 argon vs claude opus 5.5
This article was drafted with AI assistance and reviewed by a human editor.