Frontline Hotspot
Frontline Hotspot

Gemini 4 Argon: 800K lines in one reply, 77.9% on DeepSWE

Google DeepMind launched Gemini 4 Argon on September 30, 2026 (official basis): the first Gemini 4 flagship, built for long-horizon work across real-world software engineering, legal and finance knowledge work, and cyber defense. The headline change: output limit stretched from 64K to 1 million tokens (about 16x). Google-reported benchmarks put DeepSWE v1.1 at 77.9% for a new SOTA (vs GPT-6 Astra 74.1, Claude Opus 5.5 74.2), AutomationBench at 51.3% for first place, CWE-bench v1 at 68% tied-first, and GraphWalks 256k-1M at 84.2%. Weaknesses reported faithfully: FrontierSWE v2 55.0 trails Astra's 65.5 and Terminal-bench 4.0 57.4 trails Opus 5.5's 66.4 - this piece attributes the split to task shape (long-horizon wins, terminal step-by-step loses), noting Google offers no explanation. Internal cases: the 800K+ line C/C++-to-Rust migration of the Fuchsia Zircon kernel; libgav1 rewritten at 32K lines of SIMD code running 2.7x faster with frame-identical output; datacenter memory work freeing 300+ TiB. Pricing: introductory $2/$10 (cache 95% off), then $4/$20. Controlled rollout reported as-is: Fairwind Program first with 650+ partners (trusted defenders get an unguarded build), general public and paid API still locked out with no date. All scores are Google-reported.

Published October 7, 20269 min read
<!-- gemini-4-argon-release-hotspot | hotspot | Gemini 4 Argon: 800K lines in one reply, 77.9% on DeepSWE -->

On September 30, 2026, Google DeepMind published a blog post signed by Koray Kavukcuoglu, the company's Senior Vice President and Chief AI Architect, introducing Gemini 4 Argon. It is the first model in the Gemini 4 family, and Google is positioning it for a very specific kind of work: long-horizon, complex workflows such as real-world software engineering, enterprise knowledge work in legal and finance, and cybersecurity defense. The headline numbers are aggressive. The model can produce up to one million tokens in a single reply, roughly sixteen times the previous generation's cap. On DeepSWE v1.1, a long-horizon software engineering benchmark, Google reports a score of 77.9 percent, which it claims is the best result published so far. And inside Google, the model has already been used to migrate more than 800,000 lines of the Fuchsia Zircon kernel from C and C++ to Rust.

Those are striking claims. They are also, it must be said from the start, Google-reported figures. Every benchmark number in this article comes from DeepMind's own announcement, and no independent party has reproduced them yet. Just as important: most readers cannot buy Argon today. Google is running a controlled rollout through a partner program called Fairwind, and the general public and the paid API are still locked out, with no release date given. That combination, big claims plus limited access, is exactly why this release deserves a careful, honest read rather than a victory lap.

The one-million-token reply

Start with the output window, because it changes what "one reply" can mean. The previous generation of Gemini models capped output at 64,000 tokens, which was already generous by industry standards. Argon raises that ceiling to one million tokens, about sixteen times more. Google describes it as industry-leading, and on raw capacity it is hard to argue.

Why does output length matter for agentic work? When a model is refactoring a large codebase, the expensive failure mode is rarely a wrong answer on a single function. It is losing the thread halfway through a multi-thousand-file change, or forcing the orchestrating system to stitch together dozens of short replies, each one a chance to drift. A million-token reply lets the model hold an entire migration plan, the code it touches, and the test results in one continuous generation. We looked at the economics of very long single outputs in our million-token output comparison, and the conclusion there holds here: capacity alone does not guarantee quality, but insufficient capacity guarantees failure on tasks of this size.

The Google-reported benchmark table

DeepMind's announcement includes a benchmark table, and it is worth reading in full, including the rows where Argon loses. All figures below are Google-reported and have not been independently verified.

BenchmarkGemini 4 ArgonGPT-6 AstraClaude Fable 5.1Claude Opus 5.5
DeepSWE v1.1 (long-horizon software engineering)77.974.167.474.2
AutomationBench (Zapier)51.341.431.442.5
Vals Index68.9not listednot listednot listed
Vals Finance Agent v265.4not listednot listednot listed
Harvey legal agent19.65.4not listednot listed
GraphWalks 256k to 1M, BFS F184.271.8not listednot listed
Agent's Last Exam39.5not listednot listednot listed
CWE-bench v168tied firstnot listednot listed
LVBench91.7not listednot listednot listed

Three rows stand out. DeepSWE v1.1 at 77.9 percent is the flagship claim, and the margin over GPT-6 Astra at 74.1 and Claude Opus 5.5 at 74.2 is real but modest, less than four points. AutomationBench, which measures whether an agent can operate Zapier-style automation tools, shows a much wider gap: Argon at 51.3 versus 42.5 for Opus 5.5 and 31.4 for Fable 5.1. And the Harvey legal agent result is fascinating precisely because both scores look low. Argon manages 19.6 against Astra's 5.4. Astra's number is single digits, which tells you how far the whole field still is from reliable legal work, even as Google advertises a relative win.

Note also what these numbers do and do not allow. Each row compares models on the same benchmark version, which is the only kind of comparison that means anything. A score on DeepSWE v1.1 cannot be set against a score on a different benchmark, and we have kept every comparison here within the same row for that reason.

Where Argon loses

The most credible thing in DeepMind's announcement is that the weakness rows are published at all. Google's own table shows Argon trailing on three benchmarks:

  • FrontierSWE v2: Argon 55.0 versus GPT-6 Astra 65.5
  • Terminal-bench 4.0: Argon 57.4 versus Claude Opus 5.5 66.4
  • Terminal-Bench Science 0.1: Argon 57.6 versus GPT-6 Astra 68.1

An honesty line is warranted here, and Google has effectively written it for us: on benchmarks that probe harder, more recent software engineering tasks and terminal-driven scientific work, Argon is not the leader, and the gaps are double digits, not rounding errors. The pattern suggests a model tuned for very long, very structured agentic runs rather than for the sharpest single-shot performance on adversarial or fast-moving task suites. Terminal-driven suites tend to reward tight, short interaction loops with tools: many small commands, fast feedback, and little room to recover from an early mistake. Long-horizon agentic suites reward planning, consistency, and the ability to hold a coherent strategy across hours of work. A model optimized for the latter can look merely average on the former, and the Google-reported spread here is consistent with a deliberate design choice rather than a uniform capability deficit. If your workload lives in FrontierSWE v2 territory, the Google-reported numbers say Astra is currently the stronger pick. Our ongoing flagship coding model comparison keeps track of how these trade-offs shake out across the current generation.

Inside Google: the Zircon kernel and the 2.7x decoder

The benchmark table is abstract. The internal case studies are not, and they are the most vivid part of the announcement.

The flagship story is a migration of the Fuchsia Zircon kernel from C and C++ to Rust, with the largest single conversion covering more than 800,000 lines. To put that in perspective, 800,000 lines is a mid-sized operating system subsystem, the kind of code that engineering organizations staff with dedicated teams for years. Google says the converted code was deployed only after passing automated tests, simulation, and human review, which is the right reading of how this should work: the model does the grinding, humans and test suites keep the gate.

Why does a kernel migration make sense as a showcase? Zircon is the heart of Fuchsia, Google's independent operating system, and its C and C++ code carries decades of memory-management assumptions. Rust rewrites of C code usually fail for boring reasons: a pointer cast that was legal for twenty years, a lock held across a call the compiler never saw, a buffer whose size is documented only in a comment. Doing this at 800,000-line scale in machine-generated form is a claim about process as much as about the model. It also matches the million-token output window, because a conversion of this size cannot be chopped into a hundred independent prompts without losing global consistency, so the reply capacity and the migration story reinforce each other.

The second story is more technically interesting. Argon replaced about 32,000 lines of hand-tuned SIMD code in libgav1, Google's open-source AV1 video decoder, with a Rust implementation that runs 2.7 times faster than the original Rust port, with frame-identical output. The frame-identical part matters most. Video decoding is a domain where an off-by-one error produces subtle corruption rather than a crash, so bit-exact equivalence across frames is a strong correctness signal, not a marketing phrase.

There is more in the same register. Google credits Argon-driven work with freeing more than 300 TiB of memory in its datacenters, with an estimated total of 500 TiB to one PiB once fully rolled out. A quantum computing subroutine written with the model's help beat a published baseline by roughly 40 percent in minutes of work. And Google says Argon identified a sensitive information exposure vulnerability in medical software used by hospitals worldwide, one that previous frontier models had failed to find.

Treat these as Google-reported anecdotes rather than reproducible results, but the shape of the pattern is consistent: very long runs, verifiable outputs, and human or test-suite gates before deployment.

Pricing: an intro window worth watching

Google has announced pricing even though general access is not open yet. During the introductory period, Argon will cost 2 dollars per million input tokens and 10 dollars per million output tokens, with cached input priced at 95 percent off the input rate. After the introductory period ends, pricing moves to 4 dollars per million input and 20 dollars per million output.

Two observations. First, the introductory rate is aggressive for a flagship agentic model, and the cached-input discount is unusually deep; teams running repetitive agent loops against large context could see meaningful savings. Second, the post-intro price of 4 and 20 is where the real economics of the model will settle, and it roughly doubles the entry price. Anyone planning production workloads should model costs at the post-intro rate and treat the intro window as a discount on experimentation, not a baseline.

Fairwind: you cannot buy Argon yet

Here is the part that tempers everything above. Argon is not generally available. Google is running a controlled, staged rollout through the Fairwind Program, which counts more than 650 partners. Priority access goes to governments, critical infrastructure operators, and core platform teams. A notable subgroup, described by Google as trusted defenders in cybersecurity, receives a version of the model without cyber guardrails, presumably so that defensive research is not hobbled by safety filters. Google also says it is participating in a voluntary pre-release channel with the United States government.

For everyone else, including paid API customers, the door is closed, and Google has not given a date for opening it. That is an unusual posture for a consumer-facing company, and it reads as a deliberate strategy: gather real-world feedback from a curated set of sophisticated users before exposing the model broadly. The mechanics of a staged program like this matter more than they look. Partners in the priority tiers are the users most likely to wire Argon into pipelines with strong test coverage and audit trails, which means early failure modes surface in environments equipped to catch and report them. The unfiltered build for trusted defenders is a double-edged signal: it acknowledges that cyber guardrails would blunt defensive research, and it concentrates significant capability in a small group Google vets directly. The practical consequence for buyers is simple. Everything in this article, from the 77.9 to the 800,000 lines, is currently a promise backed by Google's own measurements, and the market cannot vote with its wallets until general availability. For context on how the competition is shipping, Claude's Sonnet 5.5 release took a more conventional public rollout path.

Safety claims, briefly

DeepMind's safety framing centers on agentic risk. Google claims Argon is its most resilient model against indirect prompt injection, the attack class where malicious instructions hide inside documents, web pages, or tool outputs that an agent consumes mid-task. The announced stack includes monitoring of the model's chain of thought and actions, with the ability to halt execution when monitoring flags a problem, plus hardened sandbox isolation around tool use. These are claims, not audit results, but prompt injection resilience is the right thing to be competing on for an agent model that reads untrusted content all day.

What this means for the rest of 2026

Three takeaways. First, the output window arms race is over for now: one million tokens out is the new flagship ceiling, and the interesting question shifts from capacity to what agents do with it. Second, the benchmark landscape is fragmenting into long-horizon agentic suites where Argon leads, Google-reportedly, and adversarial terminal suites where it trails, and buyers should match benchmarks to their actual workloads instead of chasing a single number. Third, the controlled rollout model may become the template for the most capable models, which means the public's first hands-on impression will arrive months after the announcement does.

One more competitive note: Argon's lead is not uncontested across the board, and open-weight models keep closing the gap on the same benchmarks. DeepSeek's V4.1 Flash open-source release, which we covered separately, is a reminder that the closed frontier no longer has the stage to itself.

Search Keywords

  • when will gemini 4 argon be publicly available
  • google argon api pricing
  • gemini 4 argon deepswe 77.9 benchmark
  • gemini 4 argon vs claude opus 5.5

This article was drafted with AI assistance and reviewed by a human editor.

This article is AI-assisted and human-edited. Last updated: 2026-10-07

Related

Frontline Hotspot

Claude Sonnet 5.5 Ships: Coding Score Leaps From 10 to 70

Claude Sonnet 5.5 shipped September 28, 2026 (US Eastern, official basis): the second model in the Claude 5.5 family, positioned as a faster, cheaper complement to Opus 5.5, live day one on the Claude platform plus AWS Bedrock, Google Cloud Vertex and Microsoft Azure, with Claude Code integrated the same day. Headline numbers: Terminal-Bench 4.0 jumps from Sonnet 5's 10.3% to 70.6% (same benchmark, different generation - nearly sevenfold); CursorBench 4.0 at 55.5% (Opus 5.5: 57.8%); OSWorld 2.1 from 57.0% to 80.1%; GDPval-AA within 2 points of Opus 5.5; output speed up over 30% and per-task cost down up to 30% on most work. Pricing strategy: unit prices unchanged versus Sonnet 5 ($2/$10, cache read $0.20) - the discount hides in token efficiency. Security: the Sonnet line gets Opus/Fable-grade cyber protections for the first time, plus an anti-extraction classifier and auto-fallback for high-risk cyber requests. Also the first Sonnet to beat Pokemon Red from screenshots alone. Competitive context: Gemini 4 Argon landed two days later (controlled release, 1M output tokens) and GPT-6.1 Astra was delayed over safety; enterprise customers are about 80% of Anthropic's business ahead of a planned IPO (Reuters). Haiku 5.5 is teased for the coming weeks - no specs published, none invented here.

Oct 5, 20269 min read
Frontline Hotspot

Gemini 3.8 Drops: Flash and the Security-First Flash Cyber

Google released Gemini 3.8 on 2026-09-02 (US) / 09-03 (Beijing) as two models: the general Flash for long-horizon engineering and agents, and the security-focused Flash Cyber for autonomous vulnerability discovery and automated patching, available only to defenders via the Fairwind Program. Official numbers: HLE-Verified 54.9%, CWE-Bench pass@1 47.2%, cross-language vuln discovery >70%, 2.6x Chrome patches, critical vulns found in <2 hours; intro pricing \$0.75/\$3.75 per million tokens.

Sep 5, 20269 min read
Frontline Hotspot

Kling 4.0 Lands in October: 30-Second 4K HDR vs Seedance

Kling 4.0 is official (September 28, 2026, per official announcements and media reports): full launch in October, with the lightweight 4.0 Flash opened the same day to black-gold annual members in a limited preview. Core specs: 4K and 1080p in 10-bit HDR, up to 30 seconds per generation, multi-round extension to 2 minutes, an 8,000-token prompt ceiling, up to 15 multimodal references (10 images + 5 videos + 7 subjects), and lip-synced multilingual output across Chinese, English, Japanese, Korean, Spanish, Portuguese, German, French and Hindi; the creation page gains timeline-based continuous creation and a canvas Agent experience. Competitive context: a direct answer to Seedance 2.5 (launched July 31 to strong reception); Citi keeps Kuaishou at Neutral (target 41 HKD) with two questions - can Kling 4.0 pull Seedance users over, and is commercial pricing competitive - against a July funding round of about USD 3 billion at a USD 18 billion post-money valuation. Pricing and the exact launch date are unpublished; this piece withholds verdicts until real output lands.

Oct 3, 20269 min read