Hardcore Reviews
Hardcore Reviews

One Million Output Tokens: Which Model Finishes the Job

A four-way comparison of coding/general models (October 5, 2026 basis; complements rather than repeats the site's late-September pricing review). The axis this time is single-response output length: Gemini 4 Argon stretched output from 64K to 1 million tokens (officially the industry's largest), changing the battlefield for long tasks. Four positions: the controlled flagship Argon (released September 30; officially first place on 13 of 18 benchmarks and a 77.9% DeepSWE v1.1 SOTA, yet in a controlled rollout the general public cannot buy; Artificial Analysis independently scores it slightly lower; introductory $2/$10 rising to $4/$20), the reigning efficiency benchmark Claude Sonnet 5.5 (September 28; Terminal-Bench 4.0 at 70.6%; $2/$10 on three clouds), the cache-economics pick GPT-6.1 Sol ($2/$0.10/$10 - cached input 95% cheaper), and the open-source reference GLM-5.3 (about 750GB of weights on ModelScope for self-hosting, domestic API $1.40/$4.40, open-source coding SOTA with CyberGym at 84.5%). Red lines honored: GLM's Terminal Bench 3.0 and Sonnet's Terminal-Bench 4.0 are different benchmark versions and are never compared directly; both official and independent assessments of Argon are presented. No benchmarks were re-run and no winner is crowned - conclusions follow the need: wait for Argon's wider access, pick Sonnet 5.5 for balanced value, Sol for cache-heavy agent pipelines, Sonnet for multi-cloud, GLM-5.3 for open self-hosting.

Published October 5, 202610 min read
<!-- million-token-output-comparison-review | review | One Million Output Tokens: Which Model Finishes the Job -->

One: A New Axis, From Unit Price to Output Ceiling

First, the division of labor: this site already ran a coding model price comparison at the end of September, lining up the input, output, and cache prices of Claude Sonnet 5.5 and GPT-6.1 Sol. This piece does not repeat that quote sheet. Instead it switches to a new axis: the single-response output ceiling. On September 30, 2026, Google shipped Gemini 4 Argon and pushed the output limit from the customary 64K tokens to one million tokens, which the company calls an industry first. For long-running tasks that changes the dimension of the battlefield: work that used to be fed across a dozen or more relay turns can now, in theory, be emitted in one pass, and the overhead of shuttling context back and forth gets rewritten along with it.

Across four slots, the controlled new flagship, the incumbent efficiency benchmark, the cache-efficiency pick, and the open-source control, this review covers four models: Gemini 4 Argon from Google, Claude Sonnet 5.5 from Anthropic, GPT-6.1 Sol from OpenAI, and GLM-5.3 as the open-source entry. The snapshot date is October 5, 2026.

The boundaries first: every comparison here is assembled from official announcements, official documentation, and media reports, not from hands-on benchmarking. With no benchmark runs of our own, this piece makes no arbitration about who is number one; it describes each model's range within its own slot and closes with conclusions tied to needs. All prices are snapshot figures and the official pricing pages remain the source of truth.

Two: The Comparison Table

DimensionGemini 4 ArgonClaude Sonnet 5.5GPT-6.1 SolGLM-5.3
AnnouncementSept 30 (ET)Sept 28 (ET)Established site coverageOpen weights already out
Output ceiling1M tokens (claimed industry first)No big number from the vendor, so none hereNot in scope of this piece, so none hereNot published
Price and cacheIntro $2/$10, cache input $0.10, later $4/$20 (official notice)$2/$10, cache read $0.20Input $2, output $10, cache $0.10Domestic API $1.40/$4.40
Benchmark framing13 of 18 official wins; independent numbers lagTerminal-Bench 4.0 at 70.6%DeepSWE v1.1 matches Astra at about one-fifth costOpen-source coding SOTA, CyberGym 84.5% overall best
AvailabilityControlled staged rollout, general users locked outLive, synced across three cloudsAvailableSelf-hostable, plus API
Open sourceNoNoNoYes, roughly 750GB in bf16
Best forEarly adopters waiting on API accessTeams wanting balanced value and multi-cloudAgent pipelines that save through cachingTeams needing self-hosting and a cheap API

The row worth staring at is availability: the model shouting the loudest number is precisely the one nobody can buy today. That sets the reading order for the next four sections. For Argon, watch the specs and the conflicting numbers. For Sonnet 5.5 and Sol, work the math that applies today. For GLM-5.3, look at the open-source bar to clear.

Three: Gemini 4 Argon, Loudest Number, Narrowest Door

Argon is the only one of the four that made its output ceiling the headline. According to the official announcement and multiple press reports on September 30, the single-response output limit rose from 64K tokens in the previous generation to one million, which Google calls an industry first. The official benchmark story is equally aggressive: first place on 13 of 18 benchmarks, with one additional tie; 77.9% on DeepSWE v1.1, a new SOTA; and 68% on CWE-bench v1, tied for first. Internal case studies hint at what million-token output enables: a 32,000-line SIMD rewrite of the libgav1 decoder running 2.7 times faster, an 800,000-line C/C++ to Rust conversion of the Fuchsia Zircon kernel, and a data center memory optimization freeing more than 300 TiB. Work at that scale is exactly what old output ceilings made unstable.

Two things must be stated plainly. First, availability: Argon ships through a controlled staged rollout, currently limited to the Fairwind program, a network of more than 650 cybersecurity partners, plus Google internal use. The next stage targets paying API customers and AI Ultra subscribers, with no date announced; ordinary users and developers are still locked out. Slotting Argon into a roadmap takes discipline: it belongs on a watchlist, not a schedule. Fairwind is aimed at cybersecurity partners, which suggests Google is handing high-risk scenarios to controlled validators first; any claim of an opening this month is speculation. Second, the numbers conflict: independent testing by Artificial Analysis shows Argon trailing its official framing, with the Agent Arena coding dimension behind Opus 4.8 and GPT-5.6-Sol. The official 13-of-18 lead and the independent lag coexist, and a careful buyer should weigh both.

On pricing, snapshot terms: an introductory $2/$10, with cache input at a 95 percent discount, that is $0.10 per million tokens, and an official notice that prices will later rise to $4/$20. In other words, even when access opens, the introductory price carries a window.

Four: Claude Sonnet 5.5, Not Chasing Big Numbers but Per-Task Efficiency

Sonnet 5.5 is the only model of the four that never made its output ceiling a selling point, and the vendor genuinely did not stress a big number, so this piece will not invent one. Its math runs differently: output speed is up more than 30 percent over Sonnet 5, and per-task token cost drops by as much as 30 percent on most work. The unit price stays flat; the total falls because fewer tokens get spent. On benchmarks, Terminal-Bench 4.0 lands at 70.6 percent against 10.3 percent for Sonnet 5. CursorBench 4.0 hits 55.5 percent, closing in on Opus 5.5 at 57.8 percent, and OSWorld 2.1 climbs from 57.0 to 80.1 percent.

The delivery story feels far more real than Argon's: it went live on September 28 across the Claude platform plus AWS Bedrock, Google Cloud Vertex, and Microsoft Azure on day one, was integrated into Claude Code the same day, and ships a zero data retention option. Pricing matches Sonnet 5 exactly: $2 input, $10 output, $0.20 cache read, snapshot terms. The release details are in our Sonnet 5.5 release coverage, and the access paths are in how to call Sonnet 5.5.

Anthropic positions it as the faster, cheaper complement to Opus 5.5 for well-scoped daily work: fixing bugs, drafting documents, slides, and spreadsheets. The benchmark detail matches: GDPval-AA sits only 2 points below Opus 5.5, and Chartography jumps from 15.6 to 61.6 percent. On safety, the Sonnet line carries network-security safeguards at the Opus tier for the first time, with high-risk cybersecurity requests able to fall back to Sonnet 5. The model ID is claude-sonnet-5-5 and the knowledge cutoff is June 2026, same as Opus 5.5. The vendor has also teased a higher-throughput, lower-cost Haiku 5.5 within weeks, with no specs published.

Five: GPT-6.1 Sol, Cache at $0.10 for Agent Pipelines

Sol's position was already established in our GPT-6.1 Sol launch piece: intelligence close to Astra, official pricing at $2 input, $10 output, and $0.10 cache, meaning cache costs 95 percent less than standard input. That figure carries the most weight in long-task scenarios, because an agent pipeline re-reads its context every turn; however large the output ceiling, the bill is often dominated by cache reads. On cost efficiency, the standing reference is DeepSWE v1.1 matching Astra at roughly one-fifth the cost, with Terminal-Bench Science at $5.47 per task.

One line of background: GPT-6.1 Astra was delayed for safety reasons, per Reuters on September 28, so the betable option on this product line today is Sol. The fuller Astra story is in our GPT-6 Astra review.

To spell out the cache math: a typical long task is a multi-turn agent loop that re-reads the system prompt, tool descriptions, and conversation history every turn. At $0.10, repeated reads cost one-twentieth of standard input, and the longer the run the bigger the saving. That is why the output ceiling and the cache column belong in one table: the ceiling decides how much one pass can hold, the cache decides what the relay costs, and together they form the real long-task bill.

Six: GLM-5.3, The Open-Source Control Row

The only open weights in this comparison belong to GLM-5.3, and its role here is the control: what the three closed models cannot give, data that never leaves your network and fine-tuning on your own terms, it does. Verified on ModelScope, the weights ship as 141 shards totaling roughly 750GB in bf16, so self-hosting demands serious GPU capacity. Skip the ops burden and the domestic API prices at $1.40/$4.40, the lowest input price of the four. Its capability is framed as open-source coding SOTA: highest open-source score on Terminal Bench 3.0, open-source SOTA on Agents' Last Exam, and 84.5% on CyberGym, the best overall. The deep dive is in our GLM-5.3 analysis.

The self-hosting bar deserves an honest count: 750GB in bf16 is the static size alone, and inference still needs headroom for activations and KV cache, which pushes tight setups toward multi-GPU or quantization. What you buy back is what no closed model offers: data inside your network, fine-tuning on your terms, and scheduling under your own budget.

One caution: benchmark versions and framings differ across vendors, so this piece performs no cross-version comparison; align on the same benchmark before ranking anything.

Seven: Conclusions Tied to Needs

With no benchmark runs of our own, this piece makes no arbitration about a single winner. The conclusions hang on what you need.

If you need long-output engineering that works today, land on Sonnet 5.5 or Sol. Argon's million tokens cannot be bought, and with no announced date, no roadmap should lean on it. The realistic options for finishing long tasks today are the two live models: Sonnet 5.5 leads on per-task efficiency and multi-cloud reach, while Sol leads on cache-cheap agent pipelines.

If you run agents and want the cache to pay, pick Sol. Cache at $0.10 pushes the cost of repeated context reads to the floor, and the DeepSWE result, matching Astra at about one-fifth the cost, is its hardest credential.

If you want balanced value with multi-cloud deployment, pick Sonnet 5.5. Same-day availability across three clouds is a genuine procurement convenience, the $2/$10 quote plus $0.20 cache reads is complete, and the speed-plus-savings framing makes it the most stable row today.

If you need open weights, self-hosting, and the lowest input price, pick GLM-5.3. Carrying 750GB of weights yourself buys data privacy and a $1.40 input price, and the open-source SOTA framing holds up.

If you want to try one-million-token output first, the only move is watching Argon's opening. The next stage covers paying API customers and AI Ultra subscribers, so prepare your budget and eval sets now and verify the official numbers the moment the door opens.

One method note to close: the four slots here are ordered by availability, not by strength. A model you can buy is the only one eligible for your schedule; the rest belong on a watchlist. That rule outranks any benchmark table.

FAQ

Q1: Can I use Gemini 4 Argon today? A: No. Per official framing, it ships through a controlled staged rollout limited to the Fairwind program, more than 650 cybersecurity partners, plus Google internal use. The next stage targets paying API customers and AI Ultra subscribers with no date announced; ordinary users and developers are still locked out.

Q2: Why are some output-ceiling cells empty for the closed models? A: Because the vendors never published or stressed them. The one million tokens for Argon is an official figure. Sonnet 5.5's vendor did not stress a big output number, and the figures for Sol and GLM-5.3 fall outside this piece's scope. Missing values are never invented here; they get filled once official numbers land.

Q3: Which of the four is cheapest on price? A: It depends on the dimension. The lowest input price is GLM-5.3 at $1.40. The lowest cache price is $0.10, shared by Argon's intro terms and Sol. Sonnet 5.5 reads cache at $0.20. Note that Argon's $2/$10 is introductory with an official notice of a later rise to $4/$20. All figures are snapshots and the official pages are the source of truth.

Q4: Which codes better, GLM-5.3 or Sonnet 5.5? A: They cannot be compared directly. GLM-5.3's score comes from Terminal Bench 3.0, while Sonnet 5.5's 70.6 percent comes from Terminal-Bench 4.0. Scores from different benchmark versions must not be put side by side. This piece only gives GLM-5.3 the qualitative open-source coding SOTA label; cross-version ranking should wait for same-benchmark data.

Q5: Are these prices final? A: No. Every price here is a snapshot from around October 5, 2026. Argon is still in its introductory window and the vendor has already announced a later increase. Always confirm against the official pricing pages before contracting.

Join the Discussion

Does your long-task workload run in relay turns or in one pass? If million-token output opened tomorrow, what would you throw at it first? Tell us your task shape and the context headaches you have hit in the comments; frequent questions will fold into future updates. If this helped, pass it to a colleague picking models right now.

This article is AI-assisted and human-edited. Last updated: 2026-10-05

FAQ

Can I use Gemini 4 Argon today?
A: No. Per official framing, it ships through a controlled staged rollout limited to the Fairwind program, more than 650 cybersecurity partners, plus Google internal use. The next stage targets paying API customers and AI Ultra subscribers with no date announced; ordinary users and developers are still locked out.
Why are some output-ceiling cells empty for the closed models?
A: Because the vendors never published or stressed them. The one million tokens for Argon is an official figure. Sonnet 5.5's vendor did not stress a big output number, and the figures for Sol and GLM-5.3 fall outside this piece's scope. Missing values are never invented here; they get filled once official numbers land.
Which of the four is cheapest on price?
A: It depends on the dimension. The lowest input price is GLM-5.3 at $1.40. The lowest cache price is $0.10, shared by Argon's intro terms and Sol. Sonnet 5.5 reads cache at $0.20. Note that Argon's $2/$10 is introductory with an official notice of a later rise to $4/$20. All figures are snapshots and the official pages are the source of truth.
Which codes better, GLM-5.3 or Sonnet 5.5?
A: They cannot be compared directly. GLM-5.3's score comes from Terminal Bench 3.0, while Sonnet 5.5's 70.6 percent comes from Terminal-Bench 4.0. Scores from different benchmark versions must not be put side by side. This piece only gives GLM-5.3 the qualitative open-source coding SOTA label; cross-version ranking should wait for same-benchmark data.
Are these prices final?
A: No. Every price here is a snapshot from around October 5, 2026. Argon is still in its introductory window and the vendor has already announced a later increase. Always confirm against the official pricing pages before contracting.

Related

Hardcore Reviews

Half Price vs One-Fifth: The Flagship Alternative Shake-Up

A value-focused comparison of flagship-adjacent coding models as of October 2026: Claude Sonnet 5.5 (official basis via machine-intelligence press: 30%+ faster, Terminal-Bench 4.0 up from 10.3% to 70.6%, priced at half of Opus 5.5) vs GPT-6.1 Sol (official basis: near-Astra intelligence at one-fifth the API price, with a paid Ultrafast speed tier) vs Claude Opus 5.5 (the reference point, $4/$20 per million tokens) vs DeepSeek V4.1 (the open-weights reference). Discipline: harnesses differ so cross-model scores cannot be compared - this piece only uses relative-to-own-flagship ratios, does not run its own evals, and does not compute Sonnet 5.5's unpublished dollar pricing. Four scenario verdicts: budget-conscious daily coding, flagship-ceiling complex work, and compliance-driven private deployment each have a winner. Complements the site's 2026-08 comprehensive and flagship-reasoning comparisons.

Oct 1, 202610 min read
Hardcore Reviews

30-Second Club: Kling 4.0, Seedance 2.5, Wan 3.0, LTX-2.5

A four-way video-model comparison (October 3, 2026 basis; complements rather than repeats the site's August five-way closed-cloud review and the Seedance 2.5 vs MiniMax H3 head-to-head). Framed by the observation that 30-second native single-pass output has become the entry ticket among flagship models, it picks one model per quadrant - cloud flagship, reigning benchmark, cloud value, open self-hosted: Kling 4.0 (October launch; 10-bit HDR, 2-minute extension, 15 references; specs announced but unverified), Seedance 2.5 (live since July 31; 30-second native output, 50 references, sub-second local editing; production deployments at XCMG, XPeng and Differential Intelligence Flight), Wan 3.0 (live since August 24; API from 0.3 CNY per second, five office document formats direct-to-video, no agent capabilities so orchestration is DIY), and LTX-2.5 (the only open weights; 66 GiB, 8-step distillation, ltx-trainer; custom community license with a USD 10 million revenue threshold). With no benchmark runs of its own, the piece refuses to crown a winner and concludes by need: Seedance for a working flagship today, Kling 4.0 to watch for pro specs, Wan 3.0 for value and batch work, LTX-2.5 for data privacy. Sora is excluded over conflicting sources and Veo unverified.

Oct 3, 202610 min read
Hardcore Reviews

Five Terminal Coding Agents: Which One Survives Your CI?

This review compares the terminal coding agent as a form factor rather than whose model is smarter: MiniMax Code CLI, Claude Code, Codex CLI, Qwen Code and Gemini CLI across seven dimensions (install, headless and CI, model freedom via BYOK, permissions and sandboxing, extension surface, open license, pricing model), with every repository number taken from a 2026-09-20 GitHub API snapshot. Key findings: model freedom is the widest gap, since only MiniMax and Qwen Code support BYOK to other vendors; the most substantial sandbox belongs to Codex CLI full-auto with the network disabled and a directory jail; Claude Code has the most mature extension surface but is closed and eats only its own model. It closes with a scenario ledger for personal daily use, unattended CI, enterprise compliance and model-swapping savings, plus three shared weaknesses: context readability, permission misjudgment and model lock-in.

Sep 20, 20268 min read