One: A New Axis, From Unit Price to Output Ceiling
First, the division of labor: this site already ran a coding model price comparison at the end of September, lining up the input, output, and cache prices of Claude Sonnet 5.5 and GPT-6.1 Sol. This piece does not repeat that quote sheet. Instead it switches to a new axis: the single-response output ceiling. On September 30, 2026, Google shipped Gemini 4 Argon and pushed the output limit from the customary 64K tokens to one million tokens, which the company calls an industry first. For long-running tasks that changes the dimension of the battlefield: work that used to be fed across a dozen or more relay turns can now, in theory, be emitted in one pass, and the overhead of shuttling context back and forth gets rewritten along with it.
Across four slots, the controlled new flagship, the incumbent efficiency benchmark, the cache-efficiency pick, and the open-source control, this review covers four models: Gemini 4 Argon from Google, Claude Sonnet 5.5 from Anthropic, GPT-6.1 Sol from OpenAI, and GLM-5.3 as the open-source entry. The snapshot date is October 5, 2026.
The boundaries first: every comparison here is assembled from official announcements, official documentation, and media reports, not from hands-on benchmarking. With no benchmark runs of our own, this piece makes no arbitration about who is number one; it describes each model's range within its own slot and closes with conclusions tied to needs. All prices are snapshot figures and the official pricing pages remain the source of truth.
Two: The Comparison Table
| Dimension | Gemini 4 Argon | Claude Sonnet 5.5 | GPT-6.1 Sol | GLM-5.3 |
|---|---|---|---|---|
| Announcement | Sept 30 (ET) | Sept 28 (ET) | Established site coverage | Open weights already out |
| Output ceiling | 1M tokens (claimed industry first) | No big number from the vendor, so none here | Not in scope of this piece, so none here | Not published |
| Price and cache | Intro $2/$10, cache input $0.10, later $4/$20 (official notice) | $2/$10, cache read $0.20 | Input $2, output $10, cache $0.10 | Domestic API $1.40/$4.40 |
| Benchmark framing | 13 of 18 official wins; independent numbers lag | Terminal-Bench 4.0 at 70.6% | DeepSWE v1.1 matches Astra at about one-fifth cost | Open-source coding SOTA, CyberGym 84.5% overall best |
| Availability | Controlled staged rollout, general users locked out | Live, synced across three clouds | Available | Self-hostable, plus API |
| Open source | No | No | No | Yes, roughly 750GB in bf16 |
| Best for | Early adopters waiting on API access | Teams wanting balanced value and multi-cloud | Agent pipelines that save through caching | Teams needing self-hosting and a cheap API |
The row worth staring at is availability: the model shouting the loudest number is precisely the one nobody can buy today. That sets the reading order for the next four sections. For Argon, watch the specs and the conflicting numbers. For Sonnet 5.5 and Sol, work the math that applies today. For GLM-5.3, look at the open-source bar to clear.
Three: Gemini 4 Argon, Loudest Number, Narrowest Door
Argon is the only one of the four that made its output ceiling the headline. According to the official announcement and multiple press reports on September 30, the single-response output limit rose from 64K tokens in the previous generation to one million, which Google calls an industry first. The official benchmark story is equally aggressive: first place on 13 of 18 benchmarks, with one additional tie; 77.9% on DeepSWE v1.1, a new SOTA; and 68% on CWE-bench v1, tied for first. Internal case studies hint at what million-token output enables: a 32,000-line SIMD rewrite of the libgav1 decoder running 2.7 times faster, an 800,000-line C/C++ to Rust conversion of the Fuchsia Zircon kernel, and a data center memory optimization freeing more than 300 TiB. Work at that scale is exactly what old output ceilings made unstable.
Two things must be stated plainly. First, availability: Argon ships through a controlled staged rollout, currently limited to the Fairwind program, a network of more than 650 cybersecurity partners, plus Google internal use. The next stage targets paying API customers and AI Ultra subscribers, with no date announced; ordinary users and developers are still locked out. Slotting Argon into a roadmap takes discipline: it belongs on a watchlist, not a schedule. Fairwind is aimed at cybersecurity partners, which suggests Google is handing high-risk scenarios to controlled validators first; any claim of an opening this month is speculation. Second, the numbers conflict: independent testing by Artificial Analysis shows Argon trailing its official framing, with the Agent Arena coding dimension behind Opus 4.8 and GPT-5.6-Sol. The official 13-of-18 lead and the independent lag coexist, and a careful buyer should weigh both.
On pricing, snapshot terms: an introductory $2/$10, with cache input at a 95 percent discount, that is $0.10 per million tokens, and an official notice that prices will later rise to $4/$20. In other words, even when access opens, the introductory price carries a window.
Four: Claude Sonnet 5.5, Not Chasing Big Numbers but Per-Task Efficiency
Sonnet 5.5 is the only model of the four that never made its output ceiling a selling point, and the vendor genuinely did not stress a big number, so this piece will not invent one. Its math runs differently: output speed is up more than 30 percent over Sonnet 5, and per-task token cost drops by as much as 30 percent on most work. The unit price stays flat; the total falls because fewer tokens get spent. On benchmarks, Terminal-Bench 4.0 lands at 70.6 percent against 10.3 percent for Sonnet 5. CursorBench 4.0 hits 55.5 percent, closing in on Opus 5.5 at 57.8 percent, and OSWorld 2.1 climbs from 57.0 to 80.1 percent.
The delivery story feels far more real than Argon's: it went live on September 28 across the Claude platform plus AWS Bedrock, Google Cloud Vertex, and Microsoft Azure on day one, was integrated into Claude Code the same day, and ships a zero data retention option. Pricing matches Sonnet 5 exactly: $2 input, $10 output, $0.20 cache read, snapshot terms. The release details are in our Sonnet 5.5 release coverage, and the access paths are in how to call Sonnet 5.5.
Anthropic positions it as the faster, cheaper complement to Opus 5.5 for well-scoped daily work: fixing bugs, drafting documents, slides, and spreadsheets. The benchmark detail matches: GDPval-AA sits only 2 points below Opus 5.5, and Chartography jumps from 15.6 to 61.6 percent. On safety, the Sonnet line carries network-security safeguards at the Opus tier for the first time, with high-risk cybersecurity requests able to fall back to Sonnet 5. The model ID is claude-sonnet-5-5 and the knowledge cutoff is June 2026, same as Opus 5.5. The vendor has also teased a higher-throughput, lower-cost Haiku 5.5 within weeks, with no specs published.
Five: GPT-6.1 Sol, Cache at $0.10 for Agent Pipelines
Sol's position was already established in our GPT-6.1 Sol launch piece: intelligence close to Astra, official pricing at $2 input, $10 output, and $0.10 cache, meaning cache costs 95 percent less than standard input. That figure carries the most weight in long-task scenarios, because an agent pipeline re-reads its context every turn; however large the output ceiling, the bill is often dominated by cache reads. On cost efficiency, the standing reference is DeepSWE v1.1 matching Astra at roughly one-fifth the cost, with Terminal-Bench Science at $5.47 per task.
One line of background: GPT-6.1 Astra was delayed for safety reasons, per Reuters on September 28, so the betable option on this product line today is Sol. The fuller Astra story is in our GPT-6 Astra review.
To spell out the cache math: a typical long task is a multi-turn agent loop that re-reads the system prompt, tool descriptions, and conversation history every turn. At $0.10, repeated reads cost one-twentieth of standard input, and the longer the run the bigger the saving. That is why the output ceiling and the cache column belong in one table: the ceiling decides how much one pass can hold, the cache decides what the relay costs, and together they form the real long-task bill.
Six: GLM-5.3, The Open-Source Control Row
The only open weights in this comparison belong to GLM-5.3, and its role here is the control: what the three closed models cannot give, data that never leaves your network and fine-tuning on your own terms, it does. Verified on ModelScope, the weights ship as 141 shards totaling roughly 750GB in bf16, so self-hosting demands serious GPU capacity. Skip the ops burden and the domestic API prices at $1.40/$4.40, the lowest input price of the four. Its capability is framed as open-source coding SOTA: highest open-source score on Terminal Bench 3.0, open-source SOTA on Agents' Last Exam, and 84.5% on CyberGym, the best overall. The deep dive is in our GLM-5.3 analysis.
The self-hosting bar deserves an honest count: 750GB in bf16 is the static size alone, and inference still needs headroom for activations and KV cache, which pushes tight setups toward multi-GPU or quantization. What you buy back is what no closed model offers: data inside your network, fine-tuning on your terms, and scheduling under your own budget.
One caution: benchmark versions and framings differ across vendors, so this piece performs no cross-version comparison; align on the same benchmark before ranking anything.
Seven: Conclusions Tied to Needs
With no benchmark runs of our own, this piece makes no arbitration about a single winner. The conclusions hang on what you need.
If you need long-output engineering that works today, land on Sonnet 5.5 or Sol. Argon's million tokens cannot be bought, and with no announced date, no roadmap should lean on it. The realistic options for finishing long tasks today are the two live models: Sonnet 5.5 leads on per-task efficiency and multi-cloud reach, while Sol leads on cache-cheap agent pipelines.
If you run agents and want the cache to pay, pick Sol. Cache at $0.10 pushes the cost of repeated context reads to the floor, and the DeepSWE result, matching Astra at about one-fifth the cost, is its hardest credential.
If you want balanced value with multi-cloud deployment, pick Sonnet 5.5. Same-day availability across three clouds is a genuine procurement convenience, the $2/$10 quote plus $0.20 cache reads is complete, and the speed-plus-savings framing makes it the most stable row today.
If you need open weights, self-hosting, and the lowest input price, pick GLM-5.3. Carrying 750GB of weights yourself buys data privacy and a $1.40 input price, and the open-source SOTA framing holds up.
If you want to try one-million-token output first, the only move is watching Argon's opening. The next stage covers paying API customers and AI Ultra subscribers, so prepare your budget and eval sets now and verify the official numbers the moment the door opens.
One method note to close: the four slots here are ordered by availability, not by strength. A model you can buy is the only one eligible for your schedule; the rest belong on a watchlist. That rule outranks any benchmark table.
FAQ
Q1: Can I use Gemini 4 Argon today? A: No. Per official framing, it ships through a controlled staged rollout limited to the Fairwind program, more than 650 cybersecurity partners, plus Google internal use. The next stage targets paying API customers and AI Ultra subscribers with no date announced; ordinary users and developers are still locked out.
Q2: Why are some output-ceiling cells empty for the closed models? A: Because the vendors never published or stressed them. The one million tokens for Argon is an official figure. Sonnet 5.5's vendor did not stress a big output number, and the figures for Sol and GLM-5.3 fall outside this piece's scope. Missing values are never invented here; they get filled once official numbers land.
Q3: Which of the four is cheapest on price? A: It depends on the dimension. The lowest input price is GLM-5.3 at $1.40. The lowest cache price is $0.10, shared by Argon's intro terms and Sol. Sonnet 5.5 reads cache at $0.20. Note that Argon's $2/$10 is introductory with an official notice of a later rise to $4/$20. All figures are snapshots and the official pages are the source of truth.
Q4: Which codes better, GLM-5.3 or Sonnet 5.5? A: They cannot be compared directly. GLM-5.3's score comes from Terminal Bench 3.0, while Sonnet 5.5's 70.6 percent comes from Terminal-Bench 4.0. Scores from different benchmark versions must not be put side by side. This piece only gives GLM-5.3 the qualitative open-source coding SOTA label; cross-version ranking should wait for same-benchmark data.
Q5: Are these prices final? A: No. Every price here is a snapshot from around October 5, 2026. Argon is still in its introductory window and the vendor has already announced a later increase. Always confirm against the official pricing pages before contracting.
Join the Discussion
Does your long-task workload run in relay turns or in one pass? If million-token output opened tomorrow, what would you throw at it first? Tell us your task shape and the context headaches you have hit in the comments; frequent questions will fold into future updates. If this helped, pass it to a colleague picking models right now.