Hardcore Reviews
Hardcore Reviews

Flagship Coding & Reasoning Showdown: Five Models Compared

In late Aug–early Sep 2026, Gemini 3.8 Flash, Qwen3.8-Max-0902, Muse Spark 1.3, Claude Fable 5.1 and GPT-5.6 Sol shipped in a tight window — a "coding agent" arms race. The review splits pricing into two philosophies: cheap workhorses (Gemini $0.75, Muse $1.25) competing on cost-per-task, and premium frontiers (Fable $10/$50, GPT-5.6 Sol $4/$20). The biggest trap is benchmark version fragmentation — Qwen uses TerminalBench 3.0, Fable uses 4.0, the rest use 2.1, and they must never sit in one comparable column; every table here respects versions. Value leaders: Muse (Intelligence Index 61 at ~$0.40/task) and Gemini (near-Opus-5 coding at 1/7 unit price); Fable 5.1 owns agentic science (Terminal-Bench-Science 52.6%) and SWE-bench Pro (81.2%). Scores are vendor/third-party; unconfirmed items flagged.

Published September 1, 202610 min read
<!-- flagship-coding-reasoning-review | review | Flagship Coding & Reasoning Showdown: Five Models Compared -->

Between late August and early September 2026, four vendors pushed new coding and reasoning flagships into the same week: Claude Fable 5.1 on September 1, then Gemini 3.8 Flash, Qwen3.8-Max-0902, and Muse Spark 1.3 on September 2. OpenAI's GPT-5.6 Sol is the exception: shipped July 9 with a promotion from August 21, yet its pricing window and capability band overlap the same decision space. Five same-tier models crammed into one window signal one thing: the coding agent has become a knife-fight. The closer the launch window, the calmer the read has to be. This article does not replay launch numbers; it re-aligns unit price, context, and benchmarks on a comparable basis, and above all flags the easiest trap to fall into: vendors are not even reporting the same Terminal-Bench version.


1. A same-window launch: the coding-agent arms race

Lay the release dates out and the intent is clear. Claude Fable 5.1 led on September 1; a day later Google's Gemini 3.8 Flash, Alibaba's Qwen3.8-Max-0902, and Meta's Muse Spark 1.3 followed almost in lockstep. Four vendors putting coding and reasoning flagships on stage inside one calendar week is unusual even for 2026. OpenAI's GPT-5.6 Sol looks like the exception: it shipped July 9 and opened a promotion on August 21, but it belongs in the same tier because its pricing window and capability positioning sit squarely inside the decision band these newcomers are fighting over.

This window is not a point-capability contest; it is the whole coding-agent product line heating up. Context windows have been pushed to the million-token range almost uniformly, cache-read prices have been cut on purpose, and every benchmark targets SWE and terminal operation. The signal is that all five are chasing the same buyers: teams that drop a model into a long-loop, multi-tool, self-correcting agent pipeline and run real engineering tasks. For readers the takeaway is direct: the unit of evaluation is not a single model but a shortlist of candidates that can survive your specific agent loop. The full pricing and capability discussion is in the Claude Fable 5.1 launch read and the agentic cost comparison review.

2. Two pricing philosophies: cheap workhorses vs premium frontiers

The five price points split cleanly into two camps. The cheap-workhorse camp (Gemini 3.8 Flash, Muse Spark 1.3) competes on cost per task. The premium-frontier camp (Claude Fable 5.1, GPT-5.6 Sol) defends price on capability ceiling. Qwen3.8-Max-0902 sits in the middle at a blended roughly five dollars per million tokens.

ModelInput $/MOutput $/MCache read $/MNote
Gemini 3.8 Flash0.753.75not listedintro price through 2026-12-31, then 1.50 / 7.50
Qwen3.8-Max-09022.006.000.17blended ~$5/M, 262K reasoning budget
Muse Spark 1.31.254.250.15Contributor tier 0.10 / 0.20, Meta trains on your data
Claude Fable 5.110.0050.000.25cache read cut 75% vs prior gen
GPT-5.6 Sol4.0020.00not listedpromo from 5/30 through 2026-11-21; over 272K input billed 2x / 1.5x

Several conditions hide behind the unit price. Gemini's 0.75 is an introductory rate that doubles at the end of 2026, so an annual budget cannot be locked at the current number. Fable's list price did not move, but cache read dropped from one dollar to 0.25 per million tokens, a 75 percent cut that is real money for multi-turn loops. Muse's Contributor tier is as cheap as 0.10 dollars input, but the cost is that Meta trains on your data, which cannot be treated as an ordinary tier on compliance grounds. GPT-5.6 Sol's promotion has a window and long context is billed at double, so long-document bills run well above the sticker.

3. The version trap: Terminal-Bench is not one benchmark

This is the single pitfall this article most wants to stop you falling into. Terminal-Bench jumped several versions in a year, and each vendor picked the version that suited its launch, so the reported version numbers do not line up:

  • Gemini 3.8 Flash, Muse Spark 1.3, and GPT-5.6 Sol report Terminal-Bench 2.1, at 89.4 percent, 88.8 percent, and 88.8 percent.
  • Qwen3.8-Max-0902 reports TerminalBench 3.0 at 29.0 percent, up from 11.3 percent, a different exam from 2.1.
  • Claude Fable 5.1 reports Terminal-Bench 4.0 at 55.8 percent, yet another version.

Writing 29.0, 55.8, and 89.4 in one row misreads them as a capability ranking, a wrong conclusion. Different version numbers mean different task difficulty and scoring, so cross-version comparison is meaningless. The correct move is either separate tables per version or an explicit "not comparable" mark. Every table below is split strictly by version; any mismatched cell is marked "different version / not comparable" rather than forced into a column.

4. Benchmarks, split strictly by version

DeepSWE v1.1 (long-horizon SWE coding)

ModelDeepSWE v1.1
Gemini 3.8 Flash73.7%
Qwen3.8-Max-090269.3%
Muse Spark 1.375.4% (highest disclosed at launch)
Claude Fable 5.1not published
GPT-5.6 Sol72.7%

Only Fable 5.1 gave no number here; Anthropic provided no SWE-bench-style figure for 5.1.

Terminal-Bench (split by version; do not compare across columns)

ModelTerminal-Bench 2.1Terminal-Bench 3.0Terminal-Bench 4.0
Gemini 3.8 Flash89.4%different versiondifferent version
Qwen3.8-Max-0902different version29.0%different version
Muse Spark 1.388.8% (xhigh)different versiondifferent version
Claude Fable 5.1different versiondifferent version55.8%
GPT-5.6 Sol88.8%different version37.3%

Note the Terminal-Bench 4.0 column compares only Fable and GPT (55.8 vs 37.3); Gemini's 4.0 mark of 19.1 is far harder than 2.1, proving versions cannot sit side by side.

SWE-bench Pro

ModelSWE-bench Pro
Gemini 3.8 Flash61.6%
Qwen3.8-Max-0902different benchmark (reports QwenSWEbench V2 = 70.0)
Muse Spark 1.3not published
Claude Fable 5.181.2%
GPT-5.6 Sol64.6%

Qwen reports its own QwenSWEbench V2 at 70.0, not SWE-bench Pro, so it cannot sit beside 61.6 and 81.2. Opus 5 scores roughly 96 on SWE-bench Verified, shown only as a magnitude hint.

HLE / HLE-Verified (reasoning)

ModelHLE
Gemini 3.8 Flash54.9%
Qwen3.8-Max-0902not published
Muse Spark 1.3not published
Claude Fable 5.160.9% (no tools)
GPT-5.6 Sol54.5%

Artificial Analysis Intelligence Index (composite)

ModelIntelligence Index
Muse Spark 1.361 (xhigh) / 62 (max)
GPT-5.6 Sol59 - 60.9
Claude Fable 5.1top tier (Opus 5 reference 63)
Gemini 3.8 Flashnot listed (coding near Opus 5)
Qwen3.8-Max-0902not listed

All scores come from vendor publications or third parties (Artificial Analysis, Code Arena). "Not published" means the vendor did not supply it and this article ran no independent test; do not read it as zero.

5. The value leaders: Muse and Gemini

If one sentence survives, it is this: the lowest cost per task goes not to the strongest model but to whoever reads price and score together. Muse Spark 1.3 reaches an Intelligence Index of 61 (xhigh) at a vendor-estimated roughly 0.40 dollars per task, the tightest squeeze of "capable enough" and "cheap" in this tier. Gemini 3.8 Flash's coding scores (DeepSWE 73.7, Terminal-Bench 2.1 89.4) already approach the previous-generation Opus 5 band while priced at one seventh of the old frontier, a textbook case of buying near-frontier coding at budget rates.

But the capability ceiling still sits with Fable 5.1. It takes 81.2 on SWE-bench Pro and 52.6 on Terminal-Bench-Science 0.1, the steadiest in this group for agentic science and long-horizon software engineering, matching its premium 10/50 tag. Qwen3.8-Max-0902 lands in the middle, blended around five dollars with generous context and reasoning budget (991K max input, 262K reasoning budget), suited to teams needing long context without paying premium. Muse and Gemini detail is in the Gemini 3.8 Flash hotspot read.

For selection, a rough cut line: budget-sensitive, high-volume, not extreme on absolute ceiling, go Muse or Gemini first; agentic science, long-horizon SWE, willing to pay for quality, Fable 5.1 remains the ceiling; need million-token context but capped budget, look at Qwen; already deep in the OpenAI toolchain, GPT-5.6 Sol on promotion still works, watch the long-context double billing.

6. A rollout checklist

ActionBasisPriority
Re-read benchmarks by version, never across columnsTerminal-Bench 2.1/3.0/4.0 not comparableHigh
Recompute the bill on your own workloadGemini year-end hike, GPT long-context double, Muse data termsHigh
Validate scores with your own evaluationall vendor or third-partyHigh
Cheap loops first Muse / Geminilowest cost per taskMedium
Premium loops use Fable 5.1leads SWE-bench Pro and agentic scienceMedium
Review Muse Contributor tier complianceMeta trains on your dataMedium

References

  • Google official model blog (Gemini 3.8 Flash, 2026-09-02): launch date, pricing 0.75 / 3.75 (intro through 2026-12-31, then 1.50 / 7.50), context 1M / 64K, DeepSWE v1.1 73.7%, Terminal-Bench 2.1 89.4%, SWE-bench Pro 61.6%, HLE 54.9%, Terminal-Bench 4.0 19.1%. Vendor-published.
  • Alibaba official release (Qwen3.8-Max-0902, 2026-09-02): pricing 2.00 / 6.00, cache read 0.17/M, blended ~$5, context 1M / 131K (991K max input, 262K reasoning budget), DeepSWE v1.1 69.3%, TerminalBench 3.0 29.0%, QwenSWEbench V2 70.0. Vendor-published.
  • Meta official release (Muse Spark 1.3, 2026-09-02): pricing 1.25 / 4.25, cache hit 0.15/M, Contributor tier 0.10 / 0.20 (Meta trains on data), context 1M, DeepSWE v1.1 75.4%, Terminal-Bench 2.1 88.8% (xhigh), Intelligence Index 61 (xhigh) / 62 (max). Vendor-published.
  • Anthropic official release (Claude Fable 5.1, 2026-09-01): pricing 10.00 / 50.00, cache read cut to 0.25/M (down 75%), context 1M / 128K, Terminal-Bench 4.0 55.8%, Terminal-Bench-Science 0.1 52.6%, SWE-bench Pro 81.2%, HLE 60.9% (no tools). Vendor-published.
  • OpenAI official release (GPT-5.6 Sol, 2026-07-09; promo 2026-08-21 to 2026-11-21): pricing 4.00 / 20.00 (promo from 5/30), over 272K input billed 2x / 1.5x, context 1.1M / 128K, DeepSWE v1.1 72.7%, Terminal-Bench 2.1 88.8%, Terminal-Bench 4.0 37.3%, SWE-bench Pro 64.6%, HLE 54.5%, Intelligence Index 59-60.9. Vendor-published.
  • Artificial Analysis Intelligence Index: Muse 61/62, GPT-5.6 Sol 59-60.9, Opus 5 reference 63, and other composite figures. Third-party.
  • Code Arena leaderboard: coding leaderboard placement reference. Third-party.
  • Tencent News, ITHome: Chinese coverage of the clustered Gemini / Qwen / Muse / Fable launches. Third-party media.
  • DataCamp, AlphaSignal, Beam AI, byteiota: third-party write-ups of version-number conventions and benchmark interpretation. Third-party media.

Frequently Asked Questions

Q1: Why do Terminal-Bench versions differ and can't be compared directly? A1: Terminal-Bench jumped from 2.1 to 3.0 to 4.0 within a year, and each step changed the task set and difficulty. Gemini, Muse, and GPT report 2.1 (89.4 / 88.8 / 88.8), Qwen reports 3.0 (29.0), and Fable reports 4.0 (55.8). Different versions mean different exams, so placing 29.0, 55.8, and 89.4 in one row misreads them as a capability ranking. Split the tables by version and mark mismatched cells "different version."

Q2: Gemini looks cheapest per token - how does the real bill work? A2: Gemini's 0.75 / 3.75 is an intro rate that doubles at the end of 2026, so an annual budget cannot assume today's price. The real bill also stacks cache hits, output length, and turn count: an agent loop spends most tokens re-reading long context, so the cache-read price is what matters. Fable cut cache read to 0.25/M, GPT doubles billing past 272K input, and Muse's Contributor tier is cheap but trains on your data. Replace "sticker price times expected tokens" with "each tier's price times your real distribution"; see the agentic cost review.

Q3: Why does Fable 5.1 have no SWE-bench score? A3: Anthropic gave no public SWE-bench-style figure for Fable 5.1, so the DeepSWE v1.1 cell reads "not published." What it did disclose is Terminal-Bench 4.0 (55.8%), Terminal-Bench-Science 0.1 (52.6%), and SWE-bench Pro (81.2%). "Not published" means the vendor did not supply it and this article ran no independent test; it is not zero, so the cell should stay blank rather than be filled with a zero when comparing across models.

Q4: Among domestic/open models, which fits a coding agent best? A4: From this group, Qwen3.8-Max-0902 and Muse Spark 1.3 are the two candidates. Qwen has the largest context (991K input, 262K reasoning budget) at a blended roughly five dollars, suiting long-context tasks; Muse has the lowest cost per task (about 0.40 dollars) and an Intelligence Index of 61, suiting budget-sensitive high-volume agents. Neither published HLE or SWE-bench Pro, so their absolute ceiling is unknown and a self-built evaluation is advised before rollout.

Q5: Can these scores go straight into a tech-selection report? A5: Not as-is. Every score comes from vendor publications or third parties (Artificial Analysis, Code Arena), none reproduced independently here, and all carry version and methodology differences. In a report, cite only version-aligned rows with a comparison point, and label them "vendor figure / third-party, not independently verified"; mark version-mismatched cells "not comparable" and never place cross-version numbers side by side. The safer path is the self-built agent evaluation approach: build a set that reflects your own workload and treat published scores as a baseline, not a conclusion.

This article is AI-assisted and human-edited. Last updated: 2026-09-01

FAQ

Why do Terminal-Bench versions differ and can't be compared directly?
Terminal-Bench jumped from 2.1 to 3.0 to 4.0 within a year, and each step changed the task set and difficulty. Gemini, Muse, and GPT report 2.1 (89.4 / 88.8 / 88.8), Qwen reports 3.0 (29.0), and Fable reports 4.0 (55.8). Different versions mean different exams, so placing 29.0, 55.8, and 89.4 in one row misreads them as a capability ranking. Split the tables by version and mark mismatched cells "different version."
Gemini looks cheapest per token - how does the real bill work?
Gemini's 0.75 / 3.75 is an intro rate that doubles at the end of 2026, so an annual budget cannot assume today's price. The real bill also stacks cache hits, output length, and turn count: an agent loop spends most tokens re-reading long context, so the cache-read price is what matters. Fable cut cache read to 0.25/M, GPT doubles billing past 272K input, and Muse's Contributor tier is cheap but trains on your data. Replace "sticker price times expected tokens" with "each tier's price times your real distribution"; see the [agentic cost review](/en/posts/agentic-cache-cost-comparison-review).
Why does Fable 5.1 have no SWE-bench score?
Anthropic gave no public SWE-bench-style figure for Fable 5.1, so the DeepSWE v1.1 cell reads "not published." What it did disclose is Terminal-Bench 4.0 (55.8%), Terminal-Bench-Science 0.1 (52.6%), and SWE-bench Pro (81.2%). "Not published" means the vendor did not supply it and this article ran no independent test; it is not zero, so the cell should stay blank rather than be filled with a zero when comparing across models.
Among domestic/open models, which fits a coding agent best?
From this group, Qwen3.8-Max-0902 and Muse Spark 1.3 are the two candidates. Qwen has the largest context (991K input, 262K reasoning budget) at a blended roughly five dollars, suiting long-context tasks; Muse has the lowest cost per task (about 0.40 dollars) and an Intelligence Index of 61, suiting budget-sensitive high-volume agents. Neither published HLE or SWE-bench Pro, so their absolute ceiling is unknown and a self-built evaluation is advised before rollout.
Can these scores go straight into a tech-selection report?
Not as-is. Every score comes from vendor publications or third parties (Artificial Analysis, Code Arena), none reproduced independently here, and all carry version and methodology differences. In a report, cite only version-aligned rows with a comparison point, and label them "vendor figure / third-party, not independently verified"; mark version-mismatched cells "not comparable" and never place cross-version numbers side by side. The safer path is the [self-built agent evaluation](/en/posts/kimi-openplatform-codex-claude-code-sop) approach: build a set that reflects your own workload and treat published scores as a baseline, not a conclusion.

Related

Hardcore Reviews

After GPT-6 Astra: How the Top Flagships Really Compare

After GPT-6 Astra, a hardcore side-by-side of four same-tier flagships: Astra, Claude Fable 5.1, Gemini 3.8 Flash, and GPT-5.6 Sol. By dimension: on reasoning and math Astra is near-saturated (FrontierMath T4 97.6%, ARC-AGI-3 99.9%) while Sol scores just 7.8% on ARC-AGI-3; for coding agents Terminal-Bench must be read by version - Astra leads 4.0 at 57.9% while Gemini tops 2.1 at 89.4% but collapses to 19.1% on 4.0; on SWE-bench Pro Fable 5.1's 81.2% is highest; on computer use Astra leads OSWorld 2.0 at 72.6%; on ExploitBench Astra hits 100%; the alignment overreach gap is the starkest at 0% (Astra) vs 48% (Sol). On value, Gemini's \$0.75/\$3.75 discount window is lowest, while Astra and Fable both sit at \$10/\$50.

Sep 4, 20269 min read
Hardcore Reviews

CodeArena: Fable 5.1 Leads, Qwen Near at 1/8 Price

CodeArena, run by LMArena, is a frontend-coding leaderboard (end-to-end web-app generation, human-preference Elo). As of 2026-09-03: Claude Fable 5.1 leads at 1765 ($40/M), Qwen3.8-Max-0902 hit 1691 on day one and now ~1688 ($5/M, reaching the front rank at one-eighth the price), Gemini 3.8 Flash sits at 1567 (cheap variant, #18), Kimi K3 ~1674; GPT-6 Astra just launched 9/3 and its coding score is pending. Takeaway: Elo measures preference not accuracy — weigh price-performance and your own needs.

Sep 5, 20269 min read
Hardcore Reviews

Closed API vs Open Weights: What Does One Image Really Cost

With ChatGPT Images 2.5 and Ant's open-source LLaDA-Image landing in the same week, text-to-image has split into closed APIs versus self-hosted open weights. This review ignores image quality and runs the cost-and-control numbers instead: five routes - closed APIs, self-hosted open weights, per-second third-party inference platforms, local consumer hardware, and domestic cloud APIs - with per-image cost projected at two volumes (100 and 10,000 images per day), plus a comparison table and scenario-based selection (hobby use, e-commerce batch, data-sensitive industries, brand-style fine-tuning, maximum quality). It flags four traps: undeclared licenses, cold starts on per-second billing, Chinese text rendering, and cross-border data transfer. Explicitly scoped apart from our 8-26 capability review of reasoning image models. Representative comparison, not hands-on benchmarking; pricing per official sites.

Sep 9, 20269 min read