Between late August and early September 2026, four vendors pushed new coding and reasoning flagships into the same week: Claude Fable 5.1 on September 1, then Gemini 3.8 Flash, Qwen3.8-Max-0902, and Muse Spark 1.3 on September 2. OpenAI's GPT-5.6 Sol is the exception: shipped July 9 with a promotion from August 21, yet its pricing window and capability band overlap the same decision space. Five same-tier models crammed into one window signal one thing: the coding agent has become a knife-fight. The closer the launch window, the calmer the read has to be. This article does not replay launch numbers; it re-aligns unit price, context, and benchmarks on a comparable basis, and above all flags the easiest trap to fall into: vendors are not even reporting the same Terminal-Bench version.
1. A same-window launch: the coding-agent arms race
Lay the release dates out and the intent is clear. Claude Fable 5.1 led on September 1; a day later Google's Gemini 3.8 Flash, Alibaba's Qwen3.8-Max-0902, and Meta's Muse Spark 1.3 followed almost in lockstep. Four vendors putting coding and reasoning flagships on stage inside one calendar week is unusual even for 2026. OpenAI's GPT-5.6 Sol looks like the exception: it shipped July 9 and opened a promotion on August 21, but it belongs in the same tier because its pricing window and capability positioning sit squarely inside the decision band these newcomers are fighting over.
This window is not a point-capability contest; it is the whole coding-agent product line heating up. Context windows have been pushed to the million-token range almost uniformly, cache-read prices have been cut on purpose, and every benchmark targets SWE and terminal operation. The signal is that all five are chasing the same buyers: teams that drop a model into a long-loop, multi-tool, self-correcting agent pipeline and run real engineering tasks. For readers the takeaway is direct: the unit of evaluation is not a single model but a shortlist of candidates that can survive your specific agent loop. The full pricing and capability discussion is in the Claude Fable 5.1 launch read and the agentic cost comparison review.
2. Two pricing philosophies: cheap workhorses vs premium frontiers
The five price points split cleanly into two camps. The cheap-workhorse camp (Gemini 3.8 Flash, Muse Spark 1.3) competes on cost per task. The premium-frontier camp (Claude Fable 5.1, GPT-5.6 Sol) defends price on capability ceiling. Qwen3.8-Max-0902 sits in the middle at a blended roughly five dollars per million tokens.
| Model | Input $/M | Output $/M | Cache read $/M | Note |
|---|---|---|---|---|
| Gemini 3.8 Flash | 0.75 | 3.75 | not listed | intro price through 2026-12-31, then 1.50 / 7.50 |
| Qwen3.8-Max-0902 | 2.00 | 6.00 | 0.17 | blended ~$5/M, 262K reasoning budget |
| Muse Spark 1.3 | 1.25 | 4.25 | 0.15 | Contributor tier 0.10 / 0.20, Meta trains on your data |
| Claude Fable 5.1 | 10.00 | 50.00 | 0.25 | cache read cut 75% vs prior gen |
| GPT-5.6 Sol | 4.00 | 20.00 | not listed | promo from 5/30 through 2026-11-21; over 272K input billed 2x / 1.5x |
Several conditions hide behind the unit price. Gemini's 0.75 is an introductory rate that doubles at the end of 2026, so an annual budget cannot be locked at the current number. Fable's list price did not move, but cache read dropped from one dollar to 0.25 per million tokens, a 75 percent cut that is real money for multi-turn loops. Muse's Contributor tier is as cheap as 0.10 dollars input, but the cost is that Meta trains on your data, which cannot be treated as an ordinary tier on compliance grounds. GPT-5.6 Sol's promotion has a window and long context is billed at double, so long-document bills run well above the sticker.
3. The version trap: Terminal-Bench is not one benchmark
This is the single pitfall this article most wants to stop you falling into. Terminal-Bench jumped several versions in a year, and each vendor picked the version that suited its launch, so the reported version numbers do not line up:
- Gemini 3.8 Flash, Muse Spark 1.3, and GPT-5.6 Sol report Terminal-Bench 2.1, at 89.4 percent, 88.8 percent, and 88.8 percent.
- Qwen3.8-Max-0902 reports TerminalBench 3.0 at 29.0 percent, up from 11.3 percent, a different exam from 2.1.
- Claude Fable 5.1 reports Terminal-Bench 4.0 at 55.8 percent, yet another version.
Writing 29.0, 55.8, and 89.4 in one row misreads them as a capability ranking, a wrong conclusion. Different version numbers mean different task difficulty and scoring, so cross-version comparison is meaningless. The correct move is either separate tables per version or an explicit "not comparable" mark. Every table below is split strictly by version; any mismatched cell is marked "different version / not comparable" rather than forced into a column.
4. Benchmarks, split strictly by version
DeepSWE v1.1 (long-horizon SWE coding)
| Model | DeepSWE v1.1 |
|---|---|
| Gemini 3.8 Flash | 73.7% |
| Qwen3.8-Max-0902 | 69.3% |
| Muse Spark 1.3 | 75.4% (highest disclosed at launch) |
| Claude Fable 5.1 | not published |
| GPT-5.6 Sol | 72.7% |
Only Fable 5.1 gave no number here; Anthropic provided no SWE-bench-style figure for 5.1.
Terminal-Bench (split by version; do not compare across columns)
| Model | Terminal-Bench 2.1 | Terminal-Bench 3.0 | Terminal-Bench 4.0 |
|---|---|---|---|
| Gemini 3.8 Flash | 89.4% | different version | different version |
| Qwen3.8-Max-0902 | different version | 29.0% | different version |
| Muse Spark 1.3 | 88.8% (xhigh) | different version | different version |
| Claude Fable 5.1 | different version | different version | 55.8% |
| GPT-5.6 Sol | 88.8% | different version | 37.3% |
Note the Terminal-Bench 4.0 column compares only Fable and GPT (55.8 vs 37.3); Gemini's 4.0 mark of 19.1 is far harder than 2.1, proving versions cannot sit side by side.
SWE-bench Pro
| Model | SWE-bench Pro |
|---|---|
| Gemini 3.8 Flash | 61.6% |
| Qwen3.8-Max-0902 | different benchmark (reports QwenSWEbench V2 = 70.0) |
| Muse Spark 1.3 | not published |
| Claude Fable 5.1 | 81.2% |
| GPT-5.6 Sol | 64.6% |
Qwen reports its own QwenSWEbench V2 at 70.0, not SWE-bench Pro, so it cannot sit beside 61.6 and 81.2. Opus 5 scores roughly 96 on SWE-bench Verified, shown only as a magnitude hint.
HLE / HLE-Verified (reasoning)
| Model | HLE |
|---|---|
| Gemini 3.8 Flash | 54.9% |
| Qwen3.8-Max-0902 | not published |
| Muse Spark 1.3 | not published |
| Claude Fable 5.1 | 60.9% (no tools) |
| GPT-5.6 Sol | 54.5% |
Artificial Analysis Intelligence Index (composite)
| Model | Intelligence Index |
|---|---|
| Muse Spark 1.3 | 61 (xhigh) / 62 (max) |
| GPT-5.6 Sol | 59 - 60.9 |
| Claude Fable 5.1 | top tier (Opus 5 reference 63) |
| Gemini 3.8 Flash | not listed (coding near Opus 5) |
| Qwen3.8-Max-0902 | not listed |
All scores come from vendor publications or third parties (Artificial Analysis, Code Arena). "Not published" means the vendor did not supply it and this article ran no independent test; do not read it as zero.
5. The value leaders: Muse and Gemini
If one sentence survives, it is this: the lowest cost per task goes not to the strongest model but to whoever reads price and score together. Muse Spark 1.3 reaches an Intelligence Index of 61 (xhigh) at a vendor-estimated roughly 0.40 dollars per task, the tightest squeeze of "capable enough" and "cheap" in this tier. Gemini 3.8 Flash's coding scores (DeepSWE 73.7, Terminal-Bench 2.1 89.4) already approach the previous-generation Opus 5 band while priced at one seventh of the old frontier, a textbook case of buying near-frontier coding at budget rates.
But the capability ceiling still sits with Fable 5.1. It takes 81.2 on SWE-bench Pro and 52.6 on Terminal-Bench-Science 0.1, the steadiest in this group for agentic science and long-horizon software engineering, matching its premium 10/50 tag. Qwen3.8-Max-0902 lands in the middle, blended around five dollars with generous context and reasoning budget (991K max input, 262K reasoning budget), suited to teams needing long context without paying premium. Muse and Gemini detail is in the Gemini 3.8 Flash hotspot read.
For selection, a rough cut line: budget-sensitive, high-volume, not extreme on absolute ceiling, go Muse or Gemini first; agentic science, long-horizon SWE, willing to pay for quality, Fable 5.1 remains the ceiling; need million-token context but capped budget, look at Qwen; already deep in the OpenAI toolchain, GPT-5.6 Sol on promotion still works, watch the long-context double billing.
6. A rollout checklist
| Action | Basis | Priority |
|---|---|---|
| Re-read benchmarks by version, never across columns | Terminal-Bench 2.1/3.0/4.0 not comparable | High |
| Recompute the bill on your own workload | Gemini year-end hike, GPT long-context double, Muse data terms | High |
| Validate scores with your own evaluation | all vendor or third-party | High |
| Cheap loops first Muse / Gemini | lowest cost per task | Medium |
| Premium loops use Fable 5.1 | leads SWE-bench Pro and agentic science | Medium |
| Review Muse Contributor tier compliance | Meta trains on your data | Medium |
References
- Google official model blog (Gemini 3.8 Flash, 2026-09-02): launch date, pricing 0.75 / 3.75 (intro through 2026-12-31, then 1.50 / 7.50), context 1M / 64K, DeepSWE v1.1 73.7%, Terminal-Bench 2.1 89.4%, SWE-bench Pro 61.6%, HLE 54.9%, Terminal-Bench 4.0 19.1%. Vendor-published.
- Alibaba official release (Qwen3.8-Max-0902, 2026-09-02): pricing 2.00 / 6.00, cache read 0.17/M, blended ~$5, context 1M / 131K (991K max input, 262K reasoning budget), DeepSWE v1.1 69.3%, TerminalBench 3.0 29.0%, QwenSWEbench V2 70.0. Vendor-published.
- Meta official release (Muse Spark 1.3, 2026-09-02): pricing 1.25 / 4.25, cache hit 0.15/M, Contributor tier 0.10 / 0.20 (Meta trains on data), context 1M, DeepSWE v1.1 75.4%, Terminal-Bench 2.1 88.8% (xhigh), Intelligence Index 61 (xhigh) / 62 (max). Vendor-published.
- Anthropic official release (Claude Fable 5.1, 2026-09-01): pricing 10.00 / 50.00, cache read cut to 0.25/M (down 75%), context 1M / 128K, Terminal-Bench 4.0 55.8%, Terminal-Bench-Science 0.1 52.6%, SWE-bench Pro 81.2%, HLE 60.9% (no tools). Vendor-published.
- OpenAI official release (GPT-5.6 Sol, 2026-07-09; promo 2026-08-21 to 2026-11-21): pricing 4.00 / 20.00 (promo from 5/30), over 272K input billed 2x / 1.5x, context 1.1M / 128K, DeepSWE v1.1 72.7%, Terminal-Bench 2.1 88.8%, Terminal-Bench 4.0 37.3%, SWE-bench Pro 64.6%, HLE 54.5%, Intelligence Index 59-60.9. Vendor-published.
- Artificial Analysis Intelligence Index: Muse 61/62, GPT-5.6 Sol 59-60.9, Opus 5 reference 63, and other composite figures. Third-party.
- Code Arena leaderboard: coding leaderboard placement reference. Third-party.
- Tencent News, ITHome: Chinese coverage of the clustered Gemini / Qwen / Muse / Fable launches. Third-party media.
- DataCamp, AlphaSignal, Beam AI, byteiota: third-party write-ups of version-number conventions and benchmark interpretation. Third-party media.
Frequently Asked Questions
Q1: Why do Terminal-Bench versions differ and can't be compared directly? A1: Terminal-Bench jumped from 2.1 to 3.0 to 4.0 within a year, and each step changed the task set and difficulty. Gemini, Muse, and GPT report 2.1 (89.4 / 88.8 / 88.8), Qwen reports 3.0 (29.0), and Fable reports 4.0 (55.8). Different versions mean different exams, so placing 29.0, 55.8, and 89.4 in one row misreads them as a capability ranking. Split the tables by version and mark mismatched cells "different version."
Q2: Gemini looks cheapest per token - how does the real bill work? A2: Gemini's 0.75 / 3.75 is an intro rate that doubles at the end of 2026, so an annual budget cannot assume today's price. The real bill also stacks cache hits, output length, and turn count: an agent loop spends most tokens re-reading long context, so the cache-read price is what matters. Fable cut cache read to 0.25/M, GPT doubles billing past 272K input, and Muse's Contributor tier is cheap but trains on your data. Replace "sticker price times expected tokens" with "each tier's price times your real distribution"; see the agentic cost review.
Q3: Why does Fable 5.1 have no SWE-bench score? A3: Anthropic gave no public SWE-bench-style figure for Fable 5.1, so the DeepSWE v1.1 cell reads "not published." What it did disclose is Terminal-Bench 4.0 (55.8%), Terminal-Bench-Science 0.1 (52.6%), and SWE-bench Pro (81.2%). "Not published" means the vendor did not supply it and this article ran no independent test; it is not zero, so the cell should stay blank rather than be filled with a zero when comparing across models.
Q4: Among domestic/open models, which fits a coding agent best? A4: From this group, Qwen3.8-Max-0902 and Muse Spark 1.3 are the two candidates. Qwen has the largest context (991K input, 262K reasoning budget) at a blended roughly five dollars, suiting long-context tasks; Muse has the lowest cost per task (about 0.40 dollars) and an Intelligence Index of 61, suiting budget-sensitive high-volume agents. Neither published HLE or SWE-bench Pro, so their absolute ceiling is unknown and a self-built evaluation is advised before rollout.
Q5: Can these scores go straight into a tech-selection report? A5: Not as-is. Every score comes from vendor publications or third parties (Artificial Analysis, Code Arena), none reproduced independently here, and all carry version and methodology differences. In a report, cite only version-aligned rows with a comparison point, and label them "vendor figure / third-party, not independently verified"; mark version-mismatched cells "not comparable" and never place cross-version numbers side by side. The safer path is the self-built agent evaluation approach: build a set that reflects your own workload and treat published scores as a baseline, not a conclusion.