Hardcore Reviews
Hardcore Reviews

Small Model Price War: Haiku 5.5 Ties Luna Only Below 100K

A small-model price-war comparison (October 8, 2026 price snapshots; five columns, four rows): Claude Haiku 5.5, GPT-6 Luna, Sonnet 5.5, and Qwen3.8-27B as the open self-hosting row. Pricing: Haiku 5.5 at $0.10/$0.50 per million (under-100K prompts), Luna $0.10/$0.50 (up to 272K), Sonnet 5.5 at $2/$10 list with cache reads halved to $0.10, and Qwen3.8-27B free weights (Apache-2.0) on your own GPU. Core finding: Haiku 5.5 matches Luna only in the short-prompt tier - above 100K tokens its $0.50 input faces Luna's $0.20, so long-document batch bills invert the picture. Benchmarks cite each model's own official numbers, labeled and never mixed across suites: Haiku 5.5 Terminal-Bench 4.0 at 39.2% and OSWorld 72.4%; no fresh Luna numbers shipped with the release; Sonnet 5.5 at 70.6% on TB4; Qwen3.8-27B has no cross-suite table. Effort settings compared (Haiku 5.5 is the first Haiku-class with the dial; Luna runs none-to-max), contexts (1M for both hosted models; Qwen 262K extensible to 1M). Verdicts by use case - high-volume subagents, batch classification, main-line coding, private deployment - with the site's standing no-stress-test-no-arbitration stance. Four same-topic deep dives linked inline.

Published October 8, 202610 min read
<!-- small-model-price-war-comparison-review | review | Small Model Price War: Haiku 5.5 Ties Luna Only Below 100K -->

On October 7, 2026, Anthropic shipped Claude Haiku 5.5 and pinned its short-prompt tier at $0.10 per million input tokens and $0.50 per million output tokens. That number did not appear in a vacuum. OpenAI's GPT-6 Luna already lists $0.10 and $0.50 for prompts up to 272K tokens, and the open-weight Qwen3.8-27B gives its tokens away under Apache-2.0 if you supply the GPU. When three supply models land on the same floor price in one season, the question that separates bills stops being "who is cheapest" and becomes "where does each price sheet bend."

This review answers that across four options: Claude Haiku 5.5, GPT-6 Luna, Claude Sonnet 5.5 as the premium anchor, and Qwen3.8-27B as the self-hosted anchor — one master table, one per-model read, one section on the prompt-length boundary where the tie breaks, and use-case verdicts rather than a crown.

Calibers first. Every price is an official page snapshot taken October 8, 2026, and every tier threshold is quoted as the vendor draws it, because they draw them in different places. Every benchmark figure is attributed to its publisher, and we do not mix suites: a score from one company's internal table is never placed next to another's as if comparable. Where something is unpublished — Anthropic has published no SWE-bench figure for Haiku 5.5 — we say so rather than borrow a substitute. Per standing policy we crown no overall winner: no head-to-head stress tests were run this batch, so verdicts are given per need.


The Master Table: Four Price Sheets, One Snapshot

The table compresses the four dimensions that move real bills most. Read the tier thresholds carefully; they are the story of this article.

DimensionClaude Haiku 5.5GPT-6 LunaClaude Sonnet 5.5Qwen3.8-27B, self-hosted
Input / output, USD per 1M tokens$0.10 / $0.50 below 100K prompt tokens; $0.50 / $2.50 at 100K or more$0.10 / $0.50 at or below 272K; $0.20 / $0.75 above 272K$2 / $10, flatWeights free under Apache-2.0; you pay for GPU time
Context window1M tokens1M tokens, 128K max outputIn Anthropic's catalog, not restated this release262,144 native, expandable toward 1M
Cache pricingNot itemized in launch sheetRead $0.01, write $0.125Read halved to $0.10No API cache concept; every call is fresh compute
Adjustable effortYes, first Haiku-tier model to offer itYes, none to maxYesreasoning_effort adjustable; thinking can be switched off

Footnote for all rows: official page snapshot, October 8, 2026. Vision is not a table row: release materials itemize differently — Luna documents image input, Sonnet 5.5 supports it, Qwen3.8-27B ships native image and video understanding, and Haiku 5.5's materials do not list vision in launch scope.


Claude Haiku 5.5: The Subagent Price Anchor

Haiku 5.5 is Anthropic's third 5.5-family release in roughly a month, pitched at high-throughput, cost-sensitive work: summarization, compression, database querying, classification, and subagent legwork inside coding agents. The full launch context — benchmarks, safety tiers, availability across the Anthropic API, Bedrock, Google Cloud, and Azure — is in what actually shipped with Haiku 5.5, so here we stay on economics.

Two numbers define the sheet. The short tier: $0.10 input and $0.50 output per million tokens, which Anthropic says undercuts the previous Haiku by around 90 percent on that band and prices input at one-twentieth of Sonnet 5.5's $2. The fine print: prompts crossing 100K tokens jump to $0.50 input and $2.50 output. Anthropic's justification is that roughly 90 percent of Haiku requests historically stayed under 100K prompt tokens — which is why the cheap tier is the headline. Your workload decides whether you live in that 90 percent.

The citable benchmarks are Anthropic-reported, from the Haiku 5.5 launch table: Terminal-Bench 4.0 at 39.2 percent, OSWorld 2.1 offline at 72.4 percent, GDPval-AA v2.1 at 1620 — restated with the vendor label and no cross-suite comparisons. One absence matters equally: no SWE-bench figure was published at launch, and we will not substitute an older Haiku's score. Haiku 5.5 is also the first Haiku-tier model with adjustable effort, trading intelligence for cost per request — a lever that compounds when the model drives cheap subagents. For browser automation, real-time customer service, and classification farms, 1M context plus a $0.10 entry rate is the aggressive part; the long-prompt tier is where it gives ground back.


GPT-6 Luna: Flat Prices and a One-Cent Cache

Luna's sheet looks deceptively similar at a glance: $0.10 input and $0.50 output per million tokens. The difference is where the tier line sits. Luna holds that rate to 272K prompt tokens — nearly three times Haiku's threshold — and its long tier above that is $0.20 input and $0.75 output, not Haiku's $0.50 and $2.50. Context is 1M tokens with a 128K output ceiling, image input is documented, and effort spans the full none-to-max range.

The quietly decisive row is cache. Luna lists cache read at $0.01 per million tokens, write at $0.125. For agentic loops that resend a stable prompt and growing history every round, one-cent read makes the repeated context almost free — no cheaper published read price exists in this field. Haiku 5.5's sheet itemizes no cache read rate and Qwen self-hosting has no cache concept, so for loop-heavy workloads Luna's effective per-round cost pulls further ahead of its headline. For wiring Luna or its sibling Sol into production, our three-cloud price list and calling guide for GPT-6 Sol and Luna covers integration.

Per benchmark discipline, we do not restate Luna's scores next to Anthropic's table; OpenAI publishes its numbers on its own pages, and mixing suites would manufacture a comparison that does not exist.


Claude Sonnet 5.5: The Premium Row Learns to Discount

Sonnet 5.5 sits in this table as the ceiling, not the competitor: $2 input and $10 output, twenty times Haiku 5.5's short-tier input rate. Two things complicate the "expensive" label. One arrived alongside the Haiku launch: Anthropic halved Sonnet 5.5's cache read from $0.20 to $0.10 per million tokens, with the company's own estimate that most agentic workloads see total costs drop about 20 percent. That is a vendor estimate, not a measurement, but a repeated prefix on Sonnet 5.5 now bills at Haiku 5.5's base input rate, narrowing the real gap inside agent loops.

Second, the premium buys capability. In Anthropic's own launch table, Sonnet 5.5 posts Terminal-Bench 4.0 at 70.6 percent and GDPval-AA v2.1 at 1840 — same suite, same vendor, labeled like the Haiku figures above — and it remains the tier for planning multi-file changes or driving an agent's main loop. Its context window is documented in Anthropic's catalog and was not restated this release, so we point to the catalog. For the rest of the 5.5 family's pricing moves, our Sonnet 5.5 release coverage has the fuller picture.


Qwen3.8-27B: Free Weights, Priced-in GPUs

The fourth row redefines "price." Qwen3.8-27B's weights are free under Apache-2.0, so its per-million-token rate is zero — and meaningless without a utilization number next to it. The hardware math from community-tracked figures: the BF16 checkpoint wants about 55.6GB of VRAM, a data-center or multi-GPU rig at full precision, while the GGUF Q4_K_M quantization fits in roughly 19GB on prosumer hardware. Monthly downloads sit around 1.03 million by llm-explorer and ggml-org calibers — no niche experiment.

You give up the API's economic machinery: no cache, no vendor SLA. You gain control — data never leaves your infrastructure, no rate limit but your own GPUs, offline batch when hardware idles, and a flexible model: 262,144 native context expandable toward 1M, adjustable reasoning_effort, a switchable thinking mode, and native image plus video understanding that no hosted small model here fully matches on paper. For deployment reality rather than datasheet, our field-tested local deployment walkthrough for the Qwen3.8 family is the practical start. The honest cost formula: divide GPU amortization and power by utilization. A saturated box beats any API price; an idling one loses to all three hosted rows.


The Long-Prompt Gap: Where the $0.10 Tie Breaks

Here is the finding this review exists to pin down. The "$0.10/$0.50, same as Luna" headline for Haiku 5.5 holds only while your prompt stays under 100K tokens. The moment it crosses, the sheets diverge sharply, in Luna's favor.

Run the arithmetic across three bands. Under 100K prompt tokens, the hosted sheets are identical: $0.10 input, $0.50 output, a genuine tie. Between 100K and 272K the asymmetry is widest: Haiku has jumped to its long tier at $0.50 input and $2.50 output while Luna still charges its base rate — input and output are five times cheaper on Luna. Above 272K, both are on long tiers, and Luna's $0.20 input against Haiku's $0.50 remains 2.5 times cheaper, with $0.75 versus $2.50 on output about 3.3 times less.

Cache widens the gap for repeated long prompts. An agent rereading a large stable prefix every round pays Luna's $0.01 per million on that prefix; Haiku 5.5's sheet itemizes no cache read rate, so that context defaults to full input price on the published numbers.

This is not a theoretical corner. Retrieval pipelines injecting large contexts, transcript summarization, repository-scale code analysis, any agent whose context grows past a hundred thousand tokens — all live above the line. The rule of thumb: under 100K tokens, pick between Haiku 5.5 and Luna on features, ecosystem, and latency, because the price tie is real; above 100K, Luna's sheet wins the per-token math outright; and if the constraint is data privacy, the price sheet was never the deciding factor.


Verdicts by Use Case, Not by Crown

No single model wins this field, and we will not pretend one does without stress tests on a fixed workload. What the price sheets support is this:

High-volume short-prompt automation — classification, extraction, routing, short-document summarization, customer service triage. Haiku 5.5 and Luna are priced identically here, so the decision moves to everything around the price: Haiku's adjustable effort and its Bedrock, Google Cloud, and Azure presence against Luna's documented image input.

Agentic loops with a stable prefix. Luna's $0.01 cache read is the lowest published read price in the field, and Sonnet 5.5's halved $0.10 read keeps the premium tier competitive once repetition dominates. Per-round cost favors Luna; capability per round favors Sonnet.

Long prompts above 100K tokens. Luna, on published numbers, and it is not close — 2.5 times cheaper on input above 272K, five times cheaper in the 100K-to-272K band where Haiku pays long-tier rates and Luna does not.

Primary coding and agent orchestration. Sonnet 5.5, with Anthropic-reported Terminal-Bench 4.0 and GDPval-AA v2.1 figures leading its own suite, ideally paired with Haiku 5.5 subagents so the cheap model absorbs high-volume legwork while the premium model plans.

Data privacy, offline batch, or burst compute. Qwen3.8-27B self-hosted, where Apache-2.0 weights and a quantized 19GB footprint make marginal token cost a hardware question rather than a pricing question.

A two-step self-test beats any verdict we could hand you: fix a representative task set and a round cap, run it against your shortlisted two, and compare per-task cost with tier crossings marked — the 100K and 272K lines explain most of what you see.


FAQ

Q1: Is Haiku 5.5 really the same price as GPT-6 Luna?

A1: Only below 100K prompt tokens, where both list $0.10 input and $0.50 output per million. Above 100K, Haiku jumps to $0.50/$2.50 while Luna stays at base rate until 272K, then moves to $0.20/$0.75 — Luna's input is 2.5 times cheaper on the published sheets.

Q2: What exactly does Haiku 5.5 cost on long prompts?

A2: At 100K or more prompt tokens, $0.50 per million input and $2.50 per million output, five times its own short-tier rates. Anthropic notes about 90 percent of historical Haiku requests stayed under 100K — the headline price is built for them.

Q3: Is self-hosting Qwen3.8-27B actually cheaper than any API?

A3: It depends entirely on utilization. The weights are free under Apache-2.0, but the BF16 checkpoint needs about 55.6GB of VRAM and the Q4_K_M GGUF about 19GB. A heavily utilized GPU beats every API rate here; an idling one loses to all.

Q4: Did Anthropic publish SWE-bench or independently verified scores for Haiku 5.5?

A4: No SWE-bench figure was published at launch. The scores cited here — Terminal-Bench 4.0, OSWorld 2.1 offline, GDPval-AA v2.1 — are Anthropic-reported from its own launch table; treat any older-model score offered as a stand-in with suspicion.

Q5: Which of the four should I pick?

A5: Match the sheet to your token profile: Haiku 5.5 or Luna under 100K tokens, Luna above 100K on per-token cost, Sonnet 5.5 for orchestration-heavy coding with Haiku subagents underneath, self-hosted Qwen3.8-27B when data must stay in-house. No overall winner is claimed without stress tests on your own workload.


References

  • Anthropic Haiku 5.5 launch and pricing page, snapshot October 8, 2026: short and long tier prices, 1M context, adjustable effort, Terminal-Bench 4.0 39.2 percent, OSWorld 2.1 offline 72.4 percent, GDPval-AA v2.1 1620, no SWE-bench published. Vendor-reported.
  • OpenAI GPT-6 Luna pricing page, snapshot October 8, 2026: tier prices at and above 272K, cache read $0.01, write $0.125, 1M context, 128K max output, image input. Vendor-reported.
  • Anthropic Sonnet 5.5 materials, snapshot October 8, 2026: $2/$10, cache read halved $0.20 to $0.10, Terminal-Bench 4.0 70.6 percent, GDPval-AA v2.1 1840, about 20 percent agentic cost reduction as a vendor estimate.
  • Qwen3.8-27B model card and community trackers (llm-explorer, ggml-org calibers): Apache-2.0 weights, BF16 VRAM about 55.6GB, GGUF Q4_K_M about 19GB, monthly downloads around 1.03 million, 262,144 native context. No cross-suite benchmark comparisons made.

This article is AI-assisted and human-edited. Last updated: 2026-10-08

FAQ

Is Haiku 5.5 really the same price as GPT-6 Luna?
Only below 100K prompt tokens, where both list $0.10 input and $0.50 output per million. Above 100K, Haiku jumps to $0.50/$2.50 while Luna stays at base rate until 272K, then moves to $0.20/$0.75 — Luna's input is 2.5 times cheaper on the published sheets.
What exactly does Haiku 5.5 cost on long prompts?
At 100K or more prompt tokens, $0.50 per million input and $2.50 per million output, five times its own short-tier rates. Anthropic notes about 90 percent of historical Haiku requests stayed under 100K — the headline price is built for them.
Is self-hosting Qwen3.8-27B actually cheaper than any API?
It depends entirely on utilization. The weights are free under Apache-2.0, but the BF16 checkpoint needs about 55.6GB of VRAM and the Q4_K_M GGUF about 19GB. A heavily utilized GPU beats every API rate here; an idling one loses to all.
Did Anthropic publish SWE-bench or independently verified scores for Haiku 5.5?
No SWE-bench figure was published at launch. The scores cited here — Terminal-Bench 4.0, OSWorld 2.1 offline, GDPval-AA v2.1 — are Anthropic-reported from its own launch table; treat any older-model score offered as a stand-in with suspicion.
Which of the four should I pick?
Match the sheet to your token profile: Haiku 5.5 or Luna under 100K tokens, Luna above 100K on per-token cost, Sonnet 5.5 for orchestration-heavy coding with Haiku subagents underneath, self-hosted Qwen3.8-27B when data must stay in-house. No overall winner is claimed without stress tests on your own workload.

Related

Hardcore Reviews

17 hours vs 16 days: which AI can actually work overnight

A continuous-autonomy comparison (October 7, 2026 basis; a different axis from the site's October 5 single-output-capacity review, which it links rather than repeats): how long can one AI work on a single task, and what does it deliver. Four rows: Ant's Ling-3.1-flash wrote a Lua-to-x86-64 compiler from scratch in about 17 hours (178/182 tests passing, self-reported) and ported a C image library to Rust in about 20 hours with an 8.015x speedup (30/30 checks); Alibaba's Qwen3.8-Max ran a 16-day autonomous build of oh-my-cli (265 commits/127 PRs/151 issues, with a public GitHub trace) plus a ~125-hour paper reproduction that beat the original method (AIME24 49.58% to 52.29%, 7,600 lines of code, 33 training rounds); Google's Argon handled the 800K+ line Zircon migration and a 2.7x libgav1 rewrite (DeepSWE v1.1 77.9%, general public still locked out); reference row DeepSeek-V4-Flash-Vision-Exp scores 59.3 on the same benchmark (MIT, self-hostable). Table columns: model / longest case / verifiability / access and cost / sourcing. An architecture section answers "Ling vs Qwen": Ling uses 7 KDA + 1 Gated MLA layers, Qwen 3 Gated DeltaNet + 1 Gated Attention, both 512-expert MoE. License red lines reported faithfully: Qwen3.8-Max open weights ship under a custom qwen3.8-max license (revenue share, threshold not finalized), not Apache 2.0 (only the 27B sibling is), and the open build is text-only with forced thinking - not the API version with vision, 1M context and tools; the "half-open" controversy is presented from both sides; 4.89TB of weights need 24 GPUs (BF16) or 16 (FP8). All numbers self-reported; DeepSWE v1.1 compared only within v1.1.

Oct 7, 202610 min read
Hardcore Reviews

One Million Output Tokens: Which Model Finishes the Job

A four-way comparison of coding/general models (October 5, 2026 basis; complements rather than repeats the site's late-September pricing review). The axis this time is single-response output length: Gemini 4 Argon stretched output from 64K to 1 million tokens (officially the industry's largest), changing the battlefield for long tasks. Four positions: the controlled flagship Argon (released September 30; officially first place on 13 of 18 benchmarks and a 77.9% DeepSWE v1.1 SOTA, yet in a controlled rollout the general public cannot buy; Artificial Analysis independently scores it slightly lower; introductory $2/$10 rising to $4/$20), the reigning efficiency benchmark Claude Sonnet 5.5 (September 28; Terminal-Bench 4.0 at 70.6%; $2/$10 on three clouds), the cache-economics pick GPT-6.1 Sol ($2/$0.10/$10 - cached input 95% cheaper), and the open-source reference GLM-5.3 (about 750GB of weights on ModelScope for self-hosting, domestic API $1.40/$4.40, open-source coding SOTA with CyberGym at 84.5%). Red lines honored: GLM's Terminal Bench 3.0 and Sonnet's Terminal-Bench 4.0 are different benchmark versions and are never compared directly; both official and independent assessments of Argon are presented. No benchmarks were re-run and no winner is crowned - conclusions follow the need: wait for Argon's wider access, pick Sonnet 5.5 for balanced value, Sol for cache-heavy agent pipelines, Sonnet for multi-cloud, GLM-5.3 for open self-hosting.

Oct 5, 202610 min read
Hardcore Reviews

30-Second Club: Kling 4.0, Seedance 2.5, Wan 3.0, LTX-2.5

A four-way video-model comparison (October 3, 2026 basis; complements rather than repeats the site's August five-way closed-cloud review and the Seedance 2.5 vs MiniMax H3 head-to-head). Framed by the observation that 30-second native single-pass output has become the entry ticket among flagship models, it picks one model per quadrant - cloud flagship, reigning benchmark, cloud value, open self-hosted: Kling 4.0 (October launch; 10-bit HDR, 2-minute extension, 15 references; specs announced but unverified), Seedance 2.5 (live since July 31; 30-second native output, 50 references, sub-second local editing; production deployments at XCMG, XPeng and Differential Intelligence Flight), Wan 3.0 (live since August 24; API from 0.3 CNY per second, five office document formats direct-to-video, no agent capabilities so orchestration is DIY), and LTX-2.5 (the only open weights; 66 GiB, 8-step distillation, ltx-trainer; custom community license with a USD 10 million revenue threshold). With no benchmark runs of its own, the piece refuses to crown a winner and concludes by need: Seedance for a working flagship today, Kling 4.0 to watch for pro specs, Wan 3.0 for value and batch work, LTX-2.5 for data privacy. Sora is excluded over conflicting sources and Veo unverified.

Oct 3, 202610 min read