On October 7, 2026, Anthropic released Claude Haiku 5.5, and the official framing leaves little to interpretation: this is the company's cheapest, fastest, most capable small model. It is the third member of the Claude 5.5 family to ship in roughly a month, following Opus 5.5 and then Sonnet 5.5, and it went live the same day on the Claude platform, AWS Bedrock, Google Cloud, and Azure. The model ID is claude-haiku-5-5.
The headline number is the price: a dime per million input tokens for most prompts, a price point that puts it in a dead heat with OpenAI's cheapest GPT-6 tier. The context window jumps from 200,000 tokens to a full one million, and Haiku 5.5 becomes the first Haiku-class model to offer adjustable effort, the same dial Sonnet and Opus users got earlier. There is one catch in the fine print, though: once a prompt crosses the 100,000-token mark, Haiku 5.5's input price climbs to five times the short-prompt rate, which is 2.5 times what GPT-6 Luna charges for the same tokens. This article walks through the pricing table and that gap, the context and effort upgrades, every published benchmark, the safety changes, and the most practical way to put the model to work. All benchmark figures below are vendor-reported unless labeled otherwise; anything unpublished is not filled in with guesses.
What Shipped: The Cheapest and Fastest Claude Yet
The official positioning is more specific than the marketing sentence suggests. Haiku 5.5 is built for high-volume, cost-sensitive tasks: summarizing documents, compressing context, running database queries, and classifying inputs. These are the jobs where a pipeline might make thousands of calls an hour, where a fivefold price difference is not a rounding error but the difference between a feature being viable or not. The second use case is structural: Haiku 5.5 is designed to run as a coding subagent alongside Opus 5.5 or Sonnet 5.5, handling lookups, file searches, and quick edits while the bigger model plans and reviews. The third is latency: per the official positioning, this is the fastest model the company has shipped, targeting real-time customer service and browser operation, the two settings where slow responses are directly felt by end users. Read the rest of the release in that light: the benchmark table is supporting evidence, and the pricing section is where the real story sits.
Pricing: A Dime Per Million Tokens, With a Catch Above 100K
Here is the full pricing structure. For prompts under 100,000 tokens, Haiku 5.5 costs 0.10 dollars per million input tokens and 0.50 dollars per million output tokens. Anthropic says this represents average savings of roughly 75 percent compared with Haiku 4.5 across workloads, and about 90 percent within that short-prompt tier specifically. For prompts at or above 100,000 tokens, the rate is 0.50 dollars per million input and 2.50 dollars per million output. To make that tiered structure land, the company shared one telling stat: by its own account, about 90 percent of requests to previous Haiku models carried prompts under 100,000 tokens. In other words, the overwhelming majority of existing Haiku traffic gets the dime rate.
Put that next to the rest of the family and the scale of the cut becomes concrete. Sonnet 5.5 costs 2 dollars per million input tokens and 10 dollars per million output, per the pricing covered in our Sonnet 5.5 launch breakdown. Haiku 5.5's input price is exactly one twentieth of that. A coding setup that routes search, summarization, and triage calls to Haiku instead of Sonnet pays a twentieth of the input cost on those calls, which is the entire argument for multi-model agent architectures. The same release window brought a second cost move: Sonnet 5.5's cached-read price was cut in half to 0.10 dollars per million tokens, with the official claim that most agentic workloads see total costs drop by about 20 percent as a result. Anthropic also added a monthly API allowance for Max and Team subscription plans, blurring the line between flat-rate subscriptions and pay-as-you-go usage.
Now the catch, and the most interesting number in this whole launch. GPT-6 Luna, OpenAI's small-model tier, is priced at 0.10 dollars per million input and 0.50 dollars per million output for prompts up to 272,000 tokens, with cached reads at 0.01 dollars and a 1 million-token context window, per the details in our GPT-6 Luna release coverage. On short prompts, the two are priced identically: dime in, fifty cents out. That is the "same price" story. But Haiku 5.5's long-prompt tier kicks in at 100,000 tokens, where input jumps to 0.50 dollars per million, while Luna holds its 0.20 dollars per million input rate beyond its own threshold. Above 100K prompts, Haiku 5.5 input is 2.5 times more expensive than Luna. The equivalence only holds under 100K; past that line, OpenAI is meaningfully cheaper on input, and Luna's penny-priced cached reads undercut anything Anthropic published for this model. For teams running document-length prompts, that single line item should drive the model choice, and our cross-vendor cost-performance comparison runs this arithmetic across all four price lists side by side.
1M Context and Adjustable Effort: Two Upgrades That Change How You Use It
The context window jumps from 200,000 tokens in Haiku 4.5 to 1 million tokens, a fivefold increase that brings the small model level with the flagship tiers. For the official use cases this is not cosmetic: a summarization pipeline that previously had to chunk a large document set can now pass it whole, and an agent doing repository-wide search can hold more of the codebase in working memory. Note how this interacts with the pricing tier above: the moment a prompt crosses 100,000 tokens, the input rate quintuples. The big window and the price cliff are two sides of the same design.
The second upgrade is adjustable effort. Haiku 5.5 is the first Haiku-class model to support the same effort dial that Sonnet and Opus models offer, letting a single request choose between spending less and thinking harder. In agent architectures this is quietly powerful: a fleet of Haiku subagents can run at low effort for routine lookups and escalate individual requests to higher effort only when a task turns out to be harder than expected, without switching models or rewriting the integration. Cost and intelligence become a per-request decision rather than an architectural one.
The Benchmarks, Properly Labeled
Every figure in this section is a vendor-reported official number, and every comparison stays within a single benchmark suite, with Haiku 4.5 as the generational baseline and GPT-6 Luna and Sonnet 5.5 as the reference points the company chose to publish.
On GDPval-AA v2.1, Haiku 5.5 scores 1620, up from 735 for Haiku 4.5, with Luna at 1437 and Sonnet 5.5 at 1840. On AA-Briefcase v1.1, the new Haiku scores 1578 versus 614 for its predecessor, against 1336 for Luna and 1824 for Sonnet 5.5. On OSWorld 2.1 in offline mode, a computer-use benchmark, Haiku 5.5 reaches 72.4 percent, up from 15.7 percent for Haiku 4.5, with Luna at 48.9 percent and Sonnet 5.5 at 83.9 percent. The jump on OSWorld is arguably the standout of the table: a 4.6-fold generational gain that moves the small model past Luna and into striking distance of the mid-tier.
On the Humanity's Last Exam, Haiku 5.5 scores 45.9 percent without tools and 57.4 percent with tools, where Haiku 4.5 managed 10.2 and 18.7 percent and Sonnet 5.5 posts 56.9 and 64.5 percent. On Terminal-Bench 4.0, the terminal coding benchmark, Haiku 5.5 scores 39.2 percent against 0.0 percent for Haiku 4.5, with Luna at 16.4 percent and Sonnet 5.5 at 70.6 percent. That 0.0 is worth pausing on: the previous Haiku did not complete a single terminal task on this suite, and the new one clears roughly four in ten. On FrontierCode 1.1 Main, Haiku 5.5 scores 46.4 percent, ahead of Luna's 42.4 percent and behind Sonnet 5.5's 52.1 percent at its Xhigh setting. On Chartography, a visual reasoning benchmark, Haiku 5.5 scores 46.4 percent without tools, up from 6.4 percent for Haiku 4.5, with Luna at 29.1 percent and Sonnet 5.5 at 61.6 percent.
One honest gap: Anthropic did not publish a SWE-bench score for Haiku 5.5, and coverage elsewhere has flagged that omission. We will not borrow Haiku 4.5's older figure or fill the slot with an estimate. If SWE-bench matters to your evaluation, that row stays blank until the vendor publishes it.
Finally, the customer data, clearly labeled as such: Asana reports that integrating Haiku 5.5 cut task completion latency by 30 percent and sped up inference per agent turn by 2.5 times. These are customer-reported figures from a launch partner, not independent measurements, and they describe Asana's own workloads.
Safety: The First Haiku With Built-In Cyber Safeguards
The safety changes in this release are easy to skim past but carry real weight. Official statements say alignment evaluations improved across the board, with fewer misaligned behaviors detected. More concretely, Haiku 5.5 is the first Haiku model to ship with built-in cybersecurity safeguards. The tier is looser than Sonnet 5.5's, but it still blocks penetration-testing-style requests, and biological safeguards match the standard set by Sonnet 5, Sonnet 5.5, and Opus 5. The same day, Anthropic opened its Cyber Verification Program with three tiers, formalizing how security researchers and red teams can request expanded cyber capabilities under oversight.
The logic mirrors what we noted when Sonnet 5.5 pushed flagship-grade protections downmarket: the small models run at the highest volumes and are the most likely to be wired into unsupervised pipelines, so protection levels have to track usage, not model tier.
The 5.5 Family Sprint and the Pre-IPO Backdrop
Zoom out and the release cadence itself is the story. Reportedly, Haiku 5.5 is the third model in the 5.5 family to ship within a single month, after Opus 5.5 and Sonnet 5.5, and the second price-related move in the space of days once the Sonnet cache-price cut is counted. Anthropic has in effect repriced the entire stack in under a week, from flagship to entry tier, while competitors hold their own lists.
The financial backdrop, per media reports: Reuters coverage ties the pacing to Anthropic's reported IPO preparations, noting a confidential SEC filing reportedly submitted on June 1, a rumored Nasdaq listing as early as October, and valuation chatter reportedly reaching around 2 trillion dollars. All of that is reported context, not confirmed fact, and none of it changes what the model does. But it frames the strategy: a pre-IPO company wants its cheapest tier absorbing the most volume, because entry-level usage growth is the line that trends up in the story told to public-market investors.
How to Put It to Work Today
The most practical entry point for most developers is the coding subagent pattern, and the economics here are the cleanest yet. Configuring a Claude Code subagent with model set to haiku now routes that agent to Haiku 5.5 automatically, with no integration change, and the input price is one twentieth of Sonnet's for the lookup, summary, and triage calls those agents spend most of their time on. Add adjustable effort on top and a subagent fleet can hold a low baseline cost while escalating individual hard requests. Our hands-on Claude Code subagent guide covers the configuration files, the model routing rules, the isolation mechanics, and the mistakes that burn tokens in production.
For choosing between Haiku 5.5 and Luna at the model-selection level, the rule of thumb falls out of the pricing section: under 100,000-token prompts, pick on ecosystem, latency, and cached-read needs, because the sticker prices are identical; at or above 100,000-token prompts, Luna's input rate is 2.5 times cheaper, and document-scale workloads should weigh that seriously.
The Takeaway
Haiku 5.5 is a deliberate, disciplined release: not a benchmark flex, but a repricing of the volume tier with a fivefold context increase and an effort dial attached. The one-twentieth-of-Sonnet input price plus the 0.0-to-39.2 percent Terminal-Bench jump is the core of the story: small-model work that was previously either too expensive or too weak now fits on one cheap rail. The 2.5-times long-prompt gap against Luna is the asterisk, and the unpublished SWE-bench row is the honest blank. Whether the scores hold up will depend on independent testing; what can be said on day one is that the price floor for a capable, fast, million-token-context model just moved, and every competitor's small-model tier now has to answer for it.
Join the Discussion
Are you routing subagents to Haiku 5.5 already, and did your agent fleet's per-task cost actually drop as much as the price list suggests? If you run document-length prompts, does the 100K price cliff change your model choice? Tell us what your first week with the model looks like in the comments.
Sources
- Anthropic official site and launch materials: positioning, pricing tiers, context window, adjustable effort, benchmark table, safety changes, and platform availability
- Reuters, October 2026: launch coverage and reported IPO preparation details, including the confidential filing, rumored listing window, and reported valuation figures
- startupfortune: noting that no SWE-bench score was published for Haiku 5.5
- Asana: customer-reported latency and inference speed figures, as cited in launch material
- GPT-6 Luna pricing from OpenAI's official pages, cross-checked against our earlier Luna coverage
- Note: all prices and benchmark scores are snapshots as of October 8, 2026, and defer to the official vendor pages; benchmark figures are vendor-reported and not independently verified by this site.