Frontline Hotspot
Frontline Hotspot

Claude Haiku 5.5: $0.10 Input, 1M Context, One Pricing Catch

Anthropic shipped Claude Haiku 5.5 on Wednesday, October 7, 2026 - the third model in its 5.5 family within a month, after Opus 5.5 and Sonnet 5.5. Pricing (official): $0.10 per million input tokens and $0.50 output for prompts under 100K tokens - roughly one-twentieth of Sonnet 5.5's $2 input rate, about 75% cheaper than Haiku 4.5 on average, and matching GPT-6 Luna's short-prompt tier; above 100K it steps up to $0.50/$2.50, making it 2.5x pricier than Luna's $0.20 input - the "same price" claim holds only for short prompts. Context jumps from 200K to 1M tokens, and it is the first Haiku-class model with an adjustable effort setting (cost vs intelligence on the same request). Official benchmark table (self-reported): GDPval-AA 1620, OSWorld 72.4%, Terminal-Bench 4.0 39.2%, FrontierCode 46.4%, benchmarked against Haiku 4.5 and GPT-6 Luna; SWE-bench is unpublished and not invented here. Launch-day price moves: Sonnet 5.5 cache reads halved to $0.10 (official: about 20% cheaper on most agentic work), plus new monthly API credits for Max and Team subscribers. The positioning is blunt: pair it with Opus/Sonnet 5.5 as coding subagents and run high-volume, latency-sensitive work like live support and browser use. Reuters frames the cadence as pre-IPO expansion (reportedly). Available day-one on Claude Platform, AWS, Google Cloud and Azure.

Published October 8, 20269 min read
<!-- claude-haiku-5-5-release-hotspot | hotspot | Claude Haiku 5.5: $0.10 Input, 1M Context, One Pricing Catch -->

On October 7, 2026, Anthropic released Claude Haiku 5.5, and the official framing leaves little to interpretation: this is the company's cheapest, fastest, most capable small model. It is the third member of the Claude 5.5 family to ship in roughly a month, following Opus 5.5 and then Sonnet 5.5, and it went live the same day on the Claude platform, AWS Bedrock, Google Cloud, and Azure. The model ID is claude-haiku-5-5.

The headline number is the price: a dime per million input tokens for most prompts, a price point that puts it in a dead heat with OpenAI's cheapest GPT-6 tier. The context window jumps from 200,000 tokens to a full one million, and Haiku 5.5 becomes the first Haiku-class model to offer adjustable effort, the same dial Sonnet and Opus users got earlier. There is one catch in the fine print, though: once a prompt crosses the 100,000-token mark, Haiku 5.5's input price climbs to five times the short-prompt rate, which is 2.5 times what GPT-6 Luna charges for the same tokens. This article walks through the pricing table and that gap, the context and effort upgrades, every published benchmark, the safety changes, and the most practical way to put the model to work. All benchmark figures below are vendor-reported unless labeled otherwise; anything unpublished is not filled in with guesses.

What Shipped: The Cheapest and Fastest Claude Yet

The official positioning is more specific than the marketing sentence suggests. Haiku 5.5 is built for high-volume, cost-sensitive tasks: summarizing documents, compressing context, running database queries, and classifying inputs. These are the jobs where a pipeline might make thousands of calls an hour, where a fivefold price difference is not a rounding error but the difference between a feature being viable or not. The second use case is structural: Haiku 5.5 is designed to run as a coding subagent alongside Opus 5.5 or Sonnet 5.5, handling lookups, file searches, and quick edits while the bigger model plans and reviews. The third is latency: per the official positioning, this is the fastest model the company has shipped, targeting real-time customer service and browser operation, the two settings where slow responses are directly felt by end users. Read the rest of the release in that light: the benchmark table is supporting evidence, and the pricing section is where the real story sits.

Pricing: A Dime Per Million Tokens, With a Catch Above 100K

Here is the full pricing structure. For prompts under 100,000 tokens, Haiku 5.5 costs 0.10 dollars per million input tokens and 0.50 dollars per million output tokens. Anthropic says this represents average savings of roughly 75 percent compared with Haiku 4.5 across workloads, and about 90 percent within that short-prompt tier specifically. For prompts at or above 100,000 tokens, the rate is 0.50 dollars per million input and 2.50 dollars per million output. To make that tiered structure land, the company shared one telling stat: by its own account, about 90 percent of requests to previous Haiku models carried prompts under 100,000 tokens. In other words, the overwhelming majority of existing Haiku traffic gets the dime rate.

Put that next to the rest of the family and the scale of the cut becomes concrete. Sonnet 5.5 costs 2 dollars per million input tokens and 10 dollars per million output, per the pricing covered in our Sonnet 5.5 launch breakdown. Haiku 5.5's input price is exactly one twentieth of that. A coding setup that routes search, summarization, and triage calls to Haiku instead of Sonnet pays a twentieth of the input cost on those calls, which is the entire argument for multi-model agent architectures. The same release window brought a second cost move: Sonnet 5.5's cached-read price was cut in half to 0.10 dollars per million tokens, with the official claim that most agentic workloads see total costs drop by about 20 percent as a result. Anthropic also added a monthly API allowance for Max and Team subscription plans, blurring the line between flat-rate subscriptions and pay-as-you-go usage.

Now the catch, and the most interesting number in this whole launch. GPT-6 Luna, OpenAI's small-model tier, is priced at 0.10 dollars per million input and 0.50 dollars per million output for prompts up to 272,000 tokens, with cached reads at 0.01 dollars and a 1 million-token context window, per the details in our GPT-6 Luna release coverage. On short prompts, the two are priced identically: dime in, fifty cents out. That is the "same price" story. But Haiku 5.5's long-prompt tier kicks in at 100,000 tokens, where input jumps to 0.50 dollars per million, while Luna holds its 0.20 dollars per million input rate beyond its own threshold. Above 100K prompts, Haiku 5.5 input is 2.5 times more expensive than Luna. The equivalence only holds under 100K; past that line, OpenAI is meaningfully cheaper on input, and Luna's penny-priced cached reads undercut anything Anthropic published for this model. For teams running document-length prompts, that single line item should drive the model choice, and our cross-vendor cost-performance comparison runs this arithmetic across all four price lists side by side.

1M Context and Adjustable Effort: Two Upgrades That Change How You Use It

The context window jumps from 200,000 tokens in Haiku 4.5 to 1 million tokens, a fivefold increase that brings the small model level with the flagship tiers. For the official use cases this is not cosmetic: a summarization pipeline that previously had to chunk a large document set can now pass it whole, and an agent doing repository-wide search can hold more of the codebase in working memory. Note how this interacts with the pricing tier above: the moment a prompt crosses 100,000 tokens, the input rate quintuples. The big window and the price cliff are two sides of the same design.

The second upgrade is adjustable effort. Haiku 5.5 is the first Haiku-class model to support the same effort dial that Sonnet and Opus models offer, letting a single request choose between spending less and thinking harder. In agent architectures this is quietly powerful: a fleet of Haiku subagents can run at low effort for routine lookups and escalate individual requests to higher effort only when a task turns out to be harder than expected, without switching models or rewriting the integration. Cost and intelligence become a per-request decision rather than an architectural one.

The Benchmarks, Properly Labeled

Every figure in this section is a vendor-reported official number, and every comparison stays within a single benchmark suite, with Haiku 4.5 as the generational baseline and GPT-6 Luna and Sonnet 5.5 as the reference points the company chose to publish.

On GDPval-AA v2.1, Haiku 5.5 scores 1620, up from 735 for Haiku 4.5, with Luna at 1437 and Sonnet 5.5 at 1840. On AA-Briefcase v1.1, the new Haiku scores 1578 versus 614 for its predecessor, against 1336 for Luna and 1824 for Sonnet 5.5. On OSWorld 2.1 in offline mode, a computer-use benchmark, Haiku 5.5 reaches 72.4 percent, up from 15.7 percent for Haiku 4.5, with Luna at 48.9 percent and Sonnet 5.5 at 83.9 percent. The jump on OSWorld is arguably the standout of the table: a 4.6-fold generational gain that moves the small model past Luna and into striking distance of the mid-tier.

On the Humanity's Last Exam, Haiku 5.5 scores 45.9 percent without tools and 57.4 percent with tools, where Haiku 4.5 managed 10.2 and 18.7 percent and Sonnet 5.5 posts 56.9 and 64.5 percent. On Terminal-Bench 4.0, the terminal coding benchmark, Haiku 5.5 scores 39.2 percent against 0.0 percent for Haiku 4.5, with Luna at 16.4 percent and Sonnet 5.5 at 70.6 percent. That 0.0 is worth pausing on: the previous Haiku did not complete a single terminal task on this suite, and the new one clears roughly four in ten. On FrontierCode 1.1 Main, Haiku 5.5 scores 46.4 percent, ahead of Luna's 42.4 percent and behind Sonnet 5.5's 52.1 percent at its Xhigh setting. On Chartography, a visual reasoning benchmark, Haiku 5.5 scores 46.4 percent without tools, up from 6.4 percent for Haiku 4.5, with Luna at 29.1 percent and Sonnet 5.5 at 61.6 percent.

One honest gap: Anthropic did not publish a SWE-bench score for Haiku 5.5, and coverage elsewhere has flagged that omission. We will not borrow Haiku 4.5's older figure or fill the slot with an estimate. If SWE-bench matters to your evaluation, that row stays blank until the vendor publishes it.

Finally, the customer data, clearly labeled as such: Asana reports that integrating Haiku 5.5 cut task completion latency by 30 percent and sped up inference per agent turn by 2.5 times. These are customer-reported figures from a launch partner, not independent measurements, and they describe Asana's own workloads.

Safety: The First Haiku With Built-In Cyber Safeguards

The safety changes in this release are easy to skim past but carry real weight. Official statements say alignment evaluations improved across the board, with fewer misaligned behaviors detected. More concretely, Haiku 5.5 is the first Haiku model to ship with built-in cybersecurity safeguards. The tier is looser than Sonnet 5.5's, but it still blocks penetration-testing-style requests, and biological safeguards match the standard set by Sonnet 5, Sonnet 5.5, and Opus 5. The same day, Anthropic opened its Cyber Verification Program with three tiers, formalizing how security researchers and red teams can request expanded cyber capabilities under oversight.

The logic mirrors what we noted when Sonnet 5.5 pushed flagship-grade protections downmarket: the small models run at the highest volumes and are the most likely to be wired into unsupervised pipelines, so protection levels have to track usage, not model tier.

The 5.5 Family Sprint and the Pre-IPO Backdrop

Zoom out and the release cadence itself is the story. Reportedly, Haiku 5.5 is the third model in the 5.5 family to ship within a single month, after Opus 5.5 and Sonnet 5.5, and the second price-related move in the space of days once the Sonnet cache-price cut is counted. Anthropic has in effect repriced the entire stack in under a week, from flagship to entry tier, while competitors hold their own lists.

The financial backdrop, per media reports: Reuters coverage ties the pacing to Anthropic's reported IPO preparations, noting a confidential SEC filing reportedly submitted on June 1, a rumored Nasdaq listing as early as October, and valuation chatter reportedly reaching around 2 trillion dollars. All of that is reported context, not confirmed fact, and none of it changes what the model does. But it frames the strategy: a pre-IPO company wants its cheapest tier absorbing the most volume, because entry-level usage growth is the line that trends up in the story told to public-market investors.

How to Put It to Work Today

The most practical entry point for most developers is the coding subagent pattern, and the economics here are the cleanest yet. Configuring a Claude Code subagent with model set to haiku now routes that agent to Haiku 5.5 automatically, with no integration change, and the input price is one twentieth of Sonnet's for the lookup, summary, and triage calls those agents spend most of their time on. Add adjustable effort on top and a subagent fleet can hold a low baseline cost while escalating individual hard requests. Our hands-on Claude Code subagent guide covers the configuration files, the model routing rules, the isolation mechanics, and the mistakes that burn tokens in production.

For choosing between Haiku 5.5 and Luna at the model-selection level, the rule of thumb falls out of the pricing section: under 100,000-token prompts, pick on ecosystem, latency, and cached-read needs, because the sticker prices are identical; at or above 100,000-token prompts, Luna's input rate is 2.5 times cheaper, and document-scale workloads should weigh that seriously.

The Takeaway

Haiku 5.5 is a deliberate, disciplined release: not a benchmark flex, but a repricing of the volume tier with a fivefold context increase and an effort dial attached. The one-twentieth-of-Sonnet input price plus the 0.0-to-39.2 percent Terminal-Bench jump is the core of the story: small-model work that was previously either too expensive or too weak now fits on one cheap rail. The 2.5-times long-prompt gap against Luna is the asterisk, and the unpublished SWE-bench row is the honest blank. Whether the scores hold up will depend on independent testing; what can be said on day one is that the price floor for a capable, fast, million-token-context model just moved, and every competitor's small-model tier now has to answer for it.

Join the Discussion

Are you routing subagents to Haiku 5.5 already, and did your agent fleet's per-task cost actually drop as much as the price list suggests? If you run document-length prompts, does the 100K price cliff change your model choice? Tell us what your first week with the model looks like in the comments.

Sources

  • Anthropic official site and launch materials: positioning, pricing tiers, context window, adjustable effort, benchmark table, safety changes, and platform availability
  • Reuters, October 2026: launch coverage and reported IPO preparation details, including the confidential filing, rumored listing window, and reported valuation figures
  • startupfortune: noting that no SWE-bench score was published for Haiku 5.5
  • Asana: customer-reported latency and inference speed figures, as cited in launch material
  • GPT-6 Luna pricing from OpenAI's official pages, cross-checked against our earlier Luna coverage
  • Note: all prices and benchmark scores are snapshots as of October 8, 2026, and defer to the official vendor pages; benchmark figures are vendor-reported and not independently verified by this site.

This article is AI-assisted and human-edited. Last updated: 2026-10-08

Related

Frontline Hotspot

Gemini 4 Argon: 800K lines in one reply, 77.9% on DeepSWE

Google DeepMind launched Gemini 4 Argon on September 30, 2026 (official basis): the first Gemini 4 flagship, built for long-horizon work across real-world software engineering, legal and finance knowledge work, and cyber defense. The headline change: output limit stretched from 64K to 1 million tokens (about 16x). Google-reported benchmarks put DeepSWE v1.1 at 77.9% for a new SOTA (vs GPT-6 Astra 74.1, Claude Opus 5.5 74.2), AutomationBench at 51.3% for first place, CWE-bench v1 at 68% tied-first, and GraphWalks 256k-1M at 84.2%. Weaknesses reported faithfully: FrontierSWE v2 55.0 trails Astra's 65.5 and Terminal-bench 4.0 57.4 trails Opus 5.5's 66.4 - this piece attributes the split to task shape (long-horizon wins, terminal step-by-step loses), noting Google offers no explanation. Internal cases: the 800K+ line C/C++-to-Rust migration of the Fuchsia Zircon kernel; libgav1 rewritten at 32K lines of SIMD code running 2.7x faster with frame-identical output; datacenter memory work freeing 300+ TiB. Pricing: introductory $2/$10 (cache 95% off), then $4/$20. Controlled rollout reported as-is: Fairwind Program first with 650+ partners (trusted defenders get an unguarded build), general public and paid API still locked out with no date. All scores are Google-reported.

Oct 7, 20269 min read
Frontline Hotspot

Claude Sonnet 5.5 Ships: Coding Score Leaps From 10 to 70

Claude Sonnet 5.5 shipped September 28, 2026 (US Eastern, official basis): the second model in the Claude 5.5 family, positioned as a faster, cheaper complement to Opus 5.5, live day one on the Claude platform plus AWS Bedrock, Google Cloud Vertex and Microsoft Azure, with Claude Code integrated the same day. Headline numbers: Terminal-Bench 4.0 jumps from Sonnet 5's 10.3% to 70.6% (same benchmark, different generation - nearly sevenfold); CursorBench 4.0 at 55.5% (Opus 5.5: 57.8%); OSWorld 2.1 from 57.0% to 80.1%; GDPval-AA within 2 points of Opus 5.5; output speed up over 30% and per-task cost down up to 30% on most work. Pricing strategy: unit prices unchanged versus Sonnet 5 ($2/$10, cache read $0.20) - the discount hides in token efficiency. Security: the Sonnet line gets Opus/Fable-grade cyber protections for the first time, plus an anti-extraction classifier and auto-fallback for high-risk cyber requests. Also the first Sonnet to beat Pokemon Red from screenshots alone. Competitive context: Gemini 4 Argon landed two days later (controlled release, 1M output tokens) and GPT-6.1 Astra was delayed over safety; enterprise customers are about 80% of Anthropic's business ahead of a planned IPO (Reuters). Haiku 5.5 is teased for the coming weeks - no specs published, none invented here.

Oct 5, 20269 min read
Frontline Hotspot

Anthropic's $2 Trillion IPO Run Starts With a $30 Trillion Pitch to Wall Street

Per the Wall Street Journal on August 25, Anthropic's IPO filing presents a total addressable market of over $30 trillion - surpassing SpaceX's $28.5 trillion to become the largest market narrative in business history. The math does not start from software or hardware sales: it tallies the total economic value and labor cost of all future work replaceable by AI models. The listing targets Fall 2026, a raise of up to $100 billion, and a valuation anchored near $2 trillion - a 2.2x jump over the $900 billion private valuation from May. The confidence: $11.6 billion Q2 2026 revenue and first positive adjusted operating profit; the risks: constrained US data center construction, overseas low-price model competition, and shrinking secondary-market tolerance. September's formal prospectus is the next hard milestone.

Aug 26, 20266 min read