Before You Write Any Code: Pick a Compliant Path
This SOP answers one concrete question: how to call Claude Sonnet 5.5 through the API, from zero to your first successful response, with the expected outcome of every step spelled out. Before the first request goes out, though, there is a gate question that matters more than any code: is your access path compliant? This section is the most important part of the article - read it fully first.
State the position plainly. Anthropic has not opened claude.ai registration or API key issuance directly to individuals in mainland China. This article does not provide, and will not teach, any method for circumventing regional restrictions: no proxies, no virtual cards, no SMS-verification workarounds, and no hint that such routes are viable. What it offers instead is the legitimate set of paths.
Path one: direct access where officially supported. If your region is on Anthropic's list of supported locations, register on the official platform and request an API key. The current supported-regions list lives on Anthropic's official pages; we do not transcribe or guess it, because these lists change and a stale copy misleads.
Path two: the three hyperscale clouds under enterprise agreements. AWS Bedrock, Google Cloud Vertex AI, and Microsoft Azure all made claude-sonnet-5-5 available in sync with Anthropic's own platform - officially published channels. Enterprises enable the model under each cloud's own account system, compliance terms, and billing rules. There is no shortcut to find, and none should be sought.
Path three: legal alternatives for mainland China developers. If you are an individual without enterprise cloud resources, the legitimate route is open-weight and domestic API models. Zhipu's GLM family and Moonshot's Kimi family publish open weights and offer APIs obtainable legally through ModelScope and official channels; our GLM-5.3 deep dive covers that flagship in detail. For many individuals this path costs less than any gray-area workaround, with zero risk.
Every step below assumes path one or two. If path three is your destination, jump to the FAQ and the link above.
On timing: claude-sonnet-5-5 was announced September 28, 2026, Eastern Time, as the second model of the Claude 5.5 series, following Opus 5.5 by a little over a week. All claims here are current as of October 5, 2026, every price is a snapshot, and the official pages always prevail.
Know the Model: ID, Pricing, and Official Positioning
Memorize the identity card first. The model ID is claude-sonnet-5-5, all lowercase with hyphens - not claude-sonnet-5.5 and not sonnet-5.5. Cloud vendors sometimes append version or region suffixes, so check each provider's model catalog before your first call; a spelling mismatch is the most common first-call failure.
Pricing, official snapshot: 2 dollars per million input tokens, 10 dollars per million output tokens, 20 cents per million cache-read tokens - unchanged from Sonnet 5. The per-token rate did not move, but Anthropic claims output speed up more than 30 percent over Sonnet 5 and cost per task down up to 30 percent for most workloads - the savings come from fewer tokens, not a cheaper rate. Where that price sits against rivals is what our late-September coding model price comparison covered; this article does not repeat the review, it solves the calling problem.
The capability numbers a caller needs: Terminal-Bench 4.0 at 70.6 percent against Sonnet 5's 10.3; CursorBench 4.0 at 55.5 percent, close to Opus 5.5's 57.8; OSWorld 2.1 up from 57.0 to 80.1 percent. In caller's language: on coding tasks that really execute in a terminal, this Sonnet stepped up a tier, and as a coding-agent backbone the gap to Anthropic's own flagship is hard to feel in daily work. Full release background sits in our Claude Sonnet 5.5 release breakdown.
The official positioning sentence: Sonnet 5.5 is the faster, lower-cost complement to Opus 5.5, suited to well-scoped everyday tasks, bug fixing, and producing documents, slides, and tables. That sentence drives the model split in Step 4.
Step 1: Choose Your Enablement Entry Point
The three legitimate paths map to three kinds of entry points. Pick one before moving on.
- Direct access, for individuals and teams in officially supported regions. Entry point: the console on Anthropic's official platform. Actions: register, complete identity and payment information, request an API key. Expected result: the console shows your key - usually displayed only once, so save it immediately - and the account page shows quota and rate tier.
- AWS Bedrock, for enterprises. Entry point: the Bedrock service in the AWS console. Actions: sign in with your corporate AWS account, locate Anthropic models in the Bedrock model catalog, request access for your account or region. Expected result: catalog status turns available and the model detail page describes regional invocation. Menu names follow AWS's official documentation; we do not invent console details.
- Google Cloud Vertex AI, for enterprises. Entry point: Vertex AI in the Google Cloud console, models in Model Garden. Actions: sign in with your corporate account, search Claude in Model Garden, enable it for your project. Expected result: the model detail page shows an enableable state and the project binding completes.
- Microsoft Azure, for enterprises. Entry point: the model catalog in Azure AI Foundry. Actions: sign in with your corporate Azure account, locate the Claude family, create a deployment as prompted. Expected result: you hold a deployment resource ready for calling.
Three shared reminders. Each cloud runs on its own account system and enterprise agreement, so qualification, regions, and billing differ - read the vendor's official pages first. Same-day availability across clouds and regions is the official claim, but listing pace varies; if the catalog search comes up empty, wait for the official announcement rather than hunting workarounds. And never buy API keys of unclear origin from third parties - the compliance risk dwarfs whatever it saves.
Step 2: Get Your First API Call Through
With a key or deployment in hand, the goal is the minimal loop of request in, answer out - elegance can wait.
- Install the official SDK. Direct users take Anthropic's official SDK - the anthropic package in Python or its Node equivalent; cloud users take their cloud's SDK or compatible interface. Expected result: the package installs and imports without errors.
- Configure credentials. Direct users put the key in an environment variable - never hard-coded, never committed. Cloud users follow their cloud's standard mechanism. Expected result: the program reads the credential at runtime.
- Send the minimal request, with the model parameter set to claude-sonnet-5-5. For direct access:
from anthropic import Anthropic
client = Anthropic() # reads ANTHROPIC_API_KEY from the environment
resp = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Explain API caching in one sentence."}],
)
print(resp.content[0].text)Expected result: the terminal prints a one-sentence answer. On a model-not-found or permission error, check the ID spelling, then whether the key's account or region has the model enabled, then your SDK version.
- Inspect the response. In a Messages API response the text lives in the content array and token usage in the usage field. Expected result: you can read input and output token counts, multiply by the official rates, and know exactly what the call cost. Check usage on every call - runaway bills start with unexamined usage fields.
- Establish your baseline. Run the same request once as simple Q&A and once against a real question from your workload, recording latency and tokens. Expected result: first-party baseline data you will use later to judge whether caching pays and whether a model upgrade is warranted.
Two frequent errors: a 429 means rate limiting - back off and retry per official guidance instead of hammering concurrently; for 400-class errors, check the messages format and max_tokens value first, since the request body is usually at fault.
Step 3: Put Caching and Zero Data Retention Into the Cost Math
That 20-cent cache-read rate is the wallet exit many callers walk past.
Caching has a clear use case: long, repeated prefixes. System prompts, reference text, and an agent's tool definitions are identical on every request; once cached, subsequent requests bill at the cache-read rate - one tenth of standard input price. In agent and batch scenarios, cache hit rate directly sets your monthly bill, and stacked with the official 30-percent-per-task saving, long-session applications save real money. Practical rule: invariant content at the front of the prompt, per-request content at the end, so the cache actually fills.
Zero data retention is officially offered with this release: once enabled, inputs and outputs are not used for training and not retained - a hard requirement for compliance-bound enterprises. Two cautions: it is an agreement-level option requested through the enablement process on the official platform or your purchased cloud, with entry point and scope defined by official pages; and enterprises on the three clouds are governed by the cloud vendor's data terms as well, so read both.
For the horizontal picture, our million-token output comparison covers long-task behavior across current models, and the GPT-6.1 Sol API walkthrough is the complete calling path for the other flagship. Read both alongside this article.
Step 4: Split Work Between Sonnet 5.5 and Opus 5.5
The official positioning draws a clean line: everyday tasks, bug fixing, documents, slides, and tables go to Sonnet 5.5, which is faster and cheaper per task; complex long tasks and harder reasoning go to Opus 5.5.
Two decision aids. First, the near-tie on CursorBench 4.0 - 55.5 versus 57.8 percent - says the gap is small for most everyday coding, so picking by budget is safe; GDPval-AA backs the same direction, with Sonnet 5.5 just 2 points below Opus 5.5. Second, Claude Code integrated Sonnet 5.5 on release day - subscribers pick it in the client and write no calling code; API users select it in request parameters. A pragmatic strategy: run Sonnet 5.5 as the default and retry the rare stuck task on Opus 5.5 - under the official positioning, the cheapest combination.
One note for enterprise evaluators: this is the first Sonnet generation carrying network-security protections at Opus level, with high-risk cybersecurity requests falling back automatically to Sonnet 5, plus a new classifier against reasoning extraction. Teams wiring the model into production should put these terms on the evaluation sheet.
If your existing code still runs an older Claude generation, our Claude Fable 5.1 API migration SOP covers the full migration checklist - that article moves you from an old model, this one starts you fresh with 5.5. Pick by your starting point.
FAQ
Q1: Can individual developers in mainland China register for Anthropic's API directly? A1: No. Anthropic has not opened registration or API issuance directly to individuals in mainland China, and this article offers no circumvention methods. Legitimate routes: enterprises enable the model through the three clouds; individuals consider legal APIs from open-weight or domestic models such as GLM and Kimi; users in supported regions register normally. The legitimate paths win on both cost and risk.
Q2: How does calling claude-sonnet-5-5 differ from the older claude-sonnet-5? A2: The model ID differs; prices are unchanged at 2, 10, and 0.2 dollars per million tokens for input, output, and cache; and the official figures claim over 30 percent faster output and up to 30 percent lower cost per task. Whether the old ID stays available follows the official announcements - do not assume it lives forever.
Q3: Do the three clouds charge the same as the official prices, and which is cheapest? A3: The official figures are Anthropic's snapshot rates. Each cloud has its own pricing, discounts, and agreements, and the actual price lives on that cloud's catalog page. This article does not rank clouds on price; enterprises should negotiate against real usage, individuals should start with small pay-as-you-go runs.
Q4: Where do I enable zero data retention? A4: It is an agreement-level enterprise option requested through the enablement process on the official platform or your purchased cloud - not a switch in an ordinary console menu. Entry point, scope, and review requirements follow the official pages; we do not invent details.
Q5: Can I use Sonnet 5.5 without writing code? A5: Yes. Claude Code integrated the model on release day, so subscribers simply select it in the client - everyday coding, bug fixing, and document writing sit squarely in its official comfort zone. The API suits developers embedding the model into their own products.
Join the Discussion
Which of the three clouds did you go with, or did you connect directly? What tripped your first call, and how much has caching shaved off your bill? Leave your enablement path and cost numbers in the comments for the next reader. If you are still torn between flagships, read the price comparison and the million-token output comparison linked above before committing - picking by need beats chasing novelty.
References
- Anthropic official site and transparency pages (primary source): release date, model ID, pricing, same-day availability across the three clouds, Claude Code integration, the zero data retention option, safety mechanisms, and the supported-regions policy
- Reuters (2026-09-28): release coverage and enterprise-customer share background
- TechNews Taiwan (2026-09-29), INSIDE, and Machine Heart: release reporting and benchmark data as relayed
This article is based on official sources and media reports (as of 2026-10-05) and is not an official partnership or promotion. All prices are snapshot values; enablement conditions, region lists, model availability, and live pricing follow the official pages of Anthropic and each cloud vendor in real time.