Field SOP
Field SOP

How to Call Claude Sonnet 5.5: Three Clouds, One Price List

A hands-on SOP for calling Claude Sonnet 5.5 (basis: October 5, 2026). It opens with the compliance red line: Anthropic does not directly offer claude.ai signup or API keys to individuals in mainland China, and this guide provides no circumvention methods (no proxies, virtual cards or SMS-verification workarounds) - only three legitimate routes: direct access from officially supported regions (list per Anthropic's live page), enterprise onboarding through the three clouds where the model shipped day one (AWS Bedrock, Google Cloud Vertex AI, Microsoft Azure), and legal alternatives for mainland developers (GLM and Kimi open weights and APIs via ModelScope and official channels). Practical coverage: the model ID claude-sonnet-5-5 (all lowercase with hyphens - the easiest spelling trap), snapshot pricing $2/$10 with $0.20 cache reads (savings come from token efficiency, not price cuts), the zero-data-retention option, the official division of labor versus Opus 5.5 (daily tasks to Sonnet, complex long tasks to Opus), and same-day Claude Code integration. Five steps each with expected outcomes: pick your entry, provision and verify the key, first call, parameters and caching, and the two-model split. All pricing and capability figures are snapshots; the official page governs.

Published October 5, 20269 min read
<!-- claude-sonnet-5-5-api-sop | sop | How to Call Claude Sonnet 5.5: Three Clouds, One Price List -->

Before You Write Any Code: Pick a Compliant Path

This SOP answers one concrete question: how to call Claude Sonnet 5.5 through the API, from zero to your first successful response, with the expected outcome of every step spelled out. Before the first request goes out, though, there is a gate question that matters more than any code: is your access path compliant? This section is the most important part of the article - read it fully first.

State the position plainly. Anthropic has not opened claude.ai registration or API key issuance directly to individuals in mainland China. This article does not provide, and will not teach, any method for circumventing regional restrictions: no proxies, no virtual cards, no SMS-verification workarounds, and no hint that such routes are viable. What it offers instead is the legitimate set of paths.

Path one: direct access where officially supported. If your region is on Anthropic's list of supported locations, register on the official platform and request an API key. The current supported-regions list lives on Anthropic's official pages; we do not transcribe or guess it, because these lists change and a stale copy misleads.

Path two: the three hyperscale clouds under enterprise agreements. AWS Bedrock, Google Cloud Vertex AI, and Microsoft Azure all made claude-sonnet-5-5 available in sync with Anthropic's own platform - officially published channels. Enterprises enable the model under each cloud's own account system, compliance terms, and billing rules. There is no shortcut to find, and none should be sought.

Path three: legal alternatives for mainland China developers. If you are an individual without enterprise cloud resources, the legitimate route is open-weight and domestic API models. Zhipu's GLM family and Moonshot's Kimi family publish open weights and offer APIs obtainable legally through ModelScope and official channels; our GLM-5.3 deep dive covers that flagship in detail. For many individuals this path costs less than any gray-area workaround, with zero risk.

Every step below assumes path one or two. If path three is your destination, jump to the FAQ and the link above.

On timing: claude-sonnet-5-5 was announced September 28, 2026, Eastern Time, as the second model of the Claude 5.5 series, following Opus 5.5 by a little over a week. All claims here are current as of October 5, 2026, every price is a snapshot, and the official pages always prevail.

Know the Model: ID, Pricing, and Official Positioning

Memorize the identity card first. The model ID is claude-sonnet-5-5, all lowercase with hyphens - not claude-sonnet-5.5 and not sonnet-5.5. Cloud vendors sometimes append version or region suffixes, so check each provider's model catalog before your first call; a spelling mismatch is the most common first-call failure.

Pricing, official snapshot: 2 dollars per million input tokens, 10 dollars per million output tokens, 20 cents per million cache-read tokens - unchanged from Sonnet 5. The per-token rate did not move, but Anthropic claims output speed up more than 30 percent over Sonnet 5 and cost per task down up to 30 percent for most workloads - the savings come from fewer tokens, not a cheaper rate. Where that price sits against rivals is what our late-September coding model price comparison covered; this article does not repeat the review, it solves the calling problem.

The capability numbers a caller needs: Terminal-Bench 4.0 at 70.6 percent against Sonnet 5's 10.3; CursorBench 4.0 at 55.5 percent, close to Opus 5.5's 57.8; OSWorld 2.1 up from 57.0 to 80.1 percent. In caller's language: on coding tasks that really execute in a terminal, this Sonnet stepped up a tier, and as a coding-agent backbone the gap to Anthropic's own flagship is hard to feel in daily work. Full release background sits in our Claude Sonnet 5.5 release breakdown.

The official positioning sentence: Sonnet 5.5 is the faster, lower-cost complement to Opus 5.5, suited to well-scoped everyday tasks, bug fixing, and producing documents, slides, and tables. That sentence drives the model split in Step 4.

Step 1: Choose Your Enablement Entry Point

The three legitimate paths map to three kinds of entry points. Pick one before moving on.

  1. Direct access, for individuals and teams in officially supported regions. Entry point: the console on Anthropic's official platform. Actions: register, complete identity and payment information, request an API key. Expected result: the console shows your key - usually displayed only once, so save it immediately - and the account page shows quota and rate tier.
  2. AWS Bedrock, for enterprises. Entry point: the Bedrock service in the AWS console. Actions: sign in with your corporate AWS account, locate Anthropic models in the Bedrock model catalog, request access for your account or region. Expected result: catalog status turns available and the model detail page describes regional invocation. Menu names follow AWS's official documentation; we do not invent console details.
  3. Google Cloud Vertex AI, for enterprises. Entry point: Vertex AI in the Google Cloud console, models in Model Garden. Actions: sign in with your corporate account, search Claude in Model Garden, enable it for your project. Expected result: the model detail page shows an enableable state and the project binding completes.
  4. Microsoft Azure, for enterprises. Entry point: the model catalog in Azure AI Foundry. Actions: sign in with your corporate Azure account, locate the Claude family, create a deployment as prompted. Expected result: you hold a deployment resource ready for calling.

Three shared reminders. Each cloud runs on its own account system and enterprise agreement, so qualification, regions, and billing differ - read the vendor's official pages first. Same-day availability across clouds and regions is the official claim, but listing pace varies; if the catalog search comes up empty, wait for the official announcement rather than hunting workarounds. And never buy API keys of unclear origin from third parties - the compliance risk dwarfs whatever it saves.

Step 2: Get Your First API Call Through

With a key or deployment in hand, the goal is the minimal loop of request in, answer out - elegance can wait.

  1. Install the official SDK. Direct users take Anthropic's official SDK - the anthropic package in Python or its Node equivalent; cloud users take their cloud's SDK or compatible interface. Expected result: the package installs and imports without errors.
  2. Configure credentials. Direct users put the key in an environment variable - never hard-coded, never committed. Cloud users follow their cloud's standard mechanism. Expected result: the program reads the credential at runtime.
  3. Send the minimal request, with the model parameter set to claude-sonnet-5-5. For direct access:
python
from anthropic import Anthropic

client = Anthropic()  # reads ANTHROPIC_API_KEY from the environment
resp = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Explain API caching in one sentence."}],
)
print(resp.content[0].text)

Expected result: the terminal prints a one-sentence answer. On a model-not-found or permission error, check the ID spelling, then whether the key's account or region has the model enabled, then your SDK version.

  1. Inspect the response. In a Messages API response the text lives in the content array and token usage in the usage field. Expected result: you can read input and output token counts, multiply by the official rates, and know exactly what the call cost. Check usage on every call - runaway bills start with unexamined usage fields.
  2. Establish your baseline. Run the same request once as simple Q&A and once against a real question from your workload, recording latency and tokens. Expected result: first-party baseline data you will use later to judge whether caching pays and whether a model upgrade is warranted.

Two frequent errors: a 429 means rate limiting - back off and retry per official guidance instead of hammering concurrently; for 400-class errors, check the messages format and max_tokens value first, since the request body is usually at fault.

Step 3: Put Caching and Zero Data Retention Into the Cost Math

That 20-cent cache-read rate is the wallet exit many callers walk past.

Caching has a clear use case: long, repeated prefixes. System prompts, reference text, and an agent's tool definitions are identical on every request; once cached, subsequent requests bill at the cache-read rate - one tenth of standard input price. In agent and batch scenarios, cache hit rate directly sets your monthly bill, and stacked with the official 30-percent-per-task saving, long-session applications save real money. Practical rule: invariant content at the front of the prompt, per-request content at the end, so the cache actually fills.

Zero data retention is officially offered with this release: once enabled, inputs and outputs are not used for training and not retained - a hard requirement for compliance-bound enterprises. Two cautions: it is an agreement-level option requested through the enablement process on the official platform or your purchased cloud, with entry point and scope defined by official pages; and enterprises on the three clouds are governed by the cloud vendor's data terms as well, so read both.

For the horizontal picture, our million-token output comparison covers long-task behavior across current models, and the GPT-6.1 Sol API walkthrough is the complete calling path for the other flagship. Read both alongside this article.

Step 4: Split Work Between Sonnet 5.5 and Opus 5.5

The official positioning draws a clean line: everyday tasks, bug fixing, documents, slides, and tables go to Sonnet 5.5, which is faster and cheaper per task; complex long tasks and harder reasoning go to Opus 5.5.

Two decision aids. First, the near-tie on CursorBench 4.0 - 55.5 versus 57.8 percent - says the gap is small for most everyday coding, so picking by budget is safe; GDPval-AA backs the same direction, with Sonnet 5.5 just 2 points below Opus 5.5. Second, Claude Code integrated Sonnet 5.5 on release day - subscribers pick it in the client and write no calling code; API users select it in request parameters. A pragmatic strategy: run Sonnet 5.5 as the default and retry the rare stuck task on Opus 5.5 - under the official positioning, the cheapest combination.

One note for enterprise evaluators: this is the first Sonnet generation carrying network-security protections at Opus level, with high-risk cybersecurity requests falling back automatically to Sonnet 5, plus a new classifier against reasoning extraction. Teams wiring the model into production should put these terms on the evaluation sheet.

If your existing code still runs an older Claude generation, our Claude Fable 5.1 API migration SOP covers the full migration checklist - that article moves you from an old model, this one starts you fresh with 5.5. Pick by your starting point.

FAQ

Q1: Can individual developers in mainland China register for Anthropic's API directly? A1: No. Anthropic has not opened registration or API issuance directly to individuals in mainland China, and this article offers no circumvention methods. Legitimate routes: enterprises enable the model through the three clouds; individuals consider legal APIs from open-weight or domestic models such as GLM and Kimi; users in supported regions register normally. The legitimate paths win on both cost and risk.

Q2: How does calling claude-sonnet-5-5 differ from the older claude-sonnet-5? A2: The model ID differs; prices are unchanged at 2, 10, and 0.2 dollars per million tokens for input, output, and cache; and the official figures claim over 30 percent faster output and up to 30 percent lower cost per task. Whether the old ID stays available follows the official announcements - do not assume it lives forever.

Q3: Do the three clouds charge the same as the official prices, and which is cheapest? A3: The official figures are Anthropic's snapshot rates. Each cloud has its own pricing, discounts, and agreements, and the actual price lives on that cloud's catalog page. This article does not rank clouds on price; enterprises should negotiate against real usage, individuals should start with small pay-as-you-go runs.

Q4: Where do I enable zero data retention? A4: It is an agreement-level enterprise option requested through the enablement process on the official platform or your purchased cloud - not a switch in an ordinary console menu. Entry point, scope, and review requirements follow the official pages; we do not invent details.

Q5: Can I use Sonnet 5.5 without writing code? A5: Yes. Claude Code integrated the model on release day, so subscribers simply select it in the client - everyday coding, bug fixing, and document writing sit squarely in its official comfort zone. The API suits developers embedding the model into their own products.

Join the Discussion

Which of the three clouds did you go with, or did you connect directly? What tripped your first call, and how much has caching shaved off your bill? Leave your enablement path and cost numbers in the comments for the next reader. If you are still torn between flagships, read the price comparison and the million-token output comparison linked above before committing - picking by need beats chasing novelty.


References

  • Anthropic official site and transparency pages (primary source): release date, model ID, pricing, same-day availability across the three clouds, Claude Code integration, the zero data retention option, safety mechanisms, and the supported-regions policy
  • Reuters (2026-09-28): release coverage and enterprise-customer share background
  • TechNews Taiwan (2026-09-29), INSIDE, and Machine Heart: release reporting and benchmark data as relayed

This article is based on official sources and media reports (as of 2026-10-05) and is not an official partnership or promotion. All prices are snapshot values; enablement conditions, region lists, model availability, and live pricing follow the official pages of Anthropic and each cloud vendor in real time.

This article is AI-assisted and human-edited. Last updated: 2026-10-05

FAQ

Can individual developers in mainland China register for Anthropic's API directly?
A1: No. Anthropic has not opened registration or API issuance directly to individuals in mainland China, and this article offers no circumvention methods. Legitimate routes: enterprises enable the model through the three clouds; individuals consider legal APIs from open-weight or domestic models such as GLM and Kimi; users in supported regions register normally. The legitimate paths win on both cost and risk.
How does calling claude-sonnet-5-5 differ from the older claude-sonnet-5?
A2: The model ID differs; prices are unchanged at 2, 10, and 0.2 dollars per million tokens for input, output, and cache; and the official figures claim over 30 percent faster output and up to 30 percent lower cost per task. Whether the old ID stays available follows the official announcements - do not assume it lives forever.
Do the three clouds charge the same as the official prices, and which is cheapest?
A3: The official figures are Anthropic's snapshot rates. Each cloud has its own pricing, discounts, and agreements, and the actual price lives on that cloud's catalog page. This article does not rank clouds on price; enterprises should negotiate against real usage, individuals should start with small pay-as-you-go runs.
Where do I enable zero data retention?
A4: It is an agreement-level enterprise option requested through the enablement process on the official platform or your purchased cloud - not a switch in an ordinary console menu. Entry point, scope, and review requirements follow the official pages; we do not invent details.
Can I use Sonnet 5.5 without writing code?
A5: Yes. Claude Code integrated the model on release day, so subscribers simply select it in the client - everyday coding, bug fixing, and document writing sit squarely in its official comfort zone. The API suits developers embedding the model into their own products.

Related

Field SOP

How to Use Kling AI: A 4.0 Flash Guide with 15 References

A hands-on SOP for Kling AI (basis: October 3, 2026). It opens by aligning expectations: full Kling 4.0 is not yet live, 4.0 Flash is limited to a small-scale preview for black-gold annual members, and pricing and credit costs are unpublished so none are invented. The five-step framework: entry and account (klingai.com or the app; verify membership and 4.0 Flash eligibility); run the minimal text-to-video loop (input, generate, review, download) to set your own baseline; an 8,000-token prompt structure in four parts (tone, storyboard, sound, constraints), borrowing the LTX-2 official prompting method (chronological, action-first, concrete cinematography within 200 words - explicitly labeled as LTX-2 guidance, not Kling documentation); a progressive drill for 15 multimodal references (single subject alignment, images plus subject, add a camera-move video, then push toward the 7-subject cap, with each asset's role spelled out); and a continuation workflow that slices content into 30-second blocks up to 2 minutes, paired with timeline-based continuous creation. Steps three through five are explicitly framed as workflows for when 4.0 reaches your account, not as available today.

Oct 3, 20269 min read
Field SOP

Self-host the open-source Dots in three commands

A self-hosting SOP for the open-source Dots: install uv (Windows PowerShell and Linux routes), then one uvx command brings it up (--openrouter-key required); open 127.0.0.1:8765 for the chat-left, live-browser-right workbench. Swap models with one --model flag across anything on OpenRouter; four advanced parameters explained in depth - --seed (one identity per seed, reproducible), --profile-dir (logins and cookies persist across runs), --proxy (timezone and language follow the exit) and --help (the full option index); invisible_playwright_mcp hands the same browser to Claude Code, Codex and Gemini CLI as an MCP server; the README's own sample task (Milan-to-Lisbon daily fare check, no-ticket means say so) demonstrates the no-fabricated-numbers prompt discipline. A closing compliance section: respect target-site terms and local law, automate only your own business and accounts you are authorized to use, no fraud/scalping/bulk sign-ups. All commands and parameters come verbatim from the README.

Oct 2, 20268 min read
Field SOP

No Ascend Hardware? Three Routes into DeepSeek's Ascend Stack

A hands-on SOP for DeepSeek's Ascend kernel stack along three routes of rising hardware demands: (1) the no-NPU reading route - clone the four repos, diff the ports against the NVIDIA-side originals, study the FlashMLA deep-dive technical report (docs/20260930-ascend-prefill-deep-dive.md), and note CUDA users can stay on the main repos since the fused kernel remains CUDA-only; (2) Ascend deployment - TileKernels installs with pip install tile-kernels (Python 3.12+, PyTorch 2.13+, TileLang 0.1.15+, Ascend 950 NPU + CANN 9.2.0+) or from source via pip install -e ".[dev]", while DeepGEMM-Ascend builds from source (git clone --recursive, ./develop.sh, pip install . --no-build-isolation on CANN 9.20); (3) DeepEP-Ascend for collective communication (git clone plus the DeepJIT submodule, ASCEND_HOME_PATH setup), honestly flagged with its missing LICENSE file and validated-stack boundary (950DT, CANN 9.2.0, torch_npu 2.13.0rc1). Includes a hardware/software requirements table, common pitfalls (scaling-factor packing, even-rank requirement, recursive submodules) and a pre-flight checklist; all commands and versions come verbatim from each repo's README.

Oct 1, 202610 min read