Field SOP
Field SOP

Change One Line of base_url: 5M Free Tokens for Your Agent

A hands-on SOP for wiring LongCat-2.5-Preview into your existing agent toolchain along three routes of rising effort: (1) web-first (sign in at longcat.ai for chat, image upload and simple agent tasks); (2) OpenAI-protocol access (point base_url at https://api.longcat.ai/openai with model ID LongCat-2.5-Preview - existing OpenAI SDKs migrate with zero code changes); (3) Anthropic-protocol access for Claude Code (three env vars: ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN, ANTHROPIC_MODEL), with Codex, OpenClaw, OpenCode and Kilo Code following the same base_url-plus-model-name swap. Includes usage tips for the 1M context and 128K output, the 5M-free-tokens offer (relayed basis), and Preview-stage caveats (no public benchmarks, unreleased weights, interface may change); validate price, latency and success rate on a small traffic slice before production.

Published September 29, 202610 min read
<!-- longcat-2-5-agent-integration-sop | sop | Change One Line of base_url: 5M Free Tokens for Your Agent -->

Meituan released its new-generation large model LongCat-2.5-Preview at the end of September (per the official announcement as relayed by AI tool coverage, dated 2026-09-28). Two spec points stand out. First, scale: a sparse-activation MoE architecture with 1.6T total parameters, of which only about 48B are activated per inference. Second, window: a native 1M-token context with a maximum output of 128K tokens. On the multimodal side, image understanding has been added and natively merged into the base model, following the DiNA (discrete native autoregressive) paradigm validated by LongCat-Next, which lexicalizes each modality into discrete tokens for unified modeling (relayed reporting). The product positioning carries on the Agentic Coding DNA of the LongCat line, with official target scenarios pointing at long-horizon autonomous operation across terminals, browsers, GUIs, spreadsheets and design tools.

That is the press-release view. For developers who want to wire the model into an existing agent toolchain, only three questions matter: what do the interfaces look like, how much code changes, and where are the traps. The answer from LongCat-2.5-Preview is that its API is compatible with both the OpenAI and Anthropic protocols, which means most mainstream agent tools can adopt it by changing configuration rather than code, pushing the migration cost down to the config-file level. This SOP splits the work into three routes with rising barriers: direct web experience, OpenAI-protocol integration, and Anthropic-protocol integration with Claude Code, each with a minimal copy-paste configuration.

Two things before we start. First, the division of labor: our earlier piece on Meituan LongCat 2 and the domestic-chip training line covers the training-side story of LongCat 2; this article is strictly about integrating the 2.5-Preview model itself, and the two pieces complement each other. Second, a piece of cold thinking that has to come first: LongCat-2.5-Preview currently has no public benchmark scores, and the official comparison table explicitly marks it as having no public benchmark. This article will therefore not cite any test scores; the only thing you can rely on is hands-on evaluation against your own tasks. More on this later.

Step Zero: Registration, API Key and the 5M-Token Bonus

All three routes share the same starting point. Register at the LongCat platform (longcat.ai/platform) and create an API key. Billing works through token packages, and new registrations receive 5 million free tokens (relayed reporting; the platform's actual rules govern).

What does 5 million tokens mean in practice? At a full 1M context, that is roughly five full-window calls. But for getting the integration running, verifying protocol compatibility, and running a small-scale evaluation against your own task set, it is plenty. Spend the free quota on evaluation first rather than production traffic, because evaluation results are what carry over, and production scaling is when you pick a package. One usage detail worth flagging: agent tasks consume far more tokens than single-turn chat, and a single tool-loop task burning tens of thousands of tokens is normal, so estimate consumption on agent assumptions, not chat assumptions.

Key hygiene follows the usual rules: never commit keys to repositories, never print them to logs, isolate per environment, rotate regularly, revoke immediately on suspicion of leakage. Every key placeholder in the examples below should be replaced with your own secret and injected through environment variables or a secrets manager, never hardcoded.

Route Selection: Match Yourself to a Lane

RouteBest forBarrierTypical use
Direct web experienceEveryoneLowest, open the site and goCapability scouting, image tasks, small trials
OpenAI-protocol integrationTeams with existing OpenAI SDK codeLow, change two linesWiring the model into an existing agent pipeline
Anthropic-protocol integrationClaude Code usersLow, set three env varsTerminal coding and long-horizon agent tasks

The decision rule: if you just want to see what the model can do, take Route 1; if your existing system runs on the OpenAI SDK or any OpenAI-compatible library, take Route 2; if Claude Code is your daily driver, take Route 3. Routes 2 and 3 are not mutually exclusive. Dual protocol means the same model ID can serve two toolchains at once, and many teams will run both and split traffic by task type.

Route 1: The Web, Zero-Configuration Scouting

The LongCat website (longcat.ai) offers a chat entry point that works right after registration. It accepts image uploads, so you can run simple agent-style tasks: paste an error screenshot and ask it to locate the bug, upload a table screenshot and have it organize the data, hand it a design mock and ask for a frontend page description. Image understanding is a structural change relative to the 2.0 era, natively merged into the base model instead of living in a separate repository, and the web client lets you verify that capability directly.

Be clear about the boundary of this form, though: for long-horizon autonomous operation, the web client can only verify so much, and real engineering validation belongs to the next two routes. Also note that web behavior and API behavior may differ; a good showing in the browser does not automatically translate to the API, and the final verdict comes from hands-on interface testing. The deliverable of this step should be a small task set: pick five to ten of the most typical tasks from your business and record how the web client handles them. When you move to the API, this becomes the seed of your evaluation set, which also gives your free-token spending a clear priority.

Route 2: OpenAI Protocol, a Two-Line Change

This is the route with the most existing code and the smallest change. The OpenAI-protocol endpoint is base_url=https://api.longcat.ai/openai, and the model ID is LongCat-2.5-Preview. If your project already uses the official OpenAI SDK or any compatible library, in theory you change exactly two things: the base_url and the API key, then set the model name to LongCat-2.5-Preview. A minimal Python example:

python
from openai import OpenAI

# Change only these two lines: base_url to the LongCat endpoint, api_key to your platform key
client = OpenAI(
    api_key="YOUR LONGCAT API KEY",
    base_url="https://api.longcat.ai/openai"
)

# Copy the model ID verbatim, watch the casing and hyphens
resp = client.chat.completions.create(
    model="LongCat-2.5-Preview",
    messages=[
        {"role": "user", "content": "Break this requirement into dev tasks with acceptance criteria"}
    ]
)
print(resp.choices[0].message.content)

Three practical reminders. First, the model ID is LongCat-2.5-Preview, copied verbatim; do not invent shorthand or aliases during the Preview stage. Second, standard OpenAI request fields such as tool calling should be used per the officially compatible surface, with actual field behavior subject to the official documentation; verify each one before launch instead of running on assumptions. Third, because this is a Preview, centralize the model ID and base_url in your configuration layer rather than scattering them through business code. Future model or endpoint changes will be a one-file diff.

Once the integration runs, resist pushing to production immediately. Run a small-traffic evaluation using the task set from step zero, focusing on three things: tool-calling success rate, instruction following under long context, and output format stability. These are where agent scenarios break most often and where models differ the most, and your own measurements beat any press release.

Route 3: Anthropic Protocol with Claude Code

If you are a Claude Code user, this route feels closest to a seamless switch. Per the official tutorial, three environment variables are all it takes:

bash
# Anthropic-protocol endpoint
export ANTHROPIC_BASE_URL="https://api.longcat.ai/anthropic"
# Platform key
export ANTHROPIC_AUTH_TOKEN="YOUR LONGCAT API KEY"
# Model ID
export ANTHROPIC_MODEL="LongCat-2.5-Preview"

claude

Each variable has one job: ANTHROPIC_BASE_URL points requests at LongCat's Anthropic-protocol endpoint, ANTHROPIC_AUTH_TOKEN carries the platform key, and ANTHROPIC_MODEL selects the model. Once configured, Claude Code traffic flows to LongCat-2.5-Preview, and tool calling, file editing and long sessions all work through the existing Claude Code workflow with no change in habits.

The first thing after switching is a smoke test: run a small task that involves tool calling and confirm the requests actually reach the new endpoint instead of a stale default config. The most direct evidence is a change in the platform-side usage stats, paired with the response behavior as a second signal. This takes a minute or two and prevents the classic accident of believing you switched when you did not.

The official tutorial also names a batch of peer tools: Codex, OpenClaw, OpenCode, Kilo Code, Hermes and others follow the same logic of changing the base_url and model name. Quick reference:

ToolChange required (official tutorial)
Claude CodeSet the three env vars ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN, ANTHROPIC_MODEL
Codex, OpenClaw, OpenCode, Kilo Code, Hermes and othersChange base_url and model name; use the endpoints and model ID above

When writing these variables into shell config files, mind the scope: prefer per-project or per-session injection over a global hardcode, so you can flip between the official model and LongCat at any time and run side-by-side comparisons. If your agent stack is still in selection, we have prior coverage to lean on: coding CLI tools in the MiniMax Code CLI resource piece, the open-source Codex harness ecosystem in this hotspot, DeepSeek-line harnesses in this breakdown, and the coding IDE landscape in this comparison.

Using 1M Context and 128K Output Without Wasting Them

The window spec is the part of 2.5-Preview most worth a dedicated section, but a big window does not mean stuffing everything in, and doing it wrong is both expensive and ineffective.

Start with where 1M context actually pays off in agent scenarios. Three cases: loading an entire repository, or most of it, in one shot for cross-file understanding; preserving full session trajectories for long-horizon tasks; and merging large batches of documents for consolidated analysis. Recognize the cost at the same time: every token you stuff in is billed, and noise unrelated to the task dilutes the model's attention on what matters, degrading agent decision quality. The operating principle is to trim context per task and treat 1M as a ceiling, not a target; only genuinely window-hungry scenarios deserve a full load.

Then the 128K maximum output. A single response of 128K tokens means large-file generation and long reports can land in one pass, and the truncation risk when an agent writes multi-file projects drops sharply. But longer output means higher human review cost, so agree on an output structure in the system prompt and make long outputs reviewable in sections rather than one monolithic block.

A cost reference: per the comparison table in our coverage, Claude Opus 5.5 is priced at 4 dollars per million input tokens and 20 dollars per million output tokens (about 29 and 143 yuan), with the same 1M context and 128K maximum output, but available only over the Anthropic protocol. LongCat-2.5-Preview bills through token packages with specifics on the platform page, so estimate a bill against your projected call volume before deciding how to split evaluation and production traffic.

Current Limitations: Three Things to Keep in Mind

First, no public benchmark. This is the cold thinking repeated throughout the piece: the official comparison table marks LongCat-2.5-Preview as having no public benchmark, and this article strictly honors that by citing no scores. For you, it means every claim of the form "it beats model X" is untrustworthy, including qualitative statements in the announcement. The only reliable method is evaluating on your own task set, and the 5 million free tokens are exactly enough for that.

Second, the weights are not released. As of 2026-09-29, a direct check of 30 repositories under the meituan-longcat GitHub organization found no LongCat-2.5-Preview repository; the organization's latest pushes are LongCat-DeepResearch (9 stars) and WBench (240 stars) from September 24, and the existing model repositories are LongCat-2.0 (565 stars, MIT), LongCat-Flash-Chat (1367 stars, MIT) and LongCat-Video (8439 stars, MIT), among others. In other words, 2.5-Preview is API-only right now; teams hoping for local deployment or weight-based fine-tuning have no path yet and can only watch for the next move. For the broader open-weights flagship landscape, see our open-source flagship comparison.

Third, everything can change during Preview. Model IDs, endpoints, parameter behavior and pricing packages may all be adjusted. The engineering countermeasure is simple: converge all LongCat-related configuration into one config module, have business code depend only on an abstract interface, and make every model upgrade or rollback a one-place change. This applies to any Preview model, not just LongCat, but it is the difference between a three-minute and a three-day integration when the change lands.

FAQ

Q1: How does LongCat-2.5-Preview score on benchmarks?

A1: There are no public benchmark scores; the official comparison table marks this explicitly, and any third-party score you see quoted is untrustworthy. Use the free registration tokens to evaluate on your own task set, focusing on tool-calling success rate and instruction following under long context.

Q2: Can I download the weights or deploy locally?

A2: Not as of 2026-09-29. The meituan-longcat organization has no 2.5-Preview repository, and the model is API-only. What is open-sourced there includes LongCat-2.0, LongCat-Flash-Chat and LongCat-Video, all under the MIT license.

Q3: How do I get the 5 million free tokens, and are they enough?

A3: Register at longcat.ai/platform to receive them (relayed reporting; the platform's rules govern). They are enough to get both protocols running and complete a small-scale evaluation; at a full 1M context that is roughly five full-window calls, so spend them on evaluation first.

Q4: How much code do I change in an existing OpenAI SDK project?

A4: In theory, two lines: set base_url to https://api.longcat.ai/openai, swap in your LongCat platform key, and set the model name to LongCat-2.5-Preview. Before launch, regress tool calling and long-context scenarios with real tasks to confirm the compatible behavior meets expectations.

Q5: Is more 1M context always better?

A5: No. Every token is billed, and irrelevant content dilutes attention and drags down agent decision quality. Treat 1M as a ceiling rather than a target, trim context per task, and fill the window only for genuinely long-horizon work such as whole-repo reasoning and extended sessions.

This article is AI-assisted and human-edited. Last updated: 2026-09-29

FAQ

How does LongCat-2.5-Preview score on benchmarks?
There are no public benchmark scores; the official comparison table marks this explicitly, and any third-party score you see quoted is untrustworthy. Use the free registration tokens to evaluate on your own task set, focusing on tool-calling success rate and instruction following under long context.
Can I download the weights or deploy locally?
Not as of 2026-09-29. The meituan-longcat organization has no 2.5-Preview repository, and the model is API-only. What is open-sourced there includes LongCat-2.0, LongCat-Flash-Chat and LongCat-Video, all under the MIT license.
How do I get the 5 million free tokens, and are they enough?
Register at longcat.ai/platform to receive them (relayed reporting; the platform's rules govern). They are enough to get both protocols running and complete a small-scale evaluation; at a full 1M context that is roughly five full-window calls, so spend them on evaluation first.
How much code do I change in an existing OpenAI SDK project?
In theory, two lines: set base_url to https://api.longcat.ai/openai, swap in your LongCat platform key, and set the model name to LongCat-2.5-Preview. Before launch, regress tool calling and long-context scenarios with real tasks to confirm the compatible behavior meets expectations.
Is more 1M context always better?
No. Every token is billed, and irrelevant content dilutes attention and drags down agent decision quality. Treat 1M as a ceiling rather than a target, trim context per task, and fill the window only for genuinely long-horizon work such as whole-repo reasoning and extended sessions.

Related

Field SOP

MiMo-V2.6 Integration SOP: Desktop, API and Local Weights

A hands-on SOP for wiring MiMo-V2.6 into your workflow along three routes of rising effort: (1) the MiMo Desktop client (install, sign in or plug in your own API key, switch MiMo-V2.6-Pro/Flash in the model list, UltraSpeed mode, screenshot-feedback iteration); (2) the MiMo open-platform API (create an app for a key, pass model name, messages, tools and multimodal inputs; for agent tasks, wire environment logs, test pass rates, screenshots and verifier feedback into an execute-check-correct loop; validate price, latency and success rate on a small traffic slice before production); (3) local weights (download from the HF collection collections/XiaomiMiMo/mimo-v26; parameter counts are unpublished, so hardware floors defer to the model cards; for RL reproduction, go through the five environment scripts in XiaomiMiMo/verl). Includes a pitfall table and a pre-launch checklist; API pricing defers to the platform documentation.

Sep 27, 202610 min read
Field SOP

Kimi Dual Protocol: One Config for Codex and Claude Code

Moonshot announced on 2026-09-02 that the Kimi API natively supports dual protocols: OpenAI Responses (api.moonshot.cn/v1) plus Anthropic Messages (api.moonshot.cn/anthropic), with kimi-k3 as the flagship model. Hands-on SOP: point Claude Code's ~/.claude/settings.json ANTHROPIC_BASE_URL to /anthropic with model kimi-k3[1m]; set Codex's ~/.codex/config.toml wire_api="responses". This turns Kimi into a unified model-routing gateway — switch the backend without touching client code. Boundaries: Responses is text+image only, kimi-k2.7-code forces thinking, and the old ANTHROPIC_API_KEY must be removed.

Sep 5, 202611 min read