Frontline Hotspot
Frontline Hotspot

1.6T Params, 48B Awake: What Meituan LongCat-2.5 Hides

Meituan shipped LongCat-2.5-Preview on 2026-09-28 (official basis relayed via AI-tool aggregators): the 1.6T-total / ~48B-active MoE and 1M-token context carry over, max output is 128K tokens; the headline change is native image understanding merged into the base model (following the DiNA discrete-native-autoregressive route validated by LongCat-Next), plus long-horizon agent tasks across terminals, browsers, GUIs, spreadsheets and design tools. The API speaks both OpenAI and Anthropic protocols, so existing agent tools connect by swapping base_url, and sign-ups get 5M free tokens (relayed basis). Cold thought: no public benchmarks (against Opus 5.5 at $4/$20 per million tokens with public evals), and as of 2026-09-29 weights are unreleased with no 2.5 repo in the GitHub org - Preview-stage figures defer to official disclosure.

Published September 29, 20269 min read
<!-- longcat-2-5-preview-hotspot | hotspot | 1.6T Params, 48B Awake: What Meituan LongCat-2.5 Hides -->

On September 28, 2026, Meituan launched LongCat-2.5-Preview, the newest entry in its LongCat model family (official figures as relayed by the Chinese AI tools directory AI-bot.cn). The most eye-catching line on the spec sheet reads: 1.6 trillion total parameters, with only about 48 billion activated per inference. A trillion-scale model in which just 48B "wakes up" for any given request is textbook MoE (mixture of experts) sparse architecture: enormous total capacity for knowledge, kept affordable by activating only a small slice of it at a time. Meituan has pushed both ends of that trade-off to a fairly aggressive extreme.

But the spec sheet is only the opening act. Three things in this Preview deserve real attention: multimodality merged natively into the base model, a positioning built around long-horizon GUI agents, and an API that speaks both the OpenAI and Anthropic protocols at once. Just as worth recording, though, are two things the release did not do: publish any benchmark scores, and release the model weights. This article walks through all five points, taking the specs at face value while keeping the conclusions deliberately cool.

One more thing to establish up front: the word "Preview" in the name is not decoration. The official positioning describes it as a preview of a new generation model, continuing the architectural line of LongCat-2.0 (official figures as relayed by AI-bot.cn). A preview means the shape is not final, that a full release with more capabilities may follow, and that any long-term judgment drawn today would be premature. First, what it gives; then, what it withholds.

1.6T Total, 48B Active: The Sparse Arithmetic

Start with the architecture math. The core idea of MoE is to split a giant model into many "expert" modules and wake only a small subset for each request. Total parameters determine how much the model knows; activated parameters determine how much it uses per call. LongCat-2.5-Preview pairs 1.6T total parameters with 48B active (official figures as relayed by AI-bot.cn), which sits at the most aggressive end of the current spectrum: capacity aimed at the trillion-parameter club, inference cost kept closer to a ten-billion-scale model.

For users, the practical meaning is simple: you do not pay a trillion-parameter inference bill for trillion-parameter knowledge. The context specs are equally striking: a native 1M token window with a maximum output of 128K tokens (official figures as relayed by AI-bot.cn). A 1M window means an entire code repository, hundreds of pages of documentation, or batches of web content can be placed in context in one shot, which is infrastructure-level groundwork for the long agent workflows discussed below.

Sparse architecture is not a free lunch, however. The wider the gap between total and activated parameters, the higher the demands on routing: which experts get woken, whether load stays balanced, and how much knowledge overlaps between experts all shape real-world behavior. The official materials do not open up these implementation details, so the honest status is "route confirmed, internals undisclosed." For ordinary users, the more practical check is consistency across repeated requests on the same task, because the variance of sparse activation tends to show up in long-tail requests rather than in peak benchmark runs.

Multimodal Merged Into the Base: The DiNA Route

In the 2.0 era, LongCat's image and video capabilities lived in separate repositories. With 2.5-Preview, image understanding has been merged natively into the base model (official figures as relayed by AI-bot.cn). On the technical route, the multimodal stack follows the DiNA (discrete native autoregressive) paradigm validated by LongCat-Next: visual and other modalities are lexicalized into discrete tokens and modeled jointly with text (relayed figures).

The value of that approach is that multimodality stops being a bolt-on "encoder plus adapter" and instead shares one vocabulary and one inference path with text. For agent scenarios this matters a great deal: understanding a screenshot, a spreadsheet, or a design draft, and understanding a natural-language instruction, should happen inside a single model rather than across a chain of stitched-together ones. Every extra hop in a stitched pipeline adds information loss and one more failure point; native unified modeling folds the whole perception-to-decision chain into a single forward pass.

There is also an easily overlooked engineering dividend: once vision lives in the base model, it can reuse the text-side infrastructure of long context, tool calling, and streaming output, without a separate serving chain maintained for images. For API users this means sending an image and sending text follow the same interface conventions, which lowers integration cost. That is why "multimodal merged into the base" deserves its place in the headline claims rather than one line in a feature list.

1M Context and GUI Agents: Built for Long Workflows

In terms of positioning, LongCat-2.5-Preview explicitly targets long-horizon autonomous operation across terminals, browsers, GUIs, spreadsheets, and design tools, continuing its Agentic Coding direction (official figures as relayed by AI-bot.cn). A long workflow means dozens or even hundreds of consecutive steps: open a page, locate a control, fill a form, verify the result, notice something wrong, roll back and retry. That chain demands far more than single-turn question answering: instruction following, visual grounding, state memory, and error recovery all have to hold, because one broken link kills the run.

GUI agents have become contested ground among flagship models because they map to the most direct commercial imagination: having a model operate software on a person's behalf, not merely answer questions. By writing long-horizon GUI operation into its product positioning, LongCat-2.5-Preview has publicly lined up at the front of that race. Positioning is positioning, though, and real-world performance is something else: no public evaluation yet proves its rank on this track, as discussed below.

It is worth stressing how compound the demands of long-horizon work are. Getting a single step right is easy; remembering the goal of step three at step fifty is hard. Recognizing a button is easy; recovering when a page fails to load, a dialog blocks the view, or a control has moved is hard. In these tasks a single error usually means restarting the whole flow, so reliable recovery matters more for usability than peak single-step performance. The best test samples for such a model are not polished demo videos but the most tedious, error-prone routines of your own business.

Dual-Protocol API: One Model, Two Dialects

The most immediately practical update is at the interface layer: the API is compatible with both OpenAI and Anthropic protocols. The OpenAI-format base_url is https://api.longcat.ai/openai and the Anthropic-format base_url is https://api.longcat.ai/anthropic, with model ID LongCat-2.5-Preview and the platform entry at https://longcat.ai/platform/ (official figures as relayed by AI-bot.cn).

Dual-protocol compatibility means a large body of existing tools can connect with zero rework. For Claude Code, the official tutorial route is simply three environment variables:

bash
export ANTHROPIC_BASE_URL=https://api.longcat.ai/anthropic
export ANTHROPIC_AUTH_TOKEN=your-token
export ANTHROPIC_MODEL=LongCat-2.5-Preview

Tools such as OpenClaw, OpenCode, Codex, Kilo Code, and Hermes connect the same way, by changing the base_url and the model name (official tutorial figures). For developers already living in the Anthropic ecosystem, this is the lowest-friction trial path: keep your toolchain, swap the backend.

Price Check: Side by Side with Opus 5.5

Billing follows token packages, and new accounts receive 5 million free tokens (relayed figures). As a reference point, the AI-bot.cn comparison table lists Claude Opus 5.5 at 4 dollars per million input tokens and 20 dollars per million output tokens (about 29 and 143 yuan respectively), with the same 1M context and 128K maximum output, but support for the Anthropic protocol only.

On paper, LongCat-2.5-Preview matches Opus 5.5 on window specifications, offers a wider protocol surface, and ships a signup grant large enough to run a small-to-medium round of realistic testing. Whether the model itself performs at the same tier is a question no public data can answer today, and that is exactly the core issue the next section takes up.

Cold Thought One: No Public Benchmarks

Here is the bucket of cold water this article owes you: LongCat-2.5-Preview currently has no public benchmark scores, and the comparison table states this explicitly with "no official benchmark" (AI-bot.cn comparison table figures).

This is not a trivial gap. By contrast, Claude Opus 5.5 has public evaluations to inspect, and flagship releases in this industry normally ship with a full set of scorecards. Every capability claim around LongCat-2.5-Preview currently lives at the level of officially relayed descriptions: the architecture is confirmed, the specs are confirmed, but the scores are simply absent. The absence of benchmarks does not mean the model is weak, but it does imply three things. First, you cannot compare it quantitatively against peer open flagship models; the cross-review of open-source flagships on this site is only possible because those scores are public. Second, every claim of "matching a given tier" must stay unverified until proven. Third, the only trustworthy evaluation left is hands-on testing, and the 5 million free tokens bring the cost of that close to zero.

The picture gets more interesting inside the Chinese flagship landscape. The Qwen3 line publishes leaderboards and lets the data speak; Kimi K3 handed open weights to the community and delegated verification to it; MiMo v2.6 proves itself through concrete scenario deployments. LongCat-2.5-Preview, by contrast, has chosen a third road: API only, no scores, no weights, pushing the burden of verification entirely onto users. The rational stance is to treat it as a strong-spec candidate awaiting verification, not as a proven front-runner.

Skipping evaluations can also be read as a pragmatic trade-off: in the agent era, public benchmarks are increasingly losing fidelity, leaderboards invite overfitting, and the gap between a gamed score and real usability keeps widening. Letting users vote with their actual workflows is a defensible logic, but it only holds if the product survives daily use. In other words, not publishing scores is not an exemption certificate; it is a wager that retention rates will do the talking.

Cold Thought Two: The Weights Are Not Out

The second open boundary is the open-source status. A GitHub API check of the meituan-longcat organization on September 29, 2026 found 30 repositories, none of them for LongCat-2.5-Preview. The organization's latest pushes, dated September 24, are LongCat-DeepResearch (9 stars) and WBench (240 stars). Existing model repositories include LongCat-2.0 (565 stars, MIT), LongCat-Flash-Chat (1367 stars, MIT), and LongCat-Video (8439 stars, MIT).

In short, 2.5-Preview is currently available only through the API, with no released weights, and anyone making adoption decisions should write that fact down. Given the family's precedent, the three model repositories above are all MIT-licensed, a later weight release is plausible. Until the official announcement lands, though, any "open weights soon" narrative is speculation. Teams that need private deployment, or that must keep data inside their own perimeter, can only watch for now or pick alternatives that have already opened their weights.

The meaning of API-only differs by audience. For individual developers and small teams, an API is in fact the fastest way to start: validate the use case with the free tokens first, then decide how deep to go, which is a sound ordering. Enterprise users should be more careful: wiring production traffic into a preview API means writing two uncertainties, availability and pricing policy, directly into the dependency chain. The safer play is to verify on side-path tasks behind a gradual rollout while keeping a fallback model one switch away. This is the other, risk-side value of dual-protocol support: the migration cost of swapping backends collapses to changing two configuration values.

Where This Fits: Division of Labor on This Site

Our earlier piece covered LongCat 2 and the domestic-chip training line, a story about training infrastructure; this article focuses on the 2.5 model itself, its architecture, interfaces, and openness, a story about the model. The two pieces cross-link and do not overlap: see Meituan LongCat 2 and the domestic chip training line for the infrastructure side.

Closing Thoughts

The spec sheet of LongCat-2.5-Preview is genuinely attractive: 1.6T total parameters, 48B active, a native 1M context, 128K maximum output, multimodality merged into the base, and a dual-protocol API. All of that is confirmed. What remains unconfirmed is capability: no benchmarks, no weights, only officially relayed descriptions.

The advice for pragmatic developers is short: spend the 5 million free tokens running your own real scenarios end to end, and that will be worth more than waiting for any leaderboard. Preview-stage models should be met with trials rather than trust, and that principle holds for every vendor. When the weights appear or the scores surface, we will come back to settle the second half of this story.

This article is AI-assisted and human-edited. Last updated: 2026-09-29

Related

Frontline Hotspot

Alibaba Shows Its Qwen4 Hand Early: Qwen3.8-Flash-Next Ships 125B Weights, But No License File

On August 26 Alibaba released Qwen3.8-Flash-Next: a multimodal MoE model that doubles as an early preview of the Qwen4 architecture - the same role Qwen3-Next once played for Qwen3.5. The main model is 125B parameters with an extra 51B of N-gram embeddings, activating just 6B per token; training costs about one ninth of Qwen3.7-Plus while delivering stronger coding and office performance. Four upgrades, unpacked: GDN compresses history while QSA uses a compressed indexer to pick important context at micro-block granularity; Gated Residual widens the residual stream into four branches; the N-gram embedding table can be offloaded to host memory and overlapped with compute via async prefetch; and the optimizer switches to Muon. Native context is 262,144 tokens, extensible to 1M with YaRN. The production Qwen3.8-Flash lists at \$0.16/\$0.47 per million tokens on QwenCloud (sources differ slightly; defer to the official site). The real open question is licensing: the GitHub repo ships no LICENSE file and its license field is None, the README simply points to the Hugging Face or ModelScope model page, and the community is already asking "why isn't it Apache 2.0?" - this article marks it unconfirmed, so verify the model page before any commercial use.

Aug 30, 20266 min read
Frontline Hotspot

DeepSeek V4.1 Flash Open Weights: The Asymmetric Design

DeepSeek open-sourced V4.1 Flash on 2026-09-10: a 552B-parameter MoE with an asymmetric Causal-Encoder-Decoder design that activates only 8B on input and 16B on output, natively multimodal, with officials citing significant KV Cache compression to cut agent-scenario cost. The API shipped alongside it - just switch the model name to deepseek-flash - and Tencent WorkBuddy, CodeBuddy plus OpenCode have integrated it fully. The model first surfaced on 9-08 as an internal preview build before being promoted on 9-10, a timeline worth noting in itself. This piece breaks down each release claim, argues the real engineering signal is not parameter count but the shift in long-context and agent cost structure implied by the asymmetric design plus 8B input activation, runs the numbers on what KV Cache compression means for accumulated multi-turn trajectories, and closes with cold takes: no published benchmark comparison, an unresolved relationship to its own V4-Flash, and concurrency and pricing still unconfirmed. Note that what shipped is model weights on HuggingFace; there is no dedicated code repository for V4.1 Flash under the official DeepSeek org.

Sep 10, 20269 min read
Frontline Hotspot

Tencent Open-Sources Hy4 Preview: A 770B Flagship That Helped Train Itself

On August 28, Tencent released and open-sourced its new flagship Hy4 preview (770B total / 49B active MoE, 78 layers): Gated DSA sparse attention + IndexCache cross-layer index reuse + iHC identity Hyper-Connections, with the README openly stating the architecture is "inspired by DeepSeek and GLM". A native MTP layer enables 3-token speculative decoding, context spans 1M tokens, and BF16+FP8 weights ship under Apache 2.0. In Tencent's internal blind eval, 163 experts scored 203 engineering tasks at 2.99/4.00, edging out GLM-5.3 (2.92) and Kimi K3 (2.94, both internal-caliber numbers). The headline is the early loop of recursive self-improvement: the model took part in automating optimization of its own training methods, data strategies, eval frameworks and low-level operators, and autonomously lifted inference end-to-end throughput by 31.8%. OpenRouter snapshot pricing: $0.834 input / $2.501 output per million tokens; free for two weeks on WorkBuddy/CodeBuddy.

Aug 29, 20266 min read