Frontline Hotspot
Frontline Hotspot

August 31, 2026: Sonnet 5 Reprices, GPT-5.4 Exits Codex, and Two Kimi Models Sunset on the Same Day

Three unrelated events landed on the same date. Claude Sonnet 5's API launch pricing expired, moving from $2 input and $10 output per million tokens to the standard $3 and $15. GPT-5.4 and GPT-5.4 mini stopped being offered to Codex users signed in with a ChatGPT account, replaced by GPT-5.6 Terra and Luna. Moonshot AI sunset kimi-k2.5 and moonshot-v1 on the same day, with kimi-k3 as the stated migration target. The layer most people miss is the second one: Sonnet 5 also changed tokenizer, so the same input now maps to 1.0x to 1.35x more tokens, which compounds with the rate change to a 65% to 100% real increase on coding workloads and close to 50% on plain text, while leaving Consumer subscriptions untouched. This piece splits the three events into sunset, replacement, and repricing, gives the urgency and response for each, and closes with a two-month expiry calendar: Claude Code limits on 2026-09-14, DashScope retiring 30-plus model IDs on 2026-10-10, deepseek-chat and deepseek-reasoner deprecated on 2026-10-24, OpenAI leaving Cursor on 2026-11-12, and the GPT-5.6 Sol promotion ending around 2026-11-21. Prices are a 2026-08-31 snapshot and source disagreements are flagged inline.

Published August 31, 20266 min read
<!-- aug31-price-sunset-deadline-hotspot | hotspot | August 31, 2026: Sonnet 5 Reprices, GPT-5.4 Exits Codex, and Two Kimi Models Sunset on the Same Day -->

Three unrelated changes landed on 2026-08-31. Claude Sonnet 5 moved off its introductory API rate of $2 input and $10 output per million tokens to the standard $3 and $15. GPT-5.4 and GPT-5.4 mini stopped being offered inside Codex to users signed in with a ChatGPT account. Moonshot AI sunset kimi-k2.5 and moonshot-v1, with kimi-k3 named as the migration target. Any one of these would read as routine lifecycle housekeeping. Stacked on the same date, they point in one direction: the period in which frontier models were priced to acquire users is closing.

Scope note: facts here come from Anthropic's pricing and tokenizer documentation at the Sonnet 5 launch (as reported by Jiqizhixin and Digital Trends), aitoolsrecap's 2026-08-31 daily summary, the Tencent AI daily brief for 2026-08-31, and TheRouter.ai's provider pricing retrieved 2026-08-24. Prices are a 2026-08-31 snapshot and move frequently; verify against your own account region's official pricing page. Items with a single source or with conflicting figures are flagged as such. Not investment advice.

What actually changed today

ChangeBeforeAfterWho is affected
Claude Sonnet 5 API$2 / $10 (intro, to 8-31)$3 / $15 (standard)API only, not subscriptions
Sonnet 5 tokenizerSonnet 4.6 tokenizerNew tokenizer, 1.0x to 1.35x tokensAll API calls
GPT-5.4 / GPT-5.4 miniServed in Codex to ChatGPT sign-in usersWithdrawn; replaced by GPT-5.6 Terra and LunaCodex users on ChatGPT login
kimi-k2.5, moonshot-v1AvailableSunsetAll callers; migrate to kimi-k3

Only the first is a price increase. The other two are a replacement and a retirement. What the three share is that code running yesterday is not guaranteed to cost the same, or to run at all, starting today.

Sonnet 5: 50 percent on the rate card, more than that on the invoice

Sonnet 5 launched on 2026-06-30 with the terms spelled out: introductory pricing of $2 input and $10 output per million tokens until 2026-08-31, standard pricing of $3 and $15 afterwards. This was announced in advance, which matters for how much grievance is warranted. Anthropic gave a two-month window; many people simply noticed it on the last day.

The layer people miss is the tokenizer. Sonnet 5 ships a new tokenizer, the same class of change that arrived with Claude Opus 4.7. Anthropic states that identical input now maps to more tokens, with the increase ranging from about 1.0x to 1.35x depending on content type, and that the introductory rate was set so that overall cost would stay roughly flat during the transition.

Compound the two and the number changes. The rate move is 50 percent. The tokenizer change adds up to 35 percent more tokens for the same content. aitoolsrecap puts coding workloads at 65 to 100 percent above where they sat last week, with text-heavy work landing closer to the headline 50 percent. If your traffic is mostly code, your experience of this increase will be considerably worse than 50 percent.

A boundary worth stating: Anthropic's position is that the introductory rate was designed to offset the tokenizer change, meaning the company considers transition-period cost roughly neutral. The 65 to 100 percent figure is a third-party calculation by content type, not an official number, and is labelled as such here.

Who is not affected

Consumer subscriptions are unaffected. Anthropic stated at launch that Sonnet 5 is the default model across Free and Pro plans and available on Max, Team, Enterprise, Claude Code, and the API. aitoolsrecap's 2026-08-31 summary says the same explicitly.

If your path is the Claude client, the web app, or Claude Code on a subscription, today's rate change will not show up on your bill. What needs recalculating is direct API usage: self-built applications, orchestration frameworks, batch pipelines, and products that resell Sonnet 5 as a substrate. For those, the useful exercise is not comparing list prices but taking last week's real call logs, applying $3 and $15, multiplying by a 1.0 to 1.35 factor, and arriving at your own unit cost before comparing against alternatives. A companion piece in this batch does exactly that comparison: see /en/posts/post-aug31-token-cost-comparison-review.

GPT-5.4 leaves Codex: a replacement, not a retirement

This one is widely misread as a model being killed. GPT-5.4 and GPT-5.4 mini left one route: Codex sessions authenticated by a ChatGPT account login. They remain reachable through the API, and they still work in Codex sessions authenticated with an API key rather than a login.

The distinction determines your remediation cost. If you run Codex on a ChatGPT login, the model serving your requests changed today, so behaviour and output style may shift and regression testing is warranted. If you already authenticate with an API key, or call the API directly, the practical impact is close to zero.

On price: GPT-5.6 Terra sits at $2 input and $12 output per million (permanent), and Luna at $0.20 and $1.20. Note also that GPT-5.6 Sol's $4 and $20 is promotional and reverts to roughly $5 and $30 around 2026-11-21. When modelling a migration, price it at the post-expiry rate. That rule applies to every limited-time rate, not just this one.

kimi-k2.5 and moonshot-v1: a hard cutoff

This one is an actual retirement. Moonshot sunset kimi-k2.5 and moonshot-v1 today, naming kimi-k3 as the migration target.

For anything still calling those IDs, this is a hard switch. Once the old IDs stop responding, calls fail rather than degrade. Reported kimi-k3 pricing varies by source: TheRouter.ai retrieved $3 input and $15 output per million with 1M context on 2026-08-24, while a page-by-page check of official Chinese pricing pages gives RMB 20 input and RMB 100 output. The gap comes down to billing region, currency, and whether cache discounts are included. This piece does not pick one number; check the official pricing page for your own account region before migrating.

There is a second expiry worth planning for now. TheRouter.ai also records a DashScope sunset wave on 2026-10-10 retiring more than 30 legacy model IDs, including qwen-turbo, qwen-vl-plus, qwen-audio-turbo, and several early Qwen3 snapshots. Today's retirement is not an isolated event. Treating model IDs as dependencies with expiry dates rather than permanent constants is the realistic posture. The full inventory and migration procedure is in this batch's SOP piece: /en/posts/model-sunset-migration-cost-sop.

Claude Code limits: announced as +25 percent, calculated by users as -17 percent

The most widely shared item today was not the price change but a correction in how a limit change was framed.

On 2026-08-29 Anthropic announced a permanent 25 percent increase to Claude Code weekly limits effective 2026-09-14, with the current 50 percent temporary boost holding until 09-13. Index the three points and the picture is plain: the pre-May baseline is 100, today is 150, and from September 14 it is 125. So the change is plus 25 percent against the old standard and minus 17 percent against what subscribers have right now.

Users worked that out publicly within hours. Anthropic then deleted the original post and republished a version stating plainly that compared to today this works out to a 17 percent reduction. A company choosing to correct its own framing rather than defend it is uncommon enough to be worth recording.

Note the effective date. The change lands on September 14, not today. If you depend on Claude Code weekly capacity for heavy development, you remain inside the 50 percent boost until September 13, and capacity will be lower than today afterwards.

The expiry calendar ahead

DateEvent
2026-09-14Claude Code permanent weekly limits take effect; 50 percent boost ends
2026-10-01OpenAI v Apple hearing
2026-10-10DashScope sunset wave retires 30+ legacy model IDs
2026-10-24deepseek-chat and deepseek-reasoner deprecated
2026-11-12OpenAI models leave Cursor following the SpaceX acquisition
~2026-11-21GPT-5.6 Sol promotional pricing ends: $4/$20 reverts to $5/$30

Read as a sequence, the signal is uncomfortable but plain. Frontier model access was priced as customer acquisition for a while, and now both flat-rate subscriptions and API rates are being recalibrated against what compute costs. That turn was always coming. The practical problem is that many budgets were built during the subsidy period, and if no headroom was left in unit price, what needs fixing now is not just migration work but the model underneath the budget.

For today, three changes map to three actions. If you call the Sonnet 5 API, recompute unit cost at $3 and $15 with a tokenizer factor applied. If you run Codex on a ChatGPT login, confirm how the new default model behaves on your tasks. If you still call kimi-k2.5 or moonshot-v1, move to kimi-k3 today; this one cannot wait for the next sprint.

This article is AI-assisted and human-edited. Last updated: 2026-08-31

FAQ

Is the Sonnet 5 price change a surprise move or a scheduled one?
Scheduled. The pricing documentation published on June 30 stated that launch pricing ran until 2026-08-31 and standard pricing would resume afterwards. Anthropic gave a two-month window; many people simply noticed on the last day. The useful lesson is that vendor pricing announcements are actionable data: register the launch-price expiry date the same way you register a dependency expiry and it stops being a surprise.
I pay for a Claude subscription. Does this affect me?
No. The change applies to API billing only; Consumer subscription allowances are unaffected. Watch the accounting boundary, though: subscription quota and API invoices are two different cost structures. Do not compare them in one table, and do not infer API cost movement from how your subscription feels. If your team uses both, model them separately with separate budget alerts.
GPT-5.4 is leaving Codex - will my API key calls break?
No. This affects Codex users signed in with a ChatGPT account, where GPT-5.6 Terra and Luna take over; API-key-authenticated calls are unaffected. The real risk for the affected group is not an error, it is a silent swap: the endpoint still works, the code still runs, but the model behind it changed. That calls for a regression pass, not waiting for something to break in production.
Where should I migrate from kimi-k2.5, and is changing the ID enough?
The stated target is kimi-k3, but the ID change is only step one. Three more things: a prompt regression, since same-vendor generation changes usually need validation of tool-call schemas and output format rather than a rewrite; tokenization, because the same prompt costs a different number of tokens and that moves both cost and context budget; and a fallback, confirming what you switch to if kimi-k3 is unavailable. Note also that kimi-k3 pricing is reported inconsistently, $3/$15 on TheRouter versus CNY 20/100 on the domestic pricing page, so defer to the official page for your account region.
What upcoming expiry dates should I be tracking?
In order: 2026-09-14 Claude Code limit change takes effect; 2026-10-01 OpenAI v Apple hearing; 2026-10-10 DashScope retires 30-plus model IDs; 2026-10-24 deepseek-chat and deepseek-reasoner deprecated; 2026-11-12 OpenAI models leave Cursor; around 2026-11-21 the GPT-5.6 Sol promotion ends, moving from $4/$20 back to $5/$30. Retired model IDs and expiring promotions are the two categories most often missed: the first breaks calls outright, the second invalidates your cost model overnight. Add this calendar to your dependency registry.

Related

Frontline Hotspot

AI Weekly 004: Seven Releases in Seven Days, but the Real Signals Are Agents, Compliance, and Cost

This week (Jul 27-Aug 2) the AI world shipped seven releases, but three signals matter more: DeepSeek-V4-Flash's post-training pushed DeepSWE from 7.3 to 54.4 (hands-on 30/30, cost under 5 fen) and Kimi K3 topped coding leaderboards; the EU AI Act August 2 deadline landed (fines up to 7% of global turnover, extraterritorial); prefix cache hits at 0.02 yuan vs 1 yuan misses make cost engineering a new skill.

Aug 2, 20265 min read
Frontline Hotspot

Kimi K3 Goes Open Source Tonight: 2.8 Trillion Parameters, World's Largest, As Yang Zhilin Closes the China-US Model Gap to 3 Months

On the evening of July 27, Moonshot AI open-sourced Kimi K3's weights: a 2.8-trillion-parameter MoE with a 1-million-token context, the world's largest open-source model, benchmarking against Anthropic's Fable 5. From the July 16 API launch to tonight's weight release, Yang Zhilin used ten days to compress the China-US model gap from 6-9 months to 3-5. Breakdown of specs, benchmarks, the comeback story, and the White House accusation.

Jul 27, 20263 min read
Frontline Hotspot

AI Weekly 003: GPT-5.6 Restricted, DeepSeek Open-Sources Inference Acceleration, Agents Shift from Chat to Work

This week's hard signals: OpenAI GPT-5.6 restricted by US regulators + self-developed Jalapeño chip, DeepSeek open-sources inference acceleration framework (A100 tasks moved to consumer GPUs, latency down 40%), Anthropic context-engineering guide, Xinliu Yuansu M-FLOW rewrites agent memory. Domestic AI carves a different track on efficiency/open-source/landing.

Jul 25, 20264 min read