After 2026-08-31, a lot of teams discovered that a model's list price is not a sufficient number. Claude Sonnet 5 moved from $2 and $10 per million tokens to $3 and $15, and it had also changed tokenizers, so identical input now maps to 1.0x to 1.35x as many tokens. The increase on your invoice is not the increase on the rate card. On top of that, Chinese providers layer peak and off-peak windows and cache-hit discounts onto their pricing, which means "which model is cheaper" cannot be answered from a single line of numbers. This piece flattens the units and computes a comparable real cost.
Scope note: Chinese model prices come from a price table compiled by checking each vendor's official pricing page individually (checked 2026-08-24, re-checked 2026-08-28 with no changes found). Overseas prices come from aitoolsrecap's 2026-08-31 summary, with TheRouter.ai's provider pricing retrieved 2026-08-24 used as cross-reference. Prices are a 2026-08-31 snapshot and move frequently; verify against the official pricing page for your own account region. Items with conflicting figures are flagged individually rather than smoothed into one number. No live benchmarking was performed; every figure here is arithmetic on published prices and excludes concurrency, rate limiting, retry, and engineering costs. Not investment advice.
The list price is the least trustworthy number
Model pricing wears at least three layers, and reading the list price means reading only the outermost one.
The first layer is time of day. DeepSeek has billed peak and off-peak since 2026-08-17: weekdays 09:00-12:00 and 14:00-18:00 are peak, off-peak is half price, and weekends are off-peak all day. The same model costs twice as much at 3pm as at 10pm. For workloads that can be scheduled asynchronously, that is a discount you collect with no engineering work at all. For anything that must respond in real time, peak is your real cost and off-peak pricing is irrelevant to you.
The second layer is cache. Most platforms charge far less for input tokens that hit a prefix cache, in some models a small fraction of the standard input rate. This variable is not held by the vendor; it is held by you. Whether your prompt structure is stable, whether system instructions sit at the very front, and whether multi-turn conversations reuse history all determine how much of that discount you actually collect. Two teams on the same model can differ three or fourfold in unit cost purely on prefix reuse.
The third layer is tokenization. Sonnet 5's tokenizer change maps identical input to 1.0x to 1.35x as many tokens, varying by content type. This is the sneakiest layer because it changes no price list anywhere; it just charges you more for the same content. It also breaks naive cross-model comparison: different models use different tokenizers, so the same Chinese corpus might be 1,000 tokens on one model and 1,300 on another. Before comparing prices per token, confirm that both sides are counting the same thing.
Fixing the unit of comparison
For comparability, this piece uses one clearly defined middle-ground metric: blended unit price = (input price + output price) / 2, which assumes input and output token volumes are equal. The 1:1 ratio is chosen not because it is typical but because it is a neutral starting point that is easy to adjust. A later section recalculates under three realistic ratios, and the ranking does change.
Three more conventions. Overseas models are priced in dollars and Chinese models in RMB with no currency conversion applied, since exchange rates move and most teams have costs and revenue in the same currency, so cross-currency comparison tends to distort rather than clarify. Cache-hit prices are listed separately rather than folded into the blended figure, because hit rates depend heavily on usage. Promotional and permanent prices are listed separately, and wherever a promotion has an expiry date, the post-expiry price is given as well, because migrations should be modelled at the post-expiry rate.
Chinese models: peak hours and cache move the spread by more than 10x
Unit: RMB per million tokens. Blended price assumes 1:1 input to output. The DeepSeek peak rule is described above.
| Model | Cache hit input | Input | Output | Blended |
|---|---|---|---|---|
| Zhipu GLM-5.3-Flash | 0.23 | 0.8 | 2.8 | 1.8 |
| Alibaba Qwen3.8-Flash | ~0.1 | 1 | 3 | 2.0 |
| Tencent Hy3 | 0.25 | 1 | 4 | 2.5 |
| Xiaomi MiMo-v2.5-Pro | 0.025 | 3.0 | 6.0 | 4.5 |
| DeepSeek V4-Flash | 0.10 | 3.0 | 9.0 | 6.0 |
| DeepSeek V4-Pro (off-peak) | not verified | 4.5 | 13.5 | 9.0 |
| Moonshot Kimi K2.6 | 1.1 | 6.5 | 27 | 16.75 |
| DeepSeek V4-Pro (peak) | 0.30 | 9.0 | 27.0 | 18.0 |
| Zhipu GLM-5.3 / 5.2 | 2 | 8 | 28 | 18.0 |
| ByteDance Seed 2.1 Pro | 1.2 | 6 | 30 | 18.0 |
| Alibaba Qwen3.8-Max | 1.5 | 12 | 36 | 24.0 |
| Moonshot Kimi K3 | 2.0 | 20 | 100 | 60.0 |
The spread is the thing to look at: 1.8 at the cheap end against 60 at the expensive end, more than thirtyfold. That is not an argument that expensive models are bad. Kimi K3 prices against its position on coding benchmarks such as Terminal-Bench, while GLM-5.3-Flash is positioned for cheap high-frequency calls. The real observation is that most teams' actual workloads do not need flagship capability, yet they pay flagship rates indefinitely.
Conflicts must be flagged item by item.
GLM-5.3 pricing is disputed across sources. The table used here gives input 8 and output 28 RMB (re-checked 2026-08-28). Two other sources give 8 and 32, and roughly 3.6 and 14.4, respectively. TheRouter.ai marks GLM-5.3 as pricing TBA and lists only 5.2 at $1.40 and $4.40. Four numbers that do not reconcile; this piece takes the page-verified and most recent one and explicitly warns against treating the resulting 18 RMB blended figure as settled.
DeepSeek V4-Pro's basis changed over time. At its 2026-08-13 GA, an institutional research note recorded uncached input at 3 and output at 6 RMB. The peak/off-peak change on 2026-08-17 moved peak to 9 and 27, with off-peak at half. If you see "V4-Pro at 6 RMB output" online, that figure predates the peak/off-peak reform: not wrong, just expired.
One entry needs a warning. The Tencent row is Hy3 (256K context, input 1, output 4). The Hy4 preview released 2026-08-28 is a separate 770B model whose official positioning is that it undercuts DeepSeek V4-Pro at peak, but no reliable official figure was obtained here, so no number is given and Hy4 is excluded from the ranking above.
Overseas models: promotional and permanent rates must be separated
Unit: USD per million tokens.
| Model | Input | Output | Blended | Nature of price |
|---|---|---|---|---|
| GPT-5.6 Luna | $0.20 | $1.20 | $0.70 | Permanent |
| GPT-5.6 Terra | $2 | $12 | $7 | Permanent |
| Claude Sonnet 5 | $3 | $15 | $9 | Permanent, plus tokenizer effect |
| GPT-5.6 Sol | $4 | $20 | $12 | Promotional, expires ~2026-11-21 |
| GPT-5.6 Sol (post-expiry) | $5 | $30 | $17.5 | Permanent |
The most commonly misread line here is Sol. At $4 and $20 it looks pricier than Sonnet 5's $3 and $15, but the two are not on the same time scale. Sonnet 5's $3 and $15 is permanent; Sol's promotion ends around 2026-11-21, after which blended cost jumps from $12 to $17.5. If a migration requires engineering effort, model the return at $17.5, not at $12. Three months of savings rarely covers a migration.
A more basic point: Luna's blended $0.70 is roughly one thirteenth of Sonnet 5's. It is not built for hard reasoning, but it is entirely adequate for high-throughput simple work such as classification, extraction, formatting, and routing. Most production traffic is mixed, and moving the simple portion off a flagship usually saves more than switching vendors does. This batch's SOP piece covers how to split that traffic: /en/posts/model-sunset-migration-cost-sop.
Where Sonnet 5 actually lands once tokenization is counted
With the tokenizer change applied, Sonnet 5's blended price is a range rather than a number.
At 1:1 with rates of $3 and $15, the blended price is $9. Applying the tokenizer factor: for workloads that are mostly natural language text, the factor approaches 1.0 and the blended price stays at $9; for code-heavy workloads near the 1.35x ceiling, it becomes $12.15. The same Sonnet 5 therefore costs between $9 and $12.15 depending on what you send it, a 35 percent spread.
That range moves its position relative to neighbours. Against Terra at $7, Sonnet 5 is about 22 percent more expensive for text-heavy work and about 42 percent more for code. Conversely, if your traffic is code-heavy, the gap between Sonnet 5 after the increase and Sol's promotional rate is narrower than the rate cards suggest, which weakens the case for migrating to Sol.
One caveat on the 1.0 to 1.35 range: it is Anthropic's "depending on content type" wording at launch, with no official breakdown of which content lands where. No further inference is made here. The cheap and accurate move is to take your own corpus, count it under the old and new tokenizers, and get your real factor. It costs almost nothing and beats any estimate.
Recalculating under three realistic load profiles
The 1:1 assumption is only a starting point. Reweighting for three common profiles changes the ordering materially.
Profile A: output-heavy, long-form generation and report writing, roughly 1:3 input to output.
| Model | Weighted blended |
|---|---|
| Zhipu GLM-5.3-Flash | 2.30 |
| Tencent Hy3 | 3.25 |
| DeepSeek V4-Pro (off-peak) | 11.25 |
| Moonshot Kimi K3 | 80.0 |
The more output dominates, the more models with cheap output pull ahead. GLM-5.3-Flash and Hy3 extend their lead, while Kimi K3's 100 RMB output rate becomes hard to justify.
Profile B: input-heavy, long-document analysis and retrieval QA, roughly 10:1.
| Model | Weighted blended |
|---|---|
| Zhipu GLM-5.3-Flash | 0.98 |
| Tencent Hy3 | 1.27 |
| DeepSeek V4-Pro (off-peak) | 5.32 |
| Moonshot Kimi K3 | 27.3 |
When input dominates, cache hit rate becomes the dominant term. In this profile the deciding factor is often not price at all but prompt engineering: whether stable system instructions sit at the front and whether multi-turn sessions reuse prefixes determines whether you pay standard input rates or cache rates.
Profile C: code workloads, where the tokenizer factor applies, using Sonnet 5 as the example.
| Model | Weighted blended | Note |
|---|---|---|
| Claude Sonnet 5 (text-heavy) | $9.0 | factor 1.0 |
| Claude Sonnet 5 (code-heavy) | $12.15 | factor 1.35 |
| GPT-5.6 Terra | $7 | no tokenizer change recorded |
The conclusion across all three profiles is consistent: there is no cheapest model, only the model that is cheapest for your particular load. Any comparison that ignores input-to-output ratio, cache hit rate, and content type is comparing list prices, not costs.
Recommendations and three hard caveats
By load type, the actionable version: route high-throughput simple work to Luna or GLM-5.3-Flash and reserve flagships for the parts that need them; schedule anything asynchronous into off-peak windows, since DeepSeek's peak-to-off-peak spread is 2x and requires no engineering change; for long-document work, optimise prompt structure for cache hits before shopping for a cheaper model; and when comparing models on code workloads, measure your own tokenizer factor first, or you are comparing two different units of measurement.
Three caveats. First, model at the post-expiry price: any promotion should be evaluated at where it lands after expiry, and Sol's jump from $12 to $17.5 is a live example. Second, old IDs expire: kimi-k2.5 and moonshot-v1 sunset today, and DashScope has another wave on 2026-10-10 retiring 30-plus model IDs, which is why model IDs belong in dependency management with expiry dates. Third, do not decide from one comparison table: everything here is arithmetic on published prices with no live benchmarking, and it excludes concurrency, rate limits, retries, and output quality differences. To choose properly, run your real traffic across the candidates for a week and let the actual invoices decide. That is the only comparison that counts.
FAQ
Q1: How do list prices that differ by 2x produce real costs that differ by 10x? A1: Because list price is only the first of four layers. Below it sit time-of-day pricing (DeepSeek's peak and off-peak rates differ by 2x), cache hits (cached and uncached input rates can differ by an order of magnitude), and tokenization (Sonnet 5 changed its tokenizer, so the same input now counts 1.0x to 1.35x more tokens). Multiplied together, two models with similar list prices can produce very different invoices. The conclusion is not that list prices are useless, it is that they are useless in isolation.
Q2: How do I actually use time-of-day pricing? A2: DeepSeek charges peak rates on weekdays from 09:00-12:00 and 14:00-18:00, halves the rate off-peak, and applies off-peak rates all weekend. Any workload that can run asynchronously, batch summarisation, offline labelling, data cleaning, regression suites that run overnight, gets a discount that requires no engineering change at all. The only precondition is that your product tolerates the delay.
Q3: What cache hit rate should I expect? A3: There is no universal number; it depends entirely on prompt structure. The rule is that stable long content (system prompt, tool definitions, document context) belongs in the prefix, and the parts that change (user input, timestamps, session state) belong at the end. Reverse that order and the hit rate approaches zero. The only way to get a real number is to measure a week of production traffic and read percentiles, not averages.
Q4: How much did Sonnet 5 actually increase, and why do sources disagree? A4: Both figures are right; they measure different things. On rates, input goes from $2 to $3 and output from $10 to $15, a 50% increase. Add the token-count inflation from the tokenizer change (official range 1.0x to 1.35x) and the compound increase lands between 65% and 100% on coding workloads, around 50% on plain text. Note also that this affects the API only; Consumer subscriptions are unaffected, so do not put the two cost types in the same table.
Q5: What do I do about promotional pricing, and should I migrate now? A5: Any migration that costs engineering time should be justified at the post-promotion price, not the promotional one. GPT-5.6 Sol's $4/$20 is promotional and reverts to roughly $5/$30 around 2026-11-21, moving the blended rate from $12 to $17.5. Three months of savings rarely covers the cost of a migration. The test is simple: if the numbers still work after the promotion ends, the migration is worth doing.