Hardcore Reviews
Hardcore Reviews

Comparing 11 Models by Real Token Cost After the August 31 Repricing: Peak Hours, Cache Hits, and Tokenizer Effects

A model's list price wears at least three more layers. Time of day: DeepSeek moved to peak and off-peak pricing on August 17, charging peak rates on weekdays from 09:00-12:00 and 14:00-18:00, halving them off-peak, and applying off-peak rates all weekend, so the same model costs twice as much at 3pm as at 10pm. Caching: prefix cache hits are billed far below standard input, and the variable sits with your prompt structure rather than the vendor. Tokenization: Sonnet 5 changed tokenizers, so the same input now maps to 1.0x to 1.35x more tokens, and the multiplier floats with content type. This comparison fixes one unit throughout, blended rate equals input plus output divided by two, assuming equal token volumes, as a neutral starting point, then recalculates under three realistic load profiles across 11 models, covering list price, cached input, peak and off-peak, and post-tokenizer position. The finding is not which model is cheapest, it is that no model is cheapest, only cheapest for your particular load: any comparison that ignores input-output ratio, cache hit rate, and content type is comparing list prices, not costs. Chinese model prices come from a page-by-page check of official pricing pages on 2026-08-24, re-confirmed on 08-28; overseas prices from a 2026-08-31 roundup. Conflicts are flagged per line. No live benchmarking was performed.

Published August 31, 20269 min read
<!-- post-aug31-token-cost-comparison-review | review | Comparing 11 Models by Real Token Cost After the August 31 Repricing: Peak Hours, Cache Hits, and Tokenizer Effects -->

After 2026-08-31, a lot of teams discovered that a model's list price is not a sufficient number. Claude Sonnet 5 moved from $2 and $10 per million tokens to $3 and $15, and it had also changed tokenizers, so identical input now maps to 1.0x to 1.35x as many tokens. The increase on your invoice is not the increase on the rate card. On top of that, Chinese providers layer peak and off-peak windows and cache-hit discounts onto their pricing, which means "which model is cheaper" cannot be answered from a single line of numbers. This piece flattens the units and computes a comparable real cost.

Scope note: Chinese model prices come from a price table compiled by checking each vendor's official pricing page individually (checked 2026-08-24, re-checked 2026-08-28 with no changes found). Overseas prices come from aitoolsrecap's 2026-08-31 summary, with TheRouter.ai's provider pricing retrieved 2026-08-24 used as cross-reference. Prices are a 2026-08-31 snapshot and move frequently; verify against the official pricing page for your own account region. Items with conflicting figures are flagged individually rather than smoothed into one number. No live benchmarking was performed; every figure here is arithmetic on published prices and excludes concurrency, rate limiting, retry, and engineering costs. Not investment advice.

The list price is the least trustworthy number

Model pricing wears at least three layers, and reading the list price means reading only the outermost one.

The first layer is time of day. DeepSeek has billed peak and off-peak since 2026-08-17: weekdays 09:00-12:00 and 14:00-18:00 are peak, off-peak is half price, and weekends are off-peak all day. The same model costs twice as much at 3pm as at 10pm. For workloads that can be scheduled asynchronously, that is a discount you collect with no engineering work at all. For anything that must respond in real time, peak is your real cost and off-peak pricing is irrelevant to you.

The second layer is cache. Most platforms charge far less for input tokens that hit a prefix cache, in some models a small fraction of the standard input rate. This variable is not held by the vendor; it is held by you. Whether your prompt structure is stable, whether system instructions sit at the very front, and whether multi-turn conversations reuse history all determine how much of that discount you actually collect. Two teams on the same model can differ three or fourfold in unit cost purely on prefix reuse.

The third layer is tokenization. Sonnet 5's tokenizer change maps identical input to 1.0x to 1.35x as many tokens, varying by content type. This is the sneakiest layer because it changes no price list anywhere; it just charges you more for the same content. It also breaks naive cross-model comparison: different models use different tokenizers, so the same Chinese corpus might be 1,000 tokens on one model and 1,300 on another. Before comparing prices per token, confirm that both sides are counting the same thing.

Fixing the unit of comparison

For comparability, this piece uses one clearly defined middle-ground metric: blended unit price = (input price + output price) / 2, which assumes input and output token volumes are equal. The 1:1 ratio is chosen not because it is typical but because it is a neutral starting point that is easy to adjust. A later section recalculates under three realistic ratios, and the ranking does change.

Three more conventions. Overseas models are priced in dollars and Chinese models in RMB with no currency conversion applied, since exchange rates move and most teams have costs and revenue in the same currency, so cross-currency comparison tends to distort rather than clarify. Cache-hit prices are listed separately rather than folded into the blended figure, because hit rates depend heavily on usage. Promotional and permanent prices are listed separately, and wherever a promotion has an expiry date, the post-expiry price is given as well, because migrations should be modelled at the post-expiry rate.

Chinese models: peak hours and cache move the spread by more than 10x

Unit: RMB per million tokens. Blended price assumes 1:1 input to output. The DeepSeek peak rule is described above.

ModelCache hit inputInputOutputBlended
Zhipu GLM-5.3-Flash0.230.82.81.8
Alibaba Qwen3.8-Flash~0.1132.0
Tencent Hy30.25142.5
Xiaomi MiMo-v2.5-Pro0.0253.06.04.5
DeepSeek V4-Flash0.103.09.06.0
DeepSeek V4-Pro (off-peak)not verified4.513.59.0
Moonshot Kimi K2.61.16.52716.75
DeepSeek V4-Pro (peak)0.309.027.018.0
Zhipu GLM-5.3 / 5.2282818.0
ByteDance Seed 2.1 Pro1.263018.0
Alibaba Qwen3.8-Max1.5123624.0
Moonshot Kimi K32.02010060.0

The spread is the thing to look at: 1.8 at the cheap end against 60 at the expensive end, more than thirtyfold. That is not an argument that expensive models are bad. Kimi K3 prices against its position on coding benchmarks such as Terminal-Bench, while GLM-5.3-Flash is positioned for cheap high-frequency calls. The real observation is that most teams' actual workloads do not need flagship capability, yet they pay flagship rates indefinitely.

Conflicts must be flagged item by item.

GLM-5.3 pricing is disputed across sources. The table used here gives input 8 and output 28 RMB (re-checked 2026-08-28). Two other sources give 8 and 32, and roughly 3.6 and 14.4, respectively. TheRouter.ai marks GLM-5.3 as pricing TBA and lists only 5.2 at $1.40 and $4.40. Four numbers that do not reconcile; this piece takes the page-verified and most recent one and explicitly warns against treating the resulting 18 RMB blended figure as settled.

DeepSeek V4-Pro's basis changed over time. At its 2026-08-13 GA, an institutional research note recorded uncached input at 3 and output at 6 RMB. The peak/off-peak change on 2026-08-17 moved peak to 9 and 27, with off-peak at half. If you see "V4-Pro at 6 RMB output" online, that figure predates the peak/off-peak reform: not wrong, just expired.

One entry needs a warning. The Tencent row is Hy3 (256K context, input 1, output 4). The Hy4 preview released 2026-08-28 is a separate 770B model whose official positioning is that it undercuts DeepSeek V4-Pro at peak, but no reliable official figure was obtained here, so no number is given and Hy4 is excluded from the ranking above.

Overseas models: promotional and permanent rates must be separated

Unit: USD per million tokens.

ModelInputOutputBlendedNature of price
GPT-5.6 Luna$0.20$1.20$0.70Permanent
GPT-5.6 Terra$2$12$7Permanent
Claude Sonnet 5$3$15$9Permanent, plus tokenizer effect
GPT-5.6 Sol$4$20$12Promotional, expires ~2026-11-21
GPT-5.6 Sol (post-expiry)$5$30$17.5Permanent

The most commonly misread line here is Sol. At $4 and $20 it looks pricier than Sonnet 5's $3 and $15, but the two are not on the same time scale. Sonnet 5's $3 and $15 is permanent; Sol's promotion ends around 2026-11-21, after which blended cost jumps from $12 to $17.5. If a migration requires engineering effort, model the return at $17.5, not at $12. Three months of savings rarely covers a migration.

A more basic point: Luna's blended $0.70 is roughly one thirteenth of Sonnet 5's. It is not built for hard reasoning, but it is entirely adequate for high-throughput simple work such as classification, extraction, formatting, and routing. Most production traffic is mixed, and moving the simple portion off a flagship usually saves more than switching vendors does. This batch's SOP piece covers how to split that traffic: /en/posts/model-sunset-migration-cost-sop.

Where Sonnet 5 actually lands once tokenization is counted

With the tokenizer change applied, Sonnet 5's blended price is a range rather than a number.

At 1:1 with rates of $3 and $15, the blended price is $9. Applying the tokenizer factor: for workloads that are mostly natural language text, the factor approaches 1.0 and the blended price stays at $9; for code-heavy workloads near the 1.35x ceiling, it becomes $12.15. The same Sonnet 5 therefore costs between $9 and $12.15 depending on what you send it, a 35 percent spread.

That range moves its position relative to neighbours. Against Terra at $7, Sonnet 5 is about 22 percent more expensive for text-heavy work and about 42 percent more for code. Conversely, if your traffic is code-heavy, the gap between Sonnet 5 after the increase and Sol's promotional rate is narrower than the rate cards suggest, which weakens the case for migrating to Sol.

One caveat on the 1.0 to 1.35 range: it is Anthropic's "depending on content type" wording at launch, with no official breakdown of which content lands where. No further inference is made here. The cheap and accurate move is to take your own corpus, count it under the old and new tokenizers, and get your real factor. It costs almost nothing and beats any estimate.

Recalculating under three realistic load profiles

The 1:1 assumption is only a starting point. Reweighting for three common profiles changes the ordering materially.

Profile A: output-heavy, long-form generation and report writing, roughly 1:3 input to output.

ModelWeighted blended
Zhipu GLM-5.3-Flash2.30
Tencent Hy33.25
DeepSeek V4-Pro (off-peak)11.25
Moonshot Kimi K380.0

The more output dominates, the more models with cheap output pull ahead. GLM-5.3-Flash and Hy3 extend their lead, while Kimi K3's 100 RMB output rate becomes hard to justify.

Profile B: input-heavy, long-document analysis and retrieval QA, roughly 10:1.

ModelWeighted blended
Zhipu GLM-5.3-Flash0.98
Tencent Hy31.27
DeepSeek V4-Pro (off-peak)5.32
Moonshot Kimi K327.3

When input dominates, cache hit rate becomes the dominant term. In this profile the deciding factor is often not price at all but prompt engineering: whether stable system instructions sit at the front and whether multi-turn sessions reuse prefixes determines whether you pay standard input rates or cache rates.

Profile C: code workloads, where the tokenizer factor applies, using Sonnet 5 as the example.

ModelWeighted blendedNote
Claude Sonnet 5 (text-heavy)$9.0factor 1.0
Claude Sonnet 5 (code-heavy)$12.15factor 1.35
GPT-5.6 Terra$7no tokenizer change recorded

The conclusion across all three profiles is consistent: there is no cheapest model, only the model that is cheapest for your particular load. Any comparison that ignores input-to-output ratio, cache hit rate, and content type is comparing list prices, not costs.

Recommendations and three hard caveats

By load type, the actionable version: route high-throughput simple work to Luna or GLM-5.3-Flash and reserve flagships for the parts that need them; schedule anything asynchronous into off-peak windows, since DeepSeek's peak-to-off-peak spread is 2x and requires no engineering change; for long-document work, optimise prompt structure for cache hits before shopping for a cheaper model; and when comparing models on code workloads, measure your own tokenizer factor first, or you are comparing two different units of measurement.

Three caveats. First, model at the post-expiry price: any promotion should be evaluated at where it lands after expiry, and Sol's jump from $12 to $17.5 is a live example. Second, old IDs expire: kimi-k2.5 and moonshot-v1 sunset today, and DashScope has another wave on 2026-10-10 retiring 30-plus model IDs, which is why model IDs belong in dependency management with expiry dates. Third, do not decide from one comparison table: everything here is arithmetic on published prices with no live benchmarking, and it excludes concurrency, rate limits, retries, and output quality differences. To choose properly, run your real traffic across the candidates for a week and let the actual invoices decide. That is the only comparison that counts.

FAQ

Q1: How do list prices that differ by 2x produce real costs that differ by 10x? A1: Because list price is only the first of four layers. Below it sit time-of-day pricing (DeepSeek's peak and off-peak rates differ by 2x), cache hits (cached and uncached input rates can differ by an order of magnitude), and tokenization (Sonnet 5 changed its tokenizer, so the same input now counts 1.0x to 1.35x more tokens). Multiplied together, two models with similar list prices can produce very different invoices. The conclusion is not that list prices are useless, it is that they are useless in isolation.

Q2: How do I actually use time-of-day pricing? A2: DeepSeek charges peak rates on weekdays from 09:00-12:00 and 14:00-18:00, halves the rate off-peak, and applies off-peak rates all weekend. Any workload that can run asynchronously, batch summarisation, offline labelling, data cleaning, regression suites that run overnight, gets a discount that requires no engineering change at all. The only precondition is that your product tolerates the delay.

Q3: What cache hit rate should I expect? A3: There is no universal number; it depends entirely on prompt structure. The rule is that stable long content (system prompt, tool definitions, document context) belongs in the prefix, and the parts that change (user input, timestamps, session state) belong at the end. Reverse that order and the hit rate approaches zero. The only way to get a real number is to measure a week of production traffic and read percentiles, not averages.

Q4: How much did Sonnet 5 actually increase, and why do sources disagree? A4: Both figures are right; they measure different things. On rates, input goes from $2 to $3 and output from $10 to $15, a 50% increase. Add the token-count inflation from the tokenizer change (official range 1.0x to 1.35x) and the compound increase lands between 65% and 100% on coding workloads, around 50% on plain text. Note also that this affects the API only; Consumer subscriptions are unaffected, so do not put the two cost types in the same table.

Q5: What do I do about promotional pricing, and should I migrate now? A5: Any migration that costs engineering time should be justified at the post-promotion price, not the promotional one. GPT-5.6 Sol's $4/$20 is promotional and reverts to roughly $5/$30 around 2026-11-21, moving the blended rate from $12 to $17.5. Three months of savings rarely covers the cost of a migration. The test is simple: if the numbers still work after the promotion ends, the migration is worth doing.

This article is AI-assisted and human-edited. Last updated: 2026-08-31

FAQ

How do list prices that differ by 2x produce real costs that differ by 10x?
Because list price is only the first of four layers. Below it sit time-of-day pricing, where peak and off-peak rates differ by 2x; cache hits, where cached and uncached input rates can differ by an order of magnitude; and tokenization, where Sonnet 5 now counts 1.0x to 1.35x more tokens for the same input. Multiplied together, two models with similar list prices can produce very different invoices. The conclusion is not that list prices are useless, it is that they are useless in isolation.
How do I actually use time-of-day pricing?
DeepSeek charges peak rates on weekdays from 09:00-12:00 and 14:00-18:00, halves the rate off-peak, and applies off-peak rates all weekend. Any workload that can run asynchronously, batch summarisation, offline labelling, data cleaning, regression suites that run overnight, gets a discount that requires no engineering change at all. The only precondition is that your product tolerates the delay.
What cache hit rate should I expect?
There is no universal number; it depends entirely on prompt structure. The rule is that stable long content (system prompt, tool definitions, document context) belongs in the prefix, and the parts that change (user input, timestamps, session state) belong at the end. Reverse that order and the hit rate approaches zero. The only way to get a real number is to measure a week of production traffic and read percentiles, not averages.
How much did Sonnet 5 actually increase, and why do sources disagree?
Both figures are right; they measure different things. On rates, input goes from $2 to $3 and output from $10 to $15, a 50% increase. Add the token-count inflation from the tokenizer change (official range 1.0x to 1.35x) and the compound increase lands between 65% and 100% on coding workloads, around 50% on plain text. Note also that this affects the API only; Consumer subscriptions are unaffected, so do not put the two cost types in the same table.
What do I do about promotional pricing, and should I migrate now?
Any migration that costs engineering time should be justified at the post-promotion price, not the promotional one. GPT-5.6 Sol's $4/$20 is promotional and reverts to roughly $5/$30 around November 21, moving the blended rate from $12 to $17.5. Three months of savings rarely covers the cost of a migration. The test is simple: if the numbers still work after the promotion ends, the migration is worth doing.

Related

Hardcore Reviews

Cache Economics: How Hit Rate Decides Your Real Agentic Bill

On 2026-09-01 Fable 5.1 cut cache read from $1 to $0.25 per million tokens (75%), and the community cheered "agents got cheaper" — but the bill is unit price times token structure: the lower cache read's share, the less the cut moves total cost. This review splits tokens into four classes (fresh input / cache write / cache read / output), gives a cost formula, and runs a sensitivity analysis across four load profiles — at 10% share the cut saves only ~7.5%, at 33% ~25%, at 60% ~45% (derived from the official reduction, not a measured bill). Verdict: unit price is only the fourth factor; hit rate, layout stability, round count and output length matter more. Six engineering preconditions lift hit rate (invariant prefix, stable layout, turn-scoped instructions, server-side history trimming, TTL by frequency, observable hit rate). Cross-vendor application needs the vendor's 2026-09-02 official snapshot across six dimensions.

Sep 1, 202610 min read
Hardcore Reviews

One compromised agent loses everything: a comparison of four credential and permission governance approaches

Credentials went from a config item to an attack surface, yet most teams' defenses are still stuck at "put the agent in a sandbox." This review splits cleanly from our sandbox-isolation comparison: the sandbox governs where code runs; credential governance governs how secrets are used, who approves actions, and whether they can leave. It contrasts four approaches — OpenClaw 2.0, OpenWorker, OpenHuman and traditional secret storage — across six lifecycle stages (store / use / approve / exfiltrate / audit / multi-agent): OpenClaw with masked requests plus an opt-in proxy allowlist; OpenWorker with hard floors, an autonomy ladder, a reviewer model and a circuit breaker, and never self-approving unattended; OpenHuman with Privacy Mode enforced in the Rust core and E2E-encrypted inter-agent comms. Secondhand data (SaaS Sentinel transcription, no primary source located) shows compromise probability 0.24 with one agent rising to 0.86 with seven — risk grows superlinearly with count, under the premise "any agent proposes, execute."

Sep 1, 202611 min read
Hardcore Reviews

What Does a 30-Second 1080P Video Actually Cost: A Six-Way AI Video Generation API Cost Comparison

Wan3.0's launch turns "what does one 30-second 1080P clip actually cost" into a question you can compute precisely. This comparison runs the money ledger across six video generation APIs: Wan3.0 at an official 1.2 RMB/s, 36 RMB for a single-segment 30-second clip (25.2 RMB discounted through 09-23); Kling 3.0 around 30 RMB but requiring 3 stitched segments; Sora 2 pro breaking 100 RMB for 30 seconds; Hailuo's 768P at just 12 RMB across 3 segments, the cheapest in the table. Our exclusive ledger exposes the single-segment duration cap as an overlooked hidden cost - segment count x gacha multiplier (15-20% below-bar rate) x stitching labor is the real price - plus a two-tier playbook (480P gacha, 1080P final render) that saves 65%. All prices tagged with official vs aggregator provenance; a representative comparison, not a stress test.

Aug 25, 20269 min read