1. Twin Arrival: A Price Sheet as the Headline
Around September 23, 2026, OpenAI quietly put two new models live on its website: GPT-6 Sol and GPT-6 Luna, per OpenAI's official disclosure as relayed by ai-bot.cn and QbitAI. Unlike the star-themed buildup that preceded the GPT-6 Astra flagship launch, this release arrived without a promotional campaign. What arrived instead was a price sheet strong enough to force the entire industry to redo its arithmetic: Sol costs 2 dollars per million input tokens and 10 dollars per million output tokens, a 50 percent cut from the previous generation GPT-5.6 Sol; Luna costs 0.1 dollars per million input tokens and 0.5 dollars per million output tokens, likewise half the previous generation's price.
Start with the names. Sol (the sun) and Luna (the moon) continue OpenAI's celestial naming tradition, in which Astra was the stars and the twin pair divides day from night. In terms of positioning, Sol is a capability-downgraded version built on the GPT-6 Astra base, officially described as distillation, with the pitch centered on "cost per unit of intelligence." Luna is the lightweight tier, inheriting Astra's multimodal capabilities and targeting high-frequency, large-scale, price-sensitive workloads. The model identifiers are gpt-6-sol and gpt-6-luna, with entry points across ChatGPT Work, Codex, and the API. The parameter counts have not been disclosed.
The availability plan is also telling: Sol is open to Plus, Pro, Business, Enterprise, and Edu subscribers, while Free and Go users can only use Luna for now. In other words, free users receive the model with the steepest discount, while those with stronger willingness to pay are steered into Sol's ecosystem. For practitioners tracking the frontier, the protagonist of this release is not a record-breaking benchmark score but a plainer question: when flagship capability is pushed down-market and prices are cut in half, what is changing in the competitive logic of the model market? In our earlier coverage of the GPT-6 Astra launch, we analyzed competition at the "capability ceiling" level; this article examines its other side, competition at the "unit cost" level.
2. The Ledger: Prices Cut in Half
Let us settle the numbers first. According to OpenAI's official disclosure as relayed by ai-bot.cn and QbitAI, the GPT-6 Sol API is priced at 2 dollars per million input tokens and 10 dollars per million output tokens, exactly half of the previous generation GPT-5.6 Sol at 4 and 20 dollars. GPT-6 Luna is priced at 0.1 dollars per million input tokens and 0.5 dollars per million output tokens, also 50 percent below its predecessor, and at a fixed price with no peak and off-peak variance.
Set against other releases from the same window, the sharpness of this cut becomes concrete (all figures below are each vendor's own official disclosure). Anthropic's Claude Opus 5.5 is priced at 4 dollars per million input tokens and 20 dollars per million output tokens, exactly double Sol; even though it cuts list prices by 20 percent, drops cached-input pricing by 60 percent (from 0.5 dollars to 0.2 dollars), and saves roughly 40 percent on total cost for typical tasks versus Opus 5, its absolute level remains a tier higher. You can read our report on Claude Opus 5 for that release's context. xAI's Grok 4.7 holds steady at 2 dollars per million input tokens and 6 dollars per million output tokens, flat against its predecessor. Placing the price sheets side by side:
| Model | Input (per million tokens) | Output (per million tokens) | Versus predecessor |
|---|---|---|---|
| GPT-6 Sol | 2 USD | 10 USD | 50 percent cheaper |
| GPT-6 Luna | 0.1 USD | 0.5 USD | 50 percent cheaper, fixed price |
| Claude Opus 5.5 | 4 USD | 20 USD | 20 percent cheaper, cached reads down 60 percent |
| Grok 4.7 | 2 USD | 6 USD | Flat |
Two layers of intent can be read from this table. The first: Sol, at less than half of Opus 5.5's output price, goes after the mid-to-high-end enterprise workflow market, and with caching and reasoning-effort control it pushes "paying for outcomes" further toward "paying for efficiency." The second: Luna's 0.1 dollar input price is no longer merely "cheaper than last generation." It carries the battle from the flagship tier straight into the lightweight tier and onto the home ground of Chinese open-weight models. For the broader arc of LLM price wars, see our price-war feature: every round of cuts redraws the boundary of "good enough and cheap."
3. Benchmark Reality: Cheaper Is Not Universally Better
A price-cut narrative is easily distorted into "stronger and cheaper at once," but the official figures, checked item by item, paint a more complicated picture. Per OpenAI's official disclosure as relayed by ai-bot.cn and QbitAI, the scorecard reads as follows.
For GPT-6 Sol: on AutomationBench, a cross-application enterprise workflow benchmark, the xhigh setting scored 33.2 percent, above Claude Opus 5's best result of 26.9 percent, with per-task cost at about 9 percent of Opus 5's; on Agents' Last Exam, spanning 55 industries, the max setting scored 56.4 percent, above Opus 5's best, at roughly 60 percent lower cost; on DeepSWE v1.1, a coding benchmark, the max setting reached 68.8 percent, just 1.1 percentage points behind Claude Fable 5's best result, at about one fifth the cost; on OSWorld 2.0, an offline computer-use benchmark, the xhigh setting scored 60.5 percent, comparable to Opus 5 at medium effort, with cost down about 80 percent; on FrontierCode, it matched Claude Fable 5.1's xhigh performance at lower cost. On factuality, the internal evaluation error count is about half that of GPT-5.6 Sol.
For GPT-6 Luna: on DeepSWE v1.1 the max setting scored 66.6 percent, approaching Opus 5 at medium effort while costing 93 percent less; on OSWorld 2.0 it exceeded GPT-5.6 Sol at medium effort at roughly one tenth the cost; on AutomationBench at high effort it improved 5.4 percentage points over GPT-5.6 Luna with cost down 58 percent; on factuality it reaches the level of GPT-5.6 Sol at about 1 percent of the cost.
Two points, however, must be stated plainly. First, on DeepSWE v1.1, Grok 4.7 scored 71.0 percent per xAI's official disclosure as relayed by Machine Heart, above Sol at 68.8 percent and Luna at 66.6 percent. On this coding benchmark, the cheaper Grok posted the highest score. It must be stressed that each vendor's scores come from its own official release with different harnesses, so cross-vendor conclusions are not warranted; still, this is enough to debunk any claim that Luna "tops the coding leaderboard." Second, and more importantly: DeepSeek V4.1 Flash's DeepSWE result is 74.2 percent, higher than Luna's 66.6 percent. What Luna wins on is price and fixed pricing, not benchmark scores. Any claim that Luna "comprehensively outclasses Chinese models" contradicts the officially disclosed numbers.
There is also a blank in the table: neither the parameter counts nor Terminal-Bench 4.0 results have been disclosed. In the same comparison set, Opus 5.5's Terminal-Bench 4.0 stands at 66.4 percent per Anthropic's official disclosure, and Grok 4.7's at 38.0 percent per xAI's official disclosure as relayed by Machine Heart, while the Sol and Luna cells can only read "not disclosed." The blank itself is a signal: OpenAI's narrative center this time is not peak capability but value for money. If your team is weighing a migration from another model, the arithmetic involved is similar to what we laid out in our Claude Fable 5.1 API migration guide: count the costs first, then move.
4. The Foundation of the 50 Percent Cut: Caching and Distillation
What sustains margins and reputation when prices are halved? The officially disclosed information points to two things: caching and downward distillation. This is the part practitioners should study most closely, because the source of the price cut is inference infrastructure efficiency, not amputated model capability.
Start with caching. Sol and Luna share one caching mechanism: higher default cache hit rates, a 90 percent discount on cached input reads, and support for "explicit breakpoints" that control where the cached prefix ends. For agentic workloads this is a tangible cost lever. Long system prompts, knowledge base documents, and multi-turn conversation histories re-enter the input side on every request; the portion that hits the cache costs one tenth of the list price, and the savings compound enough to reshape an invoice. In our earlier comparison of agentic caching costs we analyzed how caching strategy moves the bill, and a 90 percent discount sits at the aggressive end of the spectrum.
Next, reasoning-effort control. Sol supports effort tiers from low up through xhigh and max, and crucially allows switching effort mid-conversation without breaking the cache. The value of this detail is easy to underestimate. In the past, moving between effort levels often invalidated prefixes and forced cache recomputation, a hidden price hike. Now a request can stay at low effort while questions are simple, step up mid-conversation when a hard problem appears, and keep the established cache prefix intact. Efficiency and flexibility, for the first time, no longer force a trade-off.
Then comes distillation. Sol's positioning as a capability-downgraded version of the Astra base means flagship-level ability is packed into a smaller, cheaper model through distillation. Commercially, this means OpenAI did not need to train a separate top-tier model for Sol; it reuses Astra's capability assets and covers multiple price points through a same-origin, tiered lineup. The 50 percent cut comes partly from real inference savings delivered by cache hits and partly from distillation lowering the training and inference cost per unit of capability. In other words, this is not a "crippled model for a lower price" but "efficiency gains passed through into pricing." It also explains why OpenAI felt confident stating, in the same disclosure, that Sol outperforms the GPT-5.6 series on alignment metrics such as coding deception and guardrail circumvention. Being cheaper does not have to mean being less safe.
5. Ten Cents Across the Border: Hitting DeepSeek's Home Turf
Luna's most damaging weapon is not a benchmark score but where its price lands. Per DeepSeek's official disclosure as relayed by media reports, DeepSeek V4.1 Flash prices its API at 1 yuan per million input tokens off-peak and 2 yuan at peak, with output at 4 yuan off-peak and 8 yuan at peak. Luna, converted, lands at roughly 0.7 yuan per million input tokens and 3.5 yuan per million output tokens, per OpenAI's official disclosure as relayed by ai-bot.cn. It is cheaper still, and it carries no peak and off-peak variance, a fixed price around the clock.
The symbolism of this comparison outweighs the digits. Over the past two years, Chinese models, the DeepSeek line above all, have long played the role of the market's "price anchor": whenever overseas models raised prices or held their high ground, DeepSeek reminded the market with unsettlingly low pricing that intelligence can be cheap. Our report on DeepSeek V4.1 Flash going open source covers that story. With Luna, the move runs in reverse: a closed-source API is priced into DeepSeek's home territory, and lower still. For the first time, the price anchor is being redefined by an overseas closed-source vendor.
Restraint is required here, though. Cheaper does not mean better on balance: on DeepSWE v1.1, DeepSeek V4.1 Flash's 74.2 percent stands above Luna's 66.6 percent (each figure from its own vendor's official disclosure, on different harnesses, for reference only). DeepSeek also retains the moat of open weights and private deployment, something a closed API cannot replace. Luna's real threat lies in fixed, no-variance pricing combined with the enterprise trust that backs the OpenAI badge: for the long tail of applications that merely want stability, low cost, and out-of-the-box reliability, switching costs have dropped to the point where there is little reason to hesitate. The next phase of the anchor contest depends on whether Chinese models can follow on price and stability while preserving their open-source advantages, and on whether infrastructure-level efficiency competition, caching and effort control, becomes the new axis of comparison.
6. Closing: The Next Round of the Price War
Compress this release into one sentence: with its twin models, OpenAI demonstrated that a 50 percent price cut can rest not on sacrificing capability but on infrastructure efficiency, delivered through caching, distillation, and tiered pricing. Sol approaches flagship-grade practical performance at half the price; Luna, at 0.1 dollars per million input tokens, marches into DeepSeek's home territory, even though the latter still leads by 7.6 percentage points on DeepSWE. The verdict must be weighed separately on the scales of price and capability.
The advice to developers is plain: map your workload's token structure first, estimate what the cache hit rate will save you second, and only then decide which model fits. For industry watchers, the things to track are how Opus 5.5 and the flat-priced Grok 4.7 respond, and whether this round of cuts triggers another round of follow-through from Chinese model vendors. Price wars have no endgame, only rounds, and the next round will most likely be decided by efficiency rather than by leaderboards.