← All posts

Artificial intelligence

OpenAI Halved Its Token Prices Ninety Minutes After Anthropic Cut Its Own. Its Benchmark Table Argues Cost, Not Capability

OpenAI released GPT-6 Sol and Luna on Tuesday at half the API price of the models they replace, roughly ninety minutes after Anthropic cut Opus pricing by 20%. The company's own comparison table shows Sol losing to Claude on two of three headline benchmarks and argues the release on cost per task instead.

MAI
OpenAI's announcement graphic for GPT-6 Sol and Luna, the hero image from the company's release page.

OpenAI released GPT-6 Sol and GPT-6 Luna on Tuesday and halved the API price of both. By TechCrunch's count the announcement landed roughly ninety minutes after Anthropic shipped Claude Opus 5.5 with a 20% cut of its own. Two frontier labs repricing downward inside a single afternoon is not a scheduling coincidence, and it is the more interesting fact about the day than either model. The models are capable. The pricing is the announcement.

Two models, half the price

Sol is the coding and agent tier. Luna is the high-volume tier, aimed at summarisation, extraction and the clerical work that runs at scale. Both apply the training approach behind GPT-6 Astra, OpenAI's flagship, at a fraction of Astra's cost — what the company calls "advancing the frontier on cost efficiency."

Per million tokensGPT-6 SolGPT-6 Luna
Input$2$0.10
Output$10$0.50
Cached input reads90% discount90% discount

OpenAI describes both as 50% below the GPT-5.6 models they replace; The New Stack puts those predecessors at $4/$20 and $0.20/$1.20. The company attributes the cut to infrastructure rather than margin: "Improvements in caching and inference let us serve these models at lower cost." The 90% cached-read discount is the line that matters most for retrieval-heavy agents, where the same context is paid for on every step.

There is a second, quieter cost effect. OpenAI says the new models spend fewer tokens reaching an answer, with shorter and less jargon-laden responses. Anthropic made the same argument about Opus 5.5 the same morning. When both vendors claim their real price cut is larger than their published one because their model is less verbose, the rate card has stopped being the number that decides anything. Neither claim has been tested head to head by a third party yet.

The benchmark table argues cost, not capability

OpenAI's own published comparisons are worth reading closely, because of what they concede.

BenchmarkGPT-6 SolAnthropic comparisonOpenAI's cost claim
Agents' Last Exam V1 (max)56.4%Claude Opus 5: 60%60% lower cost
DeepSWE v1.1 (max)68.8%Claude Fable 5: 69.9%~80% lower cost
OSWorld 2.0 offline (xhigh)60.5%Claude Opus 5: 60.3%80% lower cost
FrontierCode 1.1matches Fable 5.1 xhighClaude Fable 5.1lower cost
AutomationBench 1.0.6 (xhigh)33.2%$0.27 per task

Read down the comparison column. Sol loses to Claude Opus 5 on Agents' Last Exam, loses to Fable 5 on DeepSWE, and ties on OSWorld. OpenAI is not claiming the best model on these tasks. It is claiming a model close enough to the best one at a fifth of the price, and it has chosen to present every row normalised by cost per task rather than by score.

That is a real change in how a frontier lab argues. The cost column used to be a footnote to the capability column. Here it is the column. Luna makes the point more bluntly: 66.6% on DeepSWE, which OpenAI says it delivers at 93% less than Opus 5. A capability gap that costs a twentieth as much is a different product decision than the same gap at parity pricing.

One number in the release deserves separate mention. OpenAI reports Sol's rate of misleading claims about its own code falling from 10.4% to 1.3% against its predecessor — an alignment result about agents that report success they did not achieve, which is the failure mode that makes long autonomous runs expensive to trust. The company notes these evaluations are deliberately adversarial and not representative of typical use.

"Pacing the frontier" lasted about a week

The competitive context makes the timing harder to read charitably. Anthropic's Dario Amodei called in mid-September for labs to pace the frontier rather than sprint it, a proposal Sam Altman publicly supported. Opus 5.5 was Anthropic's first release since that call. OpenAI's answer arrived the same afternoon, priced underneath it.

Whatever pacing the frontier turns out to mean in practice, it evidently does not include declining to reprice on a competitor's launch day. It is also notable that the mid-tier GPT-6 Terra is absent from this rollout, which 9to5Mac reads as OpenAI consolidating from four tiers to three — a simplification that makes the per-tier price comparison against Anthropic cleaner.

Who pays for the price war

Ara Kharazian, an economist at Ramp, put the dynamic plainly to Fortune:

OpenAI and Anthropic are engaged in a price war that is driving down the price of AI.

The warning attached to it is the part worth sitting with. The bull case for these companies assumes models get more expensive as they get more capable. Technology markets rarely work that way, and two vendors cutting prices on the same afternoon while capability converges within a few points of each other is the shape of a market where the product is becoming a commodity input.

For buyers that is straightforwardly good, and it is why the cost-normalised benchmark table is the right way to read the release. For the sellers, the differentiator is no longer which model is best. It is cost per completed task, and the margin on that is going in one direction.

Sources: Introducing GPT-6 Sol and Luna (OpenAI) · OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakes (TechCrunch) · OpenAI releases GPT-6 Sol and Luna — and cuts token prices in half (The New Stack) · Anthropic releases Claude Opus 5.5 and OpenAI counters with two cheaper GPT-6 models (SiliconANGLE) · What slowdown? OpenAI, Anthropic release dueling models as AI price wars heat up (Fortune) · OpenAI upgrading ChatGPT and Codex with two more GPT-6 models (9to5Mac)

Keep reading