← All posts

Artificial intelligence

Anthropic Cut Haiku's Price to Exactly Match GPT-6 Luna. The Fine Print at 100,000 Tokens Is Where They Part Company

Claude Haiku 5.5 arrives at $0.10 per million input tokens — a tenth of its predecessor, and the same base rate as OpenAI's GPT-6 Luna. The new two-tier pricing, and the labs missing from the benchmark table, complicate the headline.

MAI
Anthropic's announcement artwork for Claude Haiku 5.5, from the company's own launch page.

Anthropic released Claude Haiku 5.5 on Wednesday, and the number that matters is not a benchmark. The model's base rate is $0.10 per million input tokens and $0.50 per million output — a tenth of what Haiku 4.5 charged, and the same figure, to the cent, that OpenAI charges for GPT-6 Luna. Two labs that compete on almost nothing else have landed on an identical floor price for a small model.

The prices

Per million tokensHaiku 5.5 (≤100k)Haiku 5.5 (>100k)Haiku 4.5Sonnet 5.5
Input$0.10$0.50$1.00$2.00
Output$0.50$2.50$5.00$10.00
Cache read$0.01$0.05$0.10$0.10
Cache write$0.125$0.625$1.25$2.50

The headline is a 90% cut. The honest figure is Anthropic's own: roughly 75% lower cost on an average workload. Sonnet 5.5's cache-read price was halved to $0.10 in the same announcement, which Anthropic estimates at about 20% off a typical agent workload, and Max and Team subscribers start receiving monthly API credits this week — $100, $200, and up to $500 shared across a team.

Haiku now has a cliff

Haiku 4.5 charged one rate whatever the request size. Haiku 5.5 charges five times as much once a request passes 100,000 input tokens. Anthropic says about 90% of Haiku 4.5 requests fell under that threshold, which is how the 75% average gets built — and it also means the remaining tenth of existing traffic sees a far smaller discount. An updated tokenizer that consumes slightly more tokens per task is folded into the same estimate.

This is where the match with GPT-6 Luna stops being a match. Luna's higher tier starts at 272,000 input tokens; Haiku's starts at 100,000. The two models share a sticker price and not a cost curve, and the workloads most likely to notice are exactly the ones a cheap model gets pointed at: long-document summarization, large-context retrieval, agent loops that accumulate history. Teams pricing a migration off the headline numbers will get the wrong answer.

The comparison set stops at OpenAI

Anthropic's evaluation table sets Haiku 5.5 against Haiku 4.5, GPT-6 Luna and Sonnet 5.5, and nothing else. All of these are the company's own reported results, not independently reproduced.

EvaluationHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
GDPval-AA v2.11,6207351,4371,840
OSWorld 2.1 (offline subset)72.4%15.7%48.9%83.9%
Terminal-Bench 4.039.2%0.0%16.4%70.6%
FrontierCode 1.146.4%—42.4%52.1%
Humanity's Last Exam (no tools)45.9%10.2%—56.9%

The generational jumps are large enough to be interesting on their own — Haiku 4.5 scored zero on Terminal-Bench 4.0, and the new model clears 39%. But that 39.2% is the maximum-effort result. At the default medium effort, Anthropic's own figure is around 20%. The headline agentic number is not the configuration most developers will run.

The omission in the table is the Chinese labs. The New Stack points out that Z.ai's GLM-5.3-Flash scores 1,647 on GDPval-AA v2.1 against Haiku 5.5's 1,620, and that Alibaba's Qwen3.7 Flash is cheaper per token. Neither appears in Anthropic's comparison. A frontier lab is entitled to pick its own reference points, but a small-model release framed around being the cheapest and most capable has a weaker claim to both than the chart suggests.

What it is actually for

Anthropic is not pitching Haiku 5.5 as the thing that handles classification and routing while the big models do the work. It is pitching it as a subagent — context compaction, database queries, browser use, live support — running alongside Opus 5.5 and Sonnet 5.5. It is the first Haiku-class model with an adjustable effort setting, and the Python and TypeScript SDKs gain beta computer-use and browser-use support. Box reports 11 points of improvement over Haiku 4.5 at roughly half the latency; HubSpot reports 92.8% averaged over three runs of its CRM evaluation.

The safeguards moved in the opposite direction from the price. Cybersecurity restrictions are tighter than Haiku 4.5's — penetration testing is blocked under the standard configuration, with broader access available through Anthropic's verification programs. For a model explicitly aimed at high-volume automation, that is a deliberate trade: the cheapest model in the lineup is also the one with the least latitude on offensive security work.

The read

A 90% price cut is not a capability announcement dressed up as a pricing one. It is a pricing announcement. Small models are now the layer where the agent business is won or lost — they are the ones called thousands of times per task — and two of the leading labs have independently decided that layer is worth $0.10 per million tokens. The margin is gone on purpose. What Anthropic is buying with it is the subagent slot inside every Opus and Sonnet deployment, and the fine print at 100,000 tokens is where it hopes to earn some of that back.

Sources: Anthropic — Introducing Claude Haiku 5.5 · VentureBeat · The New Stack

Keep reading