Ten Times Cheaper Per Token, Four Times Cheaper On Average

Anthropic released Claude Haiku 5.5 on 7 October 2026 at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Haiku 4.5 cost $1 and $5, so the list price fell by a factor of ten.

Anthropic's own launch post does not claim ten times, though. It says Haiku 5.5 costs around 75 percent less to run on average, which is a factor of four. The gap is the part to plan around, because a cheaper token buys a model that reasons by default, over text that is now counted differently.

Ten divided by the new tokenizer factor of 1.25 still leaves eight for the same text, so even before reasoning the headline overstates how far a typical bill will fall.

The 100,000-Token Line And The Tokenizer Together

Above 100,000 prompt tokens Haiku 5.5 charges $0.50 for input and $2.50 for output, five times the base rate. Anthropic says prompts up to that size made up around 90 percent of requests to its previous Haiku model, so the base rate is built for the common case.

Simon Willison, who tested the model on launch day, measured that the same long prompt uses about 1.25 times as many tokens on Haiku 5.5 as on Haiku 4.5. Divide 100,000 by 1.25 and you get 80,000. A prompt that sat safely under the line at 80,000 old tokens now lands exactly on it, and anything longer pays the higher tier.

GPT-6 Luna, which Willison treats as the direct rival, charges the same $0.10 and $0.50 but holds that rate up to 272,000 tokens and then rises only to $0.20 and $0.75. For long documents the comparison flips in Luna's favour.

Where Each Model's Price Steps Up

The five price points below show how each model is billed per million tokens and where its rate changes.

ModelInputOutputRate changes
Haiku 4.5$1.00$5.00No step
Haiku 5.5, up to 100K$0.10$0.50Above 100K tokens
Haiku 5.5, over 100K$0.50$2.50Already stepped
Sonnet 5.5$2.00$10.00Cache reads halved to $0.10
GPT-6 Luna, up to 272K$0.10$0.50Above 272K: $0.20 and $0.75

Anthropic also halved the cache-read price of Sonnet 5.5 from $0.20 to $0.10 per million tokens, which it says makes Sonnet around 20 percent cheaper on most agentic work.

The performance side explains why teams will switch anyway. On Anthropic's own launch benchmarks Haiku 5.5 reaches 72.4 percent on OSWorld 2.1 against 15.7 percent for Haiku 4.5 and 48.9 percent for GPT-6 Luna. These are vendor-run figures, so test them on your own tasks.

What To Do Before You Move A Workload

Count tokens on your own prompts with the new tokenizer before you compare invoices, because the tenfold headline assumes the old count. Then run real traffic at low and medium effort. Willison's single test prompt cost 0.0936 cents at low effort and 3.3826 cents at max effort, a 36-fold spread on the same request.

Treat the new monthly API credits as a trial budget, not a discount. Anthropic's support page lists $100 a month for Max 5x, $200 for Max 20x and up to $500 pooled for Team. The credits do not roll over, they cover the API, Managed Agents and the Agent SDK but not interactive Claude Code sessions, and they do not apply to Amazon Bedrock, Vertex AI or Microsoft Foundry.

Anthropic publishes these prices in US dollars, so European and UK buyers also carry the exchange rate on every invoice.

Servola Journal

We do this for everyone trying to keep up with what technology is doing to our lives. The people who build it, and the people it happens to. The Servola Journal exists so that what we learn belongs to all of them.

Nobody pays us for this. No ads, no paywall, free to everyone. We just believe that understanding what's happening to all of us shouldn't depend on who can afford to pay for it.

If it gave you something today, tell us to keep going. Follow us, leave a like, or write a positive comment. We read every one, and they are what keeps us going.