Anthropic released Haiku 5.5 on 7 October and says it is its cheapest small model: around 75% cheaper than Haiku 4.5 on average. The price depends on prompt length: up to 100 thousand tokens, input is $0.10 per million, above that $0.50.
- 7 October 2026: Anthropic released Claude Haiku 5.5, in its words the cheapest, fastest and most capable small model it has released. It is available on all platforms, including AWS, Google Cloud and Azure, as claude-haiku-5-5.
- Prices per million tokens, for prompts up to 100 thousand / over: input $0.10 / $0.50, output $0.50 / $2.50. Haiku 4.5 cost $1 and $5. According to Anthropic, Haiku 5.5 costs around 75% less to run on average.
- On Terminal-Bench 4.0 Haiku 5.5 scores 39.2% against 70.6% for Sonnet 5.5, according to Anthropic. The same day Sonnet 5.5 cache reads got half as expensive, and Anthropic announced a monthly API credit for Max and Team subscribers, rolling out gradually from the week of 7 October.
Ten cents per million input tokens. With Haiku 4.5, that was only the price of a cache read. Now it is the price of all input, but only while the prompt is up to 100 thousand tokens. Above that, input is 50 cents, five times more.
The math is simple until you ask about length. On a short prompt, both input and output fall ten times against Haiku 4.5: from $1 to $0.10 and from $5 to $0.50. Above 100 thousand tokens the difference is two times. Anthropic says so itself, in a footnote: 90% cheaper up to 100 thousand tokens, 50% above. That is where the average "around 75%" comes from: it accounts for the share of short requests and for the new tokenizer, which makes Haiku 5.5 use slightly more tokens for the same task. Your bill depends on your prompts.
A small model is for small work. Terminal-Bench 4.0, a test of working in a terminal, gives it 39.2% and Sonnet 5.5 70.6%, according to Anthropic. The previous Haiku scored 0.0%. The jump is big. There is still a way to go to Sonnet, and Anthropic says so itself: for complex agentic coding Sonnet 5.5 and Opus 5.5 remain the better choice, while Haiku is for narrower tasks like compaction, summaries and subagent work.
The second change that day is for those who already pay for Sonnet 5.5: cache reads dropped by half, to $0.10 per million tokens. Sonnet's input and output prices do not move. According to Anthropic, this makes it around 20% cheaper on most agentic work. And if you are on Max or Team, from the week of 7 October Anthropic is gradually rolling out a monthly API credit too, usable on all models.
Before you switch models, pull from your logs how many of your prompts are up to 100 thousand tokens. With the old Haiku they were around 90% of requests, according to Anthropic, and yours may be different.