xpu liveBETA
← Back to Journal
AI Industry News

Claude Haiku 5.5: A Major API Price Cut

2 min read
Compact inference hardware illustrating Claude Haiku 5.5 efficiency
AI-generated editorial illustration; not a photograph of the reported event.

Claude Haiku 5.5 brings a substantial price cut to Anthropic’s small-model API. The company launched it on October 7 with lower token rates and new benchmark results. However, the biggest discount applies to prompts up to 100,000 tokens. Anthropic’s announcement explains the two pricing tiers.

Event date: October 7, 2026 · Sources checked: October 8, 2026

What Claude Haiku 5.5 costs

For prompts up to 100,000 tokens, input costs $0.10 and output costs $0.50 per million tokens. Haiku 4.5 charged $1 and $5 respectively, so both listed rates fall by 90%. For longer prompts, Haiku 5.5 charges $0.50 for input and $2.50 for output per million tokens: a 50% reduction against those earlier rates.

Anthropic reports stronger results than GPT-6 Luna on several published evaluations, including computer use and agentic coding. These are company-published comparisons under specific test conditions. Meanwhile, the model is available on the Claude Platform, AWS, Google Cloud and Azure. The direct API identifier is claude-haiku-5-5.

How to compare the savings

A short-request workload using one million input tokens and one million output tokens would have a $0.60 base token bill under the first tier. The same token counts at Haiku 4.5’s listed rates would cost $6. This is a simple arithmetic example; caching, tokenization, retries and other charges can change an actual invoice.

Therefore, a team should measure cost per completed task as well as token rates. A cheaper response that needs repeated corrections may save less than expected. Conversely, a small model that completes routine classification reliably could make frequent requests more affordable.

Claude Haiku 5.5 — xpu live analysis

Our view is that this launch deserves a workload-specific comparison. Start with representative summaries, extraction requests and support queries. Next, record accuracy, response time and the total tokens used, including failed attempts. Keep the input-length boundary visible in any pricing spreadsheet.

Similarly, compare the model with alternatives using the same acceptance criteria. A benchmark score can help choose a test candidate, but your own task mix determines practical value. For more complex work, test whether routing difficult requests to a larger model improves the overall result.

Finally, preserve the model version and date alongside each measurement. That makes later price changes easier to interpret and helps explain whether savings came from cheaper tokens, shorter answers or fewer retries.

Sources and further reading

Related on xpu live: How to compare AI API pricing.