Lower list price in the short tier
Compared with Haiku 4.5 at the same token counts. Actual task savings can differ because token usage may change.
CLAUDE API · VERIFIED OCTOBER 8, 2026
Estimate one request and monthly spend using Anthropic’s two prompt-length tiers. Compare Haiku 5.5 with Haiku 4.5 without sending data anywhere.
100,000 prompt tokens. This request is in the lower price tier.
Same workload, both Haiku 5.5 tiers
The threshold applies to prompt length for each request—not total monthly tokens.
| Per 1M tokens | Haiku 5.5 ≤100K | Haiku 5.5 >100K | Haiku 4.5 |
|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 |
| Output | $0.50 | $2.50 | $5.00 |
| Cache read | $0.01 | $0.05 | $0.10 |
| Cache write · 5 min | $0.125 | $0.625 | $1.25 |
Compared with Haiku 4.5 at the same token counts. Actual task savings can differ because token usage may change.
Apply the Batch API switch in the calculator to model asynchronous workloads.
The optional US-only inference setting is priced at 1.1 times the global rate.
cost = input × input_rate + output × output_rate
+ cache_read × read_rate + cache_write × write_rateEvery token quantity is divided by one million. Monthly cost multiplies the result by requests per month. Batch and regional multipliers are applied last.
“90% cheaper” is a list-price comparison for requests up to 100K prompt tokens. Anthropic says Haiku 5.5 costs around 75% less on average; actual workloads depend on prompt length and tokenization.
On the Claude Platform, prompts up to 100K use $0.10 per million input tokens and $0.50 per million output tokens. Prompts over 100K use $0.50 input and $2.50 output.
At the published list prices, it is 90% cheaper for equivalent token counts in the short tier and 50% cheaper in the long tier. Actual task-level savings can differ.
This calculator treats normal input, cache reads and cache writes as prompt tokens for each individual request. Monthly token totals do not choose the tier.
claude-haiku-5-5 on the Claude Platform. Check Anthropic’s current documentation before production use.
Yes. The published cache-read and five-minute cache-write rates are included above. Other cache durations and providers may use different prices.