Claude Haiku 5.5 vs Claude Haiku 4.5
Claude Haiku 5.5 is the lower-priced of the two on Router One at base-tier rates (Claude Haiku 5.5: request input ≤ 100,000 tokens) — about 89% less on a 1M-input + 1M-output mix. Claude Haiku 5.5 offers 1.05M of context vs 200K for Claude Haiku 4.5. Both answer on /v1/chat/completions and /v1/messages (Claude Code). Claude Haiku 5.5 is in no plan and bills the wallet; Claude Haiku 4.5 counts against the Standard models allowance on Pro, Max and Ultra.
Claude Haiku 5.5 and Claude Haiku 4.5 compared on current per-token rates, context window, and capabilities — both callable through one OpenAI-compatible endpoint with per-request cost traces.
Claude Haiku 5.5 vs Claude Haiku 4.5: rates, context window, and capabilities
| Spec | Claude Haiku 5.5 | Claude Haiku 4.5 |
|---|---|---|
| Input / 1M tokens | $0.06 | $0.30 |
| Output / 1M tokens | $0.30 | $3.00 |
| Cached input / 1M tokens | $0.006 | $0.12 |
| Context window | 1.05M | 200K |
| Capabilities | Chat, Streaming, Tool calling, Vision | Chat, Streaming, Tool calling, Vision |
| Subscription plans | Not in any plan — wallet | Standard models — Pro, Max, Ultra |
| Detail page | Claude Haiku 5.5 | Claude Haiku 4.5 |
Claude Haiku 5.5: request input ≤ 100,000 tokens: Input $0.06 / 1M tokens, Output $0.30 / 1M tokens, Cache write $0.075 / 1M tokens, Cached input $0.006 / 1M tokens; request input > 100,000 tokens: Input $0.30 / 1M tokens, Output $1.50 / 1M tokens, Cache write $0.375 / 1M tokens, Cached input $0.03 / 1M tokens. The total input tokens in each request select the tier; its rates apply to the whole request, not only tokens above the threshold.
What does 1M tokens cost on Claude Haiku 5.5 vs Claude Haiku 4.5?
Base-tier rates: For a workload of 1M input plus 1M output tokens at current rates: Claude Haiku 5.5 comes to $0.36, Claude Haiku 4.5 comes to $3.30 — Claude Haiku 5.5 is about 89% cheaper on this mix. Real workloads skew heavily toward input tokens, so weigh the input rate by your own ratio; the cached-input row above is the posted catalog rate for that model.
Switch between Claude Haiku 5.5 and Claude Haiku 4.5
The API key, base URL and endpoints stay the same — Router One serves both on /v1/messages and /v1/chat/completions, not on /v1/responses — but Claude Haiku 4.5 code may still need changes. Per Anthropic's Claude Haiku 5.5 migration guide (checked 2026-10-10), Claude Haiku 5.5 returns a 400 for a manual thinking budget (type enabled with budget_tokens; use adaptive thinking with output_config.effort instead), for sampling settings it does not accept (omit temperature, top_p and top_k) and for an assistant prefill as the last turn, and the same text counts as approximately 30% more tokens than on Claude Haiku 4.5, so recount prompts, max_tokens and cost estimates instead of comparing per-token rates alone. On a Pro, Max or Ultra plan the billing changes too: Claude Haiku 4.5 counts against the Standard models allowance, while Claude Haiku 5.5, in no plan tier as of the 2026-10-10 plan response, bills the wallet per token. Read /blog/claude-haiku-5-5-api-guide before switching the model string.
curl https://api.router.one/v1/chat/completions \
-H "Authorization: Bearer sk-your-router-one-key" \
-H "Content-Type: application/json" \
-d '{"model": "anthropic/claude-haiku-5.5", "messages": [{"role": "user", "content": "Hello"}]}'
# Same request, other model — change one string:
# "model": "anthropic/claude-haiku-4.5"FAQ
Is Claude Haiku 5.5 cheaper than Claude Haiku 4.5?
Compared at base-tier rates (Claude Haiku 5.5: request input ≤ 100,000 tokens). Input: Claude Haiku 5.5 $0.06 vs Claude Haiku 4.5 $0.30 / 1M tokens; Claude Haiku 5.5 has the lower rate. Output: Claude Haiku 5.5 $0.30 vs Claude Haiku 4.5 $3.00 / 1M tokens; Claude Haiku 5.5 has the lower rate. Total cost depends on the workload's input, output and cache usage. Claude Haiku 5.5: request input ≤ 100,000 tokens: Input $0.06 / 1M tokens, Output $0.30 / 1M tokens, Cache write $0.075 / 1M tokens, Cached input $0.006 / 1M tokens; request input > 100,000 tokens: Input $0.30 / 1M tokens, Output $1.50 / 1M tokens, Cache write $0.375 / 1M tokens, Cached input $0.03 / 1M tokens. The total input tokens in each request select the tier; its rates apply to the whole request, not only tokens above the threshold. Rates change; the /models page is the live source of truth.
Is Claude Haiku 5.5 or Claude Haiku 4.5 included in a Router One subscription?
Claude Haiku 5.5 is in no Router One plan, so every call bills the wallet per token. Claude Haiku 4.5 counts against the Standard models allowance on Pro, Max and Ultra: Pro 5,000, Max 8,000 and Ultra 25,000 requests per 30-day cycle, shared by every model in that tier. Requests beyond a plan's allowance bill the wallet per token.
Can I switch between Claude Haiku 5.5 and Claude Haiku 4.5 without changing code?
The API key, base URL and endpoints stay the same — Router One serves both on /v1/messages and /v1/chat/completions, not on /v1/responses — but Claude Haiku 4.5 code may still need changes. Per Anthropic's Claude Haiku 5.5 migration guide (checked 2026-10-10), Claude Haiku 5.5 returns a 400 for a manual thinking budget (type enabled with budget_tokens; use adaptive thinking with output_config.effort instead), for sampling settings it does not accept (omit temperature, top_p and top_k) and for an assistant prefill as the last turn, and the same text counts as approximately 30% more tokens than on Claude Haiku 4.5, so recount prompts, max_tokens and cost estimates instead of comparing per-token rates alone. On a Pro, Max or Ultra plan the billing changes too: Claude Haiku 4.5 counts against the Standard models allowance, while Claude Haiku 5.5, in no plan tier as of the 2026-10-10 plan response, bills the wallet per token. Read /blog/claude-haiku-5-5-api-guide before switching the model string.
Where do these numbers come from?
Specs and prices on this page render from the live Router One catalog — the same data as the /models page — and refresh with it. Pricing methodology is documented on /pricing-methodology. Rates on this page were read from the live catalog on 2026-10-10 (UTC) and refresh within the hour.