Skip to content

Claude Haiku 5.5 vs Claude Sonnet 5.5

Claude Haiku 5.5 is the lower-priced of the two on Router One at base-tier rates (Claude Haiku 5.5: request input ≤ 100,000 tokens) — about 95% less on a 1M-input + 1M-output mix. Both carry a 1.05M context window. Both answer on /v1/chat/completions and /v1/messages (Claude Code).

Claude Haiku 5.5 and Claude Sonnet 5.5 compared on current per-token rates, context window, and capabilities — both callable through one OpenAI-compatible endpoint with per-request cost traces.

Claude Haiku 5.5 vs Claude Sonnet 5.5: rates, context window, and capabilities

SpecClaude Haiku 5.5Claude Sonnet 5.5
Input / 1M tokens$0.06$1.20
Output / 1M tokens$0.30$6.00
Cached input / 1M tokens$0.006$0.12
Context window1.05M1.05M
CapabilitiesChat, Streaming, Tool calling, VisionChat, Streaming, Tool calling, Vision
Subscription plansNot in any plan — walletNot in any plan — wallet
Detail pageClaude Haiku 5.5Claude Sonnet 5.5

Claude Haiku 5.5: request input ≤ 100,000 tokens: Input $0.06 / 1M tokens, Output $0.30 / 1M tokens, Cache write $0.075 / 1M tokens, Cached input $0.006 / 1M tokens; request input > 100,000 tokens: Input $0.30 / 1M tokens, Output $1.50 / 1M tokens, Cache write $0.375 / 1M tokens, Cached input $0.03 / 1M tokens. The total input tokens in each request select the tier; its rates apply to the whole request, not only tokens above the threshold.

What does 1M tokens cost on Claude Haiku 5.5 vs Claude Sonnet 5.5?

Base-tier rates: For a workload of 1M input plus 1M output tokens at current rates: Claude Haiku 5.5 comes to $0.36, Claude Sonnet 5.5 comes to $7.20 — Claude Haiku 5.5 is about 95% cheaper on this mix. Real workloads skew heavily toward input tokens, so weigh the input rate by your own ratio; the cached-input row above is the posted catalog rate for that model.

Switch between Claude Haiku 5.5 and Claude Sonnet 5.5

The API key, base URL and endpoints stay the same — Router One serves both on /v1/messages and /v1/chat/completions, not on /v1/responses — and per Anthropic's docs (checked 2026-10-10) both models run adaptive thinking by default, but three things differ: the API's default effort is medium on Claude Haiku 5.5 and high on Claude Sonnet 5.5, so set output_config.effort on /v1/messages instead of relying on the default; Claude Haiku 5.5 accepts thinking disabled at high effort or below, while Claude Sonnet 5.5 turns off up-front thinking only with between_tools, likewise only at high effort or below; and Claude Haiku 5.5 does not read thinking blocks that Claude Sonnet 5.5 wrote, so switch at a task boundary. Anthropic positions Claude Sonnet 5.5 for complex agentic coding and Claude Haiku 5.5 for narrowly scoped, high-volume work such as classification, summaries and subagent tasks. As of the 2026-10-10 plan response, both models are in no plan tier and bill the wallet per token.

compare.sh
curl https://api.router.one/v1/chat/completions \
  -H "Authorization: Bearer sk-your-router-one-key" \
  -H "Content-Type: application/json" \
  -d '{"model": "anthropic/claude-haiku-5.5", "messages": [{"role": "user", "content": "Hello"}]}'

# Same request, other model — change one string:
#   "model": "anthropic/claude-sonnet-5.5"

FAQ

Is Claude Haiku 5.5 cheaper than Claude Sonnet 5.5?

Compared at base-tier rates (Claude Haiku 5.5: request input ≤ 100,000 tokens). Input: Claude Haiku 5.5 $0.06 vs Claude Sonnet 5.5 $1.20 / 1M tokens; Claude Haiku 5.5 has the lower rate. Output: Claude Haiku 5.5 $0.30 vs Claude Sonnet 5.5 $6.00 / 1M tokens; Claude Haiku 5.5 has the lower rate. Total cost depends on the workload's input, output and cache usage. Claude Haiku 5.5: request input ≤ 100,000 tokens: Input $0.06 / 1M tokens, Output $0.30 / 1M tokens, Cache write $0.075 / 1M tokens, Cached input $0.006 / 1M tokens; request input > 100,000 tokens: Input $0.30 / 1M tokens, Output $1.50 / 1M tokens, Cache write $0.375 / 1M tokens, Cached input $0.03 / 1M tokens. The total input tokens in each request select the tier; its rates apply to the whole request, not only tokens above the threshold. Rates change; the /models page is the live source of truth.

Is Claude Haiku 5.5 or Claude Sonnet 5.5 included in a Router One subscription?

Neither Claude Haiku 5.5 nor Claude Sonnet 5.5 is in a Router One plan, so every call to either bills the wallet per token.

Can I switch between Claude Haiku 5.5 and Claude Sonnet 5.5 without changing code?

The API key, base URL and endpoints stay the same — Router One serves both on /v1/messages and /v1/chat/completions, not on /v1/responses — and per Anthropic's docs (checked 2026-10-10) both models run adaptive thinking by default, but three things differ: the API's default effort is medium on Claude Haiku 5.5 and high on Claude Sonnet 5.5, so set output_config.effort on /v1/messages instead of relying on the default; Claude Haiku 5.5 accepts thinking disabled at high effort or below, while Claude Sonnet 5.5 turns off up-front thinking only with between_tools, likewise only at high effort or below; and Claude Haiku 5.5 does not read thinking blocks that Claude Sonnet 5.5 wrote, so switch at a task boundary. Anthropic positions Claude Sonnet 5.5 for complex agentic coding and Claude Haiku 5.5 for narrowly scoped, high-volume work such as classification, summaries and subagent tasks. As of the 2026-10-10 plan response, both models are in no plan tier and bill the wallet per token.

Where do these numbers come from?

Specs and prices on this page render from the live Router One catalog — the same data as the /models page — and refresh with it. Pricing methodology is documented on /pricing-methodology. Rates on this page were read from the live catalog on 2026-10-10 (UTC) and refresh within the hour.

More comparisons