Claude Haiku 5.5, which Anthropic released on October 7, 2026, is in the Router One catalog as anthropic/claude-haiku-5.5 (listed 2026-10-10), and the gateway also accepts Anthropic's own id, claude-haiku-5-5. Send "model": "claude-haiku-5-5" to https://api.router.one/v1/messages with a Router One key — or point an OpenAI-compatible client at https://api.router.one/v1 — and the request is routed, metered and logged like every other model. No Pro, Max or Ultra tier lists it as of the 2026-10-10 plan response, so every call bills per token to your wallet; the Claude Haiku 5.5 model page carries the live rates.
This guide covers what the catalog lists for the id, how to choose between Claude Haiku 5.5, Claude Haiku 4.5, Claude Sonnet 5.5 and Gemini 3.5 Flash Lite, how billing and plans apply, what Claude Code 2.1.293 changed, the first request on each endpoint, and what changes from Claude Haiku 4.5. Catalog and plan observations are dated; the live catalog is the source of truth for a new request.
Claude Haiku 5.5 at a glance
| Field | What the catalog lists (2026-10-10) |
|---|---|
| Catalog id | anthropic/claude-haiku-5.5, listed 2026-10-10 |
| Short id (alias) | claude-haiku-5-5, Anthropic's own model id |
| Context window | 1,048,576 tokens |
| Input / output | text and image in, text out |
| Capability flags | chat, streaming, tool calling, vision |
| Price lines | two, split at 100,000 prompt tokens — both on the model page |
| Endpoints | POST /v1/messages (Anthropic-native, recommended) and POST /v1/chat/completions |
| Not served on | POST /v1/responses |
| Channel | default channel only — no aws/, vertex/ or azure/ version |
| Plans | no Pro, Max or Ultra tier (2026-10-10 plan response) — wallet billing |
Anthropic's Claude Haiku 5.5 overview (checked 2026-10-10) gives Anthropic's own spec: an October 7, 2026 release, a 1M-token context window, text and image input, up to 128K output tokens on the synchronous Messages API, a June 2026 reliable knowledge cutoff, and a model "built for high-volume, latency-sensitive work such as classification, routing, extraction, and subagent tasks". Adaptive thinking is on by default, and depth is set with output_config.effort — low, medium, high, xhigh or max — with medium as the API default (effort docs); thinking: {"type": "disabled"} still turns thinking off at high effort or below. Those are Anthropic's statements; Router One vouches for the catalog entry and how the gateway routes and bills it.
Claude Haiku 5.5, Haiku 4.5, Sonnet 5.5 or Gemini 3.5 Flash Lite?
Anthropic's announcement (October 7, 2026; checked 2026-10-10) says Claude Haiku 5.5 "is designed for high-volume, cost-sensitive tasks", handling "quick and repetitive workloads (like summaries, compactions, database queries, and classification requests)", and that it "pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work". The same page draws the limit: "Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks", while Haiku 5.5 "is best suited to more narrowly scoped tasks" such as "compaction, summarization, or subagent work". Anthropic reports 39.2% for Claude Haiku 5.5 on Terminal-Bench 4.0 (Claude Haiku 4.5: 0.0%) and 72.4% on the offline subset of OSWorld 2.1 (Claude Haiku 4.5: 15.7%). Those are Anthropic's results, not Router One measurements: run your own prompts before you move a workload.
What Router One adds to the choice:
- Plan quota or wallet. As of the 2026-10-10 plan response, Claude Haiku 4.5 and Gemini 3.5 Flash Lite are in the Standard models tier of the Pro, Max and Ultra plans, so on a plan their requests draw plan quota; Claude Haiku 5.5 and Claude Sonnet 5.5 are in no plan tier and bill the wallet. Anthropic now lists Claude Haiku 4.5 as a legacy model that is "still available", and Anthropic's model deprecations page (checked 2026-10-10) dates its retirement not sooner than October 15, 2026 and promises at least 60 days' notice before a publicly released model retires.
- Posted rates and token counts. Router One posts its own rates for each model, and Claude Haiku 5.5's sit on two lines split at 100,000 prompt tokens. Per Anthropic's what's-new page (checked 2026-10-10), the same text counts as approximately 30% more tokens on Claude Haiku 5.5 than on Claude Haiku 4.5, and thinking blocks from earlier turns stay in context as input, so compare the cost of a whole task rather than per-token rates alone. The comparison pages below render the live figures.
- Switching mid-conversation. Per Anthropic's migration guide (checked 2026-10-10), Claude Haiku 5.5 reads the thinking blocks of Claude Haiku 4.5, Claude Sonnet 5, Claude Opus 4.8 and earlier models, so a conversation moved onto it from one of them keeps its reasoning. It does not read blocks from Claude Opus 5, Claude Opus 5.5, Claude Sonnet 5.5 or any Claude Fable model (the API drops those without an error), while Claude Sonnet 5.5 and Claude Opus 5.5 read the blocks Claude Haiku 5.5 writes on the Claude API and Google Cloud. Thinking blocks are also tied to the conversation, so switch at a task boundary rather than counting on the reasoning to carry over.
Three comparison pages render the spec sheets and live rates side by side:
- Claude Haiku 5.5 vs Claude Haiku 4.5 — the generation step, and wallet billing against Standard models plan quota.
- Claude Haiku 5.5 vs Gemini 3.5 Flash Lite — the cross-vendor choice; Router One serves Gemini 3.5 Flash Lite on
/v1/chat/completionsonly, so Claude Code can use only the Claude side, and as of the 2026-10-10 plan response Gemini 3.5 Flash Lite is in the Standard models tier while Claude Haiku 5.5 bills the wallet. - Claude Haiku 5.5 vs Claude Sonnet 5.5 — Haiku or Sonnet for a given workload; as of the 2026-10-10 plan response both are in no plan tier and bill the wallet.
How billing works
Claude Haiku 5.5 is priced on two rate lines. Per Anthropic's pricing docs (checked 2026-10-10), a request whose prompt is over 100,000 tokens pays the higher prices, input and output alike, and each request is priced on its own. The Router One catalog splits its own posted rates at the same 100,000-token threshold, and the model page shows both lines.
Adaptive thinking is on by default, and the model decides how much to think, steered by effort. Thinking tokens are billed as output tokens (pricing facts), also when the thinking text is not returned, which is the default display: "omitted" on this model (Anthropic's thinking docs, checked 2026-10-10). Effort is therefore a cost lever as much as a quality one. The API default is medium, and Anthropic's effort docs (checked 2026-10-10) suggest starting there for most work, including agentic coding, and using low for chat, short tool tasks and simple, high-volume requests. Measure per task before you set a default.
This guide prints no per-token figures because they go stale. The Claude Haiku 5.5 model page shows both lines, and the comparison pages above render them against Claude Haiku 4.5, Gemini 3.5 Flash Lite and Claude Sonnet 5.5.
Do subscription plans cover Claude Haiku 5.5?
Not as of the 2026-10-10 plan response: no tier of Pro, Max or Ultra lists claude-haiku-5-5, so Claude Haiku 5.5 calls bill per token to your wallet balance at the posted rates, whether or not you hold a plan. Claude Haiku 4.5 stays in the Standard models tier of all three plans, and as of the 2026-10-10 plan response Gemini 3.5 Flash Lite is in that tier too. Plan model lists change; the pricing page shows the live lists and each plan's allowance.
That matters most for Claude Code.
Claude Code: the haiku alias now sends claude-haiku-5-5
Claude Code 2.1.293 (October 7, 2026) made Claude Haiku 5.5 the default Haiku model on the Anthropic API, per the Claude Code changelog, and Claude Code treats a gateway set through ANTHROPIC_BASE_URL as the Claude API, per its gateway protocol docs (both checked 2026-10-10), so a session pointed at Router One sends claude-haiku-5-5 wherever it uses the haiku alias.
On 2026-10-10 the npm dist-tags and the native installer's channel pointers had latest on 2.1.296 and stable on 2.1.287. Per the model configuration docs (checked 2026-10-10), versions before 2.1.293 resolve haiku to Claude Haiku 4.5, so a Claude Code on the stable channel still did on that date, and the docs say to "Use v2.1.293 or later with Haiku 5.5"; claude --version shows which one you run.
The optional Haiku line goes in your shell profile — or in the env block of ~/.claude/settings.json if the one-click install wrote your setup, since the install script sets only ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN, not a model:
export ANTHROPIC_DEFAULT_HAIKU_MODEL=claude-haiku-4-5
The line is optional for everyone, on a plan or on the wallet. Per Claude Code's changelog, model configuration and subagents docs (Claude Code 2.1.296, checked 2026-10-10), from 2.1.293 the haiku alias resolves to Claude Haiku 5.5 (claude-haiku-5-5), which Router One lists as of the 2026-10-10 catalog but no plan tier lists in the 2026-10-10 plan response. Without the line, /model haiku, subagents defined with model: haiku and the built-in claude-code-guide subagent (which the subagents docs list as running on Haiku) therefore bill per token to the wallet; with it they run on Claude Haiku 4.5, which is in the Standard models tier of all three plans. The line also decides where background tasks run: per the gateway protocol docs (checked 2026-10-10), a gateway session that authenticates with ANTHROPIC_AUTH_TOKEN, as the Router One setup does, runs them on the main model unless ANTHROPIC_DEFAULT_HAIKU_MODEL pins one. The other alias pins — the Fable line for everyone, the Sonnet and Opus pins for plan holders — are in the Claude Code setup guide and Claude Code in China.
To run Claude Haiku 5.5 as the main model, use Anthropic's id rather than the catalog id: the model configuration docs (checked 2026-10-10) give it as claude-haiku-5-5 — /model claude-haiku-5-5 in a session, or claude --model claude-haiku-5-5 from your shell — and behind a custom ANTHROPIC_BASE_URL Claude Code gives each model it recognizes the context window that model has on the Anthropic API: 1M for Haiku 5.5, compacting at about 967K tokens by default. Per the same docs, Claude Code runs Haiku 5.5 at medium effort by default and cannot turn its thinking off (MAX_THINKING_TOKENS=0 has no effect on it), and a Haiku 5.5 request "costs more per token when its prompt is longer than 100K tokens". Once a session's context passes 100,000 tokens, its requests carry prompts above that split until it compacts; the model page shows both lines.
Send the first request
- Create a key. Dashboard → API Keys → Create Key. Keys look like
sk-.... For a trial, give the key amaxSpendcap: it cannot spend past that amount, and your other keys keep working (per-key cost tracking). - Keep a wallet balance. Claude Haiku 5.5 calls bill the wallet whether or not you hold a plan.
- Pick the base URL for your client. Anthropic-native SDKs and tools take
https://api.router.one(the SDK'sbase_url, orANTHROPIC_BASE_URL) and call/v1/messagesunder it; OpenAI-compatible SDKs takehttps://api.router.one/v1. - Call the model. The examples use
claude-haiku-5-5; for plain HTTP calls the catalog idanthropic/claude-haiku-5.5works too.
Messages API, streaming, with the effort level set to low: medium is the API default, and Anthropic's effort docs (checked 2026-10-10) suggest low for chat, short tool tasks and simple, high-volume requests such as this one:
curl https://api.router.one/v1/messages \
-H "x-api-key: sk-your-api-key" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-haiku-5-5",
"max_tokens": 16000,
"stream": true,
"output_config": {"effort": "low"},
"messages": [{"role": "user", "content": "Classify this support ticket as billing, bug or how-to, with one line of reasoning: The app signs me out every time I switch tabs."}]
}'
The same call from the Anthropic Python SDK (current release: pip install -U anthropic) — only base_url and the key change. A response can open with one or more thinking blocks, so read content by block type, not by position:
import anthropic
client = anthropic.Anthropic(
base_url="https://api.router.one",
api_key="sk-your-api-key",
)
with client.messages.stream(
model="claude-haiku-5-5",
max_tokens=16000,
output_config={"effort": "low"},
messages=[{"role": "user", "content": "Classify this support ticket as billing, bug or how-to, with one line of reasoning: The app signs me out every time I switch tabs."}],
) as stream:
message = stream.get_final_message()
for block in message.content:
if block.type == "text":
print(block.text)
print(message.stop_reason, message.usage)
Chat Completions, for OpenAI-compatible clients, again with an explicit max_tokens and streaming on:
curl https://api.router.one/v1/chat/completions \
-H "Authorization: Bearer sk-your-api-key" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-haiku-5-5",
"max_tokens": 16000,
"stream": true,
"messages": [{"role": "user", "content": "Classify this support ticket as billing, bug or how-to, with one line of reasoning: The app signs me out every time I switch tabs."}]
}'
Rules that hold on both endpoints:
- Set
max_tokenson every request. It covers thinking plus text (Anthropic's migration guide, checked 2026-10-10), so leave room for both rather than relying on a default: a small cap tuned for Claude Haiku 4.5 can stop after athinkingblock and before any text. On Router One's Chat Completions,max_completion_tokensis accepted as the same cap (if both are sent,max_completion_tokenswins). - Stream long work.
stream: truereturns output as it is generated (streaming guide). - Keep a healthy wallet balance. Before a request runs, the gateway reserves an estimate against your balance; when the balance cannot cover it, the gateway lowers
max_tokensto fit, which can cut a long answer short. If it cannot cover even the request's initial reservation, the gateway returns HTTP 402 before the model runs. - Leave out
temperature,top_pandtop_k. Per Anthropic's migration guide (checked 2026-10-10), Claude Haiku 5.5 returns a 400 for atemperatureother than 1, atop_pother than 0.99, anytop_k, and a request that sends bothtemperatureandtop_p; omit all three and steer with the prompt instead. - End
messageswith a user turn. Per the same guide, an assistant prefill as the last turn returns a 400 on Claude Haiku 5.5, even with thinking off: replace a format prefill with structured outputs (for classification, a tool with enum fields) and a preamble prefill with a system-prompt instruction to answer directly. - Send files inline. Router One serves no Files API: pass images as base64 content blocks, and send PDFs on
/v1/messagesas base64documentblocks, not file references.
Which endpoint. /v1/messages is the recommended path for Claude Haiku 5.5: output_config.effort, thinking with its display option and anthropic-beta headers pass through unchanged, and it is the endpoint that carries the thinking blocks you pass back in a tool loop. /v1/chat/completions handles plain chat and tool calling, but it does not forward anthropic-beta headers or replay thinking blocks across turns, so a multi-turn tool loop runs without reasoning continuity; set effort on /v1/messages with output_config.effort. Per Anthropic's migration guide (checked 2026-10-10), Claude Haiku 5.5 accepts a forced tool_choice (any or a named tool), but the response then starts with the tool call and has no thinking block; to let the model think first, keep tool_choice on auto and say in the prompt when the tool applies. For JSON, use /v1/messages with output_config.format or a strict tool (Anthropic's structured outputs docs); response_format on Chat Completions is not a reliable way to get JSON from a Claude id. /v1/messages/count_tokens on Router One returns a local estimate, not Anthropic's count; billed numbers are in each response's usage.
- Read the trace. Dashboard → Logs shows each call with model, input and output tokens, cost, latency and status (per-request observability). A 400 for a bad parameter is not retried and does not move to another model: a failed Claude Haiku 5.5 request is never silently answered by Claude Haiku 4.5 or anything else.
What changes from Claude Haiku 4.5
Anthropic's what's new page lists the breaking and behavior changes for code already running on Claude Haiku 4.5, and the migration guide turns them into a checklist (both checked 2026-10-10). The model string is not the only change. In short, as Anthropic documents them for its API:
- Manual thinking budgets return a 400.
thinking: {"type": "enabled", "budget_tokens": N}is rejected on Claude Haiku 5.5. Leavethinkingunset or send{"type": "adaptive"}, and set depth withoutput_config.effort: where Claude Haiku 4.5 ran without thinking or with a small budget, choose a lower effort level. - Thinking is on by default, and its text is omitted. A response can begin with one or more
thinkingblocks even when the request does not mention thinking, so select content blocks bytype. The blocks come back with an emptythinkingfield unless you sendthinking: {"type": "adaptive", "display": "summarized"}; Claude Haiku 4.5 returned summarized thinking instead. You can still turn thinking off with{"type": "disabled"}athigheffort or below (atxhighormaxthat returns a 400), but Anthropic calls effort the better lever. - Sampling parameters and prefill return a 400. Omit
temperature,top_pandtop_k(the rules above give the values), and endmessageswith a user turn: an assistant prefill is rejected even with thinking off. - Forced tool use is accepted, without thinking.
tool_choiceanyor a named tool is accepted, but the response starts with the tool call and carries nothinkingblock. - The same text counts as more tokens. Claude Haiku 5.5 uses the newer tokenizer of Claude 4.7 and later models: approximately 30% more tokens than Claude Haiku 4.5 for the same text, depending on the content, and large images can count as more visual tokens. Recount prompts,
max_tokenslimits and cost estimates rather than reusing Haiku 4.5's numbers. - Earlier thinking stays in context. Thinking blocks from all earlier assistant turns stay in context and count as input tokens, where Claude Haiku 4.5 kept only the latest turn's, so multi-turn conversations carry more input. Keep conversations append-only and pass thinking blocks back unchanged: replaying a block after an edit to the
systemprompt, thetoolsor an earlier message can return a 400. - Computer use moves to
computer_toolset_20260801. On the Claude API and Google Cloud, a request that declarescomputer_20250124returns a 400. - Refusals are new. A decline returns
stop_reason: "refusal"with astop_detailscategory —cyber,frontier_llm,bioorgeneral_harms— and Claude Haiku 5.5 has no server-side fallback. On Router One, retry on another model from your own code: the server-sidefallbacksparameter is rejected with a 400 before the request reaches the model.
On Router One the key, base URL and endpoints stay the same, so the switch itself is the model string, but code written for Claude Haiku 4.5 may need the changes above first: write Claude Haiku 5.5 requests to Anthropic's contract, and check stop_reason and content block types rather than assuming the shape of a reply.
From mainland China
Requests reach api.router.one from mainland China without a VPN, on the same key and base URL, Claude Code included. Top up with a card or Alipay through one hosted checkout, or with USDT/USDC on six chains (Tron, BSC, Ethereum, Polygon, Base, Arbitrum). No US credit card required. For Claude Code, see Claude Code in China and the step-by-step setup guide.
FAQ
What is the model id for Claude Haiku 5.5 on Router One? The catalog id is anthropic/claude-haiku-5.5, and the gateway also accepts Anthropic's own id, claude-haiku-5-5, for the same model on /v1/messages and /v1/chat/completions. Use claude-haiku-5-5 in Claude Code: per its model configuration docs (checked 2026-10-10), that is the model's id, and the haiku alias sends it from 2.1.293. Either id bills per token to the wallet, since no plan tier lists the model as of the 2026-10-10 plan response.
When was Claude Haiku 5.5 released, and since when can I call it on Router One? Anthropic released it on October 7, 2026. Router One listed it as anthropic/claude-haiku-5.5 on 2026-10-10, and the gateway also accepts claude-haiku-5-5; no plan tier lists it as of the 2026-10-10 plan response, so its calls bill per token to the wallet.
How much does Claude Haiku 5.5 cost through Router One? Per token, at the input and output rates on the Claude Haiku 5.5 model page, which shows two lines: per Anthropic's pricing docs (checked 2026-10-10), a request whose prompt is over 100,000 tokens is priced on the higher one, and the Router One catalog splits its own posted rates at the same threshold. Thinking bills as output. The comparison pages render the live rates against Claude Haiku 4.5, Gemini 3.5 Flash Lite and Claude Sonnet 5.5. No plan tier lists it as of the 2026-10-10 plan response, so every call bills the wallet, plan or no plan.
Is Claude Haiku 5.5 included in Router One subscription plans? Not as of the 2026-10-10 plan response: no Pro, Max or Ultra tier lists it, so its calls bill per token to the wallet on every plan. Claude Haiku 4.5 stays in the Standard models tier of all three plans, and the pricing page shows the live plan model lists.
After updating Claude Code, why do haiku-alias requests bill the wallet? Per Claude Code's changelog and model configuration docs (Claude Code 2.1.296, checked 2026-10-10), 2.1.293 and later point the haiku alias at Claude Haiku 5.5 and send claude-haiku-5-5, which is in no plan tier as of the 2026-10-10 plan response, so /model haiku, subagents defined with model: haiku and the built-in claude-code-guide subagent bill per token to the wallet. The line ANTHROPIC_DEFAULT_HAIKU_MODEL=claude-haiku-4-5 is optional for everyone: it moves those requests, and background tasks, to Claude Haiku 4.5, which is in the Standard models tier of all three plans.
Does Claude Haiku 4.5 code work unchanged on Claude Haiku 5.5? Not always. On Router One the key, base URL and endpoints stay the same, but per Anthropic's migration guide (checked 2026-10-10) Haiku 4.5 code may need changes first: replace budget_tokens thinking with adaptive thinking and an effort level, remove temperature, top_p and top_k, end messages with a user turn instead of an assistant prefill, read content blocks by type, and recount tokens, since the same text counts as approximately 30% more. The switched calls bill the wallet, since no plan tier lists Claude Haiku 5.5 as of the 2026-10-10 plan response.
Claude Haiku 5.5 or Gemini 3.5 Flash Lite? In the 2026-10-10 catalog both take text and image input with a 1,048,576-token context window, but on Router One they differ in two ways. Endpoints: Claude Haiku 5.5 is served on /v1/messages and /v1/chat/completions, Gemini 3.5 Flash Lite on /v1/chat/completions only, so Claude Code can use only Claude Haiku 5.5. Plans: as of the 2026-10-10 plan response, Gemini 3.5 Flash Lite is in the Standard models tier of the Pro, Max and Ultra plans, while Claude Haiku 5.5 is in no plan tier and bills the wallet. Request rules follow each vendor's docs, and the comparison page renders the specs and live rates side by side.
Can Codex CLI or the Responses API call Claude Haiku 5.5? No. Router One serves Claude ids on /v1/messages and /v1/chat/completions, not on /v1/responses, and Codex CLI speaks only the Responses wire format, so keep Codex on ids served there, such as GPT, and call Claude Haiku 5.5 from a Messages or Chat Completions client.
Can I turn thinking off on Claude Haiku 5.5? On the API, yes, at high effort or below: per Anthropic's thinking and effort docs (checked 2026-10-10), Claude Haiku 5.5 accepts the disabled thinking type at high effort or lower and returns a 400 for it at xhigh or max; the between_tools type that Claude Sonnet 5.5 uses returns a 400 on Claude Haiku 5.5. Anthropic calls a lower effort level the better way to trade quality for speed and cost. In Claude Code, per its model configuration docs (Claude Code 2.1.296, checked 2026-10-10), thinking cannot be turned off on this model. Thinking tokens bill as output from your wallet, so effort is the cost lever.
Next steps
- Open the Claude Haiku 5.5 model page for the live rates and endpoints.
- Compare it with Claude Haiku 4.5, Gemini 3.5 Flash Lite or Claude Sonnet 5.5.
- For Claude Sonnet 5.5, read the Claude Sonnet 5.5 API guide.
- Set up Claude Code with Claude Code in China and the Claude Code setup guide.
- Check which models each plan covers on the pricing page.
- See this month's other catalog changes in the October 2026 new-model guide.