Between the 2026-09-01 and 2026-09-05 catalog snapshots, Router One listed eight new model ids — Claude Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash, Gemini 3.1 Flash Lite and its preview twin, Gemini 2.5 Pro, Grok 4.3 and Grok Build 0.1 — while twelve ids left, and every new one is callable from the same API key on the endpoint family its detail page lists.
This is the September sibling of the August 2026 new-model guide, written the same way: what the live catalog exposes for each id (context window, input modalities, capability flags, price tiers), which comparison pages render the spec sheets side by side, and how to send a first request. The snapshot was re-read on 2026-09-07 and its id set was unchanged — 38 ids, 33 text and 5 image — but ids come and go, so the catalog is the only source of truth and this post is a dated record.
What was listed, and what left
The dates in this post are catalog snapshot dates, not vendor launch dates. We do not have a clean per-model listing day for any of the eight, so this post claims none.
| Listed between 2026-09-01 and 2026-09-05 | Catalog id | Context | Input |
|---|---|---|---|
| Claude Fable 5.1 | anthropic/claude-fable-5.1 | 1,048,576 | text, image |
| GPT-6 Astra | openai/gpt-6-astra | 1,050,000 | text, image |
| Gemini 3.8 Flash | google/gemini-3.8-flash | 1,048,576 | text, image |
| Gemini 3.1 Flash Lite | google/gemini-3.1-flash-lite | 1,048,576 | text, image |
| Gemini 3.1 Flash Lite Preview | google/gemini-3.1-flash-lite-preview | 1,048,576 | text, image |
| Gemini 2.5 Pro | google/gemini-2.5-pro | 1,048,576 | text, image |
| Grok 4.3 | grok-4.3 | 500,000 | text |
| Grok Build 0.1 | grok-build-0.1 | 500,000 | text |
No longer listed as of the same snapshot: DeepSeek V4 Pro and DeepSeek V4 Flash, GLM 5.1, GLM 5.2 and GLM 5.2 Fast, MiniMax M2.7, Doubao Seed 2.0 Lite, Doubao Seedream 5.0, Gemini 3 Pro Preview, Gemini 3.1 Pro Preview, o3 and o4-mini. A request that names one of those ids gets an error instead of a completion — the error codes page covers the unknown-model case. This post does not speculate about why an id leaves; it records that the catalog no longer lists it. GPT-5.6 Luna, whose exit and return the August post already covers, remains listed.
Claude Fable 5.1
The catalog lists Claude Fable 5.1 with the same envelope as Claude Fable 5: a 1,048,576-token context window and text plus image input. The catalog entry ships no capability flags of its own, so this post asserts none; the model page is where the flag list lives. It is a Claude-family id, so it answers on POST /v1/messages — the path Claude Code uses — and on POST /v1/chat/completions; the API compatibility facts state the family rule, and the model page lists the exact endpoints.
Two comparison pages render the live spec sheets: Claude Fable 5.1 vs Claude Opus 5 for the in-family question of what the higher rate buys against Opus 5, and Claude Fable 5.1 vs GPT-6 Astra for the cross-vendor pairing at the top of each vendor's rate card; Claude Fable 5 vs Claude Opus 5 covers the previous Fable id. The claude-*-thinking ids in the catalog are separate entries; reasoning output on them is billed at the model's posted output rate with no separate reasoning line (pricing facts).
GPT-6 Astra
GPT-6 Astra is the first GPT-6 id in the catalog. Its entry lists a 1,050,000-token context window, text and image input, and the chat, streaming, tool-calling and vision capability flags. As a GPT-family id it is served natively on POST /v1/responses — the wire format Codex CLI speaks, covered on the Codex CLI and Responses API page — and on POST /v1/chat/completions.
One billing detail belongs next to the spec sheet: Astra carries a whole-request price tier above 272,000 input tokens, the same threshold GPT-5.6 Sol carries, so its model page shows two rate lines rather than one. The tier rule is explained below. GPT-6 Astra vs GPT-5.6 Sol is the generation-over-generation comparison, and GPT-6 Astra vs Claude Opus 5 the cross-vendor one.
Gemini 3.8 Flash, Gemini 3.1 Flash Lite and Gemini 2.5 Pro
Three Gemini names arrived together, all described in the catalog as chat, tool-calling and vision models with a 1,048,576-token context window:
- Gemini 3.8 Flash is the newest Flash generation in the catalog and shares the Gemini 3.7 Flash context window and input modalities (the catalog publishes capability flags for 3.7 Flash but none yet for 3.8 Flash — the model page is authoritative), so once you have checked the flags, moving existing Flash traffic over is a
modelstring change. - Gemini 3.1 Flash Lite ships as two ids,
google/gemini-3.1-flash-liteandgoogle/gemini-3.1-flash-lite-preview, with identical published specs; the un-suffixed id is the one the Gemini 3.1 Flash Lite vs GPT-5.4 mini comparison uses. - Gemini 2.5 Pro is an older generation that is now listed. The dates in this post are catalog dates, so its presence says nothing about a vendor release; whether an older generation suits a workload is a question for your own prompt set.
Each Gemini id answers on POST /v1/chat/completions; the model page is authoritative for the endpoint list and the live rate. Access from mainland China for the whole family is on the Gemini API in China page.
Grok 4.3 and Grok Build 0.1
The catalog describes both as xAI text models — text input only, a 500,000-token context window — and lists no capability flags for either, so this post describes nothing beyond that. Grok 4.3 sits alongside Grok 4.5 and Grok 4.6; Grok 4.6 vs Grok 4.5 covers the two mainline ids that already had a comparison. Grok Build 0.1 is a new name, and the catalog entry is all we can vouch for: what the "Build" line is meant for is not something the catalog states, so the Grok Build 0.1 vs GPT-5.3 Codex Spark page pairs it with another id on spec and rate, not on a claim about purpose.
Both Grok ids carry a whole-request tier above 200,000 input tokens, the threshold Grok 4.6 also has. Access from mainland China for the family is on the Grok API in China page.
How the whole-request price tiers work
Five text ids in this snapshot list a second rate line — GPT-6 Astra and GPT-5.6 Sol above 272,000 input tokens, Grok 4.3, Grok Build 0.1 and Grok 4.6 above 200,000 — and the rule is the same for all of them: the total input tokens of one request select the tier, exactly at the threshold the lower tier still applies, and once a request is strictly above it the selected tier's rates apply to the whole request, output included, rather than only to the tokens past the line; monthly volume plays no part. The pricing methodology spells this out, the cost calculator applies it per request, and the prompt caching guide shows how to read cache write and cache read counts in usage so input is not counted twice.
Where the prices are
Not here. A per-token figure printed in a post is wrong within weeks, so the model catalog and each model page carry the live rate and the tier boundaries; the pricing page states the positioning — pricing starts as low as 10% of official provider list prices (up to 90% off) on select models — and the pricing facts state the billing rules. Rates differ by id and change over time, so read them from the page, not from memory, and never from a comparison that ranks models by a number that has since moved.
Which endpoint serves which id
Endpoint families are not interchangeable. POST /v1/chat/completions serves every chat model in the catalog; POST /v1/messages serves the currently listed Claude-family ids; POST /v1/responses is served natively for the currently listed GPT-family ids. An id sent to an endpoint that does not serve it is rejected with HTTP 400 invalid_request_error before any model is called, and the message names the path to use. For this batch that means: Claude Fable 5.1 on Messages or Chat Completions; GPT-6 Astra on Responses or Chat Completions; the Gemini and Grok ids on Chat Completions. The OpenAI-compatible API page covers the request shape, and each model page lists exactly the endpoints it serves.
How to try one in five minutes
- Create a key. Dashboard → API Keys → New key. Keys look like
sk-rk-.... If this is an experiment, give the key amaxSpendcap so a runaway loop stops at a number you chose (per-key cost tracking). - Set the base URL.
https://api.router.one/v1for OpenAI-compatible clients and SDKs;https://api.router.one(host only) for the Anthropic SDK and for Claude Code'sANTHROPIC_BASE_URL, both of which append/v1/messagesthemselves. - Send one request per family.
Chat Completions with a GPT id:
curl https://api.router.one/v1/chat/completions \
-H "Authorization: Bearer sk-your-api-key" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-6-astra",
"messages": [{"role": "user", "content": "Explain what a whole-request price tier is, in two sentences."}]
}'
Messages with a Claude id:
curl https://api.router.one/v1/messages \
-H "Authorization: Bearer sk-your-api-key" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "anthropic/claude-fable-5.1",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Explain what a whole-request price tier is, in two sentences."}]
}'
- Read the trace. Dashboard → Logs shows both calls with model, input and output tokens, cost, latency and status, so "which of the two is cheaper per solved task" becomes a filter rather than a spreadsheet (per-request observability). Retryable upstream failures are absorbed by automatic fallback across the candidate routes for the same model, and the streaming guide covers
stream: true.
Top up with a card or Alipay through one hosted checkout, or with USDT/USDC on six chains (Tron, BSC, Ethereum, Polygon, Base, Arbitrum). No US credit card required. Requests reach the catalog from mainland China without a VPN.
What to A/B first
This is not a ranking — the comparison pages render live spec sheets and rates, and your own prompt set decides. For the two highest-rate pairings, Claude Fable 5.1 vs Claude Opus 5 and GPT-6 Astra vs Claude Opus 5; to see what the GPT-6 generation changed against GPT-5.6, GPT-6 Astra vs GPT-5.6 Sol; for high-volume work, Gemini 3.1 Flash Lite vs GPT-5.4 mini; for the new xAI name against a known id, Grok Build 0.1 vs GPT-5.3 Codex Spark. Wire the winner into Claude Code (setup), Codex CLI (setup), Cursor, Zed or Continue once — they all point at the same unified LLM API gateway.
FAQ
Are DeepSeek, GLM or MiniMax models still callable through Router One? Not as of this snapshot. DeepSeek V4 Pro and Flash, GLM 5.1, 5.2 and 5.2 Fast, and MiniMax M2.7 are no longer listed as of 2026-09-05, and the re-read on 2026-09-07 showed the same id set. The live catalog at /models is the only source of truth: ids come and go, and a post cannot promise either direction.
Which endpoint do I use for GPT-6 Astra and Claude Fable 5.1? GPT-6 Astra is a GPT-family id, so it is served natively on /v1/responses and also on /v1/chat/completions. Claude Fable 5.1 is a Claude-family id, so it answers on /v1/messages and on /v1/chat/completions. Sending an id to an endpoint that does not serve it returns HTTP 400 invalid_request_error before any model is called, with the correct path named in the message.
Does the long-context tier apply only to the tokens above the threshold? No. The total input tokens of one request select the tier; exactly at the threshold the lower tier still applies, and strictly above it the selected tier's rates apply to the whole request, including output. Monthly usage does not select the tier. GPT-6 Astra and GPT-5.6 Sol carry the line at 272,000 input tokens; Grok 4.3, Grok Build 0.1 and Grok 4.6 at 200,000.
What do the new models cost through Router One? This post deliberately prints no per-token rates because they go stale. Each model page carries the live rate and tier boundaries, and pricing starts as low as 10% of official provider list prices (up to 90% off) on select models.
Is Gemini 2.5 Pro a new September 2026 model? No. Its appearance in this snapshot is a catalog listing, not a vendor release; every date in this post is a catalog snapshot date. The model page carries its live rate and endpoints.
Next steps
- Read the August 2026 guide for the previous batch — Claude Opus 5, the GPT-5.6 family and Grok 4.6.
- Pick an id on the model catalog and check its endpoints, tier boundaries and live rate; the discount qualifier is on the pricing page.
- Point Codex CLI at a GPT id with the Codex and Responses API guide, or Claude Code at a Claude id with the Claude Code setup.