Availability note (updated 2026-09-25): This is an August 2026 catalog retrospective; several models it covers have since left or returned, so check the live catalog before integrating, and see the September 2026 new-model guide for what is listed now. The dated changes are logged under Availability history at the end.
Between late July and mid-August 2026, the model catalog behind Router One's OpenAI-compatible endpoint picked up an unusually dense batch of arrivals: Claude Opus 5, the GPT-5.6 family (Sol and Terra), Grok 4.6, DeepSeek V4 Flash, and Gemini 3.7 Flash were all listed in that August snapshot, alongside Claude Fable 5 at the top of the range. This guide covers what each one is for, where they sit relative to each other on price tier — the live per-token numbers stay in the model catalog, which is the source of truth this post deliberately does not reprint — and how to A/B any two of them from one API key by changing a single model parameter.
What landed, and when
Three of the arrivals carry a verifiable listing date. To be precise about what these dates mean: they are the dates each model went live in the Router One catalog, not the vendor's own launch dates.
| Model | Listed in the Router One catalog |
|---|---|
| Claude Opus 5 | 2026-07-24 |
| DeepSeek V4 Flash | 2026-08-04 |
| Gemini 3.7 Flash | 2026-08-17 |
GPT-5.6 Sol and Terra, Grok 4.6, and Claude Fable 5 are live in the catalog as well; we do not have clean listing dates for those, so this post claims none.
Claude Opus 5
Claude Opus 5 continues the Opus line that most coding agents standardized on. The headline spec is a 1M-token context window — the whole repository slice, the test output, and the conversation history fit in one request. It effectively succeeds Claude Opus 4.8 as the coding default; if you are deciding between the two Anthropic generations, that comparison page renders both spec sheets from the live catalog. Its own successor, Claude Opus 5.5, was listed on 2026-09-24 — see the Claude Opus 5.5 API guide. Against the strongest cross-vendor alternative, Claude Opus 5 vs GPT-5.6 Sol is the pairing to study, and Claude Opus 5 vs Gemini 3.1 Pro covers the long-context rival.
Claude Fable 5
Fable 5 sits above Opus 5 in Anthropic's range — it is the top Claude tier. As of the 2026-09-24 plan response no Router One subscription tier lists Fable 5, so its calls bill to the wallet. Same 1M context; its output ceiling is on the model page and is not reprinted here. Vision and tool calling are on. Whether the premium over Opus 5 pays for itself is workload-specific — Claude Fable 5 vs Claude Opus 5 puts the two spec sheets side by side, Claude Fable 5 vs GPT-5.6 Sol is the cross-vendor flagship match-up, and Claude Fable 5 vs Gemini 3.1 Pro covers the long-context rival.
The GPT-5.6 family: Sol and Terra
OpenAI's GPT-5.6 generation ships as a family sharing one spine — roughly 1M-token context, vision, tool calling — at different price points. Sol and Terra are the two variants this post benchmarks (Luna's status is noted below):
- GPT-5.6 Sol is the flagship. One spec worth knowing before you send it a whole codebase: prompts above roughly 272K tokens bill at a separate long-context tier, so the GPT-5.6 Sol model page shows two rate lines, not one. GPT-5.6 Sol vs GPT-5.5 shows the generation-over-generation change.
- GPT-5.6 Terra is the mid-tier default for everyday work. Sol vs Terra is the "do I actually need the flagship" check. For a current GPT comparison, GPT-5.6 Sol vs GPT-5.5 renders both live catalog entries; GPT-5.4 is no longer listed as of September 9.
GPT-5.6 Luna, the family's high-volume budget variant, left the Router One catalog on 2026-08-22 and was listed again on the September 1 snapshot, then was absent again on September 9; the model catalog is the source of truth for its current availability. The budget cross-check in this post stays Gemini 3.7 Flash, covered below.
The family structure is the point: you can prototype on Sol, then walk the same prompts down to Terra and measure exactly what quality the cheaper tier costs you — the A/B section below shows how.
Grok 4.6
Grok 4.6 is the current mainline xAI model in the catalog, succeeding Grok 4.5 — Grok 4.6 vs Grok 4.5 shows what changed. Specs: 500K context (half of what the Claude and GPT flagships carry, still far beyond most workloads), at mid-tier pricing that undercuts the flagships. That makes the interesting comparisons vertical: Grok 4.6 vs GPT-5.6 Sol if you are wondering whether the flagship premium is worth it, and Grok 4.6 vs GPT-5.6 Terra for the like-for-like mid-tier fight. For stable access to the Grok family from mainland China, see the Grok API in China page.
Kimi K3
Kimi K3 joined the catalog on 2026-07-25, but it is not in the Router One catalog at the time of this update (2026-09-01); check the model catalog for current availability. DeepSeek V4 Pro and Flash, GLM-5.2 and MiniMax M2.7 also left the catalog on 2026-09-05; deepseek-v4-flash was listed again on 2026-09-10 next to the new deepseek-v4.1-flash, while V4 Pro, GLM-5.2 and MiniMax M2.7 remained unlisted on 2026-09-11.
DeepSeek V4 Flash
DeepSeek V4 Flash was listed in Router One’s catalog on 2026-08-04, rather than being a new vendor release that month. The August snapshot described it as a high-volume budget option. On 2026-09-05 both V4 Flash and V4 Pro left the catalog, so this historical role is not a current recommendation. deepseek-v4-flash was listed again on 2026-09-10 next to the new deepseek-v4.1-flash (DeepSeek V4.1 Flash model page); V4 Pro remains unlisted. Read the live rate and endpoint list from the model page rather than from the August description, and see the live catalog for alternatives.
Gemini 3.7 Flash
Gemini 3.7 Flash was the newest Flash generation in the August snapshot, listed 2026-08-17. It keeps the shape of Gemini 3.6 Flash — 1M context window, vision and tool calling, identical call shape — so moving existing Flash traffic over is a model-string change. Its catalog position is the budget band: Gemini 3.7 Flash vs Gemini 3.6 Flash renders the generation-over-generation spec sheets from the live catalog, and the Gemini 3.7 Flash vs GPT-5.4 mini cross-vendor budget pairing was retired on 2026-09-11 when GPT-5.4 mini left the catalog; DeepSeek V4.1 Flash vs Gemini 3.8 Flash is the budget-band pairing rendered live today. Capability tags and the current rate are on the Gemini 3.7 Flash model page; for access from mainland China see the Gemini API in China page.
Where they sit on price — without reprinting a rate card
A per-model price table printed in a blog post is wrong within weeks, so this post does not carry one; the model catalog lists the live per-token rate, context window, and capability tags for every model here, and each model's detail page shows its current effective price. Pricing starts as low as 10% of official provider list prices (up to 90% off) on select models.
What a post can say durably is the relative banding as of mid-August 2026:
- Flagship band: Claude Fable 5, Claude Opus 5, GPT-5.6 Sol — priced for the tasks where a wrong answer costs more than the tokens.
- Mid band: GPT-5.6 Terra, Grok 4.6 — the everyday-driver tier.
- Budget band: Gemini 3.7 Flash, DeepSeek V4 Flash — priced for volume in August. V4 Flash and GPT-5.6 Luna were absent in the September 9 catalog;
deepseek-v4-flashwas listed again on 2026-09-10 next to the newdeepseek-v4.1-flash.
Picking by workload
Coding. Claude Opus 5 and GPT-5.6 Sol are the two serious defaults; their comparison page is the starting point, and Claude Fable 5 is the escalation for the hardest tasks. Wire the same choice into your editor or CLI once — Claude Code, Codex CLI, Cursor, Zed, and Cline all point at the same gateway endpoint.
Long-horizon agent tasks. Sessions that run for hours accumulate context, so the 1M-context models — Opus 5, Fable 5, the GPT-5.6 family, Gemini 3.7 Flash — are the natural picks, with Grok 4.6 as the cost-conscious candidate when 500K is enough. Two gateway-side notes matter more than the model choice: give each agent its own key with a maxSpend cap so a runaway loop stops at a number you chose (per-key cost tracking), and let automatic fallback absorb retryable upstream failures mid-run instead of your agent's error handler.
Cheap high-volume batch. The August snapshot used DeepSeek V4 Flash as one candidate; it left the catalog on 2026-09-05 and was listed again on 2026-09-10. For a current comparison, start with DeepSeek V4.1 Flash vs Gemini 3.8 Flash or DeepSeek V4.1 Flash vs Claude Haiku 4.5 and evaluate cost per completed task.
One key, one endpoint, real A/B
Currently listed chat models use the same OpenAI-compatible endpoint, so an A/B between any two of them is a one-line change — no second account, no second SDK, no second billing relationship:
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.router.one/v1",
api_key=os.environ["ROUTER_ONE_API_KEY"],
)
PROMPT = "Refactor this function and explain the trade-offs: ..."
for model in ["anthropic/claude-opus-5", "openai/gpt-5.6-sol"]:
r = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": PROMPT}],
)
print(model, r.choices[0].message.content[:200])
Run the same prompt set through both candidates, then read the results out of the per-request trace log: every call records model, tokens, cost, latency, and status, so "which model is actually cheaper per solved task" becomes a dashboard filter instead of a spreadsheet project. That loop — swap the string, rerun, read the ledger — is the practical payoff of a unified LLM API gateway when the catalog turns over this fast.
FAQ
Do I need separate vendor accounts to try these models? No. Currently listed chat models are callable through one Router One API key on one OpenAI-compatible endpoint — switching between them is a change to the model parameter, not a new account. Requests reach the catalog without a VPN from mainland China as well.
Is DeepSeek V4 Flash a new August 2026 model? No. The DeepSeek V4 line launched in April 2026; V4 Flash is its budget high-volume tier. What is new is availability: it was listed in the Router One catalog on 2026-08-04. The dates in this post are catalog listing dates, not vendor launch dates.
What do these models cost through Router One? This post deliberately prints no per-token rates because they go stale. The models page carries the live per-model rate, context window, and capability tags, and pricing starts as low as 10% of official provider list prices (up to 90% off) on select models.
Which of the new models should I default to for coding? Claude Opus 5 and GPT-5.6 Sol are the two serious defaults, and the right answer depends on your codebase and workflow — run both against a fixed prompt set through one key and compare cost per solved task in the request traces. Claude Fable 5 is the escalation when tasks need the top Claude tier; check its context window and rate on the model page.
Availability history
- 2026-09-09: DeepSeek V4 Pro/Flash, GLM-5.2, MiniMax M2.7, GPT-5.4 and GPT-5.6 Luna were not listed in the September 9 snapshot. Gemini 3.1 Pro Preview had returned; the older Gemini 3 Pro Preview remained absent. Newer models — GPT-6 Astra, Gemini 3.8 Flash, Grok 4.3 and, between the 2026-09-01 and 2026-09-09 snapshots, Claude Fable 5.1 — were listed after August, so the tier labels in this post describe August, not today.
- 2026-09-11: GPT-5.4 mini left the catalog by 2026-09-11, so no mini GPT id is listed and the Gemini 3.7 Flash vs GPT-5.4 mini comparison page was retired; Claude Fable 5.1 had also left. DeepSeek V4.1 Flash (
deepseek-v4.1-flash) was listed on 2026-09-10, withdeepseek-v4-flashlisted again alongside it — both are served on /v1/chat/completions, on /v1/messages and natively on /v1/responses; DeepSeek V4 Pro remained unlisted. - 2026-09-25: The catalog lists 63 ids. Of the models this post covers, Claude Opus 5, Claude Fable 5, GPT-5.6 Sol and GPT-5.6 Terra, Grok 4.6, DeepSeek V4 Flash and Gemini 3.7 Flash are listed; Kimi K3, GLM-5.2, MiniMax M2.7 and the default-channel GPT-5.6 Luna are not. The historical lineup and price-tier descriptions above are not current recommendations.