> Markdown mirror of https://router.one/blog/new-llm-models-august-2026 for AI assistants and crawlers. Router One is an OpenAI-compatible LLM API gateway.
> Published: 2026-08-16 · Author: Router One Team

# New LLM Models, August 2026: Claude Opus 5, GPT-5.6, Grok 4.6

_What's new in the catalog as of August 2026 — Claude Opus 5, the GPT-5.6 family (Sol, Luna, Terra), Grok 4.6, Kimi K3, and DeepSeek V4 Flash: what each is for, and how to A/B them through one OpenAI-compatible key._

Between late July and mid-August 2026, the model catalog behind Router One's OpenAI-compatible endpoint picked up an unusually dense batch of arrivals: Claude Opus 5, the three-variant GPT-5.6 family (Sol, Luna, Terra), Grok 4.6, Kimi K3, and DeepSeek V4 Flash are all callable now, alongside Claude Fable 5 at the top of the range. This guide covers what each one is for, where they sit relative to each other on price *tier* — the live per-token numbers stay in the [model catalog](https://router.one/models), which is the source of truth this post deliberately does not reprint — and how to A/B any two of them from one API key by changing a single `model` parameter.

## What landed, and when

Three of the arrivals carry a verifiable listing date. To be precise about what these dates mean: they are the dates each model went live in the Router One catalog, not the vendor's own launch dates.

| Model | Listed in the Router One catalog |
| --- | --- |
| [Claude Opus 5](https://router.one/models/claude-opus-5) | 2026-07-24 |
| [Kimi K3](https://router.one/models/kimi-k3) | 2026-07-25 |
| [DeepSeek V4 Flash](https://router.one/models/deepseek-v4-flash) | 2026-08-04 |

The GPT-5.6 family, Grok 4.6, and Claude Fable 5 are live in the catalog as well; we do not have clean listing dates for those, so this post claims none.

## Claude Opus 5

Claude Opus 5 continues the Opus line that most coding agents standardized on. The headline spec is a 1M-token context window — the whole repository slice, the test output, and the conversation history fit in one request — with the same 8K max output as [Claude Opus 4.8](https://router.one/models/compare/claude-opus-5-vs-claude-opus-4-8), the model it effectively succeeds as the coding default. If you are deciding between the two Anthropic generations, that comparison page renders both spec sheets from the live catalog. Against the strongest cross-vendor alternative, [Claude Opus 5 vs GPT-5.6 Sol](https://router.one/models/compare/claude-opus-5-vs-gpt-5-6-sol) is the pairing to study, and [Claude Opus 5 vs Gemini 3.1 Pro](https://router.one/models/compare/claude-opus-5-vs-gemini-3-1-pro) covers the long-context rival.

## Claude Fable 5

Fable 5 sits above Opus 5 in Anthropic's range and in Router One's subscription quota tiers — it is the flagship-tier model. Same 1M context, but a much larger 32K max output, which matters for agents that emit big diffs, long structured documents, or full file rewrites in one turn: an 8K output ceiling forces continuation loops that a 32K ceiling simply absorbs. Vision and tool calling are on. Whether the premium over Opus 5 pays for itself is workload-specific — [Claude Fable 5 vs Claude Opus 5](https://router.one/models/compare/claude-fable-5-vs-claude-opus-5) puts the two spec sheets side by side, and [Claude Fable 5 vs GPT-5.6 Sol](https://router.one/models/compare/claude-fable-5-vs-gpt-5-6-sol) is the cross-vendor flagship match-up.

## The GPT-5.6 family: Sol, Luna, Terra

OpenAI's GPT-5.6 generation ships as three models sharing one spine — roughly 1M-token context, 128K max output, vision, tool calling — at three different price points:

- **GPT-5.6 Sol** is the flagship. One spec worth knowing before you send it a whole codebase: prompts above roughly 272K tokens bill at a separate long-context tier, so the [model page](https://router.one/models/gpt-5-6-sol) shows two rate lines, not one. [GPT-5.6 Sol vs GPT-5.5](https://router.one/models/compare/gpt-5-6-sol-vs-gpt-5-5) shows the generation-over-generation change.
- **GPT-5.6 Terra** is the mid-tier default for everyday work. [Sol vs Terra](https://router.one/models/compare/gpt-5-6-sol-vs-gpt-5-6-terra) is the "do I actually need the flagship" check.
- **GPT-5.6 Luna** is the high-volume budget variant, competing at the tier where [Gemini 3.6 Flash](https://router.one/models/compare/gpt-5-6-luna-vs-gemini-3-6-flash) plays.

The family structure is the point: you can prototype on Sol, then walk the same prompts down to Terra or Luna and measure exactly what quality the cheaper tier costs you — the A/B section below shows how.

## Grok 4.6

Grok 4.6 is the current mainline xAI model in the catalog, succeeding Grok 4.5 — [Grok 4.6 vs Grok 4.5](https://router.one/models/compare/grok-4-6-vs-grok-4-5) shows what changed. Specs: 500K context (half of what the Claude and GPT flagships carry, still far beyond most workloads) and 128K max output, at mid-tier pricing that undercuts the flagships. That makes the interesting comparisons vertical: [Grok 4.6 vs GPT-5.6 Sol](https://router.one/models/compare/grok-4-6-vs-gpt-5-6-sol) if you are wondering whether the flagship premium is worth it, and [Grok 4.6 vs GPT-5.6 Terra](https://router.one/models/compare/grok-4-6-vs-gpt-5-6-terra) for the like-for-like mid-tier fight.

## Kimi K3

Kimi K3 is Moonshot AI's flagship reasoning model, listed 2026-07-25 with a 1M-token context window. It arrives with real momentum in the agentic-coding community, and its catalog position is distinctive: flagship-class context at mid-tier pricing. The capability tags and current rates are on the [model page](https://router.one/models/kimi-k3); the comparison worth reading first is [Kimi K3 vs DeepSeek V4 Flash](https://router.one/models/compare/kimi-k3-vs-deepseek-v4-flash), because the two Chinese-lab models bracket the cost spectrum — K3 as the reasoning pick, V4 Flash as the volume pick.

## DeepSeek V4 Flash

One framing correction up front: DeepSeek V4 Flash is **not an August release**. The V4 line — Pro and Flash — launched in April 2026 and reset the price floor for frontier-adjacent quality; our [China model comparison](https://router.one/blog/qwen3-doubao-vs-claude-gpt-china) covers that shift in depth. What changed on 2026-08-04 is that V4 Flash became callable through Router One. Its role is the incumbent of the cheap high-volume tier: 1M context, tool calling, streaming, at the lowest price band of every model named in this post. Classification, summarization, extraction, chatbot turns at scale — this is the model you try first and only escalate from when the error rate says you must. [Doubao Seed 2.0 Lite vs DeepSeek V4 Flash](https://router.one/models/compare/doubao-seed-2-0-lite-260428-vs-deepseek-v4-flash) covers the budget-tier alternative.

## Where they sit on price — without reprinting a rate card

A per-model price table printed in a blog post is wrong within weeks, so this post does not carry one; the [model catalog](https://router.one/models) lists the live per-token rate, context window, and capability tags for every model here, and each model's detail page shows its current effective price. Pricing starts as low as 10% of official provider list prices (up to 90% off) on select models.

What a post *can* say durably is the relative banding as of mid-August 2026:

- **Flagship band:** Claude Fable 5, Claude Opus 5, GPT-5.6 Sol — priced for the tasks where a wrong answer costs more than the tokens.
- **Mid band:** GPT-5.6 Terra, Grok 4.6, Kimi K3 — the everyday-driver tier.
- **Budget band:** GPT-5.6 Luna, DeepSeek V4 Flash — priced for volume.

## Picking by workload

**Coding.** Claude Opus 5 and GPT-5.6 Sol are the two serious defaults; [their comparison page](https://router.one/models/compare/claude-opus-5-vs-gpt-5-6-sol) is the starting point, and Claude Fable 5 is the escalation for the hardest tasks, where the 32K output ceiling also pays off. Wire the same choice into your editor or CLI once — Claude Code, Codex CLI, [Cursor](https://router.one/integrations/cursor), [Zed](https://router.one/integrations/zed), and [Continue](https://router.one/integrations/continue) all point at the same gateway endpoint.

**Long-horizon agent tasks.** Sessions that run for hours accumulate context, so the 1M-context models — Opus 5, Fable 5, the GPT-5.6 family, Kimi K3 — are the natural picks, with Grok 4.6 as the cost-conscious candidate when 500K is enough. Two gateway-side notes matter more than the model choice: give each agent its own key with a `maxSpend` cap so a runaway loop stops at a number you chose ([per-key cost tracking](https://router.one/llm-cost-tracking)), and let [automatic fallback](https://router.one/llm-fallback) absorb retryable upstream failures mid-run instead of your agent's error handler.

**Cheap high-volume batch.** DeepSeek V4 Flash first, GPT-5.6 Luna as the cross-check — and [Kimi K3 vs DeepSeek V4 Flash](https://router.one/models/compare/kimi-k3-vs-deepseek-v4-flash) when a batch task turns out to need more reasoning than the budget tier delivers.

## One key, one endpoint, real A/B

Every model in this post sits behind the same [OpenAI-compatible endpoint](https://router.one/openai-compatible-api), so an A/B between any two of them is a one-line change — no second account, no second SDK, no second billing relationship:

```python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.router.one/v1",
    api_key=os.environ["ROUTER_ONE_API_KEY"],
)

PROMPT = "Refactor this function and explain the trade-offs: ..."

for model in ["anthropic/claude-opus-5", "openai/gpt-5.6-sol"]:
    r = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": PROMPT}],
    )
    print(model, r.choices[0].message.content[:200])
```

Run the same prompt set through both candidates, then read the results out of the [per-request trace log](https://router.one/llm-observability): every call records model, tokens, cost, latency, and status, so "which model is actually cheaper per solved task" becomes a dashboard filter instead of a spreadsheet project. That loop — swap the string, rerun, read the ledger — is the practical payoff of a [unified LLM API gateway](https://router.one/llm-api-gateway) when the catalog turns over this fast.

## FAQ

**Do I need separate vendor accounts to try these models?**
No. All of the models in this post are callable through one Router One API key on one OpenAI-compatible endpoint — switching between them is a change to the model parameter, not a new account. Requests reach the catalog without a VPN from mainland China as well.

**Is DeepSeek V4 Flash a new August 2026 model?**
No. The DeepSeek V4 line launched in April 2026; V4 Flash is its budget high-volume tier. What is new is availability: it was listed in the Router One catalog on 2026-08-04. The dates in this post are catalog listing dates, not vendor launch dates.

**What do these models cost through Router One?**
This post deliberately prints no per-token rates because they go stale. The models page carries the live per-model rate, context window, and capability tags, and pricing starts as low as 10% of official provider list prices (up to 90% off) on select models.

**Which of the new models should I default to for coding?**
Claude Opus 5 and GPT-5.6 Sol are the two serious defaults, and the right answer depends on your codebase and workflow — run both against a fixed prompt set through one key and compare cost per solved task in the request traces. Claude Fable 5 is the escalation when tasks need flagship-tier reasoning or its larger 32K output ceiling.

## See also

- Canonical page: https://router.one/blog/new-llm-models-august-2026
- LLM API Gateway and Routing: https://router.one/llm-api-gateway
- All blog posts: https://router.one/blog
- Models and per-model token rates: https://router.one/models (markdown: https://router.one/models.md)
- Pricing: https://router.one/pricing
- API docs (markdown): https://router.one/docs.md
- Company facts: https://router.one/facts/company.md
