> Markdown mirror of https://router.one/blog/gpt-6-astra-api-guide for AI assistants and crawlers. Router One is a unified, OpenAI-compatible LLM API gateway.
> Published: 2026-10-03 · Author: Router One Team

# GPT-6 Astra API Guide: Model ID, Codex, Plan Quota

_GPT-6 Astra on Router One: model ID, tool calls per OpenAI (Responses API only), Codex setup, and the Flagship quota on Max and Ultra (wallet billing on Pro)._

GPT-6 Astra, which OpenAI released on September 3, 2026, is in the Router One catalog as `openai/gpt-6-astra` (listed 2026-09-05). It is served natively on `POST /v1/responses` — the wire format Codex CLI speaks — and on `POST /v1/chat/completions`. Point any OpenAI-compatible SDK or client at `https://api.router.one/v1`, send `"model": "openai/gpt-6-astra"` with a Router One key, and the request is routed, metered and logged like every other model. In the 2026-10-03 plan response, GPT-6 Astra is the only model in the Flagship models tier, which only Router One's Max and Ultra plans carry; on Pro its calls bill per token to your wallet. The [GPT-6 Astra model page](https://router.one/models/gpt-6-astra) carries the live rates.

This guide covers what the catalog lists for the id, how OpenAI positions it against GPT-6.1 Sol and GPT-6 Sol, which plans cover it and how its requests are counted, what OpenAI's request rules mean for your code, how to run it in Codex CLI, and the first request on each endpoint. Catalog and plan observations are dated; the live catalog is the source of truth for a new request.

## GPT-6 Astra at a glance

| Field | What the catalog lists (2026-10-03) |
| --- | --- |
| Catalog id | `openai/gpt-6-astra`, listed 2026-09-05 |
| Context window | 1,050,000 tokens |
| Input / output | text and image in, text out |
| Capability flags | chat, streaming, tool calling, vision |
| Price lines | two: a standard line, and a whole-request line for requests strictly above 272,000 input tokens |
| Endpoints | natively `POST /v1/responses`, and `POST /v1/chat/completions` |
| Not served on | `POST /v1/messages` (Claude-family and DeepSeek ids only) |
| Channel id | `azure/gpt-6-astra` — a separately priced product of the same base model, in no plan tier |
| Plans | the Flagship models tier of Max and Ultra (2026-10-03 plan response); wallet billing on Pro |

OpenAI's [GPT-6 Astra model page](https://developers.openai.com/api/docs/models/gpt-6-astra) (checked 2026-10-03) adds what the catalog does not carry. It calls GPT-6 Astra "our most capable model for the most demanding work", lists a maximum of 922,000 input tokens and 128,000 output tokens inside the 1,050,000-token window and an April 30, 2026 knowledge cutoff, and documents `reasoning.effort` values `low`, `medium`, `high`, `xhigh` and `max`. The page marks no default effort, and OpenAI's reasoning guide (checked 2026-10-03) names the default for other GPT-6 models but not for GPT-6 Astra, so every sample below sets the effort explicitly. Per OpenAI's GPT-6 guide (checked 2026-10-03), GPT-6 Astra takes tool calls only on the Responses API, and Chat Completions serves it for requests without tools. Those are OpenAI's statements; what Router One vouches for is the catalog entry and how the gateway routes and bills it.

## GPT-6 Astra, GPT-6.1 Sol or GPT-6 Sol?

OpenAI's [GPT-6 guide](https://developers.openai.com/api/docs/guides/latest-model) (checked 2026-10-03) places GPT-6 Astra at the "highest intelligence", for "the most demanding reasoning, coding, and professional work", and GPT-6.1 Sol at "balanced speed, cost, and intelligence". OpenAI's model catalog also calls GPT-6 Astra its flagship model for complex reasoning and coding; that is OpenAI's word for its own line-up, separate from Router One's Flagship models plan tier. OpenAI's [launch post](https://openai.com/index/gpt-6-astra/) (September 3, 2026) reports that in latency simulations on OSWorld 2.0, GPT-6 Astra reached higher computer-use performance than GPT-5.6 Sol in about 47% less time per task, and its GPT-6.1 Sol launch post (September 29, 2026) reports that on DeepSWE v1.1, GPT-6.1 Sol matches GPT-6 Astra. Those are OpenAI's results, not Router One measurements: run your own prompts before you move a workload.

What Router One adds to the choice:

- **Plan quota or wallet.** In the 2026-10-03 plan response, GPT-6 Astra counts against the Flagship models allowance on Max and Ultra, while GPT-6.1 Sol and GPT-6 Sol are in no plan tier and bill the wallet. On Pro, all three bill the wallet. GPT-5.6 Sol and GPT-5.5 stay in the Premium models tier of all three plans.
- **Posted rates.** Router One posts its own rates for each id, so re-baseline cost when you switch.
- **Same request path.** All three are GPT-family ids served natively on `/v1/responses`, the Codex CLI path.

Four comparison pages render the spec sheets and live rates side by side:

- [GPT-6.1 Sol vs GPT-6 Astra](https://router.one/models/compare/gpt-6-1-sol-vs-gpt-6-astra) — Astra on the Flagship plan quota against GPT-6.1 Sol, which OpenAI's model page describes as near-Astra, billed to the wallet.
- [GPT-6 Sol vs GPT-6 Astra](https://router.one/models/compare/gpt-6-sol-vs-gpt-6-astra) — the earlier Sol model against Astra.
- [GPT-6 Astra vs GPT-5.6 Sol](https://router.one/models/compare/gpt-6-astra-vs-gpt-5-6-sol) — Astra against GPT-5.6 Sol, the model the one-click Codex install sets.
- [Claude Opus 5.5 vs GPT-6 Astra](https://router.one/models/compare/claude-opus-5-5-vs-gpt-6-astra) — the cross-vendor question; the two differ in endpoints and plan status.

## Which plans cover GPT-6 Astra?

In the 2026-10-03 plan response, GPT-6 Astra is the only model in the Flagship models tier, and only Router One's Max and Ultra plans carry that tier: Max with 150 and Ultra with 450 Flagship requests per 30-day cycle. On Max or Ultra, a GPT-6 Astra request draws 1 request from the Flagship allowance — 2 when its total input is strictly above 272,000 tokens and 4 above 512,000; the input count includes cached input once, and output does not count. Once the cycle's allowance is used, further calls bill the wallet at the posted rates. On Pro, and for the separately priced `azure/gpt-6-astra` channel id on every plan, calls bill the wallet. Plan model lists and allowances change, and the [pricing page](https://router.one/pricing) shows the live ones.

## What changes in your requests

The key, base URL and endpoints stay the same when you move to GPT-6 Astra from another GPT id. Three request rules come from OpenAI's [GPT-6 guide](https://developers.openai.com/api/docs/guides/latest-model) (migration quickstart, checked 2026-10-03):

- **Reasoning effort.** GPT-6 Astra accepts `low`, `medium`, `high`, `xhigh` and `max`, not `none`; OpenAI's reasoning guide (checked 2026-10-03) says setting `none` returns HTTP 400, and the GPT-6 guide suggests `low` in its place. Where you used `minimal`, OpenAI suggests starting at `low` and comparing results. Set it with `reasoning.effort` on Responses or `reasoning_effort` on Chat Completions.
- **Tool calls.** GPT-6 Astra takes tool calls only on the Responses API. Code that sends tools on Chat Completions moves to `/v1/responses`, where Router One serves GPT models natively.
- **Sampling parameters.** Remove `temperature`, `top_p` and `top_logprobs`, plus `logprobs` on Chat Completions; on Responses, take `message.output_text.logprobs` out of `include`. OpenAI's API changelog entry for GPT-6 Astra (September 3, 2026) says the same: no custom `temperature` or `top_p` values and no log probabilities.

## How billing works

The model page shows two rate lines, and the total input tokens of one request select which one applies. At exactly 272,000 input tokens the standard line still applies. Strictly above it, the long-context line applies to the whole request — output included — not only to the tokens past the threshold. Monthly volume plays no part. OpenAI's model page documents the same 272K threshold and the same full-request rule for its own list prices, and Router One's long-context line mirrors it.

Reasoning output is billed at the model's posted output rate, with no separate reasoning line ([pricing facts](https://router.one/facts/pricing.md)), so the reasoning effort is a cost lever as much as a quality one. The model page also lists a cached-input line; the [prompt caching guide](https://router.one/llm-prompt-caching) shows how to read cache counts in `usage` so input is not counted twice, the [pricing methodology](https://router.one/pricing-methodology) spells out the long-context rule, and the [cost calculator](https://router.one/llm-cost-calculator) applies it per request. On Max and Ultra, calls inside the Flagship allowance draw plan requests instead of wallet charges, as counted above.

This guide prints no per-token figures because they go stale. The [GPT-6 Astra model page](https://router.one/models/gpt-6-astra) shows both lines live, and the comparison pages above render the gap against the Sol models and Claude Opus 5.5.

## Codex CLI: one model line

To run GPT-6 Astra in Codex CLI, set a top-level `model` line above the `[model_providers]` table in `~/.codex/config.toml` and keep `wire_api = "responses"` on the provider:

```toml
# ~/.codex/config.toml
model = "openai/gpt-6-astra"
model_provider = "router"
model_reasoning_effort = "low"

[model_providers.router]
name = "router"
base_url = "https://api.router.one/v1"
env_key = "ROUTER_ONE_API_KEY"
wire_api = "responses"
```

Per the Codex release notes (checked 2026-10-03), Codex 0.153.1 and later can be configured with GPT-6 Astra; update with `npm install -g @openai/codex` before you switch. The sample uses the catalog id: per the openai/codex source (rust-v0.160.0, checked 2026-10-03), Codex matches a namespaced name such as `openai/gpt-6-astra` to the `gpt-6-astra` entry of its bundled model list, so its settings for GPT-6 Astra apply. Per that bundled list, Codex's default reasoning effort for GPT-6 Astra is `low` — Light in Codex's model docs, which suggest starting there — and it uses a 272,000-token context window for the model by default. The sample sets the effort explicitly because a `config.toml` written by the one-click install carries `model_reasoning_effort = "high"`.

This setup is for Max and Ultra: each GPT-6 Astra request Codex sends draws from the Flagship models allowance — 2 requests above 272,000 input tokens and 4 above 512,000. An agent session sends many requests, so a cycle's Flagship allowance can run out quickly; once it is used, further calls bill the wallet, so switch back to `model = "gpt-5.6-sol"`, which is in the Premium models tier of all three plans. Pro carries no Flagship allowance, so on Pro keep the one-click install's `model = "gpt-5.6-sol"`. The [Codex CLI in China page](https://router.one/codex-china) has the full setup, and the [Codex and Responses API page](https://router.one/codex-responses-api) explains why relays without the Responses wire format fail with Codex.

## Send the first request

1. **Create a key.** Dashboard → API Keys → Create Key. Keys look like `sk-...`. For a trial, give the key a `maxSpend` cap: it cannot spend past that amount, and your other keys keep working ([per-key cost tracking](https://router.one/llm-cost-tracking)).
2. **Keep a wallet balance.** GPT-6 Astra calls bill the wallet on Pro, beyond the Flagship allowance on Max and Ultra, and on the `azure/gpt-6-astra` channel id; top up in Dashboard → Deposit.
3. **Set the base URL** to `https://api.router.one/v1` in any OpenAI-compatible SDK or client.
4. **Call the model.** The examples use the catalog id `openai/gpt-6-astra`, set the reasoning effort explicitly and set an explicit output cap.

Responses, the native path, with the reasoning effort set explicitly:

```bash
curl https://api.router.one/v1/responses \
  -H "Authorization: Bearer sk-your-api-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-6-astra",
    "reasoning": {"effort": "medium"},
    "max_output_tokens": 25000,
    "input": "List three risks of a whole-request price tier for an agent loop."
  }'
```

Tool calls go on this endpoint. A function tool in the Responses format:

```bash
curl https://api.router.one/v1/responses \
  -H "Authorization: Bearer sk-your-api-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-6-astra",
    "reasoning": {"effort": "low"},
    "max_output_tokens": 25000,
    "tools": [{
      "type": "function",
      "name": "get_order_status",
      "description": "Look up the status of an order by its id.",
      "parameters": {
        "type": "object",
        "properties": {"order_id": {"type": "string"}},
        "required": ["order_id"]
      }
    }],
    "input": "Where is order 8812?"
  }'
```

When the output contains a `function_call` item, run the function and send its result back as a `function_call_output` item with the same `call_id`; OpenAI's reasoning guide recommends passing back the reasoning items from that turn as well.

Chat Completions, for clients that send it — without tools, as OpenAI's GPT-6 guide specifies:

```bash
curl https://api.router.one/v1/chat/completions \
  -H "Authorization: Bearer sk-your-api-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-6-astra",
    "reasoning_effort": "medium",
    "max_completion_tokens": 25000,
    "messages": [{"role": "user", "content": "List three risks of a whole-request price tier for an agent loop."}]
  }'
```

The same Responses call from the OpenAI Python SDK — only `base_url` and the key change:

```python
from openai import OpenAI

client = OpenAI(
    base_url="https://api.router.one/v1",
    api_key="sk-your-api-key",
)

response = client.responses.create(
    model="openai/gpt-6-astra",
    reasoning={"effort": "medium"},
    max_output_tokens=25000,
    input="List three risks of a whole-request price tier for an agent loop.",
)
print(response.status, response.output_text)
print(response.usage)
```

Rules that hold on both endpoints:

- **Set an output cap on every request.** Reasoning tokens count toward the output cap; OpenAI suggests reserving at least 25,000 tokens for reasoning and outputs when you start (reasoning guide, checked 2026-10-03). Per the same guide, a response that reaches the cap comes back with `status: "incomplete"`, possibly before any visible text. On Router One's Chat Completions, `max_tokens` and `max_completion_tokens` both set this cap (if both are sent, `max_completion_tokens` wins); on the Responses API it is `max_output_tokens`.
- **Set the reasoning effort.** OpenAI documents no default effort for GPT-6 Astra, so name one on every request, from `low` to `max`.
- **Leave out the sampling parameters.** No `temperature`, `top_p`, `top_logprobs` or `logprobs`, as listed above.
- **Stream long work.** `stream: true` returns output as it is generated, on either endpoint ([streaming guide](https://router.one/llm-streaming)).
- **Stay off `/v1/messages`.** Sending the id there returns HTTP 400 before any model is called, and the message says the model must be called via /v1/chat/completions. The [API compatibility facts](https://router.one/facts/api-compatibility.md) state the endpoint rule per family.

5. **Read the trace.** Dashboard → Logs shows each call with model, input and output tokens, cost, latency and status ([per-request observability](https://router.one/llm-observability)). Retryable upstream failures are absorbed by [automatic fallback](https://router.one/llm-fallback) across the candidate routes for the same model; a GPT-6 Astra request is never answered by a different model.

## Other coding tools

Coding agents send tools on almost every turn, and per OpenAI's GPT-6 guide (checked 2026-10-03) GPT-6 Astra takes tool calls only on the Responses API, so agent work with it needs a client that sends Responses requests. Before you point a tool at it, check which protocol that tool sends: the tool table in [GPT-6 Sol in Codex CLI, Cursor, Cline and OpenCode](https://router.one/blog/gpt-6-sol-coding-tools-setup) lists the request path per tool. Claude Code cannot call GPT-6 Astra: it sends Anthropic Messages requests, and `/v1/messages` serves Claude-family and DeepSeek ids only.

## From mainland China

Requests reach `api.router.one` from mainland China without a VPN, on the same key and base URL. Top up with a card or Alipay through one hosted checkout, or with USDT/USDC on six chains (Tron, BSC, Ethereum, Polygon, Base, Arbitrum). No US credit card required. Pricing starts as low as 10% of official provider list prices (up to 90% off) on select models; the [GPT API page](https://router.one/cheap-gpt-api) and each model page show the rate for a given id. For Codex CLI, see [Codex CLI in China](https://router.one/codex-china).

## FAQ

**What is the model id for GPT-6 Astra on Router One?**
openai/gpt-6-astra, the catalog id — use it on both endpoints and in Codex CLI. azure/gpt-6-astra is a separate channel id of the same base model, with its own rates, and no plan tier covers it.

**Which Router One plans include GPT-6 Astra?**
Max and Ultra, through the Flagship models tier: in the 2026-10-03 plan response, GPT-6 Astra is the only model in the Flagship models tier, and only those two plans carry it. On Pro, its calls bill per token to the wallet. The pricing page shows each plan's live allowances.

**How are GPT-6 Astra requests counted against the Flagship allowance?**
One request per call, 2 when the total input is strictly above 272,000 tokens and 4 above 512,000. Cached input counts once as part of the input, and output does not count. Calls beyond the cycle's allowance bill the wallet at the posted rates.

**Can I call tools on Chat Completions with GPT-6 Astra?**
Not per OpenAI: its GPT-6 guide (checked 2026-10-03) says GPT-6 Astra supports Chat Completions but tool calling requires the Responses API. Send tool-using requests to /v1/responses, where Router One serves GPT models natively, and keep Chat Completions for requests without tools.

**Which reasoning efforts does GPT-6 Astra accept?**
Per OpenAI's model page (checked 2026-10-03): low, medium, high, xhigh and max. OpenAI's reasoning guide says setting none returns HTTP 400, and its GPT-6 guide suggests low instead. OpenAI documents no API default for GPT-6 Astra, so set the effort on every request; in Codex, the bundled default for it is low.

**How do I use GPT-6 Astra in Codex CLI?**
Update Codex — per its release notes (checked 2026-10-03), 0.153.1 and later can be configured with GPT-6 Astra — and put a top-level model = "openai/gpt-6-astra" line above the [model_providers] table in config.toml, with wire_api = "responses" on the provider. On Max or Ultra, each GPT-6 Astra request Codex sends draws from the Flagship models allowance; once it is used, switch back to model = "gpt-5.6-sol". Pro carries no Flagship allowance, so on Pro keep the one-click install's model = "gpt-5.6-sol" (Premium models tier).

**Does the 272K price line apply to GPT-6 Astra, including in Codex CLI?**
Yes. Its model page shows two rate lines: exactly at 272,000 input tokens the standard line applies; strictly above it, the long-context line applies to the whole request, output included. Router One bills Codex CLI's requests like any other API request: on the wallet at the posted GPT-6 Astra rates, long-context line included, and on Max or Ultra a request above 272,000 input tokens draws 2 Flagship requests, 4 above 512,000.

**GPT-6 Astra or GPT-6.1 Sol?**
OpenAI places GPT-6 Astra at its highest intelligence and GPT-6.1 Sol at balanced speed, cost and intelligence (GPT-6 guide, checked 2026-10-03). On Router One the plan status differs: in the 2026-10-03 plan response, GPT-6 Astra draws from the Flagship models allowance on Max and Ultra, while GPT-6.1 Sol is in no plan tier and bills the wallet. The GPT-6.1 Sol vs GPT-6 Astra page renders the live rates side by side.

**Can Claude Code call GPT-6 Astra?**
No. Claude Code sends Anthropic Messages requests to /v1/messages, which serves Claude-family and DeepSeek ids only; a GPT id there returns HTTP 400 before any model is called.

**Do I need a ChatGPT subscription to call GPT-6 Astra?**
No. Router One bills API calls to your Router One wallet or plan. OpenAI's launch post (September 3, 2026) says GPT-6 Astra rolled out in stages to ChatGPT subscribers and to the OpenAI API; a ChatGPT plan plays no part in API calls through Router One.

**Is there no GPT-6.1 Astra?**
OpenAI's model catalog (checked 2026-10-03) has no GPT-6.1 Astra entry, and its GPT-6 guide lists GPT-6.1 Sol beside GPT-6 Astra. Router One lists GPT-6 Astra (openai/gpt-6-astra) and GPT-6.1 Sol (openai/gpt-6.1-sol); the live catalog at /models is the source of truth.

**How much does GPT-6 Astra cost through Router One?**
On the wallet, per token at the posted rates on the GPT-6 Astra model page: two lines, with reasoning billed as output. On Max and Ultra, calls inside the Flagship allowance draw plan requests instead. This guide prints no per-token figures because they go stale; the comparison pages render the live gap against GPT-6.1 Sol, GPT-6 Sol, GPT-5.6 Sol and Claude Opus 5.5.

## Next steps

- Open the [GPT-6 Astra model page](https://router.one/models/gpt-6-astra) for the live rates, both price lines and the endpoint list.
- Compare it with [GPT-6.1 Sol](https://router.one/models/compare/gpt-6-1-sol-vs-gpt-6-astra), [GPT-6 Sol](https://router.one/models/compare/gpt-6-sol-vs-gpt-6-astra), [GPT-5.6 Sol](https://router.one/models/compare/gpt-6-astra-vs-gpt-5-6-sol), [Claude Opus 5.5](https://router.one/models/compare/claude-opus-5-5-vs-gpt-6-astra), [Claude Opus 5](https://router.one/models/compare/gpt-6-astra-vs-claude-opus-5), [Claude Fable 5](https://router.one/models/compare/claude-fable-5-vs-gpt-6-astra) or [Grok 4.7](https://router.one/models/compare/grok-4-7-vs-gpt-6-astra).
- Considering a Sol model? The [GPT-6.1 Sol API guide](https://router.one/blog/gpt-6-1-sol-api-guide) and the [GPT-6 Sol API guide](https://router.one/blog/gpt-6-sol-api-guide) cover them.
- Set up Codex CLI with [Codex CLI in China](https://router.one/codex-china) and the [Codex and Responses API page](https://router.one/codex-responses-api).
- Check which models each plan covers on the [pricing page](https://router.one/pricing).
- See September's other catalog changes in the [September 2026 new-model guide](https://router.one/blog/new-llm-models-september-2026).

## See also

- Canonical page: https://router.one/blog/gpt-6-astra-api-guide
- Codex CLI China: https://router.one/codex-china
- All blog posts: https://router.one/blog
- Models and per-model token rates: https://router.one/models (markdown: https://router.one/models.md)
- Pricing: https://router.one/pricing
- API docs (markdown): https://router.one/docs.md
- Company facts: https://router.one/facts/company.md
