Skip to content
Back to Models
Text model

Grok 4.6 API pricing

xAI's Grok series — chat plus image generation.

Model IDgrok-4.6
ChatStreamingTool callingVision

Call summary

Input
$2.20$1.10/ 1M tokens
Cache write $1.10
Cache read $0.275
Output
$6.60$3.30/ 1M tokens

Applies to request input ≤ 200,000 tokens · See pricing tiers

Context Window
500K
API Endpoints
2

Grok 4.6 API endpoints

One API key — call this model through any of the 2 endpoints below.

POST/v1/chat/completionsOpenAI-compatible · Works for every model
POST/v1/responsesResponses API · Native for Codex CLI

Grok 4.6 pricing tiers

Pricing is tiered by total input length per request (cache included); once a threshold is crossed, the whole request is billed at that tier. All prices are per 1M tokens.

Tier
Standard≤ 200K
Input
$2.20$1.10Cache write: $1.10Cache read: $0.275
Output
$6.60$3.30
Tier
Long context> 200K
Input
$4.40$2.20Cache write: $2.20Cache read: $0.55
Output
$13.20$6.60

Production reliability

Privacy-safe aggregates from real Router One model calls, so you can assess stability and response latency before integrating.

Successful response share

94.06%

Success 2xx

94.06%

Rate limit 429

0.00%

Server 5xx

5.94%

TPS

53.14

tokens/s
Avg. time to first token

7.66 s

Avg. latency

37.7 s

Includes production requests ending in 2xx, 429, or 5xx. Average latency uses successful requests; time to first token and TPS require complete observations from successful streams. Public thresholds are 100 requests and 5 independent principals.

Updated Sep 30, 2026, 7:49 AM UTC

More Grok models on Router One

7 other Grok models on the same gateway and API key — each page lists endpoints, posted price, and context window.

Grok 4.6 compared with other models

Spec and price matchups against peer models.

Grok 4.6 at a glance

Grok 4.6 (model ID grok-4.6) is a Grok-series text model on Router One. Grok 4.6 is priced in whole-request tiers. Request input ≤ 200,000 tokens: Input $1.10 / 1M tokens, Output $3.30 / 1M tokens, Cache write $1.10 / 1M tokens, Cached input $0.275 / 1M tokens; request input > 200,000 tokens: Input $2.20 / 1M tokens, Output $6.60 / 1M tokens, Cache write $2.20 / 1M tokens, Cached input $0.55 / 1M tokens. The total input tokens in each request select the tier; its rates apply to the whole request, not only tokens above the threshold. Router One subscriptions include Grok 4.6 in the Mid-tier models allowance: Pro 1,200, Max 4,500, Ultra 15,000 requests per 30-day cycle. Usage beyond the allowance bills the wallet at posted rates. Within a plan allowance, a request above 200,000 input tokens counts as 2 requests. Context window: 500K tokens. Endpoints: POST /v1/chat/completions (OpenAI-compatible), POST /v1/responses (Responses API, the Codex CLI path).

Grok 4.6 FAQ

How much does Grok 4.6 cost on Router One?

Grok 4.6 is priced in whole-request tiers. Request input ≤ 200,000 tokens: Input $1.10 / 1M tokens, Output $3.30 / 1M tokens, Cache write $1.10 / 1M tokens, Cached input $0.275 / 1M tokens; request input > 200,000 tokens: Input $2.20 / 1M tokens, Output $6.60 / 1M tokens, Cache write $2.20 / 1M tokens, Cached input $0.55 / 1M tokens. The total input tokens in each request select the tier; its rates apply to the whole request, not only tokens above the threshold.

Is Grok 4.6 included in Router One subscriptions?

Router One subscriptions include Grok 4.6 in the Mid-tier models allowance: Pro 1,200, Max 4,500, Ultra 15,000 requests per 30-day cycle. Usage beyond the allowance bills the wallet at posted rates. Within a plan allowance, a request above 200,000 input tokens counts as 2 requests.

Which endpoints serve Grok 4.6?

Grok 4.6 is served on 2 endpoints — POST /v1/chat/completions (OpenAI-compatible), POST /v1/responses (Responses API, the Codex CLI path). The same Router One API key works on each.

What is Grok 4.6's context window?

Grok 4.6 accepts up to 500K tokens of context per request on Router One.

Start using Grok 4.6

Create an API key and call Grok 4.6 at $1.10 in / $3.30 out per 1M tokens (request input ≤ 200,000 tokens) on Router One — pay as you go, with per-request cost and latency visibility.