Grok 4.7 vs Grok 4.6
The listed rates and 1M-input + 1M-output workload example use base tiers, assuming every request meets: grok-4.7: request input ≤ 200,000 tokens; grok-4.6: request input ≤ 200,000 tokens. Requests above a threshold use the long-context rates below. Grok 4.7 and Grok 4.6 come to the same total on Router One for a 1M-input + 1M-output mix ($4.40), so price does not decide this pair. Both carry a 500K context window. Both answer on /v1/chat/completions and /v1/responses (Codex CLI).
Grok 4.7 and Grok 4.6 compared on current per-token rates, context window, and capabilities — both callable through one OpenAI-compatible endpoint with per-request cost traces.
Grok 4.7 vs Grok 4.6: rates, context window, and capabilities
| Spec | Grok 4.7 | Grok 4.6 |
|---|---|---|
| Input / 1M tokens | $1.10 | $1.10 |
| Output / 1M tokens | $3.30 | $3.30 |
| Cached input / 1M tokens | $0.275 | $0.275 |
| Context window | 500K | 500K |
| Capabilities | Chat, Streaming, Tool calling, Vision | Chat, Streaming, Tool calling, Vision |
| Detail page | Grok 4.7 | Grok 4.6 |
grok-4.7: request input ≤ 200,000 tokens: Input $1.10 / 1M tokens, Output $3.30 / 1M tokens, Cache write $1.10 / 1M tokens, Cached input $0.275 / 1M tokens; request input > 200,000 tokens: Input $2.20 / 1M tokens, Output $6.60 / 1M tokens, Cache write $2.20 / 1M tokens, Cached input $0.55 / 1M tokens. The total input tokens in each request select the tier; its rates apply to the whole request, not only tokens above the threshold.
grok-4.6: request input ≤ 200,000 tokens: Input $1.10 / 1M tokens, Output $3.30 / 1M tokens, Cache write $1.10 / 1M tokens, Cached input $0.275 / 1M tokens; request input > 200,000 tokens: Input $2.20 / 1M tokens, Output $6.60 / 1M tokens, Cache write $2.20 / 1M tokens, Cached input $0.55 / 1M tokens. The total input tokens in each request select the tier; its rates apply to the whole request, not only tokens above the threshold.
What does 1M tokens cost on Grok 4.7 vs Grok 4.6?
The listed rates and 1M-input + 1M-output workload example use base tiers, assuming every request meets: grok-4.7: request input ≤ 200,000 tokens; grok-4.6: request input ≤ 200,000 tokens. Requests above a threshold use the long-context rates below. For a workload of 1M input plus 1M output tokens at current rates, Grok 4.7 and Grok 4.6 both come to $4.40. Price does not separate this pair, so choose on context window, capabilities, and the latency your own request traces show; real workloads skew heavily toward input tokens, so compare the input rates by your own ratio.
Switch between Grok 4.7 and Grok 4.6 without changing code
Both models are behind the same OpenAI-compatible endpoint, so an A/B test is a one-string change — same key, same code, and every request traced with tokens, cost, and latency in the dashboard:
curl https://api.router.one/v1/chat/completions \
-H "Authorization: Bearer sk-your-router-one-key" \
-H "Content-Type: application/json" \
-d '{"model": "grok-4.7", "messages": [{"role": "user", "content": "Hello"}]}'
# Same request, other model — change one string:
# "model": "grok-4.6"FAQ
Is Grok 4.7 cheaper than Grok 4.6?
The listed rates and 1M-input + 1M-output workload example use base tiers, assuming every request meets: grok-4.7: request input ≤ 200,000 tokens; grok-4.6: request input ≤ 200,000 tokens. Requests above a threshold use the long-context rates below. Input: Grok 4.7 $1.10 vs Grok 4.6 $1.10 / 1M tokens; the rates are equal. Output: Grok 4.7 $3.30 vs Grok 4.6 $3.30 / 1M tokens; the rates are equal. Total cost depends on the workload's input, output and cache usage. grok-4.7: request input ≤ 200,000 tokens: Input $1.10 / 1M tokens, Output $3.30 / 1M tokens, Cache write $1.10 / 1M tokens, Cached input $0.275 / 1M tokens; request input > 200,000 tokens: Input $2.20 / 1M tokens, Output $6.60 / 1M tokens, Cache write $2.20 / 1M tokens, Cached input $0.55 / 1M tokens. The total input tokens in each request select the tier; its rates apply to the whole request, not only tokens above the threshold. grok-4.6: request input ≤ 200,000 tokens: Input $1.10 / 1M tokens, Output $3.30 / 1M tokens, Cache write $1.10 / 1M tokens, Cached input $0.275 / 1M tokens; request input > 200,000 tokens: Input $2.20 / 1M tokens, Output $6.60 / 1M tokens, Cache write $2.20 / 1M tokens, Cached input $0.55 / 1M tokens. The total input tokens in each request select the tier; its rates apply to the whole request, not only tokens above the threshold. Rates change; the /models page is the live source of truth.
Can I switch between Grok 4.7 and Grok 4.6 without changing code?
Yes. Both are served through the same OpenAI-compatible endpoint, so switching is changing the model string in the request — the key, base URL, and request shape stay identical.
Where do these numbers come from?
Specs and prices on this page render from the live Router One catalog — the same data as the /models page — and refresh with it. Pricing methodology is documented on /pricing-methodology. Rates on this page were read from the live catalog on 2026-09-22 (UTC) and refresh within the hour.