Router One
Back to Models
Long context

gpt-5.6-sol API pricing

OpenAI's GPT series — broad general capability, wide ecosystem, and the native models behind Codex CLI.

Model IDopenai/gpt-5.6-sol
ChatStreamingTool callingVision

Call summary

Input
$5.00$0.50/ 1M tokens
Cache write $3.13
Output
$40.00$4.00/ 1M tokens
Cache read $0.25
Context Window
1.05M
Max output
128K
API Endpoints
2
Available Configurations
4

Production reliability

Privacy-safe aggregates from real Router One model calls, so you can assess stability and response latency before integrating.

Successful response share

98.55%

Success 2xx

98.55%

Rate limit 429

0.00%

Server 5xx

1.45%

TPS

80.05

tokens/s
Avg. time to first token

10.65 s

Avg. latency

21.38 s

Includes production requests ending in 2xx, 429, or 5xx. Average latency uses successful requests; time to first token and TPS require complete observations from successful streams. Public thresholds are 100 requests and 5 independent principals.

Updated Jul 26, 2026, 4:25 PM UTC

gpt-5.6-sol API endpoints

One API key — call this model through any of the endpoints below.

POST/v1/chat/completionsOpenAI-compatible · Works for every model
POST/v1/responsesResponses API · Native for Codex CLI

gpt-5.6-sol pricing tiers

Pricing is tiered by total input length per request (cache included); once a threshold is crossed, the whole request is billed at that tier. All prices are per 1M tokens.

Tier
Standard≤ 272K
Input
$5.00$0.50Cache write: $3.13
Output
$40.00$4.00Cache read: $0.25
Tier
Long context> 272K
Input
$10.00$1.00Cache write: $6.25
Output
$45.00$4.50Cache read: $0.50