Router One
Back to Models
Long context

claude-opus-4.8 API pricing

Anthropic's Claude series — long context, dependable tool calling, and the native models behind Claude Code.

Model IDanthropic/claude-opus-4.8
ChatStreamingTool callingVision

Call summary

Input
$7.50$2.25/ 1M tokens
Cache write $4.69
Output
$37.50$11.25/ 1M tokens
Cache read $0.38
Context Window
1.05M
Max output
8K
API Endpoints
2
Available Configurations
3

Production reliability

Privacy-safe aggregates from real Router One model calls, so you can assess stability and response latency before integrating.

Successful response share

97.70%

Success 2xx

97.70%

Rate limit 429

0.00%

Server 5xx

2.30%

TPS

68.12

tokens/s
Avg. time to first token

17.22 s

Avg. latency

55.82 s

Includes production requests ending in 2xx, 429, or 5xx. Average latency uses successful requests; time to first token and TPS require complete observations from successful streams. Public thresholds are 100 requests and 5 independent principals.

Updated Jul 26, 2026, 4:25 PM UTC

claude-opus-4.8 API endpoints

One API key — call this model through any of the endpoints below.

POST/v1/chat/completionsOpenAI-compatible · Works for every model
POST/v1/messagesAnthropic-native · Native for Claude Code

claude-opus-4.8 compared with other models

Spec and price matchups against peer models.