Router One
Back to Models
Long context

gemini-3.5-flash API pricing

Google's Gemini series — very long context windows and strong multimodal input.

Model IDgoogle/gemini-3.5-flash
ChatStreamingTool callingVision

Call summary

Input
$1.50$0.45/ 1M tokens
Cache write $0.45
Output
$9.00$2.70/ 1M tokens
Cache read $0.45
Context Window
1.05M
Max output
66K
API Endpoints
1
Available Configurations
2

Production reliability

Privacy-safe aggregates from real Router One model calls, so you can assess stability and response latency before integrating.

Successful response share

98.19%

Success 2xx

98.19%

Rate limit 429

0.00%

Server 5xx

1.81%

TPS

569.65

tokens/s
Avg. time to first token

4.05 s

Avg. latency

5.85 s

Includes production requests ending in 2xx, 429, or 5xx. Average latency uses successful requests; time to first token and TPS require complete observations from successful streams. Public thresholds are 100 requests and 5 independent principals.

Updated Jul 26, 2026, 4:25 PM UTC

gemini-3.5-flash API endpoints

One API key — call this model through any of the endpoints below.

POST/v1/chat/completionsOpenAI-compatible · Works for every model

gemini-3.5-flash compared with other models

Spec and price matchups against peer models.