Back to Models
Long context
gemini-3.5-flash API pricing
Google's Gemini series — very long context windows and strong multimodal input.
Model ID
google/gemini-3.5-flashChatStreamingTool callingVision
Call summary
Input
$1.50$0.45/ 1M tokens
Cache write $0.45
Output
$9.00$2.70/ 1M tokens
Cache read $0.45
Context Window
1.05M
Max output
66K
API Endpoints
1
Available Configurations
2
Production reliability
Privacy-safe aggregates from real Router One model calls, so you can assess stability and response latency before integrating.
Successful response share
98.19%
Success 2xx
98.19%
Rate limit 429
0.00%
Server 5xx
1.81%
TPS
569.65
tokens/sAvg. time to first token
4.05 s
Avg. latency
5.85 s
Includes production requests ending in 2xx, 429, or 5xx. Average latency uses successful requests; time to first token and TPS require complete observations from successful streams. Public thresholds are 100 requests and 5 independent principals.
Updated Jul 26, 2026, 4:25 PM UTC
gemini-3.5-flash API endpoints
One API key — call this model through any of the endpoints below.
gemini-3.5-flash compared with other models
Spec and price matchups against peer models.