Back to Models
Low cost
gpt-5.4-mini API pricing
OpenAI's GPT series — broad general capability, wide ecosystem, and the native models behind Codex CLI.
Model ID
openai/gpt-5.4-miniChatStreamingTool callingVision
Call summary
Input
$0.25$0.07/ 1M tokens
Cache write $0.04
Output
$1.00$0.30/ 1M tokens
Cache read $0.00
Context Window
400K
Max output
128K
API Endpoints
2
Available Configurations
5
Production reliability
Privacy-safe aggregates from real Router One model calls, so you can assess stability and response latency before integrating.
Successful response share
99.49%
Success 2xx
99.49%
Rate limit 429
0.00%
Server 5xx
0.51%
TPS
75.16
tokens/sAvg. time to first token
6.38 s
Avg. latency
15.09 s
Includes production requests ending in 2xx, 429, or 5xx. Average latency uses successful requests; time to first token and TPS require complete observations from successful streams. Public thresholds are 100 requests and 5 independent principals.
Updated Jul 26, 2026, 4:25 PM UTC
gpt-5.4-mini API endpoints
One API key — call this model through any of the endpoints below.
gpt-5.4-mini compared with other models
Spec and price matchups against peer models.