Skip to content
Router One
Back to Models
Long context

deepseek-v4-flash API pricing

DeepSeek's models — cost-efficient reasoning and coding.

Model IDdeepseek-v4-flash
ChatStreamingTool calling

Call summary

Input
$0.60/ 1M tokens
Cache write $0.60
Output
$2.40/ 1M tokens
Cache read $0.012
Context Window
1M
API Endpoints
3

deepseek-v4-flash API endpoints

One API key — call this model through any of the 3 endpoints below.

POST/v1/chat/completionsOpenAI-compatible · Works for every model
POST/v1/messagesAnthropic-native · Native for Claude Code
POST/v1/responsesResponses API · Native for Codex CLI

Production reliability

Privacy-safe aggregates from real Router One model calls, so you can assess stability and response latency before integrating.

Successful response share

98.99%

Success 2xx

98.99%

Rate limit 429

1.01%

Server 5xx

0.00%

TPS

125.23

tokens/s
Avg. time to first token

6.2 s

Avg. latency

10.43 s

Includes production requests ending in 2xx, 429, or 5xx. Average latency uses successful requests; time to first token and TPS require complete observations from successful streams. Public thresholds are 100 requests and 5 independent principals.

Updated Sep 20, 2026, 12:44 PM UTC

More DeepSeek models on Router One

1 other DeepSeek model on the same gateway and API key — its page lists endpoints, posted price, and context window.

deepseek-v4-flash at a glance

deepseek-v4-flash is a DeepSeek-series text model on Router One. Posted rate: $0.60 in / $2.40 out per 1M tokens. Context window: 1M tokens. Accepts text; returns text. Endpoints: POST /v1/chat/completions (OpenAI-compatible), POST /v1/messages (Anthropic-native, the Claude Code path), POST /v1/responses (Responses API, the Codex CLI path).

deepseek-v4-flash FAQ

How much does deepseek-v4-flash cost on Router One?

deepseek-v4-flash is billed per token on Router One: $0.60 in / $2.40 out per 1M tokens, pay as you go.

Which endpoints serve deepseek-v4-flash?

deepseek-v4-flash is served on 3 endpoints — POST /v1/chat/completions (OpenAI-compatible), POST /v1/messages (Anthropic-native, the Claude Code path), POST /v1/responses (Responses API, the Codex CLI path). The same Router One API key works on each.

What is deepseek-v4-flash's context window?

deepseek-v4-flash accepts up to 1M tokens of context per request on Router One.

Start using deepseek-v4-flash

Create an API key and call deepseek-v4-flash at $0.60 in / $2.40 out per 1M tokens on Router One — pay as you go, with per-request cost and latency visibility.