Gemini 3.8 Flash vs GPT-5.6 Terra
Gemini 3.8 Flash is the lower-priced of the two on Router One at base-tier rates (GPT-5.6 Terra: request input ≤ 272,000 tokens) — about 23% less on a 1M-input + 1M-output mix. Both carry a 1.05M context window. Gemini 3.8 Flash answers on /v1/chat/completions; GPT-5.6 Terra on /v1/chat/completions and /v1/responses (Codex CLI). Gemini 3.8 Flash counts against the Standard models allowance on Pro, Max and Ultra; GPT-5.6 Terra counts against the Mid-tier models allowance on Pro, Max and Ultra.
Gemini 3.8 Flash and GPT-5.6 Terra compared on current per-token rates, context window, and capabilities — both callable through one OpenAI-compatible endpoint with per-request cost traces.
Gemini 3.8 Flash vs GPT-5.6 Terra: rates, context window, and capabilities
| Spec | Gemini 3.8 Flash | GPT-5.6 Terra |
|---|---|---|
| Input / 1M tokens | $0.225 | $0.25 |
| Output / 1M tokens | $1.125 | $1.50 |
| Cached input / 1M tokens | $0.225 | $0.125 |
| Context window | 1.05M | 1.05M |
| Capabilities | Chat, Streaming, Tool calling, Vision | Chat, Streaming, Tool calling, Vision |
| Subscription plans | Standard models — Pro, Max, Ultra | Mid-tier models — Pro, Max, Ultra |
| Detail page | Gemini 3.8 Flash | GPT-5.6 Terra |
GPT-5.6 Terra: request input ≤ 272,000 tokens: Input $0.25 / 1M tokens, Output $1.50 / 1M tokens, Cache write $0.25 / 1M tokens, Cached input $0.125 / 1M tokens; request input > 272,000 tokens: Input $0.50 / 1M tokens, Output $2.25 / 1M tokens, Cache write $0.50 / 1M tokens, Cached input $0.25 / 1M tokens. The total input tokens in each request select the tier; its rates apply to the whole request, not only tokens above the threshold.
What does 1M tokens cost on Gemini 3.8 Flash vs GPT-5.6 Terra?
Base-tier rates: For a workload of 1M input plus 1M output tokens at current rates: Gemini 3.8 Flash comes to $1.35, GPT-5.6 Terra comes to $1.75 — Gemini 3.8 Flash is about 23% cheaper on this mix. Real workloads skew heavily toward input tokens, so weigh the input rate by your own ratio; the cached-input row above is the posted catalog rate for that model.
Switch between Gemini 3.8 Flash and GPT-5.6 Terra
Both models answer on /v1/chat/completions, so an A/B test there is a one-string change — same key, same code, and every request traced with tokens, cost, and latency in the dashboard. Codex CLI uses /v1/responses, which Router One lists for GPT-5.6 Terra but not Gemini 3.8 Flash — so in that tool, switch the tool (or call /v1/chat/completions), not only the model string.
curl https://api.router.one/v1/chat/completions \
-H "Authorization: Bearer sk-your-router-one-key" \
-H "Content-Type: application/json" \
-d '{"model": "google/gemini-3.8-flash", "messages": [{"role": "user", "content": "Hello"}]}'
# Same request, other model — change one string:
# "model": "openai/gpt-5.6-terra"FAQ
Is Gemini 3.8 Flash cheaper than GPT-5.6 Terra?
Compared at base-tier rates (GPT-5.6 Terra: request input ≤ 272,000 tokens). Input: Gemini 3.8 Flash $0.225 vs GPT-5.6 Terra $0.25 / 1M tokens; Gemini 3.8 Flash has the lower rate. Output: Gemini 3.8 Flash $1.125 vs GPT-5.6 Terra $1.50 / 1M tokens; Gemini 3.8 Flash has the lower rate. Total cost depends on the workload's input, output and cache usage. GPT-5.6 Terra: request input ≤ 272,000 tokens: Input $0.25 / 1M tokens, Output $1.50 / 1M tokens, Cache write $0.25 / 1M tokens, Cached input $0.125 / 1M tokens; request input > 272,000 tokens: Input $0.50 / 1M tokens, Output $2.25 / 1M tokens, Cache write $0.50 / 1M tokens, Cached input $0.25 / 1M tokens. The total input tokens in each request select the tier; its rates apply to the whole request, not only tokens above the threshold. Rates change; the /models page is the live source of truth.
Is Gemini 3.8 Flash or GPT-5.6 Terra included in a Router One subscription?
Gemini 3.8 Flash counts against the Standard models allowance on Pro, Max and Ultra: Pro 5,000, Max 8,000 and Ultra 25,000 requests per 30-day cycle, shared by every model in that tier. GPT-5.6 Terra counts against the Mid-tier models allowance on Pro, Max and Ultra: Pro 1,200, Max 4,500 and Ultra 15,000 requests per 30-day cycle, shared by every model in that tier. Requests to Gemini 3.8 Flash or GPT-5.6 Terra above the model's long-context threshold count as more than one quota request. Requests beyond a plan's allowance bill the wallet per token.
Can I switch between Gemini 3.8 Flash and GPT-5.6 Terra without changing code?
On /v1/chat/completions, yes — change the model string; the key and base URL stay the same. Codex CLI uses /v1/responses, which Router One lists for GPT-5.6 Terra but not Gemini 3.8 Flash — so in that tool, switch the tool (or call /v1/chat/completions), not only the model string.
Where do these numbers come from?
Specs and prices on this page render from the live Router One catalog — the same data as the /models page — and refresh with it. Pricing methodology is documented on /pricing-methodology. Rates on this page were read from the live catalog on 2026-09-30 (UTC) and refresh within the hour.