Skip to content

Configure Immersive Translate's custom Chat Completions API

Immersive Translate is a bilingual translation browser extension: it places the translation next to the original text on web pages, and its legacy PDF and EPUB readers use the same translation service. Its built-in OpenAI service accepts your own API key and a custom API address instead of an Immersive Translate Pro membership, and that is how Router One connects: every batch of paragraphs becomes one POST to /v1/chat/completions, the endpoint that serves every chat model in the catalog, so you can translate with Claude, Gemini, DeepSeek, GPT or Grok models through one key and see the tokens and cost of each batch in Dashboard → Logs. Translation is high-volume work, so this guide covers the address to enter (https://api.router.one/v1/chat/completions works in every version), how to add catalog model names, how many requests a page produces and how to keep them under rate limits, which models suit bulk translation, and the quality check that can discard an answer. Checked against Immersive Translate 1.33.3 (the Chrome Web Store build) and 1.30.2 (the build Firefox Add-ons serves) on 2026-09-27.

Set Immersive Translate's custom API endpoint and key

In Immersive Translate, open Settings → Translation Services → OpenAI and switch the service from Pro to Custom API Key. Paste your Router One key into APIKEY. Click Expand for more custom settings and enter the Custom API interface address. The field's own help text says it accepts a base URL or a full endpoint URL, and the full Chat Completions URL is the one value every version handles the same way. Then open Set up more models, add the catalog IDs you want (next section) and select one as the Model. Click Verify service, translate a short page, and match the new entries in Dashboard → Logs by time and model before you translate a whole document:

Custom API interface address(自定义 API 接口地址)
https://api.router.one/v1/chat/completions
immersive-translate-settings
# Immersive Translate → Settings → Translation Services → OpenAI
# 沉浸式翻译 → 设置 → 翻译服务 → OpenAI
Service(服务):  Custom API Key(自定义 API Key), not Pro
APIKEY:  sk-your-router-one-key
Custom API interface address(自定义 API 接口地址), works in every version:
  https://api.router.one/v1/chat/completions
Set up more models(设置更多模型):  -all,+anthropic/claude-haiku-4.5,+deepseek-v4-flash
Model(模型):  pick one of the IDs you added, then click Verify service(点此测试服务)

Which address to enter: full URL, /v1 or bare host

Immersive Translate rewrites what you type into a request URL, and the rule changed between versions. The Chrome Web Store serves 1.33.3 (Edge Add-ons lists 1.33.1), and 1.33.3 completes almost anything; Firefox Add-ons still serves 1.30.2, which is stricter about a bare host. Enter the full URL and the version does not matter. An address that ends in /responses keeps that path, and Router One serves /v1/responses only for GPT-family models, DeepSeek IDs and Grok chat models, so stay on /chat/completions.

Which address to enter: full URL, /v1 or bare host
You enter1.33.3 (Chrome)1.30.2 (Firefox Add-ons)Verdict
https://api.router.one/v1/chat/completionsUsed as enteredUsed as enteredRecommended: works in every version
https://api.router.one/v1/chat/completions appended/chat/completions appendedWorks in 1.30.2 and later
https://api.router.oneCompleted to /v1/chat/completionsBecomes /chat/completions: HTTP 404 with a hint to add /v1Avoid

Add catalog model names with Set up more models

The OpenAI service's built-in model list holds OpenAI's own names without a vendor prefix, such as gpt-5.5 and gpt-5.4-mini, so add the exact catalog IDs yourself. Set up more models (设置更多模型) takes a comma-separated list: a name adds it (a leading + does the same), -name hides one of the built-in names, -all hides all of them, and name=Label gives a name already in the list a friendlier label. For example, -all,+anthropic/claude-haiku-4.5,+deepseek-v4-flash,+google/gemini-3.1-flash-lite leaves only those three. Alternatively, type one ID into Enter the name of the custom model (输入自定义模型名称). The ID must match /models exactly, because it is sent unchanged as the model field and an unknown ID is rejected. Current rates are on each model page; translation multiplies them by a lot of tokens, so compare them before you choose.

How many requests a page makes, and staying under rate limits

The OpenAI service sends up to 4 paragraphs per request (Maximum number of paragraphs per request) and, per the official docs, up to 10 requests per second by default, so one long article produces dozens of requests in Dashboard → Logs, each billed for its own input and output tokens. Translations are cached locally for 30 days by default, so reopening a page you already translated may send nothing new. The rate controls sit under Expand for more custom settings: Max requests per second, Max requests per minute, Maximum text length per request and Maximum number of paragraphs per request. Router One applies default request and token limits per key and per account. When a burst crosses one, it answers 429 with RATE_LIMIT_EXCEEDED or TOKEN_QUOTA_EXCEEDED and an X-RateLimit-Scope header naming the limit, and the affected paragraphs fail to translate. Lower Max requests per second (the official docs suggest about 5 for ebooks) or Max requests per minute. APIKEY also accepts several comma-separated keys, which the docs call load balancing, but keys from one Router One account still share that account's limit. If your normal volume needs more, key and account limits can be raised on request: email support@router.one with your account ID, the key's name and your expected peak.

Pick a model for bulk translation, and the quality check

Translation sends the same kind of request thousands of times, so a small, fast, non-reasoning chat model is usually the right trade. Reasonable candidates are anthropic/claude-haiku-4.5, deepseek-v4-flash, deepseek-v4.1-flash, google/gemini-3.1-flash-lite, google/gemini-3.7-flash and openai/gpt-5.6-terra. Avoid IDs that think on every request: anthropic/claude-opus-5.5 always uses adaptive thinking, and IDs ending in -thinking think by default, so each batch also bills thinking tokens at the output rate. Immersive Translate 1.33.3 adjusts some GPT requests itself: for names matching gpt-5.6 and its Sol, Terra and Luna variants, and gpt-5.1, gpt-5.2, gpt-5.4 or gpt-5.5, it sets reasoning_effort to none and drops temperature, and the match also catches prefixed IDs such as openai/gpt-5.6-terra and azure/gpt-5.6-sol. After each answer the extension compares output tokens with input tokens; when the ratio falls outside its range (0.21 to 8 by default), it discards the answer and falls back to another translation. Setting strictPrompt: true skips that check, which the docs recommend only for custom prompts that are not plain translation.

Advanced settings: temperature, timeout, rate and the output cap

Everything above can also be set as JSON under Developer settings → Edit Full User Config, inside translationServices.openai: temperature, requestTimeout in milliseconds (101000 by default for AI services), limit (requests per second), maxTextGroupLengthPerRequest, maxTextLengthPerRequest, bodyConfigs and headerConfigs for extra request fields, modelsOverrides for per-model changes, strictPrompt and langOverrides. The request path lives in baseUrlApiPath, /chat/completions by default, which is why a base URL works. One field deserves attention: 1.33.3 sends its output cap as max_completion_tokens when the host is OpenAI's or Azure's, or when the last segment of the model name starts with gpt-5, o1, o3 or o4, which includes openai/gpt-5.6-terra. Router One's Chat Completions reference documents max_tokens as the output cap, so to cap output, set it in bodyConfigs; the extension then uses max_tokens for every model. Back up your config before editing, since a JSON error makes the extension ignore it:

Edit Full User Config
{
  "translationServices": {
    "openai": {
      "temperature": 0.2,
      "requestTimeout": 60000,
      "limit": 5,
      "bodyConfigs": { "max_tokens": 2048 }
    }
  }
}

Which model ID should Immersive Translate send?

Copy the exact model ID from /models, preserving case, hyphens, and version suffixes; do not substitute a display name. Open its detail page and match the supported API endpoints, context window, and capabilities such as tool calling to the provider and features selected in Immersive Translate. A catalog listing does not mean the client can use every feature of that model. Give each client or application a dedicated API key with a maxSpend cap.

Which API protocol is Immersive Translate using?

OpenAI-compatible describes an interface format; it does not make Chat Completions (/v1/chat/completions), Responses (/v1/responses), and Anthropic Messages (/v1/messages) interchangeable. Check the installed client version, provider configuration, and actual request path against the model detail page and API compatibility fact sheet. A successful plain-text chat does not establish support for hosted tools, conversation state, or file-editing features.

Verify the Immersive Translate call in your request trace

Send a simple text request from Immersive Translate, then match its trace in Dashboard → Logs by time, model, and request_id: tokens, cost, latency, and status. Next, test streaming, tool calls, and multi-turn history separately. For failures, retain the actual request path, full error message, and request_id. If there is no matching log, check client configuration and connectivity before attributing the error to the gateway or upstream.

FAQ

Should the custom API address be the base URL or the full endpoint?

Either works in current versions, and the full endpoint works in all of them. From 1.30.2 on, the field accepts a base URL ending in /v1, and the extension appends /chat/completions; 1.33.3 even completes a bare host. The Firefox Add-ons build (1.30.2) turns a bare host into /chat/completions, which Router One answers with a 404 hint. Entering https://api.router.one/v1/chat/completions avoids the difference entirely. If translations fail right away, check this field first, then the model ID against /models.

Why does translating one web page produce dozens of requests in Logs?

Because the extension splits the page into paragraphs and sends them in small batches, up to 4 paragraphs per request by default, and it translates more as you scroll. Each batch is one billed request with its own tokens. That is normal; to see what a page costs, filter Dashboard → Logs by the key you gave the extension and add up the entries for that time window. Pages you translated before are served from the local cache for 30 days by default and cost nothing.

Immersive Translate reports 429 errors. What should I change?

Slow the extension down. It can send up to 10 requests per second by default, and a long page or an ebook can cross your key's or account's default limit. The 429 carries RATE_LIMIT_EXCEEDED or TOKEN_QUOTA_EXCEEDED and an X-RateLimit-Scope header saying which limit it was. Lower Max requests per second, to about 5 as the official docs suggest for ebooks, or set Max requests per minute. Several keys from the same account do not raise the account limit; for sustained higher volume, email support@router.one to have your limits raised.

Which model is the best value for translation?

Usually a small, fast chat model without extended thinking, because translation sends many short requests. Start with anthropic/claude-haiku-4.5, deepseek-v4-flash or google/gemini-3.1-flash-lite, translate the same page with two or three of them, and compare quality against the cost of those requests in Dashboard → Logs. Avoid anthropic/claude-opus-5.5 and IDs ending in -thinking for bulk work: they think on every request, and thinking bills as output. Current per-token rates are on each model page.

Do I need an Immersive Translate Pro membership?

Not for this setup. The extension's OpenAI service offers Pro, which uses Immersive Translate's own access, or Custom API Key, which uses yours; the official docs present them as alternatives. With Custom API Key, requests go to the address you enter and are billed by Router One per request.

Should I use the built-in OpenAI service or Add Custom Translation Service?

Either works with the same values. The built-in OpenAI service is the quickest: switch it to Custom API Key and fill in the fields above. Add Custom Translation Service (添加自定义翻译服务) creates a separate entry with its own name, which helps if you want Router One next to another provider or two entries with different models. Enter a Custom Translation Service Name such as Router One, the same Custom API interface address, your key and a model name, then click Verify service. The custom service follows the same address rules as the built-in one.

A translation sometimes disappears and a different one replaces it. Why?

That is the extension's quality check. It compares the number of output tokens with the input; when the ratio is outside its range, 0.21 to 8 by default, it treats the answer as invalid and falls back to another translation. It happens most with prompts that ask for more than translation, or with models that add commentary. Use a model that answers with the translation only, or set strictPrompt: true for the openai service in Edit Full User Config if your custom prompt is not a plain translation.

How do I set temperature, a timeout or an output limit?

Under Developer settings → Edit Full User Config, add them to translationServices.openai, for example temperature 0.2, requestTimeout 60000 and bodyConfigs with max_tokens. Router One's Chat Completions reference documents max_tokens as the output cap, and once bodyConfigs contains max_tokens the extension uses that field for every model. Back up the config first: a JSON mistake makes the extension ignore the whole block.

Which models can Immersive Translate use through the gateway?

Choose a current catalog model that supports both the endpoint and the features Immersive Translate uses. Check /models and the model detail page for the exact ID, current rates, and capabilities; a family name such as GPT or Claude is not a compatibility guarantee. A client that lists a model has only read the ID, from its own configuration or from GET /v1/models; verify an actual request too.

Models are listed, but requests fail with 400 or 404. What should I check?

Record the actual request path and error message, then check the exact model ID. A 400 can indicate invalid parameters, unsupported tools, or a model/endpoint mismatch; a 404 can indicate an incorrect path or missing resource, so it does not by itself establish that a model was retired. If the error says must be called via, use the named endpoint or select a model supported on the current endpoint. Do not add or remove /v1 or /chat/completions across all clients indiscriminately.

Does this work from Mainland China?

Yes. The gateway is reachable from Mainland China without a VPN, and the configuration is identical to the global setup.

How do I debug a 401/402/403/429?

Match the request and error message in Dashboard → Logs. For 401, check whether the key was sent and is valid; for 402, check wallet balance and maxSpend; for 403, check key permissions and access restrictions. For 429, distinguish request/token limits from upstream throttling using the error details. Keep the request_id and follow the error-codes reference.