# Connect Hermes Agent to Router One as a custom endpoint

> Markdown mirror of https://router.one/integrations/hermes-agent for AI assistants and crawlers. Router One is a unified, OpenAI-compatible LLM API gateway.
> Last updated: 2026-09-29

Hermes Agent is Nous Research's open-source (MIT) agent: an interactive terminal session with tools, plus a messaging gateway that answers from Telegram, Discord, Slack and other apps. It accepts any OpenAI-compatible endpoint, so one Router One key gives it the Claude, GPT, Gemini, Grok and DeepSeek chat models in the catalog, and every model request it sends, its side tasks included, appears in Dashboard → Logs with tokens and cost. You can add Router One through the Custom endpoint entry of the hermes model wizard, or as a named provider in config.yaml, which also lets a model run on the Anthropic Messages or Responses transport. Checked against Hermes Agent v0.21.5 (v2026.9.24, released 2026-09-24) and its docs and source on 2026-09-29.

## Install Hermes Agent 0.21.5 and create a dedicated key

Install with the official one-line installer rather than pip: as of 2026-09-29 the hermes-agent package on PyPI is still at 0.19.0 (2026-07-20), while the GitHub release and the installer are at v0.21.5. On macOS, Linux and WSL2 the installer needs only git and curl, and it sets up Python, Node.js and Hermes' other dependencies itself. Windows is supported natively: run the PowerShell command below in PowerShell, and Hermes installs under %LOCALAPPDATA%\hermes (inside WSL2 it installs under ~/.hermes, as on Linux). When it runs in a terminal, the installer ends by starting hermes setup, whose Model & Provider section is the same flow as hermes model, so you can enter the Router One values from the next section there or later. Then create a key for Hermes alone in Dashboard → API Keys with Create Key and give it a maxSpend cap: every model turn of an agent task is a billed request and Hermes' side tasks add more, so the cap is the hard stop, and a dedicated key keeps Hermes' requests together in Logs.

`terminal`

```text
# macOS, Linux, WSL2 (needs git and curl)
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
source ~/.bashrc     # or source ~/.zshrc, or open a new terminal

# Windows (native PowerShell)
iex (irm https://hermes-agent.nousresearch.com/install.ps1)

hermes --version
```

## Configure Hermes Agent to use the Router One base URL

Run hermes model in a terminal, outside any chat session, and choose Custom endpoint. At API base URL enter https://api.router.one/v1, and at API key paste the dedicated key. Hermes then checks the endpoint with GET /v1/models; Router One answers that call only when a key is present, and the reply lists every catalog ID, image-generation models included, so type the exact chat ID rather than picking a number from the list. At Select API compatibility mode choose 2, Chat Completions: api.router.one is not on Hermes' auto-detect list, and Chat Completions serves every chat model. Enter a model ID from /models such as anthropic/claude-sonnet-5, leave Context length in tokens blank for now (see the context section below) and set Display name to Router One. Hermes saves the key in its .env file as HERMES_CUSTOM_API_ROUTER_ONE_API_KEY and writes provider: custom, the base URL, the API mode and the model to config.yaml, which only references the key. To keep Router One next to other providers, or to add another transport later, define it as a named provider in config.yaml instead:

`config.yaml`

```yaml
# ~/.hermes/config.yaml — native Windows / 原生 Windows: %LOCALAPPDATA%\hermes\config.yaml
providers:
  router-one:
    api: https://api.router.one/v1
    key_env: ROUTER_ONE_API_KEY      # set in Hermes' .env / 写在 Hermes 的 .env 中
    transport: chat_completions

model:
  provider: custom:router-one
  default: anthropic/claude-sonnet-5

compression:
  threshold_tokens: 256000         # v0.21.5 default; below 200000 for Grok IDs / v0.21.5 默认值；用 Grok ID 时设到 200000 以下

# ~/.hermes/.env
# ROUTER_ONE_API_KEY=sk-your-router-one-key
```

## Pick a transport for each Router One model ID

A named provider speaks one transport: chat_completions, which Hermes also uses when the field is empty, anthropic_messages or codex_responses. Router One serves every chat model on Chat Completions, Claude-family and DeepSeek V4 IDs on Anthropic Messages, and GPT-family, DeepSeek V4 and Grok chat IDs natively on Responses, so one chat_completions entry covers every model, and a second entry only changes the wire format for one family. The api value follows the transport: Hermes' Anthropic client removes a trailing /v1 because its SDK appends /v1/messages itself, so either form works there, while the two OpenAI transports append /chat/completions or /responses to /v1. Sent over codex_responses, a Claude ID is rejected with HTTP 400 before any model runs (model '<id>' must be called via /v1/messages or /v1/chat/completions); sent over anthropic_messages, a chat ID outside the Claude and DeepSeek families gets 400 with must be called via /v1/chat/completions. Hermes sends the key as Authorization: Bearer on the OpenAI transports and as x-api-key to a third-party Anthropic host; Router One accepts both, so every entry can share one key.

`config.yaml`

```yaml
providers:
  router-one:
    api: https://api.router.one/v1
    key_env: ROUTER_ONE_API_KEY
  router-one-claude:
    api: https://api.router.one
    key_env: ROUTER_ONE_API_KEY
    transport: anthropic_messages

# switch inside a Hermes session:
#   /model custom:router-one:openai/gpt-5.6-sol
#   /model custom:router-one-claude:anthropic/claude-sonnet-5
```

| transport | api | Router One model IDs |
| --- | --- | --- |
| chat_completions (default) | https://api.router.one/v1 | Every chat model in the catalog: Claude, GPT, Gemini, Grok and DeepSeek IDs |
| anthropic_messages | https://api.router.one | Claude-family IDs, aws/claude-… and vertex/claude-… channel IDs included, plus deepseek-v4.1-flash and deepseek-v4-flash |
| codex_responses | https://api.router.one/v1 | GPT-family IDs, azure/gpt-… channel IDs included, plus deepseek-v4.1-flash, deepseek-v4-flash and Grok chat IDs |

## Set the context length and the compression point yourself

Hermes plans compression around the model's context window. A context_length you set, in the model section or per model under providers.<name>.models, always wins; left empty, Hermes tries its cache, the endpoint's /models reply and public model registries before it falls back to a built-in default. Router One's GET /v1/models reply has no field Hermes reads as the window, so the detected value may not match the Router One model page. At startup Hermes prints the window it uses as Context limit and marks a configured value as pinned. The model page on /models shows each model's full window, which is the value to enter if you set context_length. Per Hermes' configuration docs and source (v0.21.5), compression fires at the lower of two triggers: compression.threshold, 0.50 of the window (raised to 0.75 for windows under 512K), and compression.threshold_tokens, which defaults to 256000. A session with a 1M window therefore compresses at about 256,000 input tokens, and every input token up to that point is billed. Some model pages also list a long-context price line that covers the whole request once its input is above a threshold: 272,000 tokens on openai/gpt-5.6-sol and 200,000 on grok-4.7. The default 256000 is below openai/gpt-5.6-sol's 272,000 line but above grok-4.7's 200,000 line, so for Grok IDs set compression.threshold_tokens below 200000, for example 190000, if you want long sessions to compress before they reach that line; per Hermes' docs the value stays in force when you switch models. On the codex_responses transport, per Hermes' docs, GPT slugs such as gpt-5.6-sol take their window from Hermes' Codex table (272K for most) unless you set context_length, and per its source the openai/ prefix of a Router One ID does not change that.

## Reasoning effort, side tasks and embeddings

Per Hermes' docs, chat_completions requests carry reasoning_effort: medium when you have not set an effort, and a level you set with /reasoning or agent.reasoning_effort reaches the endpoint unchanged, up to max. Router One does not turn reasoning_effort into a thinking setting for Claude IDs, except on Claude Opus 5.5, where it maps the value to output_config.effort; other models may ignore it or reject the request with 400, so check the status of your first requests in Dashboard → Logs, and again after changing the level. Hermes also runs side tasks, among them session titles, context compression and image description, on your main model unless you route them elsewhere, and each one is another request on the same key. To send a task to a cheaper Router One model, set auxiliary.<task>.provider to your named provider and auxiliary.<task>.model to an exact ID, as below, or use Configure auxiliary models in hermes model. Router One serves no embeddings endpoint, so if you turn on a Hermes feature that needs an embedding model, give it another provider.

`config.yaml`

```yaml
auxiliary:
  compression:
    provider: router-one
    model: google/gemini-3.8-flash
  vision:
    provider: router-one
    model: google/gemini-3.8-flash
```

## What Hermes sends, and how to read its 4xx errors

Each model turn is one request: a task that reads files, runs commands and edits code sends one per tool round, and the side tasks above add their own, so filter Dashboard → Logs by the Hermes key. Each record shows the model, tokens, cost, status, total time and, for streamed requests that produced output, time to first token (TTFT). Read errors by status: 401 means the key is missing or wrong (with key_env, check that the variable is really in Hermes' .env); 402 means the wallet balance or the key's maxSpend is used up, and once the cap is reached the key's requests return 402 while your other keys keep working; 404 usually means a wrong path or an ID outside the catalog, so copy the ID again from /models; 400 with must be called via means the transport does not serve that ID. A turn that never reached Router One leaves no record, which points back at the local configuration.

## Which model ID should Hermes Agent send?

Copy the exact model ID from /models, preserving case, hyphens, and version suffixes; do not substitute a display name. Open its detail page and match the supported API endpoints, context window, and capabilities such as tool calling to the provider and features selected in Hermes Agent. A catalog listing does not mean the client can use every feature of that model. Give each client or application a dedicated API key with a maxSpend cap.

## Which API protocol is Hermes Agent using?

OpenAI-compatible describes an interface format; it does not make Chat Completions (/v1/chat/completions), Responses (/v1/responses), and Anthropic Messages (/v1/messages) interchangeable. Check the installed client version, provider configuration, and actual request path against the model detail page and API compatibility fact sheet. A successful plain-text chat does not establish support for hosted tools, conversation state, or file-editing features.

## Verify the Hermes Agent call in your request trace

Send a simple text request from Hermes Agent, then match its trace in Dashboard → Logs by time, model, and request_id: tokens, cost, latency, and status. Next, test streaming, tool calls, and multi-turn history separately. For failures, retain the actual request path, full error message, and request_id. If there is no matching log, check client configuration and connectivity before attributing the error to the gateway or upstream.

## FAQ

### Where does Hermes keep the Router One settings?

In config.yaml in the Hermes home directory: ~/.hermes on macOS, Linux and WSL2, and %LOCALAPPDATA%\hermes on native Windows, unless HERMES_HOME points elsewhere. The key goes in the .env file next to it. hermes config path and hermes config env-path print the two locations, and hermes config edit opens the config in your editor. The Custom endpoint wizard saves the key as HERMES_CUSTOM_API_ROUTER_ONE_API_KEY and config.yaml only references it; with a named provider, key_env names the variable. Keep both files out of dotfile repositories and shared backups.

### I set OPENAI_BASE_URL, but Hermes still does not call Router One. Why?

Per Hermes' provider docs, OPENAI_BASE_URL is honored only for its openai-api provider; a custom endpoint or a named provider reads its URL from config.yaml. Run hermes model and choose Custom endpoint, or set model.provider to custom:router-one as in the example above, then start a new session. The /model command inside a chat only switches between providers you have already configured; it cannot add one.

### Can I switch between Claude and GPT in one Hermes session?

Yes, among configured providers. /model custom:router-one:openai/gpt-5.6-sol switches to a GPT ID on the router-one entry, and /model custom:router-one-claude:anthropic/claude-sonnet-5 switches to a Claude ID on an anthropic_messages entry, if you defined one. A context_length set in the model section is dropped when you switch, so give each model you use its own context_length under providers.<name>.models.

### The model list shows image models and IDs without capability labels. Which should I pick?

GET /v1/models returns the whole catalog, including image-generation models and IDs whose model page lists no capabilities. Hermes works through tool calls, so pick a chat model whose page on /models lists tool calling; for other IDs, test one tool call first. Image-generation IDs cannot serve Hermes' chat turns.

### Can I install Hermes with pip install hermes-agent?

You would get an older release. As of 2026-09-29 the hermes-agent package on PyPI is at 0.19.0 (released 2026-07-20), while the GitHub release and the official installer are at v0.21.5 (v2026.9.24). The installer also sets up Hermes' own Python environment and dependencies, and hermes update keeps it current afterwards.

### Which models can Hermes Agent use through the gateway?

Choose a current catalog model that supports both the endpoint and the features Hermes Agent uses. Check /models and the model detail page for the exact ID, current rates, and capabilities; a family name such as GPT or Claude is not a compatibility guarantee. A client that lists a model has only read the ID, from its own configuration or from GET /v1/models; verify an actual request too.

### Models are listed, but requests fail with 400 or 404. What should I check?

Record the actual request path and error message, then check the exact model ID. A 400 can indicate invalid parameters, unsupported tools, or a model/endpoint mismatch; a 404 can indicate an incorrect path or missing resource, so it does not by itself establish that a model was retired. If the error says must be called via, use the named endpoint or select a model supported on the current endpoint. Do not add or remove /v1 or /chat/completions across all clients indiscriminately.

### Does this work from Mainland China?

Hermes' model requests to Router One work from Mainland China without a VPN, with the same configuration as elsewhere. Installing and updating Hermes is a separate matter: the installer clones the Hermes repository from GitHub and downloads its Python tooling, so it depends on your network's access to those hosts.

### How do I debug a 401/402/403/429?

Match the request and error message in Dashboard → Logs. For 401, check whether the key was sent and is valid; for 402, check wallet balance and maxSpend; for 403, check key permissions and access restrictions. For 429, distinguish request/token limits from upstream throttling using the error details. Keep the request_id and follow the error-codes reference.

## See also

- All integration guides: https://router.one/integrations
- Debug API errors in Hermes Agent: https://router.one/llm-api-error-codes
- API compatibility: endpoints and supported features: https://router.one/facts/api-compatibility.md
- Responses API setup and limits: https://router.one/codex-responses-api
- GitHub Copilot CLI setup: https://router.one/integrations/copilot-cli
- Cherry Studio setup: https://router.one/integrations/cherry-studio
- CLI setup guide: the Hermes Agent tab: https://router.one/docs/guides/cli-setup
- OpenClaw: a self-hosted assistant on the same key: https://router.one/integrations/openclaw
- Tool calling through the gateway: https://router.one/llm-tool-calling
- Hermes Agent docs: AI providers: https://hermes-agent.nousresearch.com/docs/integrations/providers
- Hermes Agent docs: installation: https://hermes-agent.nousresearch.com/docs/getting-started/installation
- Hermes Agent releases: https://github.com/NousResearch/hermes-agent/releases
- LLM API gateway overview: https://router.one/llm-api-gateway
- OpenAI-compatible API: https://router.one/openai-compatible-api
- API docs: https://router.one/docs
- Canonical page: https://router.one/integrations/hermes-agent
- Models and per-model token rates: https://router.one/models (markdown: https://router.one/models.md)
- Pricing: https://router.one/pricing
- API docs (markdown): https://router.one/docs.md
- Company facts: https://router.one/facts/company.md
