Field manual

Issue 01 · Local AI Gateway / LLM Router

OmniRoute automatic routing AI Gateway

OmniRoute setup for 290+ providers, 500+ models and 19 routing strategies. Connect Codex, check failover, compression, MCP/A2A and production limits.

OmniRoute connects 290+ providers and 500+ models through a local compatible endpoint, so Codex, Claude Code and other AI tools can share connection settings. It has 19 routing strategies, three resilience layers, CLI configurators, a compression pipeline and per-request decision information.

diegosouzapw/omniroute
stars
—
forks
—
License
—
Data as of
—
Reading time
8 min
Updated
Open original report
GitHub Stars
34.9k
Routing strategies
19
Resilience layers
3
Open-source license
MIT

01System role

Local compatible endpoints and upstream routing

OmniRoute is an AI Gateway that runs in the user's environment. It exposes compatible interfaces such as http://localhost:20128/v1 to tools and forwards requests to connected upstream providers; the gateway can convert OpenAI, Anthropic, Gemini and Responses API formats.

auto does not require a predefined combo. The system builds a virtual candidate pool from connected providers and selects a model using health, quota, cost, latency and historical performance signals; if an upstream provider fails, the routing layer can select the next available candidate.

The response header X-OmniRoute-Decision identifies the strategy, provider and routing latency. Other X-OmniRoute-* headers provide cost and usage information. These headers verify routing results and supplement dashboard status information.

  1. IDE / CLI

  2. localhost:20128/v1

  3. Auto candidate pool

  4. Resilience checks

  5. Upstream provider

  6. Decision headers

“No combo to create. Set your model to auto.”

— 官方 README · Combos

02Installation and connection

Installation, startup and connection checks

The npm package requires Node.js ≥ 22.22.2 and < 23, or ≥ 24 and < 27. After a global installation, run omniroute; the dashboard and Gateway use local port 20128.

bash
# 全域安裝並啟動 Gateway 與儀表板
npm install -g omniroute
omniroute

# 儀表板:http://localhost:20128
# OpenAI 相容 API:http://localhost:20128/v1

Provider and API verification

Connect an available provider on the dashboard's Providers page, then obtain an API key from Endpoints. Start with the model set to auto and use the model-list endpoint to check authorization and connectivity.

bash
# OpenAI 相容工具設定
Base URL: http://localhost:20128/v1
API Key:  [從儀表板 → Endpoints 複製]
Model:    auto

# 驗證授權並列出可用模型
curl http://localhost:20128/v1/models -H "Authorization: Bearer YOUR_KEY"

Docker and the Codex launcher

The official container command binds the service to loopback only and stores data in a named volume. Start Codex with omniroute launch-codex; launchers of this kind do not rewrite existing configuration files.

bash
docker run -d --name omniroute --restart unless-stopped --stop-timeout 40 \
  -p 127.0.0.1:20128:20128 -v omniroute-data:/app/data diegosouzapw/omniroute:latest

# 以目前環境啟動 Codex,不寫入設定檔
omniroute launch-codex

03Core capabilities

Nineteen strategies and supporting capabilities

19 strategies cover fixed ordering, load distribution, cost, quota, context, caching, scoring and multi-model workflows. Establish a baseline with auto, then select an explicit strategy for a reproducible requirement.

Route · 01

priority · fill-first

Fixed ordering and quota filling

priority follows list order; fill-first uses the current candidate until its conditions are met, then moves to the next.

Route · 02

weighted · round-robin

Weights and rotation

weighted distributes requests according to configured weights; round-robin cycles through candidates for predictable distribution.

Route · 03

p2c · least-used

Low-load selection

p2c compares two random candidates; least-used selects the candidate with lower recent usage.

Route · 04

random · strict-random

Random distribution

random can select again after a failure; strict-random retains the strict semantics of a single random decision.

Route · 05

cost-optimized · headroom

Cost and remaining quota

cost-optimized prioritizes unit price; headroom prioritizes remaining quota.

Route · 06

reset-window · reset-aware

Quota reset windows

Both strategies factor quota reset times into decisions, for periodic quotas and connections across multiple accounts.

Route · 07

context-relay · context-optimized

Context management

context-relay continues long conversations across models; context-optimized selects candidates according to context conditions.

Route · 08

cache-optimized · lkgp

Caching and historical performance

cache-optimized considers caching benefits; lkgp uses last-known-good performance to help select a candidate.

Route · 09

auto · fusion · pipeline

Scoring and workflows

auto scores dynamically; fusion combines results from multiple models; pipeline runs multi-stage work in sequence.

Resilience · 10

breaker · cooldown · lockout

Three layers of failure isolation

Provider circuit breakers, connection cooldowns and model lockouts address different failure scopes; model lockout is disabled by default.

Compression · 11

RTK + Caveman

Request compression

The compression pipeline can reduce tokens sent to models; run npm run eval:compression in a source checkout for offline evaluation before production use.

Protocol · 12

--mcp · A2A · connect

Agents and remote control

omniroute --mcp provides MCP; A2A supports agent interoperability; omniroute connect connects to a remote instance.

Auto model identifiers

Model identifierPrimary objective
autoBalanced default with last-known-good performance
auto/codingQuality for coding tasks
auto/fastLow latency
auto/cheapLow cost
auto/offlineRemaining quota
auto/smartQuality scoring with a retained exploration share

04Operating principles

Verifiable configuration sequence

These configuration principles come from the current default branch's README and official documentation. Each has a configuration or verification point.

TIP 01 · auto baseline

With the model set to auto, the system builds a virtual candidate pool from current connections in real time, without first writing a combo. Confirm the selected result before switching to a scenario variant or fixed strategy.

Source · docs/routing/AUTO-COMBO.md

TIP 02 · Decision headers and routing checks

Read X-OmniRoute-Decision to confirm the actual strategy, provider and routing latency. Other X-OmniRoute-* headers supply cost and usage information.

Source · Official README · Routing Decision

TIP 03 · Configuration dry run

setup-codex creates ~/.codex/<name>.config.toml. Use --dry-run first to inspect the planned content and avoid overwriting existing working settings.

Source · docs/guides/CODEX-CLI-CONFIGURATION.md

TIP 04 · Launchers for temporary integration

omniroute launch-codex and other launchers pass values needed by the current environment without writing tool configuration files. Use them to verify a single session.

Source · docs/guides/CLI-INTEGRATIONS.md

TIP 05 · Scope of the three resilience layers

Provider circuit breakers, connection cooldowns and model lockouts address different failure levels. Model lockout is disabled by default; do not assume all three layers are enabled.

Source · docs/architecture/RESILIENCE_GUIDE.md

TIP 06 · Header authorization

Use Authorization: Bearer for standard integrations. Use a compatible path containing a token only when the client cannot send custom headers.

Source · docs/guides/CLI-INTEGRATIONS.md

TIP 07 · Memory activation

Memory is disabled by default. It stores relevant data locally only after activation; check retention scope and cleanup procedures before handling sensitive content.

Source · Official README · Memory

TIP 08 · Individual security checks

Prompt injection guard defaults to warning mode; credential masker requires explicit activation. Guardrails cannot be the only security boundary in production.

Source · docs/security/GUARDRAILS.md

05Worked example

Codex connection and routing verification

The terminal transcript illustrates the operating sequence. Commands come from the official README and CLI documentation; bracketed text describes expected observations, without fixing the provider, cost or latency.

~/projects/app · OmniRoute 操作示意


$ npm install -g omniroute
$ omniroute


# [確認儀表板可由 http://localhost:20128 開啟]
# [在 Providers 連接供應商,並於 Endpoints 取得 key]


$ omniroute doctor
# [檢查 Gateway、連線與必要設定]


$ curl http://localhost:20128/v1/models -H "Authorization: Bearer YOUR_KEY"
# [確認回應包含目前可用模型]


$ omniroute launch-codex
# [launcher 啟動 Codex,不寫入設定檔]


$ codex ›
  檢查這個專案的測試失敗原因。


# [request model=auto/coding]
# [由 X-OmniRoute-Decision 核對策略、供應商與延遲]
ok: [由 X-OmniRoute-* 核對成本與使用量]

        

“One config — http://localhost:20128/v1 — and every AI IDE or CLI runs on free & low-cost models.”

— 官方 README · Quick Start

Verification points

A successful startup still requires routing checks. Confirm the model list, then inspect X-OmniRoute-Decision on an actual request for the selected provider and strategy. If compression is enabled, read X-OmniRoute-Compression and compare output quality using representative workloads.

For persistent settings, run setup-codex --dry-run to inspect the planned changes. Codex uses /v1 as its Base URL; other protocols may have different root paths, so do not copy the same value directly.

06Boundaries and limitations

Production limitations

07Next steps

Local verification and controlled deployment

After establishing a local connection, add health checks, strategy baselines, configuration management, remote access and observability in sequence. Keep reproducible verification results at each step.

Adoption sequence

1. Establish a local health baseline. After installation and Providers and Endpoints setup, run omniroute doctor, a model-list request and an actual inference request.

2. Compare Auto variants. Test auto, auto/coding, auto/fast and auto/cheap with the same workload. Record decision headers, latency, cost and task quality.

3. Persist settings and strategies. Inspect persistent settings with setup-codex --dry-run; create a persistent combo when fixed ordering, cost caps or context requirements apply.

4. Evaluate agent and remote interfaces. Enable MCP or A2A when agent control is needed; use omniroute connect for remote instances and configure least privilege and protected network endpoints.

5. Add observability. Retain X-OmniRoute-Decision, cost, usage and compression headers. Connect provider failures, fallback and quota events to existing monitoring workflows.

Three further documents

① docs/routing/AUTO-COMBO.md: Auto candidate pools, model identifiers and scoring.

② docs/guides/CLI-INTEGRATIONS.md: Setup and launcher behavior for each tool.

③ docs/architecture/RESILIENCE_GUIDE.md: Circuit breaker, cooldown and model lockout scopes.

“Self-managing model chains with adaptive scoring + zero-config auto-routing.”

— docs/routing/AUTO-COMBO.md