Local compatible endpoints and upstream routing
OmniRoute is an AI Gateway that runs in the user's environment. It exposes compatible interfaces such as http://localhost:20128/v1 to tools and forwards requests to connected upstream providers; the gateway can convert OpenAI, Anthropic, Gemini and Responses API formats.
auto does not require a predefined combo. The system builds a virtual candidate pool from connected providers and selects a model using health, quota, cost, latency and historical performance signals; if an upstream provider fails, the routing layer can select the next available candidate.
The response header X-OmniRoute-Decision identifies the strategy, provider and routing latency. Other X-OmniRoute-* headers provide cost and usage information. These headers verify routing results and supplement dashboard status information.
IDE / CLI
localhost:20128/v1
Auto candidate pool
Resilience checks
Upstream provider
Decision headers
“No combo to create. Set your model to auto.”
Installation, startup and connection checks
The npm package requires Node.js ≥ 22.22.2 and < 23, or ≥ 24 and < 27. After a global installation, run omniroute; the dashboard and Gateway use local port 20128.
# 全域安裝並啟動 Gateway 與儀表板
npm install -g omniroute
omniroute
# 儀表板:http://localhost:20128
# OpenAI 相容 API:http://localhost:20128/v1Provider and API verification
Connect an available provider on the dashboard's Providers page, then obtain an API key from Endpoints. Start with the model set to auto and use the model-list endpoint to check authorization and connectivity.
# OpenAI 相容工具設定
Base URL: http://localhost:20128/v1
API Key: [從儀表板 → Endpoints 複製]
Model: auto
# 驗證授權並列出可用模型
curl http://localhost:20128/v1/models -H "Authorization: Bearer YOUR_KEY"Docker and the Codex launcher
The official container command binds the service to loopback only and stores data in a named volume. Start Codex with omniroute launch-codex; launchers of this kind do not rewrite existing configuration files.
docker run -d --name omniroute --restart unless-stopped --stop-timeout 40 \
-p 127.0.0.1:20128:20128 -v omniroute-data:/app/data diegosouzapw/omniroute:latest
# 以目前環境啟動 Codex,不寫入設定檔
omniroute launch-codexNineteen strategies and supporting capabilities
19 strategies cover fixed ordering, load distribution, cost, quota, context, caching, scoring and multi-model workflows. Establish a baseline with auto, then select an explicit strategy for a reproducible requirement.
Route · 01
priority · fill-first
Fixed ordering and quota filling
priority follows list order; fill-first uses the current candidate until its conditions are met, then moves to the next.
Route · 02
weighted · round-robin
Weights and rotation
weighted distributes requests according to configured weights; round-robin cycles through candidates for predictable distribution.
Route · 03
p2c · least-used
Low-load selection
p2c compares two random candidates; least-used selects the candidate with lower recent usage.
Route · 04
random · strict-random
Random distribution
random can select again after a failure; strict-random retains the strict semantics of a single random decision.
Route · 05
cost-optimized · headroom
Cost and remaining quota
cost-optimized prioritizes unit price; headroom prioritizes remaining quota.
Route · 06
reset-window · reset-aware
Quota reset windows
Both strategies factor quota reset times into decisions, for periodic quotas and connections across multiple accounts.
Route · 07
context-relay · context-optimized
Context management
context-relay continues long conversations across models; context-optimized selects candidates according to context conditions.
Route · 08
cache-optimized · lkgp
Caching and historical performance
cache-optimized considers caching benefits; lkgp uses last-known-good performance to help select a candidate.
Route · 09
auto · fusion · pipeline
Scoring and workflows
auto scores dynamically; fusion combines results from multiple models; pipeline runs multi-stage work in sequence.
Resilience · 10
breaker · cooldown · lockout
Three layers of failure isolation
Provider circuit breakers, connection cooldowns and model lockouts address different failure scopes; model lockout is disabled by default.
Compression · 11
RTK + Caveman
Request compression
The compression pipeline can reduce tokens sent to models; run npm run eval:compression in a source checkout for offline evaluation before production use.
Protocol · 12
--mcp · A2A · connect
Agents and remote control
omniroute --mcp provides MCP; A2A supports agent interoperability; omniroute connect connects to a remote instance.
Auto model identifiers
| Model identifier | Primary objective |
|---|---|
auto | Balanced default with last-known-good performance |
auto/coding | Quality for coding tasks |
auto/fast | Low latency |
auto/cheap | Low cost |
auto/offline | Remaining quota |
auto/smart | Quality scoring with a retained exploration share |
Verifiable configuration sequence
These configuration principles come from the current default branch's README and official documentation. Each has a configuration or verification point.
TIP 01 · auto baseline
With the model set to auto, the system builds a virtual candidate pool from current connections in real time, without first writing a combo. Confirm the selected result before switching to a scenario variant or fixed strategy.
Source · docs/routing/AUTO-COMBO.md
TIP 02 · Decision headers and routing checks
Read X-OmniRoute-Decision to confirm the actual strategy, provider and routing latency. Other X-OmniRoute-* headers supply cost and usage information.
Source · Official README · Routing Decision
TIP 03 · Configuration dry run
setup-codex creates ~/.codex/<name>.config.toml. Use --dry-run first to inspect the planned content and avoid overwriting existing working settings.
Source · docs/guides/CODEX-CLI-CONFIGURATION.md
TIP 04 · Launchers for temporary integration
omniroute launch-codex and other launchers pass values needed by the current environment without writing tool configuration files. Use them to verify a single session.
Source · docs/guides/CLI-INTEGRATIONS.md
TIP 05 · Scope of the three resilience layers
Provider circuit breakers, connection cooldowns and model lockouts address different failure levels. Model lockout is disabled by default; do not assume all three layers are enabled.
Source · docs/architecture/RESILIENCE_GUIDE.md
TIP 06 · Header authorization
Use Authorization: Bearer for standard integrations. Use a compatible path containing a token only when the client cannot send custom headers.
Source · docs/guides/CLI-INTEGRATIONS.md
TIP 07 · Memory activation
Memory is disabled by default. It stores relevant data locally only after activation; check retention scope and cleanup procedures before handling sensitive content.
Source · Official README · Memory
TIP 08 · Individual security checks
Prompt injection guard defaults to warning mode; credential masker requires explicit activation. Guardrails cannot be the only security boundary in production.
Source · docs/security/GUARDRAILS.md
Codex connection and routing verification
The terminal transcript illustrates the operating sequence. Commands come from the official README and CLI documentation; bracketed text describes expected observations, without fixing the provider, cost or latency.
$ npm install -g omniroute
$ omniroute
$ omniroute doctor
$ curl http://localhost:20128/v1/models -H "Authorization: Bearer YOUR_KEY"
$ omniroute launch-codex
$ codex ›
檢查這個專案的測試失敗原因。
ok: [由 X-OmniRoute-* 核對成本與使用量]
“One config — http://localhost:20128/v1 — and every AI IDE or CLI runs on free & low-cost models.”
Verification points
A successful startup still requires routing checks. Confirm the model list, then inspect X-OmniRoute-Decision on an actual request for the selected provider and strategy. If compression is enabled, read X-OmniRoute-Compression and compare output quality using representative workloads.
For persistent settings, run setup-codex --dry-run to inspect the planned changes. Codex uses /v1 as its Base URL; other protocols may have different root paths, so do not copy the same value directly.
Production limitations
Local verification and controlled deployment
After establishing a local connection, add health checks, strategy baselines, configuration management, remote access and observability in sequence. Keep reproducible verification results at each step.
Adoption sequence
1. Establish a local health baseline. After installation and Providers and Endpoints setup, run omniroute doctor, a model-list request and an actual inference request.
2. Compare Auto variants. Test auto, auto/coding, auto/fast and auto/cheap with the same workload. Record decision headers, latency, cost and task quality.
3. Persist settings and strategies. Inspect persistent settings with setup-codex --dry-run; create a persistent combo when fixed ordering, cost caps or context requirements apply.
4. Evaluate agent and remote interfaces. Enable MCP or A2A when agent control is needed; use omniroute connect for remote instances and configure least privilege and protected network endpoints.
5. Add observability. Retain X-OmniRoute-Decision, cost, usage and compression headers. Connect provider failures, fallback and quota events to existing monitoring workflows.
Three further documents
① docs/routing/AUTO-COMBO.md: Auto candidate pools, model identifiers and scoring.
② docs/guides/CLI-INTEGRATIONS.md: Setup and launcher behavior for each tool.
③ docs/architecture/RESILIENCE_GUIDE.md: Circuit breaker, cooldown and model lockout scopes.
“Self-managing model chains with adaptive scoring + zero-config auto-routing.”