Local-first AI gateway routing coding agents across hundreds of providers/models with fallback, dashboard, MCP/A2A, and RTK/Caveman prompt compression.
Local-first AI gateway routing coding agents across hundreds of providers/models with fallback, dashboard, MCP/A2A, and RTK/Caveman prompt compression.
Typical uses
Centralize model access, route around quota/outages, apply configurable compression, and operate through API, CLI, dashboard, desktop/PWA, MCP, or A2A.
Why people choose it
Inference: reduces provider lock-in and manual fallback while adding one control plane for routing and token management.
License
MIT
Maintenance
active
Latest release
v3.8.50
Last meaningful update
Unknown
Maturity
established
Production readiness
development
Classification
Subject domains
LLM routingtoken compressionagent infrastructure
Task categories
LLM routingtoken compressionagent infrastructure
Project phases
implementationoperations
Secondary types
gatewayCLIMCP serverdesktop app
Select when
Need one endpoint across many model providers
Need routing/fallback/compression
Do not select when
Only one stable provider is needed
Prompt transformation is unacceptable unless disabled
Implementation complexity
medium
Setup effort
medium
Learning curve
medium
Confidence
high
Routing heuristic 78.0%
Routing specificity
4/5
Setup simplicity
3/5
Operational maturity
3/5
Composability
4/5
Local control
3/5
Security boundary clarity
5/5
First-party routing documentation
5/5
Strengths
None documented.
Poor-fit scenarios
Only one stable provider is needed
Prompt transformation is unacceptable unless disabled
Local-first AI gateway routing coding agents across hundreds of providers/models with fallback, dashboard, MCP/A2A, and RTK/Caveman prompt compression.
documentation · Compression wiki validated through public GitHub web page because repository contents API does not expose wiki pages.
Repository usage guidance
installation summary
npx; Docker; desktop
basic usage summary
Centralize model access, route around quota/outages, apply configurable compression, and operate through API, CLI, dashboard, desktop/PWA, MCP, or A2A.
Both reduce context; OmniRoute can stack RTK-style and prose compression.
Recommended combinations
Not part of a curated combination.
Routing rules
rule_006: Need one endpoint across many model providers P994
Rationale
Local-first AI gateway routing coding agents across hundreds of providers/models with fallback, dashboard, MCP/A2A, and RTK/Caveman prompt compression.