omliteroute
v1.0.14
Published
OmliteRoute — Ultra-lightweight AI Gateway & Model Router. Low-RAM (~120MB), 356 providers, auto-fallback, and zero-bloat.
Downloads
2,499
Maintainers
Readme
🪶 OmliteRoute
The Ultra-Lightweight, Low-RAM AI Gateway & Smart Model Router
356+ Providers · Auto-Fallback · RTK Compression · Low-RAM (~120MB) · Zero-Bloat
# ⚡ Install globally via npm (Recommended)
npm install -g omliteroute
omliteroute
# Or run instantly without installing
npx omliteroute💡 What is OmliteRoute?
OmliteRoute is an independent, hyper-optimized AI proxy and model router engineered specifically for laptops, low-resource VPSs, and developer workstations.
Standard AI proxy routers frequently consume 1.5 to 3.0 GB of RAM and drain battery with background scraping, unmanaged caches, and heavy WebGL animations. OmliteRoute strips away all non-routing bloatware while preserving 100% of the core AI routing engine:
- One Endpoint for Everything: Point Claude Code, Codex, Cursor, Cline, OpenCode, Hermes, or standard OpenAI SDKs to
http://localhost:20128/v1. - Zero-Downtime Resilience: Automatic fallback and account rotation across 356 AI providers (Claude, GPT, Gemini, DeepSeek, Grok, Kimi, Mistral, Ollama, etc.) whenever rate-limits (HTTP 429) or quota errors occur.
- Token Compression Built-in: Native RTK and Caveman prompt compression to save 15% to 85% on token costs.
- Featherlight Footprint: Runs comfortably in ~120–160 MB of RAM with <1% idle CPU usage.
📊 Performance Benchmark: Standard Gateway vs. OmliteRoute
| Metric | Standard Gateway | OmliteRoute 🪶 | Efficiency Gain | | :--- | :---: | :---: | :---: | | Server RAM (Idle) | 600 MB – 1.2 GB | ~120 – 160 MB | ~75% RAM Reduction | | Server RAM (Under Load) | 1.5 GB – 3.0 GB | ~250 – 400 MB | ~85% RAM Reduction | | Browser CPU (Dashboard Tab) | 15% – 25% (WebGL canvas) | < 1% (CSS-only grid) | Zero GPU / battery drain | | Database Disk Growth | Unbounded (100–500+ MB) | < 10 MB (auto-purged) | Permanent lean storage | | i18n Payload | 37.0 MB (42 languages) | 1.5 MB (en & id) | 35.5 MB freed | | Cold Boot Time | 7 – 12 seconds | ~1.5 seconds | 5x faster startup |
✨ Key Features
1. 🚀 Memory-First Architecture
- Clamped V8 Heap: Dynamic heap ceiling clamped to 512 MB (instead of 35% of total system RAM).
- Active Idle Garbage Collector: Background watcher triggers
global.gc()every 60 seconds if memory exceeds 350 MB during idle periods. - Lean SQLite Storage: SQLite cache size clamped from 64 MB down to 2 MB (
cache_size = -2048) with 16 MB memory-mapped I/O. - Rolling Call Logs:
call_logstable automatically caps itself at the latest 500 records, preventing database bloat.
2. 🌐 356+ Providers & Multi-Account Fallback
- Support for major frontier providers (Anthropic Claude, OpenAI, Google Gemini, DeepSeek, xAI Grok, Moonshot Kimi, Mistral) and local models (Ollama, vLLM, LM Studio).
- Account Pools: Add multiple accounts for the same provider; OmliteRoute rotates keys, balances quota, and auto-cools rate-limited accounts.
- Combos: Chain multiple models into a single virtual identifier (e.g.
fast-coder-> Gemini 2.5 Flash -> DeepSeek V3 -> Claude 3.5 Sonnet).
3. 🗜️ In-Flight Token Compression
- RTK Compression: Prunes tool outputs, terminal logs, and redundant JSON whitespace in agentic loops.
- Caveman Compaction: Semantic prompt compaction preserving core reasoning directives while dropping filler tokens.
4. 📱 Clean Haute Luxury Dashboard & Telegram WebApp (TWA)
- Minimalist matte dark design (
#130e1b, cards#181324, accents emerald#059669). - Zero WebGL Canvas: Home provider topology replaced with a responsive, instant-loading CSS card grid with status pills (
READY,ROUTING,RECENT,ERR). - Telegram-Ready: Pre-configured CSP and iframe headers allowing the dashboard to open directly as a Telegram WebApp (TWA) without "Internal Server Error" blocks.
📸 Dashboard & Feature Preview
🌱 Eco & Real-Time Financial Savings (MyRoute Standard)
OmliteRoute v1.0.1 integrates real-time environmental and financial telemetry directly into the Costs and Analytics dashboard, grounded in empirical data center cooling benchmarks:
- 💰 Biaya yang Dihemat (Cost Saved): Real-time financial ROI calculated against blended frontier model direct API rates ($8.50 / 1M tokens baseline). Track exactly how many dollars your free provider accounts and combos save.
- 💧 Liter Air Terhemat (Cooling Water Saved): AI data centers consume vast volumes of clean water for cooling high-density GPU clusters (~0.5L evaporated per 20–50 frontier queries). OmliteRoute tracks the literal liters of cooling water conserved via smart routing and caching (e.g.
14,700+ Liters = 29,000+ water bottles). - 🌿 Jejak Karbon Dicegah (Carbon Offset): Quantifies CO₂e emissions prevented by routing through idle capacity, local inference, and cached context windows.
- ⚡ 100% Free Routing Efficiency: Instant visibility into free vs paid token ratios so you always know your effective cost-per-token is optimized.
🧭 Complete 3-Tier Developer Navigation
Unlike bloated gateways that bury critical tools behind nested accordion folders, OmliteRoute organizes all developer utilities into 3 clear, high-density tiers:
- Homepage:
- Endpoint & Key (
/dashboard/endpoint): Base URL endpoints, tunnels, and root API access keys. - Overview (
/home): Executive telemetry strip (Active Providers, Model Catalog, RAM Footprint). - Playground (
/dashboard/playground): Test model outputs, streaming speeds, and multi-provider responses.
- Endpoint & Key (
- Gateway:
- Providers (
/dashboard/providers): Connect accounts, OAuth tokens, and API credentials across 356+ platforms. - Combo & Vision Adapter (
/dashboard/combos): Configure multi-model fallback chains, load-balancing, and multimodal bridges. - Token Saver (
/dashboard/context/settings): RTK & Caveman prompt compression rules. - CLI Tools & Orchestration: Direct command runners, Conductor agents, and task councils.
- Skills & Memory: Persistent agentic memory and tool definitions.
- MCP Server (
/dashboard/mcp): Model Context Protocol integrations (stdio, SSE, HTTP).
- Providers (
- Observe & Tools:
- Usage & Costs (
/dashboard/costs): Real-time spend, cost savings, and eco-metrics breakdown. - Quota Tracker (
/dashboard/quota): Monitor daily limits, tier resets, and account usage. - Health & Resilience (
/dashboard/health,/dashboard/resilience/connections): Live circuit breaker and latency monitoring. - Console & Audit Logs (
/dashboard/logs): High-speed rolling call logs and event inspectors. - Format Translator (
/dashboard/translator): Real-time OpenAI ↔ Claude ↔ Gemini wire-format conversion.
- Usage & Costs (
✂️ What Was Pruned to Make it "Lite"?
OmliteRoute discards all non-essential features that cause memory leaks and CPU thrashing:
- Gamification Bypassed: Zero XP calculations, level-ups, or streak writes on request hot-paths.
- 15+ Background Pollers Disabled: No Chatbot Arena ELO sync, live pricing scrapers, models.dev pollers, or live WebSocket daemons waking the CPU.
- Monaco Editor Dropped: Replaced the ~20 MB bundled VS Code editor with a lightweight dark-luxury monospace editor.
- Stripped 40 Foreign Languages: Removed 35.5 MB of unneeded JSON dictionaries, keeping clean English (
en) and Indonesian (id). - Pruned 53 Sidebar Items: Eliminated unbuilt, dead, or bloated menu items, leaving a tight 7-section core navigation.
🚀 Quick Start
1. Installation
Method A: Global via npm (Recommended)
npm install -g omliteroute
omliterouteOr run without installing: npx omliteroute
Method B: 1-Line Universal Script (Linux & macOS)
curl -fsSL https://raw.githubusercontent.com/adamhasani/omliteroute/main/install.sh | bashMethod C: Git Clone & Run
git clone https://github.com/adamhasani/omliteroute.git
cd omliteroute
npm install
./bin/omliteroute.mjsMethod D: Docker Compose (Alpine ~95MB)
docker compose -f docker-compose.lite.yml up -d2. CLI Command Reference
| Command | Description |
| :--- | :--- |
| omliteroute | Starts OmliteRoute Gateway in Lite mode (port 20128) |
| omliteroute top | Opens real-time terminal monitor (RAM, health, models, latency) |
| omliteroute --port <number> | Runs gateway on a custom port |
| omliteroute-reset-password | Resets dashboard admin password directly via terminal |
| npx omliteroute | Runs OmliteRoute on-demand without global installation |
3. Terminal Live Monitoring (omliteroute top)
Launch an interactive terminal monitor anytime without opening the browser:
omliteroute top===============================================================================
🪶 OmliteRoute — Live Terminal Monitor [Press Ctrl+C to exit]
===============================================================================
Gateway: ● ONLINE (http://127.0.0.1:20128)
Health Ping: 10 ms
Engine Mode: Omni Lite ⚡ (V8 heap: 512MB max · Idle GC: Active)
-------------------------------------------------------------------------------
📊 Operational Metrics
-------------------------------------------------------------------------------
Total AI Models: 953
SQLite Page Cache: 2 MB (clamped from 64MB)
Call Logs Rolling: Active (auto-purged to latest 500 records)
Background Sched: Pruned (15+ idle cron loops stopped)
-------------------------------------------------------------------------------
⚡ Core Endpoints Live Verification
-------------------------------------------------------------------------------
✔ GET /api/health 200 OK (Zero-overhead probe)
✔ GET /v1/models 200 OK (Dynamic catalog)
✔ POST /v1/chat/completions Ready (OpenAI / Claude stream)
===============================================================================The web dashboard is live at:
👉 http://localhost:20128
Default initial login password: CHANGEME (or configure via INITIAL_PASSWORD).
💻 Connecting Your Tools
OmliteRoute is 100% drop-in compatible with standard OpenAI and Anthropic SDKs.
cURL
curl http://localhost:20128/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "auto",
"messages": [{"role": "user", "content": "Hello world!"}]
}'Python (OpenAI SDK)
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:20128/v1",
api_key="YOUR_API_KEY"
)
response = client.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Write a Python script to sort a list."}]
)
print(response.choices[0].message.content)Node.js / TypeScript (OpenAI SDK)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "http://localhost:20128/v1",
apiKey: "YOUR_API_KEY",
});
const response = await client.chat.completions.create({
model: "auto",
messages: [{ role: "user", content: "Explain quantum computing simply." }],
});
console.log(response.choices[0].message.content);Claude Code CLI
export ANTHROPIC_BASE_URL="http://localhost:20128"
export ANTHROPIC_API_KEY="YOUR_API_KEY"
claudeCursor / Cline / Roo Code / OpenCode
Set the OpenAI Base URL in your editor settings:
- Base URL:
http://localhost:20128/v1 - API Key:
sk-omlite-...(generated in OmliteRoute dashboard -> API Keys) - Model:
auto(or choose any combo/provider model ID)
⚙️ Environment Variables
| Variable | Description | Default |
| :--- | :--- | :---: |
| PORT | Web dashboard & API port | 20128 |
| HOSTNAME | Host address to bind | 0.0.0.0 |
| OMNIROUTE_LITE | Enables memory-optimized Lite mode | 1 |
| OMNIROUTE_MEMORY_MB | V8 heap ceiling in MB | 512 |
| INITIAL_PASSWORD | Default dashboard password | CHANGEME |
| DATA_DIR | Directory for SQLite database | ~/.omniroute |
🐛 Reporting Bugs & Contributing
Found a bug, want a new provider, or have an idea to optimize OmliteRoute further?
- Open an Issue: github.com/adamhasani/omliteroute/issues
- Report a Bug: Submit a Bug Report
- Request a Feature: Submit a Feature Request
- Security Vulnerabilities: If you discover a sensitive security vulnerability, please submit it privately via GitHub Security Advisories.
When reporting bugs, running omliteroute doctor and attaching the sanitized diagnostic output helps reproduce and patch issues promptly!
📜 License & Acknowledgments
- Creator & Maintainer: Adam Hasani (@adamhasani)
- Package Registry: npmjs.com/package/omliteroute
- GitHub Repository: github.com/adamhasani/omliteroute
- License: MIT License
⭐ Star this repository if OmliteRoute saved your laptop RAM and money!
