@prismd/prismd
v0.0.23
Published
Local-first LLM gateway aggregating free/low-quota model APIs for coding agents
Maintainers
Readme
prismd
English | 简体中文 | 日本語 | 한국어 | Deutsch | Français | Español | Italiano | العربية | Türkçe
Local-first High-Availability LLM Gateway aggregating free and low-cost model APIs (OpenRouter, Groq, Cerebras, Google Gemini, NVIDIA NIM, GitHub Models, etc.) and local LLMs (Ollama), providing a stable, unified interface with automatic failover and routing for coding agents (Claude Code, Codex CLI, Cursor, OpenCode, Aider, etc.).
┌────────────────────────────────┐ ┌─────────────────────────────────────┐ ┌─────────────────────────────────────┐
│ Coding Agents (Clients) │ │ prismd Gateway (Local) │ │ Model Providers (Upstream) │
│ │ │ 127.0.0.1:8787 │ │ │
│ Claude Code (Messages API) ├──────►│ [Protocol Converter] ├──────►│ Cloud Free APIs │
│ Codex CLI (Responses API) ├──────►│ • Messages ↔ Responses ↔ Chat │ │ • OpenRouter / Groq / Cerebras │
│ Cursor / dsh (Chat API) ├──────►│ [Smart Router (free-auto)] │ │ • Google Gemini / NVIDIA NIM │
│ OpenCode / Pi / Aider ├──────►│ • Quota-Weighted & Context Check │ │ • GitHub Models / AMD │
│ │ │ [Key Pool & Circuit Breaker] │ │ │
│ │ │ • Multi-Key Round-Robin / 429 │ all │ Local Offline Fallback │
│ │ │ • Zero-Downtime Auto Fallback ├──────►│ • Ollama (qwen2.5-coder / r1) │
│ │ │ │ 429 │ • LM Studio (local GGUF models) │
└────────────────────────────────┘ └─────────────────────────────────────┘ └─────────────────────────────────────┘Key Highlights
- Unified Model Alias (
free-auto): Connect using a single alias; prismd automatically selects the best available free model. - Multi-Key Pooling & Single-Key Circuit Breaking: Hit rate limits? Configure multiple keys per provider for automatic round-robin scheduling. When a single key hits 429, only that key is cooled down while traffic seamlessly shifts to the next key.
- Opt-in Local Fallback (Ollama / LM Studio): Default aliases are cloud-only. Running a local backend? Append it to an alias queue via
config.user.json— when cloud free models hit 429 or internet drops, requests then fall back to your local models, ensuring coding agent tasks never crash. - Transparent Multi-Protocol Conversion: Full bi-directional streaming conversion across Claude Code (Messages), Codex (Responses), and Cursor/OpenCode (Chat Completions).
- Embedded Web Dashboard & Hot Reloading: Open
http://127.0.0.1:8787/uito monitor live candidate health and quota progress bars. Update configurations dynamically viaSIGHUPwithout restarting.
Support
If prismd saves you time or API costs, consider buying the author a coffee:
3-Step Quick Start
Step 1: Install and Start Gateway
# Option A: Global npm install
npm install -g @prismd/prismd # Stable release
# Or install RC preview channel (aligned with latest develop):
npm install -g @agentscraft/prismd # RC channel
# Option B: Run from source
git clone https://github.com/AgentsCraft/prismd.git
cd prismd && npm installStep 2: Initialize & Configure (Interactive Wizard)
Run the interactive setup wizard to configure your keys and clients in seconds:
prismd initThe wizard will:
- Set your local protection token (
prismd:). - Prompt you to select free providers (OpenRouter, Groq, Google Gemini, Cerebras, etc.) and enter their API keys.
- Automatically configure your coding clients (Claude Code, Codex CLI, OpenCode, Pi Agent) with safe timestamped backups (
.bak.<timestamp>) for existing configs!
(Prefer manual configuration? You can manually edit ~/.prismd/keys.yaml or set PRISMD_HOME to override the directory).
# Manual setup example: ~/.prismd/keys.yaml (recommended chmod 600)
prismd: "my-local-secret" # Local protection token (used by clients)
# Cloud Providers (supports single key or multi-key pool for round-robin):
openrouter: "sk-or-v1-xxxx"
groq:
- "gsk_key1_xxxx" # Multi-key pooling & cooldown isolation
- "gsk_key2_xxxx"
cerebras: ["csk_1_xxxx", "csk_2_xxxx"]
gemini: "AIzaSyxxxx"
nvidia: "nvapi-xxxx"
github: "ghp_xxxx" # GitHub Models personal token
amd: "amd_token_xxxx" # Optional: AMD Developer Cloud
# Local Offline Fallback:
# ollama: zero-config (automatically routes to http://127.0.0.1:11434/v1 without key)Start the gateway:
prismd
# Or from source: npm run generate:config && npm run dev📖 Provider Setup Guides: See Model Provider Integration Guides for detailed key generation and model lists for OpenRouter, Groq, Cerebras, Google Gemini, NVIDIA NIM, GitHub Models, AMD, Ollama, and LM Studio.
Step 3: Configure Your Agent
| Client | Quick Start (prismd init auto-config) | Guide |
|---|---|---|
| Claude Code | claude (configured in ~/.claude/settings.json) | Guide |
| Codex CLI | codex (configured in ~/.codex/config.toml & auth.json) | Guide |
| OpenCode | opencode (configured in ~/.config/opencode/opencode.json) | Guide |
| Pi Agent | pi (configured in ~/.pi/config.json) | Guide |
| Cursor | Settings → Models → Enable OpenAI API Key (my-local-secret)Override OpenAI Base URL: http://127.0.0.1:8787/v1Add model: free-auto | Guide |
| DeepSeek Harness (dsh) | Set base_url = "http://127.0.0.1:8787/v1" in ~/.dsh/config.tomlPRISMD_API_KEY=my-local-secret dsh --model prismd:free-auto | Guide |
| Aider | OPENAI_API_BASE="http://127.0.0.1:8787/v1" OPENAI_API_KEY="my-local-secret" aider --model openai/free-auto | Guide |
📖 Full documentation: See Client Integration Guide for detailed protocol breakdowns and advanced setups.
Features In Depth
1. Smart Routing & Automated Failover
prismd dynamically selects the optimal model candidate per request using an intelligent evaluation pipeline:
- Context Window Verification: Estimates input tokens before dispatch; automatically filters out candidates whose context window is too small, preventing 400 Context Overflow errors.
- Quota-Weighted Soft Limits: When a cloud candidate reaches 80% of its daily quota (
quotaSoftLimitRatio), it is automatically demoted to the tail of the queue, reserving remaining quota for peak requirements. - Zero-Crash Failover: If an upstream provider returns a 429 rate limit or 5xx outage, prismd transparently fails over to the next healthy candidate in the alias queue without failing the client's session.
- Default Alias:
free-auto: the single unified free queue. Prioritizes Gemini 2.0 Flash / Llama 3.3 70B. Cloud-only by default.
2. Multi-Key Pooling & Single-Key Circuit Breaking (Key Pool)
All cloud providers (Groq, Cerebras, Google Gemini, OpenRouter, NVIDIA NIM, GitHub Models, etc.) support multi-key configurations for automatic round-robin request distribution and single-key fault isolation:
~/.prismd/keys.yamlformat (YAML list or inline array):groq: - "gsk_key1_xxxx" - "gsk_key2_xxxx" cerebras: ["csk_1_xxxx", "csk_2_xxxx"] gemini: - "AIzaSy_key1_xxxx" - "AIzaSy_key2_xxxx".envor Environment Variables (comma-separated):GROQ_API_KEY="gsk_key1,gsk_key2,gsk_key3" GEMINI_API_KEY="AIzaSy1,AIzaSy2"- How it works: Requests are distributed across healthy keys via Round-Robin. When a key (e.g.
gsk_key1) receives a 429 rate limit error, only that key enters cooldown (respectingRetry-After), while subsequent requests immediately shift to the next available key (gsk_key2) or candidate model, multiplying throughput without failing requests.
3. Local LLM Fallback (Ollama & LM Studio, Opt-in)
prismd ships Ollama and LM Studio as built-in providers, but default aliases stay cloud-only — machines without a local backend get no dead candidates. Running one locally? Append it to an alias queue via config.user.json (candidate arrays replace the preset list, so keep the cloud entries you want):
{
"aliases": {
"free-auto": {
"candidates": ["gemini-2.0-flash", "cohere/north-mini-code:free", "qwen2.5-coder:7b"]
}
}
}- Ollama: Built-in zero-config provider (
http://127.0.0.1:11434/v1):ollama run qwen2.5-coder:7b - LM Studio: Supports local OpenAI-compatible server (
http://127.0.0.1:1234/v1) running GGUF models. See LM Studio Guide. - With a local candidate at the tail of the queue, requests fall back to it when cloud models are exhausted so coding agent tasks never crash midway.
4. Transparent Multi-Protocol Bridge
Full bi-directional streaming conversion across three major agent wire protocols:
- Anthropic Messages (
POST /v1/messages): Full support for Claude Code (tools, thinking blocks, SSE streams). - OpenAI Responses (
POST /v1/responses): Compatible with Codex CLI and DeepSeek Harness (dsh). - OpenAI Chat Completions (
POST /v1/chat/completions): Standard interface for Cursor, OpenCode, Pi Agent, and Aider.
5. Extensible Configuration (config.user.json)
Customize providers, register private models, or define custom model queues in config.user.json:
{
"models": {
"my-custom-model": {
"provider": "openrouter",
"contextWindow": 131072,
"maxOutputTokens": 8192,
"supportsTools": true,
"supportsReasoning": false,
"limits": { "dailyRequests": 100, "rpm": 20, "maxConcurrent": 2 }
}
},
"aliases": {
"free-auto": {
"candidates": ["my-custom-model", "gemini-2.0-flash", "qwen2.5-coder:7b"]
}
}
}Re-compile configuration with prismd generate (or npm run generate:config in source mode).
6. Dynamic Config Hot Reloading (SIGHUP)
Update routing tables, keys, or aliases without restarting the process or interrupting active streaming connections:
kill -HUP $(pgrep -f "prismd")Status & Observability
- Web Dashboard: Open
http://127.0.0.1:8787/uiin your browser:- Real-time candidate health (
healthy/rate_limited/cooldown) - Daily quota progress bars and token usage statistics
- Client usage table (last 24h): per-client × endpoint request count, success rate, P50 latency, failover count, and recent errors
- 10-language UI selector and "Reset usage" button
- Real-time candidate health (
- CLI Status & Commands:
prismd status # Display metrics table + client usage section (when gateway is live) prismd generate # Recompile ~/.prismd/prismd.json - API:
GET /v1/clientstatus— read-only JSON snapshot of the client usage window (unauthenticated, loopback-only).
Troubleshooting
- Q:
missing API key for providererror?- Verify keys in
~/.prismd/keys.yamlor.env, then runprismd generate(ornpm run generate:configin source mode).
- Verify keys in
- Q: Frequent 429s on free models?
- Add multiple keys for the provider, or append a local Ollama candidate to the alias queue (see Local LLM Fallback).
- Q: Reset daily quota counters?
- Click "Reset usage" in the Web Dashboard (
http://127.0.0.1:8787/ui) or deletedata/prismd.sqlite.
- Click "Reset usage" in the Web Dashboard (
- Q: Warning about unknown config key on startup?
- Unknown top-level keys in
config.user.jsonare tolerated with a warning and ignored. This is safe — the key may come from a newer version or a typo. Remove or correct the key to suppress the warning.
- Unknown top-level keys in

