npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@prismd/prismd

v0.0.23

Published

Local-first LLM gateway aggregating free/low-quota model APIs for coding agents

Readme

prismd

English | 简体中文 | 日本語 | 한국어 | Deutsch | Français | Español | Italiano | العربية | Türkçe

npm version npm downloads License

Local-first High-Availability LLM Gateway aggregating free and low-cost model APIs (OpenRouter, Groq, Cerebras, Google Gemini, NVIDIA NIM, GitHub Models, etc.) and local LLMs (Ollama), providing a stable, unified interface with automatic failover and routing for coding agents (Claude Code, Codex CLI, Cursor, OpenCode, Aider, etc.).

┌────────────────────────────────┐       ┌─────────────────────────────────────┐       ┌─────────────────────────────────────┐
│    Coding Agents (Clients)     │       │        prismd Gateway (Local)       │       │         Model Providers (Upstream)  │
│                                │       │          127.0.0.1:8787             │       │                                     │
│  Claude Code  (Messages API)   ├──────►│  [Protocol Converter]               ├──────►│  Cloud Free APIs                    │
│  Codex CLI    (Responses API)  ├──────►│    • Messages ↔ Responses ↔ Chat    │       │    • OpenRouter / Groq / Cerebras   │
│  Cursor / dsh (Chat API)       ├──────►│  [Smart Router (free-auto)]         │       │    • Google Gemini / NVIDIA NIM     │
│  OpenCode / Pi / Aider         ├──────►│    • Quota-Weighted & Context Check │       │    • GitHub Models / AMD            │
│                                │       │  [Key Pool & Circuit Breaker]       │       │                                     │
│                                │       │    • Multi-Key Round-Robin / 429    │  all  │  Local Offline Fallback             │
│                                │       │    • Zero-Downtime Auto Fallback    ├──────►│    • Ollama (qwen2.5-coder / r1)    │
│                                │       │                                     │  429  │    • LM Studio (local GGUF models)  │
└────────────────────────────────┘       └─────────────────────────────────────┘       └─────────────────────────────────────┘

Key Highlights

  1. Unified Model Alias (free-auto): Connect using a single alias; prismd automatically selects the best available free model.
  2. Multi-Key Pooling & Single-Key Circuit Breaking: Hit rate limits? Configure multiple keys per provider for automatic round-robin scheduling. When a single key hits 429, only that key is cooled down while traffic seamlessly shifts to the next key.
  3. Opt-in Local Fallback (Ollama / LM Studio): Default aliases are cloud-only. Running a local backend? Append it to an alias queue via config.user.json — when cloud free models hit 429 or internet drops, requests then fall back to your local models, ensuring coding agent tasks never crash.
  4. Transparent Multi-Protocol Conversion: Full bi-directional streaming conversion across Claude Code (Messages), Codex (Responses), and Cursor/OpenCode (Chat Completions).
  5. Embedded Web Dashboard & Hot Reloading: Open http://127.0.0.1:8787/ui to monitor live candidate health and quota progress bars. Update configurations dynamically via SIGHUP without restarting.

Support

If prismd saves you time or API costs, consider buying the author a coffee:

ko-fi


3-Step Quick Start

Step 1: Install and Start Gateway

# Option A: Global npm install
npm install -g @prismd/prismd              # Stable release
# Or install RC preview channel (aligned with latest develop):
npm install -g @agentscraft/prismd         # RC channel

# Option B: Run from source
git clone https://github.com/AgentsCraft/prismd.git
cd prismd && npm install

Step 2: Initialize & Configure (Interactive Wizard)

Run the interactive setup wizard to configure your keys and clients in seconds:

prismd init

The wizard will:

  1. Set your local protection token (prismd:).
  2. Prompt you to select free providers (OpenRouter, Groq, Google Gemini, Cerebras, etc.) and enter their API keys.
  3. Automatically configure your coding clients (Claude Code, Codex CLI, OpenCode, Pi Agent) with safe timestamped backups (.bak.<timestamp>) for existing configs!

(Prefer manual configuration? You can manually edit ~/.prismd/keys.yaml or set PRISMD_HOME to override the directory).

# Manual setup example: ~/.prismd/keys.yaml (recommended chmod 600)
prismd: "my-local-secret"       # Local protection token (used by clients)

# Cloud Providers (supports single key or multi-key pool for round-robin):
openrouter: "sk-or-v1-xxxx"
groq:
  - "gsk_key1_xxxx"             # Multi-key pooling & cooldown isolation
  - "gsk_key2_xxxx"
cerebras: ["csk_1_xxxx", "csk_2_xxxx"]
gemini: "AIzaSyxxxx"
nvidia: "nvapi-xxxx"
github: "ghp_xxxx"              # GitHub Models personal token
amd: "amd_token_xxxx"           # Optional: AMD Developer Cloud

# Local Offline Fallback:
# ollama: zero-config (automatically routes to http://127.0.0.1:11434/v1 without key)

Start the gateway:

prismd
# Or from source: npm run generate:config && npm run dev

📖 Provider Setup Guides: See Model Provider Integration Guides for detailed key generation and model lists for OpenRouter, Groq, Cerebras, Google Gemini, NVIDIA NIM, GitHub Models, AMD, Ollama, and LM Studio.

Step 3: Configure Your Agent

| Client | Quick Start (prismd init auto-config) | Guide | |---|---|---| | Claude Code | claude (configured in ~/.claude/settings.json) | Guide | | Codex CLI | codex (configured in ~/.codex/config.toml & auth.json) | Guide | | OpenCode | opencode (configured in ~/.config/opencode/opencode.json) | Guide | | Pi Agent | pi (configured in ~/.pi/config.json) | Guide | | Cursor | Settings → Models → Enable OpenAI API Key (my-local-secret)Override OpenAI Base URL: http://127.0.0.1:8787/v1Add model: free-auto | Guide | | DeepSeek Harness (dsh) | Set base_url = "http://127.0.0.1:8787/v1" in ~/.dsh/config.tomlPRISMD_API_KEY=my-local-secret dsh --model prismd:free-auto | Guide | | Aider | OPENAI_API_BASE="http://127.0.0.1:8787/v1" OPENAI_API_KEY="my-local-secret" aider --model openai/free-auto | Guide |

📖 Full documentation: See Client Integration Guide for detailed protocol breakdowns and advanced setups.


Features In Depth

1. Smart Routing & Automated Failover

prismd dynamically selects the optimal model candidate per request using an intelligent evaluation pipeline:

  • Context Window Verification: Estimates input tokens before dispatch; automatically filters out candidates whose context window is too small, preventing 400 Context Overflow errors.
  • Quota-Weighted Soft Limits: When a cloud candidate reaches 80% of its daily quota (quotaSoftLimitRatio), it is automatically demoted to the tail of the queue, reserving remaining quota for peak requirements.
  • Zero-Crash Failover: If an upstream provider returns a 429 rate limit or 5xx outage, prismd transparently fails over to the next healthy candidate in the alias queue without failing the client's session.
  • Default Alias: free-auto: the single unified free queue. Prioritizes Gemini 2.0 Flash / Llama 3.3 70B. Cloud-only by default.

2. Multi-Key Pooling & Single-Key Circuit Breaking (Key Pool)

All cloud providers (Groq, Cerebras, Google Gemini, OpenRouter, NVIDIA NIM, GitHub Models, etc.) support multi-key configurations for automatic round-robin request distribution and single-key fault isolation:

  • ~/.prismd/keys.yaml format (YAML list or inline array):
    groq:
      - "gsk_key1_xxxx"
      - "gsk_key2_xxxx"
    cerebras: ["csk_1_xxxx", "csk_2_xxxx"]
    gemini:
      - "AIzaSy_key1_xxxx"
      - "AIzaSy_key2_xxxx"
  • .env or Environment Variables (comma-separated):
    GROQ_API_KEY="gsk_key1,gsk_key2,gsk_key3"
    GEMINI_API_KEY="AIzaSy1,AIzaSy2"
  • How it works: Requests are distributed across healthy keys via Round-Robin. When a key (e.g. gsk_key1) receives a 429 rate limit error, only that key enters cooldown (respecting Retry-After), while subsequent requests immediately shift to the next available key (gsk_key2) or candidate model, multiplying throughput without failing requests.

3. Local LLM Fallback (Ollama & LM Studio, Opt-in)

prismd ships Ollama and LM Studio as built-in providers, but default aliases stay cloud-only — machines without a local backend get no dead candidates. Running one locally? Append it to an alias queue via config.user.json (candidate arrays replace the preset list, so keep the cloud entries you want):

{
  "aliases": {
    "free-auto": {
      "candidates": ["gemini-2.0-flash", "cohere/north-mini-code:free", "qwen2.5-coder:7b"]
    }
  }
}
  • Ollama: Built-in zero-config provider (http://127.0.0.1:11434/v1):
    ollama run qwen2.5-coder:7b
  • LM Studio: Supports local OpenAI-compatible server (http://127.0.0.1:1234/v1) running GGUF models. See LM Studio Guide.
  • With a local candidate at the tail of the queue, requests fall back to it when cloud models are exhausted so coding agent tasks never crash midway.

4. Transparent Multi-Protocol Bridge

Full bi-directional streaming conversion across three major agent wire protocols:

  • Anthropic Messages (POST /v1/messages): Full support for Claude Code (tools, thinking blocks, SSE streams).
  • OpenAI Responses (POST /v1/responses): Compatible with Codex CLI and DeepSeek Harness (dsh).
  • OpenAI Chat Completions (POST /v1/chat/completions): Standard interface for Cursor, OpenCode, Pi Agent, and Aider.

5. Extensible Configuration (config.user.json)

Customize providers, register private models, or define custom model queues in config.user.json:

{
  "models": {
    "my-custom-model": {
      "provider": "openrouter",
      "contextWindow": 131072,
      "maxOutputTokens": 8192,
      "supportsTools": true,
      "supportsReasoning": false,
      "limits": { "dailyRequests": 100, "rpm": 20, "maxConcurrent": 2 }
    }
  },
  "aliases": {
    "free-auto": {
      "candidates": ["my-custom-model", "gemini-2.0-flash", "qwen2.5-coder:7b"]
    }
  }
}

Re-compile configuration with prismd generate (or npm run generate:config in source mode).

6. Dynamic Config Hot Reloading (SIGHUP)

Update routing tables, keys, or aliases without restarting the process or interrupting active streaming connections:

kill -HUP $(pgrep -f "prismd")

Status & Observability

  • Web Dashboard: Open http://127.0.0.1:8787/ui in your browser:
    • Real-time candidate health (healthy / rate_limited / cooldown)
    • Daily quota progress bars and token usage statistics
    • Client usage table (last 24h): per-client × endpoint request count, success rate, P50 latency, failover count, and recent errors
    • 10-language UI selector and "Reset usage" button
  • CLI Status & Commands:
    prismd status      # Display metrics table + client usage section (when gateway is live)
    prismd generate    # Recompile ~/.prismd/prismd.json
  • API: GET /v1/clientstatus — read-only JSON snapshot of the client usage window (unauthenticated, loopback-only).

Troubleshooting

  • Q: missing API key for provider error?
    • Verify keys in ~/.prismd/keys.yaml or .env, then run prismd generate (or npm run generate:config in source mode).
  • Q: Frequent 429s on free models?
    • Add multiple keys for the provider, or append a local Ollama candidate to the alias queue (see Local LLM Fallback).
  • Q: Reset daily quota counters?
    • Click "Reset usage" in the Web Dashboard (http://127.0.0.1:8787/ui) or delete data/prismd.sqlite.
  • Q: Warning about unknown config key on startup?
    • Unknown top-level keys in config.user.json are tolerated with a warning and ignored. This is safe — the key may come from a newer version or a typo. Remove or correct the key to suppress the warning.