bansos-router
v0.3.1
Published
Bansos for your AI agents. Free, keyless coding models routed to every coding harness through one local endpoint. No accounts, no keys, no cost.
Maintainers
Readme
bansos-router
Free, keyless coding models for every coding harness, through one local daemon. Works without accounts or API keys.
Quick start
npm i -g bansos-routerStart the daemon and point your coding harness at it:
bansos start --bg # detached daemon on 127.0.0.1:17070
bansos setup opencode # configure any supported harnessFor pi, install the companion extension (extensions/pi):
pi install npm:pi-bansos-routerSupported harnesses for bansos setup: claude-code, aider, opencode, codex, hermes, goose, openclaw, antigravity, jcode, 9router, continue, cline, and roo. See docs/harnesses.md for harness-specific details.
Open http://127.0.0.1:17070 in your browser for the Web UI dashboard (model catalog, live latency ping, 1-click config generator, relay manager, playground).
Always-on local protections
These hold in every mode (not just strict):
- CORS only for loopback origins: a page served from a non-loopback origin gets no CORS headers, so it can neither read nor write the daemon.
- Host header validated: loopback clients must present a loopback (or
own-interface) Host header; anything else gets 403. This blocks DNS
rebinding where a hostile page resolves to
127.0.0.1. - Relay probe & mutation are SSRF-guarded: probes may only target the
active/saved relay or a public
https://host; literal loopback, private, and link-local IPs (e.g.169.254.xmetadata) are rejected unless explicitly saved. Relay URLs must behttp(s)://without credentials.
Docker
Run the daemon in a container instead of installing globally:
docker build -t bansos-router .
docker run -d --name bansos -p 17070:17070 -v bansos-data:/home/node/.bansos bansos-routerThe container runs as the non-root node user (uid 1000); state lives in
/home/node/.bansos inside the container, which is what the volume above
mounts.
Or with compose:
docker compose up -d- Image is a multi-stage build on node:22-alpine (~58 MB compressed), CLI + daemon + Web UI included, zero runtime dependencies.
- The container runs the daemon in the foreground (Docker owns the lifecycle)
bound to
0.0.0.0, with tini for correct signal forwarding and a healthcheck on/healthz. - State persists across restarts in the
bansos-datavolume; point your harnesses athttp://127.0.0.1:17070as usual. docker compose logs -freplacesbansos logsinside the container.
CLI reference
| Command | Purpose |
|---|---|
| bansos start [--bg] [--port N] [--bind H] [--unsafe-allow-non-loopback] | Run the daemon (foreground, or detached with --bg) |
| bansos logs [--activity] | Tail the daemon log in real time (every mode: daemon always logs to ~/.bansos/logs/bansosd.log); --activity prints the structured request feed shown in the web UI "Activity" tab |
| bansos stop [--all] | Stop the daemon recorded in state.json; --all stops every bansos daemon on the machine |
| bansos status [--json] | Daemon status (port, model count, alive models); reports every running daemon on the auto-bump range (17070-17090) |
| bansos models [--json] | List live catalog from /v1/models |
| bansos ping [model] [--json] | Probe live latency and rate-limit status of all models (or a specific model) |
| bansos refresh | Ask the daemon to re-run health checks now |
| bansos setup <harness...> [--model <id>] [--dry-run] [--undo] | Write, update, or undo harness configs |
| bansos relay <on\|off\|status\|url\|use\|list\|remove> | Manage relay egress |
| bansos doctor | Diagnose daemon reachability and harness config validity |
| bansos --version | Print version |
| bansos <command> --help / bansos help <command> | Per-command usage, defaults, exit codes, examples |
| bansosd | Alias for the daemon (e.g. bansosd --bg) |
Strict security mode
Strict mode is opt-in and fail-closed. Add a security block to
~/.bansos/config.json and explicitly list every provider that may receive
requests:
{
"port": 17070,
"bind": "127.0.0.1",
"security": {
"mode": "strict",
"allowedUpstreams": ["zen"],
"allowCrossProviderFailover": false
}
}With mode: "strict":
- Non-loopback binds rejected unless
--unsafe-allow-non-loopbackpassed. - Relay egress, probing, and mutation disabled across CLI, API, Web UI.
- Only exact upstreams in
allowedUpstreamsreceive requests (empty list blocks all). - Cross-provider failover disabled.
- Sensitive values and raw upstream errors suppressed from logs.
- DLP blocks pasted credentials: API keys, PATs, private keys, and credential assignments pasted into a request are blocked with HTTP 422 before they leave for an external upstream (zen/kilo/llm7). Loopback (self-hosted) gateways are exempt, since their own keys are legitimately part of the request body.
What it does
- One local daemon on
127.0.0.1:17070, started withbansos start. It speaks three wire protocols: OpenAI Chat Completions, Anthropic Messages, and OpenAI Responses (Codex CLI). - Keyless free upstreams only: OpenCode Zen, KiloCode gateway, and LLM7.
bansos setup <harness>writes config for Claude Code, Aider, OpenCode, Codex, Hermes, Goose, OpenClaw, Antigravity, JCode, 9Router, Continue, Cline, and Roo Code.- Web UI dashboard served at
http://127.0.0.1:17070/with model catalog explorer, live ping probes, 1-click harness setup generator, relay egress manager, and test playground. The catalog shows per-model capabilities (Think / Vision badges) and a live usage + activity tracker (requests, tokens, latency, error rate). - The pi extension (
pi-bansos-router) registers thebansosrprovider, so every free model shows up in pi's/modelpicker. There is also a/bansosrcommand for status. If the daemon is not running when pi starts, the extension starts it and stops it again on exit. - The catalog is health-checked: live
:freemodels replace stale seeds on a timer. Relay egress can route around rate limits, and there is a per-IP rate limiter. - If a model is rejected with
401/403/429/5xx, the daemon auto-fails over to the closest equivalent model on a different upstream (same reasoning level, context window, and effort capability), retrying up to two extra candidates before surfacing an error. Request duration (durationMs) is logged on every completion and rejection. - A model that answers
429is parked for a cooldown, so the next request starts on a fallback instead of spending another round trip to learn the same thing.Retry-Aftersets the duration when the upstream sends one, otherwise it is a minute, capped at fifteen. Parked models stay listed in/v1/modelsand are still used when nothing else qualifies. - A
401/403means the upstream will not serve that model from where the request egressed (OpenCode Zen, for instance, returns403 This service is not available in your regionwhen Cloudflare routes through a region its providers block). Retrying cannot help, so the model is parked for 30 minutes and the request fails over instead of returning the error. Sendx-bansos-no-failover: 1when you want the model's own status back. The upstream's own message is now logged alongside the status.
Available models
By default, bansos setup automatically configures intelligent defaults per harness:
- Claude Code: Maps tiers automatically (
haiku-> fast non-reasoning,sonnet-> daily reasoning,opus-> highest-capacity reasoning). - Multi-model harnesses (
opencode,goose,openclaw,continue,9router,jcode): Registers all available models (or dynamic/v1/modelsprovider) with the smart default (currentlymuse-spark-1.3-contributor-free) as primary. - Single-model harnesses (
aider,codex,hermes,antigravity,cline,roo): Same smart default. Pass--model <id>to pin a specific model. Context and max output are token counts. The seeded catalog is ~30 models and the live one changes as upstreams rotate free tiers; runbansos modelsorbansos pingto see what is alive right now.
OpenCode Zen
| model id | reasoning | vision | context | max output |
|---|---|---|---|---|
| mimo-v2.5-free | ✓ | ✓ | 200k | 32k |
| nemotron-3-ultra-free | ✓ | ✗ | 1M | 128k |
| big-pickle | ✓ | ✗ | 200k | 32k |
| nemotron-3.5-lightning-free | ✓ | ✗ | 262k | 262k |
| ling-3.0-flash-fin-free | ✓ | ✗ | 262k | 32k |
| muse-spark-1.3-contributor-free | ✓ | ✓ | 1M | 131k |
| muse-spark-1.2-contributor-free | ✓ | ✓ | 1M | 131k |
Muse Spark answers only on the Responses API (/v1/responses) and returns HTTP
500 on chat completions. bansos-router translates the upstream leg both ways, so
Muse works from any harness — chat completions, Responses, and Anthropic clients
all see their usual wire format, tool calls included.
KiloCode gateway
| model id | reasoning | vision | context | max output |
|---|---|---|---|---|
| kilo-auto/free | ✓ | ✗ | 256k | 10k |
| stepfun/step-3.7-flash:free | ✓ | ✓ | 262k | 262k |
| nvidia/nemotron-3-ultra-550b-a55b:free | ✓ | ✗ | 1M | 65k |
| nvidia/nemotron-3-super-120b-a12b:free | ✓ | ✗ | 262k | 262k |
| nvidia/nemotron-3.5-lightning:free | ✓ | ✗ | 1M | 65k |
| nvidia/nemotron-3.5-content-safety:free | ✓ | ✓ | 128k | 8k |
| liquid/lfm-2.5-2.6b:free | ✓ | ✗ | 65k | 8k |
| poolside/laguna-s-2.1:free | ✓ | ✗ | 262k | 32k |
| cohere/north-mini-code:free | ✓ | ✗ | 256k | 64k |
| poolside/laguna-xs-2.1:free | ✓ | ✗ | 262k | 32k |
| nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | ✓ | ✓ | 256k | 65k |
| inclusionai/ling-3.0-flash-sante:free | ✓ | ✗ | 262k | 32k |
| inclusionai/ling-3.0-flash-fin:free | ✓ | ✗ | 262k | 32k |
| dots-studio/dots-3-note-preview:free | ✓ | ✓ | 512k | 65k |
| thinkingmachines/inkling:free | ✓ | ✓ | 1M | 65k |
| thinkingmachines/inkling-small:free | ✓ | ✓ | 1M | 65k |
| openrouter/free | ✓ | ✓ | 200k | 65k |
LLM7
| model id | reasoning | vision | context | max output |
|---|---|---|---|---|
| codestral-latest | ✗ | ✗ | 32k | 8k |
| minimax-m2.7 | ✓ | ✗ | 180k | 32k |
| mistral-Nemo-Instruct-2407 | ✗ | ✗ | 128k | 16k |
| default | ✗ | ✗ | 128k | 8k |
| fast | ✗ | ✗ | 128k | 8k |
Notes:
- The tables above mirror the seeds in
src/upstreams/*, which were verified against each upstream's live catalog and specs on 2026-09-03 (reasoning / vision / context flags, dead promos removed). On refresh the daemon serves the live catalog, which can differ slightly as upstreams rotate free tiers. - Kilo's
:freemodels are liveness-gated: the daemon re-checks the kilo catalog on a timer and drops models that are no longer offered free. - Zen seeds are re-checked against
GET /v1/modelson refresh. The listing is not exhaustive (some servable promos are never listed), so an unlisted seed is probe-verified keyless before being kept; a retired promo (e.g.hy3-free, gone and answering 401) disappears automatically. - LLM7 models are filtered dynamically by
usage_based_only: false(tier turbo);defaultandfastare stable selectors.protier is paid-only and excluded.
Development
For contributors working from source:
git clone https://github.com/ihsan-ramadhan/bansos-router
cd bansos-router
npm install
npm run build # builds Web UI (Vite) + CLI (esbuild), outputs dist/
npm link # make `bansos`/`bansosd` available globally
npm run typecheck # tsc --noEmit
npm test # node:test
npm run dev # run the bansos CLI from source
npm run dev:daemon # run the daemon from source
npm run dev:ui # run Vite dev server for Web UIDocs
License
MIT
