routewise
v0.2.0
Published
The cost-aware LLM request router, from your terminal. Auto-routes every query to the cheapest model tier that can handle it.
Maintainers
Readme
Zero runtime dependencies. Talks to the hosted RouteWise API (or your own backend).
npm install -g routewiseQuick start
# route one query — answer on stdout, routing metadata on stderr
routewise ask "what is the capital of france"
# force a tier, or change routing sensitivity (0 economy · 1 balanced · 2 quality)
routewise ask "design a distributed rate limiter" --tier frontier
routewise ask "explain a linked list" --threshold 2
# stream tokens as they arrive
routewise stream "write a haiku about routing"
# pipe queries from a shell workflow
cat questions.txt | routewise ask --json
# interactive multi-turn chat
routewise chat
# pick a difficulty policy: emma (generic), lisa (3-tier support), kate (2-tier support)
routewise ask "I want a refund" --model lisa
# list available models + their policies, cuts and evals
routewise models
# see who you are + how much you've spent
routewise whoami
# end-to-end self-check (config, backend, providers, live ask)
routewise doctor
# score difficulty + see what each routing mode picks (no key needed)
routewise evaluate "implement a red-black tree"Getting a key
# opens the dashboard + saves your key (paste is hidden, never in shell history)
routewise login
# or, non-interactive:
routewise login rw_your-keyWhat you get per call
$ routewise ask "what is the capital of france"
Paris
[cheap score=3.11 $0.00008 · 812ms] # <- stderr, so stdout stays pipeableMetadata on stderr keeps stdout clean for grep, pipelines, and > file.
Add --json for the full structured response (tier, score, cost, latency,
cache hit, model, route reason, request log id):
$ routewise ask "convert 5 miles to km" --json
{
"response": "5 miles is approximately 8.05 kilometers.",
"routed_to": "cheap",
"cache_hit": true,
"cost_usd": 0,
...
}Readable answers
Model replies come back as Markdown, which is unreadable in a raw terminal.
routewise ask, stream and chat render replies like a good chat app:
- headings → bold (underlined for
#/##) - lists →
• item, checklists →☐/☑ - inline code →
code, fenced code blocks → indented - tables → column-aligned rows
- quotes →
│ text,---→ a rule, links →label (url)
Colors and emphasis apply when stdout is a terminal. When the output is piped
or redirected, the markdown markup is stripped but no ANSI codes are emitted
— scripts and > file stay clean. Force plain text with --no-color
(or set the NO_COLOR environment variable).
Commands
| Command | What it does |
|---|---|
| routewise ask "<q>" | Route one query. Flags: --model, --support-mode, --tier, --threshold, --bypass-cache, --no-byom, --json, --quiet, --stream |
| routewise stream "<q>" | Stream tokens as they arrive |
| routewise chat | Multi-turn REPL. /tier cheap, /json on, /reset, /exit |
| routewise models | List available models + policies (cuts, tiers, evals) |
| routewise stats | Usage, cost, savings, cache rate, tier split |
| routewise logs [--limit N] | Recent request log lines |
| routewise log <id> | Full detail (incl. response) for one log entry |
| routewise analytics | Cost analytics + per-day breakdown |
| routewise compare / calibrate | Threshold-vs-cost recommendations (JSON) |
| routewise feedback <id> <up\|down> [reason] | Submit thumbs up/down on an answer |
| routewise pricing | Model price list |
| routewise providers | Supported providers + models |
| routewise byom | Bring your own model — set/remove/list per-tier overrides |
| routewise login | Open the dashboard, save your API key (hidden paste) |
| routewise whoami | Identity + account usage snapshot (key valid?, spend, tiers) |
| routewise doctor | End-to-end self-check: config, backend, providers, live ask |
| routewise evaluate "<q>" | Difficulty score + routing per mode (public, no key) |
| routewise config | Show current config / save API key / set base URL |
| routewise version | Print version |
Configuration
Settings are read in this order:
- CLI flag:
--key <key>/--base-url <url> - Environment:
ROUTEWISE_API_KEY,ROUTEWISE_BASE_URL(orROUTEWISE_API_BASE) - Config file:
~/.config/routewise/config.json(⭐%APPDATA%\routewise\config.jsonon Windows)
routewise config set rw_xxxxxxxx # save your key
routewise config set-base https://your-backend.example.com
routewise config # view current effective config
routewise config unset key # clear the saved keyPick a model / policy
Every request is scored for difficulty and routed to the cheapest tier that can handle it. You can also choose which difficulty policy drives the routing:
| --model | Policy | Who it's for |
|---|---|---|
| emma (default) | generic | General-purpose routing |
| lisa | 3-tier support | Customer support — cheap / mid / frontier (cuts at 2.0 / 4.5) |
| kate | 2-tier support | Customer support — cheap / frontier only (single cut) |
routewise ask "My order hasn't arrived" --model lisa # support, 3 tiers
routewise ask "I was charged twice" --model kate # support, 2 tiers
routewise chat --model lisa # REPL under the same policy
routewise evaluate "I want a refund" # shows what each model picks--support-mode generic|2tier|3tier sets the raw policy id directly (wins over
--model). The routing metadata line on stderr now includes the policy, and the
stream/done JSON includes support_mode.
Bring your own model
Force a specific provider/model (and optionally your own API key) for a tier.
Set it once — it is saved locally and attached to every ask/stream
automatically. Configure one, two, or all three tiers; tiers you don't
configure keep the router's defaults. No re-entry per call. Keys are stored only
in your local config file and sent per-request; they are never stored on the
RouteWise server.
# use openai/gpt-5 with your own key for tricky queries
routewise byom set frontier --provider openai --model gpt-5 --key sk-…
# pick a groq model but bill it through RouteWise's own server keys
routewise byom set cheap --provider groq --model openai/gpt-oss-20b
routewise byom list # view overrides (keys masked)
routewise byom remove frontier # drop one tier
routewise byom remove --all # drop everything
routewise ask "…" --no-byom # one call without the overridesSaved overrides are applied only when the server routes to that tier; if no
override is saved for a tier, it uses the router's defaults. A 2-model setup?
Just set cheap and frontier. A single-model setup? Set only cheap — every
query routes there, still scored but never up-scaled.
Using it in scripts
Exit codes: 0 success, 1 on any error (auth, rate limit, all tiers failed, network).
| Shell | Pattern |
|---|---|
| Pipeline | routewise ask "$Q" --json \| jq -r .response |
| Loop | for q in $(cat qs.txt); do routewise ask "$q" --quiet; done |
The answer is written to stdout even for errors-free responses, so piping and substitution behave the way any Unix tool would.
Building from source
cd cli
npm install
npm run build # tsc -> dist/src
npm test # 48 unit tests, mock-fetch (no network)
node dist/src/cli.js --helpUsing the client as a library
The same HTTP client backs the CLI and is importable:
import { RouteWiseClient } from "routewise";
const client = new RouteWiseClient({ apiKey: "rw_…" });
const res = await client.ask({ query: "hello" });
console.log(res.routed_to, res.cost_usd);Notes
- Requires an API key (
rw_…) created through the RouteWise dashboard orscripts/create_api_key.py. Unauthenticated endpoints likepricingandproviderswork without one. - This CLI hits the hosted RouteWise API over HTTP — it does not run any
model locally. For a fully offline path, run the backend + Ollama and point
routewise config set-baseat it. - The MCP server (
routewise mcp) that reuses this client for IDE integration is planned next.
