@aditya005/jev-router
v0.5.0
Published
A tiny harness where a small router model (JEV) reads each question and picks which model answers it
Maintainers
Readme
jev-router
A small model reads your question first and decides which model should answer it.
Most harnesses send everything to one model. You pick it once, and then you either overpay on "what's the capital of Portugal" or you underthink "design our multi-region failover". jev puts a cheap classifier in front: it reads the question, picks a tier, and only then does the expensive model get woken up.
Two real runs through the same CLI, same config, nothing forced except the second tier:
$ jev --explain "what is the capital of Portugal?"
[jev] tier=low via=router model=claude-haiku-4-5-20251001 spec=latest:haiku route=327ms answer=779ms cost=$0.0001
[jev] prompt ~8 tokens
[jev] why: jev chose low, confidence 1.00
The capital of Portugal is Lisbon.
$ jev --explain --tier high "design a multi-region active-active failover strategy for a write-heavy postgres workload"
[jev] claude-opus-5 takes adaptive thinking, so the 4000-token budget is advisory only
[jev] tier=high via=forced model=claude-opus-5 spec=latest:opus route=0ms answer=147691ms cost=$0.2434The gap between those two cost lines is the whole argument. A router only has to be right often enough to be worth the 300ms it adds.
you ──▶ JEV (small router) ──▶ {"tier":"high","reason":"architecture tradeoffs"}
│
▼
low → haiku, no thinking
mid → sonnet, no thinking
high → opus, adaptive thinking ──▶ answerZero runtime dependencies. TypeScript, fetch, and the Node standard library.
Why it is built this way
A walk-in clinic puts a nurse at the front desk. Ninety seconds, no MRI, no treatment. All the nurse decides is where you go next: back out with a bandage, into an exam room, or upstairs to the specialist who will spend an hour on you. That ninety seconds is cheap and the specialist's hour is not, which is the entire economics of the thing.
Two rules fall out of that, and both shaped the code. The nurse has to be cheap, so JEV gets a
small model, a 200 token ceiling, and a prompt forbidding it from answering. And the nurse being
wrong cannot take down the clinic, so every routing failure falls through to defaultTier with a
warning instead of an exception. A routing miss costs you money or quality. It never costs you the
request.
The prompt also tells JEV to round up when it cannot decide. A hard question sent to a weak model comes back confidently wrong and you have to catch it yourself. An easy question sent to a strong model just wastes a few cents.
Install
Node 22.6 or newer, because the dev scripts run TypeScript directly through
--experimental-strip-types.
npm install -g @aditya005/jev-router # the `jev` CLI, globally
npm install @aditya005/jev-router # or as a library in a project
npx @aditya005/jev-router --models # or try it without installing anythingThe package is scoped because the unscoped jev-router was already taken on npm by an unrelated
project. The binary it installs is jev, not jev-router. If npm cannot find the package, it has
not been published yet and the source install below works either way.
git clone https://github.com/adityaarakeri/jev-router.git
cd jev-router
npm install
cp .env.example .env # fill in your keys
npm run build && npm link # puts `jev` on your PATHIssues and pull requests: github.com/adityaarakeri/jev-router.
Then the keys, in a .env next to wherever you run it. Installed from npm with no config file of
your own, jev falls back to the defaults in src/config.ts: a haiku router and Claude tiers, so
ANTHROPIC_API_KEY on its own is enough. The jev.config.json checked into this repo routes
through TypeSafe instead, so a clone wants two:
TYPESAFE_API_KEY=apikey... # the router
ANTHROPIC_API_KEY=sk-ant-... # the answering tiersRun that config with only the Anthropic key and every request still gets answered, but the router
fails and falls back to mid every time, which defeats the point. If you would rather not sign up
for a second service, point the router at Claude instead:
{ "router": { "provider": "anthropic", "model": "latest:haiku", "pin": "claude-haiku-4-5-20251001", "maxTokens": 200 } }The CLI reads .env itself on every run. It takes the nearest one at or above the working
directory and only fills in names the environment has not already set, so an exported key still
beats the file. --env-file <path> picks a different file, --no-env-file skips it. The library
entry point does not do this on your behalf; loadEnvFile() is exported if you want it.
Answers go to stdout and routing chatter goes to stderr, so piping works:
jev "summarise this file" < notes.md > answer.mdThe CLI
jev "why is my build slow?" route it, answer it, print the answer
jev --explain "..." also print tier, model, latencies and cost
jev --route-only "..." stop after the decision, print it as JSON
jev --tier high "..." skip the router, force a tier
jev --escalate "..." let JEV review the answer and redo it a tier up
jev --stats summarise the decision log
jev --models show what each "latest:" spec resolves to today
jev --strict exit 2 if the request fell back instead of routing
jev --config ./other.json "..." different config file
jev --env-file ./prod.env "..." different env file
echo "long prompt" | jev --explain read the question from stdin--route-only is the one you will live in while tuning. It calls the router and nothing else, so
you can throw fifty questions at it for a fraction of a cent before spending anything on answers.
--tier is how you get a baseline: run your question set forced to high, then run it routed, and
compare. Without that comparison you cannot tell whether the router is earning its keep.
Demos
Four scripts in demos/, each behind a one-line npm run:
npm run demo:guardrails # free, offline, no API key
npm run demo:route # routing only, about $0.0001
npm run demo:bakeoff # routing only, two routers compared
npm run demo:escalate # real answers, the expensive onedemo:guardrails is the fastest way to understand the design and it works on a plane. Every
provider is a fake, so it sends no traffic and costs nothing, and it walks the failure modes the
harness exists to survive:
scenario outcome detail
router returns prose mid (fallback=true) reason: router output was not valid JSON
router throws mid (fallback=true) reason: router call failed
router names a tier that does not exist mid (fallback=true) reason: router output was not valid JSON
prompt far larger than the router high (jev said low) ~97,508 tokens, router read 6,721 chars
bypass rule matches low (source=bypass) router called: false
answering model fails threw: 503 from the provider logged with error: 503 from the provider
router behaves mid (fallback=false) reason: ordinary debuggingNo routing failure there takes the request down with it. The one case that does throw, an
answering model dying mid-call, still writes its log row before rethrowing, so --stats cannot
quietly undercount failures.
demos/README.md covers what each one is for and what to watch in the output.
Configuration
jev.config.json at the repo root, merged one level deep over the defaults in src/config.ts. A
config file can be three lines long.
{
"defaultTier": "mid",
"router": { "provider": "typesafe", "model": "jev-latest", "maxTokens": 200 },
"tiers": {
"high": {
"provider": "anthropic",
"model": "claude-opus-5",
"maxTokens": 16000,
"thinkingTokens": 4000,
"timeoutMs": 600000,
"when": "Multi-step reasoning, architecture and tradeoff calls, tricky maths or proofs..."
}
}
}The when field is not documentation. It is pasted straight into the router's prompt as the menu
it chooses from, which means retuning the router is editing prose, not code. If too much lands on
high, tighten that line and widen mid.
thinkingTokens maps to Anthropic extended thinking, and how it gets sent depends on what the
resolved model accepts. The catalogue reports that. Older models take an explicit budget
({"type":"enabled"}) and the provider raises max_tokens above it, since the API rejects the
request otherwise. Models from 4.6 onward take {"type":"adaptive"} and reject a budget outright,
so the number becomes advisory: it says this tier should think, and the model picks the depth. A
warning tells you which happened.
Adaptive models bill reasoning as output tokens, so maxTokens on a thinking tier covers thinking
and answer together. Set it too low and the answer stops mid-sentence with nothing to say it did.
The shipped high tier is 16000 for that reason. Nothing streams here, so the whole answer has to
land inside one request, which is what the per-tier timeoutMs is for: 600000 on high, and it
overrides the top-level value for that tier only, so a slow thinking tier does not impose the same
patience on low.
| Variable | Effect |
| --- | --- |
| ANTHROPIC_API_KEY | required for any tier on the anthropic provider |
| ANTHROPIC_BASE_URL | point at a gateway or proxy |
| TYPESAFE_API_KEY | required when the router uses the typesafe provider |
| TYPESAFE_BASE_URL | point the typesafe provider somewhere else |
| OPENAI_API_KEY, OPENAI_BASE_URL | for the openai-compatible provider |
| JEV_ROUTER_MODEL | swap the router's model without editing the config |
| JEV_DEFAULT_TIER | swap the fallback tier |
| JEV_LOG_PATH | decision log path, or off to disable |
| JEV_ROUTER_PIN | override the router's pin |
| JEV_CODEX_BIN, JEV_OPENCODE_BIN | paths to the CLI binaries |
An override that disagrees with your config announces itself once on stderr, because a stale line
in a .env silently beating the config file is a miserable half hour to debug:
[jev] env overrides in effect: JEV_ROUTER_MODEL=claude-haiku-4-5-20251001 (config said jev-latest)routerTimeoutMs is separate from timeoutMs and much shorter, since a router that hangs is worse
than no router: you want it to give up and fall through long before the user notices.
Routing with TypeSafe System One
Routing is classification, not generation, and TypeSafe (docs) sells
exactly that. You hand it some state and a typed question, it hands back a probability
distribution. The typesafe provider turns a routing decision into one choice question where the
tier menu is the criteria map:
"router": {
"provider": "typesafe",
"model": "jev-latest",
"maxTokens": 200,
"pricing": { "inputPer1M": 0.042, "outputPer1M": 0 }
}One POST /v1/systemone with the digested question as state and each tier's when line as an
option description, then answers.jev.choice comes back. Because the reply is a distribution over
the tier names, there is no JSON to malform and no prefill seam to fumble, so the unparseable-reply
failure that most of this repo's fallback machinery exists for cannot happen here. A 401, a 500 or
a timeout still can, and still falls back the same way. The confidence lands in the decision
reason, which means --explain and every log row tell you how sure it was.
One run of npm run demo:bakeoff, the five questions in eval/questions.txt through both routers:
question typesafe haiku agree typesafe ms haiku ms
capital of Portugal low low yes 342ms 1251ms
is 1361 prime? low low yes 147ms 787ms
design a multi-region failover strategy for a p... high high yes 103ms 830ms
summarise the tradeoffs between optimistic and ... mid mid yes 112ms 1012ms
write a threat model for a webhook receiver tha... high high yes 151ms 794ms
router total latency spend fallbacks
typesafe 855ms $0.000087 0/5
haiku 4674ms $0.0029 0/5Five questions, one run, on one machine in one place. That is an illustration, not a benchmark, and your traffic is the only sample that means anything. Run the demo yourself before believing it.
Two things to know. TypeSafe is a router here and not a model: an answering tier on typesafe is
rejected at config load, because System One does not write prose. And latest: is rejected too,
since there is no catalogue to resolve against, so name jev-latest or pin a version like
jev-1.13.0.
Escalation runs on the router's provider, so it becomes a noul question there: a 0 to 1 belief
that the cheap answer needs redoing, escalating at 0.5 and up.
Routing locally
The openai provider talks to anything speaking the OpenAI chat-completions shape, which covers
Ollama, LM Studio and vLLM. A small local model can often classify well enough, and then the only
thing leaving your machine is the answering call.
ollama pull qwen2.5:3b
export OPENAI_BASE_URL=http://localhost:11434/v1{ "router": { "provider": "openai", "model": "qwen2.5:3b", "maxTokens": 200 } }Check it with --route-only on a handful of questions before trusting it. Small local models are
the likeliest thing here to return prose where you asked for JSON, which lands on the fallback path
every time and quietly turns your router into a constant.
Knowing whether the router still works
A router that has stopped routing looks exactly like one that is working. Every request still gets
answered, the logs look normal, and the only symptom is that everything lands on defaultTier.
Three things exist to catch that.
The decision log. One JSON line per call in .jev/decisions.jsonl. The question itself is never
written down, only a 12 character hash, so the log is safe to keep and you can still count repeats.
{"ts":"2026-09-21T04:07:37.644Z","questionId":"d8f0a5eb7c3e","questionChars":89,
"tier":"high","finalTier":"high","source":"forced","fallback":false,
"reason":"forced by caller","model":"claude-opus-5","modelSpec":"latest:opus",
"routeMs":0,"answerMs":147691,"inputTokens":32,"outputTokens":9728,
"answerCostUsd":0.24336,"totalCostUsd":0.24336,"unpricedCalls":0,"escalations":0}tier is what the router picked and finalTier is what actually answered. When they differ, an
escalation or a size move happened, and that is the row worth reading.
A few log properties worth knowing: reason is model-controlled text and gets truncated to 200
characters; size moves keep both ends in contextAdjustedFrom and contextAdjustedTo; a failed
answering call still appends a row with an error field before rethrowing; corrupt lines never
crash the reader and --stats reports how many it skipped. There is no automatic rotation, so the
file grows until you rotate it, and --stats warns once it passes about 10 MB.
jev --stats turns that file into the aggregate. The shape is real, the numbers below are
made up for the example:
decisions 184
final tier
low 61 33.2%
mid 94 51.1%
high 29 15.8%
source router=171 bypass=9 fallback=4
fallback rate 2.2% (4)
escalated 12
median route 287ms
median answer 3140ms
total cost $4.8812The fallback rate is the number to watch. Anything creeping upward means the router is failing and nobody noticed.
--strict turns a fallback into exit code 2. The answer still prints, so nothing is lost, but a
CI job or a batch script can fail loudly instead of quietly paying for the default tier forever.
total cost says n/a until you fill in pricing on the router and each tier. Nothing ships with
real prices, because a number baked in today is a wrong number in three months. Costs are computed
from reported token usage, which makes them an estimate of what you spent rather than a bill, and
they exclude anything your provider charges beyond input and output tokens.
Escalation, for questions that looked easy
Routing happens before the answer exists, so the router is judging a question by its surface. The expensive failure is the question that reads simple and is not.
With --escalate (or escalation.enabled), a cheap answer gets a second cheap call: question plus
answer, one verdict back. If the reviewer says it looks guessed or incomplete, the question is
re-answered one tier up and both attempts land in the log.
One run of npm run demo:escalate, three questions forced onto low:
question verdict why kept answer total paid
is 1361 prime? show your working low (stood) reviewer was satisfied $0.0020 $0.0020
two goroutines both read a map.. low -> mid jev scored this answer 0.59, thr.. $0.0108 $0.0153
what is the capital of Portugal? low (stood) reviewer was satisfied $0.000074 $0.000074The gap between those two cost columns is an answer you paid for and threw away. totalCostUsd
includes every attempt, so escalation never looks free in --stats, and a tier that gets bumped
half the time is more expensive than starting one tier up. If the check itself fails or returns
junk, the original answer stands, because a broken reviewer must never destroy an answer that
already exists.
Skipping the router
Every routed question pays for a router round trip, and on "thanks" that is pure waste. Bypass rules answer deterministically without calling the router at all:
"bypass": {
"maxChars": 0,
"tierForShort": "low",
"patterns": [
{ "pattern": "^(thanks|thank you|ok|cool)\\b", "tier": "low" },
{ "pattern": "\\b(architecture|threat model|race condition)\\b", "tier": "high" }
]
}Patterns are checked first, then the length rule. Both are off by default, and maxChars deserves
suspicion in particular: short and easy are different properties. "is 1361 prime?" is 17 characters
and will embarrass a small model. Start with patterns and leave maxChars at 0 unless your traffic
really is full of one-word acknowledgements.
Bypassed decisions log as "source":"bypass", so you can see how much traffic never reached the
router.
The router runs on a small model with a smaller window than the tier it routes to. Paste a 300k token log dump into a naive version of this design and it uploads the whole thing to the router, gets a 400 because it does not fit, falls back, and charges you for the upload. Even when it does fit, you just paid router input tokens on 300k tokens to save money on the answer.
So the router does not read the prompt. It reads a digest: the first headChars characters (2000
by default), a marker saying how much was cut, the last tailChars characters (1000), and a line
stating the rough size of the whole thing. The router call therefore has a fixed ceiling no matter
what arrives. A 1.35 MB prompt produces a 3.7 KB router call.
Head and tail are not arbitrary. In a real long prompt the instruction sits at one end or the other: "summarise this incident report and tell me the root cause" at the top, "what actually caused the outage?" at the bottom, 300k tokens of log lines in between. The middle is the thing being operated on. The ends are the description of the job, and the job is what decides the tier.
Size is itself a routing signal
Two mechanisms act after the router has chosen, and neither can lower a tier.
Context windows come from the Models API (max_input_tokens), or from maxInputTokens in the tier
config when the catalogue is silent. If the chosen tier cannot hold the prompt plus reserveTokens
of headroom, the request walks up to the first tier that can:
[jev] moved low -> high: prompt is roughly 337,523 tokens
[jev] why: looks simple (moved up for context size)Floors are for when you want size alone to force a stronger model even where it would technically fit:
"contextGuard": {
"headChars": 2000,
"tailChars": 1000,
"reserveTokens": 8000,
"floors": [
{ "minTokens": 60000, "tier": "mid" },
{ "minTokens": 150000, "tier": "high" }
]
}A small model with a big window will happily accept 150k tokens and then reason poorly across them. Floors are how you say "past this size, I do not care how simple the question looked".
The estimate is a heuristic, so it never refuses
Sizes come from a four-characters-per-token approximation. Good enough to choose between tiers, wrong enough that it should never veto a request: code and CJK run denser than English prose, so the estimate can miss badly in either direction.
When nothing on the menu can hold the prompt, the harness uses the widest tier, warns with both numbers, and sends it anyway.
[jev] prompt is roughly 2,000,000 tokens and the widest tier holds 1,000,000.
Sending it anyway, the API will decide.The API is the authority on what fits. A local heuristic refusing a request that would have worked
is a worse failure than a clear 400. Forced tiers get the same treatment: --tier low on a huge
prompt still moves up if low cannot hold it, with a warning, because the alternative is a
guaranteed error.
All of it lands in the log as promptTokensEst, routerTruncated and contextAdjusted, so you can
go back and count how often size rather than the router picked the tier.
Worth knowing up front: the Claude API has no claude-haiku-latest alias. Older minor versions do
get a convenience alias (claude-sonnet-4-5 points at the newest dated snapshot of that minor
version), but from the 4.6 generation onward the dateless IDs like claude-sonnet-5 are canonical
pinned snapshots, not evergreen pointers. Anthropic does not swap weights under an existing ID; a
newer model ships under a new one.
See https://platform.claude.com/docs/en/about-claude/models/model-ids-and-versions
So resolution happens client side. A tier's model takes a concrete ID or a family spec:
"high": { "model": "latest:opus", "pin": "claude-opus-5" }latest:opus asks the Models API for every claude-opus-* the key can see and picks the most
recently released one. pin is the concrete model used when that lookup cannot happen. See what
you would get right now, without spending anything on an answer:
$ jev --models
catalogue cache [anthropic]: .jev/models.json (6 min old, ttl 24h)
router jev-latest -> jev-latest [pinned]
low latest:haiku -> claude-haiku-4-5-20251001 [cache]
mid latest:sonnet -> claude-sonnet-5 [cache]
high latest:opus -> claude-opus-5 [cache]How it degrades
The catalogue is fetched once and cached to .jev/models.json for models.ttlHours (24 by
default), so you are not paying a round trip per question. When the lookup fails, resolution walks
down in order, warning at each step: fresh cache inside the TTL, then the live Models API, then
stale cache past its TTL because an old answer beats no answer, then the pin, and finally a
thrown error naming the spec if there is no pin. That last step is deliberate, since silently
substituting some other model would be worse than failing. jev --refresh-models skips straight to
the API.
The part that should make you uneasy
latest: means the model can change under you between two runs with nothing in your config
changing. Convenient for a chat tool. For an eval harness it quietly destroys comparability, since
last week's numbers came from a model you can no longer name.
Two things mitigate it. The log records both the resolved ID and the spec
("model":"claude-opus-5","modelSpec":"latest:opus"), so a run is always traceable to the concrete
model that produced it. And jev.config.example.json ships fully pinned, which is the config to
copy for anything you intend to repeat or publish. Rule of thumb: latest: for everyday use,
pinned IDs for evals.
Two smaller consequences. The pricing you filled in belongs to the model you had in mind, so a
resolution onto a newer model makes those estimates wrong until you update them. And thinking
support is read from the catalogue: a model that reports no support for extended thinking gets
thinkingTokens dropped with a warning rather than a request the API will reject, and the
catalogue also says which thinking parameter the model takes, so moving a tier onto a newer model
switches it to adaptive thinking instead of 400-ing on a budget it no longer accepts.
Local models
The same spec works against the openai provider, matching loosely on the id since local catalogues do not follow the Claude naming scheme:
"router": { "provider": "openai", "model": "latest:qwen2.5", "pin": "qwen2.5:3b" }That resolves against GET {OPENAI_BASE_URL}/models and picks the newest qwen2.5 tag Ollama
reports.
A tier can be answered by the Codex CLI or the opencode CLI instead of an HTTP API. These are subprocess agents authenticated by a subscription, not completions endpoints: jev spawns the process, writes the prompt to stdin, and reads the final message back.
"high": {
"provider": "codex",
"model": "gpt-5.4",
"pin": "gpt-5.4",
"maxTokens": 8192,
"maxInputTokens": 200000,
"when": "Multi-step reasoning, architecture and tradeoff calls...",
"exec": { "sandbox": "read-only" }
}What changes, and none of it is hidden. Both run sandboxed by default (codex with -s read-only,
opencode with --pure and no --auto) because these tools can read files and run commands, and
the bypass flags are rejected at config load. Neither reports token usage, so costUsd is
undefined and --stats prints total cost $X (excludes N unpriced calls) rather than
under-reporting; pricing on a CLI tier is rejected outright, since subscription billing has no
per-call price. Neither reports a context window either, so set maxInputTokens yourself if you
want size-based moves to work there.
They are answering tiers, not routers. A trivial prompt takes about 5 to 7 seconds through either
CLI (measured p50: codex ~6.6s, opencode ~5s), which sits at or above the default 5s
routerTimeoutMs, and config load warns if you point the router at one anyway.
Codex has no model catalogue, so a codex tier needs a concrete model or a pin. Opencode lists
models via opencode models, but those records carry no timestamps, so latest: there degrades to
"highest sorting id".
This section only matters for chat-model routers. A typesafe router returns a distribution and
skips the problem entirely.
The prompt asks for JSON and the parser tolerates fences and preamble, but neither is a guarantee. Two mechanisms sit behind it.
When the catalogue says the resolved model supports structured outputs, the router call carries a
JSON Schema (output_config.format) naming the tier enum and requiring reason, and the prefill is
dropped. The reply is then JSON by construction. The escalation reviewer carries its own schema, so
a verdict is never held to the router's shape.
Otherwise the harness prefills the assistant turn with {"tier":, so the model is not deciding
whether to open with prose: it is mid-object already and the cheapest continuation is to finish the
JSON. The provider prepends the prefill back onto the response so the parser sees a complete object
either way. This is the weaker mechanism and the reason it is no longer the default. Continuing
from a prefill that ends right before a quote, models sometimes emit a stray or doubled one
({"tier":""low"" or {"tier":low"), which does not parse and costs you the routing.
The openai provider ignores both, since support varies across local servers. Codex gets the same
schema through --output-schema. Opencode rests on the prompt alone. Whatever the mechanism, an
unparseable reply is warned about and falls back. It is never thrown.
Model JSON is extracted rather than parsed, too: a brace walker with string and escape awareness
returns every balanced object and the first one carrying a known tier wins, which survives two
objects in one reply and braces inside reason.
Transient HTTP failures (429, 500, 502, 503, 504, 529, plus network errors and timeouts) are retried
up to maxRetries times with exponential backoff and jitter, honouring retry-after when present.
Anything else fails immediately, and a failed router call still lands on the fallback path rather
than killing the request.
Timeouts are per attempt, not per request. With routerTimeoutMs: 5000 and maxRetries: 2, a
single router call can take up to about 15s plus backoff before it gives up and falls back. Budget
for the multiplier.
One consequence found the hard way: a thinking-tier answer that legitimately takes 147 seconds will
burn three full attempts against a 60s timeout before failing, and the tokens each attempt
generated are still billed. That is what the per-tier timeoutMs is for.
Using it as a library
import { ask, route, loadConfig } from "@aditya005/jev-router";
const answer = await ask("explain the CAP theorem to a new hire");
console.log(answer.decision.tier, answer.model, answer.text);
// routing only, if you want to dispatch somewhere else
const config = loadConfig();
const decision = await route(config, "is 1361 prime?");Both take injectable providers, which is how the test suite runs with zero network calls:
const answer = await ask("design a sharding scheme", {
routerProvider: async () => ({ text: '{"tier":"high","reason":"architecture"}' }),
answerProvider: async (req) => ({ text: `would have called ${req.model}` }),
});Project layout
src/types.ts shared types, no logic
src/config.ts defaults, file merge, env overrides, validation
src/http.ts POST plus retry and backoff, shared by the HTTP providers
src/provider.ts anthropic + openai-compatible calls, prefill handling, provider registry
src/providers/ typesafe.ts, codex.ts, opencode.ts
src/exec.ts process spawning for the CLI providers (stdin prompt, tree-kill)
src/jev.ts the router: prompt, schema, bypass, JSON parsing, fallback, counters
src/escalate.ts the second-pass reviewer that bumps a weak answer up a tier
src/harness.ts bypass -> route -> answer -> escalate -> log
src/log.ts JSONL decision log, reader, aggregate summary
src/cost.ts token to USD, undefined when pricing is not configured
src/models.ts "latest:opus" -> a concrete ID, with cache and pin fallback
src/context.ts prompt digest, size floors, context window fitting
src/env.ts .env discovery and parsing
src/cli.ts flag parsing and output
test/ node:test, fake providers, no network
demos/ four runnable demos, see demos/README.md| Command | What it does |
| --- | --- |
| npm run dev -- "question" | run the CLI from source, no build step |
| npm run build | compile to dist/ with declarations |
| npm run lint | type-check src, test and demos, no emit |
| npm test | run the suite against fake providers |
| npm run demo:guardrails | failure modes, offline, free |
| npm run demo:route | route the eval set and price each decision |
| npm run demo:bakeoff | the same questions through two routers |
| npm run demo:escalate | force low, review, and show what the bump cost |
CI runs lint, test and build on Node 22 and 24. Run npm ci first: npm run lint needs
typescript from node_modules, so it fails on a clean checkout until you install.
What this deliberately does not do
Worth knowing before you build on it.
- No caching. The same question routes twice and costs twice. The log stores a question hash, so a
cache keyed on
questionIdis the obvious next step. - No streaming.
ask()waits for the full response, which is also why a thinking tier needs a generoustimeoutMs. - No conversation history. Each call is a single user turn. Multi-turn changes the routing question in an interesting way, since a thread that started easy can turn hard halfway through.
- Token estimates are character based. A tokeniser would be more accurate at the cost of a dependency and a load-time hit.
- The router sees the raw question. The prompt marks it as data and the parser only accepts a known
tier name, but a determined prompt can still try to talk its way into
high. The blast radius is money, not correctness. - A
latest:spec resolves per process, not per request, and a long-lived process holds its resolution until the cache expires. - Cost figures are estimates from reported tokens, not invoices, and only as current as the prices you typed into the config.
- Routing quality is unmeasured here. You get
--route-only,--tierand the demos so you can measure it against your own question set. Until you do, any claim about savings is a guess, including the ones in this README.
License
MIT. See LICENSE.
