@rkarimabadi/model-probe
v1.1.3
Published
Probe any OpenAI-compatible endpoint and report which models are actually usable in a coding agent (tool calling, tool round-trip, streaming).
Downloads
856
Maintainers
Readme
model-probe
Point it at any OpenAI-compatible endpoint with an API key. It tells you which
models you can actually use inside a coding agent — not just which names are
listed in /v1/models.
Why this exists
A model list is not a capability list. On a real gateway we probed, /v1/models
returned 355 models; only 16 were permitted for the key, 332 returned
403 you are not allowed to use X model, and 7 were retired (410).
Worse, listing access is not agent access. Of those 16, several answered chat
requests perfectly but silently ignored the tools parameter — even with
tool_choice: "required". A coding agent cannot use them: it never gets a tool
call back. Blind trust in /v1/models produces a config that looks fine and
fails the moment the agent tries to edit a file.
model-probe runs a five-stage live probe per model and classifies the result.
Two gateways, two completely different traps
Both were probed by this tool, both with the same key.
Gateway A — 355 models listed. Only 16 accessible: 332 returned
403 you are not allowed to use X model and 7 were retired (410). Of the 16,
four answered chat perfectly but silently ignored the tools parameter, even
with tool_choice: "required" — claude-sonnet-4, gpt-5, x-ai/grok-4, and
gemini-2.5-pro. A coding agent cannot use those: it never gets a tool call back.
Only 3 backends were genuinely agent-ready. Also, eight different names
(gpt-5.1, gpt-5.2, o3-mini, o1-preview, …) were all the same gpt-5.
Gateway B — 24 models listed, all a newer generation, and every single one
unusable — for an entirely different reason. A naive client gets
415 Unsupported Media Type on all 24 and reports 24 broken models. The real
cause: the gateway is Fastify-based and rejects a bare application/json request
body, accepting only the parameterised form. Once the client sends
application/json; charset=utf-8, all 24 resolve to 402 payment-required with
three distinct billing reasons — the key is valid, the endpoint is healthy, and
the account is the blocker.
Neither failure mode is visible in /v1/models. Both are detected automatically.
Requirements
Node.js >= 18.17 (uses native fetch and node:util parseArgs).
Zero runtime dependencies — nothing to install.
Install
npm install -g @rkarimabadi/model-probeOr run it without installing:
npx @rkarimabadi/model-probe probe --include 'gpt|coder'From a clone of this repo:
npm link # or: npm install -g .The command is then available everywhere:
model-probe --version
model-probe doctorQuick start
# 1. Check auth + reachability. No paid completion calls.
model-probe doctor --base-url https://your-gateway.example/v1 --api-key YOUR_KEY
# 2. Probe a focused slice first
model-probe probe --include 'gpt|coder|deepseek|claude'
# 3. Then sweep everything and keep the evidence
model-probe probe --yes --out report.json --csv report.csvThe base URL does not need the /v1 suffix — pass https://your-gateway.example
and the prefix is auto-detected.
Commands
| Command | Purpose |
|---|---|
| probe | Probe models and classify coding-agent readiness |
| doctor | Verify config, auth source, and endpoint reachability (read-only) |
| list | Print model ids from /v1/models (discovery, no probing) |
| init | Store base URL + API key in ~/.model-probe/config.json |
| request | Raw escape hatch to any path on the gateway |
Full option list: model-probe help.
probe
model-probe probe [--model ID]... [--include REGEX]... [--exclude REGEX]...
[--all] [--limit N] [--concurrency N] [--no-stream]
[--out FILE] [--csv FILE] [--emit-config] [--provider-id ID]
[--json] [--verbose] [--quiet] [--fail-if-none] [--yes]--model IDprobes exactly that id and skips the catalog fetch (repeatable).--include/--excludefilter catalog ids by regex (repeatable).--allalso probes non-chat models (embeddings, OCR, speech, safety classifiers) that are skipped by default.- Refuses to probe more than 200 models unprompted; pass
--yes, or narrow with--include/--model/--limit.
request
model-probe request --path /models
model-probe request --method POST --path /chat/completions --data '{"model":"..."}' --allow-writeRead-only by default: any method other than GET/HEAD requires --allow-write.
How a model is judged
Five stages, each a real request to /chat/completions:
| Stage | What it does | What it proves |
|---|---|---|
| reach | 1-token completion (max_tokens: 1) | The model exists and the key may use it |
| tools (auto) | Sends a get_weather tool with tool_choice: "auto" | The model will call tools unprompted |
| tools (required) | Retries with tool_choice: "required" | Runs only if auto failed — catches models that need to be forced |
| roundtrip | Replays assistant.tool_calls + a role: "tool" result | The gateway accepts tool-result messages and the model answers from them |
| stream | stream: true, counts SSE chunks | The agent can stream tokens |
A tool call only counts if the response contains a real tool_calls entry
with a function name. Many gateways echo the request's tools array back; that
is not a tool call.
Non-chat models are filtered out of probe by default (and shown by
list --annotate) so you do not waste requests on embeddings and OCR models.
Declarations are evidence, not truth
When a catalog declares tool support, model-probe skips the tool stages only
if it declares false, and always verifies a declared true. That asymmetry is
deliberate, and one gateway proved why:
| Declared tools: true | Actually honoured tool calls |
|---|---|
| 138 models | 45 (33%) |
Trusting the declaration would have reported 138 usable models on that gateway
where 42 exist. Everything else came back chat-only — reachable, chat-capable,
and useless to an agent.
The same principle applies to non-chat filtering: a declared supports_chat:
false, api, category, or supported_endpoints is trusted (it is cheap and
removes work), but the name heuristics are only a fallback, and an unrecognised
category is probed rather than guessed away. Wrongly excluding a working model is
the worst failure this tool can have.
Verdicts
| Verdict | Meaning |
|---|---|
| agent-ready | Reachable, emits real tool calls, and consumes a tool result. Usable in a coding agent. |
| tools-partial | Emits tool calls, but the tool-result round-trip failed |
| chat-only | Reachable but ignores tools entirely — fine for chat, useless for an agent |
| rate-limited | HTTP 429 carrying a rate-limit message — transient; re-run |
| payment-required | HTTP 402, or a 401/403/429 whose body names a balance, quota, or spend-cap problem. The key is valid but the account cannot pay for this model. Not a capability limit, and never retried or paced |
| denied | HTTP 401/403 with a JSON error — this key may not use this model. Also used when a free tier is licensed only to the vendor's own client ("free tier can only be used in OpenCode") |
| gateway-blocked | HTTP 403 with an HTML body — an edge/WAF block, not a permission denial. Usually means the request shape is wrong |
| media-type-rejected | HTTP 415, or a 400 reporting an unparsed JSON body, that survived retries and every request shape |
| not-found | HTTP 404, or a 400 saying Unknown or unsupported model — listed but not provisioned for this account |
| not-chat | HTTP 400 saying the model takes only image/audio/document input, or is a realtime model the text endpoint refuses |
| gone | HTTP 410 — the provider retired the model |
| unreachable | Transport, TLS, or timeout failure (after one timeout escalation) |
| error | Any other upstream failure |
Ranking score (0–6): tool calling 3, tool-result round-trip 2, streaming 1. Results sort by verdict, then score, then latency.
Content-type negotiation
Some gateways fail to parse a bare application/json body and only accept the
parameterised form. model-probe sends the bare type first and, on that
rejection, rotates the whole run through four ordered {content-type, accept}
shapes and replays — without spending a retry. The chosen shape is reported on
stdout and as endpoint.contentType / endpoint.shapesTried in --json.
The rejection does not always look like a media-type error. Two real forms seen on gateways here:
415 {"code":"FST_ERR_CTP_INVALID_MEDIA_TYPE","message":"Expected request with `Content-Type: application/json`"}
400 {"fieldErrors":{"messages":["Invalid input: expected array, received undefined"]}}The second is the dangerous one: a 400 complaining that messages is missing,
for a request that definitely sent messages. Keying only on 415 left an entire
259-model gateway reporting every model as error. A 400 whose body reads like
an unparsed body now counts as recoverable too.
Negotiation is used rather than hardcoding the parameterised form because most gateways accept the bare type, and the tool should not change the request shape unless it has to.
The shape that works is locked in. As soon as a request with a body parses,
that shape becomes the run's preferred shape and every later request starts from
it — rotation never moves past it. Only a request that actually carried a body
may vote: a bodyless GET /models says nothing about how a body must be encoded,
and letting it vote pins the run to the wrong shape and fails every later POST.
If the preferred shape later starts being rejected, the preference is dropped and
the shapes are re-probed rather than leaving the run permanently stuck.
Cold starts
Serverless backends scale to zero, so the first request to a model pays its
load time. A model that looks dead at a 45s timeout can answer at 90s — one
gateway's cold model needed 86s, and reporting it as unreachable was a
false negative that hid a working model.
Timed-out requests are therefore retried once with the timeout doubled (floored at 120s, capped at 300s) before a model is called unreachable. Recoveries and non-recoveries are counted separately, because "escalated" and "recovered" are very different things:
COLD STARTS: 6 request(s) timed out and answered once the timeout was raised.
Serverless backends scale to zero, so a slow first answer is not a dead model.
2 request(s) still did not answer after the timeout was raised — those models are
hanging or far too slow, not cold-starting. Raise --timeout further to keep waiting.This applies to the reach probe — the first contact, which is the call that pays the cold start — not to every stage.
Billing blocks
HTTP 402 does not mean the model is broken, so it is reported separately rather than as a generic error — otherwise you go hunting for a capability problem that does not exist. Different model families are gated differently, so the reasons are grouped and counted:
BLOCKED BY BILLING — the key is valid, the balance/credit tier is not:
12x Insufficient balance
9x Daily check-in required to use free models. Please visit ... to check in.
3x Anthropic models are not available with the free 500,000 Telegram bonus credits.A billing failure does not always arrive as 402. Some gateways return it as HTTP 429 with a quota message, e.g.:
{"error":{"code":"1113","message":"Insufficient balance or no resource package. Please recharge."}}Classifying that as rate-limited is actively harmful in two ways: it tells the
user to slow down and retry, which can never work, and it makes the client pace
and retry a permanent block — one gateway went from ~0.4s per model to 31–70s
per model. So a 429 is classified by its body: a balance/quota message
becomes payment-required, and a genuine throttle message
("Too many requests after rate-limit", "Global rate limit exceeded") stays
rate-limited. Billing-classified responses are never retried or paced.
summary.billing[] carries the same data in --json. A tier that needs a daily
check-in is a five-second fix rather than a dead end, and the tool tells you which
is which.
Alias groups
Gateways often serve many model names from one backend. The report groups them,
so you learn that gpt-5.1 is really gpt-5 — and there are two independent
sources for this:
- Observed routing — the
SERVED BYcolumn, taken from the model name the gateway echoes back. This catches aliases that the catalog never admits to. - Declared aliases — some catalogs ship an
aliasesarray stating the other ids the same model answers to (qwen3.8-max == qwen/qwen3.8-max). The report prints these underDECLARED ALIASES, which is useful when writing config, where the exact id form matters.
Because catalog metadata is this valuable, --model still consults /models
rather than skipping it — otherwise a named model would lose its declared
aliases, context window, and tool support.
JSON policy
--json writes a single stable envelope to stdout. Diagnostics, warnings,
and the progress line go to stderr, so stdout stays pipeable.
Success (probe):
{
"ok": true,
"command": "probe",
"version": "1.0.0",
"generatedAt": "2026-01-01T00:00:00.000Z",
"endpoint": {
"baseUrl": "https://gateway.example/v1",
"apiBase": "https://gateway.example/v1",
"apiBaseResolved": true,
"baseUrlSource": "flag",
"apiKeySource": "env",
"authStyle": "bearer"
},
"summary": {
"catalogCount": 355,
"probed": 299,
"skipped": 56,
"counts": { "agent-ready": 5, "chat-only": 11, "denied": 276 },
"agentReady": 5,
"recommended": [{ "model": "openai/gpt-oss-20b", "score": 6, "latencyMs": 519, "served": "openai/gpt-oss-20b" }],
"aliasGroups": [{ "served": "gpt-5", "aliases": ["gpt-5", "gpt-5.1"], "verdict": "chat-only" }]
},
"skipped": [{ "id": "mistral-embed", "reason": "non-chat:embed" }],
"models": [
{
"model": "openai/gpt-oss-20b",
"served": "openai/gpt-oss-20b",
"reach": { "ok": true, "status": 200, "latencyMs": 519, "error": null, "message": null },
"capabilities": { "tools": true, "toolChoiceRequired": false, "roundtrip": true, "streaming": true, "doneSentinel": true },
"detail": {},
"toolCall": { "id": "call_1", "name": "get_weather", "arguments": "{\"city\":\"Tehran\"}", "origin": "tool_calls" },
"verdict": "agent-ready",
"score": 6
}
]
}Error (any command):
{ "ok": false, "error": { "code": "catalog_unavailable", "message": "GET .../models failed: HTTP 401" } }Error codes: invalid_option, missing_base_url, catalog_unavailable,
empty_selection, confirmation_required, write_not_allowed, bad_data,
unhandled.
list --json returns { ok, command, apiBase, count, ids[], skipped[] }.
request --json returns { ok, command, url, method, status, latencyMs, error, body }.
doctor --json returns { ok, checks[], auth, endpoint, ... } — with the key
masked, never in full.
The API key never appears in any output or report file. A test asserts this.
Exit codes
| Code | Meaning |
|---|---|
| 0 | Command ran successfully |
| 1 | Error (unreachable endpoint, failed request, failed doctor check) |
| 2 | Usage error (bad option, missing endpoint, confirmation required) |
| 4 | --fail-if-none and no model was agent-ready |
--fail-if-none makes probe usable as a CI gate.
Auth
Precedence, highest first:
--base-url/--api-keyflagsMODEL_PROBE_BASE_URL/MODEL_PROBE_API_KEYenvironment variables~/.model-probe/config.json(written byinit, override path withMODEL_PROBE_CONFIG)
Flags leak into shell history and process listings — prefer the env vars or
init for regular use.
Non-bearer gateways: --auth-style x-api-key|api-key|none, or send anything with
--header "X-Api-Key: ...". An explicit --header always wins over the resolved key.
Self-signed endpoints: --insecure disables TLS verification and says so on stderr.
Emitting a provider config
model-probe probe --model openai/gpt-oss-20b --emit-config --provider-id mygatewayPrints a ready-to-paste MiMoCode / OpenCode provider block for the top-ranked
model, with apiKey left as <YOUR_API_KEY> so the block stays safe to share or
commit.
What this does not prove
- Not a benchmark. It measures protocol compatibility, not answer quality. A
model can be
agent-readyand still be bad at coding. - Point-in-time. Gateways change routing. Re-run when behavior looks odd.
- No cost or limit data. Context window, output limits, and pricing are not
published by
/v1/models, so they are deliberately not guessed. - A single probe run can be unlucky. One
504observed in one run does not mean a model is broken; check--verbosestage detail and re-run.
Development
npm test # node:test — 127 tests, no network to the outside worldTests cover the pure decision logic — tool-call detection (including the "prose mentions tools" and "echoed empty array" false positives), verdict classification, scoring and ordering, catalog filtering, alias grouping, billing grouping, CSV escaping, and the provider-config shape — plus a real local HTTP server that proves the 415 content-type promotion replays correctly, sticks for later requests, and stays bounded when the gateway never accepts anything.
