npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@rkarimabadi/model-probe

v1.1.3

Published

Probe any OpenAI-compatible endpoint and report which models are actually usable in a coding agent (tool calling, tool round-trip, streaming).

Downloads

856

Readme

model-probe

Point it at any OpenAI-compatible endpoint with an API key. It tells you which models you can actually use inside a coding agent — not just which names are listed in /v1/models.

Why this exists

A model list is not a capability list. On a real gateway we probed, /v1/models returned 355 models; only 16 were permitted for the key, 332 returned 403 you are not allowed to use X model, and 7 were retired (410).

Worse, listing access is not agent access. Of those 16, several answered chat requests perfectly but silently ignored the tools parameter — even with tool_choice: "required". A coding agent cannot use them: it never gets a tool call back. Blind trust in /v1/models produces a config that looks fine and fails the moment the agent tries to edit a file.

model-probe runs a five-stage live probe per model and classifies the result.

Two gateways, two completely different traps

Both were probed by this tool, both with the same key.

Gateway A — 355 models listed. Only 16 accessible: 332 returned 403 you are not allowed to use X model and 7 were retired (410). Of the 16, four answered chat perfectly but silently ignored the tools parameter, even with tool_choice: "required" — claude-sonnet-4, gpt-5, x-ai/grok-4, and gemini-2.5-pro. A coding agent cannot use those: it never gets a tool call back. Only 3 backends were genuinely agent-ready. Also, eight different names (gpt-5.1, gpt-5.2, o3-mini, o1-preview, …) were all the same gpt-5.

Gateway B — 24 models listed, all a newer generation, and every single one unusable — for an entirely different reason. A naive client gets 415 Unsupported Media Type on all 24 and reports 24 broken models. The real cause: the gateway is Fastify-based and rejects a bare application/json request body, accepting only the parameterised form. Once the client sends application/json; charset=utf-8, all 24 resolve to 402 payment-required with three distinct billing reasons — the key is valid, the endpoint is healthy, and the account is the blocker.

Neither failure mode is visible in /v1/models. Both are detected automatically.

Requirements

Node.js >= 18.17 (uses native fetch and node:util parseArgs). Zero runtime dependencies — nothing to install.

Install

npm install -g @rkarimabadi/model-probe

Or run it without installing:

npx @rkarimabadi/model-probe probe --include 'gpt|coder'

From a clone of this repo:

npm link          # or: npm install -g .

The command is then available everywhere:

model-probe --version
model-probe doctor

Quick start

# 1. Check auth + reachability. No paid completion calls.
model-probe doctor --base-url https://your-gateway.example/v1 --api-key YOUR_KEY

# 2. Probe a focused slice first
model-probe probe --include 'gpt|coder|deepseek|claude'

# 3. Then sweep everything and keep the evidence
model-probe probe --yes --out report.json --csv report.csv

The base URL does not need the /v1 suffix — pass https://your-gateway.example and the prefix is auto-detected.

Commands

| Command | Purpose | |---|---| | probe | Probe models and classify coding-agent readiness | | doctor | Verify config, auth source, and endpoint reachability (read-only) | | list | Print model ids from /v1/models (discovery, no probing) | | init | Store base URL + API key in ~/.model-probe/config.json | | request | Raw escape hatch to any path on the gateway |

Full option list: model-probe help.

probe

model-probe probe [--model ID]... [--include REGEX]... [--exclude REGEX]...
                  [--all] [--limit N] [--concurrency N] [--no-stream]
                  [--out FILE] [--csv FILE] [--emit-config] [--provider-id ID]
                  [--json] [--verbose] [--quiet] [--fail-if-none] [--yes]
  • --model ID probes exactly that id and skips the catalog fetch (repeatable).
  • --include / --exclude filter catalog ids by regex (repeatable).
  • --all also probes non-chat models (embeddings, OCR, speech, safety classifiers) that are skipped by default.
  • Refuses to probe more than 200 models unprompted; pass --yes, or narrow with --include / --model / --limit.

request

model-probe request --path /models
model-probe request --method POST --path /chat/completions --data '{"model":"..."}' --allow-write

Read-only by default: any method other than GET/HEAD requires --allow-write.

How a model is judged

Five stages, each a real request to /chat/completions:

| Stage | What it does | What it proves | |---|---|---| | reach | 1-token completion (max_tokens: 1) | The model exists and the key may use it | | tools (auto) | Sends a get_weather tool with tool_choice: "auto" | The model will call tools unprompted | | tools (required) | Retries with tool_choice: "required" | Runs only if auto failed — catches models that need to be forced | | roundtrip | Replays assistant.tool_calls + a role: "tool" result | The gateway accepts tool-result messages and the model answers from them | | stream | stream: true, counts SSE chunks | The agent can stream tokens |

A tool call only counts if the response contains a real tool_calls entry with a function name. Many gateways echo the request's tools array back; that is not a tool call.

Non-chat models are filtered out of probe by default (and shown by list --annotate) so you do not waste requests on embeddings and OCR models.

Declarations are evidence, not truth

When a catalog declares tool support, model-probe skips the tool stages only if it declares false, and always verifies a declared true. That asymmetry is deliberate, and one gateway proved why:

| Declared tools: true | Actually honoured tool calls | |---|---| | 138 models | 45 (33%) |

Trusting the declaration would have reported 138 usable models on that gateway where 42 exist. Everything else came back chat-only — reachable, chat-capable, and useless to an agent.

The same principle applies to non-chat filtering: a declared supports_chat: false, api, category, or supported_endpoints is trusted (it is cheap and removes work), but the name heuristics are only a fallback, and an unrecognised category is probed rather than guessed away. Wrongly excluding a working model is the worst failure this tool can have.

Verdicts

| Verdict | Meaning | |---|---| | agent-ready | Reachable, emits real tool calls, and consumes a tool result. Usable in a coding agent. | | tools-partial | Emits tool calls, but the tool-result round-trip failed | | chat-only | Reachable but ignores tools entirely — fine for chat, useless for an agent | | rate-limited | HTTP 429 carrying a rate-limit message — transient; re-run | | payment-required | HTTP 402, or a 401/403/429 whose body names a balance, quota, or spend-cap problem. The key is valid but the account cannot pay for this model. Not a capability limit, and never retried or paced | | denied | HTTP 401/403 with a JSON error — this key may not use this model. Also used when a free tier is licensed only to the vendor's own client ("free tier can only be used in OpenCode") | | gateway-blocked | HTTP 403 with an HTML body — an edge/WAF block, not a permission denial. Usually means the request shape is wrong | | media-type-rejected | HTTP 415, or a 400 reporting an unparsed JSON body, that survived retries and every request shape | | not-found | HTTP 404, or a 400 saying Unknown or unsupported model — listed but not provisioned for this account | | not-chat | HTTP 400 saying the model takes only image/audio/document input, or is a realtime model the text endpoint refuses | | gone | HTTP 410 — the provider retired the model | | unreachable | Transport, TLS, or timeout failure (after one timeout escalation) | | error | Any other upstream failure |

Ranking score (0–6): tool calling 3, tool-result round-trip 2, streaming 1. Results sort by verdict, then score, then latency.

Content-type negotiation

Some gateways fail to parse a bare application/json body and only accept the parameterised form. model-probe sends the bare type first and, on that rejection, rotates the whole run through four ordered {content-type, accept} shapes and replays — without spending a retry. The chosen shape is reported on stdout and as endpoint.contentType / endpoint.shapesTried in --json.

The rejection does not always look like a media-type error. Two real forms seen on gateways here:

415  {"code":"FST_ERR_CTP_INVALID_MEDIA_TYPE","message":"Expected request with `Content-Type: application/json`"}
400  {"fieldErrors":{"messages":["Invalid input: expected array, received undefined"]}}

The second is the dangerous one: a 400 complaining that messages is missing, for a request that definitely sent messages. Keying only on 415 left an entire 259-model gateway reporting every model as error. A 400 whose body reads like an unparsed body now counts as recoverable too.

Negotiation is used rather than hardcoding the parameterised form because most gateways accept the bare type, and the tool should not change the request shape unless it has to.

The shape that works is locked in. As soon as a request with a body parses, that shape becomes the run's preferred shape and every later request starts from it — rotation never moves past it. Only a request that actually carried a body may vote: a bodyless GET /models says nothing about how a body must be encoded, and letting it vote pins the run to the wrong shape and fails every later POST. If the preferred shape later starts being rejected, the preference is dropped and the shapes are re-probed rather than leaving the run permanently stuck.

Cold starts

Serverless backends scale to zero, so the first request to a model pays its load time. A model that looks dead at a 45s timeout can answer at 90s — one gateway's cold model needed 86s, and reporting it as unreachable was a false negative that hid a working model.

Timed-out requests are therefore retried once with the timeout doubled (floored at 120s, capped at 300s) before a model is called unreachable. Recoveries and non-recoveries are counted separately, because "escalated" and "recovered" are very different things:

COLD STARTS: 6 request(s) timed out and answered once the timeout was raised.
Serverless backends scale to zero, so a slow first answer is not a dead model.

2 request(s) still did not answer after the timeout was raised — those models are
hanging or far too slow, not cold-starting. Raise --timeout further to keep waiting.

This applies to the reach probe — the first contact, which is the call that pays the cold start — not to every stage.

Billing blocks

HTTP 402 does not mean the model is broken, so it is reported separately rather than as a generic error — otherwise you go hunting for a capability problem that does not exist. Different model families are gated differently, so the reasons are grouped and counted:

BLOCKED BY BILLING — the key is valid, the balance/credit tier is not:
   12x  Insufficient balance
    9x  Daily check-in required to use free models. Please visit ... to check in.
    3x  Anthropic models are not available with the free 500,000 Telegram bonus credits.

A billing failure does not always arrive as 402. Some gateways return it as HTTP 429 with a quota message, e.g.:

{"error":{"code":"1113","message":"Insufficient balance or no resource package. Please recharge."}}

Classifying that as rate-limited is actively harmful in two ways: it tells the user to slow down and retry, which can never work, and it makes the client pace and retry a permanent block — one gateway went from ~0.4s per model to 31–70s per model. So a 429 is classified by its body: a balance/quota message becomes payment-required, and a genuine throttle message ("Too many requests after rate-limit", "Global rate limit exceeded") stays rate-limited. Billing-classified responses are never retried or paced.

summary.billing[] carries the same data in --json. A tier that needs a daily check-in is a five-second fix rather than a dead end, and the tool tells you which is which.

Alias groups

Gateways often serve many model names from one backend. The report groups them, so you learn that gpt-5.1 is really gpt-5 — and there are two independent sources for this:

  • Observed routing — the SERVED BY column, taken from the model name the gateway echoes back. This catches aliases that the catalog never admits to.
  • Declared aliases — some catalogs ship an aliases array stating the other ids the same model answers to (qwen3.8-max == qwen/qwen3.8-max). The report prints these under DECLARED ALIASES, which is useful when writing config, where the exact id form matters.

Because catalog metadata is this valuable, --model still consults /models rather than skipping it — otherwise a named model would lose its declared aliases, context window, and tool support.

JSON policy

--json writes a single stable envelope to stdout. Diagnostics, warnings, and the progress line go to stderr, so stdout stays pipeable.

Success (probe):

{
  "ok": true,
  "command": "probe",
  "version": "1.0.0",
  "generatedAt": "2026-01-01T00:00:00.000Z",
  "endpoint": {
    "baseUrl": "https://gateway.example/v1",
    "apiBase": "https://gateway.example/v1",
    "apiBaseResolved": true,
    "baseUrlSource": "flag",
    "apiKeySource": "env",
    "authStyle": "bearer"
  },
  "summary": {
    "catalogCount": 355,
    "probed": 299,
    "skipped": 56,
    "counts": { "agent-ready": 5, "chat-only": 11, "denied": 276 },
    "agentReady": 5,
    "recommended": [{ "model": "openai/gpt-oss-20b", "score": 6, "latencyMs": 519, "served": "openai/gpt-oss-20b" }],
    "aliasGroups": [{ "served": "gpt-5", "aliases": ["gpt-5", "gpt-5.1"], "verdict": "chat-only" }]
  },
  "skipped": [{ "id": "mistral-embed", "reason": "non-chat:embed" }],
  "models": [
    {
      "model": "openai/gpt-oss-20b",
      "served": "openai/gpt-oss-20b",
      "reach": { "ok": true, "status": 200, "latencyMs": 519, "error": null, "message": null },
      "capabilities": { "tools": true, "toolChoiceRequired": false, "roundtrip": true, "streaming": true, "doneSentinel": true },
      "detail": {},
      "toolCall": { "id": "call_1", "name": "get_weather", "arguments": "{\"city\":\"Tehran\"}", "origin": "tool_calls" },
      "verdict": "agent-ready",
      "score": 6
    }
  ]
}

Error (any command):

{ "ok": false, "error": { "code": "catalog_unavailable", "message": "GET .../models failed: HTTP 401" } }

Error codes: invalid_option, missing_base_url, catalog_unavailable, empty_selection, confirmation_required, write_not_allowed, bad_data, unhandled.

list --json returns { ok, command, apiBase, count, ids[], skipped[] }. request --json returns { ok, command, url, method, status, latencyMs, error, body }. doctor --json returns { ok, checks[], auth, endpoint, ... } — with the key masked, never in full.

The API key never appears in any output or report file. A test asserts this.

Exit codes

| Code | Meaning | |---|---| | 0 | Command ran successfully | | 1 | Error (unreachable endpoint, failed request, failed doctor check) | | 2 | Usage error (bad option, missing endpoint, confirmation required) | | 4 | --fail-if-none and no model was agent-ready |

--fail-if-none makes probe usable as a CI gate.

Auth

Precedence, highest first:

  1. --base-url / --api-key flags
  2. MODEL_PROBE_BASE_URL / MODEL_PROBE_API_KEY environment variables
  3. ~/.model-probe/config.json (written by init, override path with MODEL_PROBE_CONFIG)

Flags leak into shell history and process listings — prefer the env vars or init for regular use.

Non-bearer gateways: --auth-style x-api-key|api-key|none, or send anything with --header "X-Api-Key: ...". An explicit --header always wins over the resolved key.

Self-signed endpoints: --insecure disables TLS verification and says so on stderr.

Emitting a provider config

model-probe probe --model openai/gpt-oss-20b --emit-config --provider-id mygateway

Prints a ready-to-paste MiMoCode / OpenCode provider block for the top-ranked model, with apiKey left as <YOUR_API_KEY> so the block stays safe to share or commit.

What this does not prove

  • Not a benchmark. It measures protocol compatibility, not answer quality. A model can be agent-ready and still be bad at coding.
  • Point-in-time. Gateways change routing. Re-run when behavior looks odd.
  • No cost or limit data. Context window, output limits, and pricing are not published by /v1/models, so they are deliberately not guessed.
  • A single probe run can be unlucky. One 504 observed in one run does not mean a model is broken; check --verbose stage detail and re-run.

Development

npm test        # node:test — 127 tests, no network to the outside world

Tests cover the pure decision logic — tool-call detection (including the "prose mentions tools" and "echoed empty array" false positives), verdict classification, scoring and ordering, catalog filtering, alias grouping, billing grouping, CSV escaping, and the provider-config shape — plus a real local HTTP server that proves the 415 content-type promotion replays correctly, sticks for later requests, and stays bounded when the gateway never accepts anything.