npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

smallm

v0.5.0

Published

Find the right small language model (SLM) for a task, based on given context. Live rule-based matching against the HuggingFace API.

Readme

smallm

Find the right small language model (SLM) for a task, based on given context.

smallm does a live lookup against the HuggingFace public API and runs a rule-based scoring engine (task match, size fit, context window, domain tag) to rank candidates — no embeddings, no local registry, no config files.

Install

npm install smallm

Usage

import { findModels } from "smallm";

const results = await findModels({
  task: "chat",           // "summarize" | "classify" | "extract" | "chat" | "code" | "translate" | string
  contextLength: 8192,    // required — hard filter, in tokens
  maxParamsB: 7,           // optional — hard filter, "under 7B"
  domain: "medical",       // optional — used in scoring only
  hardware: "gpu-low",     // optional — required (along with maxLatencyMs) to enforce latency
  maxLatencyMs: 2000,       // optional — v0.2: real hard filter when benchmark data exists
  scoringMode: "rule",       // optional — v0.3: "rule" (default) | "embedding" | "hybrid"
  limit: 5,                   // optional — defaults to 5
});

console.log(results);
// [
//   {
//     name: "org/some-7b-chat-model",
//     provider: "huggingface",
//     paramsB: 6.5,
//     contextWindow: null,          // see "Known limitations" below
//     reasonWhy: "Strong task match, sufficient context window, and comfortably within your parameter limit.",
//     score: 87,
//     scoringMode: "rule"
//   },
//   ...
// ]

Multi-provider (v0.4)

// Query a locally running Ollama installation instead of (or alongside) HuggingFace.
await findModels({ task: "chat", contextLength: 4096, providers: ["ollama"] });
await findModels({ task: "chat", contextLength: 4096, providers: ["huggingface", "ollama"] });

// Structured hardware — widened, not replacing the original string enum.
await findModels({ task: "chat", contextLength: 4096, hardware: { type: "gpu", vramGB: 8 } });
await findModels({ task: "chat", contextLength: 4096, hardware: "gpu-low" }); // still works

Prerequisite for providers: ["ollama"]: Ollama must be running locally (ollama serve, or the desktop app) and reachable at http://localhost:11434. If it isn't, smallm doesn't throw or crash the whole call — it logs a warning and simply returns no Ollama candidates, so a mixed ["huggingface", "ollama"] query still returns your HuggingFace results.

Important: HardwareSpec.vramGB filtering uses an estimated footprint (~2GB VRAM per 1B params, a common fp16 rule of thumb) — not a measured or authoritative figure. It's meant to rule out obviously-too-large models, not to guarantee something fits.

Scoring modes (v0.3)

// Default — identical to v0.1/v0.2 behavior. Matches on exact task/pipeline_tag mapping.
await findModels({ task: "chat", contextLength: 4096, scoringMode: "rule" });

// Scores purely on text similarity between your task string and each model's
// available text (id, pipeline_tag, tags) — useful for free-text tasks that
// don't map neatly onto the fixed task enum.
await findModels({ task: "pull key dates and amounts from scanned invoices", contextLength: 4096, scoringMode: "embedding" });

// Blends both: hybridScore = ruleScore * 0.60 + embeddingScore * 0.40 (locked weights — rule score still dominates).
await findModels({ task: "chat", contextLength: 4096, scoringMode: "hybrid" });

Multi-provider + structured hardware (v0.4)

// Query Ollama's locally installed models instead of (or alongside) HuggingFace.
// Requires a running local Ollama installation — see "Ollama prerequisite" below.
await findModels({ task: "chat", contextLength: 4096, providers: ["ollama"] });

await findModels({ task: "chat", contextLength: 4096, providers: ["huggingface", "ollama"] });

// Structured hardware spec — widens (doesn't replace) the original string enum.
await findModels({
  task: "chat",
  contextLength: 4096,
  hardware: { type: "gpu", vramGB: 8 }, // excludes models estimated to not fit
});

// Old string-enum form still works exactly as before:
await findModels({ task: "chat", contextLength: 4096, hardware: "gpu-low" });

Ollama prerequisite: providers: ["ollama"] (or including "ollama" in a multi-provider list) expects a local Ollama installation running at http://localhost:11434. If it's not running or not installed, smallm doesn't throw or crash the whole query — it logs a warning and simply contributes zero Ollama candidates, so a mixed ["huggingface", "ollama"] query still returns HuggingFace results normally. Ollama itself is never installed, started, or asked to download models by this package — it only queries what's already available locally.

vramGB filtering is a heuristic, not a measured constraint. There's no standard "how much VRAM does model X need" figure available from either provider's listing API, so smallm estimates it from paramsB using a commonly-cited rule of thumb (~2GB VRAM per 1B parameters, roughly fp16 inference). A model quantized to 4-bit would actually need much less — this estimate is conservative/pessimistic in that case. Treat vramGB filtering as a rough guide, not a guarantee.

Handling errors (v0.2)

import { findModels, ValidationError, RateLimitError, HFApiError } from "smallm";

try {
  await findModels({ task: "chat", contextLength: 4096 });
} catch (err) {
  if (err instanceof ValidationError) {
    // your query was malformed — fix it, retrying won't help
  } else if (err instanceof RateLimitError) {
    // HuggingFace rate-limited us — smallm already retried 3x with backoff before this threw
  } else if (err instanceof HFApiError) {
    // some other non-2xx from HuggingFace
  }
}

API

findModels(query: ModelQuery): Promise<ModelMatch[]>

The single public entry point. Validates the query, fetches live candidates from HuggingFace (cached, retrying transient failures), applies hard filters, scores and ranks the survivors, and returns the top N.

ModelQuery

| Field | Type | Required | Notes | |---|---|---|---| | task | "summarize" \| "classify" \| "extract" \| "chat" \| "code" \| "translate" \| string | yes | | | contextLength | number | yes | Hard filter (tokens) | | hardware | "cpu" \| "gpu-low" \| "gpu-high" \| { type: "cpu" \| "gpu"; vramGB?: number } | no | String enum: required alongside maxLatencyMs to enforce latency filtering (v0.2). Object form (v0.4): vramGB enables the estimated-footprint hard filter. | | domain | string | no | Used in scoring, not filtering | | maxParamsB | number | no | Hard filter, e.g. 7 = "under 7B" | | maxLatencyMs | number | no | v0.2: real hard filter when a benchmark entry exists for the (model, string-enum hardware) pair. No entry = not enforced. | | scoringMode | "rule" \| "embedding" \| "hybrid" | no | v0.3: which scoring strategy to use. Defaults to "rule". | | providers | ("huggingface" \| "ollama")[] | no | v0.4: which registries to query. Defaults to ["huggingface"]. | | limit | number | no | Defaults to 5 | | cacheOptions | { dir?: string; ttlMs?: number } | no | v0.2: configure the file-based cache. Defaults: OS temp dir + /smallm-cache, 5-minute TTL |

ModelMatch

| Field | Type | Notes | |---|---|---| | name | string | Provider-specific model id | | provider | "huggingface" \| "ollama" | v0.4: widened from the literal "huggingface" | | paramsB | number \| null | null if size couldn't be detected | | contextWindow | number \| null | null if unknown (see limitations) | | reasonWhy | string | Generated from the scoring breakdown | | score | number | 0–100 | | scoringMode | "rule" \| "embedding" \| "hybrid" | v0.3: which mode produced this result's score |

Errors (v0.2)

All thrown errors extend SmallmError:

  • ValidationError — bad ModelQuery input. Never retried.
  • HFApiError — non-2xx response from HuggingFace (base class; carries .status).
  • RateLimitError — HFApiError subclass specifically for HTTP 429.

findModels retries transient HuggingFace failures (429 and any 5xx) up to 3 times with exponential backoff (500ms / 1000ms / 2000ms) before throwing.

How scoring works

Hard filters run before any scoring — a model that fails a hard filter never gets scored, regardless of scoringMode or which providers it came from:

  • maxParamsB, contextLength, task (MVP)
  • maxLatencyMs when a benchmark entry exists for the (model, string-enum hardware) pair (v0.2)
  • hardware.vramGB when an estimated footprint exceeds it (v0.4, heuristic — see "Known limitations")

Survivors are scored. In "rule" mode (the default) that's four weighted components:

  • Task match — 50%
  • Context window fit — 20%
  • Size fit — 15%
  • Domain match — 15%

"embedding" mode scores purely on text similarity instead; "hybrid" blends both 60/40 (rule/embedding). See "Scoring modes (v0.3)" above.

Unknown metadata (e.g. undetectable param count, no benchmark entry, unreachable Ollama) is never punished or excluded — it gets a neutral mid-range score for that component only, or is simply not filtered on.

Known limitations

  • Embedding mode uses lexical, not semantic, similarity. The v0.3 "embedding"/"hybrid" scoring modes are backed by a small, local, dependency-free character-trigram hashing scheme (see src/embeddings.ts) — not a downloaded neural embedding model. It's genuinely offline and free, but two phrases that are semantically close yet share few characters (e.g. "extract key-value pairs from invoices" vs. "structured data extraction") will score lower than a real embedding model would give them.
  • vramGB filtering is a heuristic, not a measured constraint. Estimated from paramsB at ~2GB/1B params (roughly fp16). Quantized models need much less — this estimate skews conservative in that case. See "Multi-provider + structured hardware (v0.4)" above.
  • Benchmark data is illustrative, not measured. The shipped src/data/benchmarks.json contains a small set of placeholder latency numbers for demonstration and testing — not real measured inference latency. Replace it with genuinely benchmarked data before relying on maxLatencyMs exclusions in production.
  • contextWindow is currently always null for every provider — neither HuggingFace's list endpoint nor Ollama's /api/tags returns it, and the pipeline doesn't do a per-model detail fetch. This means the context-length hard filter rarely excludes anything, and the context-fit score is always neutral. Tracked as a possible future addition.
  • Ollama candidates never have a pipeline_tag, so the task hard filter always lets them through (same "unknown, don't exclude" rule as everywhere else) — task relevance for Ollama results relies entirely on scoring, not filtering.
  • Sub-billion-parameter models (e.g. 125M) aren't detected by the size regex and report paramsB: null.
  • File-based cache only — no database engine. Cache entries are per-process-agnostic (survive restarts) but there's no cross-machine sharing.

CLI (v0.5, optional)

smallm is library-first, but ships an optional CLI as a thin wrapper around findModels(). Every flag maps 1:1 to a ModelQuery field — the CLI contains no scoring, filtering, or provider logic of its own; it only parses flags and formats output. If you can do it via findModels(), you can do it via the CLI, and nothing else.

npx smallm find --task summarize --context 4096
npx smallm find --task chat --context 8192 --max-params 7 --json
npx smallm find --task chat --context 4096 --providers huggingface,ollama --hardware-type gpu --hardware-vram 8

Default output is a table:

NAME                                       PROVIDER     PARAMS(B)  SCORE  MODE       REASON
------------------------------------------ ------------ ---------- ------ ---------- --------------------------------------------------
org/some-7b-chat-model                     huggingface  6.5        87     rule       Strong task match, sufficient context window, and…

Pass --json for the raw ModelMatch[] instead — identical to what findModels() returns programmatically:

npx smallm find --task chat --context 4096 --json | jq '.[0].name'

Flags

| Flag | Maps to | Notes | |---|---|---| | --task <task> | task | Required | | --context <tokens> | contextLength | Required | | --max-params <billions> | maxParamsB | | | --domain <domain> | domain | | | --hardware <cpu\|gpu-low\|gpu-high> | hardware (string form) | | | --hardware-type <cpu\|gpu> + --hardware-vram <gb> | hardware (HardwareSpec form) | Takes precedence over --hardware if both given | | --max-latency <ms> | maxLatencyMs | | | --scoring-mode <rule\|embedding\|hybrid> | scoringMode | | | --providers <list> | providers | Comma-separated, e.g. huggingface,ollama | | --limit <n> | limit | | | --cache-dir <path> / --cache-ttl <ms> | cacheOptions | | | --json | — (output mode only) | Not a ModelQuery field — controls formatting, not what's queried |

Run npx smallm find --help (or with no arguments) to see this from the terminal. Invalid input (a missing --task, a bad --scoring-mode value, etc.) isn't validated by the CLI itself — it's passed straight to findModels(), which throws the same typed errors (see "Handling errors" above) that the CLI then prints cleanly with a non-zero exit code.

Development

npm install
npm run build   # compile TypeScript -> dist/ (also chmods dist/cli.js executable)
npm test         # run the unit + integration test suite
npm link          # optional: try the CLI locally as the `smallm` command

Changelog

See CHANGELOG.md.