npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

skills-evals

v0.2.0

Published

Validate, trigger-test, and regression-test your agent artifacts — Claude/Copilot skills (SKILL.md), Copilot instructions & custom agents, Claude custom agents, Cursor rules, and prompt files.

Readme

skills-evals

Know when your agent skills stop working.

skills-evals validates, trigger-tests, and regression-tests every agent artifact in your repo — Claude/Copilot skills (SKILL.md), Copilot instructions and custom agents, Claude custom agents, Cursor rules, and prompt files — so you find out in CI when a skill stops triggering or behaving as intended after your codebase (or the skill) changed.

Zero dependencies. Node ≥ 18.17. Compatible with Anthropic skill-creator's evals.json schema.

📖 Full documentation: ahnafyy.github.io/skills-evals (or browse docs/)

Quick start

npx skills-evals list                    # see what it found in your repo
npx skills-evals init                    # scaffold eval case stubs
npx skills-evals run                     # tiers 1+2 (free, deterministic, CI-safe)
npx skills-evals run --update-baseline   # snapshot for regression detection — commit it
npx skills-evals behavioral my-skill     # tier 3 — a real agent proves the skill still works

Or let an agent set it up for you

This repo ships an installable setup-skills-evals skill that walks any agent (Claude Code, GitHub Copilot, Cursor, …) through the whole setup — it inventories your artifacts, asks what you want to test, fills in the eval cases from your answers, and wires up CI plus your local runner (npm, Gradle, Make, or a .sh).

It's a native Agent Skill (open standard) — no CLI or middleman. Because it sets up other repos, install it once, globally, so it's available everywhere:

# Claude Code — personal skill, works in every repo
curl -fsSL https://raw.githubusercontent.com/ahnafyy/skills-evals/main/skills/setup-skills-evals/SKILL.md \
  --create-dirs -o ~/.claude/skills/setup-skills-evals/SKILL.md
# GitHub Copilot: same file at ~/.copilot/skills/setup-skills-evals/SKILL.md

Open any repo and ask your agent: "set up skills-evals in this repo". To commit it into a single project instead (and the exact path per agent), see skills/setup-skills-evals/.

What it discovers

| Format | Files | Routing | | --- | --- | --- | | Skills (Claude & Copilot) | **/SKILL.md | description | | Copilot instructions | .github/copilot-instructions.md, **/*.instructions.md, AGENTS.md | always / applyTo globs | | Copilot custom agents | **/*.agent.md, .github/agents/*.md | description | | Claude custom agents | .claude/agents/*.md | description | | Cursor rules | .cursor/rules/**/*.mdc | alwaysApply / globs / description | | Prompt files | **/*.prompt.md | manual |

The three tiers

| Tier | What it checks | Where | Cost | | --- | --- | --- | --- | | 1. Structural | Frontmatter, naming, description limits, "use when" triggers, glob validity | CI (validate, run) | Free | | 2. Trigger & routing | Positive prompts rank their artifact top-k; negative prompts don't; no two descriptions near-collide | CI (run) | Free | | 3. Behavioral | An agent following the artifact satisfies its expectations[] | Scheduled CI (behavioral) | A cheap model run |

Tiers 1 and 2 are free and deterministic, so they gate every PR. Tier 3 is where the real value is — it's the only tier that proves an agent following your skill actually does the right thing, and it belongs on a schedule (e.g. nightly/weekly) so drift in the underlying codebase or model surfaces on its own. Point it at a cheap, fast model: most of the signal is did the agent take the right actions, which small models judge fine.

Tier 2 is a deterministic lexical approximation of routing (stemmed TF-IDF over descriptions, ranked within each kind's pool — skills compete with skills, Claude agents with Claude agents). It can't judge semantics — that's Tier 3's job — but it catches the two failure modes that dominate real trigger bugs: a description missing the vocabulary users actually say (false negative), and an over-broad description that outranks the right artifact (false positive). Glob-routed artifacts (applyTo, Cursor globs) are tested with file paths instead of prompts.

Regression detection (the point of all this)

skills-evals run --update-baseline   # snapshot; commit .skills-evals/baseline.json
skills-evals run                     # every later run diffs against the snapshot

The baseline stores every trigger outcome, collision pairs, the rank-1 rate, and a content hash of each artifact. On later runs:

  • a previously-passing trigger that now fails on an unchanged artifact → error (catalog drift: another artifact now wins that routing);
  • the same failure on a changed artifact → warning (expected churn — review, then re-baseline);
  • a new description collision → error;
  • a rank-1 rate drop → warning.

So when a teammate adds a new skill whose description hijacks your skill's prompts — or an edit quietly breaks routing — CI fails with the exact prompt that regressed.

Eval case format

One file per artifact: evals/cases/<name>.json. The evals[] array is Anthropic skill-creator's evals.json schema verbatim, so its tooling works against these files unmodified.

{
  "artifact": "test-driven-development",
  "kind": "skill",
  "trigger": {
    "positive": [
      { "prompt": "Write a failing test for this bug before fixing it", "top_k": 1 }
    ],
    "negative": [
      { "prompt": "Draft a commit message for these changes", "owner": "commit-messages" }
    ]
  },
  "evals": [
    {
      "id": 1,
      "prompt": "Fix the reported rounding bug in the invoice totals, test-first.",
      "expected_output": "A failing test demonstrating the bug, a minimal fix, full suite passing",
      "files": ["invoice-app/"],
      "expectations": [
        "A failing test is written and shown failing before the fix",
        "The implementation is the minimum needed to pass"
      ],
      "trust_level": "provisional"
    }
  ]
}
  • positive prompts are realistic user asks that should route to this artifact (top_k defaults to 3; tighten to 1 for a signature ask). Don't copy the description — that games the eval.
  • negative prompts belong to a different artifact. Declaring that artifact in owner turns the negative into a real pairwise routing test: the owner must outrank this artifact, preventing vacuous passes.
  • For glob-routed artifacts use { "path": "src/app/index.ts" } entries instead of prompts.
  • trust_level: "provisional" marks a behavioral eval without fixtures; its results are a sanity check, not evidence.

Behavioral evals (Tier 3)

This is the tier that matters most: a real agent follows the artifact and a grader checks it did the right thing. Run it on a schedule so codebase and model drift surface on their own, and point it at a cheap, fast model to keep it running often.

skills-evals behavioral test-driven-development --dry-run   # print the plan, no model call
skills-evals behavioral test-driven-development             # execute + grade
skills-evals behavioral my-skill --adapter copilot --grader claude

Each eval runs in a throwaway workspace (fixtures from files[] materialized out of evals/fixtures/), captures the execution trace, and grades the trace — not the model's final prose — against expectations[]. The trace is fenced as untrusted data in the grader prompt and piped over stdin; grader output is validated as JSON before being written to .skills-evals/results/ (gitignored) in skill-creator's grading.json shape.

Executor adapters

| Adapter | Binary | Trace | Fidelity | | --- | --- | --- | --- | | claude (default) | claude | stream-json with tool calls | Full — grades what the agent did | | copilot | copilot | final response text | Degraded — grades conservatively | | cursor | cursor-agent | final response text | Degraded — grades conservatively |

Register your own:

const { registerAdapter } = require('skills-evals');
registerAdapter({
  name: 'my-runtime',
  traceKind: 'text',
  run({ prompt, systemPrompt, cwd, tools, timeoutMs }) { /* return trace */ },
  judge(graderPrompt, { timeoutMs }) { /* return raw grading */ },
});

GitHub Action

- uses: ahnafyy/[email protected]
  with:
    root: .
    # optional — run tier 3 in CI (needs the adapter CLI + credentials):
    # behavioral: my-skill, my-other-skill
    # adapter: claude|copilot|cursor   (default: config file, then claude)

Fails the job on any error-level finding, including baseline regressions. Tiers 1+2 need no secrets; see CI docs for adapter selection.

Programmatic API

const {
  discover, validateCatalog, loadCases, runTriggerEvals,
  buildBaseline, loadBaseline, diffBaseline, runBehavioral, loadConfig,
} = require('skills-evals');

const config = loadConfig(process.cwd());
const artifacts = discover(config.root, { exclude: config.exclude });
const tier1 = validateCatalog(artifacts);
const tier2 = runTriggerEvals({ artifacts, cases: loadCases(config.casesDir), config });

Configuration

Optional skills-evals.config.json at the repo root:

{
  "casesDir": "evals/cases",
  "fixturesDir": "evals/fixtures",
  "topK": 3,
  "collisionWarn": 0.5,
  "collisionError": 0.75,
  "minPositive": 3,
  "minNegative": 2,
  "minEvals": 1,
  "exclude": ["examples/**"],
  "behavioral": {
    "adapter": "claude",
    "tools": "Read,Glob,Grep,Edit,Write,Bash"
  }
}

Documentation

The full docs live at ahnafyy.github.io/skills-evals — a static site rendered straight from the markdown in docs/:

Prior art

What skills-evals adds: multi-format artifact discovery (Copilot + Claude + Cursor), deterministic TF-IDF trigger routing with per-kind pools, path-based trigger tests for glob-routed artifacts, trace-graded behavioral evals with pluggable executor adapters, and baseline snapshots with change-aware regression diffing.

License

MIT