npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

skillfid

v0.1.2

Published

Evaluate how faithfully agent skills apply their source documentation

Readme

skillfid

skillfid is a CLI for developers who build agent skills from documentation. It tests whether a skill helps an agent answer questions grounded in that source, revealing missing or imprecise guidance that reviewing the skill alone can miss. Use it to compare skill revisions against a reusable closed-book baseline and identify source-backed improvements.

Canopy skill evaluation report showing an 83.91% skill-assisted score and actionable diagnoses

Explore the Canopy result

Canopy, is a synthetic distributed build cache used to demonstrate the complete workflow.

Inspect the included evaluation and trace its findings to the recorded data without making a model call:

The original skill scored 83.91%, compared with 0% closed book. After source-backed improvements, v2 scored 100% against the same dataset and baseline. Both results used gpt-5.6-sol as subject and judge with three trials per question. Scores are specific to the dataset and evaluation configuration.

Run the complete Canopy flow

Complete the requirements and installation first. Then build an immutable dataset from the Canopy corpus:

skillfid --progress human dataset build \
	--corpus ./examples/canopy/corpus \
	--output-dir ./.work/canopy/datasets \
	--work-dir ./.work/canopy/dataset-work

Set DATASET to the published path printed by the command, then verify it locally:

DATASET=./.work/canopy/datasets/<dataset-id>

skillfid dataset verify --dataset "$DATASET"

Create the reusable closed-book baseline:

skillfid --progress human eval baseline \
	--dataset "$DATASET" \
	--output-dir ./.work/canopy/baselines \
	--work-dir ./.work/canopy/baseline-work

Evaluate the original skill with explicit invocation. The run reuses the compatible baseline automatically:

skillfid --progress human eval run \
	--dataset "$DATASET" \
	--skill ./examples/canopy/skill \
	--baseline-dir ./.work/canopy/baselines \
	--skill-invocation explicit \
	--output-dir ./.work/canopy/runs \
	--work-dir ./.work/canopy/eval-work

Set RUN to the output path, then generate the report locally:

RUN=./.work/canopy/runs/<run-id>

skillfid eval report \
	--run "$RUN" \
	--dataset "$DATASET" \
	--title "Canopy cache operations"

To evaluate the improved skill, rerun eval run with --skill ./examples/canopy/skill-v2. See the Canopy walkthrough for the checked-in artifacts and reproduction paths. If a model-backed command is interrupted, use the resume command printed by the CLI.

Requirements

Before running skillfid, make sure that you have:

  • Node.js 24 or later
  • GitHub Copilot CLI authenticated for Copilot access

Check them with node --version and copilot --version. The default model for both subject and judge is gpt-5.6-sol; you can override either per command. Evaluation runs use an isolated repository-local profile and do not expose authentication tokens to agent tools.

Installation

Install skillfid from npm:

npm install --global skillfid

For a minimal model-backed smoke test, use the one-fact Arbor example.

Command reference

Build an immutable, oracle-calibrated dataset from a Markdown corpus:

skillfid dataset build --corpus ./corpus --json

Every generated question must produce an oracle answer that receives a stable, perfect criterion-level judgment before publication. The same answer is judged independently three times. Unanimous results are accepted; disagreement triggers two more judgments, and only a 4/5 result is accepted. A 3/2 split is unstable and blocks publication.

The dataset stores every judgment and its consensus in calibrations.jsonl; the corpus remains the sole ground truth. Datasets older than schema v6 lack the combined structural and integrity proof and must be rebuilt.

After changing the Copilot runtime, model, judge, or harness, recalibrate without repeating corpus inventory or question extraction:

skillfid dataset recalibrate \
	--dataset ./datasets/<dataset-id> \
	--output-dir ./datasets \
	--json

Recalibration copies documents, knowledge, evidence, questions, verification, coverage, and audit records unchanged. It reruns only the oracle answer and independent consensus judgments for each question. The oracle answer is generated once and held fixed across all judge repeats. Every question must still receive a stable, perfect calibration before publication.

The result is a new immutable dataset whose manifest records sourceDatasetId. Continue interrupted recalibration with the same command and --resume.

Verify a dataset locally without rerunning extraction:

skillfid dataset verify --dataset ./datasets/<dataset-id> --json

Measure the reusable closed-book baseline:

skillfid eval baseline \
	--dataset ./datasets/<dataset-id> \
	--json

Evaluate a skill using the latest exactly compatible baseline:

skillfid eval run \
	--dataset ./datasets/<dataset-id> \
	--skill ./path/to/skill \
	--json

Skill activation is automatic by default. To invoke the discovered project skill as /skill-name in every skill-condition prompt, use explicit invocation:

skillfid eval run \
	--dataset ./datasets/<dataset-id> \
	--skill ./path/to/skill \
	--skill-invocation explicit \
	--json

--skill-invocation accepts auto or explicit. It changes only the skill condition. The resolved mode is recorded in the run manifest; the baseline does not change.

Oracle calibration is a dataset publication gate, not an evaluation condition or model-specific ceiling. Baseline and skill model settings inherit from the dataset by default, but both commands may select another model configuration.

A compatible baseline must match the dataset ID, subject model, judge model, reasoning effort, trial count, evaluator version, and Copilot CLI version exactly. Evaluation fails with an actionable message when none exists.

Both commands default to three trials per question. Use --trials <count> consistently to override that default. When storing baselines outside ./baselines, use matching --output-dir on eval baseline and --baseline-dir on eval run.

Generate a self-contained HTML report:

skillfid eval report \
	--run ./runs/<run-id> \
	--dataset ./datasets/<dataset-id> \
	--title "My skill"

The report defaults to <run>/report.html. Use --output <file> to choose another location. Report generation is local and makes no Copilot calls.

Execution and recovery

Run skillfid --help for the complete command reference, including JSON schemas, prerequisites, and exit codes. Primary output goes to stdout, while progress and errors go to stderr.

Copilot calls have a 600-second timeout and one fresh-session retry by default. Configure them with --timeout <seconds> and --timeout-retries <count>.

Dataset builds and evaluations run independent work concurrently. The default is 10; set --concurrency <count> to any positive integer. Answers and judgments are separate recovery checkpoints, so interruption after answering does not require generating that answer again.

Matching operations reuse completed work by default. Resume selects the latest matching incomplete operation, including one originally started with --fresh. Validated jobs are stored in <work-dir>/operations.sqlite using SQLite WAL and retained after completion.

Use --fresh to start from scratch without deleting earlier state. When an interactive operation is interrupted, the CLI prints the exact resume command, so you do not need to reconstruct it.

Inspect retained operations without modifying the journal:

skillfid operation status --work-dir .work/eval --json
skillfid operation status --work-dir .work/eval --operation-id <id> --json

All commands support --progress auto|human|agent|json|quiet. Auto selects an in-place display on a TTY and bounded agent snapshots otherwise. Human mode shows current work and elapsed time, followed by an estimate and a compact result. JSON progress is emitted as JSON Lines on stderr without changing final stdout.

Artifacts and scoring

Each baseline stores its closed-book answers and judgments with an exact compatibility manifest. Each skill run embeds those baseline records alongside fresh skill answers. You will find the answers in answers.jsonl, criterion-level judgments in judgments.jsonl, failure analysis in diagnoses.jsonl, and aggregate scores in summary.json.

Embedding baseline records keeps reports self-contained without repeating closed-book inference. The summary retains build-time calibration metadata for compatibility; calibration is not an evaluation condition or model call.

HTML reports preserve that separation. Scores come from all recorded trials, and uplift is shown in percentage points. Only diagnoses with concrete file targets appear as recommended work; the evidence view retains every trial answer and its failed-criterion rationale.

Isolation and safety

Every question, trial, and condition runs in a fresh non-resumed Copilot SDK session with its own filesystem workspace. Subject sessions share one client, and judge sessions share another. Conversation and workspace state are not reused.

Evaluation calls deny shell execution, file writes, and URL access. This stops the closed-book baseline from searching external sources while preserving local read access for skill files. Service, authentication, and validation failures surface immediately while completed journal checkpoints remain resumable.

See DESIGN.md for the architecture, scheduler behavior, and scoring model.

Development

Run the tests:

npm test

Ferryline is the synthetic distributed build-cache scenario from GitHub Next's Knowledge Compressor article. Reproduce its calibration from a repository checkout; the extraction script requires internet access to fetch the article:

npm run calibration:extract
npm start -- dataset build \
	--corpus .work/js-calibration/article/corpus \
	--output-dir .work/js-calibration/datasets \
	--json
npm run calibration:compare -- \
	.work/js-calibration/datasets/<dataset-id> \
	.work/js-calibration/article/reference-questions.json

Support and contributing

Report defects and request features in GitHub Issues. Read CONTRIBUTING.md before opening a pull request.

License

Licensed under the MIT License.