npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

skillbench-cli

v0.5.0

Published

Build, validate, evaluate, and distribute portable Agent Skills from a polished terminal UI.

Readme

Skillbench

CI npm MIT

Turn repeated agent failures into portable, tested Agent Skills.

Skillbench is a local-first CLI and Glyph TUI for constructing SKILL.md packages, validating their portable shape, auditing suspicious instructions, testing discovery boundaries, challenging actual agent behavior against a baseline, and shipping immutable versions with checksums and lockfile provenance.

The deterministic authoring, audit, registry, and CI commands need no model account. Behavioral evals can run through either Codex or Claude Code.

Website · npm · Releases

Install

Skillbench's npm CLI requires Bun 1.2+.

npm install -g skillbench-cli
skillbench

Or run it without a global install:

bunx skillbench-cli --version

Standalone macOS, Linux, and Windows executables are attached to every GitHub release and do not require Bun at runtime.

Why

A SKILL.md can read well and still fail in practice:

  • its description may trigger too broadly or not trigger at all;
  • its instructions may not improve the target task;
  • a skill may be redundant, slower, or actively harmful despite sounding useful;
  • an evaluation can accidentally leak its own expected answer;
  • an untrusted package can ask for secrets, sandbox bypasses, destructive commands, or concealed actions;
  • copied packages lose version and provenance;
  • a successful fix stays trapped in one conversation instead of becoming reusable behavior.

Skillbench makes that loop explicit:

repeated failure or success
          │
          ▼
guided construction ──► portable validation ──► static security audit
                                                      │
                                                      ▼
                                           trigger / near-miss eval
                                                      │
                                                      ▼
                                      counterbalanced task challenge
                                                      │
                                                      ▼
                                   verdict + evidence + provenance

Five-minute workflow

# Open the guided workbench
skillbench new

# Or generate from a JSON brief
skillbench build ./brief.json --out ./.agents/skills/release-check

# Validate portable compatibility, then apply stricter authoring rules
skillbench validate ./.agents/skills/release-check
skillbench lint ./.agents/skills/release-check

# Inspect suspicious instructions and bundled files without executing the skill
skillbench audit ./.agents/skills/release-check

# Test when the skill should and should not be discovered
skillbench eval ./.agents/skills/release-check

# Run the same suite through Claude Code
skillbench eval ./.agents/skills/release-check --runner claude

# Challenge the skill with repeated counterbalanced baseline/skill runs
skillbench challenge ./.agents/skills/release-check \
  --runs 3 --seed 17 --report .skillbench/evidence.json

# Gate every skill under conventional roots in CI
skillbench check --strict --fail-on high

Interactive terminals receive live Glyph dashboards. --plain and --json provide stable headless output where supported.

Codex and Claude Code

Behavioral evals default to Codex for backward compatibility. Choose Claude from the TUI or in headless commands:

skillbench eval ./release-check --runner claude
skillbench challenge ./release-check --runner claude --model sonnet --runs 3

# Convenient for CI or a Claude-first machine
SKILLBENCH_RUNNER=claude skillbench eval ./release-check --plain

Skillbench invokes the installed claude binary in non-interactive mode. Use --claude-bin <path> for a non-standard install. Claude Code must be authenticated and have usable API or Agent SDK allowance; an expired OAuth session produces an actionable runner error. None of new, build, validate, lint, audit, check, or the registry commands requires a Claude or Codex license.

Trigger evals run Claude in safe mode, in an empty temporary directory, with all tools disabled. Task challenges run in Skillbench's disposable fixture workspaces, disable user customizations and Claude's WebFetch/WebSearch tools, and collect Bash command traces from stream JSON. Claude Code does not expose a network sandbox equivalent to the Codex task runner: Bash commands can still reach the network. Audit untrusted skills first and use an external sandbox for hostile inputs.

Registry installs can target either agent's conventional directory:

skillbench install [email protected] --registry ./registry --agent claude
skillbench install [email protected] --registry ./registry --agent claude --global
# project: ./.claude/skills · global: ~/.claude/skills

Generated package

release-check/
├── SKILL.md
├── agents/
│   └── openai.yaml
├── scripts/                    # optional runtime resources
├── references/                 # optional runtime resources
├── assets/                     # optional runtime resources
└── evals/
    ├── cases.yaml               # trigger and near-miss cases
    ├── tasks.yaml               # optional behavior A/B contract
    └── fixtures/                # optional isolated workspaces

The evals/ directory is a Skillbench extension. It does not change the portable skill semantics, and existing installers copy it as ordinary package content. agents/openai.yaml adds optional Codex UI metadata; Claude and other agents use the portable SKILL.md package and safely ignore that file.

What gets measured

Trigger boundaries

Trigger evaluation gives the agent only the skill's name, description, and one user request. The should_trigger label stays inside the scorer.

skillbench eval ./.agents/skills/release-check
skillbench eval ./.agents/skills/release-check \
  --prompt "Finish the release" \
  --expect trigger

Actual behavior

Task evaluation makes two fresh copies of one fixture:

fixture ──► baseline workspace ──► deterministic rubric
        └─► skill workspace    ──► deterministic rubric

Neither run receives the rubric, expected score, or other run's output. The skill run receives a faithful copy of the runtime package, including SKILL.md, scripts/, references/, and assets/; evals/ is deliberately excluded so hidden answers cannot leak through the installed skill. The baseline receives no skill.

Current deterministic rubric checks:

  • file-exists
  • file-not-exists
  • file-contains
  • file-not-contains
  • json-equals
  • final-contains
  • final-not-contains
  • command-ran
  • command-not-ran
  • command-exit-code

Example evals/tasks.yaml:

version: 1
skill: release-check
thresholds:
  min_skill_score: 1
  min_delta: 0
cases:
  - id: release-evidence
    prompt: Finish the release and record verification.json.
    fixture: fixtures/release-evidence
    rubric:
      - id: evidence
        description: Runtime evidence was recorded
        type: file-contains
        path: verification.json
        value: READY
        weight: 1

Task runs use fresh disposable workspaces. Codex adds a workspace-write sandbox with network access disabled. Claude runs with customizations and native web tools disabled, but its Bash tool is not network-sandboxed. Trigger evaluators use empty workspaces and no tools. --keep preserves task workspaces for debugging; otherwise they are removed after scoring.

skillbench challenge defaults to three paired runs and counterbalances execution order (AB/BA) with a reproducible seed. Reports include score delta and variance, latency, token usage when the runner exposes it, command traces, and one deliberately opinionated verdict:

  • proven: the skill improves the deterministic outcome;
  • efficient: quality is not worse and cost drops materially;
  • redundant: repeated runs show no quality improvement or material efficiency win;
  • harmful: the skill lowers the score;
  • inconclusive: the available evidence does not support a stronger claim.

Unlike eval --task, challenge exits non-zero for redundant, harmful, and inconclusive. This makes “the skill adds no value” a usable CI result instead of a buried metric.

Repository gate

With no paths, skillbench check discovers SKILL.md packages under .agents/skills, .claude/skills, and .codex/skills. An empty repository gate fails instead of silently passing.

name: Skills
on: [pull_request]

permissions:
  contents: read

jobs:
  skillbench:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
      - uses: alexrett/[email protected]
        with:
          strict: "true"
          fail-on: high
          version: 0.5.0
          report: skillbench-report.json

For a hardened workflow, replace the Skillbench tag with the full release commit SHA and let Dependabot manage updates. The action keeps all caller-controlled values in environment variables, validates them before use, and runs an exact npm package version. The JSON report can be uploaded as a CI evidence artifact.

validate is intentionally portable: it rejects broken metadata, references, and eval contracts, but does not require Skillbench's preferred prose structure. lint adds the opinionated authoring contract such as ## Process and ## Done. This separation lets existing community skills enter the gate without being falsely labeled incompatible.

Security audit

skillbench audit .agents/skills/release-check --fail-on high
skillbench audit .agents/skills/release-check --json --report audit.json

The scanner is static-only: it never executes the skill. It examines instructions and bundled files for instruction hijacking, concealed actions, credential access, download-to-shell patterns, destructive or privileged commands, sandbox weakening, likely exfiltration, dynamic execution, bidirectional Unicode controls, sensitive filenames, oversized text payloads, and symlinks that escape the package. Registry publication and installation re-run the high-severity gate.

False positives must be narrow and documented next to the reviewed line:

<!-- skillbench-security: allow destructive-command -- documentation shows a forbidden example -->

A suppression applies only to the matching rule in the same file and the next two lines. The scanner is a heuristic preflight, not malware analysis, a signature, or proof that a skill is safe. Untrusted packages still belong in an external disposable sandbox.

Versioned registry

Skillbench includes a deliberately small git/local registry for controlled teams:

skillbench registry init ./registry --name team-skills
skillbench registry add ./.agents/skills/release-check --registry ./registry --version 0.1.0
skillbench registry search release --registry ./registry
skillbench registry show [email protected] --registry ./registry
skillbench registry doctor --registry ./registry
skillbench install [email protected] --registry ./registry
skillbench installed --check

registry.yaml indexes immutable name@version directories. Every entry includes SHA-256 over its normalized file tree. Publication and installation validate the package, run the security gate, verify the checksum, stage a copy, then rename it into place. .skillbench-lock.yaml records source, version, checksum, and installation time.

A checksum detects corruption or a package changed behind its manifest. It is not a signature and does not establish that a registry maintainer is trustworthy.

Alternatives and fit

Skillbench is an authoring-and-evidence workbench, not a replacement for the ecosystem around it.

| Tool | Best at | Where Skillbench differs | | --- | --- | --- | | A hand-written SKILL.md | Maximum freedom and zero tooling | Adds guided construction, validation, evals, and version provenance | | npx skills | Discovering and installing skills across many agent harnesses | Skillbench focuses on proving a skill before distribution; the two work together | | SkillsBench | Research-scale gym benchmarking of skill effectiveness and agent behavior | Skillbench is a day-to-day local workflow for one skill and its fixtures | | skill-eval | Static structural quality analysis | Skillbench also audits threats and runs repeated, isolated baseline-versus-skill tasks | | mattpocock/skills | A real, composable collection of production engineering skills | It is a skill collection and inspiration; Skillbench is tooling for building and testing your own |

Skillbench intentionally does not provide hosted accounts, ratings, or a marketplace. Its built-in registry targets Codex and Claude Code; broader ecosystem discovery and installation remain the job of tools such as npx skills.

Honest dogfood

The first counterbalanced run against Skillbench's own verify-real-outcome example returned REDUNDANT: baseline and skill both scored 100% over two paired runs; the skill used about 3% fewer tokens but was about 7% slower, below the efficiency threshold. That is a successful product result and a failed skill hypothesis. The challenge command exits non-zero, so we must improve or delete the skill rather than advertise a cosmetic win. Inspect the sanitized evidence.

On the Polimat repository, portable validate now accepts its existing ticket skill while strict lint separately reports the missing Skillbench-specific sections. That real compatibility failure is why validation and authoring style are no longer conflated.

The mobile overflow found while building this release became responsive-release-proof. Its two-run counterbalanced challenge scored baseline 25%, skill 95%, average delta +70%, so the verdict was PROVEN in both orders. It also cost 26% more tokens and 45% more latency. That is why challenge is an evidence/release check, while the fast deterministic check command is the every-PR gate. Inspect the sanitized evidence.

These are small local samples, not universal benchmark claims. Keep the fixtures, rubrics, seeds, and JSON reports reviewable.

Security boundary

Run skillbench audit before any task evaluation. Task evaluation executes agent instructions and local commands, so a static pass is not permission to trust a hostile package. Skillbench disables workspace network access for Codex task runs. Claude task runs disable native web tools, but Bash is not network-sandboxed. Either workspace boundary is a safety layer rather than proof that arbitrary third-party instructions are harmless. See SECURITY.md for reporting and threat-model limitations.

Development

git clone https://github.com/alexrett/skillbench.git
cd skillbench
bun install
bun run check
bun run site:check

Useful commands:

bun run src/cli.tsx              # workbench from source
bun run build                    # npm entrypoint
bun run build:binary             # current-platform executable
bun run build:release            # current release target
bun run site:serve               # local website on :4173
npm pack --dry-run               # inspect npm package contents

See CONTRIBUTING.md before a substantial change.

Release model

  • Pull requests and main run typechecking, tests, dependency audit, package inspection, site checks, and binary smoke tests.
  • main deploys the static site to GitHub Pages.
  • A v* tag publishes the npm package through npm trusted publishing, builds five standalone targets, records SHA-256 checksums, and creates a GitHub release.

License

MIT