npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@gbesse/jev-eval

v0.1.0

Published

Reproducible evaluation, calibration, threshold optimization, and CI regression gates for typed probabilistic decision models.

Readme

@gbesse/jev-eval

Reproducible evaluation and calibration for System One-compatible decision models.

What it measures

  • Accuracy and accepted accuracy after confidence gating.
  • Coverage: the fraction of decisions accepted automatically.
  • Binary or multiclass Brier score from the full probability distribution.
  • Expected calibration error (ECE).
  • Mean, p50, and p95 end-to-end latency.
  • Metrics sliced by question type and user-defined tags.
  • A cost-optimal review threshold for configurable false-decision and human-review costs.

The runner writes every completed case to results.jsonl before starting another. Interrupted runs resume by case id. Labels remain in the evaluator and are never included in provider requests.

Install and run

npm install @gbesse/jev-eval
jev-eval validate --config eval.yaml
jev-eval run --config eval.yaml

Use provider.type: typesafe for the official endpoint, http for a compatible gateway, or recorded for deterministic offline fixtures. With TypeSafe, set TYPESAFE_API_KEY; the key name is configurable with apiKeyEnv.

See the repository examples/eval.yaml and examples/data/eval.json. A dataset is a JSON array or JSONL file. Each case contains id, state, questions, expected, and optional tags, group, or recordedResponse.

CI gates

gates can require minimum accuracy, coverage, and accepted accuracy or maximum Brier and ECE. A failed gate exits with status 1. Provider errors also fail the gate instead of disappearing from the denominator.

gates:
  minAccuracy: 0.90
  minAcceptedAccuracy: 0.98
  maxEce: 0.08

report.json is stable, machine-readable output. report.html is a standalone human report. Use jev-eval compare to calculate candidate-minus-baseline changes.

Production methodology

Keep near-duplicates and paraphrases in the same group when splitting datasets. Pin the resolved model version. Evaluate option-order rotations and adversarial text as separate cases. Choose review costs before inspecting the held-out test set. A low ECE on a small dataset is not evidence of general calibration.