npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@jev-harness/eval

v0.5.0

Published

Evaluation and calibration toolkit for TypeSafe's Jev (System One) decision model.

Readme

@jev-harness/eval

Evaluation and calibration toolkit for TypeSafe Jev — the System One decision model. The README of @jev-harness/core says it plainly: calibration is not correctness, so validate on your own labeled data. This package is the tooling for that validation: run a labeled dataset through your questions and get accuracy, calibration, and a recommended threshold per question.

Install

npm install @jev-harness/eval

Quick start (CLI)

export TYPESAFE_API_KEY=...
jev-eval --dataset cases.jsonl --out report.json

Example output:

Jev eval — 24 cases, model jev-latest, 2026-09-20T10:00:00.000Z
  requests 24  input tokens 3120  output tokens 0  est. cost $0.000131

touches_auth (noul, 24 scored)
  accuracy 91.7%  precision 90.0%  recall 94.7%  f1 92.3%
  brier 0.0612  auc 0.972  ece 0.0380
  suggested threshold (max F1): 0.55  (f1 92.3%, youdenJ 0.842)
  reliability: [0.0–0.1) n=3 avgP=0.04 avgY=0.00 | [0.8–0.9) n=5 avgP=0.84 avgY=0.80 | [0.9–1.0) n=14 avgP=0.96 avgY=1.00

est. cost uses Jev's input pricing ($0.042/Mtok); output tokens are free.

Dataset format

JSONL — one case per line, optionally carrying a shared questions map:

{"questions": {"touches_auth": {"type": "noul", "instructions": "Does this change affect authentication or session security?"}}, "state": {"diff": "changed login redirect"}, "label": {"touches_auth": true}}
{"state": {"diff": "bumped a dev dependency"}, "label": {"touches_auth": false}}

JSON — one object:

{
  "questions": {
    "risk": {
      "type": "score",
      "instructions": "Risk level",
      "criteria": ["None", "Low", "Moderate", "High", "Critical"]
    }
  },
  "cases": [{ "id": "c1", "state": { "diff": "..." }, "label": { "risk": "High" } }]
}

A bare array of cases works too — pair it with --questions questions.json.

Cases may also carry three optional benchmark fields:

| Field | Meaning | | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | | slice | Which benchmark slice the case belongs to ("obfuscation", "steering", "false-positive-trap", …). The report breaks the metrics down per slice. | | pair | Invariance group id. Cases sharing one describe the same action in different words; the report measures how far their probabilities diverge. | | note | Why the case is labeled the way it is — benchmark cases carry their rationale so the labels can be audited later. |

Labels are plain values, coerced per question type:

| Type | Accepts | | -------- | --------------------------------------- | | noul | true/false, "yes"/"no", 0/1 | | choice | the criteria key as a string | | score | the level name, or its 0-based index |

Cases with a missing or non-coercible label are skipped (counted, not fatal).

What you get per question

| Type | Metrics | | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | noul | accuracy, precision, recall, F1 at your threshold, Brier, rank AUC, ECE, reliability diagram, and a threshold sweep with the max-F1 (Youden tiebreak) recommendation | | choice | top-1 accuracy, multiclass Brier, calibration of confidence against being right | | score | MAE in level units, within-1 rate, Pearson correlation with the label |

The suggested threshold is a starting point from your data — keep final thresholds and side effects in your code, and prefer a threshold that matches the cost asymmetry of your workflow (a destructive-gate veto and a skill hint should not share one).

For noul questions the report also breaks the confusion matrix down per slice, with Wilson 95% intervals on recall and precision, and summarizes paraphrase invariance (maxΔp across pair groups). Slices exist because a healthy aggregate hides a broken slice: adversarial cases are a small share of a set, so losing all of them barely moves the total.

Benchmarking a decision

A single labeled set answers "does this work?". It cannot answer "does it still work on inputs I did not tune against?", which is the question that matters once a threshold ships. Build the set in two splits and keep them apart:

  • dev / tuning — question wording and thresholds may be iterated against it.
  • holdout — never used to choose wording. It is the generalization number, and it is the one worth quoting.

The repo's own gates are measured this way, in golden/:

| dataset | split | cases | questions | | ------------------------------- | ------------ | ----- | --------------------------------------- | | destructive-gate.json | dev / tuning | 78 | destructive (tool call) | | destructive-gate.holdout.json | holdout | 89 | destructive (tool call) | | merge-gate.json | dev / tuning | 41 | destructive + secret_leak (PR diff) |

Between them they cover clear cases, obfuscated commands (MITRE T1027.010 techniques: quoting, command substitution, wrappers, globs, variable indirection, encoded payloads), injected-steering text that argues for its own classification, distractor context, false-positive traps (dry runs, kill -0, writes to /dev/null), placeholder-vs-real credentials, and paraphrase pairs.

Holdout hygiene. A holdout is only worth what its discipline is worth. Use it to report, then treat it as a regression reference: repeated inspection, re-tuning against it, or reusing it to pick wording all leak it back into development and turn it into a second, quieter training set. When a case in it surfaces a miss, do not rewrite the set — record the miss, and add fresh cases for the next unbiased estimate. Each dataset here carries a comment saying whether it has already been inspected.

Two helpers keep the comparison honest when the sets are small:

  • wilsonInterval — confidence interval for a proportion, which keeps its coverage near 0/1 where the normal approximation does not.
  • mcnemarTest — exact paired test between two versions on the same cases; it only looks at where they disagree, so it answers "did this wording change actually help?" instead of "are the totals different?".

Neither is a substitute for more labeled data. Report the interval, then add cases.

A note for case authors

The API sits behind Cloudflare, which answers some shell-injection-shaped payloads with a 403 challenge before Jev ever sees them. Observed so far: IFS word-splitting (expanding IFS to rebuild the spaces in a command), and a runtime's shell-exec helper called inline. Quoting, command substitution, wrappers, globs, variable indirection, eval, base64/hex payloads, rev/xxd decoding and ANSI-C quoting all pass. Such a case cannot be measured through the public endpoint — probe the payload once before adding it, and keep the technique out of the set if the edge refuses it.

Keep literal payloads out of shipped files too. The same class of filter sits in front of the npm registry, so a literal payload quoted in a packaged README makes npm publish answer 403 Forbidden for that one package while every other package in the same release publishes fine — describe the technique instead of quoting the bytes.

Programmatic use

import { loadDataset, runEval, formatReport } from "@jev-harness/eval";

const dataset = loadDataset("cases.jsonl");
const report = await runEval({ apiKey: process.env.TYPESAFE_API_KEY! }, dataset, {
  concurrency: 4,
});
console.log(formatReport(report));

CLI options

| Flag | Default | Purpose | | -------------------- | ------------------------- | -------------------------------------------------- | | --dataset <path> | (required) | JSON or JSONL dataset | | --questions <path> | — | Question map JSON, merged under per-line questions | | --out <path> | — | Also write the full report as JSON | | --model <name> | jev-latest | Pin the model under test | | --base-url <url> | https://api.typesafe.ai | API override | | --timeout-ms <n> | 15000 | Per-request timeout | | --concurrency <n> | 4 | Cases in flight | | --no-sweep | off | Skip the noul threshold sweep | | --no-fail | off | Exit 0 even when cases failed (exploratory runs) | | --sweep-steps <n> | 20 | Sweep resolution |

Failing cases never abort the run: they are listed in the report. The CLI exits 0 when every case passed, 1 when cases failed (--no-fail opts out for exploratory runs), and 2 on usage errors such as a missing key or an unreadable dataset.

CI usage

Run the golden set on PRs that touch question wording, thresholds, or the model pin, and diff the report — calibration drift shows up as a changed suggested threshold or rising ECE before your users notice.

License

Apache-2.0.