npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@cycgraph/evals

v0.3.4

Published

Automated eval harness and quality-assurance gatekeeper for @cycgraph/* packages.

Readme

@cycgraph/evals

Regression-test harness for agent workflows. Deterministic + LLM-as-judge assertions, multi-sample evaluation, baseline drift gates.

License: Apache 2.0 Node.js

📚 Documentation  ·  🏃 Running evals  ·  📖 Assertions reference  ·  📐 Drift and baselines


Quality-assurance gate for the @cycgraph/* packages. Detects when a change in one package silently degrades the reasoning, schema compliance, or observable behavior of another, and tells you whether the regression is real or just sample noise.

This README is the quick-start + API at-a-glance. For concepts (drift gates, baseline persistence, sample stability), recording workflows, and extension recipes, see the Eval Harness section of the docs site.

What it gives you

  • 54 golden trajectories across 3 suites (orchestrator, memory, context-engine) with stable IDs and provenance.
  • Two assertion tracks:
    • Deterministic — pure library calls (no LLM): segmentation, dedup, budget, subgraph, conflict detection, etc.
    • Semantic — LLM-as-judge with three built-in rubric metrics (answer_relevancy, faithfulness, logical_coherence). Five reference-free metrics (instruction_following, output_quality, safety, compression_fidelity, qa_answerability) are exposed but not yet wired into a default suite.
  • Multi-sample evaluation — distinguishes flaky LLM responses from genuine regressions.
  • Baseline persistence — compares each run against the prior committed state and flags regressions that hide under the absolute drift ceiling.
  • Recording infrastructure — re-record any trajectory by running the input through the real System-Under-Test; goldens become observable behavior, not hand-authored intent.
  • Tag-routed dispatch — branching / supervisor / retry / etc. trajectories pick the right SUT graph automatically.
  • Efficacy + bench runners — sibling CLIs (evals:efficacy, bench) that measure absolute extraction/compression quality rather than drift.
  • Telemetry insights — deterministic, offline detectors over recorded runs that produce a ranked list of findings with supporting evidence (buildInsightsReport).
  • Knob sweeps — enumerate finite knob values from a finding, measure each against the same recorded prefix, and decide on the evidence (enumerateSweeps, decideSweep); no model proposes anything.

Quick start

Run the deterministic track (no LLM, <1s)

npm run evals --workspace=packages/evals -- --deterministic-only

Runs every library-level test across context-engine and memory. The orchestrator suite is semantic-only, so it doesn't appear in deterministic runs. Suitable for PR-time gating.

Run the full semantic gate (CI mode)

OPENAI_API_KEY=sk-... npm run evals:ci --workspace=packages/evals

Uses GPT-4o as the judge with 3 samples per metric and the OpenAI provider. Reports per-suite drift, flaky tests, and baseline delta.

Re-record goldens

# Memory + context-engine — no LLM needed
npx tsx packages/evals/scripts/record-goldens.ts --suite memory
npx tsx packages/evals/scripts/record-goldens.ts --suite context-engine

# Orchestrator — requires Anthropic key, real LLM calls
npx tsx packages/evals/scripts/record-goldens.ts --suite orchestrator

# Preview routing without running anything
npx tsx packages/evals/scripts/record-goldens.ts --suite memory --plan-only

# Actually overwrite the SQLite dataset
npx tsx packages/evals/scripts/record-goldens.ts --suite memory --commit

A dry-run writes golden/recording-diff-<suite>.json with old vs new for every trajectory. Inspect that before passing --commit.

Compare against a baseline

npm run evals --workspace=packages/evals -- --deterministic-only --baseline

The first run with --baseline creates golden/baselines/main-latest.json. Subsequent runs compare against it and exit with code 2 if any suite regressed by more than the noise floor (default 1 percentage point), even when the absolute drift ceiling hasn't been crossed.

CLI flags

| Flag | Type | Default | Purpose | |------|------|---------|---------| | --mode | local \| ci | local | Picks provider (Ollama / GPT-4o) + default concurrency | | --suite | suite name | (all) | Restrict to a single suite | | --samples | int | 1 local, 3 ci | Number of judge samples per semantic test | | --deterministic-only | flag | false | Skip the semantic track entirely (library checks only) | | --baseline | flag | false | Compare against persisted baseline; persist on pass | | --baseline-noise-floor | float | 1 | Min pp delta to count as a regression | | --sut-model | string | claude-sonnet-4-6 | Model for the orchestrator SUT | | --provider | anthropic \| openai \| ollama | (per mode) | Override the judge provider the mode would pick | | --commit | string | (auto) | Short git SHA stamped onto a new baseline snapshot |

Exit codes

| Code | Meaning | |------|---------| | 0 | Drift gate passed, no baseline regression | | 1 | Drift gate failed OR a suite failed to load | | 2 | Baseline regression detected, drift gate passed |

Configuration

| Variable | Required | Default | Purpose | |----------|----------|---------|---------| | OPENAI_API_KEY | CI only | — | GPT-4o judge API key | | ANTHROPIC_API_KEY | Recording, --provider anthropic | — | Claude API key for orchestrator recording and the Anthropic judge | | OLLAMA_BASE_URL | Local only | http://localhost:11434 | Ollama endpoint | | OLLAMA_MODEL | Local only | llama3:8b-instruct-q4_K_M | Local judge model | | EVAL_DRIFT_CEILING | No | 5.0 | Drift % gate threshold |

Judge concurrency is a provider option (maxConcurrency), not an env var: 2 for Ollama, 8 for OpenAI and Anthropic.

API at a glance

Assertions

import {
  // Structural — schema-level checks on tool calls
  assertToolCallStructure, assertTrajectoryStructure,
  // Deterministic — pure numeric/set/stability checks
  assertGreaterThanOrEqual, assertLessThanOrEqual,
  assertContainsAllKeys, assertSetEquals, assertStable, assertEqual,
  // Semantic — built-in LLM rubric metrics
  ANSWER_RELEVANCY, FAITHFULNESS, LOGICAL_COHERENCE, BUILT_IN_METRICS,
  // Reference-free — score without a comparison output
  INSTRUCTION_FOLLOWING, OUTPUT_QUALITY, SAFETY,
  COMPRESSION_FIDELITY, QA_ANSWERABILITY, REFERENCE_FREE_METRICS,
} from '@cycgraph/evals';

Telemetry insights + knob sweeps

import {
  buildInsightsReport, formatInsightsReport, DETECTORS,   // findings from recorded runs
  buildWorkflowProfile,                                    // per-node cost/latency profile
  enumerateSweeps, decideSweep,                            // knob sweeps over a finding
  planCombination, decideCombination,                      // combine winning knob values
} from '@cycgraph/evals';

Multi-sample semantic evaluation

import { evaluateMetricMultiSample, ANSWER_RELEVANCY } from '@cycgraph/evals';

const result = await evaluateMetricMultiSample(
  { input, actualOutput, expectedOutput },
  ANSWER_RELEVANCY,
  callJudge,
  { samples: 3, threshold: 0.8 },
);
// { median, stdDev, samples, stable, passed, reasoning }

Baseline persistence

import {
  snapshotFromDrift, writeBaseline, loadBaseline,
  compareBaseline, formatBaselineDelta,
} from '@cycgraph/evals';

const snapshot = snapshotFromDrift({ drift, driftCeiling: 5, commit: 'abc1234' });
writeBaseline(snapshot);
const delta = compareBaseline(snapshot, loadBaseline());
console.log(formatBaselineDelta(delta));

Recording

Recording runs through the scripts/record-goldens.ts script, which drives each input through the real System-Under-Test and rewrites the golden dataset. See Re-record goldens above for usage. The SUT layer it uses (runOrchestratorSut, runMemorySut, runContextEngineSut, the build*Graph builders, planForTrajectory, the retry-tool fixtures) is not part of the package barrel; it lives under src/sut/ and is consumed by the recording script via relative imports.

Dataset

import {
  loadGoldenTrajectories, loadManifest, listAvailableSuites,
  writeGoldenDataset, createSqliteBuffer, applyMigrations,
} from '@cycgraph/evals';

Runner

import { runEvals } from '@cycgraph/evals';

const result = await runEvals({
  mode: 'local',
  deterministicOnly: true,
  baseline: true,
  samples: 3,
});
// { drift, raw, suiteLoadErrors, baselineDelta?, baselineLoadError?, flakyTests? }

Golden dataset

Trajectories are stored as compressed SQLite (.sqlite.gz) under golden/data/, indexed by golden/manifest.json with sha256 checksums. The manifest is the source of truth for what's recorded; SQLite blobs are the data.

golden/
├── manifest.json               # Versioned index with sha256 — points at the live files
├── data/
│   ├── orchestrator-v3.sqlite.gz    # older -v1/-v2 files are retained alongside
│   ├── memory-v3.sqlite.gz
│   └── context-engine-v3.sqlite.gz
└── baselines/                  # (gitignored) per-run baseline snapshots
    └── main-latest.json

Schema migration — when a tool signature changes in a sibling package, scripts/migrate-golden.ts applies ordered transforms (rename / remove / add-required) to keep trajectories in sync without manual replay.

Architecture

            ┌────────────────────────────┐
            │     runEvals(config)       │
            └─────────────┬──────────────┘
                          │
       ┌──────────────────┼──────────────────┐
       ▼                  ▼                  ▼
  Deterministic    SUT-driven Semantic    Baseline
   (static         (runSutDispatch →       (load → compare
    registry)       evaluateMetricMulti)    → write on pass)
       │                  │                  │
       └────────┬─────────┘                  │
                ▼                            │
         computeDrift()                      │
                ▼                            │
         DriftReport ◄───────────────────────┘
                ▼
         formatReport() → stdout + GH annotations

Both tracks are commit-coupled: the deterministic track runs library code in-process, and the SUT-driven semantic track runs each trajectory through runSutDispatch against the real packages, then hands the observed output to the judge. When samples > 1, the semantic track runs N independent judge samples per metric and flags tests with inconsistent outcomes as flaky, which is distinct from genuine drift.

Development

# Unit tests for the harness itself
npm test --workspace=packages/evals

# Build
npm run build --workspace=packages/evals

# Type check
npm run lint --workspace=packages/evals

# Efficacy matrix (extraction/compression quality vs labeled corpora)
npm run evals:efficacy --workspace=packages/evals

# Compression bench (context engine vs reference compressors; --smoke for a quick pass)
npm run bench --workspace=packages/evals

Covers assertions, dataset I/O, schema migration, SUT dispatch, multi-sample evaluation, baseline persistence/comparison, and runner integration.

Related

Contributing

Issues and PRs welcome on GitHub. See CONTRIBUTING.md.

License

Apache 2.0.