@cycgraph/evals
v0.3.4
Published
Automated eval harness and quality-assurance gatekeeper for @cycgraph/* packages.
Readme
@cycgraph/evals
Regression-test harness for agent workflows. Deterministic + LLM-as-judge assertions, multi-sample evaluation, baseline drift gates.
📚 Documentation · 🏃 Running evals · 📖 Assertions reference · 📐 Drift and baselines
Quality-assurance gate for the @cycgraph/* packages. Detects when a change in one package silently degrades the reasoning, schema compliance, or observable behavior of another, and tells you whether the regression is real or just sample noise.
This README is the quick-start + API at-a-glance. For concepts (drift gates, baseline persistence, sample stability), recording workflows, and extension recipes, see the Eval Harness section of the docs site.
What it gives you
- 54 golden trajectories across 3 suites (
orchestrator,memory,context-engine) with stable IDs and provenance. - Two assertion tracks:
- Deterministic — pure library calls (no LLM): segmentation, dedup, budget, subgraph, conflict detection, etc.
- Semantic — LLM-as-judge with three built-in rubric metrics (
answer_relevancy,faithfulness,logical_coherence). Five reference-free metrics (instruction_following,output_quality,safety,compression_fidelity,qa_answerability) are exposed but not yet wired into a default suite.
- Multi-sample evaluation — distinguishes flaky LLM responses from genuine regressions.
- Baseline persistence — compares each run against the prior committed state and flags regressions that hide under the absolute drift ceiling.
- Recording infrastructure — re-record any trajectory by running the input through the real System-Under-Test; goldens become observable behavior, not hand-authored intent.
- Tag-routed dispatch —
branching/supervisor/retry/ etc. trajectories pick the right SUT graph automatically. - Efficacy + bench runners — sibling CLIs (
evals:efficacy,bench) that measure absolute extraction/compression quality rather than drift. - Telemetry insights — deterministic, offline detectors over recorded runs that produce a ranked list of findings with supporting evidence (
buildInsightsReport). - Knob sweeps — enumerate finite knob values from a finding, measure each against the same recorded prefix, and decide on the evidence (
enumerateSweeps,decideSweep); no model proposes anything.
Quick start
Run the deterministic track (no LLM, <1s)
npm run evals --workspace=packages/evals -- --deterministic-onlyRuns every library-level test across context-engine and memory. The orchestrator suite is semantic-only, so it doesn't appear in deterministic runs. Suitable for PR-time gating.
Run the full semantic gate (CI mode)
OPENAI_API_KEY=sk-... npm run evals:ci --workspace=packages/evalsUses GPT-4o as the judge with 3 samples per metric and the OpenAI provider. Reports per-suite drift, flaky tests, and baseline delta.
Re-record goldens
# Memory + context-engine — no LLM needed
npx tsx packages/evals/scripts/record-goldens.ts --suite memory
npx tsx packages/evals/scripts/record-goldens.ts --suite context-engine
# Orchestrator — requires Anthropic key, real LLM calls
npx tsx packages/evals/scripts/record-goldens.ts --suite orchestrator
# Preview routing without running anything
npx tsx packages/evals/scripts/record-goldens.ts --suite memory --plan-only
# Actually overwrite the SQLite dataset
npx tsx packages/evals/scripts/record-goldens.ts --suite memory --commitA dry-run writes golden/recording-diff-<suite>.json with old vs new for every trajectory. Inspect that before passing --commit.
Compare against a baseline
npm run evals --workspace=packages/evals -- --deterministic-only --baselineThe first run with --baseline creates golden/baselines/main-latest.json. Subsequent runs compare against it and exit with code 2 if any suite regressed by more than the noise floor (default 1 percentage point), even when the absolute drift ceiling hasn't been crossed.
CLI flags
| Flag | Type | Default | Purpose |
|------|------|---------|---------|
| --mode | local \| ci | local | Picks provider (Ollama / GPT-4o) + default concurrency |
| --suite | suite name | (all) | Restrict to a single suite |
| --samples | int | 1 local, 3 ci | Number of judge samples per semantic test |
| --deterministic-only | flag | false | Skip the semantic track entirely (library checks only) |
| --baseline | flag | false | Compare against persisted baseline; persist on pass |
| --baseline-noise-floor | float | 1 | Min pp delta to count as a regression |
| --sut-model | string | claude-sonnet-4-6 | Model for the orchestrator SUT |
| --provider | anthropic \| openai \| ollama | (per mode) | Override the judge provider the mode would pick |
| --commit | string | (auto) | Short git SHA stamped onto a new baseline snapshot |
Exit codes
| Code | Meaning | |------|---------| | 0 | Drift gate passed, no baseline regression | | 1 | Drift gate failed OR a suite failed to load | | 2 | Baseline regression detected, drift gate passed |
Configuration
| Variable | Required | Default | Purpose |
|----------|----------|---------|---------|
| OPENAI_API_KEY | CI only | — | GPT-4o judge API key |
| ANTHROPIC_API_KEY | Recording, --provider anthropic | — | Claude API key for orchestrator recording and the Anthropic judge |
| OLLAMA_BASE_URL | Local only | http://localhost:11434 | Ollama endpoint |
| OLLAMA_MODEL | Local only | llama3:8b-instruct-q4_K_M | Local judge model |
| EVAL_DRIFT_CEILING | No | 5.0 | Drift % gate threshold |
Judge concurrency is a provider option (maxConcurrency), not an env var: 2 for Ollama, 8 for OpenAI and Anthropic.
API at a glance
Assertions
import {
// Structural — schema-level checks on tool calls
assertToolCallStructure, assertTrajectoryStructure,
// Deterministic — pure numeric/set/stability checks
assertGreaterThanOrEqual, assertLessThanOrEqual,
assertContainsAllKeys, assertSetEquals, assertStable, assertEqual,
// Semantic — built-in LLM rubric metrics
ANSWER_RELEVANCY, FAITHFULNESS, LOGICAL_COHERENCE, BUILT_IN_METRICS,
// Reference-free — score without a comparison output
INSTRUCTION_FOLLOWING, OUTPUT_QUALITY, SAFETY,
COMPRESSION_FIDELITY, QA_ANSWERABILITY, REFERENCE_FREE_METRICS,
} from '@cycgraph/evals';Telemetry insights + knob sweeps
import {
buildInsightsReport, formatInsightsReport, DETECTORS, // findings from recorded runs
buildWorkflowProfile, // per-node cost/latency profile
enumerateSweeps, decideSweep, // knob sweeps over a finding
planCombination, decideCombination, // combine winning knob values
} from '@cycgraph/evals';Multi-sample semantic evaluation
import { evaluateMetricMultiSample, ANSWER_RELEVANCY } from '@cycgraph/evals';
const result = await evaluateMetricMultiSample(
{ input, actualOutput, expectedOutput },
ANSWER_RELEVANCY,
callJudge,
{ samples: 3, threshold: 0.8 },
);
// { median, stdDev, samples, stable, passed, reasoning }Baseline persistence
import {
snapshotFromDrift, writeBaseline, loadBaseline,
compareBaseline, formatBaselineDelta,
} from '@cycgraph/evals';
const snapshot = snapshotFromDrift({ drift, driftCeiling: 5, commit: 'abc1234' });
writeBaseline(snapshot);
const delta = compareBaseline(snapshot, loadBaseline());
console.log(formatBaselineDelta(delta));Recording
Recording runs through the scripts/record-goldens.ts script, which drives each input through the real System-Under-Test and rewrites the golden dataset. See Re-record goldens above for usage. The SUT layer it uses (runOrchestratorSut, runMemorySut, runContextEngineSut, the build*Graph builders, planForTrajectory, the retry-tool fixtures) is not part of the package barrel; it lives under src/sut/ and is consumed by the recording script via relative imports.
Dataset
import {
loadGoldenTrajectories, loadManifest, listAvailableSuites,
writeGoldenDataset, createSqliteBuffer, applyMigrations,
} from '@cycgraph/evals';Runner
import { runEvals } from '@cycgraph/evals';
const result = await runEvals({
mode: 'local',
deterministicOnly: true,
baseline: true,
samples: 3,
});
// { drift, raw, suiteLoadErrors, baselineDelta?, baselineLoadError?, flakyTests? }Golden dataset
Trajectories are stored as compressed SQLite (.sqlite.gz) under golden/data/, indexed by golden/manifest.json with sha256 checksums. The manifest is the source of truth for what's recorded; SQLite blobs are the data.
golden/
├── manifest.json # Versioned index with sha256 — points at the live files
├── data/
│ ├── orchestrator-v3.sqlite.gz # older -v1/-v2 files are retained alongside
│ ├── memory-v3.sqlite.gz
│ └── context-engine-v3.sqlite.gz
└── baselines/ # (gitignored) per-run baseline snapshots
└── main-latest.jsonSchema migration — when a tool signature changes in a sibling package, scripts/migrate-golden.ts applies ordered transforms (rename / remove / add-required) to keep trajectories in sync without manual replay.
Architecture
┌────────────────────────────┐
│ runEvals(config) │
└─────────────┬──────────────┘
│
┌──────────────────┼──────────────────┐
▼ ▼ ▼
Deterministic SUT-driven Semantic Baseline
(static (runSutDispatch → (load → compare
registry) evaluateMetricMulti) → write on pass)
│ │ │
└────────┬─────────┘ │
▼ │
computeDrift() │
▼ │
DriftReport ◄───────────────────────┘
▼
formatReport() → stdout + GH annotationsBoth tracks are commit-coupled: the deterministic track runs library code in-process, and the SUT-driven semantic track runs each trajectory through runSutDispatch against the real packages, then hands the observed output to the judge. When samples > 1, the semantic track runs N independent judge samples per metric and flags tests with inconsistent outcomes as flaky, which is distinct from genuine drift.
Development
# Unit tests for the harness itself
npm test --workspace=packages/evals
# Build
npm run build --workspace=packages/evals
# Type check
npm run lint --workspace=packages/evals
# Efficacy matrix (extraction/compression quality vs labeled corpora)
npm run evals:efficacy --workspace=packages/evals
# Compression bench (context engine vs reference compressors; --smoke for a quick pass)
npm run bench --workspace=packages/evalsCovers assertions, dataset I/O, schema migration, SUT dispatch, multi-sample evaluation, baseline persistence/comparison, and runner integration.
Related
@cycgraph/orchestrator— the system under test@cycgraph/memory— knowledge-graph SUT@cycgraph/context-engine— compression SUT- Orchestrator's internal
runEval— lightweight per-graph assertion framework (different from this package's regression harness) examples/eval-gated-learning/— runnable demo of the eval-gated retention loop (poisoned lessons evicted on outcome evidence)
Contributing
Issues and PRs welcome on GitHub. See CONTRIBUTING.md.
