@cef-ai/eval
v0.1.0
Published
Workflow evaluations for CEF: datasets, experiments, scoring and comparison over an injected store and workflow target.
Maintainers
Keywords
Readme
@cef-ai/eval
Evaluations for CEF workflows: build datasets of cases from scratch or from real runs, run them as experiments against a workflow version, and compare results and run metrics against a baseline. Warn-only — nothing here blocks a deploy.
Browser-safe (no Node built-ins). ROC and the cef eval CLI use the same
library over the same storage layout, so either reads what the other wrote.
Vocabulary
| Term | What it is |
| --- | --- |
| Dataset | The set of cases for one workflow; versioned. Names are unique per workflow. |
| Case | One workflow.start payload (input) and what the run's Result should be (expected). |
| Version | A frozen cut of a dataset: every live case pinned at one revision by sha256. |
| Experiment | One run of a dataset version on a workflow version. |
| Repeats | How many times each case runs in an experiment (default 1). |
| Run | One case × one repeat inside an experiment — a RunRecord with output, scores and metrics. |
| Baseline | The experiment others are compared to (Header.baselineId, setBaseline). |
| Result | What a workflow declares as its output: an output node's params.result. A run's workflow.completed carries exactly those fields; that is what cases are scored against. |
Declaring a Result
{ "id": "result", "kind": "output", "params": {
"result": { "category": "={{ $json.category }}", "confidence": "={{ $json.score }}" },
"resultTypes": { "category": "string", "confidence": "number" },
"outcome": "={{ $json.category }}"
} }Cases and scoring
{
"id": "refund-angry",
"input": { "text": "I want my money back NOW" },
"expected": {
"category": "billing",
"confidence": ">= 0.8",
"tags": { "$contains": "refund" },
"reply.language": { "$oneOf": ["en", "en-GB"] }
},
"limits": { "maxDurationMs": 20000 }
}One score per expected field (eval/field@1), addressed by dot path. A plain
value is deep equality (an object is a subset match), a string like ">= 0.8"
is a numeric band, and matchers are $eq $oneOf $contains $regex $between $gte
$lte $exists $approx/$tol. A matcher naming several operators needs all of them
to hold ({ "$exists": true, "$gte": 1 }); $between is [min, max] with
min ≤ max. A field the output does not have fails. A case passes
when every judged field score passes; with no expectations it is unscored. A run that
failed or timed out fails. The summary's passRate (shown as Accuracy) and
fields (per expected field: passed / failed / notApplicable / rate) roll the
scores up. Limits (maxDurationMs maxTokens maxSteps maxCost,
header-wide or per case) are scored as eval/limit@1, counted as breaches, and
never fail a case.
$regex and trust. Patterns run in JavaScript's backtracking engine,
synchronously, inside scoring — nothing can time one out, so a catastrophic
pattern freezes the CLI or the ROC tab. Cases come from the shared bucket, so
scoring refuses (fails the field with unsafe $regex refused: …) a pattern
longer than 256 characters or one that quantifies a group which itself contains
a quantifier ((a+)+, (\w+\s?)*). That is a heuristic, not a proof — an
overlapping alternation such as (a|a)* passes it. Treat write access to a
dataset like write access to the workflow it tests: anyone who can write cases
can make scoring slow.
Scoring against the version's Result
An experiment records the Result the workflow version declares
(Experiment.resultFields, from resultFieldsFromGraph(graph) — the graph
object, its JSON text, or a manifest's params.graph; several output nodes
are unioned). An expected field that Result does not declare is scored n/a
(pass: null, "not in this version's Result"), and one whose expectation cannot
match the declared type is n/a too ("type changed: expected number, Result
declares string") — the dataset is older than the version, which is not a
regression. checkResultCompatibility(resultFields, cases) reports the same up
front: missing, typeChanged, and unexpectedNew (declared, never expected).
Storage
On an injected EvalStore (list, get, put, putIfAbsent; run
storeConformance against an implementation). In the agent service's bucket:
workflows/<wf>/datasets/<dataset>/header.json Header (+ baselineId)
workflows/<wf>/datasets/<dataset>/cases/<caseId>/r<n>.json write-once revisions
workflows/<wf>/datasets/<dataset>/cases/<caseId>/deleted.json tombstone
workflows/<wf>/datasets/<dataset>/archived/<caseId>.json archive flag
workflows/<wf>/datasets/<dataset>/versions/v<n>.json {version, header, cases:[{id,rev,sha256}], createdAt, note?}
workflows/<wf>/experiments/<expId>/meta.json Experiment (dataset@version, workflowVersion, resultFields, status, summary)
workflows/<wf>/experiments/<expId>/runs/<caseId>.<repeat>.json RunRecord
workflows/<wf>/experiments/<expId>/attempts/a<n>.json attempt claim (write-once)Experiments are siblings of the datasets: an experiment names the
dataset@version it ran. Every dataset function takes (store, workflowId,
dataset, …), every experiment function (store, workflowId, expId, …).
Storage written by the first release (evals/datasets/<dataset>/…, experiments
inside the dataset) is copied into this layout by migrateLegacyLayout(store)
(cef eval migrate): putIfAbsent per key, legacy keys left in place, a
evals/MIGRATED.json marker written once every dataset was copied.
needsLegacyMigration(store) says whether to run it.
Dataset versions
Versions are cut automatically: ensureDatasetVersion(store, wf, dataset)
returns the latest version when the working set (live cases at their latest
revisions) is what it pins, and otherwise cuts the next one with the note
auto: <n> added, <m> edited, <k> removed. runExperiment calls it when no
datasetVersion is given. cutVersion remains for explicit cuts (cef eval
push --cut); versionHistory lists each version with what changed and how
many experiments ran it.
compareExperiments compares only the cases present on both sides at the same
revision (comparedOn), and reports the rest as datasetDiff: {added, removed,
edited}.
Running experiments
import { runExperiment, resultFieldsFromGraph } from "@cef-ai/eval";
const exp = await runExperiment({
store, workflowId: "ticket-triage", dataset: "triage", // datasetVersion: 3 — default: ensureDatasetVersion
workflowVersion: "1.4.0", live: false,
resultFields: resultFieldsFromGraph(graph),
repeats: 3, concurrency: 4, timeoutMs: 120_000,
target, // WorkflowTarget: start / waitForEnd / metrics / pinVersion?
onProgress: ({ done, total }) => console.log(done, total),
});Each run publishes into its own context, eval-<expId>-<caseId>-<repeat>-a<attempt>.
A non-live version is run through target.pinVersion, which routes only this
experiment's contexts (prefix eval-<expId>-) to it. The meta is written first,
each run record as it finishes, the summary last; pass experimentId to resume.
A resume is a new attempt (attempt on the meta and on each run it writes), so
a run the dead attempt had already started never shares a context with its
replacement. The attempt is claimed write-once (putIfAbsent on
experiments/<expId>/attempts/a<n>.json, n upward until the claim sticks)
before anything is published, so two concurrent resumes of one experiment get
different attempts and never publish into the same contexts.
CLI walkthrough
export CEF_AS_PUBKEY=0x… # the agent service; names the bucket
export CEF_KEYSTORE=~/owner.json # owner wallet (or CEF_SECRET_PHRASE)
export CEF_KEYSTORE_PASSWORD=…
export CEF_ACCESS_TOKEN=… # CLI access token from ROC, for the orchestrator
# In the workflow's folder: --workflow defaults to the workflow cef.config.ts
# (or dist/) declares, --dir to ./datasets.
cef eval create triage --max-duration-ms 30000
cef eval pull triage # → datasets/triage/{header.json,cases/*.json,.versions.json}
$EDITOR datasets/triage/cases/refund-angry.json
cef eval push triage # new revisions; --cut "note" cuts a version explicitly
cef eval run triage # latest dataset version (cut if the cases changed) × the live workflow version
cef eval run triage --workflow-version 1.5.0 --repeats 3 --baseline exp-mgfz8c1k-a1b2
cef eval run triage --dataset-version 2 # an older dataset version
cef eval experiments triage
cef eval compare exp-mgfz8c1k-a1b2 exp-mgg0p2xq-9zt4
cef eval migrate # once, for a bucket written by the first releaseThe wallet's S3 credential is minted once, kept for its hour in
~/.cef/eval-credentials.json (owner-only; CEF_EVAL_CREDENTIAL_CACHE=<file>
moves it, =off disables it) so each command does not mint another against the
gateway's per-wallet cap, re-minted before it expires during a long run, and
re-minted once if the gateway refuses it.
Store credentials are taken in this order: explicit --access-key/--secret
($CEF_EVAL_KEY/$CEF_EVAL_SECRET, an S3 key issued for the owner's wallet);
then the owner's wallet, minting as above. $CEF_DDC_ACCESS_TOKEN, the token
cef push uses, cannot open dataset storage: it is limited to the agent
service's registry bucket, while datasets live in the gateway bucket
as-<pubkey>, a separate DDC bucket, and it cannot be delegated to the gateway.
With only that token set, cef eval says so and names the two options above.
run records the version's Result — the live deployment's graph when the
version is live, else the project's own graph of exactly that version — and
warns before starting about cases expecting fields it no longer returns. The
baseline stays on the platform: pull leaves it out of header.json, push
keeps the bucket's, and run compares against it when --baseline is not given.
Every command takes --json. push writes a new revision only for a case that
changed and refuses one edited in the bucket since the last pull (--force to
write over it). A repo case that is deleted or archived in the bucket is refused
before anything is written; --restore undeletes/unarchives it and writes the
repo's version (--force does not). pull and push --prune only remove what
the last pull wrote (.versions.json): a case file the pull never wrote is
reported as local-only and kept, and a case added in the bucket after the pull
is never pruned. pull records each case's content hash in .versions.json and
refuses, before writing anything, to overwrite or remove a case file edited since
the last pull/push — one matching neither that hash nor the bucket's current
content — unless --force. Both pull and push refuse a dataset dir whose
header.json names another workflow.
