evidence-layer
v0.3.0
Published
Checks that turn a review's assertions into claims that could be proved wrong.
Maintainers
Readme
evidence-layer
Checks that turn a review's assertions into claims that could be proved wrong.
A review - written by a person, or generated by a model that just wrote the change it is reviewing - can assert anything at all about what it read and what it ran, in prose that reads identically whether or not any of it happened. A README asserts a test count, a build size, a set of pinned versions; all true when written, none of them with a clock. This package closes that gap for an ordinary Node project: one executable whose commands each take one kind of claim and make it falsifiable, plus a journal that gives the whole layer a clock.
Falsifiable does not mean fragile or fakeable. It means testable: you can name, in advance, what would prove the claim wrong. A claim that survives every possible observation is not strong, it is empty.
Install
npm install --save-dev evidence-layerRequires Node 22 or newer. Zero runtime dependencies.
Node 20 is not supported, and the reason is a date rather than a taste: it reached end of life on 2026-04-30 and receives no security patches. A package whose entire subject is claims with a clock does not get to quietly support a runtime that stopped having one.
AI-driven development, and the vibe-coding case
A coding agent writes the change and then writes the account of the change. Both come out of the same model, in the same fluent prose, and the account reads the same whether the suite ran or not. "I ran the tests and they pass" is a plausible next sentence, which is all a language model needs in order to emit it. Plausibility is not observation, and nothing in the surrounding text tells you which one you got.
That is not an accusation of lying. It is cheaper than lying: gathering evidence is off the trajectory the writer is already on, and a sentence costs nothing. The identical failure occurs in human reviews - it is only that nobody could previously generate a thousand of them a week.
Vibe coding sharpens the problem to a point. The whole trade is accepting code you did not read, which leaves the model's summary as the only thing standing between you and the diff. An unchecked summary is not a compressed version of reading the code; it is a different artifact that happens to be made of the same words.
What this layer changes, concretely:
- Gathering moves off the writer.
gatherEvidence()runs the expensive commands before the review is written, with something that is not the reviewer, and writes the results intoevidence.local.mdas artifact blocks a review can cite. It refuses to run when the working tree is not at the commit you asked about - an agent that drifted a commit mid-session otherwise produces confident evidence about a different codebase. - Every claim carries an address.
VERIFIED[tests-a3f91c]names the exact artifact it rests on, andclaimsresolves that id, checks it was collected at the commit under review, and warns when one artifact is quietly backing four unrelated claims. A bare tag any writer can satisfy for free carries no information; an id is either in the document, at that commit, or it is not. - The receipt proves the agent read this repository. Stale context
windows, files renamed three commits ago, paths reconstructed from a
half-remembered tree:
receiptchecks the review's Access Receipt against git's own output, and backs every cited file with that file's verbatim first line. The baseline is checked by equality againstmerge-base origin/<target> <head>, never ancestry, because a stale base fails in the reassuring direction - the diff comes out larger, so nothing looks missing. - Honest stays cheap, which is the only reason any of it gets used. Gathering writes a skeleton with the ids already substituted, so citing an artifact means leaving a block in place, and asserting something unbacked means deleting part of the template. A convention that costs the writer more than the shortcut loses to the shortcut every time, model or not. A run its time limit killed is suggested as "did not complete", never as a failure it never got to report.
- Agent review gets a scorecard instead of a vibe. A commit fixing a bug a
review let through adds
Missed-By: <review or PR>to its own message, andgovernanceextracts those into the journal on every run. "Is the agent's review catching anything real?" becomes a count over time - a floor, never a rate (src/core/misses.tssays why). - The output is machine-readable, so the loop can close. Exit codes are a
contract, every check reports the same five levels, and
governanceprints a delimited JSON block on every run, clean ones included. An agent can run the check, read its own grade, and fix the finding before a human opens the PR. - Artifact ids do not depend on your terminal. Captured output is
ANSI-stripped before it is hashed, because
FORCE_COLORis set in most CI environments and inherited by every child process. An id that changes with the terminal is not an address.
What it does not do is judge whether an artifact's output actually supports
the sentence next to it. That needs reading output for meaning - a probabilistic
judgement, and putting one inside a deterministic checker reintroduces exactly
the circular validation this layer exists to avoid. See what this does not
do below, and
examples/plain-node-project/reviews/shotgun-artifact.md
for the fixture that passes on purpose to prove the limit is real.
The one rule everything here is graded against
The honest path must be cheaper than the convenient one - fewer tokens,
fewer separate decisions, less distance from the trajectory the author (human or
model) is already on. Nothing here relies on anyone wanting to be careful.
Anything that made correctness more expensive than a shortcut was rejected or
rebuilt before shipping - see src/core/skeleton.ts's file header for the one
concrete near-miss.
The commands
One executable, five commands, six checks - journal is the one command with
subcommands:
npx evidence-layer claims review.md # is every claim backed by evidence?
npx evidence-layer receipt review.md # was this review written against this repo?
npx evidence-layer ci-summary output.txt # render captured output as a CI summary
npx evidence-layer governance # is anything actually running the checks?
npx evidence-layer journal report # what the checks have actually found
npx evidence-layer journal tag <id> real # mark a finding real or false| command | the claim it makes falsifiable |
| --- | --- |
| claims | "I verified this" - every VERIFIED[id] tag must address a real artifact collected at the commit under review |
| receipt | "I read this repository" - the review's Access Receipt must match the repository's own git output, and a cited file must be backed by that file's own verbatim first line |
| ci-summary | nothing; it renders a captured run for a CI job summary, and reports unverifiable when it has nothing to read |
| governance | "this is enforced" - every invariant a decision record declares must be reachable from a git hook or a CI workflow, and no governance exception may be past its expiry; on the way, it records every Missed-By: commit trailer into the journal (--no-journal to skip) |
| journal report | "these checks are useful" - real findings, false findings and misses, counted over time rather than assumed |
| journal tag | as above, at the moment of fixing, by whoever fixed it |
A receipt quotes each cited file in a table - one row per file, the first
non-blank line in a code span, | escaped as \|:
**Sample integrity**:
| File | First non-blank line |
|---|---|
| `src/a.ts` | `import { b } from './b'` |The one-line form from 0.1.x (**Sample integrity**: `path` -> `line`) is
still read, so existing reviews keep verifying.
The colon-separated names these checks had before they were a package
(check:claims, check:receipt, ci:summary, journal:report,
journal:tag) still work as aliases.
Gathering is deliberately not a subcommand. It runs whatever your project's own commands are, at your project's own gates and timeouts, so it is a library call in a script you own rather than a CLI flag pretending to know them.
The loop, in a project with an agent in it
// scripts/evidence.mjs - run before the review is written, not after
import { gatherEvidence } from 'evidence-layer'
gatherEvidence({ repo: process.cwd(), base: 'origin/main' })node scripts/evidence.mjs # artifacts + a pre-filled claim skeleton
npx evidence-layer claims review.md # does every VERIFIED[id] resolve, at this commit?
npx evidence-layer receipt review.md # was this written against this checkout?
npx evidence-layer governance # and is anything actually running the above?Step one carries the weight, and it is the step an agent will skip on its own, because it is off the path. Put it in the script the agent is told to run, not in the instructions it is told to follow.
The last step belongs in CI rather than on the machine that wrote the change. Every check here can be run by whoever wrote the diff, which makes a passing local run a self-report - and anyone who skipped the gathering step skips the verifying step just as easily. Same commands, different party.
Exit codes are a contract
| code | meaning | | --- | --- | | 0 | clean | | 1 | at least one error | | 2 | bad usage - never a pass | | 3 | flaky - non-zero, but not the same claim as "the change broke something" |
Five outcomes, not two: pass, warning, error, unverifiable, flaky.
unverifiable is deliberately neither a pass nor an error. A checker that
reports success when it could not run manufactures confidence out of nothing.
Findings print one status icon per level, coloured unless NO_COLOR is set to a
non-empty value - no-color.org's exact rule, so an
inherited opt-out can still be cancelled with NO_COLOR=. Colour is not gated
on isTTY: sandboxed sessions and CI log viewers have no terminal attached and
render ANSI perfectly well.
What "zero-config" means, precisely
claims, receipt, ci-summary, journal, governance, exceptions and misses
need no configuration at all - they operate on markdown text, git output and
package.json, which every project already has. Call gatherEvidence({ repo,
base }) with nothing else and it plans pnpm lint, pnpm typecheck, a
lockfile check and pnpm test - the three scripts an ordinary Node project
already has names for, plus one fact pnpm-lock.yaml can answer about itself -
and reports the working tree's own status alongside them.
What was never claimed to be zero-config, because it cannot be: any check
that asserts a fact specific to one project - a route existing, a dependency
pinned to an exact version, a build fitting a size budget. Those are facts about
that project, not about Node, and belong in that project's own script. This
repository's scripts/check-docs.ts and scripts/check-deps.ts are the worked
examples: one checks the claims this README makes, the other checks that
every dependency this package declares is actually used, and both live here
rather than in the package because nobody else's README or manifest makes
those exact claims.
Using it as a library
import { lintClaims, checkReceipt, parseReceipt, gatherEvidence } from 'evidence-layer'
import { parseTestCount } from 'evidence-layer/adapters/node-default'Everything the CLI does is exported. src/core/ has no knowledge of any
project's stack; src/adapters/node-default.ts is the one file allowed to know
it is running on Node and how a handful of common test reporters format their
output.
The portability test
A claim that a package works on "any project" is exactly the kind of claim this
package's own claims check would tag UNVERIFIED if nobody had checked it.
examples/plain-node-project/ is a deliberately
unrelated fixture - plain JavaScript, no TypeScript, no framework, no monorepo,
its own unconnected package.json - that
test/portability.test.ts exercises
readInvariants, reachableCommands, lintClaims, checkReceipt and
renderSummary against. Clone it, run the layer against it, and see the checks
fail on purpose: examples/plain-node-project/README.md
is the guided version.
The boundary is a test, not an intention
Nothing under src/core/ may import anything but Node builtins and its own
siblings. That is checked by test/boundary.test.ts on
every run, because a layer separation enforced by good intentions is worth
nothing. Node builtins are fine there: this is CLI tooling, expected to shell
out and touch the filesystem, not a sandboxed engine.
What this repository checks about itself
The package is its own first consumer. pnpm check:all runs this repository's
own checks and then checks that something actually runs them - an invariant no
git hook or CI workflow reaches is a finding, because a rule with no trigger is
decoration.
The suite is 210 tests across 18 files, and that sentence is checked
too - check:docs runs the suite and compares. Change the number and watch it
fail; a check nobody has seen fail is indistinguishable from a check that
cannot fail.
pnpm evidence --base origin/master # gather this repository's own evidence bundle
pnpm check:docs # do the claims on this page still hold?
pnpm check:deps # is every dependency this package declares actually used?
pnpm check:ladder # did this commit climb the decision ladder it says it did?
pnpm check:all # every check, plus: does anything run them?
pnpm journal:report # what the checks have found here, over timeCI runs the suite on every supported Node version and, in a separate job with no
package manager and no install, unpacks the built dist/ and runs it on bare
node. If that job passes, "zero runtime dependencies" is a checked fact rather
than a sentence on this page.
What this does not do
- It does not make findings correct. Access, gathered evidence and grounded claims are necessary, not sufficient. None of the three makes a claim true.
- It does not stop someone determined to fake it. It removes the default failure, which is the one that actually happens.
- It does not judge taste. "This will be hard to maintain" is a legitimate
review comment and always will be, untagged - prose outside a
**Claim**:line is never graded. It just must not be dressed as a check. - The reachability scan reads literal command strings, so a command assembled at
runtime is invisible to it unless the invariant self-discloses that fact and
accepts
unverifiableinstead of a guess. Blind spots that are written down can be calibrated against; blind spots that are not, cannot. - The misses journal measures a floor, not a rate. A miss nobody traced back to the review that let it through stays invisible, so it supports a trend and a case-by-case reading, never a claimed percentage of defects caught.
Releasing
pnpm release # patch: 0.1.0 -> 0.1.1
pnpm release minor # 0.1.0 -> 0.2.0
pnpm release 1.0.0 # exactly that
pnpm release --dry-run # run every check, print the plan, change nothingOne command: it checks, bumps, commits, tags and pushes. The tag is what triggers publication - nothing is published from a laptop, because npm only generates a provenance attestation inside CI, and a local publish would quietly ship a version weaker than every other version.
Before it touches anything it checks that the release is cut from the default
branch, that the tree is clean, that the commit is the one the remote has, that
the tag is free, that the version has never been published, and that
pnpm check:all --enforce passes. Every one of those exists because it went
wrong during the first release of this package.
Docs
docs/evidence-layer.md- the full tour, in the order the checks are meant to be rundocs/decisions/- why each piece exists, with the invariants each record claims are enforceddocs/governance-exceptions.md- how an exception gets an expiry date instead of becoming permanentdemo/- the four demos the talk runs, in slide order, with the receipt fixtures and the stale-baseline refs they need
