@geonosis/ratchet
v3.1.0
Published
Debt as a number that may only shrink — one ratchet, pluggable counters.
Maintainers
Readme
@geonosis/ratchet
Through the front door: geonosis ratchet — the metapackage pins this and every other kit tool at
ONE version, and passes the exit code through unchanged.
Debt as a number that may only shrink. It measures a repo's counters, compares each against a committed baseline, and fails only on regression — so a repo can adopt a gate the day it is written, with the existing debt named rather than forgiven, and nothing new can land behind it.
Install
pnpm add -D @geonosis/ratchetgeonosis-ratchet --prove # first: prove every counter CAN report a finding
geonosis-ratchet # then: every counter
geonosis-ratchet --tier fast # only the counters in that tier
geonosis-ratchet --exclusive # one heavy run at a time on this machine
geonosis-ratchet --cwd apps/webExit 1 when a number grew or moved in a way the run cannot vouch for, 2 when a counter could not measure at all, 0 otherwise.
--prove is the first thing to run after installing
A counter that reads 0 because the tree is clean and a counter that reads 0 because it can no longer read the tool are the same number. The second is worse than no gate at all: it reports green, and when it shrinks the ratchet writes the lie into the baseline.
This is not hypothetical. On one day, 2026-08-29, five gates were found reporting green while
measuring nothing: this package's own oxlintErrors read 0 under the --format=unix its config
asked for, and had since it was written; a lint rule sat at "error" for months with no
configuration to fire on; a parity run said PASS about rules it never exercised; a test runner
exited 0 over failing suites for weeks; and a git rebase | tail hid a conflict behind a pipe's
exit status.
So every counter ships a probe: a known-bad input it must be able to read.
geonosis-ratchet --prove PROVEN oxlintErrors: reads 1 on a planted finding
PROVEN typecheckErrors: reads 1 on a planted finding
PROVEN lawLineCount: reads 3 on a planted finding
prove PASS — every counter read the finding its probe planted.Each configured counter's probe plants its input in a temp directory — never in your repo — and the
counter runs there exactly as it runs in a real measurement. The tools come from your repo's
node_modules/.bin, so what is proven is the toolchain you will actually measure with.
PROVEN <key>: reads N on a planted finding— the gate has been seen red.CANNOT FAIL <key>: read 0— the input was planted and the counter still read nothing.CANNOT MEASURE <key>: <error>— the probe threw, or the counter ships none.
The run stops at the first counter that reads 0 or throws, and exits 2. Put it in front of the ratchet in whatever script your CI runs:
{ "verify": "… && geonosis-ratchet --prove && geonosis-ratchet" }A counter you wrote yourself takes a probe beside its run: { input: (dir) => void, command?:
(dir) => string, params?, expect: number }. expect is what the planted input is worth, and it is
at least 1 — a probe that plants nothing proves nothing.
A probe proves the command, not the regex
A shipped probe plants an input in the shape the counter's OWN default reads. archViolations
plants a scanner that prints ✗ … lines and exits non-zero; sumOfCounts plants a line its default
match counts. That proves the command is spawned, its output is captured and its exit code is told
apart from its findings — and it proves nothing whatever about a match YOUR config supplies,
because the planted input was written for the default one.
This was measured, not imagined: a consumer's mis-escaped match read 0 from a real scan, --prove
said PROVEN off the shipped plant, and the ratchet lowered the baseline — ten pieces of debt erased
with a congratulatory message (#140).
So the counters that take a reading param — archViolations (match), sumOfCounts (match),
duplication (headings), todos (markers) — require a sample from the config that supplied it:
{
"counter": "archViolations",
"command": "node scripts/scan-architecture.mjs",
"match": "^VIOLATION ",
"probe": { "sample": "VIOLATION a\nVIOLATION b\nfine\n", "expect": 2 }
}Without the probe block the line reads UNPROVEN, naming the param and its value. With it, the
sample is run through YOUR match and must count expect — a sample your match cannot read is a
refusal that names the match, not a zero.
A baseline is not proof the gate still gates
The ratchet compares a number against a number. Swapping the tool that produces it — a vendored lint
plugin for a published one, one version for the next — can hold every number and still have stopped
firing, because a rule that fires nowhere counts zero exactly like a rule with nothing to find. Check
that separately, with @geonosis/lint-parity:
npx geonosis-lint-parity \
--corpus node_modules/@geonosis/oxlint-plugin-biological-architecture/corpus \
--a oxlintrc.old.json --b oxlintrc.new.json --out proofs/reach
npx geonosis-lint-parity --a oxlintrc.old.json --b oxlintrc.new.json --out proofs/parity -- apps packagesReach first — parity between two configs that both fire nothing is perfect parity — then the tree.
--exclusive: one heavy run at a time on a machine
Three sessions on one laptop each started their own verify, and a 25-minute run took two hours and
came back with failures that were about the load, not the code. "Announce before anything heavy" is
a habit that works right up until the moment somebody is concentrating; --exclusive is the same
agreement as code.
geonosis-ratchet --exclusive --exclusive-timeout 1800It takes ~/.cache/geonosis/heavy.lock before measuring anything, writing the pid, cwd and start
time into it, and gives it back in a finally and on SIGINT/SIGTERM. A second run waits instead of
contending, and says every 15 seconds who has the lock and for how long. A holder whose process is
gone — killed, crashed, a lid closed mid-run — is stale, and is taken over with a printed note
rather than making everyone wait for a pid that no longer exists. The wait is bounded:
--exclusive-timeout seconds (default 1800) and then exit 2, because a run that hangs silently
looks exactly like a run that is broken.
Put it on the script your sessions actually run, so nothing has to remember:
{ "verify": "… && geonosis-ratchet --prove && geonosis-ratchet --exclusive" }GEONOSIS_HEAVY_LOCK moves the lock file — for a sandbox with no writable home, or for a test.
Config
geonosis.ratchet.json at the repo root names the counters and how to run them:
{
"baseline": "gate-baseline.json",
"counters": [
{ "counter": "oxlintErrors", "command": "npx oxlint --format=unix ." },
{ "counter": "typecheckErrors", "command": "npx tsc --noEmit" },
{ "counter": "cloneCount", "command": "npx jscpd .", "tiers": ["full"] },
{ "counter": "lawLineCount", "path": "CLAUDE.md" }
]
}key— the baseline key this entry writes. Defaults tocounter; required when one counter is used twice. Two entries writing one key is a config error, not a race.tiers— the tiers this entry belongs to. Absent means every tier, so a repo that never passes--tieris unaffected. Names are free-form;fastandfullare convention.okExits— the exit codes this entry's command finishes with when it has a reading. Each counter declares its own tool's set ([0, 1]for a linter that exits 1 over findings,[0, 1, 2]for tsc,[0]where nothing is declared); an entry overrides it for a command it knows better. Any other exit is "the tool did not run" and the counter is refused — never a number. So is a non-zero exit whose output is a Node stack trace, whatever the set says: a counter command that threw once had its trace's line numbers summed as findings (#18).
gate-baseline.json holds one number per key (plus any _comment, which is preserved):
{ "oxlintErrors": 21, "typecheckErrors": 0, "cloneCount": 40, "lawLineCount": 137 }A number that grew fails. A number that shrank rewrites the baseline down, in place, so the win is locked in by the same commit that earned it. Equal passes.
A number that fell while the measurement behind the baseline changed is UNPROVEN: the fall is
not banked, the baseline is left exactly as it was, and the run does not pass. A rise is
recoverable — it fails, someone looks, and the baseline's own header sanctions raising a number with
the reason in the commit. A fall is not: once banked, the lower number is the ceiling from then on,
and a fall that came from the instrument seeing less is never found. So the refusal names the way to
accept it — confirm the fall is real, then lower the key in gate-baseline.json by hand, with the
reason in the commit. "The measurement changed" is what gate-baseline.identity.json records beside
each number: the command the reading was taken with, the params the counter was configured with, the
build that took it, and a digest of every settings file its tool reads — a rule switched off in
.oxlintrc.json moves the instrument as surely as a new command does. The tool counters carry their
own reads globs (oxlint's configs and ignore files, tsconfig*.json, oxfmt's, jscpd's, knip's,
turbo.json); an entry's reads replaces them. A counter configured by params alone — one that
runs no command — is reconfigured by editing its options, so the options are part of what a fall is
vouched against; they are compared as configuration rather than as text, and the same options in
another key order are the same measurement. A fall whose identity is UNKNOWN — nothing recorded, an
unreadable file, a number edited by hand — is refused the same way: absence vouches for nothing. A
run that meets a hand-edited number exactly records the measurement behind it, so the next fall is
judged.
A counter outside the requested --tier is not measured and says so —
SKIP cloneCount: not measured by --tier fast. It never counts as held, never reaches the exit
code, and its baseline survives a rewrite another counter earned: printing OK for a census nobody
took is how a red main goes unnoticed for three commits.
A counter whose command cannot run is a hard error, never a silent zero. A gate that cannot measure
has not passed. Nor is a number believed on its own: oxlintErrors reads the finding lines AND the
summary oxlint printed, and raises when they disagree.
The warnings-only window, and expectFormat
The last-resort backstop under the oxlint counters — output in no shape we know and a non-zero exit means the format moved, so refuse — has one gap, and it is worth naming rather than pretending otherwise. oxlint exits 0 when a run found only warnings. On such a run the backstop cannot fire: an unrecognised format reads as zero warnings, and a zero that only shrinks is a zero the ratchet banks by rewriting the baseline down.
--deny-warnings is not the fix. It makes warnings exit non-zero by making them errors, which
moves every warning into oxlintErrors and changes what both numbers mean — a bigger lie than the
one it patches, told to the whole baseline.
The fix is to say which shape you asked for:
{ "counter": "oxlintWarnings", "command": "npx oxlint --format=unix .", "expectFormat": "unix" }With expectFormat set, output that matches none of the three shapes and is not empty once npm's
own chatter is stripped (> … script echoes, npm warn …, npm notice … update nags, blank
lines) is refused even on a clean exit. npm ERR! is not stripped — that is npm saying the
command never ran, and it must never read as zero. Unset is the default and changes nothing, so a repo that never opts in keeps exactly
the behaviour it had. Valid values: unix, agent, default.
When a number grew, up to ten lines of that counter's command output print under the
<-- REGRESSED line, indented — so the report says what grew, not only that something did. A
counter that reads a file rather than running a command (lawLineCount) prints nothing extra.
The lines are the ones the counter says its number came from, not the tail of the run. oxlintRule
cites the findings attributed to ITS rule; oxlintErrors cites errors and oxlintWarnings
warnings. The tail was wrong for exactly the run that needs it most: a consumer's +2 on one rule
printed two unrelated warnings as its evidence, because those were the last lines oxlint wrote
and the two real errors sat higher up. A counter written outside this package that names nothing —
or that recognises nothing in a run — falls back to the tail, as before.
testFailures reads the runner's summary, never the exit code
The counter looks for the runner's own count and refuses when it cannot parse one. It never
reads the process exit code, because a test runner exiting 0 over a red suite is commoner than
anyone expects: @cloudflare/vitest-pool-workers 0.22 on vitest 4.1 exited 0 with failing tests on
every workerd suite of a consumer, and every gate that trusted the exit code reported green over red
for weeks. A refusal is loud and stops the run; a trusted 0 is silent and banks the red as a win.
The lines it reads are the runner's summary lines, whole, and nothing else — two dialects, both captured from the binary rather than guessed at:
| Runner | The line | Read as |
| --- | --- | --- |
| vitest 3.2.7 · 4.1.11 | Tests 1 failed \| 1 passed (2) | 1 |
| vitest 3.2.7 · 4.1.11 | Tests 3 passed (3) · Tests 2 skipped (2) | 0 |
| vitest 3.2.7 · 4.1.11 | Tests no tests (a suite that threw before collecting) | refused |
| bun 1.4.0 | 1 fail on its own line | 1 |
| bun 1.4.0 | 0 fail | 0 |
Nothing else in the output counts. Test Files 1 failed (1) is a count of files, not of tests;
(fail) one [11.71ms] is bun naming one; naming 4 failing test(s) is some wrapper counting its
own findings. All three used to be read as a failure count — a green run regressing a baseline,
the false RED mirroring the false green above (backlog #44). A run that prints no summary line at
all is unmeasured, and the refusal names the two dialects so a third runner's output is a message
rather than a wrong number. Summaries add up: a command that invokes the runner twice prints
two, and reading the first and stopping banks the second suite's failures as a win.
For the same reason, give every test package its own testFailures entry, each with its own
key. One entry over one workspace measures one workspace; the suites it does not run are not zero
failures, they are unmeasured — and unmeasured reads exactly like green.
{ "counter": "testFailures", "key": "testFailuresApi", "command": "bun --cwd apps/api test", "tiers": ["full"] }examples/workers-app.ratchet.json is four test entries for that reason: apps/web plus the three
workerd packages that were previously outside every gate.
…or the runner's JSON report, which is stronger
report: "vitest-json" reads the runner's own machine-readable answer instead of its prose:
{
"counter": "testFailures",
"key": "testFailuresApi",
"command": "cd apps/api && bunx vitest run --reporter=json --outputFile={report}",
"report": "vitest-json"
}{report} is replaced with a path in a temp directory the counter makes and removes; give the entry
a reportPath instead when the report belongs somewhere your CI already collects. A command in this
mode with neither is refused, naming what is missing — a run whose report goes nowhere is a run
nobody can read.
The counter then reads numFailedTests, and refuses rather than returning a number when:
- the file is not there — a crash before the reporter wrote is not a pass, and it prints no summary line either, so the summary parser had nothing to refuse on;
- the file is not JSON, or has no
numFailedTestsin it; - the report says
success: falseand names 0 failing tests — the shape a pool that dies mid-run writes. The run did not finish, so there is no number to bank.
The summary mode stays the default; nothing changes for an entry that does not ask for a report.
--prove proves all three: testFailures ships one probe per reading mode — summary (vitest),
summary (bun), vitest-json — and prints them by name (PROVEN testFailuresApi (vitest-json)),
because proving the mode nobody configured says nothing about the mode they did, and a summary
dialect nobody proved is a dialect nobody has been shown to read.
oxlintRule counts a warned rule twice, on purpose
A rule parked at "warn" as ratcheted debt appears in two numbers: once inside oxlintWarnings,
and once under its own oxlintRule key, which measures it at "error" through a strict temp copy of
the config. So arming a rule at warn raises both, and fixing one finding lowers both.
That is not double-counting the total; the second key exists so the debt is visible per rule instead
of hidden inside a lump sum that a different rule's warning could mask. A Medusa storefront today: oxlintWarnings
12 → 135 when no-raw-html-atoms was armed, with its own key at 123.
--prove also proves the lock
--exclusive is a claim about the machine, so it is measured on the machine, every --prove:
PROVEN exclusive: two runs of 400ms serialised, the second starting 161ms after the first finishedThe self-test runs two children of this CLI over --hold <ms> — an instrument that takes the lock,
says when it started, waits, says when it finished, and gives the lock back — and requires the second
to have started after the first finished. Anything else prints CANNOT FAIL exclusive: interleaved
and exits 2. It uses a lock file of its own in a temp directory, so proving the lock never takes the
real one out from under the runs it exists to serialise.
It is here because the lock did not work for two releases and every gate was green throughout:
openSync(path, 'wx') is atomic about the NAME and not about the holder, and a run polling in that
window read an unparsable lock, called it stale, and took it. A lock nobody has watched fail has not
been shown to work — and this one guards the running time of every other gate.
Running it where it will actually run
Counters run in the caller's environment. They inherit the shell the ratchet was started in, nothing more. If a child command needs
NODE_OPTIONS— a TypeScript shim, a loader — put it on the script that invokesgeonosis-ratchet, not on the counter's own line, and not only in your interactive shell.NODE_OPTIONSin the parent script + a counter that shells out topnpm= the counter dies. The nestedpnpmINHERITS the option. A Medusa storefront preloads a TypeScript-5 shim throughNODE_OPTIONSin itslintscript; run the ratchet from inside that script and the nested pnpm goes looking for a.pnpmfile.mjsthat is not there and exits non-zero. The ratchet refuses — correctly, a command that cannot run is never a silent zero — but nothing in the message is near the cause. Run the ratchet outside that script, or unsetNODE_OPTIONSfor the nested call:{ "counter": "archViolations", "command": "env -u NODE_OPTIONS pnpm --silent verify:arch" }geonosis-doctor --only driftasks for this shape by name: a manifest script carryingNODE_OPTIONSbeside a counter whose command shells topnpm.A path in a counter's command must be absolute, or
--provecannot run it. A probe runs in a scratch directory with none of your repo in it, so a loader or shim named relatively —node --import ./scripts/ts5-shim.mjs …— resolves to nothing there and the counter comes back "command did not run". The command is the same string in both places; write the path as$PWD/scripts/ts5-shim.mjsand it works in the repo and in the probe alike.A
typecheckErrorscounter needs the same precondition CI gives it: build the workspace packages first. In a fresh worktree thedist/*.d.tsfiles do not exist yet, andtscreports a false +N of missing-module errors that has nothing to do with the change under test.pnpm build && geonosis-ratchetis the shape; a baregeonosis-ratchetin a clean checkout is not.The counter no longer reports that number. When tsc's TS2307 or TS2305 diagnostics name a package of this workspace — read from
pnpm-workspace.yaml, orworkspacesin the rootpackage.json— it refuses:geonosis-ratchet: typecheckErrors: workspace packages not built: @shop/shared — build them before measuring typecheckErrorsMixed output refuses too. Counting the readable half would produce a number that is neither the debt nor the build error, banked into the baseline as though somebody had measured it. An unresolved module that is NOT a workspace package —
lodash— is counted as before: that is a dependency to fix, not a build to run.A release-age cooldown will refuse a package published minutes ago. If pnpm's
minimumReleaseAgeblocks@geonosis/*, exclude the scope for that project — pnpm writesminimumReleaseAgeExcludeintopnpm-workspace.yaml— rather than lowering the cooldown. The cooldown is protecting every other dependency in the tree; the exclusion is scoped to the one you chose to trust.Bumping to a new version: keep BOTH exclusion pairs — old and new — through the install, then drop the old one and re-verify. The policy check is applied to the entries in the LOCKFILE, not only to what you asked for, so removing the old pair first makes the resolver walk back through the version you are leaving and refuse it — an install that fails for the version you are replacing, not the one you are taking. The order that works, learned under load on a Medusa storefront's 0.2.0 bump:
# both pairs in pnpm-workspace.yaml's minimumReleaseAgeExclude pnpm install grep -c "@geonosis/.*@<old>" pnpm-lock.yaml # must be 0 before the old pair goes # now delete the old pair pnpm install --frozen-lockfileAn exclusion left behind is a cooldown quietly off for a package nobody is watching any more, so the deletion belongs in the same commit as the bump.
That policy check walks every entry in the lockfile and looks hung. It is not. ~2,300 entries took ~90 seconds on a loaded laptop with no output while it worked. Wait for it rather than killing the install and re-running into the same wait.
Counters
Every counter takes its command from the config, so the toolchain stays the repo's own.
| id | counts | params |
| --- | --- | --- |
| oxlintErrors / oxlintWarnings | findings under any shape oxlint prints — --format=unix, the compact agent format, the graphical default — cross-checked against the tool's own summary. ruledSeparately names rules ratcheted under their own oxlintRule key, which this count leaves out: turning a stricter rule on means one new oxlintRule entry (adopted once, then shrink-only) while the aggregate holds. A rule named here that no oxlintRule entry counts is refused | command, expectFormat, ruledSeparately |
| oxlintRule | one named rule's findings, counted only on lines the run reported as findings and refused when none of them attributes itself readably; with config, after forcing the rule to error in a temp copy — "warn", "off" and the ["off", { … }] array form alike — so debt cannot grow behind a downgrade | rule, config, command, expectFormat |
| typecheckErrors | error TS occurrences, refusing when TS2305/TS2307 name a workspace package of this repo — an unbuilt sibling is a missing build, not debt | command |
| testFailures | the runner's own failure summary — or, with report: "vitest-json", numFailedTests out of the JSON report the command wrote; throws when nothing is readable, when the report is absent, and when the report failed with nothing failing | command, report, reportPath |
| unformattedFiles | paths --list-different names that exist on disk | command |
| cloneCount | jscpd's Found N clones, with ignore kept out of the scan — ["**/package.json", "**/CHANGELOG.md"] by default, because npm fixes a manifest's shape and a release tool writes a changelog, so neither is duplication anyone can pay down. It is passed as jscpd's -i, which REPLACES the ignore of a -c config rather than adding to it: exclusions kept in .jscpd.json belong here instead, and "ignore": [] scans everything. formats names the jscpd formats read — ["javascript", "jsx", "typescript", "tsx"] by default, because prose and captured tool output are not code anyone deduplicates; "formats": [] reads every format | command, ignore, formats |
| knipIssues | the total under EVERY heading knip's compact reporter emits — the shape <Title> (N), not a list of titles, so a heading the kit has not heard of is debt rather than a silent zero. headings narrows it, and every printed heading the narrowing leaves out must be named under excusedHeadings with a reason or the run refuses | command, headings, excusedHeadings |
| boundaryIssues | N issues found from a boundary scan | command |
| archViolations | lines matching a marker your own architecture scan prints, refusing a non-zero exit that printed none of them — a scan that could not run is not a clean scan | command, match |
| sumOfCounts | the total of one capture group across a per-file census (grep -rc) | command, match |
| lawLineCount | the lines of the law file — a ceiling that can only come down | path |
| probelessRules | the rules an oxlint config ENABLES that the plugin it loads declares no probe() for. Those are the ones geonosis-doctor's exercised can only answer UNJUDGED about, which #137 makes a WARN naming the kit as owner — and a WARN the kit carries across releases is a downgraded rule by another name. The plugin is imported from the tree it is pointed at, and one that cannot be loaded is a refusal, never a zero | config, plugin |
| runtimeCodeShipped | 0 when a change shipped runtime code, 1 when it shipped none | command, patterns |
| disabledCiJobs | lines of if: false across the workflow files — a job switched off to get a release through, still off. A condition that merely mentions false is not one. No workflows directory at all reads 0 | dir |
| bundleBytes | one integer out of whatever your sizing command printed, separators and all; with match, the group that pattern names rather than the last integer, refusing when it matches nothing. Tolerates. | command, match, tolerance |
| fastTierMs | finishedAt − startedAt from the gate report geonosis-verify wrote, refusing a report of another tier rather than timing the wrong gate. Tolerates. | report, tier, tolerance |
| testsWithoutRunner | workspaces holding *.test.*, *.spec.* or __tests__/ with no test script — the suites nobody runs, which read exactly like suites that pass | script |
| packagesWithoutTypecheck | workspaces with no typecheck script | script |
| walkFindings | the defects in the report geonosis-walk wrote, over every page; with classes, only those classes, refusing a class the walk does not have. A missing or unparsable report is a refusal — the walk writes none when it could not run | report, classes |
| orphanTodos | debt markers outside the plan graph: one naming no plan, and one naming a plan that is not in the plans directory. Leading zeroes are a spelling, so (021) finds 21-…md. Fixtures, dependencies and build output are not read | markers, plans, roots |
| commentLines | the source lines carrying a comment, as oxc-parser lists them: a * inside a template literal or JSX text is not a comment, a // why after code is one, and a line is counted once however many comments it holds. A file the parser refuses, or a .vue, .svelte or .astro file, stops the count and is named. Fixtures, dependencies and build output are not read, and ignore names the generated trees whose directory is not called dist — a path relative to the repo root, empty by default, because a number that moves with whether the build has run is measuring the build | roots, ignore |
| suppressionCount | the eslint-disable, oxlint-disable, biome-ignore, @ts-expect-error and @ts-ignore markers in code files — law 3 as a number that may only shrink, for a repo that arrives carrying hundreds and cannot adopt it as a wall on day one | roots |
| adoptionSeamLines | the lines of the files geonosis.json → adoption.seams names: what a consumer still writes to hold a floor's shape, which adoption brings down. No seams declared reads 0 | — |
| floorSurface | the names a floor package's published .d.ts files export, read off the entries its exports map promises — every door is a line to learn and a promise to keep. A map promising no declaration, or one that is not built, is a refusal, never an empty surface | dir |
| doctorWarnings | the distinct WARNs of geonosis-doctor --json — one per check and rule, however many configs repeat it — refusing output with no findings list: a doctor that could not examine the tree is not a clean one. Its FAILs are the doctor's own exit code, not this number | command |
| bumpDurationMs | how long the last geonosis update took, from the bump-report.json it wrote; a repo that never ran one is a refusal, not a bump that took no time. Tolerates. | report, tolerance |
| timeToGreenDoctorMs | how long that bump took to reach a green doctor — the part a consumer feels. A bump whose doctor was red, or never ran, has no time-to-green. Tolerates. | report, tolerance |
| openBacklogRows | the rows of a register not marked done. Report it; never ratchet it — it rises every time an adoption is measured honestly, and a shrink-only gate over it would teach a repo to file less | path, done |
tolerance, and the counters that accept one
Most counters count findings, where one more is one too many. bundleBytes, fastTierMs,
bumpDurationMs and timeToGreenDoctorMs measure something that moves on a dependency patch or a
slower machine nobody chose, and a gate that fails on +40 bytes is a gate that gets switched off
within a week. tolerance is the fraction of the baseline such a number
may drift up before the ratchet calls it growth:
{ "counter": "bundleBytes", "command": "…", "tolerance": 0.1 } // +10 % is noise, +11 % is debtOnly a counter that declares tolerates accepts one, and only those four do.
Anything else stops the run:
geonosis-ratchet: "oxlintErrors" does not accept a tolerance — only a counter that measures a
quantity declares oneThat refusal is the point. A band around a findings count is not a tolerance, it is the gate turned
off from the config file: "tolerance": 100 on oxlintErrors and the run still prints PASS. It is
law 2 — never downgrade a rule — wearing a friendlier name, and it would be the easiest edit in the
repo to get past a reviewer. Whether a number is a quantity or a count is the counter's to know, not
the config's, so the counter declares it.
It forgives noise upward, never downward: a shrink of any size is still a shrink and still lowers
the baseline, or the number would stop following the artefact down. A tolerance that is not a
non-negative number stops the run and says which entry it was — one that silently became NaN
would make every comparison false, which reads exactly like a counter that can never grow.
The geonosis Claude Code plugin puts any write that introduces or raises one in front of a human
(permissionDecision: "ask"): a band somebody chose and a band somebody's agent chose are not the
same thing.
Tolerance is also why a rewrite lowers only the numbers that shrank. While held implied
now === baseline, writing every measured key back was a harmless no-op; with a tolerance it is
not, and a counter that grew inside its tolerance would have had that growth laundered into its new
floor by the first unrelated win.
Two example configs from real repos live in examples/ in the repository.
Programmatic use
import { COUNTERS, formatReport, runRatchet } from '@geonosis/ratchet'
const result = await runRatchet({ counters: COUNTERS, cwd: process.cwd(), tier: 'fast' })
process.stdout.write(formatReport(result))counters is a plain array, so a repo can add its own { id, run } beside the built-ins.
Two things this cannot see about itself
A baseline is lowered in place when a number shrinks; that is what locks a win in, and it is
also what lets a branch write any number it likes and stay green on every gate it runs. And a
workspace whose suite no testFailures entry covers is unmeasured, which reads exactly like green.
@geonosis/doctor asks both from outside — the
baseline at HEAD against another ref, and which workspaces have a test script nothing reads a report
from:
npx geonosis-doctor --baseline-against origin/main --strictApache-2.0.
The pairing rule (#136)
A scoped fast tier (changed files only) is safe exactly when the full tier is TOTAL: every hard cap
— max-lines, bundle bytes, suppression counts — needs a full-tier counter watching the whole tree,
or it is a cap in prose that a file can sit over indefinitely. oxlintRule, suppressionCount,
bundleBytes and friends exist to be that counter.
