npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@dzhechkov/harness-core

v0.8.46

Published

Shared harness logic - skill loading, additive apply, and the init/sync/verify/doctor operations.

Downloads

5,631

Readme

@dzhechkov/harness-core

Shared logic for the DZ harness — the engine behind @dzhechkov/harness-cli and any other consumer.

One artifact for signing and publishing

Sign, publish and the sibling-drift gate now pack with one function, packArtifact: it pins workspace dependencies, removes prepublishOnly for packing, restores the original package.json, and inventories the resulting tarball. Signing hashes the extracted artifact. A blocked drift names up to five changed files, the remaining count and the inventory source (pack-artifact) in both the CLI message and audit detail. This addresses the 2026-09-24 incident in which different packers produced LICENSE/package.json drift and an unnamed “1 file” refusal. Supported platforms are Linux and macOS with GNU/BSD tar; Windows is not claimed, and macOS verification is manual rather than an OS-matrix CI claim. Staging is synchronous within one process; concurrent packing of the same package by multiple processes needs external coordination (no cross-process lock is added).

Generation guard for backups

The repository helper .claude/helpers/generation-guard.cjs compares backup counts without a build step. Its verdicts are first-generation (no previous count), ok (equal or growing), collapsed (any decrease without a receipt), authorized (a decrease with a matching receipt), and receipt-invalid (a present receipt is malformed or does not match the new count). Every generation decision prints a line with previous, next, and verdict.

For an intentional shrink, write an operator receipt with this shape:

{"at":"2026-09-05T04:41:00Z","reason":"harmonize dedup","before":739,"after":18}

at is an ISO-8601 timestamp, reason is non-empty, counts are non-negative integers, before >= after, and after must equal the new count. A malformed receipt refuses even an equal, growing, or first generation. There is no environment bypass.

scripts/backup-backlog.sh reads <source>/shrink-receipt.json and compares the source count with the cloned backlog-manifest.json before removing the cloned backlog files. A refusal exits 6, preserves that generation, and pushes nothing. An authorized push records shrinkReason in the manifest and renames the source receipt to shrink-receipt.<stamp>.used.json. DZ_BACKLOG_BACKUP_SRC selects the source directory (default .dz/backlog); it does not disable judgement. A missing or unrunnable guard exits 3.

.claude/helpers/brain-checkpoint.cjs export --json exports to .agentic-qe/aqe.rvf.tmp, reads the previous count from aqe.rvf.meta.json, and judges before renaming the RVF and its idmap. Exporter failures, missing pattern counts, and refused generations preserve the previous RVF. Refusals return exported:false, reason, previous, and next with exit 0; an invalid receipt also reports verdict and why. The receipt is aqe.rvf.shrink-receipt.json; authorization renames it to aqe.rvf.shrink-receipt.used.<stamp>.json. Successful exports write {at, patterns, bytes} to the sidecar, and verify --json includes its patterns when available.

An authorized shrink follows consume → swap → finalize: first rename the receipt to aqe.rvf.shrink-receipt.consuming.json, then replace the RVF/idmap and write meta, then rename the receipt to aqe.rvf.shrink-receipt.used.<stamp>.json. receiptConsumed:false with reason:'receipt-consume-failed' means the previous generation was preserved; receiptConsumed:true records that authorization was spent. A swap failure returns exported:false and receiptSpent, the receipt path after a best-effort rename to aqe.rvf.shrink-receipt.spent-on-failure.<stamp>.json (the consuming path if that rename fails). The swap is not atomic across the RVF, idmap, and meta files. A finalize failure still returns exported:true, with receiptConsumed:true and receiptFinalized:false, and logs the retained receipt path. Non-authorized runs keep their existing JSON keys. Each receipt authorizes exactly one shrink attempt; a failed attempt requires a fresh receipt instead of restoring the spent authorization.

Test execution

npx vitest run uses two projects and returns one combined verdict: parallel runs the ordinary suites concurrently, while serial runs process-spawning and real-time suites one file at a time. The serial paths in test/serial-suites.txt are regenerated from test/serial-suites-census.test.ts, which scans test sources for process and timing markers, including execSync( and execFile(, and fails when the list and census differ.

The census reads markers from comment-stripped code: the test-side helper test/helpers/serial-census.ts (maskTsComments / maskReport / spawnMarkersIn, shared with the harness-cli twin by relative import; it needs the typescript devDependency, so it is NOT a src/ export) asks the TypeScript parser where the comments are and blanks them to same-length spaces, keeping line terminators; string, template and regex contents stay verbatim (markers such as dist/bin.js live in strings by design). FAIL-SAFE (AM-5): a file the parser cannot parse is treated as all code (parseOk:false), so the only possible error is an unmasked comment that over-serializes one suite — never a hidden spawn. dist/bin.js counts only when the same file also holds an import STATEMENT of child_process at a line start (ESM import … from or const … = require(; a fixture string that spells one is data) or of a test helper — a relative import whose path carries a helpers/ segment at any depth (./helpers/x, ../../helpers/x, ./util/helpers/x; not ./helpersX/, not a bare helpers package) — otherwise it is a fixture path, i.e. data. MEASURED 2026-09-25 (backlog ffe4e076): the raw scan had listed 5 comment-only suites in harness-cli and 2 data-only suites here as serial; two of the flipped files spawn through a production helper and are now named by that helper's call (spawnRoundCodex(, defaultIntegrationProcessPort.run() instead of by a comment.

Full-suite worker ceiling (CORE_MAX_WORKERS, vitest.config.ts)

The root test block caps maxWorkers at CORE_MAX_WORKERS (2, minWorkers: 1), so npx vitest run with no flags is safe by default. The ceiling lives on the root test block, not on the parallel project's poolOptions — an earlier version of this config set poolOptions.forks.maxForks on the parallel project instead, and that was a false guarantee: a fix-round measurement (2026-09-16) compared process names (node (vitest N), polled from /proc/<pid>/cmdline) over the same 30-file parallel set and saw names vitest 1..vitest 7 (14 workers observed) under the project-level poolOptions, against never more than vitest 1/vitest 2 under either a --maxWorkers=2 CLI flag or maxWorkers on the root test block. Vitest 3.2.4 simply does not honour poolOptions.forks.maxForks set on a project the way it honours the root-level knob (or the equivalent CLI flag) — full proof in features/core-suite-memory-ceiling/07_code_changes/change_manifest.md, section "Фикс-раунд 1".

The number itself is not "8 minus a guess" — it is arithmetic from a MEASURED trace (see the comment above the constant in vitest.config.ts): the memory pressure is NOT the embedding model loaded inside the vitest worker (that guess is REFUTED), it is a CHILD process that some tests spawn per file (an embedding daemon, or a dz teach invocation). On 2026-09-16 three full-suite runs were traced with the same instruments, and each number below says WHICH run produced it, because two of the three runs did not have a ceiling that actually bound:

| run | ceiling | tree peak | minimum free | result | |---|---|---|---|---| | 06:24–06:29 | --maxWorkers=2 CLI flag (binds) | 5721 MB | 2995 MB | green, 7018 passed, 281 s | | 06:46–06:49 | poolOptions on the parallel project (does NOT bind — effectively unbounded) | 8207 MB | 804 MB | green, but see below | | 07:01–07:06 | root test.maxWorkers (binds), no flags | 5383 MB | 3804 MB | green, 333 files, 7024 passed, 275 s |

The last row is the profile of the shipped configuration — the command a person actually types. The middle row is the DEFECT being measured, not this configuration, and it is where the per-process-class peaks come from: a vitest process 2324 MB, an embedding-daemon child spawned by a test 2321 MB, a child dz teach --from-json 1786 MB, with up to 4 daemon children alive at once (3917 MB combined). Those per-class numbers are real, but quoting them as the memory profile of the 2-worker run would be a misattribution — a Codex round-2 finding, fixed here.

The load-bearing fact stays: the memory is NOT the embedding model inside the vitest worker, it is in the CHILD processes the tests spawn, and the ceiling bounds how many worker-plus-child pairs are alive together. CORE_MAX_WORKERS must not be raised without a fresh trace of the last shape.

vitest.config.ts also fails LOUD at config-load time if CORE_MAX_WORKERS is ever set to something other than a positive integer (0, negative, or fractional) — a Codex fix-round finding that a ceiling accepting those values is not a ceiling at all.

test/mutation-registry.json's maxWorkers is kept equal to this same constant (test/suite-worker-ceiling.test.ts reddens if either drifts from the other).

findExactLesson(records, text, domain?) finds the earliest lesson whose trimmed, whitespace-collapsed text matches exactly (case-sensitive), optionally within one metadata domain, and reports whether that existing lesson is quarantined.

Карантин: источник правды — лексический стор; зеркало — проекция; dz vector reindex пересобирает зеркало из лексических записей и восстанавливает паритет меток. Прямая правка зеркала не меняет авторитетное состояние и при следующем перестроении будет утрачена.

countLearningStoreRowsReadonly() reports the mirror as TWO figures, deliberately not one. vectorRows stays the whole mirror (lessons + backlog ideas + book units) because the store guard reads it as an integrity signal against a recorded high-water mark; narrowing it would present a healthy store as a collapse. vectorLessonRows counts mirrored LESSONS only (dz-teach and dz-learning) and is the figure comparable with lexicalRows. It survives the metadata fallback path whenever task_type is still readable, and is absent when the mirror cannot be decomposed — absent means "unknown", which the panel must report rather than treat as agreement. If a readonly guard count meets another SQLite writer, the guard reports busy: the write is not refused, health is explicitly not measured for that run, and the high-water mark does not move. statuslineData().patternMirror is absent on equal counts AND when there is no mirror at all — nothing to compare, and a permanently-lit indicator on every project without a vector tier would carry no information. It is { state: 'different', lexical, vector } on divergence, and { state: 'unavailable' } only when the mirror EXISTS but cannot be read or decomposed: that is a tool failure and it is worth saying out loud. statuslineData().brainKuCounts carries the KU volume of each brain source in brain order; an empty array means the volumes could not be listed, never that the sources are empty.

Core boundary checks

Run npx vitest run test/core-boundary.test.ts from this package. Rule A scans top-level src/*.ts (excluding *.generated.ts) for process.argv and process.exit in code using TypeScript's AST. Comments and literal text are excluded; expressions inside template interpolations are code. The scanner is internal to this test and is not exported from index.ts.

rule A: debt remains red while existing violations remain; it has no exceptions. rule A: scanner separately compares the measured locations with the pinned requirements list. The corrected list contains two debts: brain.ts:995 and integration-probe-worker.ts:304. Measured with npx vitest run test/core-boundary.test.ts -t 'rule A: (scanner|debt)': the scanner test passes and the debt test fails with those locations. integration-probe-worker.ts:308 uses process.exitCode, and setup.ts:172 is template text; neither violates rule A. No debt is repaired by this change.

The IO ratchet pins 57 files / 66 imports in test/core-boundary-ratchet.json, measured with npx vitest run test/core-boundary.test.ts -t 'IO ratchet'. It counts import declarations, import-equals, dynamic imports and require(...) for node:fs, node:child_process, node:https, their bare forms (fs, child_process, https) and subpaths (such as node:fs/promises), excluding mentions inside comments and strings. The expanded set was remeasured and the totals remain unchanged on this tree. Both totals may decrease; neither may increase. Updating the baseline requires an explicit edit; tests never rewrite it. A source file that cannot be read aborts the measurement instead of counting as zero.

Subpath membership has a separate test, IO imports: subpath-only source belongs to the IO set: the growth ratchet alone cannot detect an undercount. The current src/loop-lint.ts has no imports from the configured IO modules, so it cannot serve as a live witness for subpath membership. The causal probe replaces specifier.text === name || specifier.text.startsWith(`${name}/`) with specifier.text === name in a temporary copy of the scanner and runs npx vitest run test/core-boundary.test.ts -t 'IO imports: subpath-only|IO ratchet'. The membership test fails, the ratchet passes; restoring the scanner makes both pass.

Rules B and C are measurement only: the test prints direct child_process/https imports in harness-cli/src/cli.ts; a pure directory has not been designated, so rule C is not measurable. Neither rule is enforced by assertions.

Mutation entries rule-a-tokenizer-not-regex and ratchet-refuses-unreadable were both PROVEN (respectively 2 and 1 failing tests under mutation) after fresh core and CLI builds:

npm run build && npm --prefix ../harness-cli run build
node ../harness-cli/dist/bin.js mutation-gate --only rule-a-tokenizer-not-regex,ratchet-refuses-unreadable --test-cmd "npx vitest run test/core-boundary.test.ts -t '^(?!.*rule A: debt)'" --json

Only the intentionally red debt test is excluded from that gate's baseline. The scanner accuracy test remains included.

Shared Markdown masking

maskMarkdown in src/markdown-masker.ts is the single implementation used by amendment-trace and swarm-brief; the standalone plan-completeness gate carries a byte-identical .mjs copy. It blanks fenced blocks and HTML comments while preserving UTF-16 offsets and newlines. It is not a complete CommonMark parser. Reader policies remain explicit: amendment-trace restores unclosed blocks; swarm-brief and K2 hide them through EOF. Brief also retains list barriers and nested-comment ambiguity diagnostics through the line callbacks. Its inlineComments policy also retains comments after prose, while paired backtick runs on the same line protect code-span delimiters.

The four-space indented-code gap is not closed for amendment-trace or K2 in this feature. Its single implementation address is src/markdown-masker.ts. Contrary to the original plan's premise, swarm-brief already masked indented code; its indentedCode option preserves that behavior. The default remains off. Future work must decide the other readers' policy at this one address. Regenerate every gate copy from this source and run npx vitest run test/markdown-masker.test.ts from this package; byte equality is tested, including the installed and packaged gate locations.

Text-mangling warnings (text-mangling.ts)

detectMangledText(text, kind) is a pure, I/O-free detector exported with MangleKind and MangleSymptom. It returns possible empty-substitution-hole, dangling-arrow, empty-brackets and short-for-kind symptoms in offset order, with UTF-16 code-unit offsets and excerpts of at most 40 code units. The trimmed-length thresholds are 40 for teach and 20 for backlog. Literal backticks and $( are not symptoms by themselves. These are advisory warnings, never refusals; the CLI shares the detector across teach and backlog add/edit, including file input.

Memory index guard

checkMemoryIndex({ indexText, files, limits? }) checks a memory index against its sibling Markdown files without I/O: a 24 000-byte UTF-8 budget, 200 UTF-16 code units per line, broken and duplicate links, unindexed files, and unsupported hooks. Hook support is a heuristic: at least half of the Unicode letter/number tokens of length four or more in the hook's first clause (before ;) must occur in the linked file, ignoring case; hooks with fewer than three tokens are not judged. The named defaults can be overridden through limits. test/memory-index-guard.test.ts always runs pure fixtures; its live check reads roam/claude-state/memory/MEMORY.md and sibling *.md files, prints measured bytes, lines and finding count, and requires no findings. When that index is absent (as in CI), the live case skips with the full path and reason in its title and one warning line. Neither the checker nor the live test writes to memory. Support is judged on stems from stem.ts: находку/находка → находк and подписью/подпись → подпис, while проверил ≠ проверка and teacher ≠ teach remain distinct. Measured 2026-09-24 on the live index: the stem rule exposed one hook the substring rule had matched inside other words (fixed in the memory file), then 0 findings with a minimum support of 0.5 on two lines (a stale "five" from the 23.09 report was corrected after QE re-measured it).

Store-guard mark pruning (store-guard-prune.ts)

planStoreGuardPrune(entries, {exists, tmpDirs}) is the pure half behind dz store-guard --prune. It sorts every mark file in ~/.dz-store-guard into four buckets and computes the reclaimable bytes: stale-temp (the recorded project path is gone AND lies under a temp root — the only bucket the command ever deletes), live, gone-outside-tmp and unreadable. The classification order is fixed — unparseable first, then existence, then the temp-root test — and "under" is separator-aware, so /tmpfoo is not under /tmp.

Two input hazards are handled in this module because they decide whether files are deleted, and the pure half is where that is testable:

  • a degenerate temp root ('/', '') authorizes nothing and is dropped. TMPDIR=/ makes os.tmpdir() return the single character / — node strips a trailing slash only when the path is longer than one character — and a prefix test against / would match every absolute path, deleting exactly the gone-outside-tmp marks the bucket exists to protect;
  • the caller passes both the raw and the canonical form of each temp root, because a mark records the project root as resolved, not as realpath'd; on a machine where the temp root is itself a symlink a canonical-only list matches nothing and the command silently does nothing.

Both are pinned by tests in test/store-guard-prune.test.ts; the deletion side, the dry-run default and the symlinked-directory refusal live in the CLI and are pinned in harness-cli.

Per-turn admission debt (session-retro.ts)

The engine behind dz retro and dz retro --scan-tail: it turns a session transcript into events, detects recurring PROCESS rakes, and folds an admission debt — an error narrated in chat with no dz teach behind it.

Public surface used by the CLI: streamSessionEvents, parseSessionJsonl, detectProcessRakes, buildRetro, renderRetro, renderDrill, retroLessonText, foldAdmissionDebt, runRetroTailScan, retroSentinelIsFresh, renderRetroDebtDirective, findLatestTranscript, resolveScanTailTranscript (new), plus the types SessionEvent, RetroPendingSentinel, TailScanOutcome and ScanTailSource (new), and the constants RETRO_DEBT_MARKER, RETRO_PENDING_FILE, RETRO_SCAN_STATE_FILE, RETRO_SCAN_LOCK_NAME, PROCESS_SIGNATURES, DEFAULT_DRILL_THRESHOLD; since retro-debt-sentinel-per-session also the pure resolver retroSessionPaths(dzDir, sessionId) (with safeSessionDirName, sessionIdFromTranscript, the type RetroSessionPaths) and the constants RETRO_SESSION_DIRNAME, RETRO_SESSION_PENDING_BASENAME, RETRO_SESSION_STATE_BASENAME, RETRO_SESSION_STALE_MS.

One directory per session (retro-debt-sentinel-per-session). The sentinel and the scan bookmark live at .dz/retro/<session>/pending.json and .dz/retro/<session>/scan-state.json, where <session> is the transcript basename without .jsonl (Claude Code names the transcript after the session id, so it equals the hook payload's session_id); an id that is not a plain token — path-like, empty, over 128 chars, . or .. — is replaced by the first 32 hex of its sha256, so nothing can escape .dz/retro/. A scan writes and removes files only inside its own dir; the pre-feature branch that unlinked a sentinel of ANOTHER transcript as "foreign" is gone — with two sessions in one worktree it deleted the neighbour's LIVE debt, and the shared bookmark made every turn a bounded re-scan from 0 (MEASURED, backlog 58f3c56fbb9d6893). TailScanOutcome now carries sessionId and pendingPath. Legacy adoption, once: the flat .dz/retro-pending.json / retro-scan-state.json (the names RETRO_PENDING_FILE / RETRO_SCAN_STATE_FILE still export) are adopted by a scan only when they name THIS transcript and the session has no per-session copy yet, and are unlinked only after the per-session copies are committed; a stranger's, or a nobody's (no transcript field), is left byte-identical. Bounded sweep: after each scan, .dz/retro/*/ dirs whose scan-state.json is older than RETRO_SESSION_STALE_MS (7 days) are removed — own dir excluded, undatable dirs kept, at most 64 entries examined, every error swallowed. The recall hook reads only <session root>/.dz/retro/<session_id>/pending.json; a payload without a string session_id yields '' before any fs call — the hook guesses no session (test/retro-sentinel-per-session.test.ts, registry id retro-sentinel-path-is-per-session).

SessionEvent carries an optional toolUseId — tool_use.id on a call, tool_result.tool_use_id on its result — which is the pairing key the debt fold needs to tell WHICH command a result belongs to.

Load-bearing properties, each pinned by a test that goes RED when the property is mutated out (test/session-retro.test.ts, test/retro-scan-tail-source.test.ts, registry ids in test/mutation-registry.json):

  • Block order survives parsing. An admission text block is emitted BEFORE the tool_use of the same message, so admitting and teaching in one turn reads as settled (retro-p1-1-block-order).
  • Only an executed Bash teach can pay. A tool_result echoing the phrase, an echo/grep decoy, or a non-Bash tool call never settles anything (retro-p1-2-bash-only-teach).
  • A newline is a command boundary. The dominant field form (DZ=…\n$DZ teach … --project $B) pays; every decoy still stays armed (retro-adr5-newline-boundary).
  • The RECEIPT settles, not the command text. A teach with a tool_use_id is registered by its call and cleared only by that call's own result carrying a line dz teach prints on a real write — so exit 0\ndz teach "never runs" pays nothing (retro-r1-teach-receipt-required).
  • Every awaiting teach is retained. Two parallel teach calls cannot cancel each other out; if both results come back receipt-less the debt stays armed (retro-r3-retain-awaiting-teaches).
  • An admission is a confession, not a bug-fix report. The Russian branch is an allowlist of verb and adverb forms, so «Я ошибку валидации исправил» — a NOUN in a completion report — arms nothing (retro-r3-admission-verb-only).
  • The scan never guesses its transcript. resolveScanTailTranscript takes an explicit flag, then a positional path, then the Stop hook's stdin transcript_path; with none it returns {path: null, reason} rather than the newest file on disk.
  • No lost update. The whole read→fold→write tail-scan transaction runs under the retro-scan named lock beside the store it guards; contention advances nothing (retro-p1-3-unlocked-scan).

Lesson payoff (bandit re-rank)

lesson-bandit.ts / lesson-payoff.ts add a payoff axis to lesson recall: a Beta posterior per (domain, lesson) that answers "has this lesson ever actually helped?" — the question neither cosine similarity nor SAFLA-delta asks.

Public surface: contextKeyFor, classifySignal, makeRewardEvent, recordReward, recordExposures, payoffTermsFor, narrowBanditReport, banditStats, renderBanditHealth, resolveBanditConfig, and the vendored LessonBandit engine.

Load-bearing properties, each pinned by a test that goes RED when the property is mutated out:

  • Disarmed by default. memory.learning.banditRerank absent ⇒ the module is not even loaded and ranking is byte-identical.
  • A view is not a reward. kind:'recall-hit' is recorded as an EXPOSURE; only an explicit confirmation moves the posterior.
  • Quarantine-closed. Quarantined lessons never receive trial impressions unless banditExploration is armed explicitly — that flag weakens an existing guarantee, so it ships off.
  • Bounded. The term is added, never assigned, and capped; similarity still selects the candidates.
  • No lost update. State writes go through a named lock; the reproducer test asserts both halves.
  • Unicode-scoped. Domain keys keep letters in any script — Cyrillic and CJK domains stay distinct instead of sharing one posterior.

The engine is vendored (215 lines, MIT, zero imports) rather than imported: its upstream path is not in agentdb's exports map, and a ranking feature that quietly stops ranking looks exactly like one that works.

Per-stage Codex model matrix

With primary: 'codex', the budget table always selects models by stage. The optional RoutingEnv.complexityTier (S, M, L, or XL) affects only planning; the workflow twins read it from args.tier. Omitting the tier selects S/M planning: flagship (workhorse in eco). Explicit model overrides retain their existing precedence.

| Stage | Normal / hybrid Codex axis | Eco Codex axis | |---|---|---| | Router | Terra · medium | Luna · medium | | Requirements | Sol · medium | Terra · medium | | Research (evidence collection) | Terra · medium | Luna · medium | | ADR | Astra · high | Sol · high | | QCSD / ideation | Sol · high | Terra · high | | DDD | Sol · high | Terra · high | | Architecture | Astra · high | Sol · high | | Plan · S/M | Sol · high | Terra · high | | Plan · L/XL | Astra · high | Sol · high | | Code | Sol · high | Terra · high | | QE | Claude Sonnet (independent family) | Claude Sonnet (independent family) | | Fleet | Sol · high | Terra · high |

CODEX_TIERS names roles: premium (gpt-6-astra) for consequential decisions; flagship (gpt-5.6-sol) for direct work; workhorse (gpt-5.6-terra) for evidence; high-volume (gpt-5.6-luna) for mechanics. Eco lowers each selected Codex tier by one level; these are capability assignments, not measured prices. Claude-primary cells and the cross-family QE rule retain their existing behavior.

Experiment envelope (feature-adr-envelope.ts, ADR-001 envelope-before-dispatch)

The feature-adr conveyor writes an OUTCOME per run (grade, some tokens) but never used to write the DECISION behind it — what kind of task this was, how big, at what priority, which model arms the router considered, which one it picked, and who judged it. buildExperimentEnvelope assembles that as one plain-data object, built exactly ONCE per run (right after the Step-0 router, once tier and taskKind are known, and before Step 1 dispatches anything), then threaded byte-for-byte into every place the run reports itself:

  • every autowritten run-cost ledger row (.dz/feature-adr/run-cost-ledger.jsonl, field envelope);
  • every captured training pair (.dz/fa-training/<slug>/<stage>.jsonl, field envelope, alongside the narrower legacy budgetMode — not instead of it);
  • the round state opened via dz round open --envelope <json>, copied into the round's ledger row by closeRound on dz round close.

Shape (ExperimentEnvelope): schema:1, runId, attempt (integer ≥ 1), taskKind (one of feature|bugfix|refactor|tooling|docs|research), tier (S|M|L|XL), priority (speed|balance|quality|unset), treeSha (40-hex or null + treeShaReason), arms ({mode: string[], stages: {stage: string[]}} — what the routing tables OFFERED), chosen ({mode, stages: {stage: spec}} — what was actually resolved), policy ({name, version, propensity}), evaluator ({family, model, source: 'planned'|'actual'}). validateExperimentEnvelope(value) returns {ok:true} or {ok:false, reason} naming the FIRST invalid field.

The writer refuses an automated row without one (FR-5, D2). In run-records.ts, decideRecordWrite for kind:'ledger' refuses (exit 2) an auto:true row that carries no envelope, and refuses ANY row (auto or manual) whose present envelope fails validation. A manual row without auto/envelope is unaffected — the old shape still writes exactly as before (C-3). Read it back with jq '.envelope' .dz/feature-adr/run-cost-ledger.jsonl.

args.priority (FR-4). A learning-stratum LABEL, one level above budget/deliveryGate — an explicit knob always wins over the preset:

| priority | budget preset | deliveryGate | |---|---|---| | speed | eco | false | | balance | normal | false | | quality | normal | true | | unset (default) | whatever args.budget says | whatever args.deliveryGate says |

PRIORITY_PRESETS + resolvePriority(raw) + applyPriorityPreset(priority, explicit) live next to BUDGET_PRESETS in feature-adr-routing.ts; an unknown priority is a startup error naming the valid list, never a silent unset. Setting priority alone (no other routing knob) turns routing on.

roundId — the round's join key, as a field (backlog c60cc857)

Every row closeRound writes carries roundId, the round's own identity (round-<slug>-<n>-<closedAt>). The same string is still the first token of note, because the confirming re-read of the ledger greps for it there — but a key that lives inside prose can only be joined by parsing prose, and with a --note present the field reads <marker> | <note>, so an equality match on note misses it.

MEASURED on the COMMITTED ledger at 488f596c: the round's own identity was in a field 0 times out of 99 round rows — while sitting inside note prose 99 times out of 99.

One command prints every LEDGER COUNT above, in this order: rows; round rows; of those, rows with roundId; rows whose note starts with the round marker; rows with taskId; rows with stateId; rows carrying none of runId/taskId/stateId; rows with runnerId. It answers 436 99 0 99 28 76 22 99.

git show 488f596c:.dz/feature-adr/run-cost-ledger.jsonl | node -e \
  'const L=require("fs").readFileSync(0,"utf8").split("\n").filter(Boolean).map(JSON.parse),R=L.filter(r=>r.stage==="round"),h=(r,f)=>r[f]!==undefined&&r[f]!==null&&r[f]!=="";console.log(L.length,R.length,R.filter(r=>h(r,"roundId")).length,R.filter(r=>/^round-.+-\d+-\d{8}T/.test(r.note||"")).length,R.filter(r=>h(r,"taskId")).length,R.filter(r=>h(r,"stateId")).length,R.filter(r=>!h(r,"runId")&&!h(r,"taskId")&&!h(r,"stateId")).length,R.filter(r=>h(r,"runnerId")).length)'

The command names the commit rather than HEAD on purpose: the ledger is tracked and keeps growing, so a HEAD form would print different numbers every week and the claim would stop being checkable.

What the other id fields are, and why they do not answer this section's question: taskId names the TASK a round serves, stateId its state file, runnerId the host that wrote the row. None of them is the round's ledger identity.

This paragraph needed five corrections, every one from cross-family review (Codex gpt-5.6-sol), and they are kept because the corrections teach more than the result. (1) The first draft counted an uncommitted working copy — 457/119/117 — which cannot be audited from the repository. (2) The second said 97 rows "carried no join key in any field" while measuring only a missing runId. (3) The third quoted 28/76/22 that the shown reproducer did not print. (4) The fourth called those 22 rows "no id field at all" when runnerId is present on every one. (5) The fifth added that runnerId claim without adding its count to the command. Each time the NUMBER was right and the SENTENCE was wider than the predicate that produced it — the same defect this repo keeps finding in code verdicts, wearing prose. Six rounds on one paragraph is itself the finding: a sentence added to JUSTIFY a number is a new claim, and it needs its own predicate or it does not belong — the sixth round caught the phrase "every number this section states", which the ledger command cannot cover because this correction list is not a ledger count. Its evidence is git log on this file, not that command, and reviewing this paragraph stopped here on purpose: the repo's own rule is that three rounds on one mechanism mean you write the boundary instead of taking a fourth.

RoundLedgerRow is the shape this version WRITES, not the shape the file holds. The ledger is append-only, so rows written before a field existed do not have it — roundId included. Read a ledger line as unknown and put it through validateClosedRoundLedgerRow (or your own guard); never type a parsed historical line with this interface and dereference a later field as guaranteed.

QE self-report, arm fields, and their honest limits (instrument-round-b, ADR-001 D1/D2, fix-round-1)

The .claude/workflows/feature-adr.js pipeline script (not a harness-core module — a workflow the Step-8 QE stage runs) stamps three MORE trusted-by-construction fields onto every autowritten run-cost ledger row, on top of the envelope above:

  • aqeInvoked/aqeInvokedSource/aqeEvidence — whether the QE agent actually invoked a live agentic-qe tool this pass, SELF-REPORTED by the agent (aqeInvokedSource:'qe-self-report'), never derived from MODE (a run's intent, not an event). Unreported ⇒ honest null/'not-reported'. fix-round-1 (Codex r1 MEDIUM finding 6): aqeEvidence is trimmed; blank-after-trim keeps aqeInvoked but sets aqeEvidence:null, aqeEvidenceStatus:'blank'; a non-blank value is checked against a SOFT shape (contains mcp__agentic-qe__, or starts with aqe ) and marked aqeEvidenceStatus:'recognized'/'unrecognized' — the value is KEPT either way, never discarded. NAMED LIMIT: a fabricated tool name that happens to match the soft shape is indistinguishable from a truthful one. This mechanism can catch an obviously wrong shape; it cannot catch a convincing lie. That is the owner-selected self-report design (no external observer inside the agent's own tool calls exists), not a gap this round could close.
  • arm/propensity/armSource — copies of envelope.chosen.mode/envelope.policy.propensity, promoted to top level so a reader never parses the nested envelope; armSource:'invalid-envelope' (fix-round-1, Codex r1 HIGH finding 5) when the envelope object resolves to NEITHER field, honestly distinct from the working 'envelope' label. The per-stage extra object every autorow call site passes is FORBIDDEN from carrying arm/propensity/armSource/envelope — stripped before the spread, then envelope: ENVELOPE is restated and the arm fields re-derived from that exact value — so the row's nested envelope and its top-level arm fields can never disagree, even under an extra that tries to smuggle its own.

What it provides

Evidence-gated companion integrations

runInit reads and aggregates adjacent INTEGRATIONS.json manifests once per run. Requested components always produce one of two explicit outcomes: a receipt-backed emission or a named refusal. The current measured admission set is one cell—Claude Code 2.1.235, project MCP—qualified by a non-executing live registration probe. The other 19 cells refuse by stable reason code. A pending/committed .dz/integrations-ownership.json journal prevents an observed user value or forged ledger from becoming overwrite authority; AgentDB setup uses this same writer and adopts only its known historical shape.

--allow-integrations <sha256:…> binds consent to the exact aggregate. --no-integrations is an explicit skills-only short circuit. --no-verify cannot authorize emission. A Claude Pending approval observation is registered but ready: false.

| Module | Exports | Purpose | |---|---|---| | skills | loadSkillFromDir, listSkills, listSkillsDetailed, describeSkillLoadFailure, formatSkillLoadFailures, formatSkillApplyFailures, discoverSkillIds, walkFiles, isSkillJunkFile, SKILL_JUNK_DIRS, SKILL_JUNK_FILES | Read skill directories into CanonicalSkill objects. Two listing functions, deliberately: listSkills THROWS on the first unloadable skill and always will — it is a published export, and silently turning it into a skip-and-collect function would downgrade every unknown third-party consumer from fail-closed to fail-silent without their consent (an incomplete catalogue reported as complete); a pinned regression test asserts it still throws. listSkillsDetailed is the total variant callers ask for BY NAME: it returns {skills, failures} with a per-id try/catch, so one unparseable SKILL.md never hides the ones after it (order-independence is the tested property — the offender first, middle or last yields the same counts). Every failure is NAMED — describeSkillLoadFailure is the single place a pathless parser throw becomes {id, absolute path, verbatim reason, first line}, because the parser is handed only TEXT and can never supply a path. formatSkillLoadFailures renders that list for stderr in one of two modes chosen by the CALLER (absolute paths for dz list/dz sync; relative-to-package for dz install, where a node_modules/** path is not actionable). Symlinks and junk (feature skills-walk-symlinks-and-junk): walkFiles, the asset-discovery loop loadSkillFromDir and getSkillInfo both run on, resolves every symlink with statSync before deciding whether it names a file or a directory — a Dirent from readdirSync answers false to BOTH isDirectory() and isFile() for a symlink entry, so trusting those two checks alone silently drops every symlinked asset (MEASURED: 2 of 4 fixture assets vanished, exit 0, before this fix). A symlink whose target cannot be stat'd is a broken symlink; a directory (reached directly or through a symlink) whose realpath is already on the current ANCESTOR chain ends the walk there instead of recursing — the guard tracks the recursion path, not every directory ever visited, so two non-cyclic aliases of one directory (alias1 -> shared, alias2 -> shared) are both walked under their own logical paths (Codex r2, lead fix) (walk-guards-cycles, its own dedicated mutation entry as of fix-round 1 — the symlink-resolution mutation alone cannot prove the guard, because disabling symlink-following ALSO stops any cycle from ever being reached), which is what stops an a -> .. cycle from hanging. Containment (fix-round 1, lead item AM-8): a symlink is followed only when its RESOLVED target's real path lies within the skill directory's own real path — assets/secret -> /etc/hostname, or a relative -> ../../.. that escapes upward, is refused with reason 'symlink escapes the skill directory' and never bundled, whether the escaping target is a file or a directory; only a .. path COMPONENT counts as an escape — a file legitimately named ..asset is inside the root (Codex r2, lead fix). SKILL.md itself gets the same check BEFORE it is read: a SKILL.md that is a symlink escaping the skill directory makes the whole skill REFUSED with a named error (it is mandatory, so it cannot merely be skipped); an in-tree SKILL.md symlink still loads (Codex r2 CRITICAL, lead fix) (an escaping directory is not recursed into either — nothing beneath it is walked). What counts as junk IS THE PUBLISHED CONTRACT (fix-round 1 HIGH-1 — Codex's finding that this contradicts "never drops a legitimate skill asset" is REFUTED-BY-CONTRACT, not a bug: a skill cannot ship an asset under one of these exact names, on purpose or by accident, and that is the deliberate trade this design makes, not an oversight to be widened into content-sniffing): directories __pycache__, node_modules, .git, __MACOSX, .pytest_cache, .mypy_cache (SKILL_JUNK_DIRS); files named exactly .DS_Store or Thumbs.db, or matching *.pyc, *.pyo, *.swp, *.swo, or .#* (SKILL_JUNK_FILES + isSkillJunkFile) — a trailing ~ (editor backup) is deliberately NOT on the list: it is the one pattern a legitimate asset name can plausibly end with (notes~), and the list is conservative by contract — a false positive would silently drop a real asset (Codex r2, lead decision). This is NOT a whitelist — any other file (including a skill author's own notes.local.txt) is kept as a real asset; filtering someone else's files by name is not this list's job. Every junk entry, broken symlink, escaping symlink, and detected cycle is counted and NAMED, never silently dropped: walkFiles returns {files, skipped} where skipped is {path, reason}[] and reason NAMES the matched pattern ('junk file (*.pyc)', 'junk directory (__pycache__)', not a bare 'junk file' — fix-round 1 HIGH-1(b)), and loadSkillFromDir threads that list onto its CanonicalSkill result as an optional skipped field (present only when something was actually skipped, so every existing consumer that only reads the CanonicalSkill shape is unaffected). An unreadable directory (readdirSync throwing — fix-round 1 MEDIUM-3) is also a named 'unreadable directory (<errno>)' skip, never a throw out of loadSkillFromDir. dz install sums the junk-tagged entries across the installed package's skills and prints one line — skills: skipped N junk entr(y|ies) (…) — only when N > 0; the count is ENTRIES, not files (a skipped junk directory is one entry regardless of how many files sit underneath it, since walkFiles never descends into it to count those), and a directory path in the list is shown with a trailing / (fix-round 1 MEDIUM-4) | | apply | applyEmitResult | Write an adapter EmitResult to disk — additively | | repo-boundary | isRepoBoundary, RepoBoundaryIo | A repository boundary is a .git directory with a real HEAD file or a worktree gitdir: redirect; an empty or unrelated .git entry is not a boundary, so dz run from a directory such as /tmp with a stray empty .git no longer treats it as a project root (and no longer creates a .dz store there). Store-scoped locks stay at <root>/.dz/locks/<name>.lock, a pure function of the root — and since lock-never-seeds-store they REFUSE (StoreAbsentError) instead of creating that .dz. | | targets | TARGETS, TargetName, isTargetName, resolveTargetName, TARGET_ALIASES, TARGET_NAMES_SORTED, formatTargetProblem, formatTargetAliasNote, normalizeTargetToken | --target name → platform adapter, plus the resolution layer in front of it. isTargetName/TARGETS/TARGET_NAMES are UNCHANGED: boundaries.json names isTargetName as the scanned --target validation boundary, and every resolution ends in exactly that guard — the boundary is routed THROUGH, never relocated. resolveTargetName is total and pure, with fixed precedence: exact canonical → normalised canonical (case/padding/separators: Claude_Code, claudecode) → an explicit TARGET_ALIASES row → unique normalised prefix → Levenshtein ≤ 3 strictly better than the runner-up → nothing. Aliases ACCEPT; prefix and Levenshtein only SUGGEST — an alias row is an owner decision recorded in DATA (adding one is one line and zero control flow), while a fuzzy match is a guess, and installing to the wrong target on a guess is worse than one round-trip. An ambiguous prefix (co → codex/copilot) is terminal with NO suggestion, for the same reason. formatTargetProblem renders the two-line refusal, keeping the literal --target must be one of: substring that shipped assertions pin | | agents-policy | POLICY_SOURCES, extractPolicyBlocks, renderPolicySections, detectPolicyDrift, measureAgentsMdBudget | Pure anchored policy extraction, 12-hex source stamps, drift classification and Codex project-doc byte-budget measurement. The stamps prove source/target synchronization only; they do not prove that a runtime read or obeyed the text | | sign | listPackFiles, listSignablePackFiles, verifyManifest, verifySbomAgainstManifest, packAllowlistFromPackageJson, isInsidePackAllowlist, OUTSIDE_FILES_REASON | Shared node_modules/.git exclusions; verify sees MORE than sign (smuggled symlinks still fail). The sweep scans the whole directory and skips nothing; a swept path outside package.json.files (directory entries, exact files, main/bin, * globs; root package.json and readme/copying/license/licence bare or .<ext> not ending in ~/$ always inside per npm-packlist, CHANGELOG* not — pnpm omitted it, measured) is still a failure but reads present in the directory but not signed — outside package.json.files, the packer would not ship it, and VerifyResult.outsideFiles lists such paths. After authenticating the Ed25519 manifest, verification derives the canonical CycloneDX document from those signed entries and requires the no-follow sbom.json read to match it exactly. Current/v3 signing refuses malformed, duplicate-key, or precision-losing root package.json JSON and preserves object order throughout exports, imports, and typesVersions, so condition-order entry-point changes cannot hide behind packer-noise canonicalisation. Readers retain v1/v2 compatibility | | guard | evaluateGuard, resolveRules, scanSecrets, scanSecretsChunked, SECRET_SCAN_OVERLAP_BYTES, DEFAULT_RULES, parsePnpmLockImporters | Declarative HARD/SOFT constraint engine behind dz guard (publish/teach/consolidate pre-flight; fail-closed). The SOFT signature-fresh rule warns before publish when a changed pack no longer verifies against its signed .dz-manifest.json; all manifest, key, and filesystem reads remain in the CLI fact gatherer. The SOFT lockfile-in-sync rule compares each workspace package's @dzhechkov/* dep specs against the specifier pnpm-lock.yaml records for that importer — the ERR_PNPM_OUTDATED_LOCKFILE CI break, caught at publish. Its lockfile reader (parsePnpmLockImporters) is a pure RECOGNISE-OR-REFUSE parser (no YAML dependency): it reads only the lockfileVersion: 9+ importer layout and returns undefined for a legacy v5/v6 file, a truncated one, or any shape that leaves an importer with zero specifiers — because a half-parse reports every real dependency as "not recorded". The rule FAILS OPEN on that undefined (no violation) and is pinned SOFT-only via SOFT_ONLY_RULES, so no config can turn a parser that admits uncertainty into a publish blocker. The HARD licence-hold rule (+ LICENCE_HOLD_PENDING_MARKER) is the machine side of a declared licence precondition (package.json.licenseHold, ADR-001 hermes-claude-adaptation): silent while the pack stays private:true (the npm layer refuses it), it HARD-blocks publish the moment the pack becomes publishable with the hold unsatisfied — LICENSE absent/empty or still carrying the <!-- PENDING: grant placeholder, no Grant-Confirmation: <url> line, empty THIRD_PARTY_NOTICES, or a non-SPDX license field | | pack-inventory | listPublishInventory, gatherPublishSecretFacts, readFileChunks, parseNpmPackListing, looksBinarySample, NUL_RATIO | Publish inventory from a supplied tarball, a content-hash cache, or the live packer (npm pack --dry-run --json), with a named fallback-walk on packer failure. Scans the inventory as streams, retaining every coverage gap by name; files are never skipped for size. The binary skip is a RATIO, not "one NUL" (feature pack-inventory-binary-sniff-ratio): looksBinarySample(first 8 KiB) is binary iff NUL bytes exceed NUL_RATIO = 0.01 of its length; empty ⇒ text. ONLY NULs count — the rule it replaces skipped on a NUL and nothing else, so any wider criterion (other control bytes, UTF-8 validity) would skip files the old rule scanned: a latin-1 source with a secret, a text with \x01 separators — a fail-open regression (fix round 1, F1). UTF-16 (about half NULs) stays binary; a PNG with few NULs is scanned — harmless, the scan is read-only. MEASURED 2026-09-21 (backlog d3841a3b): the old includes(0) sniff skipped a 19 628-byte TypeScript source with ONE NUL at offset 4383 — together with its dist twin — from the no-secrets scan, fail-open. One NUL in 8 KiB of text is now scanned; UTF-16 and NUL-padded binaries are still skipped with the binary label. Mutation-defended: binary-sniff-one-nul-is-text | | slop-lint | slopLint, parseSlopRegistry, validateSlopLintConfig, DEFAULT_SLOP_CONFIG, BUNDLED_SLOP_REGISTRY_URL | Pure deterministic EN/RU lexical-density and structural-style analysis behind advisory dz lint. It excludes protected Markdown, requires at least two distinct registered marker IDs in one paragraph, divides marker hits by max(visibleWords, wordFloor), and reports bullet walls or registered three-adjective stacks independently. Under the default 4/2/25 policy, the distinct-ID floor owns paragraphs through 50 words and density is the dilution cap from 51 words onward. The core performs no file, network, clock, locale, or process I/O; policy/config failures are typed diagnostics rather than empty clean results. | | stem | tokenize, stemToken, stems | Zero-dependency EN/RU word-form normalisation (light suffix stripping applied to BOTH sides of a match) behind registry search and recommend, so «анализы» finds «анализ»; a RU topic dictionary maps Russian queries onto catalogue topics, and an unmapped topic is reported as a miss rather than silently widened. | | course-staleness | classifyCourseStaleness, CourseStalenessState, CourseStalenessInput, CourseStalenessResult | Pure tutorial/package parity classifier. It distinguishes S0 SHIPPED, S3 TUTORIAL_STALE, S4 PACKAGE_BEHIND, malformed/mismatched/unknown registry inputs, and—load-bearing—E2 UNSTAMPED; an absent source stamp can never collapse into shipped. The caller supplies registry facts, so classification performs no file, process, clock, or network I/O. | | backlog + backlog-embed | dedupIdea, classifyDedup, dedupPairBand, dedupEmbedText, lexicalContainment, ensureBacklogEmbedForm, recordAbsorption, alignIdea, spinRoulette, readGoalMapDetailed, parseEffort, ensureBacklogGitignored, harmonizeBacklog, transitionIdeas, checkTransition, resolveIdPrefix, IDEA_TRANSITIONS | The Smart Backlog engine behind dz backlog: status lifecycle via transitionIdeas (ship/drop/reopen against the IDEA_TRANSITIONS table — unique-short-prefix resolution, idempotent ship/drop no-ops, non-idempotent reopen, all-or-nothing fail-closed batches, line-preserving atomic JSONL rewrite that keeps every non-status byte of untouched records); content-addressed idea records in .dz/backlog/ideas.jsonl, semantic dedup over the REUSED agentdb vector namespace (dz-backlog — no second store), weighted-max GoalMap alignment, and a seeded weighted roulette. Dedup is TWO-SIGNAL since the register-inflation fix (MEASURED 2026-08-11 on the real 105-idea store: full-length embeds INVERTED the signal on long texts — genuine paraphrases 0.35–0.61 vs topically disjoint long-RU pairs up to 0.9195): dedupEmbedText embeds a bounded 400-char excerpt (backlog-embed.ts, one form shared by query/mirror/reindex so vectors can never split spaces; ensureBacklogEmbedForm re-mirrors v1 stores once, batched), and dedupPairBand requires a cosine-threshold DUPLICATE to also share subject vocabulary (lexicalContainment ≥ 0.3, else demoted to RELATED with the pair reported — the 0.941 register-only absorption) while promoting a same-idea re-capture at a different length (containment ≥ 0.95, cosine ≥ 0.75) to a subset duplicate; recordAbsorption keeps every absorbed text in absorbed.jsonl so a wrong verdict is reversible (mutation-defended: backlog-dedup-demotion-corroboration, hand-verified 6 red). Every band/weight decision is a PURE function. classifyDedup also reports the top-1 match id alongside the cosine (the calibration surface for the 0.92 duplicate band — observational, it never moves the verdict); readGoalMapDetailed returns the entries the defensive reader DROPPED with a reason AND the fields it REPAIRED with their raw values (a weight clamped before validation made the validator's out-of-range branch dead code), so goals --validate can never report a vacuous "valid (0 goals)" nor hide a weight: 7; parseEffort returns a printable note for every clamp; ensureBacklogGitignored gitignores the store on first write (raw ideas are private prompt-class content) — atomically, preserving the file's dominant EOL, recognising every plain spelling of an existing rule (/.dz/, .dz/**, …) via backlogIgnoreStatus, and obeying a ! negation as an explicit user opt-out instead of overriding it | | no-stubs | scanStubs, checkNoStubs, scannableStubPath, STUB_MARKERS, STUB_PHRASES, STUB_SCAN_EXTENSIONS | Pure unfinished-stub scanner behind the SOFT no-stubs publish rule (backlog 0b403a0106103901, Karpathy-Michaels rule XI): bare markers (TODO/FIXME/HACK/XXX/PLACEHOLDER) case-SENSITIVE with hard word boundaries (hackathon/todos/a marker inside a hash never fire; MEASURED: relaxing case doubles this repo's hits and adds only prose) + the implement later phrase case-insensitive. SCOPE = the CHANGE-SET (the working-tree git status --porcelain -uall diff — -uall so a brand-new untracked DIRECTORY is scanned file-by-file instead of collapsing to one invisible ?? newdir/ line; .gitignore semantics unchanged), never the whole tree — MEASURED: a tree-wide scan is 32+25 hits of mostly ancient legitimate markers, i.e. noise that gets a gate switched off. Markdown gets PROSE scoping (fenced blocks + backticked spans are QUOTES, not stubs). Waiver-with-REASON only, per line (no-stubs: <reason>) or per path (.dz/guard.json stubWaivers, the feature-adr-setup --guards shape); a reasonless waiver is REFUSED as its own finding and exempts nothing. Self-exemption is STRUCTURAL: every marker in the module and its tests is assembled from string fragments, so the gate's own source scans clean — a tested property, not a path skip. Fail-open on missing evidence (no change fact / ungathered contents ⇒ nothing reported) but never fail-SILENT: skipped scannable files (deleted/oversize/unreadable/beyond the file cap) surface as ONE aggregate notes entry in the GuardResult + audit record — information that can never move the verdict. KNOWN LIMITS are documented at the top of no-stubs.ts instead of implied away (whole-line inline waiver token = layer-4 auditability defence; reason QUALITY not judged; boolean fence model, not CommonMark; git-quoted paths undecoded; TS-monorepo extension allowlist; worktree-not-index reads; exact-string config-waiver paths). Mutation-defended (no-stubs-bare-marker-fires observed 10 red, no-stubs-skipped-note-emitted observed 2 red) | | feature-adr-setup (P3) | renderGuardsConfig, renderGuardsRunner | Scaffolds deterministic guard tests into a TARGET project: guards.config.json + a zero-dependency check.mjs runner (loc-cap, secret-scan, frozen-file sha256 pins, waivers-with-reasons) — dz feature-adr-setup --guards | | usage | computeUsage, TOKEN_WEIGHTS, readUsageLimits, deriveUsageCalibration | Read-only Claude usage ESTIMATE behind dz usage. Tokens are COST-WEIGHTED input-equivalents (input 1x, cache-write 1.25x / 1h 2x, cache-read 0.1x, output 5x) — a flat sum is 89-99.7% cache-read (MEASURED) and tracks conversation length, not work. Scans subagent transcripts too (<session>/subagents/*.jsonl), follows no symlinks, reads only regular files (symlinked FILES and DIRECTORY components alike are skipped), and caps the walk BY RECENCY so a huge history cannot discard current usage. pct stays null while limits are unconfigured — an unconfigured estimate is never dressed up as a number | | cost-ledger | deriveCostLedger, buildCostLedger, verifyCostLedgerReport, stageCostAggregates, renderCostLedger, writeCostLedgerJsonl, COST_LEDGER_SCOPE | Per-stage cost ledger behind dz usage --by-stage. A feature-adr run reports ONE number; this joins the workflow's own stageLabel() strings to the per-agent transcripts the harness already writes, so a run becomes an itemized receipt. POST-HOC DERIVER, not a writer — no workflow edit, and a KILLED run is still derivable. The invariant: accounted + unaccounted === runTotal and accounted + doubleAttributed === Σ stages, RAW integer equality (rounding happens exactly once, per sample, at extraction), re-derived from the emitted report by verifyCostLedgerReport — the writer clamps, the verifier enforces. A mismatch is a NAMED defect (Unaccounted, DoubleAttributed, ForeignSample, MissingStageTranscript, MalformedRecord), never a rounding remainder, so epsilon defaults to 0. The run total comes from the run's transcript DIRECTORY LISTING, NOT the record's own totalTokens — that field is exactly Σ workflowProgress[].tokens in 29 of 29 recorded runs (MEASURED), so an invariant against it can never fail. Both sides share ONE estimator with dz usage (weightedTokensOf). stageCostAggregates is a pure feed-forward reader for auto-cost routing that EXCLUDES non-reconciling runs (now also INCOMPLETE_INVENTORY runs — the !== 'BALANCED' gate already excludes it, no second branch to forget); wiring it into routing is deliberately out of scope. HONEST SCOPE, printed by every surface: local transcript ESTIMATES, not billed amounts — it catches ATTRIBUTION errors, NOT pricing errors; hasKnownPricing marks rows priced by the sonnet-class fallback. INSUFFICIENT_DATA is a distinct verdict, never collapsed into BALANCED. measurement-integrity (ADR-001 D1/D2): every row also carries stageCanonical (the verbatim stage classified against feature-adr-stage-canon.ts's one ordered table — see that module below — 'unknown' when no rule matches, never silently folded into infra) plus attempt/attempts (a label repeated N times in one run is N separate rows, each tagged attempt: i of attempts: N, instead of one row silently summing them). The report gains byCanonicalStage (every canonical stage + unknown + unattributed, {tokens, agents, attempts}) and reconciliation.orphanTranscripts ({count, tokens, ids, method} — transcripts present in the run directory with NO workflowProgress[] entry; method: 'per-transcript' is an exact sum over each orphan's own samples, 'count-fallback' is the best estimate when only ids are known). A run whose ONLY problem is a named orphan (nothing else defective) reports verdict INCOMPLETE_INVENTORY — outranks BALANCED, outranked by DEFECT — never the old Unaccounted/DEFECT pair that used to swallow the orphan into the generic bucket | | feature-adr-stage-canon | CANONICAL_STAGES, STAGE_LABEL_RULES, canonicalStage | measurement-integrity ADR-001 D1: the canonical stage taxonomy — 11 pipeline stages (router, requirements, research, adr, ideation, ddd, architecture, plan, code, qe, fleet) + infra for the bookkeeping/plumbing labels around them. canonicalStage(label) classifies ONE verbatim stageLabel() string against ONE ordered prefix table (first match wins; a label · model suffix is matched on the part before ·) and returns {stage, label, known} — the input label is NEVER rewritten, only classified next to it. An unrecognised label is {stage:'unknown', known:false}, never silently infra. The completeness fixture (test/feature-adr-stage-canon.test.ts) is 47 labels copied verbatim from a live recorded run (wf_5a7755c7-f92) — Step 0's assessment counted 48 on the same record; a live reproducer counted 47, and the one-label gap does not change which prefixes are needed. Pure — no filesystem, no clock; the core-boundary ratchet pins it at zero node:fs imports | | codex-rollouts | parseCodexRollout, matchCodexRollouts | measurement-integrity ADR-001 D3: a pure reader for Codex CLI rollout logs (~/.codex/sessions/YYYY/MM/DD/rollout-<ts>-<uuid>.jsonl) — 130 of 156 recorded Codex ledger rows carry tokens: null even though the spend is sitting on disk, because the pipeline dispatches codex exec without an explicit session id. parseCodexRollout(text, fileName?) extracts {id, cwd, model, startedAt, endedAt, totals} from one file's TEXT (never opens a file itself — the CLI does that); it accepts BOTH the schema Step 0 documented (type:"token_count", payload.info.total_token_usage) AND the schema actually observed live on this machine 2026-09-16, cli_version 0.154.0 (type:"token_usage_record", payload.usage; model on turn_context, not session_meta) — a reader that understood only a shape nothing on disk still emits would fail at the exact thing it exists to fix. matchCodexRollouts(rollouts, {from, to, cwd?, model?}) joins a stage's time window to the rollout that produced its spend by INTERVAL OVERLAP, never "nearest in time" (two reviews back to back would misattribute) — 0 matches is {status:'none'}, 1 is {status:'one', rollout}, >1 is {status:'ambiguous', candidates}, never a first-pick. Pure — the core-boundary ratchet pins it at zero node:fs imports | | compounding | mulberry32, bootstrapDelta, decidePromotion, assembleCompoundingReport, assembleLessonToRuleFunnel | Pure learning-loop payoff engine behind dz compounding: seeded deterministic bootstrap (conservative nearest-rank lower-95), promotion that refuses non-finite/malformed input and anything under 5 samples per arm, and dz-native measurements (pool write-only ratio, guard trajectory by RATE, replay readiness over unique untruncated prompt events). Its lesson-to-rule funnel reports UTC calendar-month eligible → attempted → accepted → executions counts from prospective promotion-run and anchored guard-audit evidence. Zero alone is not a finding: only a non-empty predecessor followed by an empty named successor in three consecutive measured months produces one; unavailable evidence remains NOT MEASURED with its reason. Compaction keeps the newest query-bearing rows verbatim and aggregates ONLY the rest (read totals are invariant across compactions). Also reports EVENT-CHAIN health of the evidence logs it computed from (evidenceLogs in, instrumentation.chains out) — verified / defect kinds / uncovered pre-chain prefix, with no logs handed in producing no line at all rather than a vacuous "clean" | | event-chain | fnv1a32, nextChainFields, appendChainedLines, chainRewrite, guardedRewrite, verifyEventChain, EVENT_CHAIN_SCOPE | Pure hash-chain over the two learning-evidence logs (.dz/recall-usage.jsonl, .dz/guard-audit.jsonl): each appended record carries seq + prevHash (FNV-1a over the previous line AS WRITTEN, so key order cannot make writer and verifier disagree), derived from the LAST LINE ONLY so a per-prompt hook stays O(1). verifyEventChain names eight classes — BrokenLink, DuplicateSeq, NonMonotonicSeq, TornTail, DoubleCounted, LedgerImbalance, MalformedLedger, ClaimInterrupted. The last four exist because a rewriter must not be able to certify itself: the compaction ledger's arithmetic (Σ weight + dropped === source, dropped ∈ [0, source]) is enforced with NO clamps in the verifier (the clamp belongs to the writer), a damaged ledger line is a defect rather than a silently-disabled check, and a claim that never reached its throughSeq — because the segment restarted or the file ended — is reported instead of escaping through the discontinuity. guardedRewrite is the concurrency guard for any whole-file rewrite: exclusive lock, plus a re-read of the live file after computing the new text and BEFORE the rename, so a concurrent append aborts the attempt and is folded into a bounded retry rather than overwritten (it narrows the read→rename window; it cannot close it, and says so). Records written before chaining existed stay LEGAL and are counted as an uncovered preChainPrefix; an unreadable tail never blocks a write (fresh MARKED segment — an unreadable t