letspecify
v0.23.0
Published
Engineering governance for spec-driven projects: scaffold a governance repo, verify the model and run the quality gates. The model decides, the tool materialises.
Maintainers
Readme
letspecify
Engineering governance for spec-driven projects: Agentic Model Driven + Spec Driven Development, with human control at the risky points and end-to-end traceability.
Code is not regenerated from the spec. The spec lives next to the code and the traceability binds them — spec-anchored, not spec-as-source. That is a decision, not a limitation: what this project verifies is that what is declared and what is written still say the same thing.
Design principle: the model decides, the tool materialises. The topology of the product repos is an architecture decision (PRIN-001: specification before code) derived from the foundational document during the governance bootstrap — it is not settled by a CLI flag.
Commands
| Command | When |
|---|---|
| init | A new project: governance precedes the code. |
| adopt | The code already exists: governance arrives afterwards and adopts it where it is. |
| materialize | Create what the model declares and does not yet exist (init -r and adopt call it too). |
| model | Project the governance to JSON, so another tool can consume it without parsing YAML. |
| check | Verify the model: that everything declared parses, that every identifier cited exists, which number is taken and where, and that the traceability closes from both ends. |
| test | Test the corpus where check tests the model: the documents themselves — front matter, sections, the status of what they cite, the graph between them, the vocabulary. Then twelve properties over the corpus as a whole, which no single document can fail. Thirty-four checks, no language model, and a baseline so it can be adopted on a corpus written before it. |
| critique | The third layer, and the engine calls no model. It publishes what a critic needs and validates the verdict that comes back — six defect types, each one what a deterministic layer measured itself unable to decide. Two commands with a harness in between. |
| gates | Run the quality gates the model declares, and compose their aggregates. |
| spec-levels | Publish the three specification depth levels: what sections each one requires, and the criterion handed to whoever writes the spec. Needs no project. |
| spec-level | Which of the three governs a branch, and where it comes from. With --set, declare it. |
| index | The reverse index: from a file to the fragments of specification that govern it. Generated, versioned, byte-stable. |
| tasks | What is still to be done, and what can be done at once. Every task carries a kind — development or governance — and the pending set comes out as waves: a wave is what MAY run in parallel, and it names the tasks that own the same file and therefore cannot. Answers; never schedules. |
| task-status | Records the state of one task, in the register that declares it. The writing sibling of tasks, and a command of its own so that query's promise to run no version control stays unconditional. It writes one field, commits nothing, and judges neither the words you use nor the transition you make. |
| skills | The instruction document a role is given, and the register of what the model declares. With --for <role> it publishes the document on stdout — a provenance line per skill in force, then their bodies — and ITS BYTES ARE THE PROMISE: one model yields one document, because nobody downstream opens the file. With --json it publishes the manifest instead, never the bytes. It only reads: no file written, no git, no network. |
| launch-plan | The plan for launching a piece of work: which role-threads, which branch each repository is owed, which roots each role writes. Declares; runs no git. |
| launch-instructions | The second phase: writes each role's start-up file, once the threads it must name exist. |
| scaffold | The skeleton of a test, from an acceptance criterion. It writes to stdout unless told otherwise, and the skeleton FAILS on purpose. |
| upgrade | The project catches up with the engine. Every release since 0.13.0 shipped a migration written in prose and none shipped a way to run it; this runs them, oldest first, and records which engine the project is on. Invasive on purpose, and safe by being reviewable: it refuses a dirty tree, names every file before it asks, and copies what it cannot merge before regenerating it. |
| drift | Which parts of a governance nobody exercises: cold specs over hot code, policies nobody cites, gates with no checker. Counts; never judges. |
letspecify init -n <name> -f <foundational.md>
Generates, in the current folder (or in --dir):
<target>/
├── <name>-engineering/ # the governance repo
│ ├── FOUNDATIONAL.md # your foundational document, imported with its provenance
│ ├── BOOTSTRAP.md # a guide for an agent to complete the model + the topology
│ ├── project-model.yaml # root manifest (principles, conventions)
│ ├── model/ # 12 modules: goals, domains, agents, policies…
│ ├── specs/ # global (WHAT) + e2e/ops/obs; materialize creates the domain ones
│ ├── features/ issues/ logs-analysis/ docs/ traceability/
│ └── .claude/settings.json
└── src/
└── README.md # the product repos are declared by the modelWithout -r, domains.yaml and repositories.yaml are born minimal (only
governance + e2e/operations/observability) with commented templates for
registering more, and agents.yaml ships the generic fleet with no domain agents
(they are recruited during bootstrap). The expected flow:
- Open
<name>-engineering/with Claude Code and runBOOTSTRAP.md. - The agent proposes the domain/repo topology from
FOUNDATIONAL.md(as an ADR) → the human approves it → it is recorded in the model. letspecify materializecreates what has been declared.
With -r backend,frontend[,infra] (a shortcut for a topology you already know)
the model is born pre-filled with those domains/repos/agents and is materialised
on the spot.
letspecify adopt [-s <repo1,repo2>] [-n <name>] [-f <foundational.md>]
For code that already exists. It inspects the repos given with -s (or the
ones it discovers in the current folder), creates <name>-engineering/ next to
them and registers them in the model where they are:
<workspace>/
├── <name>-engineering/ # the governance repo, with adoption/ and ADOPTION.md
├── my-api/ # local_path ../my-api (untouched)
├── my-web/ # local_path ../my-web (untouched)
└── my-infra/ # local_path ../my-infra (untouched)It does not write a single byte inside the adopted repos: it only reads them.
No git init, no commits, no new files. Everything generated lives in the
governance repo.
Without -n, the name is derived from the folder (the repo itself if there is
only one, the workspace if there are several). Without -s, the repos in the
current folder are discovered (or the current folder itself, if it is a repo).
-d only decides where governance is born; it fails if that would leave it
inside an adopted repo.
What it prepares, and what it deliberately does not:
| It generates | Content |
|---|---|
| adoption/inventory.yaml | Facts observed per repo: stack, detected build/test commands, languages, size, entry points, markers (docker/k8s/CI/tests), git state. It is the agent's context manifest: it saves exploring blind. |
| model/repositories.yaml | The real repos with their local_path, default_branch, remote and build_cmd/test_cmd already filled in. |
| model/domains.yaml | One provisional domain per repo (provisional: true), with the interactions matrix derived from the observed roles (frontend→backend→infra). |
| model/agents.yaml | One domain agent per repo, autonomy: L1, bounded to its repo. |
| model/policies.yaml | An ADOPTION block (POL-ADOPT-01..05): specify as you touch, no opportunistic refactors, tell observed from inferred, gates calibrated against reality. |
| ADOPTION.md | The adoption guide for the agent; it replaces BOOTSTRAP.md. |
| adoption/GAP-ANALYSIS.md | Where a spec is missing, what diverges from the foundational document, what debt there is. |
| FOUNDATIONAL.md | With -f, the imported document. Without -f, a template marked PENDING RECONSTRUCTION from the code. |
Each repo's role (backend/frontend/infra/mobile/library) is a
hypothesis with its evidence, not a decision: the adoption bootstrap confirms
it with the human. The same goes for the domains — a repo is not necessarily a
domain, and ADOPTION.md explains when to merge or split them.
The rule that makes adopting inherited code viable is specify as you touch
(POL-ADOPT-01): an area's spec is written the first time a feature modifies it,
not all at once on day 1. --dry-run shows the full plan without writing
anything.
letspecify adopt --refresh [-d <repo-engineering>]
Photographs the already-registered repos again and rewrites only
adoption/inventory.yaml. It does not touch the model (those decisions are
already yours) or the repos.
letspecify materialize [--dry-run] [--no-git] [-d <repo-engineering>]
Runs from the root of the governance repo (or with -d). It reads
model/repositories.yaml and model/domains.yaml and creates what the model
declares and does not yet exist:
- Product repos: a folder at their
local_pathwith a README (the inherited governance rules) +.gitignore+git init+ an initial commit. - Domain spec folders:
specs/<dir>/with theirSPEC-<PREFIX>-000-template.md(the prefix comes from the domain'sspec_prefixfield, or from the suffix of its id).
It is idempotent and non-destructive: whatever exists is skipped (an existing
repo is registered in the model with its local_path, not recreated).
--dry-run shows the plan without touching anything.
From a worktree of the governance repo
local_path is relative to the governance repository, not to the directory
-d receives. When that directory is a worktree of governance — the normal
case once every unit of work lives in its own tuple of worktrees — model,
check, gates, materialize and adopt --refresh all resolve against the
main clone, obtained with git rev-parse --git-common-dir. Without it every
local_path starting with .. came out as missing, and materialize would have
created the repos inside .worktrees/ (BUG-2026-013).
Over a directory that is not a git repo — a legitimate case: model works over a
plain folder — resolution falls back to the directory received, as always. And the
governance files themselves (specs, features) are read and written in the tree
the command runs from: that is the version being edited.
letspecify check [-d <repo-engineering>] [--json] [--strict] [--branches]
A deterministic lint of the governance repo. It points; it does not fix, and
there is no --fix — a lint that fixes governance DECIDES, and that is not a
tool's job.
| Rule | What it looks at |
|---|---|
| R1 | Everything the model declares parses. |
| R2 | Every identifier cited — RN, CA, ADR, RISK, POL, CONTRACT — exists. |
| R3 | Duplicate identifiers, and with --branches, which number is taken in the unmerged branches. |
| R4 | Every declared criterion has a row in the matrix, and every feature with criteria appears in the traceability. |
| R5 | The tests the matrix says verify each criterion — do they still exist? |
| R6 | The graph of the work can be true. Tasks: a depends_on naming a task no tasks.yaml declares, a cycle, a kind outside development and governance, a corrects naming a task nobody declares or the task itself, and one task identifier declared twice — all errors. Features: the relations a feature declares in derives_from — see below. |
What R6 says about the relations between features. A feature declares, in its own
manifest, the features it derives from: a list under derives_from: at the root, each
entry naming one feature, one kind and a why. <F> below is the declaring feature,
<T> the feature it names; the file is always <F>'s feature.yaml.
| Severity | Message |
|---|---|
| error | <F> derives from "<T>", and no feature declares it |
| error | <F> declares a derivation that names no feature |
| error | <F> declares that it derives from itself, and a derivation joins two features |
| error | <F> declares a derivation from "<T>" with no kind; a kind is one of continues, supersedes_part, depends_on |
| error | <F> declares a derivation from "<T>" of kind "<K>", which is none of continues, supersedes_part, depends_on |
| warning | <F> declares the derivation <K> "<T>" <n> times, in <fields>: one relation is written once |
| warning | <F> declares derives_from under traceability, and it belongs at the root of the manifest; it was read all the same |
Nothing is reported about a relation's why, about a relation naming a feature whose
manifest could not be read (that manifest is already an R1 finding), or about one
naming an identifier two features declare (already an R3 finding). Nothing is reported
about depends_on_features on its own either: it is read as depends_on and never
judged, and it only counts towards a duplicate when the same relation is also written in
derives_from.
And model --json publishes them. Every feature carries derivesFrom, always
present and an empty list when the feature declares no relation. Each entry is
{ feature, kind, why, declaredIn }: the kind is one of continues, supersedes_part
or depends_on, and declaredIn names the field it was read from — derives_from,
traceability.derives_from, depends_on_features or
traceability.depends_on_features. A derives_from entry is published exactly as
written, defects included, so check can report them. An entry of depends_on_features
becomes a relation of kind depends_on only when it opens with the identifier of a
feature the model declares; anything else — a word, an identifier no manifest under
features/ declares — is not a relation and is published as nothing. The identifier is
looked up, never matched by its prefix: if you keep an idea's or a bug's manifest under
features/, relations to it are read like any other. No relation is ever read out of a
manifest's text.
The shape's version does not move: a field was added.
Every feature also carries derivedBy: the same relations, seen from the other end.
It is always present and an empty list when nothing derives from the feature. Each entry
is { feature, kind, why, declaredIn }, where feature names the feature that
derives — the later one — and kind, why and declaredIn are copied from its
declaration. So declaredIn names a field of the deriving feature's manifest, the one
the relation was read from, and never a field of the feature carrying the entry: that
manifest is not read for it and is never edited. derivedBy is computed from the
published derivesFrom on every run and written nowhere, so removing a declaration from
the later manifest removes it from the earlier one's view. Nothing is filtered: a
relation check reports is still shown, as declared, and a key an earlier manifest
writes about its own successors is not read.
Two severities, and only one decides the exit code: error breaks the build,
warning does not — unless you pass --strict. The asymmetry is deliberate: a
noisy lint gets turned off, and a lint that is off protects nothing.
R3 --branches exists because reserving an identifier by looking only at your own
working copy is how a collision happens. It walks the references of every branch,
including the ones not merged yet.
letspecify test [-d <repo-engineering>] [--json] [--strict] [--rule <code,...>]
check looks at the model. This looks at the documents: their front
matter, their sections, the status of what they cite, the graph they draw between
themselves, and the vocabulary the glossary fixes. Then at the corpus as a
whole. Twenty-two rules and twelve properties, no language model anywhere.
| Family | What it decides |
|---|---|
| STR-01…06 | The document is shaped like the kind of document it says it is. A template owes the field names of the kind it is a template for — that is the row that catches a template teaching the wrong shape to everything written from it. |
| REF-01…06 | What a document points at exists and still stands: a parent that resolves, a decision nobody superseded, a piece of work, a repository, an owning agent. |
| GRA-01…03 | No cycle in the declared graph, nothing in force depending on something superseded, no specification outside every declared domain path. |
| IDX-01…04 | Against the committed reverse index, never a fresh one: stale scopes, annotations naming nothing, criteria with no code, scopes that govern nothing. |
| LEX-01…03 | Forbidden synonyms, non-canonical spellings, and glossary entries nobody uses. |
What it does not decide is as important as what it does. Identifier
uniqueness, dangling citations and whether the traceability closes are check's
R1…R5 and stay there. Two readers of one question drift apart.
A rule that could not look is not a clean rule. With no reverse index the four
rules that read it come back not evaluated, naming the artefact — never as
silence. Same for a glossary that declares no term.
Every rule has a stable code, and a project switches one off — or raises a warning
to an error — in model/spec-rules.yaml, never on the command line. Every run
prints what was switched off and the reason: a rule that is off and invisible
is off for good.
letspecify test --write [-d <repo-engineering>] [--dry-run]The baseline is what makes this adoptable on a corpus written before it.
--write freezes what is failing today, keyed by rule and by an address that does
not move when a document is reordered; from then on only new findings change
the exit code, and every run publishes how many are frozen, how many no longer
occur and how many are new. Only a person writes it — a reporting run never
touches it — and --rule cannot be combined with --write, because a partial run
cannot prune a whole baseline.
And twelve properties, which no single document can fail
A rule reads one document. A property holds over the corpus, and the twelve below are the questions that only make sense when you have all of it — the property-based testing idea, pointed at a specification instead of at a function.
| Property | What it holds |
|---|---|
| SYM-01 | No group of criteria is without a guard among them — the unwanted-behaviour case that groups of happy paths quietly omit. |
| SYM-02 | A criterion that holds while a state holds has one describing how that state is left. |
| CON-01 | Two fragments governing one file do not demand opposite things. |
| CON-02 | Two near-duplicate criteria do not differ in a negation or a number — the copy-paste that changed one word. |
| STA-01 | Every declared state is reachable from the initial one. |
| STA-02 | Every state with no way out is declared terminal. |
| STA-03 | Every declared state has a way in, except the initial one. |
| STA-04 | Every workflow phase names a declared state. |
| ENT-01 | No two criteria overlap beyond the declared threshold. |
| VER-01 | A criterion carries exactly one obligation. |
| VER-02 | A criterion is written in one of the five EARS forms. |
| VER-03 | A criterion uses no term the model declares imprecise. |
Alongside them a metrics block, with no threshold applied to any of it — duplication, asymmetry, term divergence, their composite as a single entropy number, and verifiability. Each is printed with its denominator, so redundancy is something you watch move over releases rather than something you notice one day:
METRICS — measurements, and no threshold is applied to any of them:
duplication 0.0127 (2/158)
asymmetry 0.3481 (55/158)
term_divergence 0.0000 (0/5)
entropy 0.1203 over duplication + asymmetry + term_divergence
verifiability 0.7405 (117/158)
158 criteria · 12403 pairs comparedA component that cannot be computed is named and dropped from the composite rather than counted as zero — an entropy over two of three components says so.
They warn and do not block — every one of them, at birth. A property earns the
right to fail a build by being right for long enough on a real corpus, which is
recorded per project in model/spec-rules.yaml and nowhere else; there is no
command-line flag for it. The thresholds a property compares against live in that
same file, with the measurement that chose them written beside each one.
A property that reads zero on its own corpus is still worth shipping, and this
is the argument for it: SYM-02 had one state-driven criterion to look at here and
was called close to vacuous. The first foreign corpus it saw, it found a real gap.
letspecify critique --request --spec <SPEC-...> | --verdict <path> | --summary
The third layer. test decides what a document and a corpus can be checked for
without understanding a sentence; this is what is left over — and only what is
left over. Six types, and each names the measurement or the recorded refusal it
answers:
| type | what it is | why a rule cannot decide it |
|---|---|---|
| CRT-AMB | lexical ambiguity | the lexical rules only see terms the glossary declares |
| CRT-UNV | unverifiable criterion | VER-01…03 test the FORM of a criterion and stop there |
| CRT-IMP | implicit requirement | an undeclared absence is invisible to everything deterministic |
| CRT-ERR | missing error behaviour | pairing a criterion to its guard was measured and refused |
| CRT-SCP | undeclared scope | the index checks declared scopes, never implied ones |
| CRT-CTR | contradiction with a cited fragment | CON-01/02 see a negation or a cardinal and no more |
THE ENGINE CALLS NO MODEL. Not here, not behind a flag. --request writes what
a critic needs — the fragments, the fragments the document CITES, the closed
catalogue, the schema of the answer, and how many independent passes the depth
level asks for. Something else runs the critic that many times. --verdict reads
the answer back.
It is two commands with a step in between, and that will surprise you the first
time. What it buys: letspecify gains no network client, no credential and no
retry policy, and its exit code still answers about a file rather than a service.
The verdict is validated before a single field of it is read, against
CONTRACT-VERDICT-001 and never by trusting a runtime’s structured-output mode. A
runtime that promises schema-shaped output is a runtime asking to be believed, and
a verdict that is believed is not a verdict. Eight steps in a fixed order, and the
report names the one that refused a document — the difference between a task for
whoever wrote the harness, one for whoever configured the model, and one for the
critic.
A finding outside the six rejects the whole verdict. Not downgraded, not logged advisory, not passed through: there is nowhere to put it. So does a finding with no remedy, and so does provenance that omits the model, the family, the provider, the quantisation, the catalogue version or the passes.
Only what a strict majority of the passes contains is emitted. Identity is the type and the subject, never the prose — two passes word one gap differently. What did not recur is a defect of the CRITIC and is recorded as such. Every remedy survives, unmerged: three readings of one gap beat a smoothed average of them.
A specification with no accepted verdict is not critiqued, which is never the same as nothing found.
--summary reads the dossier and publishes, per defect type, how many distinct
fragments it has been raised against. That is the evidence a type is promoted on —
out of the model’s reach and into a cheaper layer — printed whether or not anybody
is looking.
letspecify gates [-d <repo-engineering>] [--json] [--fast] [--flow <WF-...>] [--write]
Runs the runners model/quality-gates.yaml declares and composes the aggregate
gates, naming their causes.
The model declares the runner; the tool runs it. Adding a gate means
declaring it, not touching the engine. Three shapes: engine (a capability of
the CLI itself), exec (a command, and its exit code), repo-suite (the
test_cmd or lint_cmd repositories.yaml already declares).
Six verdicts, and three of them mean "not checked":
| | |
|---|---|
| green / red | it was checked, and this is the result |
| no-runner | nobody wrote a checker — neither green nor red |
| unreachable | a checker exists and could not run HERE — a precondition this machine does not meet |
| waived | a human granted an exception, with a reason and an expiry |
| not-applicable | the gate has nothing to do with this work (applies_to) |
The two "not checked" answers are kept apart because they ask for different
things: no-runner asks for code, unreachable asks for another machine. A red
asks for a fix. Collapsing any two of them loses the one that says what to do
next.
A checker that could not look is not a green checker, and an aggregate with a
no-runner entry cannot come out green. It is what stops "five of six checked and
nobody knows about the sixth" from reading as approved.
letspecify upgrade [-d <repo-engineering>] [--json] [--write] [--yes]
The project catches up with the engine. Every release since 0.13.0 shipped a
migration written in prose — 0.13.0 opens with "a new project is now born clean, and
an existing one has one migration to do", 0.14.0 gives its own the heading "Your
migration, in one place" — and none shipped a way to run it. This runs them.
It reads metadata.engine in project-model.yaml, the engine that generated or last
upgraded this project, and applies the migrations above it oldest first, moving the
record last. An absent record is not corruption: it means older than every
migration, which is the state of every project made before the field existed.
letspecify upgrade # from which version to which, and every change. Writes nothing.
letspecify upgrade --write # announces the consequences, confirms, applies all of them
letspecify upgrade --write --yes # unattended: says so rather than being asked (CI)--write lists every file and every key, then asks — a question put before the
consequences are on screen is one nobody can answer — and it wants the whole word yes,
not y. With no terminal to ask it refuses (exit 2) and names --yes: reading EOF
and calling it no strands every pipeline on a silent no-op, and calling it yes lets a
redirected stdin rewrite a model nobody was watching. --yes is how being unattended
stops being an accident and becomes something somebody typed.
It is invasive on purpose. It writes into files a person edits, because a command that stopped at the first hand-written line would leave the project where it found it — six of thirty-four rules blind, two metrics absent, exit code 0. What makes that safe is not restraint but reviewability:
- it refuses to run on a modified working tree, so everything it writes lands in one
git diffyou read before you commit; - it names every file and every change before it asks;
- what it cannot merge —
letspecify.mdis prose, and prose has no correct merge — it copies to.letspecify/backup/<from>-to-<to>/before regenerating it. A copy that cannot be taken abandons the run.
A migration owns the shape (keys added, renamed, moved); the project owns the
values it already declared. And a migration that carries vocabulary reads
metadata.language: it writes the block in force for the language the engine
distributes, and inert for any other. An English modal measured over a corpus that
is not in English would turn an honest not evaluated into a green rule measuring
nothing.
check, test and gates say so in their header when a project is behind, naming both
versions. It moves no exit code: being behind is a fact about versions, not a finding
about the corpus.
Not called repair — index --repair already holds that word for a narrow,
unambiguous rewrite, and drift holds another.
The skills a project declares — model/skills.yaml
A skill is an instruction document addressed to one or more roles, and it is not a
briefing. The briefing in model/agents.yaml is prose you edit in place, with no
version, no fingerprint and no way of travelling anywhere but into a role's own file. A
skill is versioned: its versions are written once and never edited, exactly one of
them is in force at a time, and the pointer moves by an edit somebody reviews. That
is why it lives in a module of its own — check is tolerant one file at a time, and a
register grown inside a block of prose would take the launch sets down with it.
init creates the module and registers it in project-model.yaml; upgrade gives both
to a project made before it existed, writing only where the file is absent. The engine
ships no skill and seeds none. The register arrives empty, and that is a declaration
rather than a gap: this catalogue is the project's and open, which is the opposite of the
specification depth levels, whose catalogue is the engine's and closed.
Two bounds are declared in the module and the engine carries neither: max_lines,
how long a skill may run to, and forbidden, the tools, assistants and runtimes a skill
may not name — the same document is handed to every runtime, so one that names a runtime
is wrong everywhere else. Both are yours to re-choose, and a bound you delete is not
defaulted to a number nobody chose.
letspecify skills --for <role> publishes the instruction document that role is given.
It only reads: it writes no file, runs no version control and opens no network
connection.
letspecify spec-levels [--json] [--lang <id>]
The catalogue of the three depth levels — how much specification a piece of work owes — with the sections each one requires and the criterion an agent reads.
| Level | Requires |
|---|---|
| minimal | objective · acceptance_criteria |
| standard | the above plus constraints · edge_cases · error_behaviour · references |
| exhaustive | the above plus invariants · rejected_alternatives · consequences |
Going up a level ADDS sections; it never changes the shape of the document. They are not three templates: it is one schema with three cuts, and that nesting is checked on every invocation rather than trusted.
The table is the engine's and is not configurable: there is no fourth level, and no project file can add one. That is the whole point — the criterion is the same everywhere.
It needs no governance repository and always exits 0 when it could answer,
inside a governed project or outside one.
letspecify spec-level [-d <dir>] [--branch <name>] [--json]
Which level governs a branch, resolved in two steps:
does the branch declare its own? yes -> that is its level (origin: branch)
| no
the default of its LANE (origin: lane)
| none matches
nothing, which is a normal answer (origin: none)Both live in model/spec-levels.yaml, which init and materialize create.
Where the level comes from matters as much as the level: "standard because
your branch says so" and "standard because your lane does" lead to different
decisions when somebody wants to change it.
--set <level> declares it, which modifies the model. Use --dry-run first
to see exactly what would change; the write is atomic, and it never commits —
driving git is yours.
letspecify tasks [--pending] [--id <FEAT-…|TASK-…>] [--kind <k>] [--json]
The field this reads had been published and read by nobody. depends_on shipped in
the projection and in the template, and nothing in the engine looked at it: no cycle was
detected, no dependency naming a task that does not exist was reported, and nothing
worked out which pending tasks were free of each other.
$ letspecify tasks --pending
47 declared · 10 pending · 37 done pending: 8 development, 2 governance
wave 1 · 2 tasks, nothing between them
TASK-FEAT-2026-007-02 development * todo repo-letspecify
TASK-FEAT-2026-009-01 governance todo repo-engineering
...
* the kind was resolved from the repository, not declared by the task.Every task has a kind, resolved in the two steps spec-level uses: what the task
declares, otherwise the role of its repository — the one role that governs and holds
no product code means governance, every other role means development — and the answer
always says which of the two answered. A task naming no repository the model declares
gets no kind rather than a plausible default, because a plausible default here would
never be checked.
That fallback is what makes the field addable to a project that already has tasks: none of them has to be rewritten to repeat what the model states one file away, and no migration ships with this.
The flow is waves. The first is every pending task waiting on nothing still pending; each one after it is the tasks whose remaining dependencies all sit behind it. A wave is an upper bound on parallelism — it says what MAY run at once, assigns nothing, mints no thread, creates no branch and runs no git.
And a wave names what a dependency graph cannot see. Two tasks that wait on nothing can still be unable to run together, because they write the same file. Where two tasks in one wave declare the same resource in the same repository, the collision is named and neither task is moved — which of the two goes first is not the engine's to decide. The comparison is over the strings the model declares, so a collision through a glob is one this does not see, and the answer says so.
A task in a cycle is given no wave, because inventing an order for it would be
inventing the decision the cycle is missing. It is still exit 0: a cycle and a
dangling dependency are defects of the model, and check (rule R6) is what turns
those into an exit code. A query answers; a verification decides.
--id is resolved by looking it up, not by its prefix: a task id narrows to that
task, an artefact id to its tasks, and anything else is no-such-artifact. An empty
answer says which of three things it was — nothing left to do, an artefact that breaks
down into no task, or no such artefact — and succeeds in all three.
Asking about one feature answers both ends of its derivations. The answer of
--json always carries an artifact key. It is null unless --id named a declared
artefact — so it is null with no --id, with a task's id, and with an id nobody
declares. Otherwise it is { id, title, status, derivesFrom, derivedBy }, and the two
lists are copied from the projection, entry for entry, never recomputed. It is filled
even when no task is shown: --pending on a feature whose tasks are all done answers
tasks: [] and empty: nothing-pending, and artifact.derivedBy is still full.
--pending and --kind narrow tasks and do not narrow it. empty keeps its meaning —
it describes the tasks shown, not the whole answer. The human form prints the block
before Nothing to show, one line per relation, says nothing declared rather than
omitting an empty list, and marks a relation read from depends_on_features. It is one
level each way: a derived feature's own derivations are one more tasks --id away.
letspecify index [-d <repo-engineering>] [--json]
The reverse index: given a file, which fragments of specification govern it?
A loop of agents that opens on a small change used to explore the repository to find that out — every session, from cold, at a cost proportional to the size of the project rather than to the size of the change. Now it reads one document:
$ letspecify index --neighbourhood lib/render.js
repo-example:lib/render.js
CA-34 declared SPEC-ENG-004
RN-21 declared SPEC-ENG-004The artefact is traceability/spec-index.yaml: generated, versioned, sorted and
byte-stable — two runs over the same tree produce the same bytes, so it is
reviewed in a pull request like anything else. It therefore carries no date:
when it was generated, git answers.
| flag | what it answers |
|---|---|
| (none) | coverage, with its denominator, and what is stale |
| --write | generates the artefact. The only path that writes it |
| --neighbourhood <paths> | the hot query: the fragments governing those files |
| --impact <fragment> | the reverse: the files a fragment reaches |
| --ungoverned | tracked files no link reaches |
| --unimplemented | declared fragments no file reaches |
| --verify [--base <ref>] | compares a fresh generation against the committed artefact, bounded to the files the change touches. What a gate runs |
Two sources create links and one never does. A domain spec declares
governs: (paths) and implements: (fragments) in its front matter; a bare
@CA-12 beside the code is the fine one. Git creates none — it only
classifies against the adoption baseline and detects renames.
That baseline is what makes adoption survivable: in a project whose governance arrived after its code, a file untouched since the baseline is not yet indexed — inherited debt, not a defect — while one touched since and still unlinked is ungoverned. Only the second is ever expected to reach zero.
No query regenerates, and an absent artefact answers with exit 0: most
repositories have no index, and without one the loop is slower and no more.
Options
| Option | Description |
|---|---|
| -n, --name | Project name (kebab-case). Required for init; for adopt it is derived from the folder if omitted. |
| -f, --foundational | Path to the foundational .md document. Required for init; optional for adopt (if missing, it is rebuilt from the code). |
| -s, --source | adopt: existing repos to adopt, comma-separated. Without it they are discovered from the current folder. |
| -r, --repos | Shortcut for a known topology: backend,frontend,infra. Without it, the bootstrap decides. |
| -d, --dir | init: target folder / adopt: where to create the governance repo (default: next to the adopted repos) / materialize, model, check, gates: the governance repo. |
| --owner | Owner of the project-model (default: git config user.email). |
| --refresh | adopt: photograph the registered repos again and rewrite only adoption/inventory.yaml. |
| --json | model/check/gates: emit the JSON instead of the human summary. |
| --strict | check: warnings change the exit code too. |
| --branches | check: also look at the other branches for numbering. |
| --fast | gates: skip the runners declared cost: slow. |
| --flow <WF-...> | gates: declare the workflow, for gates with applies_to. |
| --write | gates: append the verdicts to traceability/gate-results.yaml. |
| --pending | tasks: only what is not done. |
| --id <id> | tasks: narrow to one FEAT-/BUG-/TASK- / task-status: the task whose state is recorded (required) / launch-: the task the launch is for. |
| --set <v> | spec-level: the level to declare / task-status: the state to record (required) / launch-: the role-set. |
| --kind <k> | tasks: only development or only governance work. |
| --for <role> | skills: the role whose document is published. Bound to a role and to nothing else. |
| --dry-run | adopt/materialize/spec-level --set/task-status: show what would happen without writing anything. |
| --no-git | Do not initialise git repos or make initial commits. |
| --force | Write even if the target exists (adds/overwrites, never deletes). |
Exit codes
CONTRACT-ENGINE-001 §2. They are the signal; stdout is not parsed to decide
whether something worked.
| | |
|---|---|
| 0 | success |
| 1 | failure (for check, errors were found; for gates, a blocking gate is red; for task-status, the model is defective or the file could not be written; for skills, the module is defective or the skills a role is owed cannot be resolved) |
| 2 | usage error (arguments) |
| 3 | the directory is not a governance repo |
| 4 | the identifier names nothing this model declares — for task-status, no declared task; for skills --for, a role no launch set declares |
3 exists because "wrong arguments" and "this has no governance" are opposites
for a caller: the second is a normal answer — most repositories in the world have
no governance — and the first is the caller's own bug.
4 exists for the same reason one step in: this identifier is not in this model
is a normal answer about a well-formed request, so folding it into 2 would
tell a caller to fix arguments that are fine, and folding it into 1 would tell it
the model is broken when the model is intact.
A code means one thing everywhere. The four above keep their meanings in every command; where an outcome none of them names has to be told apart, the command declares a code of its own in its specification rather than leaving a caller to discover it by running the command and watching. One number never carries two meanings across commands.
Installation
npm install -g letspecifyNode >= 16. No dependencies except js-yaml.
From the repository, to develop:
npm install && npm linkLanguage
The engine speaks English, and so does the governance it generates: a
governance project created by letspecify init or letspecify adopt is born in
English.
What it reads is another matter: check, gates and model understand a
governance repo written in any language, because they recognise artefacts by
directory and by identifier, never by prose. The reference case is the project
this tool was built for, whose governance repo is in Spanish and is checked with
exactly the same findings as before the engine was translated.
Maintaining the templates
The templates live in templates/ with {{VAR}} placeholders:
templates/engineering/— the shared skeleton of the governance repo (bothinitandadoptuse it).templates/adopt/— rendered on top of the previous one duringadopt:ADOPTION.md(which replacesBOOTSTRAP.md),adoption/and theFOUNDATIONAL.mdtemplate to be rebuilt.templates/repo/andtemplates/src/— the product repos created bymaterializeand byinit.
The files that depend on the topology (domains.yaml, repositories.yaml,
tasks.yaml, domain agent blocks) are generated in lib/repos.js. Their
generators work on "units": the REPO_DEFS archetypes for init -r, or the ones
lib/adopt.js builds from real repos (with their own repoName, localPath,
buildCmd…). lib/inspect.js is the reconnaissance of an existing repo
(read-only) and lib/materialize.js the materialisation from the model. The
names _gitignore/_gitattributes/_gitkeep are renamed to dotfiles on copy.
A placeholder with no value makes generation fail on purpose (lib/render.js),
so that a broken template cannot silently produce a half-finished project. If you
add one, give it a value on all three paths: init without -r, init -r
(buildContext in lib/init.js) and adopt (buildAdoptContext in
lib/adopt.js).
Tests
npm test # node --test test/Node >= 20 to run them: node --test <directory> does not accept a directory
until 20. The CLI itself works from 16, which is what engines declares — they
are two different requirements and it is worth not confusing them.
They cover the inspection heuristics, how adopt resolves its sources and
destination, that the adopted repos are left untouched, and the regression of
init with and without -r (valid YAML and no unresolved placeholders in all
three modes).
License
MIT © kaerumotoharu-sudo
