npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

musubix3

v0.1.22

Published

Evidence-driven specification development skills for GitHub Copilot CLI

Readme

musubix3

Latest release v0.1.22 · GitHub Copilot CLI only · Node.js ≥20 · TypeScript · MIT

日本語

What changed from musubix2 to musubix3 (Japanese)

GitHub Copilot can plan, generate, edit, test, and review software. musubix3 adds repository-local specifications plus deterministic, fail-closed checks over the evidence required by the repository's configured quality profile: requirements → constitution → design/ADRs → implementation → traceability → quality evidence.

Learned from musubix2's concepts, rebuilt cleanly in three workspaces. No artifact compatibility or migration is promised. This repository does not guarantee correctness simply because IDs are linked or requirements are satisfiable.

Why GitHub Copilot alone is not enough

Copilot is the implementation engine. It understands a request, explores the repository, proposes a plan, edits files, runs tools, and explains the result. That is necessary, but a successful conversation is not durable proof that:

  • the implemented behavior matches an explicit, measurable requirement;
  • every requirement is connected to design, code, and an authoritative test;
  • Red really failed before Green passed without the test being rewritten;
  • test, graph, formal, and quality results still match the current source;
  • a requirement change propagated through all affected artifacts in order;
  • a policy was not weakened merely to make the final gate pass.

Conversation text such as “tests passed” or “implementation complete” is not enough because it can become stale, omit scope, or disappear outside the repository. Copilot should remain responsible for reasoning and development; musubix3 makes the completion criteria persistent and machine-checkable.

| GitHub Copilot provides | musubix3 complements it with | |---|---| | Planning, coding, refactoring, and tool execution | Repository-local SDD Skills that require explicit requirements, design decisions, implementation links, and completion conditions | | Test generation and runner execution | Structured TEST identities, native-report normalization, and verified Red/Green/Refactor evidence | | Explanations of what changed | Typed requirement → design → code → test traceability and bidirectional impact analysis | | Repository exploration | Deterministic Code Graph indexing, unresolved-local dependency diagnostics, and architecture gates | | Suggestions for constraints and invariants | Optional Z3/Lean consistency checks and model-to-passing-test correspondence | | A session-level completion report | Freshness, fingerprints, input-stability checks, protected policy baselines, attestations, and a fail-closed readiness gate |

musubix3 does not replace Copilot, add another coding agent, or claim that formal satisfiability proves implementation correctness. Copilot performs the development; musubix3 records the specification, checks the evidence, rejects stale or incomplete required evidence, and leaves a reviewable answer to “why does this repository's configured policy consider the change ready?”

Quick start

There are two ways to load musubix3's skills into a project — pick exactly one per project, never both:

  • npm install (this section, recommended): add musubix3 as a project dev dependency, then run npx musubix3 init to copy the skills into .github/skills/ and scaffold starter .musubix/ artifacts. Version and upgrades are tracked in your project's package.json/lockfile (npm install ...@latest + npx musubix3 upgrade); see Upgrade. Recommended because the skills and starter SDD artifacts are committed to the repository itself, so every contributor and CI job gets the same reproducible setup without each of them running a separate plugin install.
  • copilot plugin install (native plugin/marketplace route): register musubix3 directly with Copilot CLI's own plugin manager — no npm dependency is added to the project, and no files are copied into .github/skills/. Upgraded with copilot plugin update musubix3 (or the marketplace equivalent). See Distribution options for the exact commands.

The rest of this section documents the npm install route.

For a reproducible project-local installation:

npm install --save-dev --save-exact musubix3@latest
npx --no-install musubix3 --version
npx --no-install musubix3 init --dry-run
npx --no-install musubix3 init
copilot

For a one-time evaluation without pinning subsequent Skill-driven CLI runs:

npx musubix3@latest --version
npx musubix3@latest init --dry-run

Install the exact local dependency before relying on generated Skills in continued development; their commands intentionally use the repository-local npx --no-install musubix3 executable.

To build the repository itself:

git clone https://github.com/nahisaho/musubix3.git
cd musubix3
npm install
npm run build
node dist/packages/cli/src/main.js --help

Ask Copilot: “Use sdd-change to add this feature and propagate it through the specification, implementation, traceability, and quality gate.” All eight skills instruct Copilot to follow your input language (English/Japanese).

init (install alias) copies repository-local skills and creates starter SDD artifacts. It preserves existing files, merges a cache ignore rule, and is idempotent. --force replaces only named bundled/managed paths; it does not delete unrelated files. Review its dry-run first. No Copilot global settings, MCP, LSP, hooks, or project instructions are overwritten. Symbolic-link write targets and paths escaping the project are refused.

The starter is deliberately not release-ready: replace its example, implement and test it, and configure real check commands before expecting the gate to pass.

Upgrade

upgrade is available starting with 0.1.14 (not in 0.1.13 or earlier — check with npm view musubix3 versions).

For the npm-installed route (init/repository-local skills):

npm install --save-dev --save-exact musubix3@latest
npx --no-install musubix3 upgrade --dry-run
npx --no-install musubix3 upgrade

upgrade refreshes only the bundled .github/skills/sdd-* files that differ from the installed package version. It never creates, replaces, or deletes .musubix/config.json, .musubix/policy-baseline.json, .musubix/constitution.md, ADRs, feature artifacts, evidence, or .gitignore — those are yours to keep customizing. It is idempotent: it compares against the currently installed package's skill files, so running it again without first installing a newer musubix3 version reports every skill file as unchanged. Review --dry-run first, same as init.

init --force remains available for replacing bundled/managed paths more broadly (including .musubix/config.json, constitution.md, and the starter feature's requirements.md/design.md), but it overwrites customized content in those files, so prefer upgrade for routine version bumps.

For the native plugin route:

copilot plugin update musubix3

For the native marketplace route:

copilot plugin marketplace update musubix3-marketplace
copilot plugin update musubix3

These delegate to Copilot's own plugin manager, which re-fetches plugin.json/the marketplace catalog; they do not touch .musubix/ at all (only the npm-installed route manages .musubix/ artifacts).

Distribution options

Choose one skill-loading route to avoid duplicate skill names.

Native plugin (direct)

Pick exactly one of the following (not both):

copilot plugin install ./musubix3            # local clone, from its parent directory
copilot plugin install nahisaho/musubix3     # published GitHub repository, no local clone needed

plugin.json at the repository root is the source of truth and references .github/skills/. The plugin contains skills, not an agent runtime or background services. Git installs do not compile/install the npm engine: build the clone or install the npm package separately when running npx musubix3 commands.

There is a third way to run the same copilot plugin install, useful only if you want a durable local path instead of an ephemeral npx cache — do this instead of, not in addition to, the two commands above:

npm install --save-dev --save-exact musubix3@latest
npx --no-install musubix3 plugin-install

plugin-install only calls copilot plugin install <absolute-package-root> for you; it does not edit Copilot internals, and it is unrelated to the npm-installed skill-copy route in Quick start — it never runs init and never copies files into .github/skills/. Do not run both this and npx musubix3 init in the same project.

Native marketplace

copilot plugin marketplace add nahisaho/musubix3
copilot plugin install musubix3@musubix3-marketplace
# Local development:
copilot plugin marketplace add ./musubix3

The catalog is .github/plugin/marketplace.json; its plugin source is .. These flows use the native plugin interface.

Repository-local skills / npm installer

Run npx --no-install musubix3 init after installing the exact local dependency, or copy .github/skills/sdd-* there yourself. Start Copilot in that trusted project. init --root <dir> targets another project; --feature <slug> changes the starter directory and ID prefix. Installing another feature does not reset existing configuration.

Skills and native boundaries

| Skill | Purpose | |---|---| | sdd-change | End-to-end feature/change/fix propagation and completion gate | | sdd-requirements | Six controlled EARS forms and measurable constitution | | sdd-design | Explicit responsibilities/interfaces/constraints, ADRs, diagrams | | sdd-implementation | Native editing with requirement-linked code and tests | | sdd-traceability | Generated coverage, dangling links, bidirectional impact | | sdd-quality | Actual verification commands, policy, readiness evidence | | sdd-knowledge | Local artifact/Git retrieval, not conversational memory | | sdd-formal-codegraph | Optional consistency checks, compiler graph, architecture |

Use native Copilot for planning, editing, research, review, security review, memory, code navigation/LSP, MCP management, and subagent/fleet/task coordination. musubix3 does not implement those services, generic code/test generation, orchestration, a scheduler, MCP server, Claude support, an interactive REPL, or a resident watcher. status, query, impact and --changed provide one-shot value. Skills may combine neural proposals from Copilot with symbolic checks; there is no separate “neurosymbolic AI” model or claims of learned verification.

Workflow

Natural-language requests such as “develop/build/create/implement X” activate sdd-change as the mandatory first Skill. It elicits and validates requirements and design before implementation; only an explicit request to implement existing approved artifacts may enter sdd-implementation directly. The implementation Skill fails closed when those validated artifacts are absent or invalid. When material context is missing, requirements elicitation asks one highest-priority question at a time and waits for the answer; it does not batch questions or finalize requirements while blockers remain.

  1. Use native planning/research to establish intent and measurable acceptance.
  2. Record change-record CHANGE-ID impact, edit/validate requirements, run a rubber-duck review of the requirements document and fix every finding (repeat review/fix until zero issues remain), record the requirements checkpoint, then obtain explicit artifact-bound human requirements approval before design.
  3. Design explicit components, record trade-offs in ADRs, run a rubber-duck review of the design document/ADRs and fix every finding (repeat until zero issues remain), record design, then obtain explicit artifact-bound human design approval before implementation.
  4. Write an annotated behavior test, record a structured failing tdd red, then record the change red checkpoint.
  5. Implement the minimum change, record implementation, run passing tdd green, then record green and refactor. For a change with multiple requirements, red/implementation/green may instead be recorded once per requirement subset (an independent batch per subset), completing each requirement's Red-Implementation-Green loop before moving to the next, instead of completing every requirement's Red before any Implementation.
  6. Add trace annotations, build graphs, inspect impact and fix missing coverage.
  7. Configure real checks and run the candidate gate --changed. Run a rubber-duck review of the release/quality evidence summary and fix every finding (repeat until zero issues remain). After every required non-approval check passes, obtain explicit human release approval, rerun the gate/status, and only then commit, push, publish, or deploy.

Any AI-generated documentation deliverable (requirements, design, ADRs, the CHANGE document, or release/quality evidence summaries) goes through this rubber-duck review/fix loop until zero issues remain before the corresponding human approval step is requested.

npx musubix3 requirements validate .musubix/features/example/requirements.md --json
npx musubix3 constitution validate --json
npx musubix3 approval prepare requirements --json
npx musubix3 approval record requirements --approver "Requirements Owner" --artifact-sha256 "$REVIEWED_HASH" --confirm
npx musubix3 design validate .musubix/features/example/design.md --json
npx musubix3 design c4 .musubix/features/example/design.md
npx musubix3 approval prepare design --json
npx musubix3 approval record design --approver "Design Owner" --artifact-sha256 "$REVIEWED_HASH" --confirm
npx musubix3 change-record CHANGE-0001 design --requirement REQ-EXAMPLE-001
npx musubix3 tdd red TEST-EXAMPLE-001 --requirement REQ-EXAMPLE-001 --command test
# Implement the minimum behavior without changing the test.
npx musubix3 tdd green TEST-EXAMPLE-001 --requirement REQ-EXAMPLE-001 --command test
npx musubix3 tdd refactor TEST-EXAMPLE-001 --requirement REQ-EXAMPLE-001 --command test
# A change declaring REQ-EXAMPLE-001 and REQ-EXAMPLE-002 may record red/implementation/green
# once per requirement subset instead of once for the whole change:
npx musubix3 change-record CHANGE-0001 red --requirement REQ-EXAMPLE-001
npx musubix3 change-record CHANGE-0001 implementation --requirement REQ-EXAMPLE-001
npx musubix3 change-record CHANGE-0001 green --requirement REQ-EXAMPLE-001
npx musubix3 trace build
npx musubix3 trace check --strict --json
npx musubix3 graph index
npx musubix3 graph impact src/service.ts
npx musubix3 gate --changed --json
npx musubix3 approval prepare release --json
npx musubix3 approval record release --approver "Release Owner" --artifact-sha256 "$REVIEWED_HASH" --confirm
npx musubix3 gate --changed --json
npx musubix3 approval validate --json
npx musubix3 status --json

A requirement already satisfied as a side effect

tdd red rejects recording a Red phase for a test that is already passing (TDD_TARGET_RESULT). This is correct: it is not a bug when a correct, general implementation written for one requirement also happens to satisfy a separate, not-yet-implemented requirement (for example, a general capacity-aware assignment algorithm that also correctly handles a "no resource available" edge case required by a different requirement). There is deliberately no --already-satisfied-by-style declaration to bypass Red for this case: a human declaration that a requirement is "already satisfied elsewhere" is not measured evidence, and accepting one would let an unimplemented or incorrectly implemented requirement pass gate on an unverified claim.

The recommended practice is to prove the Red genuinely, the same way as any other requirement, by briefly and deliberately narrowing the scope of the already-correct shared implementation so the new requirement's test fails, recording tdd red, then restoring the correct implementation for tdd green:

# The shared implementation is already correct; temporarily narrow it
# (e.g. re-add the special case the general algorithm already subsumes)
# so TEST-EXAMPLE-002 genuinely fails.
npx musubix3 tdd red TEST-EXAMPLE-002 --requirement REQ-EXAMPLE-002 --command test
# Restore the correct (already-written) implementation; no other code changes.
npx musubix3 tdd green TEST-EXAMPLE-002 --requirement REQ-EXAMPLE-002 --command test

Keep the narrowing edit local and revert it in the same step as recording Green; do not leave the weakened code committed at any point.

A brand-new module for a from-scratch requirement

tdd red requires a collected and executed failing test, not merely a test file that exists: if the test imports a source module that does not exist yet, the runner fails at collection time (0 tests run) instead of failing an assertion, and tdd red correctly reports this as an invalid Red (TDD_REPORT_INVALID). For a vitest/jest adapter, the thrown error now appends the underlying suite collection-failure message (for example Cannot find module '../src/service.js') after No annotated TEST-* identities were found, to make this cause visible immediately instead of requiring a separate investigation.

The recommended practice for a genuinely new module is the same stub-then-correct technique as above: create a compiling implementation file with a deliberately wrong body first, so the runner can collect and execute the test (and it fails on the wrong behavior, not on a missing import), then implement the correct behavior for Green:

# src/service.ts does not exist yet; create it with an intentionally wrong body
# so the test can be collected and genuinely fails on assertion, not on import.
npx musubix3 tdd red TEST-EXAMPLE-003 --requirement REQ-EXAMPLE-003 --command test
# Replace the stub body with the correct implementation.
npx musubix3 tdd green TEST-EXAMPLE-003 --requirement REQ-EXAMPLE-003 --command test

Command reference

All analysis commands accept --root <directory> and --json. Files resolve relative to that root; access outside it is refused. plugin-install uses the installed package root. Exit codes: 0 successful operation, 1 rejected validation/gate or requested solver failure, 2 usage, I/O or malformed config. status is informational (exit 0 even if not ready); inspect gate.ready.

| Command | Behavior | |---|---| | init [--dry-run] [--force] [--feature slug] | Preserve-first skills/artifacts installation; install alias | | upgrade [--dry-run] | Refresh bundled skill files only; never touches config, constitution, or feature artifacts | | plugin-install | Invoke native Copilot installer (no internal config edits) | | requirements validate <file> | IDs, priorities, declared/detected EARS pattern | | requirements scaffold <slug> [--title <text>] | Create .musubix/features/<slug>/requirements.md from a fixed placeholder template; never overwrites an existing file | | constitution validate [file] | Versioned principles and measurable rule definitions | | design validate <file> | Fields, global requirement IDs, existing ADR references | | design scaffold <slug> | Create .musubix/features/<slug>/design.md from a fixed placeholder template; never overwrites an existing file, and does not require a pre-existing requirements.md | | design c4 <file> | Mermaid component/dependency diagram from explicit fields | | approval prepare <requirements\|design\|release> [--diff-only] | Display the exact deterministic manifest and hash for human review; --diff-only additionally reports changedFiles (only paths whose content differs from the last recorded approval for that stage; all current paths with diffOnlyBaseline: "none" when no prior approval exists) without altering artifactSha256/artifacts | | approval record <stage> --approver <name> --artifact-sha256 <hash> --confirm [--fast-reapprove --own-files <path...>] | Record approval only if the reviewed hash is still current. --fast-reapprove --own-files <path...> re-records a stale approval after the reviewer has manually confirmed that a rebase only touched files outside the --own-files list they name: it requires a prior approval by the same approver and rejects (naming the paths) if any listed own file itself changed, if any listed path is unknown to both the prior and current manifest, or if the own-files list is empty; --own-files without --fast-reapprove is rejected. The system verifies only the paths the reviewer lists — it cannot detect an own file the reviewer forgot to list, so the reviewer remains accountable for supplying a complete own-files set | | approval validate | Report each approval as approved, missing, or stale; never infer approval from validation | | trace build | Generate global trace snapshot and feature copies | | trace check [--strict] | Dangling IDs, stale inputs/paths, mandatory coverage | | trace impact <id-or-path> | Bidirectional breadth-first traversal with explanation paths | | graph index [--changed] | Compiler imports, declarations and best-effort call targets | | graph impact <symbol-or-path> | Conservative reverse-import closure; path#name disambiguates | | graph cycles | Strongly connected components; exit 1 when cycles exist | | graph gate | Fresh index + architecture rule/cycle checks | | knowledge build | Markdown and bounded Git evidence index | | knowledge query <text> [--limit 10] | Deterministic TF-IDF/cosine results and staleness flag | | formal generate <file> [--format both\|smt2\|lean] | Reproducible solver inputs with SHA-256 evidence | | formal doctor | Probe Z3, Lean, and lake env lean availability and versions | | mutation doctor | Probe language-aware local mutation engines and show setup recommendations | | formal check <file> [--solver auto\|none\|z3\|lean] | Check the explicit Boolean/conditional/numeric/temporal/transition model | | model-correspondence validate | Revalidate Formal JSON → generated trace → authoritative passing test evidence (run evidence refresh first to generate its evidence file) | | evidence refresh [--changed] | Regenerate derived evidence through the same fail-closed gate pipeline | | evidence unlock --recover | Recover only a demonstrably dead same-host evidence-writer lock. Live, cross-host, malformed, PID-reused, unsupported-platform, and indeterminate owners are refused without mutation. | | evidence merge --incoming <directory> [--dry-run] | Merge the current root's valid order/TDD/change/waiver history with another valid project history. Base records remain first; exact duplicates are deduplicated; ambiguous payloads, chronology inversions, and post-Quality batch additions fail without writes. Dry-run performs the same planning and validation. The incoming directory is never modified. | | evidence merge --recover | Recover an interrupted evidence merge from its journal. Prepared transactions roll back; committed transactions verify and roll forward. If recovery reports EVIDENCE_MERGE_RECOVERY_UNSAFE, back up .musubix/evidence, restore or verify order.json, tdd.json, changes.json, and change-waivers.json from a trusted source, quarantine merge journal/temporary files, then rerun validation. | | mutation validate | Revalidate requirement-scoped schema-v1 killed-mutant evidence | | mutation identity <REQ-ID> <TEST-ID> <sourcePath> <operator> <line> <column> | Print the deterministic MUT-* identity a mutation report must declare | | tdd validate | Validate persisted Red/Green/Refactor order, fingerprints, durations, and hash-chain evidence | | tdd red\|green\|refactor <TEST-ID> --requirement <REQ-ID> --command <name> | Execute and record a verified TDD phase. Recording the project's first tdd cycle (any red call while .musubix/evidence/tdd.json has zero cycles) activates gate's tdd coverage evaluation project-wide, for every mandatory requirement, not just the ones touched by the current change; each uncovered requirement then surfaces as TDD_REQUIREMENT_UNCOVERED. If tdd was not already required for another configured reason, that same first cycle is also what makes the check required. approval record release always runs the full (non---changed) gate, so it is blocked by any resulting TDD_REQUIREMENT_UNCOVERED diagnostics. tdd migrate cannot be used to bulk-onboard previously-uncovered requirements: it only reuses already-covered valid Green-backed evidence. tdd red prints/returns a TDD_ADOPTION_PROJECT_WIDE warning (in a warnings array, separate from diagnostics) the moment this first cycle is persisted, listing every other still-uncovered mandatory requirement. | | tdd migrate <test-id> / tdd migrate <old-id> <new-id> | Relink already-covered evidence without fabricating a fresh Red/Green cycle. tdd migrate <test-id> preserves the existing fingerprint-migration mode for a single covered test whose current declaration still matches the stored evidence under the superseded fingerprinting algorithm. tdd migrate <old-id> <new-id> relinks a pure identifier rename when the current authoritative new-id declaration differs from old-id only in its @id annotation and adapter-matched test identity string, the non-test sourceFingerprint is unchanged, and the renamed test remains the sole authoritative declaration at the same test path. | | workflow-record <skill> <phase> --status <status> | Record a compact self-reported workflow declaration | | workflow waiver record <code> --skill <skill> --phase <phase> --recorded-at <timestamp> [--index <n>] --approver <name> --reason <text> --confirm | Record an audited, bounded downgrade of one declaration-scoped workflow reconciliation diagnostic (WORKFLOW_SKILL_NOT_INVOKED, WORKFLOW_INVOCATION_ORDER, WORKFLOW_INVOCATION_INCOMPLETE, WORKFLOW_INVOCATION_FAILED, or WORKFLOW_INVOCATION_REUSED) to a non-blocking waived status; a paired WORKFLOW_BINDING_MISSING diagnostic sharing the same declaration scope is downgraded together with it. It cannot waive WORKFLOW_INVOCATION_UNVERIFIED, which only workflow-verify having actually run this session can resolve | | workflow waiver record-all --approver <name> --reason <text> --confirm | Bulk variant of workflow waiver record: waives every currently-outstanding waivable declaration-scoped diagnostic across the whole reconciliation report in one all-or-nothing call, instead of one record invocation per diagnostic. Rejects (recording nothing) if any bulk waiver precondition fails first — malformed waiver evidence, an invalid waiver chain, a blank --approver/--reason, or a WORKFLOW_INVOCATION_UNVERIFIED diagnostic (which workflow-verify must resolve first; no per-declaration or bulk waiver can substitute for workflow-verify never having run). Each waived declaration-scoped diagnostic's paired WORKFLOW_BINDING_MISSING is downgraded together with it, identically to the existing single-record command. With zero remaining candidates (already waived, or none present) it succeeds idempotently and records nothing | | workflow-sanitize <copilot.jsonl> <output-file> [--session-id <uuid>] | Remove messages and non-Skill tool data before review or strict verification | | workflow-verify <copilot.jsonl...> [--strict] [--session-id <uuid>] | Reconcile Skill events; compatible mode accepts multiple transcript files (concatenated in chronological session order) so declarations whose invocation occurred in an earlier Copilot CLI session can be reconciled; --strict/--session-id still require exactly one file | | attestation oidc-audience --key-id <id> [--public-key-file <pem>] | Derive the GitHub custom audience that authorizes a signing key | | attestation payload --provider <name> --run-id <id> --key-id <id> [--public-key-file <pem>] [--github-oidc-token-file <jwt>] | Emit canonical unsigned CI payload for external signing | | attestation verify | Verify static-key or GitHub OIDC-authorized Ed25519 provenance | | change-record <CHANGE-ID> <phase> --requirement <REQ-ID...> [--allow-unchanged] [--dry-run] | Record ordered artifact/TDD fingerprints for a staged change. impact/requirements/design/quality require the change's full requirement ID set; red/implementation/green also accept a proper non-empty subset, recorded as an independent per-requirement batch, so a multi-requirement change can be completed with an interleaved per-requirement Red-Implementation-Green loop instead of one global batch. When multiple batches contain the same requirement, validation uses the batch with the latest Red order; a later incomplete batch supersedes older evidence and must be completed rather than silently falling back. Rejects with a non-zero exit and a stable *_UNCHANGED_AT_RECORD diagnostic (leaving changes.json/order.json untouched) when a phase's fingerprint is byte-identical to the immediately preceding phase's, since that mistake is otherwise only caught much later by trace check --strict/gate, at which point the append-only evidence store makes it unfixable. --allow-unchanged bypasses this check for the requirements phase only (the one case sdd-change documents as legitimate: a defect fix that intentionally leaves its requirement unchanged) and persists a durable marker so later validation does not re-flag it. --dry-run previews the exact success/rejection outcome, including every existing check, without persisting anything. Each recorded phase/batch stores both order (the verified, gate-checked logical append sequence from order.json — the only field guaranteed correct and monotonic per change) and recordedAt (an independently captured wall-clock timestamp with no ordering guarantee relative to order); gate reports a non-blocking CHANGE_RECORDEDAT_OUT_OF_ORDER warning when a change's recordedAt values disagree with its order sequence | | change quality-recover [--json] | Recover an interrupted atomic Quality refresh. A repeated full-set Quality after a newer complete corrective batch retains earlier checkpoints in schema-version-2 qualityHistory; incomplete/current or unnecessary refreshes fail with stable CHANGE_QUALITY_REFRESH_* errors. | | config lint | Report configured commands whose args reference repository-relative paths that do not exist | | config scaffold | Propose native test-command entries for detected Go/Rust/Maven/Python/Node toolchains without writing .musubix/config.json | | gate [--changed] [--feature <name>] | Fresh full checks plus actual configured commands; persist evidence. --feature scopes requirements/design/trace/tdd/change-history/change-completeness checks to one feature as a diagnostic view; never a substitute for the repository-wide gate | | status | Artifact counts and readiness/staleness summary |

Evidence writer coordination

Commands that mutate evidence or generated project state use one fail-fast lock at <realpath(project-root)>/.musubix/evidence/.writer-lock.json. This includes init/upgrade writes, gate and evidence refresh, TDD/change/workflow/approval recording, trace/graph/knowledge/formal generation, and evidence merge/recovery. Complete owner metadata is staged, flushed, and published with a same-directory exclusive hard link, so the canonical lock is never visible as empty or partial. Filesystems that cannot provide this atomic publication fail with EVIDENCE_WRITER_LOCK_ACQUIRE_FAILED; there is no non-atomic fallback. When Windows directory synchronization rejects EPERM, EINVAL, or ENOTSUP after successful staging-file synchronization and atomic publication or verified release, musubix3 treats only that directory-entry durability operation as an unsupported capability. File synchronization, publication, metadata, unlink, open/close, every other error, and every non-Windows platform remain fail-closed.

Coordinated readers such as status, approval preparation/validation, trace and graph inspection, knowledge queries, TDD validation, attestation payload/verify, and merge dry-run fail immediately with EVIDENCE_WRITER_LOCKED when an unrelated owner exists. They do not wait or hold a read lease. Nested analysis operations in one owner async context reuse the lock; child processes and worker threads do not inherit that context. A configured command that recursively runs a musubix3 writer or coordinated reader against the same root therefore fails fast and must be removed from that command chain. Programs that write project files directly outside musubix3 are not coordinated.

An abrupt exit can leave the lock. Run evidence unlock --recover; automatic recovery is currently Linux-only and requires matching hostname, boot identity, PID namespace, and a demonstrably absent owner PID. Live, PID-reused, cross-host, malformed, changed, unsupported, or indeterminate locks are left untouched with the exact path and inspection guidance. There is no force mode. Because the required identity probes are unavailable, automatic recovery remains inspection-only on Windows and macOS. The canonical lock and .writer-lock.<transactionId>.json staging files are ignored and excluded from generated inputs. After confirming no related process is active, an abandoned staging file may be removed by its exact path.

Writer-lock checks precede merge-journal checks. When both an abandoned writer lock and pending merge journal exist, run evidence unlock --recover first and evidence merge --recover second. Unlock recovery never reads or modifies the merge journal.

Evidence-history merge

evidence merge requires both roots to have individually valid order.json, tdd.json, and changes.json; change-waivers.json is optional. The incoming real path must be disjoint from the current root, remains read-only, and the current worktree must already contain each newly introduced incoming change document. Duplicate comparison canonicalizes JSON after removing only rebuilt sequence/order/hash-link fields. The result keeps base history first and then incoming relative order, so wall-clock recordedAt warnings can remain even when logical order is valid. If base Quality precedes an incoming Red or Implementation batch, merge the histories before recording Quality and record Quality again against the merged history.

Use --dry-run for the same conflict, candidate, stale-waiver, and supersession analysis without writes. A successful merge can make waiver snapshots stale and inactive or select a later authoritative waiver; explicitly re-approve those waivers and rerun normal gate and status. Only order.json, tdd.json, changes.json, and change-waivers.json are merged. Workflow, approval, quality, formal, mutation, correspondence, performance, attestation, and native test-report evidence must use their existing regeneration or conflict-resolution workflows. EVIDENCE_MERGE_CANDIDATE_INVALID reports each underlying file/entity diagnostic. Use --recover after interruption; for EVIDENCE_MERGE_RECOVERY_UNSAFE, back up .musubix/evidence, restore or verify all four targets from a trusted source, quarantine merge-owned journal and temporary files, and rerun structural validation.

--changed reads staged, unstaged, untracked and renamed/deleted paths from Git. It reports affected files but conservatively recomputes all deterministic checks and executes all configured commands. This avoids unsafe incremental skips. There is no daemon, polling loop or background service. Workflow evidence automatically requires workflow. TDD evidence automatically requires tdd; a .musubix/changes/CHANGE-*.md document additionally requires tdd, change-history, and change-completeness, even when omitted from requiredChecks.

Artifact schema (v1)

.github/skills/sdd-*/SKILL.md
.musubix/
  config.json
  constitution.md
  features/<slug>/
    requirements.md
    design.md
    trace.json                 # generated; never hand-edit
  decisions/ADR-0001.md
  evidence/
    quality.json                # actual gate report, initially skipped
    workflow.json               # declarations plus optional strict transcript/session evidence
    tdd.json                    # append-only phase hash chain and TDD cycles
    changes.json                # staged change checkpoints
    order.json                  # shared monotonic TDD/change chronology ledger
    performance.json            # deterministic operation-budget observations
    model-correspondence.json   # Formal model → trace → fresh passing-test correspondence evidence
    mutation.json               # fresh requirement-scoped mutation executions
    attestation.json            # optional externally signed CI provenance
  cache/                       # ignored; generated indexes and solver inputs

Requirements and design

---
schemaVersion: 1
feature: auth
---
## REQ-AUTH-001: Reject expired sessions
Priority: must
Type: functional
Pattern: event-driven
Statement: When a session expires, the system shall reject the request.
Acceptance: An expired-session request produces HTTP 401.
Formal: {"kind":"conditional","condition":"session.expired","consequence":"request.rejected"}

## DES-AUTH-001: Session guard
Responsibilities: Reject requests whose session has expired.
Interfaces: guard(request) returns a principal or HTTP 401.
Constraints: Do not log session tokens.
Requirements: REQ-AUTH-001
ADRs: ADR-0001
Depends-On: DES-AUTH-002

Put these entries in their respective requirements.md / design.md; declare every dependency as another component. A component that genuinely has no architecturally significant decision may write ADRs: none — <reason> (a concrete, non-placeholder reason) instead of a real ADR reference; a bare none, an empty field, or a placeholder reason (TODO/TBD/N/A/未定) still fails validation, now with DES_ADR_EXEMPTION_REASON when a "none" marker is present without a usable reason. Each requirement has one controlled statement. Accepted priorities are must (default), should, may. Requirement types are functional (default) and non-functional. Formal: is optional strict single-line JSON. Supported kinds are conditional, numeric (integer comparison), temporal (withinMs plus optional nonnegative afterMs), and transition (from/event/to). Numeric units ms/s/min share an exact duration dimension, while bytes/kib/mib share an exact size dimension. Other units and incompatible dimensions remain separate. It models only the declared fields. A non-functional requirement may also declare Performance: {"counter":"visitedNodes","max":100,"testId":"TEST-AUTH-002"}; the named passing test must report that integer operation counter. IDs use uppercase REQ-, DES-, CODE-, TEST-, a feature prefix, and ≥3 digits. ADRs use ADR- plus ≥4 digits. IDs must be globally unique.

Six EARS patterns: “The system shall …”; “When …, the system shall …”; “While …, …”; “If …, then …”; “Where …, …”; combined distinct Where/While/When clauses. Japanese controlled forms are documented in README-ja.md. These are syntax checks, not natural-language understanding; arbitrary prose is deliberately rejected.

Source/test trace annotations are read from language comments. JS/TS uses parser-aware comment locations; Haskell supports -- and {- ... -}, Lua supports -- and --[[ ... ]], and Visual Basic supports apostrophe comments including ''' XML documentation. These scanners exclude string literals; other languages require their supported line or block comments. Python annotations must use consecutive # comments; annotation-like text in docstrings is ignored with TRACE_ANNOTATION_IN_PYTHON_DOCSTRING guidance:

/** @id CODE-AUTH-001
 * @implements REQ-AUTH-001
 * @design DES-AUTH-001
 */
export function guard() { /* actual implementation */ }

/** @id TEST-AUTH-001
 * @verifies REQ-AUTH-001
 */
// Real behavior test follows.

One block comment per entity; comma/space-separated targets. @design is optional. This doc comment and native test-report ID matching are two independent requirements that must both be satisfied for tdd red/tdd green to succeed. The @id TEST-*/@verifies REQ-* comment only makes the test discoverable as a kind: 'test' trace-graph node, so the CLI can resolve its requirement link. Separately, the configured adapter (or tddReport) must match that same ID inside the executed native test report, using its own adapter-specific mechanism — see the adapter reference table below. A test with only the doc comment but a report the adapter cannot match to that ID fails Red/Green (the report never contains a recognized TEST-* entry); a test whose report matches but lacks the doc comment fails earlier with Annotated test ID not found: <id> (not present in the trace graph). Both failure modes look similar but have different causes — check the doc comment first, then the adapter's matching rule below. In PHP, use plain /* ... */ blocks rather than /** ... */ PHPDoc: PHPDoc reserves @implements for generic type declarations, so PHPStan/Psalm report phpDoc.parseError on requirement-ID lists inside doc comments. musubix3 reads either form. Mandatory implementation coverage may be direct or through a linked design; tests must directly verify a requirement. Links alone are not semantic proof. Each feature's trace.json holds the complete repository snapshot, including cross-feature edges and input SHA-256 fingerprints; copies intentionally agree. The cache is preferred when present; feature snapshots support cache-free checks. Rebuild when inputs change; stale impact queries are rejected.

Constitution and configuration

---
version: 1.0.0
---
## PRINC-001: Evidence first
### RULE-001: No missing trace coverage
Metric: trace.errors
Limit: 0

Supported metrics are requirements.errors, design.errors, trace.errors, graph.violations, formal.errors, formal.modeledFraction, tests.annotatedIds, tests.executedIds, commands.failures and commands.skipped. Every rule declares a nonnegative numeric upper bound. constitution validate checks the definition; only gate measures it. Unavailable evidence is skipped, never a measured zero.

Example .musubix/config.json (adapt command arguments to your own project):

{
  "schemaVersion": 1,
  "language": "auto",
  "qualityProfile": "custom",
  "commands": [
    { "name": "typecheck", "command": "npm", "args": ["run", "typecheck"], "required": true, "timeoutMs": 120000 },
    {
      "name": "test",
      "command": "npm",
      "args": ["test", "--"],
      "adapter": "vitest",
      "required": true,
      "timeoutMs": 120000
    }
  ],
  "requiredChecks": ["requirements", "design", "constitution", "trace", "graph", "commands"],
  "thresholds": { "design": 1, "implementation": 1, "tests": 1 },
  "formal": { "solver": "none", "minModeledFraction": 0, "timeoutMs": 12000 },
  "mutation": { "mode": "compatible" },
  "tdd": { "redPreflightCommands": [] },
  "approval": { "mode": "required" },
  "workflow": {
    "mode": "compatible",
    "maxAgeSeconds": 3600,
    "maxFutureSkewSeconds": 60,
    "maxEventSkewMs": 1000,
    "maxTranscriptBytes": 250000000,
    "maxTranscriptLineBytes": 2000000
  },
  "attestation": {
    "mode": "local",
    "maxAgeSeconds": 3600,
    "maxFutureSkewSeconds": 60,
    "trustedPublicKeys": [],
    "githubOidc": { "mode": "off" }
  },
  "codeGraph": { "mode": "compatible" },
  "architecture": {
    "forbidCycles": true,
    "rules": [
      { "name": "domain-isolation", "from": "src/domain/**", "disallow": ["src/ui/**", "npm:express"] }
    ]
  }
}

Config is validated strictly; misspelled keys, invalid bounds and duplicate commands fail closed. Globs support *, **, ?; external imports use npm:. A command may declare an optional cwd (a project-root-relative directory) so it runs from a service subdirectory instead of the project root — useful for a polyglot monorepo where a language's tooling expects to run from its own package/module directory. config lint reports CONFIG_CWD_INVALID for a cwd that escapes the project root or does not exist, and checks that command's repository-relative path arguments (CONFIG_ORPHANED_PATH) against its own cwd instead of the project root. Evidence and report paths (tddReport/testReport/mutationReport, .musubix/config.json itself) are always resolved from the project root regardless of cwd; only the spawned process's own working directory changes. qualityProfile is custom by default. minimal preserves the core SDD gate, recommended also requires strict Code Graph, TDD and structured test identities, and release requires the complete formal, mutation, workflow, change, performance and CI-attestation checks. Stronger profiles reject missing or weakened settings rather than silently filling in evidence. approval.mode is required in newly initialized projects and requires current requirements, design, and release approvals. Existing schema-v1 configs that omit approval load in compatible mode. Approval files store the stage, approver, approvedAt, per-artifact SHA-256 values, and a deterministic manifest SHA-256. Run approval prepare before review and pass that exact hash to approval record; an intervening change is rejected and later changes become stale. Release recording recomputes the gate and requires every required non-approval check to pass rather than trusting cached quality evidence. approval prepare <stage> --diff-only additionally reports changedFiles: only the artifact paths whose content differs from that stage's last recorded approval (not a live git diff), computed from the same deterministic manifest used for artifactSha256 — the full hash and artifacts map are always still returned unchanged, so integrity verification is never weakened. diffOnlyBaseline is "approved" when compared against a prior recorded approval, or "none" (with every current path listed) when this stage has never been approved before. With --domain, the comparison is scoped to that domain's prior approval, matching normal domain-scoped approval semantics. If the last recorded approval file for that stage exists but fails schema validation, approval prepare --diff-only fails fast with the same Invalid <stage> approval evidence. error approval validate reports, rather than silently treating corrupted evidence as "no prior approval". Without --diff-only, non-JSON output is unchanged; with --diff-only and no --json, the console summary prints only the changed paths and a short baseline note instead of the full artifact listing. The approver string is explicit local evidence, not authenticated identity; repositories that require independent identity must also use protected review, CODEOWNERS, or CI/OIDC controls. Local approval evidence records explicit intent but does not cryptographically authenticate the approver; protect release authorization with repository review, CODEOWNERS/branch protection, or CI/OIDC attestation. approval prepare release's manifest excludes a fixed, source-hardcoded scratch-file convention from its untracked candidates: any untracked path at or beneath .musubix/scratch/, and any untracked file whose basename matches *.scratch.<ext> (for example debug.scratch.json). Write ad hoc inspection/debug output using this convention so it never changes the release-manifest hash on repeated inspection. This exclusion is a fixed, version-controlled list, not a configurable setting: it applies only to untracked candidates (a tracked file named this way is still included), and only as a final filter on the manifest's output, never before the structural nested-workspace/generated-directory exclusion scan, so it cannot be widened to quietly hide a real tracked change. Use tdd.redPreflightCommands to reference plain configured formatter commands; they must pass before Red captures the authoritative test fingerprint. Conventional .venv and venv Python environments containing a regular pyvenv.cfg, plus generated __pycache__/ directories, are excluded from source snapshots and Code Graph indexing. Standalone .pyc/.pyo files remain tracked because Python can execute source-less bytecode modules. Gradle .gradle/, Dart .dart_tool/, SwiftPM .build/, Zig .zig-cache//zig-out/, and .NET .dotnet/ CLI homes are excluded only when their parent contains the corresponding project manifest. Arbitrary same-named source directories remain tracked. codeGraph.mode defaults to compatible, where unresolved computed import()/require() calls remain warnings. Set it to strict to make those diagnostics gate-blocking errors. A trusted strict policy baseline prevents downgrading the project back to compatible mode. Coverage bounds are fractions [0,1] of mandatory requirements. Bare trace check --strict always requires full coverage; the aggregate gate uses config thresholds. The reported value is link coverage, not proof; with zero mandatory requirements it is null (not applicable). Add test-identities to requiredChecks to require every annotated TEST-* ID to be reported passed by a fresh structured report from a successful configured command. Add formal to enforce the configured solver and minimum modeled fraction. Every requirement containing explicit Formal: JSON automatically requires model-correspondence: its current formal constraint and generated trace must lead to at least one authoritative TEST-* that passed in a fresh structured command report. Missing, changed, unlinked, or stale evidence fails closed. .musubix/evidence/model-correspondence.json is only produced by evidence refresh; running model-correspondence validate before that file exists fails with MODEL_CORRESPONDENCE_MISSING, whose message and the command's own --help output both name evidence refresh as the fix. Add tdd to require complete Red-Green cycles. Red must be an observed nonzero test result; Green/Refactor must pass with the same configured command and unchanged test file. Test names/output must contain their TEST-* ID. language records project preference; skills follow input language. Machine diagnostic codes are stable English; human status labels include Japanese.

.musubix/policy-baseline.json records minimum required checks, coverage, architecture, formal policy, mutation and approval modes, workflow strict/session/freshness settings, CI-required attestation and strict OIDC identity/key binding, and required command names. A baseline that requires TDD Red preflights must also include their normalized commands definitions, which prevents replacing a trusted formatter while retaining its name. Weakening is rejected; changing the baseline in gate --changed requires independent approval. Protect the baseline with review/CODEOWNERS.

Only run trusted configuration: gates execute its commands with inherited environment, no shell interpretation and bounded time/output. Required command failure or skip blocks readiness regardless of requiredChecks.commands. A structured test command also fails if it reports no executed tests or any skipped, failed, or errored test, even when its process exits zero; this prevents integration suites from silently passing without their dependencies. Optional command failures are nonblocking unless a constitution rule rejects the measured count. No configured commands is skipped, not passed. Adapter test-ID declaration reference (each mechanism is distinct; see the dual-requirement note above for how this relates to the @id/@verifies doc comment):

| Adapter | ID-declaration mechanism | Worked example | | --- | --- | --- | | vitest / jest | ID as a substring anywhere in the test title | it('TEST-APP-001 rejects empty input', () => { ... }) | | pytest | ID in the underscore-form test function name | def test_TEST_APP_001(): ... | | go-test | ID as the trailing suffix of the test/subtest name | func TestTEST_APP_001(t *testing.T) { ... } | | cargo | ID as the trailing suffix of a Rust test identifier | fn test_app_001() { ... } | | junit | Exact @Tag("TEST-APP-001") plus an ID-bearing method name or @DisplayName | @Tag("TEST-APP-001") @Test void test() { ... } | | dotnet (xUnit) | ID inside [Fact(DisplayName = "...")] | [Fact(DisplayName = "TEST-APP-001 rejects empty input")] |

junit is the only adapter that requires a separate tag annotation in addition to an ID-bearing name/@DisplayName; every other adapter matches the ID directly from the test's own name/title string.

TDD commands require either command-specific tddArgs plus a tddReport, or a built-in vitest, jest, pytest, go-test, cargo, junit, or dotnet adapter. The junit adapter drives the Java JUnit Platform Console launcher, not arbitrary JUnit-XML producers; runners such as PHPUnit that only emit JUnit XML need explicit tddArgs/tddReport (or testReport) configuration instead. A custom musubix-json report is a single-line-safe JSON document {"schemaVersion":1,"tests":[{"id":"TEST-APP-001","status":"passed"}]} where status is passed, failed, skipped or error, and an optional "operations":{"counter":12} map carries deterministic performance counters. Explicit custom configuration takes precedence. Adapters derive targeted arguments and normalize native JSON/JSONL/XML into musubix-json. Vitest/Jest reports may contain unrelated skipped tests; targeted TDD selects only the requested ID. pytest requires the JSON-report plugin and underscore-form test names such as test_TEST_APP_001. Go uses a TEST-* subtest name, Cargo uses a Rust identifier such as test_app_001. JUnit methods should carry an exact @Tag("TEST-APP-001") and an ID-bearing method name or @DisplayName; the normalizer reads both testcase attributes and JUnit Platform display-name output. Surefire/Failsafe testcase elements are parsed whether they are self-closing or carry <system-out>/<system-err> children, so framework logging such as the Spring Boot banner does not hide passing identities. xUnit tests use [Fact(DisplayName = "TEST-APP-001 ...")] so TRX preserves the identity. Before each phase, musubix3 deletes the previous report, creates any required report parent directory, and requires a fresh musubix-json document containing exactly the selected test. Its status must be failed during Red and passed during Green/Refactor; skipped, error, missing and malformed reports fail. A non-test project input must change before Green. Identical phase output reused by different tests is rejected. Every phase is also appended to a SHA-256-linked immutable record chain. TDD and change checkpoints additionally share a persisted monotonic order ledger, which is authoritative for Red/Green boundaries; wall-clock timestamps are informational. Legacy chronology without order evidence fails with an explicit migration diagnostic. Missing, reordered, altered or orphaned records invalidate the evidence. Superseded cycles are never replaced by a newer recording: if any cycle lacks test-scoped provenance or a valid Red/Green, move .musubix/evidence/tdd.json aside and re-record every cycle from a clean Red baseline. There is deliberately no partial prune command, and hand-editing the evidence is unsupported.

Deterministic mutant identities come from musubix3 mutation identity <REQ-ID> <TEST-ID> <sourcePath> <operator> <line> <column>, which prints the exact MUT-* value a schema-v1 mutation report must declare. mutation validate reads .musubix/evidence/mutation.json; a configured mutationReport is converted into that file by the gate, so validating before a gate run reports absent rather than passing evidence.

CI executes isolated native contracts for Vitest, Jest, pytest with pytest-json-report, Go test, Cargo test, and the pinned JUnit Platform Console. The .NET adapter consumes standard TRX and is additionally validated by unit contracts and the C# application experiment. Each fixture contains an unrelated failing test, proving that the generated selector executes only the requested identity and that the real native report normalizes correctly. Jest is development-only; the Python, Go/Rust, and Java/JUnit tooling is provisioned only in CI and is not shipped as a package runtime dependency.

Change checkpoints fingerprint only implementation files linked to each changed requirement, plus their Code Graph dependencies. An unrelated source change cannot satisfy the implementation phase. The automatic change-completeness gate checks each CHANGE-ID for classified functional/non-functional requirements, measurable Acceptance criteria, concrete design responsibilities/interfaces/ constraints, an existing ADR, linked code, authoritative annotated tests, bounded TDD and trace edges. Each CHANGE document must contain a Requirements: line enumerating exactly the chronology's normative requirement IDs.

Structured test results may add "operations":{"visitedNodes":42}. A declared deterministic performance budget automatically requires the performance gate; elapsed time alone cannot satisfy it. For every observation, performance.json records a gate-generated run identity and SHA-256-linked provenance covering the configured command name, executable and rendered arguments, report path/source, fresh report bytes, test ID/status, counter/value, and process status/exit code. Validation re-reads persisted file, directory, and captured stdout reports and rejects missing or altered reports, record mutation, configuration drift, duplicate counter sources, non-passing tests, and results not produced by a successful configured command. The signed performance head hashes stable semantic fields while the JSON retains run/execution IDs, timestamps, report hashes, and chained provenance, so an equivalent gate can be rerun after signing without invalidating the signature. This provenance also gates CHANGE completeness and status freshness. Native runner reports do not expose application operation counters, so projects with such budgets must also configure an instrumented musubix-json report. Native adapters and custom reports can coexist; only the instrumented report that actually emits the named operation counter can prove the performance budget.

Mutation evidence uses a configured command with "mutationReport":{"format":"musubix-mutation-json","path":"..."}. Each fresh schema-v1 mutant record carries a deterministic MUT-<hash> identity (derivable with the mutationIdentity helper exported from musubix3/analysis), must-functional requirement ID, authoritative test ID, source/test paths and SHA-256 fingerprints, operator, one-based line/column, and killed|survived|skipped|error status. The gate adds command, rendered-argument, report, process, and exit provenance to mutation.json. Supplied evidence must cover every must functional requirement with a current linked killed mutant. Duplicate/conflicting, non-killed, stale, unlinked, altered-report, and configuration-drift evidence is rejected. mutation.mode defaults to compatible (absence is allowed); set it to strict and protect it plus the mutation command in the policy baseline for release. No mutation engine dependency is bundled. For Python, mutation doctor recommends removing existing __pycache__ directories before each run, then using python -B -m mutmut and python -B -m pytest to avoid new bytecode. Mutation and model-correspondence semantic heads are included in attestations and their underlying provenance is revalidated.

Quality evidence records required flags, actual exits/output, metrics, timestamps and input fingerprints. Changed-run paths, HEAD and impacts survive a later full gate. workflow-record stores a self-reported Skill/phase/status and optional command SHA-256 without storing command text. workflow-verify imports only Skill invocation metadata from a Copilot JSONL log and binds every completed declaration one-to-one, in order, to a distinct completed successful tool call. Use workflow-sanitize first when the source transcript contains messages, non-Skill tool arguments, or output that should not enter review evidence. Each Skill invocation must therefore record exactly one final workflow outcome; multi-phase chronology belongs in change-record, not duplicate workflow events. Incomplete, failed, reused, out-of-order and stale bindings fail. Set "workflow":{"mode":"strict"} or pass --strict to additionally require valid JSON on every nonempty line, valid event timestamps, consistent one-to-one tool start/completion lifecycles, and exactly one successful final terminal state. Supported terminal formats are result with exitCode: 0, or the current Copilot CLI lifecycle format with one session UUID and a final session.shutdown whose data.shutdownType is routine. The shutdown format also accepts one or more intermediate routine shutdowns when each is immediately followed by session.resume; workflow-sanitize retains those lifecycle boundaries. Mixed terminal formats, unmatched shutdown/resume transitions, multiple session identities, abnormal shutdowns, and trailing events fail closed. Compatible mode alone accepts more than one transcript path; supplied files are concatenated in ascending order of each file's earliest event timestamp (not command-line order), while each file's own internal event order is preserved untouched — this lets declarations whose invocation occurred in an earlier Copilot CLI session be reconciled by additionally supplying that session's transcript file. A toolCallId that appears as a tool start in more than one supplied file is rejected. Copilot CLI writes each session's JSONL transcript to ~/.copilot/session-state/<sessionId>/events.jsonl; this path is an internal detail of the installed Copilot CLI version, not a musubix3-owned contract, so confirm it against your installed version rather than assuming it is stable across releases. The terminal sessionId, exit code, event count, terminal timestamp, raw source hash and canonical transcript hash are persisted. workflow.expectedSessionId or --session-id rejects substitution with a different caller-declared session. Strict verification also bounds terminal transcript age and future clock skew with workflow.maxAgeSeconds and workflow.maxFutureSkewSeconds. Unrelated concurrent events—and even timestamps produced by different execution clocks—may be non-monotonic, so strict mode uses JSONL source order for causal tool/result lifecycles rather than imposing a timestamp sort. Pair/terminal clock skew is enforced only when maxEventSkewMs is explicitly supplied; terminal age and future-skew policies remain independently enforced. Verification remains streaming and resource-bounded: transcripts default to 100,000,000 bytes, while workflow.maxTranscriptBytes can explicitly raise the limit up to 1,000,000,000 bytes. Individual JSONL records default to 1,000,000 bytes and can be bounded up to 10,000,000 with workflow.maxTranscriptLineBytes. Protect both chosen bounds in the policy baseline so they cannot be widened silently. If project inputs change while a gate is running, input-stability reports each added, modified, or deleted path with before/after SHA-256 values. Standard Cargo/Maven target/, manifest-scoped .NET bin/ and obj/, and project-local .nuget/packages/ output are excluded, but source-like generated inputs remain fail-closed. Write command-generated reports under .musubix/evidence/native/ rather than the tracked source tree, otherwise a command that writes its own report during the gate invalidates input stability. Built-in adapters own their targeting and report arguments. A legacy leading Cargo/Go test subcommand is merged safely; conflicting report flags such as --json-report are rejected. Changes during a gate fail input stability; later source/config changes make status stale. Attestations older than maxAgeSeconds, or issued farther in the future than maxFutureSkewSeconds, fail. Local mode explicitly reports unsigned evidence.

ci-required supports two deliberately distinct trust models:

  • Static trusted-key mode (githubOidc.mode: "off"): keyId must select a configured Ed25519 public key. The signature covers repository, Git HEAD, CI provider/run ID, evidence heads, and the non-generated workspace snapshot.
  • GitHub OIDC strict mode: configure githubOidc.mode: "strict" and a custom audience base. The verifier discovers GitHub's issuer metadata and JWKS, verifies the RS256 JWT, and checks issuer, bound audience, exp/nbf/iat, repository, commit sha, and run_id, plus optional workflow and ref. keyBinding: "public-key" authorizes an attestation-carried ephemeral Ed2551