@raditia/craftkit
v1.56.0
Published
AI coding skills that auto-sync across Claude Code, Cursor, Gemini CLI, and Codex CLI
Maintainers
Readme
craftkit v1.56.0
One repo of AI coding skills that auto-syncs across Claude Code, Cursor, Gemini CLI, and Codex CLI. Pull once and every AI tool gets the same workflows, rules, and commands.
Table of contents
- Why bother? · token savings with RTK + Caveman + Ponytail
- Install
- How it works
- Using the workflows
- Just say what you want
- Enforcement gates · the hooks that stop an unrouted edit
- Agent dashboard · live view of running agents, opt-in
- Dynamic workflows ·
/parallel-review,/parallel-ship,/parallel-build - How the classifier picks agents
- Sequential fallback ·
/review,/ship,/build - Planning pipeline: /define ·
/interview→/spec→/test-cases→/plan - Experimental: /team-build · agent-teams build
- Cross-model: /cross-review · Claude and Codex review, then check each other
- Fix, tests, and PR message
- Grill, research, and handoff · stress-test plans, delegate reading, hand off sessions
- Scoring a run: /eval · weighted correctness %, judged and ledgered
- Skills reference
- Agents reference
- Architecture (EVPMR)
- Model routing
- Managing skills
- Tooling · RTK, Caveman, Ponytail, Karpathy Guidelines
- Design notes · why things work the way they do
- Changelog
Why bother?
AI coding sessions are expensive. Two things drain tokens fast: verbose shell output the AI has to read, and verbose AI responses you have to read. This repo ships two compression layers that cut both.
RTK: compresses what the AI reads (shell output)
Shell commands like git diff and jest dump noise before the signal. RTK filters it out before it reaches the AI.
── WITHOUT RTK (38 tokens) ──────────────────────────────────────────
On branch feature/checkout-flow
Your branch is ahead of 'origin/feature/checkout-flow' by 3 commits.
(use "git push" to publish your local commits)
Changes not staged for commit:
(use "git add <file>..." to update staging area)
modified: src/checkout/ViewCheckout.tsx
modified: src/checkout/PresenterCheckout.ts
Untracked files:
src/checkout/__tests__/ViewCheckout.test.tsx
── WITH RTK (6 tokens) ──────────────────────────────────────────────
M src/checkout/ViewCheckout.tsx
M src/checkout/PresenterCheckout.ts
? src/checkout/__tests__/ViewCheckout.test.tsx~84% reduction on a single call. Across a full session (git diff, tsc, jest, lint) it compounds to 60-90% savings on AI input tokens.
Caveman: compresses what you read (AI output)
The caveman plugin strips filler, hedging, and pleasantries from every response. Same findings, fewer words.
── WITHOUT CAVEMAN (~65 tokens) ─────────────────────────────────────
Sure! After carefully reviewing the code, I can see that there's
actually an issue in the ViewCheckout component. It looks like
there's a useState hook being used directly in the View layer,
which basically violates the EVPMR architecture pattern. You'll
want to move that state logic into the Presenter layer instead.
── WITH CAVEMAN (~18 tokens) ────────────────────────────────────────
[ERROR] ViewCheckout.tsx:14: useState in View layer.
Why: violates EVPMR.
Fix: move to PresenterCheckout.ts.~72% reduction per response. Full review sessions with reasoning and multi-step output: 40-60% output savings.
Ponytail: compresses what the AI generates (code output)
The ponytail decision ladder enforces YAGNI before any code is written. Before generating code, the AI stops at the first rung that holds: does this need to exist? is it in stdlib? is it a native feature? is an installed dep enough? can it be one line? Only then: minimal code. Deliberate shortcuts are marked with ponytail: comments naming their ceiling and upgrade path.
── WITHOUT PONYTAIL ─────────────────────────────────────────────────
// custom retry logic with exponential backoff + jitter
class RetryManager {
private attempts = 0;
async execute<T>(fn: () => Promise<T>, maxRetries = 3): Promise<T> { ... }
private calcDelay(attempt: number): number { ... }
}
── WITH PONYTAIL ────────────────────────────────────────────────────
// ponytail: no retry lib, inline for now. ceiling: >3 callers → extract.
const withRetry = (fn, n = 3) => fn().catch(e => n > 0 ? withRetry(fn, n-1) : Promise.reject(e));The six-tag ponytail rubric (delete: stdlib: native: yagni: shrink: narrate:) lives in karpathy-guidelines, so code is written against the same list review scores it by. Every code-writing turn checks its own diff before reporting done, and findings are applied as deletions at the named file:line, never as rewrites.
80-94% code reduction on over-engineered solutions. Pairs with /ponytail-review (a diff), /ponytail-audit (the whole repo) and /ponytail-debt (deferred shortcuts and removable flags).
flag-safety applies the same marker idea to feature flags: with the flag OFF, behavior must match the pre-change code across code paths, persisted state, API contracts and analytics. Each flag branch carries a flag: comment with the key, the OFF behavior and when to remove it. /build and /parallel-build check it while writing; the review agents and /parallel-ship check it again.
Combined impact
| Layer | Compresses | Typical savings | |-------|------------|-----------------| | RTK | Shell output → AI input | 60-90% on dev operations | | Caveman | AI output → your reading | 40-60% on prose responses | | Ponytail | Code generated | 80-94% on over-engineered solutions | | Together | All directions | 50-80% total session cost |
Typical feature review session without compression: ~40,000 tokens. With RTK + Caveman + Ponytail: ~8,000-20,000 tokens.
Install
Option A, npm (version pinning + rollback):
npm install -g @raditia/craftkitPin a version or roll back:
npm install -g @raditia/[email protected]Option B, git (auto-update on git pull):
git clone [email protected]:raditia/craftkit.git ~/craftkit
cd ~/craftkit
bash install.shinstall.sh wires up the post-merge and post-rewrite hooks and runs the first sync. Merge and rebase pulls then add, update, and remove installed skills automatically. Existing git installations should rerun bash install.sh once to add the rebase hook.
Requirements: bash 3.2+, curl. macOS ships bash 3.2 by default.
Upgrading from ≤ v1.23.0: the next sync uninstalls the retired Copilot and Crush integrations automatically. Copilot @ agents it wrote into your other repos may be committed there, so sync prints those paths and leaves them for you.
Contributing to craftkit itself: see CONTRIBUTING.md. Short version:
there is no application build (the product is instructions); behavioral fixtures run through check.sh, so check.sh is the gate and a second
consecutive sync.sh must report no work.
bash check.sh # content integrity, exit 0 required before commit
bash sync.sh # distribute; a second consecutive run must report no workHow it works
Every git pull triggers a sync that installs rules, skills, commands, and agents into each AI tool:
---
config:
theme: base
look: classic
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
themeVariables:
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
fontSize: 16px
primaryColor: "#FFFFFF"
primaryBorderColor: "#C1C4C6"
primaryTextColor: "#242628"
lineColor: "#A2A6A8"
clusterBkg: "#F5FBFF"
clusterBorder: "#F0F1F2"
titleColor: "#707577"
edgeLabelBackground: "#FFFFFF"
flowchart:
curve: basis
wrappingWidth: 240
nodeSpacing: 30
rankSpacing: 40
---
flowchart TB
subgraph S1["STAGE 1 · CONTENT"]
c1("rules/")
c2("skills/")
c3("commands/")
c4("agents/")
c5("partials/<br/>spliced in, never alone")
end
subgraph S2["STAGE 2 · DISTRIBUTE"]
pull("git pull<br/>post-merge / post-rewrite hooks") --> sync("sync.sh<br/>idempotent · check.sh gates every change")
end
subgraph S3["STAGE 3 · TOOLS · one adapter each"]
t1("Claude Code")
t2("Cursor")
t3("Gemini CLI")
t4("Codex CLI")
end
subgraph S4["STAGE 4 · SESSION GATES"]
gates("Claude Code: routing and enforcement hooks<br/>Codex: rule loader + routing + verify-on-stop hooks<br/>Cursor/Gemini: advisory rules")
end
subgraph S5["STAGE 5 · OUTPUT"]
out("Verified change, shipped")
end
c1 & c2 & c3 & c4 & c5 --> pull
sync --> t1 & t2 & t3 & t4
t1 & t2 & t3 & t4 --> gates
gates --> out
classDef n fill:#FFFFFF,stroke:#C1C4C6,color:#242628
classDef key fill:#D1F0FF,stroke:#0A9AF2,color:#242628
classDef ok fill:#FFFFFF,stroke:#029D24,color:#029D24
class c1,c2,c3,c4,c5,pull,t1,t2,t3,t4 n
class sync,gates key
class out ok
style S1 fill:transparent,stroke:transparent
style S2 fill:transparent,stroke:transparent
style S3 fill:transparent,stroke:transparent
style S4 fill:transparent,stroke:transparent
style S5 fill:transparent,stroke:transparentFive namespaces, one source of truth:
| Directory | Loaded | Invoked |
|-----------|--------|---------|
| rules/ | Every session, automatically | Never, since they are always present |
| skills/ | On demand | Slash command or natural language |
| commands/ | On demand | Slash command or natural language |
| agents/ | Spawned by an orchestrator | native spawning with the named profile (Claude and Codex) |
| partials/ | Only as a splice into a skill, command or agent | Never, since it ships inside its host file (every tool for a skill or command, Claude and Codex for an agent) |
At runtime: Gateway, Orchestrators, state
sync.sh is the Distributor: it runs at install time and never while you work. At runtime three
layers act inside each AI tool:
| Role | What | Where |
|------|------|-------|
| CraftKit Gateway | Claude uses the routing, loader, guard, and exit hooks. Codex loads applicable rule bodies at SessionStart, injects routing guidance on each prompt, and checks verification at Stop. | hooks/; Claude and Codex use native hooks. Cursor and Gemini get advisory text. Codex and Cursor discover skills from ~/.agents/skills/ |
| Orchestrators | Run a workflow: resolve the feature once in Phase 0, pass the slug down, spawn skills and agents | commands/*.md |
| Skills and agents | Do one job; skills reach Figma and Lark through the host's MCP client | skills/, agents/ (Claude and Codex) |
The Gateway does not see MCP calls. Per-feature state lives in the repo, never in a tool:
---
config:
theme: base
look: classic
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
themeVariables:
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
fontSize: 16px
primaryColor: "#FFFFFF"
primaryBorderColor: "#C1C4C6"
primaryTextColor: "#242628"
lineColor: "#A2A6A8"
clusterBkg: "#F5FBFF"
clusterBorder: "#F0F1F2"
titleColor: "#707577"
edgeLabelBackground: "#FFFFFF"
flowchart:
curve: basis
wrappingWidth: 240
nodeSpacing: 30
rankSpacing: 40
---
flowchart TB
subgraph R1["INSTALL TIME"]
D("Distributor<br/>sync.sh + adapters/")
end
subgraph GW["GATEWAY · hooks/ · Claude Code and Codex"]
direction LR
L("Loader<br/>SessionStart") ~~~ R("Router<br/>UserPromptSubmit") ~~~ G("Guards<br/>PreToolUse") ~~~ X("Exit gates<br/>Stop")
end
subgraph R3["ORCHESTRATORS · commands/*.md"]
O("/define · /parallel-build<br/>/build · /team-build<br/>/parallel-review · /parallel-ship<br/>/fix · /ship<br/>Phase 0: resolve slug once<br/>+ approved test cases, pass down")
end
subgraph R4["WORKERS"]
direction LR
S("Skills · skills/*<br/>/spec · /test-cases · /plan<br/>/fe-test · /eval") ~~~ A("Agents · agents/*.md<br/>cold reviewers<br/>Claude and Codex") ~~~ M("MCP via the host's client<br/>Figma · Lark")
end
subgraph R5["STATE · in the repo, per feature"]
direction LR
ST1("docs/planning/<slug>.md<br/>intent + sources: pointers") ~~~ ST2("docs/planning/<slug>.tests.md<br/>test cases · repo is master") ~~~ V("Published view<br/>Excel export")
end
R1 -->|sync| GW
GW -->|routes each prompt| R3
R3 --> R4
R4 --> R5
classDef n fill:#FFFFFF,stroke:#C1C4C6,color:#242628
classDef key fill:#D1F0FF,stroke:#0A9AF2,color:#242628
classDef ok fill:#FFFFFF,stroke:#029D24,color:#029D24
class D,O,S,A,M,V n
class R,L,G,X key
class ST1,ST2 ok
style R1 fill:transparent,stroke:transparent
style GW fill:#D1F0FF,stroke:#C1C4C6
style R3 fill:transparent,stroke:transparent
style R4 fill:transparent,stroke:transparent
style R5 fill:transparent,stroke:transparentWhere files land per AI tool
| Tool | Always-on (rules/) | On-demand (skills/ + commands/) | Agents (agents/) |
|------|----------------------|--------------------------------------|--------------------|
| Claude Code | ~/.claude/CLAUDE.md (managed block) | ~/.claude/commands/<name>.md → /<name> | ~/.claude/agents/<name>.md |
| Cursor | ~/.cursor/rules/*.mdc (alwaysApply) | ~/.agents/skills/<name>/SKILL.md (shared native skills, local only) | n/a |
| Gemini CLI | ~/GEMINI.md (managed block) | ~/GEMINI.md (managed block), and also lists the shared ~/.agents/skills/ | n/a |
| Codex CLI | ~/.codex/AGENTS.md (short managed block); ~/.codex/hooks.json loads applicable full rules from ~/.craftkit/codex/rules/ at session start | ~/.agents/skills/<name>/SKILL.md (shared native skills, including workflows) | ${CODEX_HOME:-~/.codex}/agents/<name>.toml |
Codex and Cursor read skills from ~/.agents/skills/, so they share one install there. Cursor does not copy that folder to Cloud Agents, so CraftKit skills reach local Cursor sessions only. Gemini CLI reads it too, but keeps its full ~/GEMINI.md block: as native skills, workflows would load only on demand, behind a consent prompt on every activation. CraftKit's named specialists install for both Claude and Codex. Codex profiles carry the live injected instructions, use a read-only sandbox, and inherit the configured model with medium reasoning effort. Parallel workflows use native spawning up to the available concurrency, queue excess workers, and collect every result. Sequential fallback applies when spawning is unavailable. Unowned or symlinked Codex profiles are preserved and reported as collisions.
If a skill name already belongs to another install in ~/.agents/skills/, sync leaves that directory untouched, warns with its path, and continues installing the other skills. Remove or rename the conflicting directory if you want CraftKit's version of that skill.
Retired: GitHub Copilot and Crush (supported through v1.23.0). Neither has a headless entry point, so neither can join cross-tool agent fan-out (more).
Using the workflows
Just say what you want
Natural language routes to the right command automatically. No slash commands required.
"plan this feature" → /define (interview → spec → test-cases → plan, checkpoint-gated)
"review this" → /parallel-review
"build this feature" → /parallel-build
"ship this" → /parallel-ship
"fix this bug" → /fix
"write tests for this" → /fe-test · /android-test · /ios-test (by platform)
"generate PR message" → /pr-message
"poke holes in my plan" → /grill (also: "grill this", "stress-test my design")
"research X for me" → /research (background agent, primary sources)
"hand this session off" → /handoff (also: "summarize for the next agent")
"connect our docs repo" → /context-source (also: "which docs repos are connected", "disconnect the old docs repo")Platform is not inferred. On every prompt hooks/craftkit-routing.js resolves it from cwd and injects the answer:
---
config:
theme: base
look: classic
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
themeVariables:
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
fontSize: 16px
primaryColor: "#FFFFFF"
primaryBorderColor: "#C1C4C6"
primaryTextColor: "#242628"
lineColor: "#A2A6A8"
clusterBkg: "#F5FBFF"
clusterBorder: "#F0F1F2"
titleColor: "#707577"
edgeLabelBackground: "#FFFFFF"
flowchart:
curve: basis
wrappingWidth: 240
nodeSpacing: 30
rankSpacing: 40
---
flowchart TD
S["every prompt"] --> W["walk up from cwd"]
W --> C{"marker at this level?"}
C -->|"none"| U["parent directory"]
U --> W
C -->|"settings.gradle"| A["Android<br/>MVP"]
C -->|"Podfile · Package.swift<br/>*.xcodeproj"| I["iOS<br/>MVVM-C"]
C -->|"package.json + React dependency"| R["RN / web<br/>EVPMR"]
C -->|"other package.json"| N["Node / tooling<br/>project conventions"]
C -->|"two or more<br/>at one level"| M["mixed<br/>union both<br/>agent sets"]
A --> INJ["inject platform into the prompt"]
I --> INJ
R --> INJ
N --> INJ
M --> INJNearest ancestor wins, so "write tests for this" in an Android repo resolves to /android-test, never /fe-test.
Enforcement gates: hooks that refuse
Routing context is only text: an agent can read it, announce the right skill, and hand-roll the work anyway. These hooks close that gap. The four gates can stop a call; the other four only inject context, rewrite a command, or notify.
Codex installs craftkit-codex.js into ~/.codex/hooks/ and registers SessionStart, UserPromptSubmit, PreToolUse and PostToolUse on Bash, and Stop in ~/.codex/hooks.json. It loads applicable rule bodies, using a concise Codex routing section instead of the full Claude-oriented rule. An explicit $skill or leading /command request points to the native skill file without injecting a second body. Verification requires a completed successful command against an unchanged snapshot from start to finish and the latest edited state, including committed edits; parallel hook results append without overwriting peers. Unknown or unfinished results do not count. Direct foreground commands and && chains are supported; commands that mask failures with pipelines or semicolons require a separate check invocation. The gate covers check.sh, Node type/lint checks, Gradle lint/tests, and iOS tests/SwiftLint. Codex requires a one-time /hooks review and trust of the new definitions before they run. Native skill activation is not exposed as a stable hook event, so the gateway cannot prove that a skill body was followed; verification is the enforced part. CRAFTKIT_GATE=off disables it.
| Hook | Event | What it does |
|------|-------|--------------|
| craftkit-routing.js | UserPromptSubmit | Injects the routing table, platform, model tiers, and any locally installed skills. Advisory |
| gate-skill-first.js | PreToolUse on Edit\|Write\|MultiEdit\|NotebookEdit | Asks before a source edit in a session that never invoked a skill, naming the skills that fit the file. Once per turn |
| gate-verify-on-stop.js | Stop | Blocks a turn that edited source but ran no verification command, and names the command (check.sh if present, else typecheck + lint) |
| gate-announce-honored.js | Stop | Blocks a reply that says Running /<skill> with no Skill call, or carries no routing declaration at all |
| gate-read-size.js | PreToolUse on Read | Refuses a whole-file read over 800 lines and points to the bulk-read agent or an offset/limit read. Denies rather than asks |
| craftkit-platform-rules.js | SessionStart | Loads platform:-scoped rules only where the cwd matches, so EVPMR laws stay out of Kotlin and Swift sessions |
| craftkit-update-check.js | SessionStart | Tells you when a newer craftkit is on npm. Asks the registry at most once a day (1.5s timeout, cached in ~/.craftkit-state/update-check) and stays silent on any failure |
| craftkit-read-cap.js | PreToolUse on Bash | Rewrites a bare cat F / rtk read F over 800 lines to rtk read -m 800. Leaves piped commands alone |
Shared helpers (not registered as hooks): craftkit-transcript.js finds the current turn for every gate, craftkit-platform.js detects the platform, craftkit-filesize.js holds the 800-line threshold, and craftkit-drift.js answers "has this changed since baseline" as clean, drifted, or cannot-verify.
How the gates behave:
- Fail open. Unreadable transcript, bad stdin, or no gate command in the project: the call passes.
ask, notdeny, so you keep the override. The Read gate is the exception, because anaskis silently approved under auto-accept.- Subagents are checked at the parent. Their edits pass inside the agent; the parent's Stop gate picks them up from
git status, the same way it catches edits made withsed -ior a heredoc. - Skill gate is per session, Stop gates are per turn. A turn continuing already-routed work ("apply the fixes") isn't asked again.
- A downgrade refuses to sync.
sync.shwon't run from a checkout older than the installed version. - Plugin skills (
<plugin>:<name>) aren't enumerated by the announce gate and pass.
| Escape hatch | Turns off |
|--------------|-----------|
| CRAFTKIT_GATE=off | The three edit/Stop gates and platform-rules injection |
| CRAFTKIT_READ_GATE=off | gate-read-size.js |
| CRAFTKIT_READ_CAP=off | craftkit-read-cap.js |
| CRAFTKIT_ALLOW_DOWNGRADE=1 | The sync downgrade guard |
| CRAFTKIT_UPDATE_CHECK=off | craftkit-update-check.js |
Removing a Claude hook from _CRAFTKIT_HOOKS uninstalls it on the next sync. The reasoning behind each behavior above, with the measurements: design notes.
Agent dashboard (opt-in)
A live terminal view of what your agents are doing: the main session (model, effort, context, cost), a box for each running subagent that appears when it starts and disappears when it finishes (model, tokens used, elapsed time, tool calls, current action), a total of the tokens all its subagents used, and a session log. Model and tokens come from Claude Code's own subagent transcripts; Codex boxes show the model its events carry, and no tokens, because its transcript format is not a stable interface. It covers Claude Code and Codex sessions, several at once: a strip at the top lists every session active in the last 30 minutes and not ended (tool, project folder name, a short session id so two sessions in one project differ, running subagents, last event), numbered in the order they started. ←/→ or 1-9 switch and pin the view, a goes back to following the newest, and ccdash <n> opens a window pinned to session n: the number is turned into that session's id before the window opens, so it keeps watching that session when the strip renumbers. Inside tmux a pinned view opens as a side-by-side pane; on macOS each is its own window. A session with no tool call yet, or none in 30 minutes, is not listed.
Example, from sample data (two sessions, the second pinned, three subagents running and one finished):
CLAUDE CODE AGENT TREE · agentic-skills · a91c4f · live · pinned
1 ○ Codex booking-web b7e210 idle live
▸ 2 ● Claude agentic-skills a91c4f 3 running live
←/→ or 1-9 switch · a follow newest
╭──────────────────────────────────────────────────────────╮
│ Opus 5.5 · main session │
│ effort high ctx ███░░░░░░░ 38% $2.41 │
│ last: Read scripts/dashboard.py │
╰──────────────────────────────────────────────────────────╯
│
├─ 3 subagent(s) running · 1 done · subagents used 169.1k tokens
╭────────────────────────────╮ ╭────────────────────────────╮ ╭────────────────────────────╮
│ code-quality │ │ adversarial │ │ ponytail-review │
│ ◐ running 2m14s │ │ ◐ running 1m58s │ │ ◐ running 1m31s │
│ sonnet-5-5 │ │ opus-5-5 │ │ sonnet-5-5 │
│ 41.8k tok · 1.6k out │ │ 98.7k tok · 2.4k out │ │ 19.1k tok · 700 out │
│ 1 tool call(s) │ │ 1 tool call(s) │ │ 1 tool call(s) │
│ Read sync.sh │ │ Read ccdash │ │ Grep agent_usage │
╰────────────────────────────╯ ╰────────────────────────────╯ ╰────────────────────────────╯
── session log ───────────────────────────────────────────────────────────────────────────
22:40:02 main Read scripts/dashboard.py
22:40:09 code-quality started
22:40:14 code-quality Read sync.sh
22:40:21 ponytail-revi… started
22:40:30 ponytail-revi… Grep agent_usage
22:40:36 adversarial started
22:40:44 adversarial Read ccdash
22:41:03 bulk-read started
22:41:40 bulk-read finished
9 events · esc / q to closeCRAFTKIT_DASHBOARD=1 bash sync.sh # turn on; the choice persists across later syncs
ccdash # open it, following the newest session; esc or q closes it
ccdash 2 # open pinned to session 2 (one window per session: ccdash 1, ccdash 2, ...)
CRAFTKIT_DASHBOARD=0 bash sync.sh # turn off; removes everything it installedInstalled with npm, set the same variable on the install; later upgrades keep the choice:
CRAFTKIT_DASHBOARD=1 npm install -g @raditia/craftkit # turn on
CRAFTKIT_DASHBOARD=0 npm install -g @raditia/craftkit # turn offThen start a new Claude session so the hooks load, and in Codex run /hooks once to trust the new one.
It is off by default because the logger writes the first 60 characters of every tool call's file path or command to ~/.craftkit/agent-tree/events/. Those files are private to you (0600), common credential shapes (auth headers, --password x, token=, *_SECRET_*=, -u user:pass, https://user:pw@, sk-/ghp_/xoxb-/AKIA tokens) are masked before writing (shape-based, so it narrows exposure rather than guaranteeing none), the logger deletes files older than 7 days, and turning the dashboard off deletes the directory. Off installs nothing. CRAFTKIT_DASHBOARD also takes on/off, true/false, yes/no; anything else is warned about and ignored.
| Piece | Where | What it does |
|-------|-------|--------------|
| craftkit-agent-log.js | Claude and Codex, on SubagentStart, SubagentStop, PostToolUse, SessionEnd | Appends one line per event to the session's log and keeps a file per running subagent, deleted when it stops and cleared when the session ends. Fails open. Codex needs a one-time /hooks trust |
| craftkit-statusline.js | Claude statusLine | Shows model, effort, context, cost and running subagents, and saves them for the dashboard. If you already have a status line it is kept: saved to ~/.craftkit-state/statusline.json, run first with the same input, and this is appended after it. Turning the dashboard off puts yours back exactly |
| scripts/dashboard.py, scripts/ccdash | ~/.craftkit/bin, linked into ~/.local/bin when that exists | The viewer and its launcher. ccdash opens a tmux popup inside tmux (a side pane before tmux 3.2), a separate iTerm or Terminal window when run from a chat (! ccdash, macOS), or full screen in a plain terminal. Without a terminal it prints one snapshot |
Cost, measured on a 20,000-event log: the logger and status line each finish in under 0.1s, but they are a node start per tool call and per status refresh (5s) while the dashboard is on. The open dashboard idles near 0% CPU and ~20 MB, because it keeps one state per session and reads only the lines added since its last frame. A subagent quiet for 30 minutes (one long build or test run looks the same) is labelled quiet rather than dropped. One killed before its SubagentStop fires is cleared when its session ends, or hidden after a day if the session was killed too. Turning the dashboard off while a session is open makes that session report a missing hook on each tool call until it restarts.
Dynamic workflows (default)
Build, review, and ship use dynamic parallel execution: a classifier detects the platform (RN/web, Android, iOS), reads your actual diff, selects only the agents that matter, and runs them concurrently. Test-only diffs skip deep review entirely. Every command below works on all three platforms; only the gates and the agent set change.
Each workflow is drawn in five lanes (You, Hooks, Main agent, Sub-agents, Result). Box styles:
---
config:
theme: base
look: classic
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
themeVariables:
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
fontSize: 16px
primaryColor: "#FFFFFF"
primaryBorderColor: "#C1C4C6"
primaryTextColor: "#242628"
lineColor: "#A2A6A8"
clusterBkg: "#F5FBFF"
clusterBorder: "#F0F1F2"
titleColor: "#707577"
edgeLabelBackground: "#FFFFFF"
flowchart:
curve: basis
wrappingWidth: 240
nodeSpacing: 30
rankSpacing: 40
---
flowchart LR
y("You<br/>your prompt") ~~~ h("Hook<br/>runs automatically") ~~~ m("Skill<br/>on the main agent") ~~~ a("Sub-agent<br/>read-only, parallel") ~~~ v("Verdict")
classDef you fill:#FFFFFF,stroke:#707577,color:#242628
classDef hook fill:#D1F0FF,stroke:#0A9AF2,color:#242628,stroke-dasharray:4 3
classDef main fill:#FFFFFF,stroke:#0A9AF2,color:#242628
classDef sub fill:#FFFFFF,stroke:#029D24,color:#242628
classDef verdict fill:#0A5C2C,stroke:#0A5C2C,color:#8BE200
class y you
class h hook
class m main
class a sub
class v verdict/parallel-review
Triggered by:
"review this"/"help me review"/"code review"/"LGTM check"
---
config:
theme: base
look: classic
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
themeVariables:
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
fontSize: 18px
primaryColor: "#FFFFFF"
primaryBorderColor: "#C1C4C6"
primaryTextColor: "#242628"
lineColor: "#A2A6A8"
clusterBkg: "#F5FBFF"
clusterBorder: "#F0F1F2"
titleColor: "#707577"
edgeLabelBackground: "#FFFFFF"
flowchart:
wrappingWidth: 260
---
swimlane-beta LR
subgraph you["YOU"]
ask("review my changes")
end
subgraph hooks["HOOKS"]
route("routing hook<br/>platform detected")
end
subgraph main["MAIN AGENT"]
gates("fast gates<br/>classify diff")
tst("tests<br/>in background")
syn("synthesis<br/>dedupe · rank")
end
subgraph subs["SUB-AGENTS"]
rev("reviewers in parallel<br/>read-only")
end
subgraph result["RESULT"]
v("READY TO MERGE<br/>BLOCKED · INCOMPLETE")
end
ask --> route --> gates
gates --> rev
gates --> tst
rev --> syn
tst --> syn
syn --> v
gates -.->|gate fails| v
classDef you fill:#FFFFFF,stroke:#707577,color:#242628
classDef hook fill:#D1F0FF,stroke:#0A9AF2,color:#242628,stroke-dasharray:4 3
classDef main fill:#FFFFFF,stroke:#0A9AF2,color:#242628
classDef sub fill:#FFFFFF,stroke:#029D24,color:#242628
classDef verdict fill:#0A5C2C,stroke:#0A5C2C,color:#8BE200
class ask you
class route hook
class gates,tst,syn main
class rev sub
class v verdict/parallel-ship
Triggered by:
"ship this"/"prepare for PR"/"is this ready?"/"get this ready to merge"
---
config:
theme: base
look: classic
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
themeVariables:
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
fontSize: 18px
primaryColor: "#FFFFFF"
primaryBorderColor: "#C1C4C6"
primaryTextColor: "#242628"
lineColor: "#A2A6A8"
clusterBkg: "#F5FBFF"
clusterBorder: "#F0F1F2"
titleColor: "#707577"
edgeLabelBackground: "#FFFFFF"
flowchart:
wrappingWidth: 260
---
swimlane-beta LR
subgraph you["YOU"]
ask("ship this")
end
subgraph hooks["HOOKS"]
route("routing hook<br/>platform detected")
end
subgraph main["MAIN AGENT"]
gates("fast gates<br/>classify diff")
tst("tests + coverage<br/>RN/web ≥ 93%")
syn("synthesis<br/>test-case trace")
end
subgraph subs["SUB-AGENTS"]
rev("reviewers + ponytail<br/>perf · a11y · adversarial")
end
subgraph result["RESULT"]
v("READY TO MERGE<br/>BLOCKED · INCOMPLETE")
end
ask --> route --> gates
gates --> rev
gates --> tst
rev --> syn
tst --> syn
syn --> v
gates -.->|gate fails| v
classDef you fill:#FFFFFF,stroke:#707577,color:#242628
classDef hook fill:#D1F0FF,stroke:#0A9AF2,color:#242628,stroke-dasharray:4 3
classDef main fill:#FFFFFF,stroke:#0A9AF2,color:#242628
classDef sub fill:#FFFFFF,stroke:#029D24,color:#242628
classDef verdict fill:#0A5C2C,stroke:#0A5C2C,color:#8BE200
class ask you
class route hook
class gates,tst,syn main
class rev sub
class v verdict/parallel-build
Triggered by:
"build feature X"/"implement X"/"create a new screen"
---
config:
theme: base
look: classic
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
themeVariables:
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
fontSize: 18px
primaryColor: "#FFFFFF"
primaryBorderColor: "#C1C4C6"
primaryTextColor: "#242628"
lineColor: "#A2A6A8"
clusterBkg: "#F5FBFF"
clusterBorder: "#F0F1F2"
titleColor: "#707577"
edgeLabelBackground: "#FFFFFF"
flowchart:
wrappingWidth: 260
---
swimlane-beta LR
subgraph you["YOU"]
ask("build feature X")
end
subgraph hooks["HOOKS"]
route("routing hook<br/>platform detected")
end
subgraph main["MAIN AGENT"]
impl("context · scaffold<br/>implement · gates")
tst("write tests<br/>while agents run")
syn("synthesis<br/>consensus · unique")
end
subgraph subs["SUB-AGENTS"]
rev("reviewers in parallel<br/>picked by classifier")
end
subgraph result["RESULT"]
done("DONE · BLOCKED<br/>INCOMPLETE")
end
ask --> route --> impl
impl --> rev
impl --> tst
rev --> syn
tst --> syn
syn --> done
impl -.->|gate fails| done
classDef you fill:#FFFFFF,stroke:#707577,color:#242628
classDef hook fill:#D1F0FF,stroke:#0A9AF2,color:#242628,stroke-dasharray:4 3
classDef main fill:#FFFFFF,stroke:#0A9AF2,color:#242628
classDef sub fill:#FFFFFF,stroke:#029D24,color:#242628
classDef verdict fill:#0A5C2C,stroke:#0A5C2C,color:#8BE200
class ask you
class route hook
class impl,tst,syn main
class rev sub
class done verdictHow findings become one verdict
Every review agent reads the same files in the same message and shares nothing, so agreement between them is evidence:
---
config:
theme: base
look: classic
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
themeVariables:
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
fontSize: 16px
primaryColor: "#FFFFFF"
primaryBorderColor: "#C1C4C6"
primaryTextColor: "#242628"
lineColor: "#A2A6A8"
clusterBkg: "#F5FBFF"
clusterBorder: "#F0F1F2"
titleColor: "#707577"
edgeLabelBackground: "#FFFFFF"
flowchart:
curve: basis
wrappingWidth: 240
nodeSpacing: 30
rankSpacing: 40
---
flowchart LR
h1["1 · INDEPENDENT REVIEW<br/>same files · one message<br/>no shared state"] ~~~ h2["2 · MERGE"] ~~~ h3["3 · RANK BY AGREEMENT"] ~~~ h4["4 · VERDICT"]
a1("code-quality")
a2("platform review")
a3("platform a11y")
a4("ponytail · performance")
adv("adversarial<br/>argues against<br/>shipping")
dd("Deduplicate<br/>by file:line<br/>skipped agent →<br/>coverage-gap<br/>warning")
k1("CONSENSUS<br/>2+ agents<br/>independently · fix first")
k2("Standard<br/>one agent<br/>normal confidence")
k3("UNIQUE<br/>uncorroborated<br/>kept, lower confidence")
k4("Contradiction<br/>state both<br/>judge by evidence<br/>never average")
k5("BLIND SPOTS<br/>what the whole<br/>panel missed")
v1("READY TO MERGE<br/>no errors, gates pass")
v2("BLOCKED (list)<br/>an error, a failed gate<br/>or a missing test case")
v3("INCOMPLETE<br/>an agent failed to run<br/>never ready")
note["UNVERIFIED claims<br/>cannot back an ERROR,<br/>so an unproven finding<br/>cannot block a merge"]
a1 & a2 & a3 & a4 --> dd
dd --> k1 & k2 & k3 & k4
adv -.-> k5
k1 & k2 & k3 & k4 & k5 --> v2
k1 ~~~ v1
k4 ~~~ v3
k5 ~~~ note
classDef n fill:#FFFFFF,stroke:#C1C4C6,color:#242628
classDef key fill:#D1F0FF,stroke:#0A9AF2,color:#242628
classDef ok fill:#FFFFFF,stroke:#029D24,color:#029D24
classDef done fill:#0A5C2C,stroke:#0A5C2C,color:#8BE200
classDef warn fill:#FEF5FC,stroke:#FA9EB4,color:#8B1842
classDef warnline fill:#FFFFFF,stroke:#FA9EB4,color:#242628
classDef info fill:#D1F0FF,stroke:#D1F0FF,color:#024590
classDef dash fill:#FFFFFF,stroke:#C1C4C6,color:#242628,stroke-dasharray:4 3
classDef quiet fill:transparent,stroke:transparent,color:#707577
class a1,a2,a3,a4 ok
class adv,k5 warnline
class dd key
class k1,v1 done
class k2 n
class k3 dash
class k4,v2 warn
class v3 info
class note,h1,h2,h3,h4 quietHow the classifier picks agents
The classifier reads your actual changed files, not just filenames, and selects only the agents that apply. Irrelevant agents are skipped entirely.
RN / web (EVPMR) agents selected:
──────────────────────────────────────────────────────────
View*.tsx → code-quality + fe-review + fe-a11y
Presenter*.ts → code-quality + fe-review
Model*.ts → code-quality (type/correctness focus)
Entry*.tsx or Resource*.ts → fe-review
View or Presenter + /parallel-ship → + fe-performance
Android (MVP)
──────────────────────────────────────────────────────────
*Activity/Fragment/Widget.kt, layout → code-quality + android-review + android-a11y
*Presenter.kt, *ViewModel.kt → code-quality + android-review
*Repository/Interactor/UseCase.kt → code-quality
Dagger *Module/*Component.kt → android-review
Presenter/VM/adapter + /parallel-ship → + android-performance
iOS (MVVM-C)
──────────────────────────────────────────────────────────
*ViewController/View/Cell.swift → code-quality + ios-review + ios-a11y
*ViewModel.swift → code-quality + ios-review
*Fetcher.swift → code-quality + ios-performance
*Contract/Factory/Coordinator.swift → ios-review
All platforms
──────────────────────────────────────────────────────────
any non-test src + build/ship → + ponytail-review (over-engineering)
3+ architecture layers changed → + adversarial (devil's advocate)
auth / payment / credential paths → code-quality (security emphasis)
intent file resolves under docs/planning/ → code-quality (spec conformance: diff vs planned acceptance criteria)
test files only → agents skipped, gates onlyFive example diffs and what the classifier picks for /parallel-review:
| Example diff | Agents picked | Why |
|---|---|---|
| ViewCheckout.tsx + PresenterCheckout.ts | code-quality · fe-review · fe-a11y | a View changed, so a11y joins |
| ModelCheckout.ts only | code-quality (type-safety focus) | targeted findings, no EVPMR or a11y noise |
| __tests__/ViewCheckout.test.tsx only | none: gates only (tsc + lint + test) | tests only, so Phase 2 is skipped and costs no agents |
| Entry + View + Presenter + Model | code-quality · fe-review · fe-a11y · adversarial | 3+ layers changed, so adversarial argues against merging |
| CheckoutFragment.kt + CheckoutPresenter.kt (Android) | code-quality · android-review · android-a11y · android-performance | same command, native agents; gates are gradlew lint + testGeneralDebugUnitTest |
Sequential fallback
When you want a lightweight, single-pass run, use the explicit slash command.
| Command | When to prefer |
|---------|---------------|
| /review | Quick sanity check, small diff |
| /ship | Simple pre-merge gate, tests already passing |
| /build | Scaffold-only, no parallel validation needed |
They are also the automatic substitute wherever subagents can't be spawned:
---
config:
theme: base
look: classic
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
themeVariables:
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
fontSize: 16px
primaryColor: "#FFFFFF"
primaryBorderColor: "#C1C4C6"
primaryTextColor: "#242628"
lineColor: "#A2A6A8"
clusterBkg: "#F5FBFF"
clusterBorder: "#F0F1F2"
titleColor: "#707577"
edgeLabelBackground: "#FFFFFF"
flowchart:
curve: basis
wrappingWidth: 240
nodeSpacing: 30
rankSpacing: 40
---
flowchart TD
N["build / review / ship intent"] --> Q{"can this context<br/>spawn subagents?"}
Q -->|"yes"| P["/parallel-build<br/>/parallel-review<br/>/parallel-ship"]
Q -->|"no: you are a subagent<br/>(no Agent tool), or a session<br/>instruction disables spawning"| T["twin<br/>/build · /review · /ship<br/>/team-build also to /build"]
T --> S2["announce the command actually run;<br/>name the lost validation axis once"]
P --> RUN["execute"]
S2 --> RUNcheck.sh verifies each twin exists and is mapped in both the rule and the routing hook.
Planning pipeline: /define, before you build
/define runs /interview (de-fuzz the ask) → /spec (PRD) → /test-cases (QA cases from Figma/Lark, approved by you) → /plan (tasks), pausing for your approval after each, so a bad spec can't quietly turn into bad tasks. It offers /ideate when the approach is open and plan-roaster before build. The result goes into docs/planning/<slug>.md, which every execution skill reads.
---
config:
theme: base
look: classic
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
themeVariables:
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
fontSize: 16px
primaryColor: "#FFFFFF"
primaryBorderColor: "#C1C4C6"
primaryTextColor: "#242628"
lineColor: "#A2A6A8"
clusterBkg: "#F5FBFF"
clusterBorder: "#F0F1F2"
titleColor: "#707577"
edgeLabelBackground: "#FFFFFF"
flowchart:
curve: basis
wrappingWidth: 240
nodeSpacing: 30
rankSpacing: 40
---
flowchart TB
subgraph ROW[" "]
direction LR
subgraph DEF[" "]
direction TB
hd("Define<br/>/define") --- d1("interview") --- d2("spec") --- d3("test-cases") --- d4("plan")
end
subgraph BLD[" "]
direction TB
hb("Build<br/>/parallel-build") --- b1("platform routing") --- b2("scaffold + implement") --- b3("type + lint gates") --- b4("parallel agents + tests")
end
subgraph REV["in parallel · read-only"]
direction TB
hr("Review<br/>/parallel-review") ~~~ r1("code-quality") ~~~ r2("platform reviewers") ~~~ r3("a11y reviewers") ~~~ r4("adversarial (3+ layers)")
end
subgraph SHP[" "]
direction TB
hs("Ship<br/>/parallel-ship · pre-merge") --- s1("coverage gate<br/>RN/web ≥ 93%") --- s2("perf + ponytail agents") --- s3("feature-flag rollback") --- s4("test-case trace<br/>cases approved in Define<br/>missing one = BLOCKED")
end
DEF --> BLD --> REV --> SHP
end
subgraph CHK["YOUR CHECKPOINTS"]
direction LR
k1("approve every phase<br/>continue · edit · stop") ~~~ k2("read the verdict<br/>DONE · READY TO MERGE<br/>BLOCKED · INCOMPLETE") ~~~ k3("act on the findings<br/>review agents are read-only") ~~~ k4("open the PR, merge<br/>opt-in /adr · /docs · /eval")
end
subgraph FIX["SEPARATE PATH FOR BUGS · /fix"]
direction LR
f1("failing test first") --- f2("isolate") --- f3("hypothesize") --- f4("fix") --- f5("regression test")
end
ROW ~~~ CHK ~~~ FIX
classDef n fill:#FFFFFF,stroke:#C1C4C6,color:#242628
classDef key fill:#D1F0FF,stroke:#0A9AF2,color:#242628
classDef ok fill:#FFFFFF,stroke:#029D24,color:#029D24
classDef quiet fill:transparent,stroke:transparent,color:#707577
class d1,d2,d4,b1,b2,b3,b4,r1,r2,r3,r4,s1,s2,s3,f2,f3,f4 n
class hd,hb,hr,hs key
class d3,s4,f5 ok
class k1,k2,k3,k4 quiet
style f1 fill:#FFFFFF,stroke:#0071CE,color:#242628
style ROW fill:transparent,stroke:transparent
style DEF fill:transparent,stroke:transparent
style BLD fill:transparent,stroke:transparent
style REV fill:transparent,stroke:transparent,color:#029D24
style SHP fill:transparent,stroke:transparent
style CHK fill:transparent,stroke:#C1C4C6,stroke-dasharray:3 3
style FIX fill:transparent,stroke:transparent,color:#0071CEIt stops at a reviewed plan. /adr and /docs come later, offered at the end of /parallel-ship once the code is final. Each planning skill also runs on its own. To challenge a plan you already have, use /grill (interactive) or the plan-roaster agent (one shot).
Experimental: /team-build, agent teams
Built on Claude Code's experimental agent teams. Requires
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1. Explicit/team-buildonly, since saying "build feature X" still routes to/parallel-build.
Your session becomes a team lead: it plans the work, then spawns teammates that build different files at the same time and message each other directly, working off a shared task board.
| | /parallel-build (default) | /team-build (experimental) |
|---|---|---|
| Who writes the code | The main session, one file at a time | Two implementer teammates, in parallel |
| Helpers | One-shot reviewers that report back once | Persistent teammates that claim tasks and message each other |
| Coordination | None needed | Shared task board; finishing one task unblocks the next |
| Model split | One session model | Lead on escalated (opus), teammates on everyday (sonnet) |
| Token cost | ~1× | ~5× |
| Best for | Most features | Larger multi-file features where parallel implementation pays for the overhead |
- One file, one owner for the whole build, so nobody overwrites anyone. Teammates ask each other directly (the Presenter owner asks the Model owner about a type) instead of going through the lead.
- Staged spawn. The reviewer and tester start only once there is something to review or test.
---
config:
theme: base
look: classic
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
themeVariables:
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
fontSize: 18px
primaryColor: "#FFFFFF"
primaryBorderColor: "#C1C4C6"
primaryTextColor: "#242628"
lineColor: "#A2A6A8"
clusterBkg: "#F5FBFF"
clusterBorder: "#F0F1F2"
titleColor: "#707577"
edgeLabelBackground: "#FFFFFF"
flowchart:
wrappingWidth: 260
---
swimlane-beta LR
subgraph you["YOU"]
ask("/team-build<br/>explicit only")
end
subgraph lead["LEAD · escalated model"]
pre("preflight<br/>teams on · lead model")
plan("plan + scaffold<br/>one file, one owner")
ver("verify integration<br/>typecheck · lint · tests")
end
subgraph team["TEAMMATES · everyday model"]
impl("impl-a ‖ impl-b<br/>message each other")
rt("reviewer, then tester<br/>staged spawn")
end
subgraph result["RESULT"]
rep("report<br/>+ verdict")
end
ask --> pre --> plan --> impl --> rt --> ver --> rep
classDef you fill:#FFFFFF,stroke:#707577,color:#242628
classDef hook fill:#D1F0FF,stroke:#0A9AF2,color:#242628,stroke-dasharray:4 3
classDef main fill:#FFFFFF,stroke:#0A9AF2,color:#242628
classDef sub fill:#FFFFFF,stroke:#029D24,color:#242628
classDef verdict fill:#0A5C2C,stroke:#0A5C2C,color:#8BE200
class ask you
class pre,plan,ver main
class impl,rt sub
class rep verdictWorks on RN/web, Android and iOS; the task board follows each platform's file layout.
- Claude Code only. Teams are a harness feature, so on other tools the preflight falls back to
/parallel-buildor/build. - ~5× the tokens of a solo build.
- Teammates don't survive
/resume. An interrupted build restarts coordination from the task board.
Full workflow: commands/team-build.md.
Cross-model: /cross-review
Claude and Codex review the same diff without seeing each other's work, then each answers the other's findings once. Your session reads both rounds and keeps what the evidence supports. Explicit /cross-review only.
---
config:
theme: base
look: classic
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
themeVariables:
fontFamily: "-apple-system, BlinkMacSystemFont, Segoe UI, Helvetica, Arial, sans-serif"
fontSize: 18px
primaryColor: "#FFFFFF"
primaryBorderColor: "#C1C4C6"
primaryTextColor: "#242628"
lineColor: "#A2A6A8"
clusterBkg: "#F5FBFF"
clusterBorder: "#F0F1F2"
titleColor: "#707577"
edgeLabelBackground: "#FFFFFF"
flowchart:
wrappingWidth: 260
---
swimlane-beta LR
subgraph you["YOU"]
ask("/cross-review<br/>explicit only")
end
subgraph script["cross-review.sh · fail closed"]
pre("preflight<br/>both CLIs · auth allowlist<br/>tracked diff ≤ 256 KiB")
chk("validate replies<br/>every peer finding answered<br/>tree unchanged")
stop("could not run<br/>no same-model fallback")
end
subgraph panel["PANEL · read-only, repo only"]
r1("round 1, blind<br/>claude -p ‖ codex exec")
r2("round 2, once<br/>AGREE · DISPUTE · CANNOT-VERIFY")
end
subgraph host["HOST · your session"]
adj("adjudicate by evidence<br/>read the cited lines")
end
subgraph result["RESULT"]
rep("findings + provenance<br/>then /eval")
end
ask --> pre --> r1 --> r2 --> chk --> adj --> rep
pre -->|"any check fails"| stop
chk -->|"bad format · tree changed"| stop
classDef you fill:#FFFFFF,stroke:#707577,color:#242628
classDef main fill:#FFFFFF,stroke:#0A9AF2,color:#242628
classDef sub fill:#FFFFFF,stroke:#029D24,color:#242628
classDef stop fill:#FFFFFF,stroke:#D1292E,color:#242628,stroke-dasharray:4 3
classDef verdict fill:#0A5C2C,stroke:#0A5C2C,color:#8BE200
class ask you
class pre,chk,adj main
class r1,r2 sub
class stop stop
class rep verdict| Step | What happens |
|---|---|
| Preflight | Both CLIs present, both auth methods on the allowlist for this project, tracked diff captured and under 256 KiB. Any miss stops the run before anything is sent |
| Round 1 | Both CLIs get the same prompt and diff, blind to each other (claude -p --restricted --strict-mcp-config with Read/Grep/Glob, codex exec --ignore-user-config -s read-only) |
| Round 2 | Each marks every one of the other's findings AGREE, DISPUTE or CANNOT-VERIFY, and may withdraw its own |
| Validate | Replies must parse and answer every peer finding; the tracked diff and untracked path list must be unchanged since the start |
| Synthesis | The host keeps consensus, settles disputes by reading the cited lines, labels the rest |
- Both CLIs are required. If either is missing or fails, the run stops with
cross-review could not run: …instead of quietly becoming a same-model review. - Only one critique round. More rounds make the models drift toward agreement rather than evidence.
- Account safety. Before any diff is sent,
~/.craftkit/cross-review-allowed-authmust hold aproject=/absolute/pathline for this repository (one line per approved project) plus the exact methods allowed for both CLIs, for exampleclaude=claude.aiandcodex=Logged in using ChatGPT.CRAFTKIT_CROSS_REVIEW_POLICYpoints at a different file. The CLIs expose the login method, not the account, so confirm the signed-in accounts are approved. Auth checks and panelist calls clear API keys, alternate endpoints and provider override variables; Codex also ignores user configuration. - Scope. Tracked changes since the merge base, including working-tree edits. Untracked files are never sent, only counted, since an unignored
.envis exactly the file nobody meant to share. Binary diffs name the file without reviewable contents. - Layout-tolerant, content-strict. Code fences, a preamble, wrapped lines and lowercase severities are accepted; a line that does not parse, or a critique that skips a peer finding, stops the run with the raw replies kept for inspection.
- Reproducible. Each run keeps the diff, the exact prompts, both CLI versions and the commit in a unique owner-only directory under
~/.craftkit-state/cross-review/. - Panelists run with
CRAFTKIT_PANELIST=1(the script refuses to start inside one, and the routing hook stays silent) andCRAFTKIT_GATE=off, so the Stop gates cannot block a headless reviewer.
Script: scripts/cross-review.sh, installed to ~/.craftkit/bin/. Workflow: commands/cross-review.md.
Fix, tests, and PR message
"something is broken" / "fix this bug" / "this crashes"
/fix → fe-context → reproduce → isolate → fix → regression test
"write tests" / "add tests" / "coverage is low" → resolves platform first
/fe-test → RN/web: write tests for all changed paths, enforce ≥93% coverage
/android-test → JUnit + MockK Presenter tests (no fixed coverage bar)
/ios-test → Quick + Nimble ViewModel specs (no fixed coverage bar)
"generate PR message" / "draft a PR" / "what should my PR say"
/pr-message → read diff → write title + summary + goal + changes + coverage → humanize (if installed) → copy to clipboardGrill, research, and handoff
Three general-purpose skills adapted from mattpocock/skills (MIT). All natural-language routed; none auto-run.
"poke holes in my plan" / "grill this" / "stress-test my design" → needs an EXISTING plan
/grill → map plan as design tree → ask whole frontier per round (each ❓ with ➡️ recommended
answer) → sub-agents fetch facts, you only decide → done when frontier empty.
Side effects: resolved terms → docs/glossary.md · hard-to-reverse decisions → offers /adr
"research X for me" / "find out how the Y API works" / "dig into the docs"
/research → background agent reads PRIMARY sources only (official docs, source code, specs)
→ cited Markdown note in the repo → you keep working meanwhile
"hand this off" / "summarize this session for the next agent" / "wrapping up for today"
/handoff → handoff doc in OS temp dir: goal, verified state, decisions + why, ordered next
steps, suggested skills. Links existing artifacts by path, never duplicates. Secrets redacted.Picking the right interrogator:
| You have | You want | Use |
|----------|----------|-----|
| A fuzzy new ask, no plan | Requirements extracted | /interview (one question at a time) |
| An existing plan/decision | It challenged, interactively | /grill (frontier rounds) |
| A finished plan doc | A cold second opinion, one shot | plan-roaster agent |
Scoring a run: /eval
Reviews say what is wrong. /eval says how much was right, as one number you can track across runs.
"score this run" / "how correct was that" / "what is our success rate"
/eval → gathers diff + PLANNING acceptance criteria + gate results
→ spawns eval-judge (cold) → five criteria, each 0-5, each weighted
→ recomputes the weighted sum in awk (judgment is the model's, arithmetic is not)
→ appends a row to docs/evals/ledger.md → derives the running success rate| Criterion | Weight | Scored on | |---|---:|---| | Spec conformance | 35 | Every PLANNING acceptance criterion actually implemented | | Correctness | 25 | Edge cases, error paths, no crash or data-loss path | | Pattern adherence | 20 | EVPMR / MVP / MVVM-C contract holds | | Verification | 15 | Tests cover changed paths and pass; type + lint clean | | Simplicity | 5 | Ponytail rubric: nothing to delete |
Correctness % = Σ (score / 5 × weight), so 5·4·4·3·5 is 85.0%. Bands: ≥90 PASS · 75-89 PASS WITH GAPS · <75 BLOCKED.
The scorer is built so a number can't hide a failure:
- Floors beat the band. Any criterion at 0, or Spec conformance / Correctness at 2 or below, is
BLOCKEDwhatever the total. Otherwise "perfect except it doesn't do what was asked" still scores 85%. - Unscorable is a gap. With no intent file, Spec conformance is
n/aand the verdict isINCOMPLETE, reported out of the remaining 65 points instead of reweighted upward. - The success rate is computed from the ledger each time, never stored.
- Optional hard gate: put a
**Threshold:**line in the ledger header and/evaluses it instead of 90, which is the shape for CI./evalitself only reports; it never blocks a merge or reverts.
Skills reference
Always-active rules
Loaded automatically on every session. Never invoke these; they're always present.
| Rule | Enforces |
|------|---------|
| fe-rules | EVPMR layer constraints, TypeScript strict, module-over-barrel imports, styling tokens, React correctness, tracking |
| flag-safety | Flag OFF stays behavior-identical: code paths, persisted state, API contracts, analytics. flag: marker, both states tested |
| grounding | Claims that drive action carry provenance: [verified: how], [from context.md @sha], [UNVERIFIED]. An [UNVERIFIED] claim cannot back an [ERROR] finding or an edit. Cold agents review handed content only; staleness reports cannot-verify, never clean |
| karpathy-guidelines | Think before coding, simplicity, surgical changes, goal-driven, read before write, tests verify intent, checkpoint after steps |
| using-agent-skills | Skill routing (mandatory gate: classify before every response, announce match or "No skill matched."), model selection, severity labels, parallel classifier, model for judgment only, surface conflicts |
Frontend skills, on demand
Use when a task is narrower than a full workflow.
| Skill | When to use | Escalate if |
|-------|-------------|-------------|
| fe-context | Derive the branch's change context from the diff, emitted into the turn, no file written | Diff spans > 10 interdependent files |
| fe-scaffold | Create a new 5-file EVPMR module | Novel architecture outside EVPMR |
| fe-review | EVPMR pattern review only | Architectural conflicts with non-obvious resolution |
| fe-patterns | Props drilling, shared state placement (Context in Model, provider in Entry), composition patterns, hooks discipline | Novel state architecture |
| fe-performance | Waterfall elimination, bundle size, re-renders | Lighthouse regressions with non-obvious root cause |
| fe-a11y | Labels, roles, focus management, reduced motion, for RN & Next.js | Complex focus flows spanning multiple routes |
| fe-design | Visual design that reads as AI-generated: default gradients and glass, template layouts, decorative filler, invented dashboard numbers, responsive breaks. Adapted from anti-slop (MIT) | Whether a technique earns its place is contested |
| fe-test | Write/improve tests, enforcing ≥93% coverage. RN/web only, since native goes to /android-test / /ios-test | Can't reach 93%, root cause unclear |
Native mobile skills, on demand
Native mobile does not use EVPMR. For single-screen work, read a real sibling screen instead of deriving context. For concrete module names, add a project override at <repo>/.claude/skills/<name>/; the same name shadows the global skill inside that repo.
The *-review, *-a11y and *-performance skills are also the live source for the matching cold agents (via craftkitInject), so editing the skill updates the agent on the next sync.
Android: MVP + Core framework, Dagger, Gradle Dynamic Feature Modules:
| Skill | When to use | Escalate if |
|-------|-------------|-------------|
| android-patterns | Architecture reference: MVP layers, DI, module split, navigation | Novel state/effect orchestration |
| android-scaffold | Scaffold a new screen (View/Presenter/ViewModel + Dagger wiring) | Outside the Core MVP contract |
| android-review | Review a diff against the MVP contract | Architectural conflict, non-obvious resolution |
| android-a11y | TalkBack labels/state, touch targets, Compose semantics | Complex focus flows across screens |
| android-performance | Main-thread/coroutine, RecyclerView, recomposition, leaks | Jank/leak with non-obvious root cause |
| android-test | JUnit + MockK Presenter tests (Turbine for Flow) | Path unreachable without production refactor |
| android-context | Branch-scoping doc for multi-screen work | Multi-module cross-feature -api changes |
iOS: MVVM-C, Bazel + CocoaPods, Quick + Nimble:
| Skill | When to use | Escalate if |
|-------|-------------|-------------|
| ios-patterns | Architecture reference: MVVM-C, Fetcher, Coordinator, Dependency-struct DI | Novel state/effect orchestration |
| ios-scaffold | Scaffold a new screen (Contract/VC/View/ViewModel/Factory/Fetcher) | Outside the MVVM-C contract |
| ios-review | Review a diff against the MVVM-C contract | Architectural conflict, non-obvious resolution |
| ios-a11y | VoiceOver labels/traits, focus, Dynamic Type, red
