agent-harness-kit
v0.24.0
Published
Solo-dev harness engineering kit for Claude Code, with experimental Codex and Kiro CLI runtime rendering.
Downloads
232
Maintainers
Readme
agent-harness-kit
The infrastructure layer that makes AI agents production-ready.
Solo-dev harness engineering kit for Claude Code, with an experimental Codex-readable runtime surface. One command, ~30 minutes, and your hobby project gets the patterns that took OpenAI from prototype to 1M lines of agent-generated code: layered architecture, structural tests, garbage collection, review subagents, JSON feature tracking, and pre-completion checklists — without the enterprise overhead.
Honest Expectations
What this kit DOES differentiate from bare claude-cli (anecdotal + design-level):
- ✅ Opinionated CLAUDE.md template (50–80 lines) so context isn't blown on style
- ✅ 33 skills that codify Hashimoto/OpenAI rituals
- ✅ 10 read-only review subagents for cheap second-opinion passes and mandatory done-claim advice
- ✅
.harness/feature_list.json+.harness/project/state.json+ ADR template + shared project memory for solo-scale planning hygiene - ✅ Solo-dev cost defaults (~$2/day) and per-run budget enforcement
What it does NOT measurably differentiate (5 consecutive null benches, May 2026):
- ❌ Structural enforcement on happy-path 1-shot tasks. When seed code shows the layer pattern, claude-cli follows it — the boundaries lint has nothing to catch. We measured 0/6 ui→repo violations across bare and kit arms on the
ts-layeredfixture.
Where the structural test MIGHT still earn its keep (untested, listed for honesty, not as a claim):
- Long multi-turn sessions where pattern context drifts
- Adversarial "make it fast" pressure that tempts shortcuts
- Greenfield code with no existing pattern to follow
- Weaker model substrates (haiku, gpt-4o-mini)
Use the lint as a safety net, not as the reason you adopted the kit.
Overhead (measured, and shrinking)
The kit's costs are constant; its benefits are conditional. So it measures its
own footprint on every release (npm run report:harness-tax) and a CI guard
(check:harness-tax) refuses any change that makes these numbers regress:
| What a fresh init costs you | Value |
|---|---|
| Fixed per-session context (CLAUDE.md, no eager imports) | ~2.9 KB / ~710 tokens |
| Default hook commands (Claude) | 7 |
| Hooks that fire per Edit | 1 |
| Stop hook | scoped, warn-by-default on the normal lane |
Heavy governance is opt-in, not default. See docs/harness-engineering.md for
the "why".
Installation
Option A: One-line install (recommended)
curl -sL https://raw.githubusercontent.com/tuanle96/agent-harness-kit/main/install.sh | bashIf the interactive prompt exits with aborted by user at Project name in a piped shell, rerun with defaults:
curl -sL https://raw.githubusercontent.com/tuanle96/agent-harness-kit/main/install.sh | bash -s -- --yesOr run the initializer directly so the prompt owns the terminal input:
npx agent-harness-kit initUpgrade existing installation:
curl -sL https://raw.githubusercontent.com/tuanle96/agent-harness-kit/main/install.sh | bash -s -- --upgradeUpgrade is non-destructive: user-edited managed files get sidecars, and
user-owned config is only patched for missing compatibility defaults such as
project memory, task/evidence contracts, failure-learning signal bundles,
orchestration/model-routing policy, and release-readiness review promotion
gates. Use npx agent-harness-kit upgrade --plan to write an impact artifact
under .harness/upgrades/<runId>/plan.json before applying, and
npx agent-harness-kit upgrade --explain <changeId> to inspect a single
planned change with category, diff summary, and rollback notes.
To delegate install or upgrade to an AI agent, generate a versioned onboarding prompt instead of hand-writing a checklist:
npx agent-harness-kit prompt install --runtime codex
npx agent-harness-kit prompt upgrade --runtime claude,codexThe prompt tells the agent to inspect the repo, run doctor, apply init/upgrade,
verify runtime hooks, and prove .harness/memory/current-summary.md is injected
by SessionStart. The matching machine gate is:
node .harness/scripts/check-runtime-surface.mjs --runtime=codexOption B: Scaffold into existing repo
npx agent-harness-kit initFor the experimental Codex surface:
npx agent-harness-kit init --runtime codex
npx agent-harness-kit init --runtime claude,codexOption C: Install as Claude Code plugin
/plugin marketplace add tuanle96/agent-harness-kit
/plugin install agent-harness-kit@agent-harness-kit-marketplaceSupport Matrix
| Stack | Adapter | Preset | Dev command | Status |
| ------------------------------ | ------------ | ----------- | -------------------------------------- | ------ |
| Next.js 14 + TypeScript | typescript | nextjs | npm run dev | v0.1 |
| Express | typescript | node-api | node ./src/server.js | v0.1 |
| Fastify | typescript | node-api | node ./src/server.js | v0.1 |
| NestJS | typescript | node-api | npm run start:dev | v0.1 |
| FastAPI | python | fastapi | uvicorn app.main:app --reload | v0.1 |
| Django | python | django | python manage.py runserver | v0.1 |
| Flask | python | flask | flask --app app run --debug | v0.1 |
| Go | go | none | go run ./cmd/... | v0.4 |
| Rust | rust | none | cargo run | v0.4 |
| Swift | swift | none | swift run | v0.7 |
| Kotlin | kotlin | none | ./gradlew run | v0.7 |
CLI Commands
agent-harness-kit init # scaffold a repo (interactive)
agent-harness-kit init --yes # accept all detected defaults
agent-harness-kit init --runtime codex # render AGENTS.md + .codex/.agents + .harness/* only
agent-harness-kit init --runtime claude,codex # render both runtime instruction files
agent-harness-kit init --runtime antigravity # render GEMINI.md + .agents + .harness/* only
agent-harness-kit init --runtime claude,antigravity # render both Claude and Antigravity surfaces
agent-harness-kit init --pack nextjs-saas # apply Next.js SaaS governance defaults
agent-harness-kit init --pack api-backend,python-data # apply multiple policy packs
agent-harness-kit upgrade # non-destructive upgrade, preserves user edits
agent-harness-kit upgrade --plan # write .harness/upgrades/<runId>/plan.json without applying managed file changes
agent-harness-kit upgrade --apply --yes # explicitly apply the planned upgrade behavior
agent-harness-kit upgrade --explain <changeId> # explain one change from the latest plan artifact
agent-harness-kit doctor # diagnose runtime + run harness preflight checkers
agent-harness-kit doctor --strict # also run the full readiness gate
agent-harness-kit doctor --runtime codex # diagnose the Codex surface
agent-harness-kit prompt install --runtime codex # print an AI-agent onboarding prompt
agent-harness-kit prompt upgrade --runtime claude,codex # prompt an agent to upgrade + verify memory/hooks
agent-harness-kit explain --bypass <fingerprint> # explain bypass audit coverage and repair path
agent-harness-kit context-query "How does task evidence validation work?" --scope scripts --lane normal --json
agent-harness-kit state doctor # diagnose SQLite state health
agent-harness-kit state migrate --dry-run # preview state schema migrations
agent-harness-kit state export --redact # export shareable redacted state JSON
agent-harness-kit state prune --older-than=30d --dry-run # preview retention cleanup
agent-harness-kit state explain <runId> # inspect matching trace/story/session rows
agent-harness-kit report --html # write .harness/reports/harness-dashboard.html
agent-harness-kit strictness set strict # migrate readiness gates to a stricter tier
agent-harness-kit pack validate --pack nextjs-saas # validate one policy pack manifest + rule examples
agent-harness-kit pack publish --pack nextjs-saas --dry-run --json # review a publishable pack bundle
agent-harness-kit --version
node .harness/scripts/harness-readiness.mjs --strict # generated project release/readiness gate
node .harness/scripts/harness-state.mjs init # initialize SQLite operational state
node .harness/scripts/harness-state.mjs trace-quality --strict # score trace/friction quality
node .harness/scripts/strictness.mjs plan --tier=release # preview compiled readiness gates
node .harness/scripts/runtime-parity-report.mjs # publish Claude/Codex parity scorecard
node .harness/scripts/runtime-conformance.mjs # publish runtime adapter conformance results
node .harness/scripts/check-trace-corpus.mjs # validate sanitized public trace corpus
node .harness/scripts/pr-annotations.mjs --github-annotations # write PR annotations, Markdown, and SARIF
node .harness/scripts/check-policy-packs.mjs # validate installed policy packs
node .harness/scripts/check-stable-schemas.mjs --strict # validate schema compatibility policy
node .harness/scripts/check-evidence-attestation.mjs --strict # validate replayable evidence sidecars
node .harness/scripts/policy-pack-publish.mjs --pack nextjs-saas --dry-run # create pack publish plan
node .harness/scripts/check-session-isolation.mjs --strict # active-task worktree isolation gate
node .harness/scripts/prepare-session-worktree.mjs --task <id> # create isolated task worktree
node .harness/scripts/orchestration-contract-from-task.mjs <id> # derive workflow contract from task contractcontext-query works without external services. Install srcwalk for stronger
structural matches:
npm install -g srcwalkUse --require-srcwalk only when your dev image or CI should fail closed if
that binary is missing.
Dependency Footprint
Runtime dependencies are intentionally split by surface:
| Dependency | Why it is present | Impact if missing |
| ---------- | ----------------- | ----------------- |
| commander | CLI command routing (init, upgrade, doctor) | CLI cannot start |
| @inquirer/prompts | Interactive init/upgrade prompts | Interactive mode fails; --yes paths still avoid most prompts |
| @clack/prompts | Polyglot setup selector with cancel handling | Polyglot-root setup falls back poorly |
| react + ink | Rich polyglot onboarding renderer only, not the hot scaffold path | Smart setup loses the app map UI; core render/upgrade logic still does not depend on React state |
| handlebars | Template rendering | init/upgrade cannot render scaffold files |
| picocolors | CLI diagnostics | Output loses structured color but behavior is otherwise unchanged |
Optional peer dependencies are adapter tooling, not core runtime:
| Peer dependency | Used by | When missing |
| --------------- | ------- | ------------ |
| ts-morph | TypeScript structural runner | npm run harness:check fails with an explicit install message |
| eslint-plugin-boundaries | TypeScript ESLint defense-in-depth config | ESLint boundary config cannot run, but the ts-morph runner remains the primary gate |
| dependency-cruiser | Optional TypeScript dependency graph checks | Dependency-cruiser reports are unavailable; structural runner still enforces layer direction |
The TypeScript init path patches these peer tools into the target repo's devDependencies non-destructively.
Deep docs
Kept lean on purpose. The detail lives next to the code:
docs/what-ships.md— full inventory (33 skills, agents, scripts) + directory structure.docs/configuration-reference.md— every.harness/config.jsonfield.docs/philosophy.md— the 5 axioms.docs/harness-engineering.md— why harness engineering, and the 2025–2026 trend.
License
MIT
Contributing
Issues and PRs welcome at github.com/tuanle96/agent-harness-kit
Found a bug? Open an issue.
Have a pattern from your own harness? Submit a PR with the skill or hook.
Want to add a language adapter? Check docs/adding-an-adapter.md and the existing TypeScript/Python/Go/Rust/Swift/Kotlin adapters.
