@nathapp/nax
v0.83.4
Published
AI Coding Agent Orchestrator — loops until done
Maintainers
Readme
nax
AI Coding Agent Orchestrator — loops until done.
Give it a spec. It writes tests, implements code, verifies quality, and retries until everything passes.
Why nax
nax is an orchestrator, not an agent — it doesn't write code itself. It drives whatever coding agent you choose through a disciplined loop until your tests pass.
- Agent-agnostic — runs its own in-process native agent by default, or drives Claude Code, Codex, Gemini CLI, OpenCode, or any ACP-compatible agent
- TDD-enforced — acceptance tests must fail before implementation starts
- Loop until done — verify, retry, escalate, and regression-check automatically
- Monorepo-ready — per-package config and per-story working directories
- Extensible — plugin system for routing, review, reporting, and post-run actions
- Language-aware — auto-detects Go, Rust, Python, TypeScript from manifest files; adapts commands, test structure, and mocking patterns per language
- Semantic review — LLM-based behavioral review against story acceptance criteria; catches stubs, placeholders, and out-of-scope changes
- Adversarial review — LLM-based adversarial code review that probes for input handling, error paths, and abandoned implementations
- Context curator — deterministic post-run analysis that proposes additions/deletions to context.md and rules files, preventing context drift
- Guarded agent commands — agent-authored shell commands run inside an OS sandbox by default, are adjudicated by a per-stage bash approval mode, and can pause for interactive approval that you can remember and later revoke (
nax approvals)
Install
npm install -g @nathapp/nax
# or
bun install -g @nathapp/naxRequires: Bun 1.3.7+ (nax runs on the Bun runtime even when installed through npm; CI pins Bun 1.4.0). Git must be initialized.
The default native agent needs provider credentials — run nax auth login <provider> or, for CI, set the provider's environment variable (a stored credential takes precedence).
Quick Start
cd your-project
nax init # Create .nax/ structure
nax setup # Optional: LLM-analyze the repo and write .nax/config.json
nax features create my-feature # Scaffold a feature
# Write your spec, then plan + run
nax plan -f my-feature --from spec.md
nax run -f my-feature
# Or plan and run in one command
nax run -f my-feature --plan --from spec.mdSee docs/ for full guides on configuration, test strategies, monorepo setup, and more.
How It Works
(plan →) acceptance setup → route → execute → verify → review (semantic + adversarial) → escalate → loop → regression gate → acceptance- Plan (optional) — Generate
prd.jsonfrom a spec file using an LLM - Acceptance setup — Generate acceptance tests; assert RED before implementation
- Route — Classify story complexity and select model tier (fast → balanced → powerful)
- Context — Gather relevant code, tests, and project standards per story
- Execute — Run agent session (native in-process agent by default, or an ACP agent such as Claude Code, Codex, Gemini CLI)
- Verify — Run scoped tests; rectify on failure before escalating
- Review — Run lint + typecheck + semantic review + adversarial review; autofix before escalating
- Escalate — On repeated failure, retry with a higher model tier
- Loop — Repeat steps 3–8 per story until all pass or a cost/iteration limit is hit
- Regression gate — Run full test suite after all stories pass
- Acceptance — Run acceptance tests against the completed feature
CLI Reference
| Command | Description |
|:--------|:-----------|
| nax init | Initialize nax in your project |
| nax setup | Analyze the repo and generate .nax/config.json via LLM |
| nax features create | Scaffold a new feature directory |
| nax features list | List all features and story status |
| nax features resolve | Resolve a feature name and its spec source |
| nax plan | Generate prd.json from a spec file (--decompose <storyId> splits an existing story) |
| nax spec lint | Check a spec's machine-extracted sections before planning |
| nax run | Execute the orchestration loop (--compare for a multi-agent bake-off, --schedule to defer) |
| nax resume | Resume an interrupted run from its checkpoint |
| nax precheck | Validate project readiness |
| nax status | Show live run progress |
| nax logs | Stream or query run logs |
| nax runs | List recorded run metadata (nax runs show <run-id>) |
| nax replay | Reconstruct a post-mortem timeline for a previous run |
| nax accept | Override failed acceptance criteria |
| nax unlock | Release a stale lock from a crashed nax process |
| nax generate | Generate agent context files (CLAUDE.md, AGENTS.md, …) from .nax/context.md |
| nax prompts | Assemble or initialize prompts |
| nax context | Inspect context-engine artifacts and feature fragments |
| nax rules | Lint, export, or migrate the canonical rules store (.nax/rules/) |
| nax detect | Detect test-file patterns and optionally persist them |
| nax agents | List available coding agents |
| nax auth | Manage provider credentials for the native agent (login, import, list, rm) |
| nax approvals | List or revoke remembered command approvals (list, rm) |
| nax trust | Trust or untrust project folders (list, add, rm, check) |
| nax mcp lock | Pin configured MCP servers' tool surface to .nax/mcp-lock.json |
| nax config | Display the effective merged config (--explain, --diff); nax config profile manages config profiles |
| nax curator | Inspect, commit, or garbage-collect curator proposals |
| nax routing calibrate | Propose complexity→tier mapping adjustments from run history |
| nax plugins list | List installed plugins |
| nax migrate | Move generated content from .nax/ to the output directory (~/.nax/<project>/) |
For full flag details, see the CLI Reference.
Configuration
.nax/config.json is the project-level config. Key fields:
{
"agent": {
"protocol": "hybrid", // "acp" | "native" | "hybrid" — which transports are permitted
"default": "native" // In-process nax-ai agent; or an ACP agent such as "claude"
},
"execution": {
"maxIterations": 20, // Note: `nax run -m <n>` overrides this only when the flag is passed
"permissionProfile": "unrestricted", // "unrestricted" | "safe" | "scoped"
"storyIsolation": "shared", // "shared" | "worktree"
"bashApproval": "raw", // "raw" | "gated" | "escalate" — how agent Bash commands are adjudicated
"sandbox": {
"enabled": true, // OS sandbox around agent-authored Bash / Exec commands (on by default)
"network": { "allowedDomains": ["registry.npmjs.org"] } // Omit for unrestricted, [] for no network
},
"commandInterceptor": {
"provider": "rtk", // Token-reducing proxy for the Git tool
"enabled": true, // Off by default — opt in per project
"git": { "verbs": ["log", "diff"] } // Only these subcommands are rewritten
}
},
"tdd": {
"strategy": "auto" // How to write tests (see Test Strategies)
},
"routing": {
"strategy": "keyword" // "keyword" | "llm"
},
"quality": {
"commands": {
"test": "bun test", // Root test command
"lint": "bun lint", // Optional linter
"typecheck": "bun typecheck" // Optional type checker
}
},
"mcp": {
"servers": {
"codebase-memory": {
"command": "codebase-memory-mcp", // stdio MCP server binary (client only)
"args": [], // Optional server args
"stages": ["run"], // Attach in these pipeline stages ("*" = all)
"allowedTools": ["search_graph"] // Optional: subset of locked tools that is grantable
}
}
}
}mcp attaches external Model Context Protocol (MCP) servers as tool providers — nax is a client only, never an MCP server. The server id is the tool-name namespace: the codebase-memory server advertises its tools as codebase-memory__search_graph, codebase-memory__trace_path, and so on. Before any of those tools are grantable, run nax mcp lock at the project root: it connects every enabled server once, pins the advertised tool surface (name + input-schema hash) to .nax/mcp-lock.json, and that lockfile is committed like bun.lock. stages is the attachment control — a server's tools attach only to the listed pipeline stages, and an empty list attaches nowhere. MCP tool reach follows the permission profile: unrestricted advertises every attached server's tools, scoped only what the stage's Mcp(...) rules admit, and safe none at all. allowedTools narrows which locked tools are grantable; omitted means every locked tool is.
execution.commandInterceptor rewrites the Git tool's argv through rtk so log and diff output reaches the model compressed. It is confined to the Git site: user-authored quality.commands and acceptance.command are never wrapped. It fails open — if the rtk binary is missing the call runs as plain git.
Both features are native-agent only. An ACP agent (claude, codex, opencode, gemini) brings its own tools, so nax's Git tool is never invoked and no MCP tool is advertised. A project on "protocol": "acp" can hold a complete, valid config for both and get zero effect, with no error. The built-in defaults (agent.protocol: "hybrid", agent.default: "native") let both take effect once configured (the interceptor itself is off by default); a config that switches to an acpx agent does not.
See MCP & Command Interception for setup, verification and troubleshooting, and the Configuration Guide for the full schema.
execution.bashApproval decides how an agent's Bash command is adjudicated (ADR-030). The default raw is a pass-through — no per-segment grant matching or root containment, only a best-effort screen that refuses a parseable command naming a nax-owned file (.nax/config.json, a feature prd.json, the queue-control files). gated matches each segment against the stage's single Bash(...) allow rule, and escalate turns a denial the gate could not adjudicate into an interactive approval prompt; under either, a stage without a Bash(...) rule never gets the tool, and nax warns about such inert stages at run start. Approvals you choose to remember are kept per project and managed with nax approvals list / nax approvals rm; execution.approvalTimeout (default 600000 ms) bounds how long a prompt waits before denying.
execution.sandbox wraps agent-authored Bash and RunCommand exec commands in an OS sandbox (backend srt, on by default): writes are confined to the repository root, system temp directories and package-manager caches, credential files are unreadable, and network.allowedDomains optionally limits network access. When the sandbox is enabled but unavailable on the machine, raw Bash is refused rather than run unsandboxed — switch the stage to gated/escalate or set sandbox.enabled: false. execution.commandSafety.shadow optionally attaches a loopback shadow classifier that scores every agent command and records the result without ever deciding anything.
These settings govern nax's own Bash and RunCommand tools, so, like MCP and the interceptor, they take effect for the native agent; an ACP agent runs commands under its own tooling.
See Sandbox & Command Safety, Approvals, The Bash Tool and Permissions.
Key Concepts
Project trust
nax runs repository-controlled code — project plugins, context plugin providers, hooks, MCP servers and the quality / test / acceptance / setup commands — unsandboxed on the host. Before a gated command (nax run, nax plan, nax mcp lock, …) acts on a project, that project's folder must be trusted, and a refusal exits 2 with run: nax trust add <root>.
nax trust add "$PWD" # trust this checkout (asks for confirmation)
nax trust list # what is trusted, and which entry covers the cwd
nax trust check --json # machine-readable answer for a scriptTrust is hierarchical and stored once per machine (~/.nax/trust.json): an entry for a folder covers that folder and every folder beneath it, so trusting a parent covers the clones, worktrees and new projects you create under it later — you do not re-grant trust per checkout. Trusting / or your home directory would cover everything you own and is refused unless you pass --force. In CI, grant trust once at the top of the job with nax trust add "$PWD" --yes, which is a no-op (exit 0) on a cache that already covers the path.
See nax trust for the full surface.
Test Strategies
nax supports five test strategies. When tdd.strategy is "auto" (default), the planner selects the strategy per story based on complexity and content — security-critical stories always get three-session-tdd regardless of complexity.
| Strategy | Sessions | When to use |
|:---------|:---------|:------------|
| three-session-tdd | 3 | Expert stories and security-critical code (auth, tokens, RBAC) — strict isolation: test-writer cannot touch src/, implementer cannot touch tests |
| three-session-tdd-lite | 3 | Complex stories — relaxed isolation: test-writer may add minimal src/ stubs |
| tdd-simple | 1 | Simple and medium stories — single session, TDD discipline (red → green → refactor) |
| test-after | 1 | Exploratory / prototyping — implement first, add tests after |
| no-test | 0 | Config-only, docs, CI, dependency bumps — requires noTestJustification |
See Test Strategies Guide for the full routing decision tree and security override rules.
Story Decomposition
An oversized story is split into smaller sub-stories with nax plan -f <feature> --decompose <storyId> — a plan-time operation, not a mid-run stage. The precheck story-size gate (precheck.storySizeGate) flags stories over its thresholds. The parent is marked decomposed and the sub-stories, with dependency ordering, are added to the PRD.
See Story Decomposition Guide.
Regression Gate
After all stories pass, nax runs the full test suite once. If it fails, nax maps each failing test file back to the story that introduced it and runs a targeted rectification cycle for that story. A suite timeout is accepted as a pass by default (execution.regressionGate.acceptOnTimeout). If failures remain, the affected stories are marked regression-failed, the run status becomes failed, and the on-final-regression-fail hook fires.
Parallel & Isolated Execution
With --parallel <n>, nax runs up to n stories whose dependencies are already satisfied at the same time, each in its own git worktree (0 = auto, based on CPU cores). Both modes run the regression gate once at the end; only sequential runs attribute a regression to a story and rectify it — in parallel mode it is reported for a human.
Even in sequential mode, stories can be isolated in per-story git worktrees (execution.storyIsolation: "worktree") to prevent cross-story state leakage.
Monorepo Support
Per-package context files, per-package test commands, and per-story working directories are supported. Initialize with nax init --package packages/api. Package config files live at .nax/mono/packages/<pkg>/config.json.
See Monorepo Guide.
Hooks
Lifecycle hooks, defined in .nax/hooks.json (project) and ~/.nax/hooks.json (global) — not in config.json — fire at key points (on-start, on-story-complete, on-all-stories-complete, on-complete, on-final-regression-fail, and more). Use them to trigger deployments, send notifications, or integrate with external systems.
See Hooks Guide.
Plugins
Extensible plugin architecture for prompt optimization, custom routing, code review, and reporting. Plugins live in .nax/plugins/ (project) or ~/.nax/plugins/ (global). Post-run action plugins (e.g. auto-PR creation) can implement IPostRunAction for results-aware post-completion workflows.
See Plugin System.
Agents
The default agent is native: nax drives the model in-process over @nathapp/nax-ai (no CLI binary), using the built-in models.native Anthropic tier map, so it needs Anthropic credentials (nax auth or the provider's environment variable) unless you override the map. Every other agent is reached via ACP (Agent Client Protocol) — a JSON-RPC protocol that provides persistent sessions, exact token/cost reporting, and multi-turn session continuity.
| Agent | Binary | Notes |
|:------|:-------|:------|
| Native (nax-ai) | — (in-process) | Default. agent.default: "native" |
| Claude Code | claude | Set agent.default: "claude" |
| OpenCode | opencode | Set agent.default: "opencode" |
| Codex | codex | Set agent.default: "codex" |
| Gemini CLI | gemini | Set agent.default: "gemini" |
| Pi Coding Agent | pi | Set agent.default: "pi" (via the pi-acp bridge) |
| Aider | aider | Known name with no dedicated ACP adapter entry (generic defaults) |
| Any ACP-compatible | — | See acpx agent docs |
See Agents Guide and the Context Engine Guide for agent-portable context configuration.
Troubleshooting
| Problem | Solution |
|:--------|:---------|
| Precheck blocks with "Uncommitted changes detected" | Commit or stash your changes — nax's own runtime files are ignored by the check |
| HOME env warning | Set HOME to an absolute path — nax warns if it contains ~ |
| ACP sessions leaking | Upgrade to nax v0.48+ and ensure .nax/acp-sessions.json is gitignored |
| Monorepo packages misclassified | Ensure .nax/mono/packages/<pkg>/config.json is set up per package |
| Agent Bash refused with "sandbox unavailable" | The OS sandbox is on by default and raw Bash requires it; the message names why the sandbox probe failed (e.g. unsupported platform, or a container where the sandbox cannot enforce). Fix the environment, set the stage's bashApproval to gated/escalate, or set execution.sandbox.enabled: false |
| Acceptance tests regenerating every run | Check acceptance-meta.json — stale fingerprints indicate outdated story context |
See the Troubleshooting Guide for more.
Credits
nax is inspired by Relentless — the same "keep trying until done" philosophy, applied to AI agent orchestration.
ACP support is powered by acpx from the OpenClaw project.
License
MIT
