railhead
v0.4.0
Published
Unattended, dependency-ordered ticket execution via opencode — implement, verify, review, commit, one ticket at a time.
Maintainers
Readme
Railhead
Drive opencode over a queue of tickets, unattended: each ticket is implemented, verified, reviewed, and committed before the next begins, in the order the plan emitted. A contracts index keeps every judging phase's context O(ticket), so a small local model or a cloud model can run an arbitrarily large project without holding it all at once.
How it works
A plan is sliced into an ordered ticket queue; each ticket runs a gate and lands as one commit before the next begins. Optional review gates can add corrective tickets — or replan the remainder — before the run moves on.
per ticket, in plan order
BUILD durable opencode session resumed across checkpoints
VERIFY your build/test commands; must be green
SMOKE optional launch command stays up (panic / not-found signatures fail)
REVIEW read-only, diff-scoped critique of the change
(light preset: only when the build was stressed)
COMMIT "<NN> — <title>", then the contracts index is updated
VISUAL optional post-commit screenshot review; completes (correctives
included) before the next ticket starts
at group checkpoints and run end
GOAL optional judges the integrated build against the goal and design docs
STRUCTURAL optional judges accumulated source for architectural drift
RUN END end-of-run visual/goal/structural passes, then report.md- SMOKE runs when the planner emitted a
$SMOKElaunch command and the framework is recognized; a process still running at the timeout passes. - Retries feed findings back into the same builder session.
[BLOCKER]findings always retry (then hard-fail);[MAJOR]retry per cadence. Any[BLOCKER]can become a corrective ticket that runs the full gate inline before the originating review passes; a failing goal review can request a replan of the remaining tickets. - Tickets run strictly in the order the planner emitted; the frontier is the first ready ticket. Committed tickets never re-open.
- The builder is one durable opencode session resumed across checkpoints, so the author of the code receives review feedback directly (ADR 0047).
- Gates are always fresh and diff-scoped — reviewers never inherit the builder's context — and phases can push
LEARNED:facts into.railhead/learnings.mdfor later prompts. - The ledger (
.railhead/<run-id>/) records state and the raw event stream per phase; a crashed or stopped run resumes where it left off. - Failures escalate through a three-rung ladder (retry → restart worker → diagnose/fail), with step, stall, and degraded-target guards for unattended runs.
Install
npm install -g railhead # or: npx railhead <command>Requires Node 20+, opencode on PATH, and a configured model (local or hosted). From a clone: npm install && npm link.
Quick start
railhead build "a CLI that parses RSS feeds" # interview → PLAN.md → tickets → run
railhead fix "paddles don't move" # bug report → reproduction questions → fix ticketrailhead init (run automatically on first use) asks for a model per seat and probes each one for context limit, vision, and reasoning. opencode models lists what is available.
Building a product, not a demo (ADR 0051)
A product is not one shot: its vision exists first, an MVP proves it, and it grows feature by feature over months. Railhead's build loop is exactly what one feature needs — planning, building, reviewing, fixing until good — so it doubles as the feature engine:
railhead product "a hiking log my family actually opens; first a rough MVP, then search" # condense (or steer) the product arc
railhead feature # build the next roadmap step, unattended, overnight
railhead status # next morning: the arc's steps + run outcome
railhead fix "...bug report..." # anything the morning found- The product arc (
docs/product.md) is the durable artifact: decided vision, traits, workflows, stack, and an ordered roadmap of steps that all future feature runs steer against. Hand-editable; one condense call, then a sharpening interview (depth picker,--sharpen/--no-sharpen) questions the MVP cut and each step's outcome; answer:done(or press Ctrl-D) to end the interview early — answers so far still apply. Every answer is appended to.railhead/plan-latest/interview.jsonlas it is given, so a provider failure loses nothing;railhead product --answers .railhead/plan-latest/interview.jsonl "<what you were steering>"then finishes the revision from the recorded answers alone. The full prose is shown and adoption is explicit. railhead featurederives the first eligibletodostep into a structured feature prompt (or takes"<one feature>"/--step N), grounds it in the repo's contracts/digest/learnings (and the previous attempt's report when reopening), plans it in feature posture — no scaffold, contracts/charter/stack are decided — and builds it through the full gate stack. On finish the step is markedbuiltand the arc update is committed on the run branch.- The human gate is enforced — a
builtstep blocks the steps after it until you test it and answer the morning prompt (done/leave/reopen <feedback>), or override with--step N; a resumed run still knows which step it builds. Merge therun/<step-NN-title>branch to carry the arc update. - Consistency across builds is machinery, not hope: the decided stack and the coherence charter ride every feature plan (a held charter owns the look — a craft ticket is only added for a genuinely new surface), the run's plan docs live in a stable per-step namespace (project docs are never overwritten), and the goal review judges the feature against the step itself — later steps are out of scope by design.
- A red baseline refuses to start —
railhead featurechecks the project's verify is green first (runrailhead fixbefore the feature), and a feature plan copies that suite rather than extending the global gate. The plan that seeds an empty verify list is exempt (greenfield step 1 has no suite to be green yet — ADR 0053).
build generates the plan, a bounded clarifying interview (preset-gated) revises it — every answer is persisted to .railhead/plan-latest/interview.jsonl as it is given, and :done/Ctrl-D ends the interview early with the answers collected so far — and then, unless -a/--auto, the final plan is written to PLAN.md for you to read and request changes; accepting it decomposes the plan into tickets. Resolved vocabulary lands in CONTEXT.md, hard decisions in docs/adr/, and the design narrative and architecture in docs/design.md / docs/architecture.md. fix asks only about reproduction — the planner reads the code itself.
Commands
railhead init git init + default railhead.json
railhead build "<prompt>" [flags] turn a description into tickets, then run them
railhead fix "<bug report>" [flags] turn a bug report into a fix ticket, then run it
railhead run <tickets-dir> run a queued ticket set (auto-resumes interrupted runs)
railhead resume [<run-id>] continue a stopped/interrupted run
railhead status [<run-id>] live summary of a run
railhead next [<run-id>] next actionable ticket(s) and what blocks them
railhead log [<run-id>] [<phase>] readable transcript of a phase from the ledger
railhead reset [--hard] abandon the latest interrupted run (--hard drops its commits)
railhead diagnose screenshots [--model M] check that a model can take a screenshot and read it backbuild/fix flags: [--model M] [-a|--auto] [-c] [--full|--medium|--light|--none] [--verbose] [--yolo], plus per-gate overrides (--review, --vision, --goal, --structural, --sharpen). run accepts [--plan M] [--exec M] [--review M] [--visual M] [--extract M] [--goal-model M] [-m N] [--pause-on-failure] [--quiet|--verbose] [--fresh] and the same gate overrides. -a alone means --light.
build is the greenfield posture (its plan opens by scaffolding the project). In a repo that already has tracked code it asks whether to switch to feature posture instead — under -a it refuses with the command to run — and --greenfield forces build through for a genuine rebuild.
Configuration
{
"verify": ["npm run typecheck", "npm test"],
"smoke": [], // launch commands for the smoke phase (usually seeded by the planner)
"max_retries": 3,
"max_review_retries": 3,
"max_attempts": null, // absolute retry cap (null = max_retries * 3)
"max_phase_steps": 50, // kill a phase stuck in a tool loop
"verify_timeout_sec": null, // kill a hung verify command (null = 600s)
"stall_timeout_sec": null, // kill a phase with no output (null = 3600s)
"max_step_model_sec": null, // kill a single step stuck in model time (null = 3600s)
"request_ceiling_tokens": null, // largest request a phase may send (null = the model's configured opencode limit.context)
"context_guard": "telemetry", // "kill" stops a phase crossing 95% of its ceiling; default logs only
"sharpen_max_rounds": 6, // interview round cap (0 disables the interview)
"infra_backoff_sec": [60, 300, 900, 1800],
"model": {
"plan": "deepseek/deepseek-v4-flash",
"implement": "deepseek/deepseek-v4-flash",
"review": null, // null = falls back to implement
"visual": null, // vision-capable; null = falls back to review
"goal": null, // null = falls back to visual → review → implement
"extract": null // cheap model, single-shot structured output
},
"code_review": { "mode": "medium", "trigger": "smart" }, // "trigger": "always" reviews every ticket
"visual_review": { "mode": "off", "round_wall_sec": null },
"goal_review": { "mode": "off" }, // add "checkpoint_action": "advisory" for advisory goal checkpoints
"structural_review": { "mode": "off" },
"provider": { // optional: probe the model server before each phase
"base_url": "http://127.0.0.1:8080",
"health": {
"url": "/status", // absolute, or a path resolved against base_url
"timeout_sec": 5,
"pass": { "path": "ready", "equals": true }
}
},
"checkpoint_granularity": "product" // "ticket" | "group" | "product"
}Models are independent — a strong model can plan while a cheap one executes — and null falls back down the chain.
Model seats
| Seat | Tier floor | Notes |
|------|-----------|-------|
| plan | 27B+ | Shapes the ticket graph; don't cheap out |
| implement | 27B+ | Most multi-step reasoning |
| review | 27B+, separate model recommended | Judgment work; should not be weaker than implement |
| goal | 27B+ | Judges the integrated build; shapes remaining work |
| visual | 27B+ with vision | Screenshot review |
| extract | 9B OK | The one seat where a cheap model is endorsed |
Railhead warns (never blocks) when review/goal is weaker than implement, or when any judgment seat parses below 27B.
Reviews
Every gate has a cadence mode: full | medium | light | off. Presets choose defaults; per-gate flags override. The code gate additionally carries a firing trigger — always, or smart (the --light default: review only after context stress). An interactive run (build/fix with no preset and no per-gate flag) asks each gate's cadence — code review, visual, goal, structural — with the light-preset defaults (the code question offers smart).
| Preset | Code review | Visual | Goal | Structural | Interview |
|--------|-------------|--------|------|------------|-----------|
| --full | per-ticket, BLOCKER+MAJOR retry | per-ticket + run-end | checkpoints + run-end | checkpoints + run-end | on |
| --medium | per-ticket, BLOCKER+MAJOR retry | run-end | checkpoints | checkpoints | on |
| --light (default) | smart: per-ticket only after context stress; BLOCKER+MAJOR retry | run-end | corrective checkpoints + run-end | run-end | off |
| --none | off | off | off | off | off |
Notes:
- Smart code review (
--light, the default): the per-ticket review is spent only where the build showed context stress or an unusual event — a compaction, more than one build attempt, a builder-session restart, a spec reconciliation, a$BLOCKED/unverified exit, or a mid-run replan. A clean single-pass ticket skips the review phase;report.mdrecords the coverage ("N run, M skipped") and any skipped ticket a later group/run-end gate flagged, so a green run never implies scrutiny it did not get. A skipped review is "not run", never a pass.--review medium|full(or"trigger": "always"inrailhead.json) reviews every ticket. - Severities:
[BLOCKER]always retries (up to the cap, then hard-fail).[MAJOR]retries through the budget inmedium/full; inlightit gets one corrective attempt, then soft-passes. Minor findings never retry. - Mid-run vs run-end: goal and structural fire at group checkpoints as well as at run end; visual fires per-ticket (under
full) and at run end. When goal review fires at run end it takes visual's whole-app seat — the goal + design-doc frame is stronger. - Corrective tickets:
[BLOCKER]findings generate corrective tickets that run the full gate inline before the originating review may pass. - Visual review runs the app, captures screenshots with a vision model, and judges them against the acceptance criteria.
fixasks the same cadence questions asbuild— a bug fix is not auto-escalated to per-ticket visual. - Vision is measured, not declared. A probe has the seat model read a generated PNG;
build/run/fix/resumerefuse to start a vision gate on a blind model. - Halt: any phase can write
.railhead/STOP(contents = reason) to stop the run for a human.resumerefuses until the file is deleted. - Ctrl-C is a request, not a kill. The first press finishes the ticket in flight's gate and stops at the commit boundary; the second stops immediately. Either way
railhead resumecontinues with no gate left owed.
Live output
implement ── step ────────────────────────────────
implement → bash ls -la [completed exit 0]
implement ✔ tool-calls · in 477, out 123, cache 8304Tool calls, step markers, and errors stream by default, plus a heartbeat with elapsed time, step count, and peak context. --verbose adds the model's reasoning and the exact prompts; --quiet keeps only the heartbeat. railhead log replays any phase afterwards.
Ledger
Every run writes to .railhead/<run-id>/:
state.json— full ticket state, the resume sourceevents/<NN>-<phase>.jsonl— raw opencode event stream per phasereport.md— end-of-run summary with review history and context telemetry
Contracts index
After each commit, a diff-only extract pass updates railhead.contracts.json with the files and symbols that changed. Later tickets receive exact pointers instead of "go explore the codebase" — this is what keeps per-ticket context O(ticket).
Bench
npm run bench runs the engine/model trials and writes bench-results/ (gitignored): a JSON per trial plus a running results.md comparison table, so a future engine change (Splash vs llama.cpp vs oMLX vs LM Studio) is measured the same way every time. Every mode makes real model calls; none of it is part of npm test. Scratch projects default to the system temp dir — keep them outside this repo, or opencode anchors the project at the nearest package.json upward and the model sees railhead's own sources.
npm run bench -- --mode plan --project <dir> [--fresh] --prompt "<goal>" --model <provider/model>
npm run bench -- --mode e2e --project <dir> --tickets <dir> --model <provider/model>
npm run bench -- --mode fixture --model <provider/model>plan— planning only (runPlan); records wall time, ticket count, and the plan phases' first-step prefix cache.e2e— a fullrailhead runon a prepared ticket set; records wall time, commits, goal/visual verdicts, and per-phase first-step cache.fixture— buildsscripts/fixtures/seeded-defect/(a counter app whose tests encode a defect its README forbids) and asserts a gate still names the discrepancy — the judge-quality check.
Setting up opencode
The one critical setting is the per-model context limit — opencode cannot infer it for local models:
// ~/.config/opencode/opencode.jsonc
{
"provider": {
"lmstudio": {
"options": { "baseURL": "http://127.0.0.1:1234/v1", "apiKey": "not-needed" },
"models": {
"prism-ml/bonsai-27b": {
"capabilities": { "limits": { "max_context_window_tokens": 32768, "max_output_tokens": 8192 } }
}
}
}
},
"model": "lmstudio/prism-ml/bonsai-27b"
}Set request_ceiling_tokens in railhead.json to bound the largest single request (prompt + output reserve) a phase may send. Leave it unset and Railhead uses the model's configured opencode limit.context — the effective limit opencode reports, config overrides included. That is where the safety margin belongs: on a warm KV server, set limit.context below the server's hard limit (e.g. 80000 when the hard limit is 100000) and both opencode's compaction and Railhead's ceiling see the same number. Only set request_ceiling_tokens below the window; never mirror the server's full capacity. Crossings are logged, not killed, unless you set context_guard: "kill". After a run, report.md shows whether any ticket's peak context approached the ceiling.
Set finite request timeouts. A provider configured with timeout: false, headerTimeout: false, and chunkTimeout: false lets a wedged request run until Railhead's stall guard ends it (~an hour). Railhead warns at startup when a model one of the run's seats actually uses sits behind that triple — a provider only other tools touch stays quiet. Pair it with the optional provider.health probe above: an unhealthy server then fails a phase in seconds — through the normal failure ladder — instead of stalling. The probe is operator-declared and vendor-neutral (an HTTP URL plus a pass rule over the JSON body); no server is auto-detected, and with no provider block nothing changes. A restarted provider only colds the prompt cache — the base session persists and the next phase re-warms it.
Design principles
- Context is O(ticket), not O(project) — fresh, diff-scoped judging phases plus the auto-extracted contracts index are the seam that makes long unattended runs possible; the builder is one session that compacts.
- No gate overlaps the builder — review phases run in sequence with implementation, so a reviewer never sees a half-edited worktree (ADR 0046).
- Green at every step — every commit passed verify first, so the suite must be a baseline (typecheck, lint, existing tests), never future-feature assertions.
- Verify, then trust — every commit has passed its gate; the ledger is the audit trail.
Decisions and trade-offs are recorded in docs/adr/.
