@open-tomato/rafa
v0.42.0
Published
Alpha. Runs a plan through the ralph loop: one Claude Code session per task, one commit per task, then a pull request and a CI wait.
Maintainers
Readme
rafa
Run a plan through the ralph loop: one Claude Code session per task, one commit per task, then a pull request and a wait for CI, from one terminal.
What guides it:
- Tools before judgement. Whatever code can decide, code decides: the next task, the branch, the version, whether a pull request is green. The agent is asked for the work itself and little else, which leaves less to chance.
- Repeatable wherever it can be. The steps around the agent are scripts with tests, each run once and in a fixed order, so everything that is not the agent's own writing comes out the same on a second run, on another machine and after a resume.
- Your tools, not ours. rafa drives what you already use (git, the GitHub CLI, GitHub Issues) instead of rebuilding it, and keeps each of those behind a small interface. Swapping one is meant to be an add-on, not a migration: other issue trackers are on the way, starting with Linear, and support for coding agents beyond Claude Code is being specified.
Alpha. rafa is built with rafa, in the open, and it is not finished: commands, files and defaults still change between versions, and the Roadmap below says what is and is not there yet. Feedback and bug reports are welcome at github.com/open-tomato/rafa/issues; the order things are being built in is the pinned Roadmap issue.
Before you run it
Read this once; rafa also says it the first time you start a run. Use
rafa plan risk to see what the plan may do on this machine and under
your accounts.
Every Claude Code session rafa starts runs with
--dangerously-skip-permissions. That is what lets a plan run unattended, and it means a task can edit, delete and run anything your user account can, without asking. There is no switch for it yet. The notice namesloop.settingSourcesas the run resolved it: each session loads its settings from those scopes, and permission rules and hooks in a scope left out do not reach the session. With the default,project,local, the scope left out isuser; with all three loaded, no scope is left out.A run acts under your accounts. It commits, pushes its branch, opens a pull request with
gh, waits for CI, and may file the blockers and unrelated bugs it meets as issues on the project's tracker.It spends your Claude usage. Every Claude Code session rafa starts counts against your Claude plan's limits, or is billed when you run Claude Code with an API key.
budget=on a task caps that task. The commands that start sessions are marked 🪙 in this README:| Command | Spends usage | | --- | --- | |
rafa plan create| 🪙 one planning session | |rafa loop start| 🪙 one session per task, one for the wrap-up and up toloop.wrapUp.retriesmore when it opens no pull request, repair sessions while CI is red, the sessions of up to--retry(loop.retries) more passes of the loop in a row after the stops it retries, and under--continueone read-only decision session per decision | |rafa pr triage --resolve| 🪙 runs a small fixed plan through the loop; without--resolve, nothing | |rafa skill backfill --propose| 🪙 one session per batch of skills; without--propose, nothing | |rafa skill search| 🪙 onehaikusession reading the twelve best-ranked files, one per kind with--all; with--no-model, nothing | |rafa agent search| 🪙 the same asrafa skill search, over agent definitions | |rafa next| 🪙 when the step it runs is one of the above; it asks before each step | |rafa epic close| 🪙 one verification planning session and one session per check; nothing while a member is open | |rafa stretch start| 🪙 one session per operator it starts: the engineer, watchtower and analyst, or the one--rolenames; with--dry-run, nothing | |rafa stretch item| 🪙 one planning session, unless a plan of the issue is there already, and the loop it starts; with--dry-run, nothing |The same 🪙 marks appear in
rafa --helpat all levels and inrafa describeoutput.Everything else reads files, git and GitHub and spends nothing.
rafa loop resumestarts no session itself, but it lets a paused run go on spending. rafa calls no model API directly: all of it goes through theclaudecommand.Your machine may go to sleep in the middle of a run. rafa does not keep it awake, so the operating system's power settings can suspend a loop mid-task. The task carries on when the machine wakes, if its session survived, and its recorded time includes the time asleep, which then reads as an outlier in
rafa effort report --trend. On a machine you leave running loops, turn sleep off in its power settings, or run the loop under the operating system's own sleep lock, which lasts as long as the command it wraps:systemd-inhibit --what=sleep:idle --why="rafa loop" rafa loop start --plan=<plan>caffeinate -i rafa loop start --plan=<plan>The first is Linux's, the second macOS's. On a managed device, check its power policy before you do either: some organisations report, or block, software that keeps a device awake. An opt-in
--keep-awakeflag is planned (#641).So run it in a repository, on a branch and on a machine where all of that is acceptable: a container or a disposable checkout is a good first home. Nothing here is a sandbox.
On a terminal, rafa loop start and rafa plan create print these
notices and ask Continue? [y] yes [d] yes, and do not show this again
[N] cancel. Answering d records it in ~/.rafa/notices.json
({"dismissed": ["alpha", "danger"]}); delete the file to see them
again. Without a terminal they are printed as warnings and the run goes
on.
Install
The package is @open-tomato/rafa on npm (publishConfig names
https://registry.npmjs.org/ for the scope as well as in general, so a
machine that maps @open-tomato to another registry still publishes
there, and npm publish builds first through prepack). It installs
globally under either package manager:
npm i -g @open-tomato/rafabun add -g @open-tomato/rafaEither one puts the rafa binary on PATH. The package declares no
runtime dependency — dependencies, peerDependencies and
optionalDependencies are all absent — so the install resolves nothing
beyond the package itself. The binary and every export run under bun,
which engines names: see Runtime below.
How to use it
The cycle is always the same five steps, and most commands end by naming the next one, so you rarely have to remember it.
Set the project up, once. In the repository you want to work on:
rafa init rafa doctor rafa doctor --deepinitwrites what is missing of.rafa/and~/.rafa/, theirconfig.yamlfiles, and directoriesspecs/,plans/,runs/,effort/,instincts/under.rafa/. On a GitHub repository it also offers to set up the board (labels, the spec issue template, a pinned Roadmap issue).doctorruns every checkloop startruns before its first session, and starts nothing: the prerequisites your config and the plan name,ghand its login when the repository is on GitHub, and the install itself.Write a spec. A spec says what you get, where things stand, the design, what can go wrong, the tasks the plan must carry, and how you will know it is done. Either a file,
.rafa/specs/my-feature.md, or an issue opened from the "Spec" template and labelledspec:ready. docs/specs-and-roadmap.md has the template and a prompt for drafting one.Turn the spec into a plan. One Claude Code session reads the spec and the repository and writes a checklist the loop can parse:
rafa plan create --spec=.rafa/specs/my-feature.md # 🪙 rafa plan create --issue=42 # 🪙 the spec is issue #42's body rafa plan create --next # 🪙 the first undone line of the Roadmap issue rafa roadmap # see the roadmap as a table rafa plan show my-feature # read it before you run itThe plan lands in
.rafa/plans/PLAN-<stub>.md, with aPREREQUISITES-<stub>.mdbeside it when something has to be true before the run. Read both. A plan is plain markdown: edit a task, drop one, add one.Run it.
rafa loop start --plan=.rafa/plans/PLAN-my-feature.md # 🪙On
mainit offers to createfeat/<stub>from the latest base and run there. Then, per task: one Claude Code session, the task's report stored, one commit. A task that reportsblockedis marked and the run stops with the reason; fix what it names and start again, and the blocked task goes first. After the last task a wrap-up session syncs with the base, commits a release fragment when the project has them (the plan's level and notes), pushes, opens the pull request and waits for CI, spending repair sessions on a red one.From another terminal:
rafa loop status,rafa loop pause(after the running task),rafa loop stop(now).Land it and look at what it cost.
rafa pr current # number, title, checks, URL rafa pr triage # why is it red, and is the fix simple (🪙 only with --resolve) rafa pr merge [--skip-checks] # asks y/N, merges, switches to the base, pulls, deletes branches rafa release settle # folds fragments to a version and pushes rafa release tag # tag the released version (optional, depends on config) rafa effort collect && rafa effort reportThen step 2 again, or
rafa plan create --next. Whenrafa nextis used, it runs settle and optionally tag as part of the workflow.
Release fragments and settling
A branch no longer owns a version number. Instead, the wrap-up commits a
release fragment (a small markdown file naming the plan, level and notes),
and rafa release settle assigns the version on the base branch after
merges. This lets several branches merge in any order without a version
collision: whichever machine settles first releases all waiting fragments
under one version.
The pull request shows a forecast of what version the branch would get if
merged now. The fragment is stored under release.fragments (default
.changes/), one file per plan. rafa doctor warns when fragments are
waiting and suggests running settle. rafa pr list marks forecasts as
(base moved) if the base branch changed since the forecast was written.
Every project starts with fragments settling to the base branch and tagging by hand. You can configure settle to open a pending release pull request instead (for protected base branches), and have settle tag the commit automatically (for CI ownership of releases). Four keys, all optional:
# Empty config: fragments in .changes/, semver-by-level, settle pushes, tags by hand
# (This is the default for every new project.)
release:
tag: settle # settle tags the commit it pushed
release:
settle: pr # protected base branch: one pending release PR
pr:
versionCollision: refuse
release:
settle: pr
tag: settle
pr:
versionCollision: refuseThe keys are: release.fragments (path, default .changes/), release.strategy
(fold strategy, default semver-by-level), release.settle (push or pr,
default push), release.tag (manual or settle, default manual),
release.publishCommand (the publish line rafa release tag prints and never
runs, default npm publish). pr.versionCollision controls what happens when a merge would collide
(default report). Read context/release.md for the full design,
docs/ci-release-settle.md for automating settle in CI, and
rafa release settle --help for the settle command.
Tracking effort and learning from skills
Each run records what it cost in time, tokens and work, and which skills were offered, invoked and worked. Key commands for the effort record:
rafa effort collect # gather and store session logs
rafa effort report # per-plan summary: sessions, tasks, costs
rafa effort import <file> # merge another device's store into this onerafa effort report shows the per-plan tables (sessions by status and
outcome), the task report tallies, and any runs whose preflight halted. To
see which skills earned their place and which are ignored:
rafa effort report --skills
rafa effort report --skills --plan=rafa-24-know-which-skills-earnWith --skills, the report shows per plan and per resolver (the method
that picked which skills to offer: planner by task declaration, tag
by ranking, or none):
- Each skill: offered count, invoked count, reported in the outcome,
recurrence count, and its signal:
earning(worked with no failure recurrence),recurring(invoked but a failure recurred),unmeasured(invoked with no failure strings declared), orignored(offered but never invoked). - Each lesson: injected count, recurrence count, and whether it recurred.
- M1 and M2: how many task lines used skills you offered, and how many of the skills you offered were actually invoked.
- Plan CI: the settled result of the pull request's checks when it landed.
The report opens with a fixed line explaining that it shows co-occurrence, not causation: a skill can be invoked and its failure recur for reasons the skill does not cover.
To see whether task sessions are getting more expensive over time, and whether that follows the plan or the calendar:
rafa effort report --trend # last 3 days against the 14 before, then the loops
rafa effort report --trend --loops=10 --by=effortThe trend report has two parts. Trend compares the recent task
sessions with a baseline: median and p90 of minutes, output tokens,
cache-read tokens and turns per task, one row per day, and the sessions
over the outlier line (baseline median plus three robust standard
deviations). Loops lists one row per plan's loop, newest first: its
wall-clock and summed task time, number of tasks, minutes per task (min,
max, average, median), and average tokens and turns per task. Under the
rows, a drift reading tests each per-task figure across the loops in the
order they ran (the Mann-Kendall trend test). no steady trend means the
cost moves with the plan; rising means it grows whatever the plan.
To see everything at once, the loops running now included:
rafa effort collect && rafa effort dashboard
rafa effort dashboard --output=json # one key per widget, for a dashboardThe dashboard prints the running loops with their tasks done over total
and three estimates of the time left: by task (the plan's average task
session), by progress (the time since the plan first started, per task
done), and this session (what rafa loop status prints). By progress
counts the time between tasks too, so it is the one to compare against
the clock. Then come the trend, the loops, each plan's skills M1 and M2,
and the totals of every stored session.
Choosing how the effort store travels
An effort store can travel between a project's devices in different ways, each
suited to different setups. Set effort.sync in .rafa/config.yaml:
| Setup | Strategy | Why |
|---|---|---|
| One laptop | local | Nothing to sync |
| Two personal devices, occasional | file | No server, no account; copy a file |
| Personal projects, frequent | git | History and access from git, works offline |
| A company or several teams | service | One hub, live status, shared rules |
| Advanced, exploratory | p2p | Devices talk directly (spike) |
Configuration examples:
# Default: no sync, each device keeps its own store
# (nothing set, or effort: {sync: local})
# File exchange with no server
effort:
sync: file
# Hub-based sync (requires module and hub.url)
effort:
sync: service
hub:
url: https://hub.example.comOnly local and file ship with rafa; git, service and p2p need a module.
Starting a second device
To run rafa on a second device for the same project, bring the effort store over
consistently using rafa effort copy and rafa effort import:
- On the new device, check for an existing
.rafa/effort/store and set it aside (rename or back up). - From the primary device, make a consistent copy:
Or userafa effort copy --to=<path>sqlite3 .rafa/effort/effort.sqlite ".backup <path>", but never a plain file copy during a write. - Transfer the copied store and
.rafa/config.yamlto the new device's.rafa/directory. The exported file holds private text — task descriptions, findings, and notes — meant only for your own devices. - After both devices have recorded effort, bring them back together with:
rafa effort import <copied-store-file>
Which agents and skills are involved
A task line may end with a declaration, for example
{agent=tdd-guide skills=dev-planner effort=medium}. The planner
writes it; you can change it.
agent=routes the task to a Claude Code subagent. The planner picks by the task's SHAPE: implementation, tests, prose, a red build, a review. Agents and skills come from three tiers in order: project (.claude/agentsand.claude/skills), rafa (bundled with the package inbundled/agentsandbundled/skills), and user (~/.claude/agentsand~/.claude/skills). Earlier tiers shadow later ones; this repository's routing table is incontext/workflow.md, and yours is whatever your project holds.- The rafa tier is served by default. Turn
tiers.rafa: offin.rafa/config.yamlto load only project and user items. Usetiers.agents: {name: false}ortiers.skills: {name: false}to turn a single item off entirely. When two tiers hold different items under the same name, the loop refuses it unless a config pin (for example,tiers.agents: {name: project}) declares which tier to use. A skill pin must name the copy Claude Code loads: it loads a user skill over a project skill, and both over the rafa copy. So arafapin is refused while the project or a loaded user tier holds a different copy, and aprojectpin while a loaded user tier does; pin the tier the refusal names, or delete or rename the copy it names.loop startrefuses a plan that names an agent or a skill it cannot resolve before any session is paid for. - User-tier items are invisible unless
loop.settingSourcesincludesuser.rafa agent listshows what a run sees,rafa agent vendor <name>copies one in or updates it, andrafa skill check .claude/skills --project=.refuses a skill an agent could not follow (a path that does not resolve, a missing field) before it costs a task. - The plan format itself is a skill,
dev-planner, shipped in the rafa tier and used when the project has none of its own. - Skills for this task. A task may declare the skills it needs with
skills=skill1,skill2on its line. At dispatch, one of three resolvers picks which skills to offer:planneruses the task'sskills=declaration exactly, in order;tagignoresskills=and ranks every enabled skill against the task text to find the top 3 scoring above a floor of zero;noneoffers no skills. The resolver is set bytask.skillsin.rafa/config.yaml(defaulting toplanner), and the--skills-resolver=flag overrides it. A compact skill index is listed in the planner's prompt so the planner sees what skills are available before writing the task. When a task is dispatched, the prompt gains a "Skills for this task" section with the skills the resolver picked, or stays unchanged if none were picked. The dispatch record stores the resolver name and which skills were offered. - Lessons from earlier tasks. Up to 5 lessons from the learning store
are added to a task's prompt as a "Lessons from earlier tasks" section
when
task.lessonsison(the default). The lessons are blessed passages from earlier task findings that recurred or reached a confidence floor. When neither skills nor lessons appear, the prompt stays unchanged from before. model=,effort=,tools=andbudget=on a task line set the session's model, reasoning effort, tool list and spending cap when no agent decides them.
Every command has help at three levels (rafa --help,
rafa loop --help, rafa loop start --help), and
rafa describe --output=json is the same roster for a tool or an agent.
Operators (alpha)
Alpha: tested on rafa's own development, and may become a feature. Expect the files and their steps to change between versions.
An operator is a Claude Code agent that runs rafa itself, the way a
person does, instead of being one of the agents a loop hands its tasks
to. rafa ships its operators under bundled/operators/ in the package,
and no loop is ever served one.
The first is the stretch agent, rafa-stretch-engineer. One session
of it is a stretch: it sweeps the board, groups duplicate bugs,
proposes a bucket of up to 10 issues ranked bugs first, and after you
approve the bucket runs it one loop at a time on an integration branch,
stretch/<n>, checking effort, bugs, delivery and conflicts between
items. It stops a second time for you to merge the integration branch
into main, settles one version, and writes a report with the gaps it
hit and what it would change about itself, which waits for your yes. A
second session, rafa-stretch-watchtower, watches it and its loops
without writing anything, and alerts you when one of them needs you.
Nothing is linked into ~/.claude. Each stretch runs on its own copy of
the operators, made under .rafa/stretch/<n>/operators/ when it starts
and loaded with claude --plugin-dir, so a stretch already running is
never changed by a pull or an update of rafa.
Allow what they run without a prompt in the project's
.claude/settings.local.json, and keep every push away from main:
{
"permissions": {
"allow": [
"Bash(rafa:*)",
"Bash(setsid nohup env RAFA_OUTPUT=events rafa loop start:*)",
"Bash(git fetch:*)",
"Bash(git push origin origin/main:refs/heads/stretch/:*)",
"Bash(gh pr create:*)",
"Bash(claude agents:*)",
"Bash(tail:*)",
"Bash(grep:*)",
"Edit(.rafa/config.yaml)",
"Read(~/.claude/projects/**)"
],
"deny": [
"Bash(git push --force:*)",
"Bash(git push -f:*)",
"Bash(git push origin main:*)",
"Bash(git push origin HEAD:main:*)"
]
}
}Then start the engineer, the watchtower and the analyst in one tmux
session, stretch-<project>-<n>, from the project's main checkout, with
Remote Control on so the bucket question and the alerts reach you away
from the machine. From a rafa checkout beside the project:
bash ../rafa/scripts/stretch/stretch.sh start --remote-controlThe agents read loops in the compact output, which you can use on your
own too: RAFA_OUTPUT=events rafa loop start … (or --output=events)
prints one rafa· line per loop event and nothing else. Every other
command prints as text under it.
rafa· task 3/9 start "Group duplicate bugs"
rafa· task 3/9 done 12m 340k tokens
rafa· wrap-up session
rafa· no pr no open pull request for feat/rafa-485Configuration
Settings live in .rafa/config.yaml in the project, and in
~/.rafa/config.yaml for every project on the machine; a setting in the
project's file outranks the same one in yours. rafa init writes both
with every setting commented out at its default, so uncomment a line,
with its section line, to change it.
Tab indentation in .rafa/config.yaml is refused with rafa's own error
on every Bun version, before the parser sees the file. The error names
the line that carries a tab; use spaces only.
effort
One setting says how long a command waits for another rafa process that
is writing the effort store before it gives up with SQLITE_BUSY, and
another names how the store travels between a project's devices.
| Key | Default | What it sets |
|---|---|---|
| effort.busyTimeoutMs | 5000 | the wait in milliseconds, a whole number from 1 to 60000 |
| effort.sync | local | how the store travels: local, file, git, service or p2p |
There is no unlimited wait, since a lock wait with no end can hang a
loop; 0, a fraction and a quoted "5000" are refused.
effort:
busyTimeoutMs: 10000hub
Three settings say how a device reaches a rafa-hub under
effort.sync: service.
| Key | Default | What it sets |
|---|---|---|
| hub.url | unset | the hub's http or https address; required with effort.sync: service |
| hub.tokenSecret | unset | the name the hub token is stored under in the secret store, never the token |
| hub.timeout | 3s | how long one request to the hub may take, whole seconds from 1s to 30s |
A config naming service with no hub.url in either file is refused,
and so is a URL carrying a user name or password. 0s, a negative or
bare number and a timeout past 30s are refused.
effort:
sync: service
hub:
url: https://hub.example.org
tokenSecret: rafa-hub-tokencleanup
Three settings shape what rafa cleanup lists.
| Key | Default | What it sets |
|---|---|---|
| cleanup.staleDays | 30 | the age in days past which a branch is listed as Stale |
| cleanup.worktreeIdleDays | 7 | the idle days past which a worktree is listed |
| cleanup.keep | [] | glob patterns naming branches that are never listed |
Both day counts take a whole number above zero.
Besides the cleanup.keep matches, rafa cleanup never lists the
branch checked out, the base branch (pr.base, else the branch
origin/HEAD names, else main), and the branch origin/HEAD names
even when pr.base names another, so a repository that merges into an
integration branch never sees its default branch offered for deletion.
cleanup:
staleDays: 30
worktreeIdleDays: 7
keep: ["release/*"]status
One setting turns off the line rafa prints on stderr, before a command that runs inside a project, when something is new since the last one.
| Key | Default | What it sets |
|---|---|---|
| status.notice | true | whether that one line, naming rafa status or rafa cleanup, is printed |
It takes true or false as written; a quoted "false" is refused.
status:
notice: falseclaims
Two settings control how claims work when two devices work on one project. Devices claim issues to prevent collisions: no two can plan or start the same issue, and stale or lost claims are recovered without losing work.
| Key | Default | What it sets |
|---|---|---|
| claims.staleAfter | 3d | how long a rafa:claimed claim stands before takeover is allowed: a duration such as 3d, 1w, or disabled |
| claims.ahead | off | whether to reserve one issue ahead on the roadmap: off or allow |
claims.staleAfter takes a duration or the word disabled; 0, a fraction,
and a quoted value are refused. When disabled, nothing is ever stale and
claims must be released by hand. claims.ahead enables claim-ahead as an
opt-in per run with --claim-ahead on rafa next --roadmap and plan create
--next.
# Default: claims are made, staleAfter is 3d, no claim ahead
# (nothing set, or claims: {})
# Set staleness only
claims:
staleAfter: 3d
# Disable staleness and enable claim ahead
claims:
staleAfter: disabled
ahead: allowloop and the guard
A loop guards itself: every turn, it watches the checkout's branch and HEAD and halts if either changes externally (a branch switch in another terminal, a pull that moved the base). The work stays committed and nothing is lost; without the guard the edits would move with the checkout. The guard runs on every loop and never interferes with the loop's own commits.
Worktree use cases. loop start --as-worktree runs the loop in a new
git worktree alongside your main checkout, so you stay on main in your
terminal and can keep working — reviewing, pulling or merging other PRs —
while the loop runs beside you. This is a git worktree, the second desk on
the same repository, with its own branch and its own working directory. Both
share one .rafa/ and one effort store. The worktree is created under
.rafa/worktrees/ by default, or in the directory loop.worktreeDir names
when configured; a run started again with the flag reuses the worktree it
left there. A loop without --as-worktree runs in the current checkout,
which is unchanged. The guard stops a loop if you switch away from it; with
a worktree, you are free to switch the main checkout to anything, and the
loop stays on its branch. Two loops run in two worktrees, each with its own
branch, both under the same .rafa/, allowing parallel work on two epics at
once. Merging a PR while a loop runs no longer offers rafa self-update right
away; pr merge names the running loop and asks you to update after it stops,
since the binary is what the loop spawns its sessions from, and swapping it
mid-run changes the engine while driving.
Edge case: the 2026-09-29 incident. A loop switched the main checkout to
its branch; 48 seconds later, git checkout main and git pull in another
terminal moved it back, and the edits followed. Both branches were at the
same commit, so git switch <loop-branch> recovered it. With the guard the
loop halts instead, keeping the work and the branch where it is while you
sort out what happened.
| Key | Default | What it sets |
|---|---|---|
| loop.worktreeDir | .rafa/worktrees | the directory where --as-worktree creates worktrees, absolute or relative to the project root |
| dangerous.selfUpdateDuringLoop | false | whether rafa self-update runs while a loop of this project is live |
Both take a string and a boolean as written; quoted values are refused.
# empty: loops run in the current checkout, guarded
# --as-worktree puts them in .rafa/worktrees (the default)
loop:
worktreeDir: ../rafa-loops # worktrees beside the repository
dangerous:
selfUpdateDuringLoop: true # update the binary even while a loop runsSpecs, issues and the roadmap
You can plan from a local file and never touch a board. When you want the queue, the specs and their history in one shared place, rafa works from GitHub Issues:
- A spec is an issue, opened from a template with six headings (what
you get, starting position, design, what can go wrong, tasks,
definition of done).
rafa init --boardsets up the template, the labels and a pinned "Roadmap" issue. - The roadmap is one ordered task list in that pinned issue.
rafa plan create --nexttakes the first line that is neither done nor already being worked on, and stops rather than skipping ahead when that issue is not ready.rafa plan create --issue=<n>plans from one issue directly. Epic lines on the roadmap are descended into when they are open,horizon:nowand not done, making their specs the walk's focus;rafa epicsshows one epic's specs as a table. - Epics group specs into features. An epic is a bigger issue with an acceptance criteria, an estimate and an ordered checklist of its specs. Specs carry the epic's label to join it, and rafa never plans across two epics. Projects with no epics work exactly as before; an epic is purely optional.
- A few labels carry the state:
type:spec,spec:ready(a person says a plan may be made from it),spec:needs-work(details pending, or the planner's review found gaps and listed them),type:bugwithneeds-triagefor what a run files on its own,type:epicandepic:<slug>(the epic and its member specs). - Safety, in short. An issue's text ends up in an agent's prompt, so
rafa plans only from an issue whose author can write to the
repository, that a maintainer has labelled
spec:ready, that is complete, and that holds no local path or token. Comments are never read into a plan, rafa never appliesspec:readyby itself, and machine-specific failures a run meets stay off your public tracker. Specs in an epic may not carry twoepic:labels;rafa issue readyrefuses the second one. - Boards and switching. Most projects stay on one board. A project
with several teams can hold one board per team, each with its own
owner and owned folders listed in
Owner:andOwns:lines, and tied to the repository'sCODEOWNERSfile.rafa switch <n>moves to that board's epic and position,rafa board listshows all boards, andrafa statusprints where you stand. A switch made by hand re-homes; pass--no-rehometo keep your home and go back to it later. The position file.rafa/position.jsonholds your current place, where you were, and where you came from — the same two-slot pattern ascd -andgit checkout -, plus a home anchor. - Hopping between epics with
rafa next --roadmap. When your epic's work is blocked by an issue in another epic,rafa next --roadmapreaches for that blocker, works it, and comes home. You hold two slots: home (your anchor) and current (the work in hand). Reaching for a third would mean letting go of home, so rafa halts if a blocker is itself blocked. The pattern keeps unattended work safe: a hop into another team's epic asks for that team's review before it changes their code. Solo projects see a dry hop — the useful part without the permission check. One team hops between its own epics. Several teams open a pull request when a hop crosses to another board, andrafa statussayswaiting on #C (owner review)until the team reviews and approves. The hop record (.rafa/hop.json) tracks where home is and what rafa works. Whenrafa nextfinishes, it deletes the record and returns home. - Board relationships: a choice of modes. A board tracks relationships
— which issues block which — in one of two ways.
labels(the default) usesspec:blockedlabels andBlocked by:lines, carrying no cost when a board is read.nativeuses GitHub's built-in parent and blocked-by links, showing them in GitHub's UI and costing about 8 API points per read of 5,000 per hour. Solo projects and those new to rafa start withlabels. Chosenativewhen the relationships should show in GitHub, or when nothing should clear a blocker by hand. Config it asboard: relationships: native, then runrafa init --boardto move from one mode to the other and remove the old marks. Read docs/specs-and-roadmap.md under "Board relationships" for the full choice and upgrade path. - GitHub project as a board mirror.
rafa init --board --projectcreates a GitHub project in your repository, kept in step with your issues by rafa. The project mirrors the issues' labels, pull requests, close state, and the roadmap checklists: one column per issue (Backlog, Triage, Needs work, Ready, Blocked, Claimed, In development, Waiting for approval, In review, Done, Cancelled), a row per open issue and closed issue with a plan or in review. Five fields track Status, Horizon, Rank, Blocked by and Progress; every command that changes an issue's state refreshes its project row, andrafa board syncrepairs any drift made outside rafa. Optional per repository: runrafa init --boardwithout--projectto skip the project, orrafa board syncadds missing issues and refreshes every item on an existing project. Read context/board-project.md for the full design and the warnings. - Other trackers. GitHub Issues is what works today. Linear support is being ported from the project rafa grew out of, as an optional add-on in a later version. For anything else, open or upvote a request in the issues.
The full guide, with the spec template explained, a prompt for drafting a spec, epics, boards and every gate in order, is docs/specs-and-roadmap.md.
From a checkout
To install dependencies:
bun installTo run the loop from a checkout (no arguments prints the help):
bun src/rafa.ts loop start --plan=<file>Every command but init, the help and describe runs inside a rafa
project: the nearest directory at or above the working directory holding
.rafa/config.yaml. Outside one it prints the rafa init hint and exits
1; bun src/rafa.ts init --yes sets up the git toplevel as one.
bun run build writes dist/, which the rafa bin and the package's
exports point at.
This project was created using bun init in bun v1.3.14. Bun is a fast all-in-one JavaScript runtime.
Running rafa from a snapshot
Run the global rafa from a copy of the build, never from this
checkout's dist/: bun run build opens with rm -rf dist, so a loop
running from dist/ has its runner replaced by the first task that
builds. bun run snapshot in this checkout, or rafa self-update run
inside it, builds, copies dist/ into ~/.rafa/runtime/<version>/ with
the version from package.json, points ~/.rafa/bin/rafa at the copied
cli.js, and prints the path the link resolves to, exiting 0. Both
install the same way (src/runtime/install.ts). The bin is not in
~/.bun/bin, where bun link in this checkout re-points rafa at the
checkout's dist/ with no message, so put ~/.rafa/bin on PATH ahead
of ~/.bun/bin; both warn when it is not, and rafa doctor checks it.
Both exit 1 before building anything while a plan tracker in plan.dir
(.rafa/plans unless .rafa/config.yaml names another) still holds an
open or blocked task, and name every such tracker. rafa self-update
runs only inside a project, so rafa init the checkout first, and it
also exits 1 while a loop of the project is live, naming each loop's
branch and pid, unless dangerous.selfUpdateDuringLoop is true in the
config; --force does not override that wait. They exit
2 when they could not run: a package.json that is not rafa's or cannot
be read, a config or tracker they could not read, a loop record
rafa self-update could not read, or a build, copy or link that failed.
A version is installed once. Both also exit 1, before building, when
~/.rafa/runtime/<version>/ is already there, naming that directory and
the version package.json gave: a loop may be running from it, and a
second build under the same version would swap its runner. Raise the
version to install beside it. bun run snapshot --force and
rafa self-update --force install over it anyway, replacing the
directory whole, so nothing the old build left is kept: the replacement
is copied beside the directory and renamed into its place, and only then
is the old one removed, so no reader meets a half-replaced runtime. The
link lands by a rename too, so a shell meets the old link or the new one
and never none.
Bringing a project up to the installed rafa
rafa self-update moves the rafa every project on the device runs, and
leaves what each project holds as it was set up. rafa update current,
run inside a project, brings that project to the installed rafa: it
creates the .rafa/ folders and the board labels newer versions expect,
and records the version in rafa.lock at the repository root, a small
JSON file meant to be committed. It works within a patch range: the
installed version must share the major and minor of the one the lock
records. Below 1.0.0 it also crosses newer minors, so a lock at 0.34.1
moves to 0.36.0, and a project with no lock is adopted. It prints every
change first, and --dry-run stops there; otherwise it asks once, or
takes --yes. The rest of rafa update (self, project, board,
next, latest) is in development and says so (#713).
Runtime
The build targets bun, and engines names bun alone: there is no
engines.node, because no node version runs the binary. The package
root, ./cli and ./store import bun:sqlite, which node's ESM loader
refuses before any module code runs, so only ./plan, ./ports and
./learning load under node at all. module points at
./dist/index.js, the same build the root of exports names. The package ships no type declarations:
exports names no types, and a TypeScript consumer gets TS7016 under
strict.
Roadmap
What rafa does today and what is planned, in the order it is being built. ✅ is shipped, ⬜ is next; the order and each issue's state live in the pinned Roadmap issue, and a line here is ticked by the change that finishes the feature.
- ✅ Run a plan task by task: one Claude Code session per task, one commit per task, then a pull request and a wait for CI
- ✅ Plans that carry their own background, so each task reads only the part of the plan it needs
- ✅ Every task reports back what it did, what it found and what blocked it, and rafa keeps the record
- ✅ See what each plan cost: sessions, tokens and commits per plan
- ✅ Install once, set up any project with
rafa init - ✅ Checks before a run: a missing tool or key stops the run with its name, before any session is paid for
- ✅ A spending cap per task
- ✅ One consistent command line, with help at every level that people and agents can both read
- ✅ Stop, pause, resume and check on a running plan
- ✅ Blockers and unrelated bugs found along the way are filed as issues, once each
- ✅ rafa can safely work on its own code and update itself
- ✅ Usable as a library inside other services, not only as a command
- ✅ A health check for skills: one format, and a checker that refuses a broken skill before an agent can follow it
- ✅ The agents a plan needs are checked before the run, and copied into the project with one command
- ✅ Plan straight from the issue board, or from whatever is next on the roadmap
- ✅ Review, merge and clean up pull requests from the command line
- ✅ A failing pull request is diagnosed, and fixed when the fix is simple
- ✅ Release fragments on branches, settled to a version on the base branch after merges, with forecasts in the PR and protection for multiple concurrent branches
- ✅ Start a plan from the main branch and rafa makes the branch for you
- ✅ One command takes you to the next step: merge, clean up, plan, branch, start
- ✅ Every command that spends Claude usage says so in its help
- ✅ Before a run, see what it can do on this machine and under your accounts
- ✅ See every skill and agent a run would use, browse them, and ask about them in plain language
- ✅ The roadmap in one table: what is ready, what blocks it, and what already has a plan, a branch or a pull request
- ✅
rafa doctor --deep: what a loop session and its subagents can actually reach — settings,PATH, providers and the tools its stack needs - ✅ Verified and stamped references in specs: each reference the spec names is extracted by pattern, verified against its target, and fingerprinted so changes are caught when the spec is refreshed
- ✅ Clean up merged, stale and unpushed branches and idle worktrees
- ✅
rafa status: everything in one snapshot, and one line about what changed since you last looked - ✅ The right skills reach the right task, chosen when the plan is written
- ✅ Know which skills earn their place and which are ignored
- ✅ rafa learns from its own runs: what one task works out is handed to the tasks that need it later
- ✅ Every change to an epic is a command that leaves a trail:
rafa epic new,defer,promote,move,closewith a verification gate,cancel, and the end-of-epic lines inrafa next - ⬜ Skills and lessons shared across projects and machines
- ⬜ Config as code: a typed
rafa.config.tsholding your settings, your passes and your flows, with today's behaviour as the default - ⬜ Every check rafa runs has a class you can see, and your workflow can move the rest: one question per decision, and a dry run that walks the whole flow
- ⬜ Every outside call doubled in tests, every outcome produced, and each past incident kept out for good
- ⬜ Find skills and agents that overlap or contradict, and refine one without losing the original
- ⬜ A retrospective: evidence, independent conclusions, a ranked action plan
- ⬜ Work on several issues or specs at the same time
- ⬜ Team retrospective and a project status check in server mode
- ⬜ Feedback from outside projects reaches the rafa board through triage
- ⬜ Add-ons: install a tracker, an output or a set of skills (Linear, Obsidian and others) without changing rafa
- ⬜ A live terminal dashboard
- ⬜ Change how rafa works without changing rafa: settings, prompts and steps live in your project
- ⬜ Other coding agents: run a plan without Claude Code
- ⬜ Run a plan with enforced permissions instead of
--dangerously-skip-permissions - ⬜
rafa doctor --security: an outside scan of your Claude Code setup, with what rafa itself does stated first
Attribution
rafa stands on other people's work and on earlier work of ours. NOTICE carries the licences; this is the story.
- The Ralph technique. Running an agent in a plain loop, one fresh session per step over a plan kept on disk, is the "Ralph" technique described by Geoffrey Huntley in Ralph Wiggum as a "software engineer". The name "ralph loop" in this project is a nod to it. What rafa adds around the loop (plans with declared routing, structured task reports, preflight, effort records, the pull request and CI stage) is ours; the idea of the loop is not.
- Open Tomato. rafa's first loop, its issue-tracker port, its CLI event format and the learning design below come from projects in the open-tomato organisation. Most of them are not public yet; this section will link the specific repositories as they are published.
- Loop implementation. The code rafa started from was imported from
marcostomatti/template-agentic-research(Apache-2.0), where the loop had grown its plan and report parsers, effort collection and tracker. - Instinct model. The instinct record (trigger, action, confidence,
evidence, scope) is adapted from the
continuous-learning-v2skill ofaffaan-m/everything-claude-codeby Affaan Mustafa (MIT). - Shared learning. The design for merging what separate runs learn
(one record per lesson, a rule for conflicting lessons, promotion on
recurrence) follows Open Tomato's hive-learning design. rafa now records
lessons in
.rafa/instincts/— one lesson per task finding that carries a resolution — and pushes them after each task's report. Lessons that meet a confidence floor are blessed and handed to tasks that need them later; lessons that recur enough are promoted into the pages that own their subjects. Sharing lessons across projects and machines is planned for a future version.
License
Apache-2.0; see LICENSE. NOTICE names the works rafa builds on and their licences. rafa drives Claude Code and is not affiliated with or endorsed by Anthropic.
