worksteps
v0.0.2
Published
Durable, agent-driven workflows for work across processes and sessions.
Readme
Overview
Worksteps is a workflow library for agent runners. The runner invokes Worksteps, performs the next available tasks, and continues until the run completes. Worksteps records tasks, leases, and outputs in a durable store so work can resume safely across processes and sessions.
Features
- Durable runs: resume work across processes and sessions from a SQLite database
- Agent tasks: delegate typed prompts or execute commands directly
- Matrices: expand work across values, axes, and repeated trials
- Dependencies: order steps and pass typed outputs downstream
- Concurrency: use fenced leases to prevent duplicate work
- Retries and timing: reclaim expired work, back off, and schedule eligibility
- Small runtime model: invoke a workflow to perform one task
Install
Set up Worksteps:
npx worksteps setupSetup installs the agent plugin where supported and installs the global skill for other agent runners.
After installation, describe the work normally. The skill tells the agent when to use Worksteps, so future prompts do not need to mention it.
Usage
Use Worksteps through your agent
After setup, prompt your agent with the outcome that you want:
/worksteps Audit every package for deprecated APIs and continue until every package has been checked.The agent runner writes a workflow in code mode, invokes it with npx worksteps, performs
delegated tasks, and continues until the workflow completes.
Write a workflow
Create audit.ts:
// audit.ts
import { Run, Workflow, z } from 'worksteps'
Workflow.create('audit', {
inputs: { areas: ['architecture', 'tests'], directory: '.' },
})
.step('inspect', {
matrix: (c) => c.inputs.areas,
output: z.object({ findings: z.array(z.string()) }),
run: (c) => Run.prompt(`Inspect ${c.inputs.directory} for ${c.matrix} problems.`),
})
.step('report', {
needs: ['inspect'],
run: (c) => ({ findings: c.outputs.inspect.flatMap((result) => result.findings) }),
})
.serve()Run a workflow
Invoke the workflow file to start and select a run:
npx worksteps ./audit.tsUse the returned id in subsequent commands. Recover it later with npx worksteps list:
npx worksteps status --id audit_7f3a
npx worksteps claim inspect/architecture --id audit_7f3a
npx worksteps claim inspect/tests --id audit_7f3aThe agent repeats the displayed commands until the workflow reports status: complete. Worksteps
records the workflow path so later commands can reopen the file. Keep the file in place while the
run is active.
Walkthrough
Runners
The value returned by run determines how a task is performed:
run: async () => ({ files: await relevantFiles() }) // the function performed the work
run: () => Run.exec('pnpm', ['test', '--run']) // run an argv without a shell
run: (c) => Run.prompt(`Inspect ${c.matrix}.`) // delegate or call an executorRun.exec captures { code, stdout, stderr }. Exit code 0 succeeds by default. Use ok when a
command uses another exit code as data, and output to transform the capture:
run: (c) =>
Run.exec('git', ['grep', '-l', c.inputs.term], {
ok: (code) => code <= 1,
output: (result) => ({ files: result.stdout.trim().split('\n').filter(Boolean) }),
})Run.prompt requires an output schema so delegated submissions can be validated.
Do not perform effects before returning a descriptor. The run callback is evaluated again after
a retry or reclaim. Workflow execution is at-least-once, so use c.step.id as the idempotency key
for side effects.
See workflow.step, Run.exec, and
Run.prompt.
Matrices and tasks
matrix expands a step into one task per value:
matrix: (c) => c.inputs.areas // c.matrix is one area.step('trial', {
matrix: {
model: ['sol', 'terra'],
task: (c) => c.inputs.tasks,
exclude: [{ model: 'terra' }],
},
repeat: 3,
output: z.object({ passed: z.boolean() }),
run: (c) =>
Run.prompt(`Use ${c.matrix.model} to complete ${taskPrompt(c.matrix.task)}.`),
})Selectors use readable values, such as trial/sol/write-greeting#2. Object values use the first
available identity field among id, slug, key, name, path, file, and title. Use key
to define another stable identity.
Write if and key after matrix. TypeScript infers their matrix value type; placing
them first degrades if, key, and run to unknown.
A matrix can read an upstream output. These tasks remain deferred until the dependency
completes and the values become known.
Dependencies and outputs
needs names the steps that must settle before a step runs and the outputs that it can read:
.step('report', {
needs: ['inspect'],
run: (c) => ({ count: c.outputs.inspect.length }),
})c.outputs.<step> is an array for a matrix step and a single value for a step without a matrix.
Steps not named in needs are not available.
when controls the dependency gate:
'success'waits for successful dependencies and is the default.'always'runs after every dependency settles.'failure'runs when every dependency settles and at least one fails.
Use c.results.<step> with 'always' or 'failure' to inspect each task's status, output,
and error. A terminal dependency failure skips downstream success-gated steps so the run can
settle.
See needs, output,
when, and workflow context.
Repeat until a condition holds
Use until when the number of rounds depends on previous results. The step adds one task at
a time and stops when the predicate holds:
.step('round', {
limit: 12,
output: z.object({ found: z.array(z.string()) }),
until: (c) => quiet(c.results.round) >= 2,
run: (c) => {
const seen = new Set(c.results.round.flatMap((result) => result.output?.found ?? []))
return Run.exec('vitest', ['run', '--reporter=json'], {
ok: () => true,
output: (result) => ({ found: newFailures(result.stdout, seen) }),
})
},
})The current task is excluded from c.results.round. Selectors are round#1, round#2, and
so on. Reaching limit or throwing from the predicate settles the step and records the reason in
the event feed.
Runs and addressing
A run ID derives from the workflow name, definition, resolved inputs, and generation. Starting again with the same inputs creates a new generation without changing earlier rows.
npx worksteps status --id audit_7f3a # inspect the run
npx worksteps claim inspect/tests --id audit_7f3a # perform one task
npx worksteps result --id audit_7f3a # read completed outputsRead stored runs without loading or executing their workflow declarations:
npx worksteps list
npx worksteps show audit_7f3aSee the CLI guide for every command.
Leases, retries, and timing
Each claim holds a fenced lease. An expired lease is treated as a crash and reclaimed on the next claim. A late result from the replaced worker is rejected.
.step('deploy', {
after: '2026-08-10T02:00:00Z',
backoff: '30s',
maxParallel: 1,
retries: 2,
timeout: '15m',
run: () => Run.exec('wrangler', ['deploy']),
})Nothing fires on a clock. after and backoff determine when work becomes claimable, but the
agent runner or its scheduler must invoke the workflow again.
maxParallel limits active leases within one step. Other tasks and steps can still run
concurrently.
See after, backoff,
retries, timeout, and
maxParallel.
Stores and events
Store.sqlite() creates one database per workflow. Stores open
lazily, so --help and --schema do not create a database.
const workflow = Workflow.create('audit', {
store: Store.sqlite({ path: '.worksteps/audit.db' }),
})| Store | Status | Purpose |
| ------------- | ----------- | -------------------------------------------------- |
| SQLite | Available | User-global durable storage through Store.sqlite |
| Cloudflare D1 | Coming soon | Managed SQLite storage on Cloudflare |
| PostgreSQL | Coming soon | Shared storage across processes and hosts |
One Store instance can open many workflow scopes. SQLite maps each scope to a separate file by default. Shared stores such as PostgreSQL and D1 can partition one physical database by the same workflow ID.
The SQLite schema contains runs, declarations, step_deps, steps, events, and selection.
The append-only events table is a changefeed for auditing and future user interfaces.
API Reference
See the API reference and CLI guide.
License
MIT
