dynamic-workflow
v0.1.0
Published
Deterministic control flow over non-deterministic agent leaves — a flexible, observable orchestration framework over local coding-agent CLIs.
Maintainers
Readme
Dynamic Workflow
A flexible, observable agent-orchestration framework over local coding-agent CLIs (Claude Code, Codex, OpenCode, Pi). Written in Bun + TypeScript.
What it's for
Orchestration is deterministic control flow over non-deterministic leaves.
Your control flow — phases, barriers, pipelines, loops, fan-out — is ordinary
TypeScript. The only non-deterministic thing is an agent() call: a coding-agent
CLI process whose output you can't predict.
The one capability the built-in Agent/Workflow tools can't offer: mix engines
in a single run. engine is a field on each leaf, so one script fans Claude +
Codex + OpenCode + Pi leaves concurrently.
Install
Bun-only. The kernel is built on Bun.spawn, Bun.FileSink, and
Bun.CryptoHasher — it does not run on Node.
bun add dynamic-workflow # from npm
bun add github:TSHOGX/dynamic-workflow # straight from gitTwo entry points:
import { createRun, agent, phase, barrier, pipeline, loop, withRun } from "dynamic-workflow";
import { adversarialVerify, judgePanel, loopUntilDry } from "dynamic-workflow/recipes";TypeScript source ships as-is (no build step) — Bun imports .ts from
node_modules directly.
The engines are peer tools, not dependencies: a leaf shells out to whichever
CLI its engine names, so claude / codex / opencode / pi must be on
PATH and already authenticated for the engines you actually use. The package
itself has zero runtime dependencies.
Where runs land
Journals land under ~/.dynamic-workflow/runs/. Override the root with
DYNAMIC_WORKFLOW_RUNS_ROOT — per-project run history is just:
export DYNAMIC_WORKFLOW_RUNS_ROOT="$PWD/.runs"Pass workflow and every run keeps its own journal while sharing one resume
cache — the recommended shape:
await runWorkflow({ workflow: "diary" }, body); // run id generated
await runWorkflow({ workflow: "diary", resume: true }, body); // reuses the cache~/.dynamic-workflow/runs/diary/
cache/ shared by every run of "diary"
runs/2026-08-10T09-16-39-487-97b64add/ one journal per execution
runs/2026-08-10T09-16-34-442-e2756de2/That split exists because one identifier cannot do both jobs. leafKey
deliberately excludes the run id, so with a single flat id the id is the cache
identity: reuse it and you keep your cache but overwrite your history (the journal
is truncated on open), randomize it and resume: true silently becomes a no-op
against a cache that is always empty. So workflow is stable and owns the cache —
resume means "redo this workflow, skip the leaves whose inputs didn't change",
which was never about one particular past run — while the run id is unique and owns
the journal. Generated ids are <utc-timestamp>-<uuid8>, so lexical order is
chronological order.
Passing runId alone keeps the original flat <run-id>/ layout, unchanged.
Inspect runs with the bundled CLI (dwf is a shorter alias):
bunx dwf ls # every run, newest first
bunx dwf ls diary # just one workflow's runs
bunx dwf view diary # the most recent run of "diary"
bunx dwf view diary/<run-id> # one exact run
bunx dwf view <run-id> --once # flat run id, one-shot renderAn unrecognized ref lists what is available instead of tailing a path that will never exist.
If you adopt a workflow name that already existed as a flat run id, the old
journal ends up inside the new workflow directory. It is real history, so it stays
listed as <workflow>/legacy rather than disappearing — while a bare <workflow>
always resolves to the workflow's own runs, never to that pre-migration journal
(which can be newer by mtime than the runs that replaced it).
How it runs leaves
One resident Bun process Bun.spawns each CLI and reads its stdout JSON event
stream line-by-line in real time — no tmux, no polling. The session id rides on
the stream, so each run is recorded against its engine's own transcript
(no-clobber history). Per-engine differences (OpenCode's export-from-SQLite vs
the others' per-session JSONL) are erased by the engine adapters.
Shared opencode server pool (on by default)
A standalone opencode run is a full runtime per leaf (~530MB peak, ~12 threads
measured). Fan out many of them and the box saturates CPU context-switching
between fat processes that are mostly idle waiting on a remote model API — lots
of overhead, no useful work. So createRun runs a small pool of long-lived
opencode serve processes and points each opencode leaf at one via
opencode run --attach <url>. The leaf becomes a thin client (~176MB measured)
and the heavy runtime is shared across the run.
The pool starts lazily on the first opencode leaf, so claude/codex-only runs
never spawn a server. It's on by default — just remember to await
backend.close() when the run ends so the servers are torn down. Size or
disable it:
Port reservation. opencode 1.17.x ignores
serve --port 0(it falls back to a fixed 4096), so asize > 1pool would collide and all but one server would die "before announcing a URL". The pool reserves a distinct free port per server (bind:0, release, pass it explicitly) and retries on the rare reserve-then-bind race. Seesrc/backend/server-pool.tsfor the details.
await createRun({ runId, opencodePool: { size: 4 } }); // size the pool
await createRun({ runId, opencodePool: false }); // opt out (per-leaf runtimes)For very high fan-out, the pi engine is a lighter alternative to opencode
leaves — ~142MB vs ~530MB peak per process and ~3.7x faster at 8-way concurrency
(measured), so it scales further on the same box without needing a server pool.
Pure mode (on by default)
Each leaf also sheds local startup weight by default — skipping plugins, hooks,
auto-memory, and user config that trusted orchestration leaves rarely need. The
single pure field on a spec maps to each engine's minimal-startup switch:
| engine | flag | skips |
|---|---|---|
| opencode | --pure | external plugins |
| claude | --bare | hooks, LSP, plugin-sync, auto-memory, CLAUDE.md auto-discovery (keeps session save) |
| codex | --ignore-user-config | ~/.codex/config.toml (auth still uses CODEX_HOME) |
| pi | --no-extensions --no-skills --no-context-files --no-prompt-templates | extension/skill/context/template discovery (keeps session save) |
The trade-off is real: a leaf that needs a hook, the project's CLAUDE.md
conventions, or a custom model provider defined in ~/.codex/config.toml must
opt back into the full environment with pure: false:
agent({ engine: "codex", cwd, prompt, pure: false }); // load my config.tomlEphemeral mode (on by default)
Separate from pure: ephemeral controls whether the leaf's session is
written to disk. It's on by default — trusted orchestration leaves shed
the session-write I/O, which matters under high fan-out. Opt out when you need
the transcript:
agent({ engine: "claude", cwd, prompt, ephemeral: false }); // keep the session for dynamic-workflow view| engine | flag | supported |
|---|---|---|
| claude | --no-session-persistence | yes |
| codex | --ephemeral | yes |
| pi | --no-session | yes |
| opencode | — | no — the leaf still persists; a note.cap warns rather than dropping silently |
ephemeral is spec-level only (no createRun default) and is not part of
the resume cache key — it changes only where output lands, not the leaf's
result. When on, transcriptRef is meaningless (no file is written) even if a
session id still appears on the event stream. Because it defaults on, opencode
leaves warn once each by default (no silent caps); set ephemeral: false to
silence it when you accept persistence.
Sampling params (temperature / topP)
Sampling travels with the workflow: set it once on createRun and every
opencode leaf in the run uses it uniformly.
const { ctx, journal, backend } = await createRun({
runId: "my-run",
sampling: { temperature: 0.9, topP: 0.95 }, // GLM-style, run-wide
});Engine support is uneven, and the framework is honest about it rather than pretending otherwise:
| engine | temperature / topP | how |
|---|---|---|
| opencode | ✅ honored | agent-level config via OPENCODE_CONFIG_CONTENT (no CLI flag exists) |
| claude | ❌ ignored | no flag, env, or setting for it |
| codex | ❌ ignored | reasoning models reject the config key |
| pi | ❌ ignored | no CLI/config knob (upstream declines it — earendil-works/pi#1837 → "build it as an extension") |
How the run-level value reaches opencode:
- Pooled leaves (the default) inherit it from the shared server: the pool
boots its
opencode serveprocesses with the sampling config on every built-in agent, and the model request runs in the server, so every attached leaf uses it. Sampling genuinely lives on the workflow's servers. - Pool-off / standalone leaves get the same config injected as
OPENCODE_CONFIG_CONTENTper process, so behavior is identical either way.
Per-leaf overrides and the other rules:
- Per-leaf override. A leaf may still set
temperature/topPon its own spec; it wins over the run default for that leaf. Because the override must reach the model request (which under--attachruns in the server with only the run-level config), an override leaf automatically runs standalone, skipping the pool. - No silent drop. Set sampling on claude/codex/pi and the leaf still runs,
but the value is ignored and a
note.capwarning is emitted (same no-silent-caps rule as top-N truncation). - Resume. The effective sampling (per-leaf override ?? run-level) is part of the resume cache key — changing temperature re-runs opencode leaves instead of replaying a stale cached result.
Quick start
import { runWorkflow, agent, phase, barrier } from "dynamic-workflow";
const results = await runWorkflow({ runId: "my-run" }, async () =>
phase("survey", () =>
barrier([
() => agent({ engine: "opencode", model: "opencode/north-mini-code-free", cwd: "/tmp", prompt: "…" }),
() => agent({ engine: "claude", model: "sonnet", cwd: "/tmp", prompt: "…" }),
() => agent({ engine: "codex", cwd: "/tmp", prompt: "…" }),
() => agent({ engine: "pi", cwd: "/tmp", prompt: "…" }),
])
)
);runWorkflow owns the run scope: it enters withRun for you and flushes the
journal and closes the backend in a finally — including when your body throws,
in which case the error still propagates. Forgetting backend.close() leaks the
shared opencode server pool's processes, so that teardown belongs in the library,
not in a rule you have to remember.
Use createRun directly when you genuinely need to own the pieces — a long-lived
context, custom sinks, or several runs sharing one backend:
const { ctx, journal, backend } = await createRun({ runId: "my-run" });
try {
await withRun(ctx, async () => { /* … */ });
} finally {
await journal.flush();
await backend.close();
}The engines are peer tools, so a leaf fails clearly if its CLI is missing — checked once per engine, lazily on first use, because probing all four up front would false-alarm anyone who only installed the one they use:
engine "codex" needs `codex` on PATH, but it was not found. Install it or fix
PATH — the engines are peer tools, not dependencies of this package.Watch it live in another pane:
bun run src/cli.ts view my-run # tail-follow; --once for a one-shot renderRun the cross-engine demo:
bun run examples/mixed-fanout.tsRun the dynamic deep-research demo (every fan-out width decided at runtime):
bun run examples/deep-research.ts "your question here"
bun run src/cli.ts view deep-research # watch the dynamic tree fill in livePrimitives
| primitive | what it does |
|---|---|
| agent(spec) | one normalized leaf run. engine is a field → engines mix freely. Non-throwing — check result.status. |
| pipeline(items, ...stages) | default for multi-stage. Items flow independently, no barrier between stages (wall-clock = slowest single chain). A throwing stage drops that item to null. |
| barrier(thunks) | the only named "parallel": run all, await all, null for throwers, never rejects. Use only when you need the whole set (dedup, total-count, compare-all). |
| phase(name, body) | observability grouping — not a barrier. |
| loop(body, {until, maxRounds}) | unbounded discovery/convergence. until: dry(k) / count(n) / budget() / predicate(fn). maxRounds is a required backstop. |
Concurrency is the scheduler's default — emitting many agent() calls runs
them in parallel up to the gates. You never request parallelism; you only mark
the sync point (barrier).
Scheduler & gates
A global gate (default min(16, cores−2)) plus per-engine sub-gates
(claude / codex / opencode / pi all default 8) capped below global — a
leaf needs both a global slot and its engine's slot. A shared token-budget pool
powers budget()-terminated loops.
An engine with no entry in engineLimits gets defaultEngineLimit (8), not
the global limit — a custom engine still gets a real gate rather than being able
to occupy every slot.
Custom engines
engine is an open string, and adapters are supplied per run — so you can add
another CLI (or an HTTP-backed engine) without forking:
import { createRun } from "dynamic-workflow";
import type { Engine, RunEvent } from "dynamic-workflow";
const myCli: Engine = {
name: "my-cli",
// Declare what the engine can actually do. The kernel reads these instead of
// hardcoding engine names, so a custom engine gets the same handling as a
// built-in. Both default to false when omitted — the conservative answer, so an
// ignored request earns a `note.cap` warning instead of vanishing.
supportsSampling: true,
supportsEphemeral: false,
// `ctx.sampling` carries the run-level default; the adapter decides how (and
// whether) to apply it. Returning `env` is how an engine passes config that has
// no CLI flag.
buildArgv(spec, ctx) {
const cmd = ["my-cli", "--json", "--cwd", spec.cwd];
if (spec.model) cmd.push("--model", spec.model);
const temp = spec.temperature ?? ctx?.sampling.temperature;
if (temp !== undefined) cmd.push("--temperature", String(temp));
cmd.push("--", spec.prompt);
return { cmd, promptVia: "arg" };
},
// Map one line of the CLI's stdout onto the normalized event vocabulary.
parseLine(line): RunEvent | null {
const e = JSON.parse(line);
return e.type === "text" ? { kind: "text", text: e.text } : null;
},
};
const { ctx, journal, backend } = await createRun({
runId: "with-my-engine",
engines: [myCli],
scheduler: { engineLimits: { "my-cli": 4 } },
});Extras merge after the built-ins, so passing an adapter named claude
replaces the built-in one. Registration is an explicit option rather than a
module-level registerEngine() on purpose: one mutable registry shared by two
concurrent runs in the same process is the same class of bug as a shared context
slot.
Cancellation
Pass an AbortSignal and the whole run stops — queued leaves never start,
in-flight children are killed:
const ac = new AbortController();
const { ctx, journal, backend } = await createRun({ runId: "r", signal: ac.signal });
// ... later, from a Ctrl-C handler or a cancelled request:
ac.abort();A cancelled leaf resolves with status: "cancelled" rather than throwing, so
agent() stays non-throwing and barrier/pipeline keep their shape. That
status is deliberately distinct from "failed": the leaf never got to answer, so
"we stopped asking" is not the same event as "the model failed" — and a leaf whose
child was killed mid-flight reports cancelled too, not the incidental process
failure.
Cancelled leaves are never written to the resume cache. Caching a non-answer
would make a later resume: true replay it as though it were a real result,
permanently.
Omitting signal changes nothing.
Two kinds of retry
These are deliberately separate:
- Transient / infrastructure retry (in
DirectBackend) — when a spawn fails with a transient signal likeSQLITE_BUSY/database is locked(OpenCode's session-store lock under concurrent spawns), the backend re-spawns with exponential backoff + jitter. It also adds a small pre-spawn jitter (default ≤60ms) so many leaves don't grab the lock in the same instant. This does not consume the user'sretriesbudget — it's below the agent layer. Tune vianew DirectBackend(engines, { spawnJitterMs, transientRetries, transientBackoffMs }). - Schema retry (in
agent()) — re-prompts with a correction when output fails schema validation, consumingspec.retries. A semantic, user-level retry.
Structured output (opt-in)
Off by default — a leaf returns its natural-language .text. Set schema only
when code consumes the result:
const r = await agent({
engine: "opencode", cwd: "/tmp",
prompt: "Name a primary color.",
schema: { type: "object", required: ["color"], properties: { color: { type: "string" } } },
retries: 1,
});
r.structured; // { color: "red" } (validated; retried-with-correction on mismatch)Reading results
agent() never throws — under fan-out one bad leaf must not abort the run, so
failure is a value you inspect:
const r = await agent({ engine: "pi", cwd, prompt, schema: BRIEF });
if (r.status !== "ok" || !r.structured) return null; // the hand-written checkThat line conflates three questions (did it run, is there a payload, is it the
shape I expect) and its null loses the reason. These helpers name the checks
without changing agent()'s contract:
import { structured, structuredOr, textOr, okResults } from "dynamic-workflow";
const o = structured<Brief>(r);
if (!o.ok) console.warn(o.reason, o.error); // "failed" | "timeout" | "cancelled" | "unstructured"
else use(o.value);
const brief = structuredOr<Brief>(r, EMPTY); // fallback, no branching
const prose = textOr(r, "(unavailable)"); // for leaves with no schema
const kept = okResults(await barrier(tasks)); // drops nulls AND failed leaves"unstructured" is worth its own reason: the leaf ran fine but produced nothing
parseable, which points at the prompt rather than the engine. And okResults
exists because barrier yields (T | null)[] and a failed leaf is a non-null
result with a bad status — filtering correctly needs both checks.
Observability & resume
Every state transition is one NDJSON line in runs/<run-id>/journal.ndjson —
the single source of truth. The live view, cost accounting, and resume all
derive from it. createRun({ resume: true }) reuses cached leaf results for
unchanged inputs (keyed by phase/identity/engine/model/prompt + sampling,
where identity is the leaf's cacheKey when set, else its positional index),
so a re-run skips the unchanged prefix instantly. Set spec.cacheKey to a
stable per-task id when a leaf can land at a different position across runs
(varying fan-out, filtered items) — the cache should index "which task is
this", not "the Nth task to run". Keep control flow deterministic (no
Date.now() / Math.random()) or resume keys drift.
Recipes
Quality patterns are compositions of the primitives, not kernel code — see
recipes/: adversarialVerify, perspectiveVerify,
judgePanel, loopUntilDry, completenessCritic, multiModalSweep (parallel
agents each searching a different way — by container / content / entity / time —
blind to each other, merged + deduped, since one angle won't find everything).
Layout
src/
backend/{types,direct}.ts contract + the only backend (Bun.spawn)
backend/server-pool.ts shared `opencode serve` pool (--attach clients)
engines/{opencode,claude,codex,pi}.ts per-engine argv + event normalization
scheduler.ts gates + budget pool
primitives.ts agent/barrier/pipeline/loop/phase
schema.ts opt-in structured output
journal.ts NDJSON writer + result cache + resume keys
view.ts live phase-folded tree renderer
cli.ts `dynamic-workflow view`
index.ts public surface + createRun()
recipes/ quality-pattern cookbook
examples/ mixed-fanout, deep-research, step1-smokeRuns live outside the package (~/.dynamic-workflow/runs by default) so an
installed copy never writes into node_modules.
Tests
bun test # 100 tests across scheduler/primitives/journal/schema/recipes/backend/engines
bun --bun tsc --noEmitTests use mock/scripted backends so they're fast and deterministic; the
examples/ exercise the real CLIs.
One test file is gated behind a real model and credentials — it covers the pooled-leaf SSE path (the exact scenario fake-pool tests can't catch):
OPENCODE_E2E=1 bun test src/backend/attach.e2e.test.ts