llm-agent-factory
v0.1.2
Published
Typed framework for durable LLM agents: pipeline builder, FSM engine, durable checkpoints, and adapters for headless CLI providers (Claude Code, Codex CLI)
Maintainers
Readme
llm-agent-factory
Typed framework for durable LLM agents — a pipeline builder with compile-time contracts, an FSM engine, crash-safe checkpoints, and adapters for headless CLI providers (Claude Code, Codex CLI). Runs on your existing subscription — no API keys.
Subscription-based, not API-key-based. Agents run through the CLI tools you already pay for — Claude Code (Pro/Max) and Codex CLI (ChatGPT Plus/Pro) — driven headlessly. There is no ANTHROPIC_API_KEY/OPENAI_API_KEY to provision and no per-token billing: usage draws on your subscription's limits, and the built-in usage-limit waiting resumes runs when the window resets.
State machine, not prompt chain. You assemble a pipeline from typed nodes; it compiles into a finite state machine that owns all control flow. The LLM reasons only inside LLM nodes — it never chooses transitions. Every step is checkpointed, so a crashed or rate-limited run resumes exactly where it stopped without re-paying for completed LLM calls.
definePipeline() ──► compilePipeline() ──► StateMachine.run()
nodes executors + checkpoints,
(typed contracts) transition table outcome journal,
audit logWhy
- 🧩 Compile-time pipeline contracts — each node declares what it
requiresandprovides(zod schemas + inferred types). Wrong node order does not compile, and the missing key is named in the type error (ContractError<"plan">). - 🛡️ Runtime validation too — the same zod schemas reject corrupted resumes and nodes lying about their output.
- 💾 Durable by design — "at-least-once execution, exactly-once recording": a checkpoint before every node, completed outcomes journaled and replayed on resume, expensive work inside a node split into durable steps.
- 🔁 Retries, gates, and repair loops as primitives — declarative budgets instead of hand-rolled
try/catchwebs; exhaustion escalates to a human, never loops forever. - 🔌 Headless CLI providers — run agents on your Claude Code or Codex CLI subscription (no API key): all CLI flags encapsulated, usage-limit waiting with probing built in.
- 🧪 Testable —
fakeSpawnfor provider tests without real CLIs,directStepRunnerfor node unit tests without the engine. - 🚫 Zero domain logic, zero runtime deps — your app supplies context, seeds, and failure schemas; the only peer dependency is
zod.
Table of contents
- Installation
- Quick start
- Core concepts
- Application extension points
- API overview
- Testing your nodes
- License
Installation
npm install llm-agent-factory zodRequires Node.js ≥ 22. Pure ESM. zod is a peer dependency — node contracts are your zod schemas, so the instance must be shared.
Quick start
import { z } from 'zod';
import {
Claude,
compilePipeline,
createAuditWriter,
defineNode,
definePipeline,
makeRunId,
saveCheckpoint,
StateMachine,
} from 'llm-agent-factory';
// 1. A configured LLM instance — injected into nodes, never picked by them.
const sonnet = new Claude({ model: 'sonnet', cwd: process.cwd() });
// 2. Nodes with dual contracts: types checked at compile time, zod at runtime.
const planNode = defineNode({
name: 'plan',
requires: z.object({ task: z.string() }),
provides: z.object({ plan: z.string() }),
llmLabel: sonnet.label,
async run(ctx) {
const res = await sonnet.invoke({
prompt: `Make a step-by-step plan for: ${ctx.task}`,
mode: 'read-only',
cwd: process.cwd(),
timeoutMs: 10 * 60 * 1000,
});
return { ok: true, provides: { plan: res.resultText } };
},
});
const executeNode = defineNode({
name: 'execute',
requires: z.object({ plan: z.string() }), // ← must be provided upstream, or it won't compile
provides: z.object({ result: z.string() }),
async run(ctx) {
return { ok: true, provides: { result: `did: ${ctx.plan}` } };
},
});
// 3. Assemble and compile: pipeline → executors + declarative transition table.
const pipeline = definePipeline<{ task: string }>('demo', 'task')
.then(planNode)
.then(executeNode)
.build();
const compiled = compilePipeline(pipeline, {});
// 4. Run with durable hooks — every node boundary is checkpointed.
const runId = makeRunId('demo');
const runsDir = `./runs/${runId}`;
const ctx = {
// engine bookkeeping (RunContextCore) …
runId, runsDir, seq: 0, pipeline: pipeline.name,
options: { verbose: false, settingOverrides: {}, settingDefaults: {} },
nodeRetries: {}, retryHints: {}, repairAttempts: 0,
priorAttemptSummaries: [], usedModels: {},
// … plus your seed
task: 'summarize the repo',
};
const machine = new StateMachine(compiled.executors, compiled.transitions, {
saveCheckpoint: (draft) =>
saveCheckpoint(runsDir, { ...draft, pipeline: { name: pipeline.name, hash: pipeline.hash } }),
appendAudit: createAuditWriter(runsDir),
});
const { finalState } = await machine.run(compiled.startState, ctx);Core concepts
Nodes
A node is the unit of work. It carries a dual contract:
- type-level —
Req/Provgenerics inferred from the zod schemas; the pipeline builder accumulates provided keys, so a node whoserequiresisn't satisfied upstream is a compile error naming the missing key; - runtime —
requiresvalidates the incoming context (rejects corrupted resumes),providesvalidates the node's own output before it is folded into the context.
Nodes never reach into config or CLI flags: the engine resolves behavior settings (CLI overrides → node opts → config defaults) and passes them in opts.settings. A node never picks its model either — a configured LLM instance is injected via its factory.
Outcomes
A node returns one NodeOutcome; the compiler turns the semantics into guard edges of the state machine:
| Outcome | What the engine does |
| ------------ | --------------------------------------------------------------------------------------- |
| ok | validate provides with zod, fold into context, go to the next node |
| retry | re-run the same node within its retries budget (detail becomes the next attempt's hint); exhaustion → NEEDS_HUMAN |
| repair | route to the gate's repair node within its budget; after the fix, return to the start of the gate; exhaustion → NEEDS_HUMAN |
| stop | clean completion → DONE |
| needsHuman | escalate → NEEDS_HUMAN |
| fatal | crash → FAILED |
Gates and repair loops
A gate is a chain of check nodes (they validate, they don't extend the context). A failing check routes to the gate's repair node; after a successful fix the loop re-enters the gate from the start:
const pipeline = definePipeline<{ taskId: number }>('build-and-verify', 'task')
.then(fetchNode)
.then(implementNode)
.gate(lintNode, testNode) // checks run in order; any may report `repair`
.repair(fixNode, { budget: 3 }) // at most 3 fix attempts per run
.build();The loop is type-safe by construction: a repair node provides nothing (Prov = {}), so the context entering the gate is identical before and after a fix. The engine guarantees the repair node lastFailure and priorAttemptSummaries.
Reusable segments are defineFragment() — same accumulating contracts, and a fragment can't start with repair or end with an open gate, by construction.
Durability
The discipline is "at-least-once execution, exactly-once recording":
- Checkpoint before every node (v2, atomic tmp+rename, includes the pipeline name and structure hash). An interrupted run resumes at the interrupted node; resuming with a structurally modified pipeline is rejected by hash — changing a model, settings, or budget does not break resume.
- Outcome journal (
outcomes.jsonl): a node'sokoutcome is recorded before it is folded into the context. On resume the completed node is replayed from the record — the LLM call is not paid for twice. Non-ok outcomes are deliberately not journaled. - Durable steps inside a node: split expensive work with
opts.step.do(name, fn)— a crash between steps doesn't repeat completed ones, and the optionalrecordpredicate keeps rate-limited results out of the record. - Audit trail: every transition appended to
transitions.jsonl; bounded retries around durable writes.
LLM instances
An instance is a concrete tool + model + permissions, assembled once and handed to node factories:
import { Claude, Codex } from 'llm-agent-factory';
const opus = new Claude({ model: 'opus', effort: 'high', cwd: repoRoot });
const codex = new Codex({ model: 'gpt-5-codex', permissions: { mode: 'workspace-write' }, cwd: repoRoot });Claudedrives headless Claude Code (claude -p),Codexdrives Codex CLI (codex exec) — both authenticate through the CLI's own login (your subscription), so the framework never touches an API key.- The provider-neutral
LlmPermissionsis translated per adapter (claude → allowed/disallowed tools; codex → sandbox/writable roots);effortlikewise. invokeWithArtifacts()wraps a call with prompt/response artifacts, per-provider session affinity, and a wait-for-usage-limit loop with cheap probing and silent-stall protection.parseStructured()extracts the last```jsonblock and validates it with your zod schema.
Application extension points
The core knows nothing about your domain. You supply:
- Run context —
interface MyContext extends RunContextCore { … }; pin it once with a wrapper:const defineMyNode = (def) => defineNode<…, MyContext>(def). - Pipeline seeds —
definePipeline<Seed>(name, seedKind); the starting context comes from your shell. - Terminal executors —
compilePipeline(pipeline, { terminalExecutors })for summaries and run indexes (noop stubs otherwise). - Failure schema — the engine needs only
RepairFailureBase; your full failure taxonomy stays structurally compatible. - Paths — no assumed repo layout:
runsDir, lessons root, etc. arrive as parameters. - Runs index —
updateRunsIndex<E extends RunsIndexEntryBase>with your own entry fields.
API overview
| Area | Key exports |
| --- | --- |
| Pipeline builder | defineNode, definePipeline, defineFragment, compilePipeline |
| FSM engine | StateMachine, createAuditWriter, TERMINAL_STATES |
| Durability | saveCheckpoint, loadCheckpoint, createOutcomeJournal, createStepRunner, updateRunsIndex |
| LLM instances | Claude, Codex, invokeWithArtifacts, parseStructured, renderTemplate |
| Providers (low-level) | createProviderRegistry, createClaudeProvider, createCodexProvider, runCli |
| Test doubles | makeFakeSpawn, FakeChild, directStepRunner |
| Utilities | makeRunId, writeArtifact, runCommand, withRetries, withFileLock, waitForLimitReset, loadLessons, git helpers |
Everything is exported from the single barrel — import { … } from 'llm-agent-factory' — with full TypeScript declarations.
Testing your nodes
- Node unit tests: call
node.run(ctx, { attempt: 1, settings: {}, step: directStepRunner })— no engine, no filesystem. - Provider tests:
makeFakeSpawnfakes the CLI process, so adapter behavior (flags, parsing, timeouts, rate-limit detection) is testable withoutclaude/codexinstalled.
npm run build # tsc → dist/
npm test # vitest: engine, builder, checkpoint, providersLicense
MIT © George Mkrtchyan
