npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

llm-agent-factory

v0.1.2

Published

Typed framework for durable LLM agents: pipeline builder, FSM engine, durable checkpoints, and adapters for headless CLI providers (Claude Code, Codex CLI)

Readme

llm-agent-factory

Typed framework for durable LLM agents — a pipeline builder with compile-time contracts, an FSM engine, crash-safe checkpoints, and adapters for headless CLI providers (Claude Code, Codex CLI). Runs on your existing subscription — no API keys.

npm version license: MIT node >= 22 types: TypeScript

Subscription-based, not API-key-based. Agents run through the CLI tools you already pay for — Claude Code (Pro/Max) and Codex CLI (ChatGPT Plus/Pro) — driven headlessly. There is no ANTHROPIC_API_KEY/OPENAI_API_KEY to provision and no per-token billing: usage draws on your subscription's limits, and the built-in usage-limit waiting resumes runs when the window resets.

State machine, not prompt chain. You assemble a pipeline from typed nodes; it compiles into a finite state machine that owns all control flow. The LLM reasons only inside LLM nodes — it never chooses transitions. Every step is checkpointed, so a crashed or rate-limited run resumes exactly where it stopped without re-paying for completed LLM calls.

definePipeline() ──► compilePipeline() ──► StateMachine.run()
     nodes              executors +            checkpoints,
  (typed contracts)   transition table       outcome journal,
                                              audit log

Why

  • 🧩 Compile-time pipeline contracts — each node declares what it requires and provides (zod schemas + inferred types). Wrong node order does not compile, and the missing key is named in the type error (ContractError<"plan">).
  • 🛡️ Runtime validation too — the same zod schemas reject corrupted resumes and nodes lying about their output.
  • 💾 Durable by design — "at-least-once execution, exactly-once recording": a checkpoint before every node, completed outcomes journaled and replayed on resume, expensive work inside a node split into durable steps.
  • 🔁 Retries, gates, and repair loops as primitives — declarative budgets instead of hand-rolled try/catch webs; exhaustion escalates to a human, never loops forever.
  • 🔌 Headless CLI providers — run agents on your Claude Code or Codex CLI subscription (no API key): all CLI flags encapsulated, usage-limit waiting with probing built in.
  • 🧪 Testable — fakeSpawn for provider tests without real CLIs, directStepRunner for node unit tests without the engine.
  • 🚫 Zero domain logic, zero runtime deps — your app supplies context, seeds, and failure schemas; the only peer dependency is zod.

Table of contents

Installation

npm install llm-agent-factory zod

Requires Node.js ≥ 22. Pure ESM. zod is a peer dependency — node contracts are your zod schemas, so the instance must be shared.

Quick start

import { z } from 'zod';
import {
  Claude,
  compilePipeline,
  createAuditWriter,
  defineNode,
  definePipeline,
  makeRunId,
  saveCheckpoint,
  StateMachine,
} from 'llm-agent-factory';

// 1. A configured LLM instance — injected into nodes, never picked by them.
const sonnet = new Claude({ model: 'sonnet', cwd: process.cwd() });

// 2. Nodes with dual contracts: types checked at compile time, zod at runtime.
const planNode = defineNode({
  name: 'plan',
  requires: z.object({ task: z.string() }),
  provides: z.object({ plan: z.string() }),
  llmLabel: sonnet.label,
  async run(ctx) {
    const res = await sonnet.invoke({
      prompt: `Make a step-by-step plan for: ${ctx.task}`,
      mode: 'read-only',
      cwd: process.cwd(),
      timeoutMs: 10 * 60 * 1000,
    });
    return { ok: true, provides: { plan: res.resultText } };
  },
});

const executeNode = defineNode({
  name: 'execute',
  requires: z.object({ plan: z.string() }), // ← must be provided upstream, or it won't compile
  provides: z.object({ result: z.string() }),
  async run(ctx) {
    return { ok: true, provides: { result: `did: ${ctx.plan}` } };
  },
});

// 3. Assemble and compile: pipeline → executors + declarative transition table.
const pipeline = definePipeline<{ task: string }>('demo', 'task')
  .then(planNode)
  .then(executeNode)
  .build();

const compiled = compilePipeline(pipeline, {});

// 4. Run with durable hooks — every node boundary is checkpointed.
const runId = makeRunId('demo');
const runsDir = `./runs/${runId}`;

const ctx = {
  // engine bookkeeping (RunContextCore) …
  runId, runsDir, seq: 0, pipeline: pipeline.name,
  options: { verbose: false, settingOverrides: {}, settingDefaults: {} },
  nodeRetries: {}, retryHints: {}, repairAttempts: 0,
  priorAttemptSummaries: [], usedModels: {},
  // … plus your seed
  task: 'summarize the repo',
};

const machine = new StateMachine(compiled.executors, compiled.transitions, {
  saveCheckpoint: (draft) =>
    saveCheckpoint(runsDir, { ...draft, pipeline: { name: pipeline.name, hash: pipeline.hash } }),
  appendAudit: createAuditWriter(runsDir),
});

const { finalState } = await machine.run(compiled.startState, ctx);

Core concepts

Nodes

A node is the unit of work. It carries a dual contract:

  • type-level — Req/Prov generics inferred from the zod schemas; the pipeline builder accumulates provided keys, so a node whose requires isn't satisfied upstream is a compile error naming the missing key;
  • runtime — requires validates the incoming context (rejects corrupted resumes), provides validates the node's own output before it is folded into the context.

Nodes never reach into config or CLI flags: the engine resolves behavior settings (CLI overrides → node opts → config defaults) and passes them in opts.settings. A node never picks its model either — a configured LLM instance is injected via its factory.

Outcomes

A node returns one NodeOutcome; the compiler turns the semantics into guard edges of the state machine:

| Outcome | What the engine does | | ------------ | --------------------------------------------------------------------------------------- | | ok | validate provides with zod, fold into context, go to the next node | | retry | re-run the same node within its retries budget (detail becomes the next attempt's hint); exhaustion → NEEDS_HUMAN | | repair | route to the gate's repair node within its budget; after the fix, return to the start of the gate; exhaustion → NEEDS_HUMAN | | stop | clean completion → DONE | | needsHuman | escalate → NEEDS_HUMAN | | fatal | crash → FAILED |

Gates and repair loops

A gate is a chain of check nodes (they validate, they don't extend the context). A failing check routes to the gate's repair node; after a successful fix the loop re-enters the gate from the start:

const pipeline = definePipeline<{ taskId: number }>('build-and-verify', 'task')
  .then(fetchNode)
  .then(implementNode)
  .gate(lintNode, testNode)          // checks run in order; any may report `repair`
  .repair(fixNode, { budget: 3 })    // at most 3 fix attempts per run
  .build();

The loop is type-safe by construction: a repair node provides nothing (Prov = {}), so the context entering the gate is identical before and after a fix. The engine guarantees the repair node lastFailure and priorAttemptSummaries.

Reusable segments are defineFragment() — same accumulating contracts, and a fragment can't start with repair or end with an open gate, by construction.

Durability

The discipline is "at-least-once execution, exactly-once recording":

  • Checkpoint before every node (v2, atomic tmp+rename, includes the pipeline name and structure hash). An interrupted run resumes at the interrupted node; resuming with a structurally modified pipeline is rejected by hash — changing a model, settings, or budget does not break resume.
  • Outcome journal (outcomes.jsonl): a node's ok outcome is recorded before it is folded into the context. On resume the completed node is replayed from the record — the LLM call is not paid for twice. Non-ok outcomes are deliberately not journaled.
  • Durable steps inside a node: split expensive work with opts.step.do(name, fn) — a crash between steps doesn't repeat completed ones, and the optional record predicate keeps rate-limited results out of the record.
  • Audit trail: every transition appended to transitions.jsonl; bounded retries around durable writes.

LLM instances

An instance is a concrete tool + model + permissions, assembled once and handed to node factories:

import { Claude, Codex } from 'llm-agent-factory';

const opus = new Claude({ model: 'opus', effort: 'high', cwd: repoRoot });
const codex = new Codex({ model: 'gpt-5-codex', permissions: { mode: 'workspace-write' }, cwd: repoRoot });
  • Claude drives headless Claude Code (claude -p), Codex drives Codex CLI (codex exec) — both authenticate through the CLI's own login (your subscription), so the framework never touches an API key.
  • The provider-neutral LlmPermissions is translated per adapter (claude → allowed/disallowed tools; codex → sandbox/writable roots); effort likewise.
  • invokeWithArtifacts() wraps a call with prompt/response artifacts, per-provider session affinity, and a wait-for-usage-limit loop with cheap probing and silent-stall protection.
  • parseStructured() extracts the last ```json block and validates it with your zod schema.

Application extension points

The core knows nothing about your domain. You supply:

  1. Run context — interface MyContext extends RunContextCore { … }; pin it once with a wrapper: const defineMyNode = (def) => defineNode<…, MyContext>(def).
  2. Pipeline seeds — definePipeline<Seed>(name, seedKind); the starting context comes from your shell.
  3. Terminal executors — compilePipeline(pipeline, { terminalExecutors }) for summaries and run indexes (noop stubs otherwise).
  4. Failure schema — the engine needs only RepairFailureBase; your full failure taxonomy stays structurally compatible.
  5. Paths — no assumed repo layout: runsDir, lessons root, etc. arrive as parameters.
  6. Runs index — updateRunsIndex<E extends RunsIndexEntryBase> with your own entry fields.

API overview

| Area | Key exports | | --- | --- | | Pipeline builder | defineNode, definePipeline, defineFragment, compilePipeline | | FSM engine | StateMachine, createAuditWriter, TERMINAL_STATES | | Durability | saveCheckpoint, loadCheckpoint, createOutcomeJournal, createStepRunner, updateRunsIndex | | LLM instances | Claude, Codex, invokeWithArtifacts, parseStructured, renderTemplate | | Providers (low-level) | createProviderRegistry, createClaudeProvider, createCodexProvider, runCli | | Test doubles | makeFakeSpawn, FakeChild, directStepRunner | | Utilities | makeRunId, writeArtifact, runCommand, withRetries, withFileLock, waitForLimitReset, loadLessons, git helpers |

Everything is exported from the single barrel — import { … } from 'llm-agent-factory' — with full TypeScript declarations.

Testing your nodes

  • Node unit tests: call node.run(ctx, { attempt: 1, settings: {}, step: directStepRunner }) — no engine, no filesystem.
  • Provider tests: makeFakeSpawn fakes the CLI process, so adapter behavior (flags, parsing, timeouts, rate-limit detection) is testable without claude/codex installed.
npm run build   # tsc → dist/
npm test        # vitest: engine, builder, checkpoint, providers

License

MIT © George Mkrtchyan