@demystify/agent-harness
v0.1.0
Published
The loop that drives an agent-kernel plan to its end, resumably, on a host that can be killed at any moment. At-least-once by design: the cursor advances only on a recorded success, and every step carries a stable `runId:stepId` idempotency key so the hos
Maintainers
Readme
@demystify/agent-harness — the loop that drives a plan, on a host that can be killed
@demystify/agent-kernel plans and records. It is structurally incapable of
sending, and that stays. This is the layer that executes: it asks the kernel
for the next step, invokes the host's handler, and records what came back.
Zero runtime dependencies — not even on the kernel, whose types are mirrored rather than imported. Node ≥ 22, ESM, TypeScript strict. No keys, no database, no env vars, nothing to mock.
Why this exists rather than a for loop
Because of two properties a for loop does not have, both learned the expensive
way on Vercel.
1. It refuses to START a step it does not expect to finish
A serverless platform kills a function at its limit. It does not ask where you are. If the kill lands mid-step, that step re-runs on the next tick with a half-applied effect behind it — the row written, the outcome unrecorded, the ledger disagreeing with the world.
Stopping cleanly between steps costs nothing at all.
a `for` loop with a timeout this harness, with deadlineAt
─────────────────────────── ────────────────────────────
s1 ✔ ─ s2 ✔ ─ s3 ▓▓▓ KILL s1 ✔ ─ s2 ✔ ─ ⏸ paused
↑ ↑
killed mid-step: the WhatsApp the deadline refused to start
message went out, the outcome s3. Nothing is in flight. The
was never recorded, and the next next tick resumes at s3 — an
tick sends it again. ordinary resumption, not a repair.const result = await executeRun(kernel, runId, declaration, {
handlers,
deadlineAt: Date.now() + 50_000, // under a 60s function limit
});
result.status; // "paused" — the run is still running, and still resumablepaused is this package's word, not the kernel's: the run never left running
and its claim is intact. The next tick claims the same (agent, tenant, period),
is told acquired: false, and continues the run it was handed.
If your steps are slow and roughly predictable, reserve time for them:
{ deadlineAt: Date.now() + 50_000, reserveMs: 8_000 }
// with 5s left, a step that usually takes 8s is NOT started2. Every step carries a stable idempotency key
The cursor advances only when a step is recorded as succeeded. A process killed after a step was handed out but before its outcome was written leaves the cursor where it was — so the next tick is handed the same step again.
That is at-least-once, and it is deliberate. The alternative loses work silently. Exactly-once across a process boundary is a lie; at-least-once with a stable key is the honest version of the same promise.
So every handler is given the thing it needs to be idempotent:
ctx.idempotencyKey; // "run_7f3a:notify-owner" → `${runId}:${stepId}`Identical across every re-execution of that step, and different for every other step in the product. Derived, never generated: no clock, no counter, no uuid — a key drawn from any of those is a different key on the retry, which is the same as having no key at all. A handler that writes a row keys on it. A handler that ignores it will double-write, and that is a review question rather than a silent one.
Install
npm install @demystify/agent-harnessRun an agent loop with nothing else installed
No database, no keys, no kernel. This file runs as-is:
import { executeDag } from "@demystify/agent-harness";
const steps = [
{ id: "fetch", kind: "invoice.list" },
{ id: "score", kind: "risk.rank", estimateMinor: 400 },
{ id: "draft", kind: "message.draft", estimateMinor: 300 },
];
// The scheduler is YOURS. Here: a chain. In production, usually
// `(state) => readyStepIds(plan, state)` from @demystify/workflow.
const order = ["fetch", "score", "draft"];
const ready = (state: { succeeded: readonly string[] }) => {
const next = order.find((id) => !state.succeeded.includes(id));
return next ? [next] : [];
};
const result = await executeDag({
runId: "run_7f3a",
steps,
ready,
deadlineAt: Date.now() + 50_000,
handlers: {
"invoice.list": (ctx) => console.log("fetching", ctx.idempotencyKey),
"risk.rank": () => ({ spentMinor: 380 }),
"message.draft": () => {
// ↑ YOUR function. This is where a send would live, if there were one.
return { spentMinor: 290 };
},
},
});
result.status; // "succeeded" | "failed" | "blocked" | "halted" | "paused"
result.state; // { succeeded: [...], failed: [...] } — persist this to resume
result.spentMinor;result.state is the whole of what has to be persisted. Pass it back in and the
run continues where it stopped.
With the kernel, for a ledger that survives the process
import { AgentKernel, MemoryRunStore } from "@demystify/agent-kernel";
import { executeRun } from "@demystify/agent-harness";
const kernel = new AgentKernel({ store: new MemoryRunStore() });
const declaration = { agent: "collections", capabilities: [] };
await kernel.claimRun({
runId: "run_7f3a",
agent: "collections",
tenantKey: "org_abc", // opaque; the kernel never parses it
period: "2026-08-19",
plan: { steps },
ceiling: { amountMinor: 5_000, currency: "INR" },
});
const result = await executeRun(kernel, "run_7f3a", declaration, {
handlers,
deadlineAt: Date.now() + 50_000,
});The harness takes a structural KernelLike — nextStep, recordStep,
cancel, getRun. An AgentKernel satisfies it, and so does your own ledger.
There is no import between the two packages and no version lock: the kernel's
types are mirrored in src/types.ts, exactly as @demystify/workflow mirrors
its Step. That the mirror still fits was checked with tsc against
@demystify/[email protected]'s published declarations, not assumed.
What it guarantees
| # | Guarantee | How |
|---|---|---|
| 1 | A step that would exceed the deadline is never STARTED | deadlineAt (+ optional reserveMs) is checked before the handler is invoked, not after. The kill lands between steps. |
| 2 | The run stays resumable after a pause | Status is still running, the cursor has not moved, nothing is in flight. |
| 3 | At-least-once, honestly | The cursor advances only on a recorded success. Nothing here pretends to be exactly-once. |
| 4 | runId:stepId, identical on every re-execution | Derived from two arguments. test/no-transport.test.ts forbids every source of entropy in src/. |
| 5 | A failed step does not advance the cursor | It is recorded as failed, with the reason the handler threw, and the run halts there. |
| 6 | A failed ledger WRITE is not a failed step | It propagates, leaving the run resumable — see "changed from the original". |
| 7 | A kind with no handler fails CLOSED | Never skipped. Skipping lets a partial plan report success. |
| 8 | Spend is what the handler reports, never the estimate | And a step that overshot its estimate is named in overspentSteps. |
| 9 | Money is integer minor units | No float exists in the code. |
| 10 | Tenancy is an opaque string | The harness never parses it; it never sees it. |
| 11 | No clock in the decision path but the injected one | Date is named in exactly one file, and it is a reference, not a call. |
| 12 | Structurally incapable of sending | See below. |
Audit and approvals: ports, not dependencies
@demystify/audit-chain and @demystify/approvals fit these shapes. So does an
array, a console.log, or a Postgres INSERT in your own code. This package
depends on neither, which is what keeps it installable on its own.
await executeRun(kernel, runId, declaration, {
handlers,
audit: { record: (event) => chain.append(event) },
approvals: { check: async (req) => await askAHuman(req) },
});An approval has three answers and no fourth:
| decision | what happens |
|---|---|
| approved | the step runs |
| pending | the run pauses, resumable — cursor untouched, handler never invoked. The next tick asks again. |
| denied | linear: the run is cancelled with the approver's reason (not a step failure). DAG: the step is marked failed, so a when: "failed" compensation edge can fire. |
The audit port fails open: a throwing sink never aborts a run that is
otherwise fine. It is not silent about it either — every failure comes back on
result.auditFailures. Fail closed on safety, fail open on telemetry.
What this is NOT
Not an LLM agent loop. There is no model call here, no prompt, no provider,
no retry-on-refusal. It drives a kernel plan. If a step calls a model, that
happens in your handler, in your file. The test suite fails the build if a
provider name ever appears in src/.
Not a transport. Same rule as agent-kernel, for the same reason:
There is no sender here. No
send, nodispatch, no chokepoint.
A chokepoint is still a place where sending happens, so it is still a target. The stronger property is that nothing here holds a reference through which anything can leave the process.
One assertion in the no-transport suite is deliberately inverted, and it is
documented in the file. The kernel forbids any export that accepts a callable,
because for a package that only plans and records, a callback is a transport.
The harness cannot inherit that: driving a host-supplied handler is the entire
product. So — exactly as @demystify/evals did for its candidate function — the
assertion is replaced by its inverse:
- the handler must be required, never defaulted: this package never supplies a function of its own for a step to run;
- a kind with no handler must fail the run, never be skipped;
- everything else still holds — no
fetch, no socket, no provider SDK, no endpoint, no credential, no export whose name implies an outbound effect.
Not exactly-once. See above. Anyone promising it across a process boundary is selling something.
Not a durable execution runtime. Temporal, DBOS, Inngest and Restate solve surviving machines, and solve it properly. This survives a tick — it stops cleanly and resumes from a persisted cursor. If you need cross-machine durability, run one of those underneath.
Not a queue, a worker pool, or a retry policy. executeDag performs a ready
batch one step at a time, checking the deadline before each. How many to run at
once, on what pool, with what backoff, is yours.
Known limits
reserveMsis a heuristic. A step that overruns its reservation is still killed mid-flight, and only the idempotency key saves it.- The key is
runId:stepId, unescaped. If your run ids can contain:, the key is ambiguous. Hash or prefix them before they reach the kernel. - No parallel fan-out.
executeDagis sequential within a batch, on purpose. - The DAG loop keeps no
runningset. A step handed out and killed before its outcome was written must be offered again; arunningmarker that outlives the process is exactly what would stop that. - The harness cannot check that your handler is idempotent. It can only give it the key.
Testing
pnpm test # 130 tests, 99% statementsCoverage thresholds are enforced by pnpm test, not merely configured. Every
guarantee was verified red-green — the guard was removed and the suite
confirmed to fail, not merely to pass with it present:
| mutation | what went red |
|---|---|
| canStartStep always returns true | 13 tests, incl. "never invokes the handler for the step it refused to start" and "leaves the run RESUMABLE" |
| key becomes runId:stepId:Date.now() | 14 tests, incl. "is IDENTICAL when the same step is re-executed" |
| a thrown handler recorded as ok: true | 9 tests, incl. "a failed step does not advance the cursor" |
| deadline checked before nextStep | "still records a ceiling halt rather than reporting a pause forever" |
| a kind with no handler is skipped | 4 tests, incl. "never reports success for a partially-run plan" |
| audit port no longer fails open | "finishes the run even though every audit write throws" |
| recordStep moved back inside the try | "re-offers a step that was begun but never recorded — the mid-step kill" |
| DAG records a failed step as succeeded | 5 tests, incl. the when: "failed" compensation branch |
| DAG ignores the injected clock | 3 deadline tests |
Prior art — this is an extraction, not an invention
The loop, the pause semantics, the runId:stepId key, the fail-closed handler
lookup, the "spend is 0 when a handler throws" rule and the iteration bound all
come from Finocket's lib/ai/agent-executor.ts, which has run this in
production on Vercel against real Indian SMB traffic. Finocket offered it for
extraction rather than letting a third runtime be written, and reviewed the
shape. The comments explaining the serverless failure modes are theirs, kept
close to verbatim because they are the reasoning, not decoration.
Changed from the original, deliberately: the success-path recordStep was
moved outside the try that wraps the handler. In the original, a ledger write
that failed was caught and re-recorded as a step failure — halting the run
with a reason describing the database, for work that actually succeeded. Left
outside, the write failure propagates, the cursor stays put, and the next tick
re-offers the same step with the same idempotency key: the case at-least-once was
designed for.
Added beyond the original: the optional DAG loop, the audit and approvals
ports, reserveMs, pauseCause, auditFailures, and ctx.deadlineAt.
Everything Finocket relies on returns the same shape it always did.
MIT.
