agentfootprint
v9.87.1
Published
The explainable agent framework — backtrack a wrong answer to the exact context that caused it (evidence, not guesses). Built on footprintjs.
Maintainers
Keywords
Readme
Map, Walker, Trace, Fold, Lens
Five words for the five jobs every file in this library does.
- Map — what could happen: the chart an
AgentorLLMCalldeclares, the skills and the edges between them. Still; a run never changes it. - Walker — what is happening: the ReAct loop, one stage at a time, with its cursor and its guards.
- Trace — what happened and why: the typed event stream and the commit log, written as the run happens, never reconstructed after.
- Fold — what may be claimed now: everything derived from that record — the evidence corpus, the reachable set, the influence ranking, the context ledger — each computed by one owner and handed out detached.
- Lens — what this reader is served: for an agent that means the wire — the system prompt, the messages and the tool list the model actually gets, plus every sentence composed into them. A Lens may leave something out; it may never say the run holds nothing when it does.
Every folder under src/ opens its README with its role. The map of the whole tree, with the laws and worked examples, is docs/design/map-walker-trace-fold-lens.md.
The new error class
For decades, software had two kinds of errors — and developers never needed deep domain knowledge to fix either:
| Error class | Where the bug lives | How you find it |
|---|---|---|
| Infrastructure — crash, timeout, 500 | the system | infra logs, monitoring |
| Business logic — wrong branch, wrong math | the code | stack trace, debugger, console.log |
| Contextual — wrong tool chosen, wrong fact believed, stale memory trusted | what the model was given | nothing. Until now. |
Agents introduced the third class. The code is correct, the infra is healthy, the answer even reads well — and the run is still wrong, because something influenced the model:
| The model… | because… | |---|---| | picked the wrong tool | two descriptions read nearly alike — it chose between twins | | believed a wrong "fact" | a tool returned it, or an injected fact planted it | | followed the wrong instruction | the wrong skill / steering fired — or fired one iteration too early | | answered from the past | a previous turn or stale memory bled into this one |
Classical logs can't explain any of it: they record what the code did, never what the context did. The debugging question changed — no longer "what did my code do?" but "who influenced the model?"
The idea
If contextual errors live in what the model was given, then the run itself must be structured so context is evidence — every injection, read, write, decision, and tool call recorded connected, the moment it happens. Not logs you grep. Evidence you ask.
Quick start — runs offline, no API key
npm install agentfootprint footprintjsimport { Agent, defineTool } from 'agentfootprint';
import { mock } from 'agentfootprint/providers';
const weather = defineTool({
name: 'weather',
description: 'Get current weather for a city.',
inputSchema: {
type: 'object',
properties: { city: { type: 'string' } },
required: ['city'],
},
execute: async ({ city }: { city: string }) => `${city}: 72°F, sunny`,
});
const agent = Agent.create({
provider: mock({ reply: 'I checked: it is 72°F and sunny.' }),
model: 'mock',
})
.system('You answer weather questions using the weather tool.')
.tool(weather)
.build();
const result = await agent.run({ message: 'Weather in Paris?' });
console.log(result); // → "I checked: it is 72°F and sunny."For production, import a real provider from agentfootprint/providers and swap it in — anthropic(...) / openai(...) / bedrock(...) / gemini(...) / ollama(...). Only the import line changes; the agent code stays the same. (The vendor-SDK providers live on the agentfootprint/providers subpath so the main agentfootprint barrel stays free of optional peer-dep requires; mock, browserAnthropic, and browserOpenai are on the main barrel.)
Run against a local model — the free rung between the mock and the bill
No cloud account, no API key, no vendor SDK, $0 per token:
import { Agent } from 'agentfootprint';
import { ollama } from 'agentfootprint/providers';
const agent = Agent.create({ provider: ollama('llama3.2'), model: 'llama3.2' }).build();
// → talks to http://localhost:11434 (run `ollama pull llama3.2` first)This is the step that makes "the test run and the production run are the same code path" more than a slogan. A mock proves your control flow; it can't tell you whether a real model calls your tool, or what it makes of a tool description you wrote in a hurry. A local model can — and because it's free, you'll actually check before you pay.
ollama() talks Ollama's native API directly, so there's nothing to install on this side, streamed calls report real token counts (so .compaction() and cost budgets work), and when it can't work it says why in words that contain the fix — ollama serve when nothing is listening, ollama pull <model> when the model isn't there, never a raw connection error and never a hang.
For llama.cpp's llama-server, vLLM, Together or Groq, use openai({ baseURL: 'http://localhost:8080/v1', apiKey: 'not-needed', defaultModel: '…' }) — any server speaking the OpenAI Chat Completions API, same Agent code either way. Full recipes: Ollama guide · OpenAI-compatible endpoints.
Then add context
A real agent carries more than one prompt and one tool: facts about the user, always-on rules, skills that unlock on demand. Declare each piece — the framework decides when it fires and which slot it lands in, and every piece is born tracked:
import { defineFact, defineSteering, defineSkill } from 'agentfootprint/context';
const agent = Agent.create({ provider, model })
.system('You are a support agent.')
.fact(defineFact({ // data the model should know — always on
id: 'user-profile',
data: 'Name: Maya · Plan: Pro · Customer since 2022',
}))
.steering(defineSteering({ // rules the model must follow — always on
id: 'refund-policy',
prompt: 'Never promise a refund before checking the policy tool.',
}))
.skill(defineSkill({ // guidance the LLM loads when it asks
id: 'billing',
description: 'Use for refunds, charges, billing questions.',
body: 'When handling billing: confirm identity first, then…',
tools: [refundTool],
autoActivate: 'currentSkill', // ...and scope its tools to that window too
}))
.build();Same shape for .instruction() / .memory() / .rag() / raw .injection() — they're all the one primitive, Injection = slot × trigger × cache. The full model ↓
The SkillMap — declare which skills connect
When an agent has several skills, the routing between them is a thing you declare, not prose you hope the model follows. The official names, in one sentence: you declare the SkillMap; the agent is the SkillWalker; the recording carries both.
- DECLARE —
defineSkillMap(a permanent alias ofskillGraph— same function, both names forever): the skills, the edges between them, entry matchers as data (match:), and — since 9.51.0 — guards as data on route edges (guard:). - ATTACH —
.skillGraph(map). There is no walker class to construct: the agent IS the SkillWalker, moving its cursor over your map. - WATCH — every recording carries the map (
skill.graph_declared, guard conditions included) and the walk (cursorMoveon every iteration, with the evidence that opened or closed each guarded hop). The SkillGraph debugger renders both.
The walker moves by exactly three movers:
| mover | who decides | on the record |
|---|---|---|
| llm | the model picks via read_skill, bounded by the gate | by: 'model-pick'; refusals as skill.rejected |
| guard | your data decides — a when / onToolReturn / onToolStatus / guard: edge fires | by: 'route', with cursorMove.guard evidence when a data guard judged it |
| linear | no choice — a hand-off that fires every time its source finishes | by: 'route' |
A guard: is the data twin of the when predicate (at most one of the two per edge): conditions over the hop (toolName, status, iteration, …) and over the tool result's own JSON fields, in footprintjs's filter grammar (eq/ne/gt/gte/lt/lte/in/notIn, all ANDed):
const map = defineSkillMap()
.entry(triage, { match: { keywords: ['check', 'account'] } })
.route(triage, escalation, {
onToolReturn: 'assess_risk',
guard: { riskLevel: { in: ['high', 'critical'] }, score: { gte: 0.7 } },
})
.route(escalation, wrapup, { onToolReturn: 'close_case' })
.build();Because the guard is data: the build proves contradictions (guard-unsatisfiable — a guard that can never pass refuses to build, naming the conflict), map.toMermaid() captions the edge (on assess_risk when riskLevel in [high, critical] AND score ≥ 0.7), and every evaluation that decides a hop — taken or refused — puts its per-condition evidence on the record (cursorMove.guard / cursorMove.guardsClosed), so "why didn't my edge fire?" is a lookup, not a debugging session. Runnable end to end · The five-minute walkthrough →
When a skill's playbook is really a sequence, declare it as data and the framework walks it — offering one step's tool at a time (banner-led: [Step 2 of 6 — confirm the duplicate charge]), with skip_step to put a decline on the record and one teaching nudge if the model stops early. Your other tools stay available throughout — a declared order, not a cage:
defineSkill({
id: 'refund',
description: 'Handles refunds end to end, by declared procedure.',
body: 'Follow the refund procedure. Every step says why it exists.',
tools: [findOrder, checkHistory, issueRefund, fileReceipt],
steps: [ // the procedure, as data (9.18.0)
{ tool: 'find_order', note: 'find the order before touching money' },
{ tool: 'check_history', note: 'confirm the duplicate charge' },
{ tool: 'issue_refund', note: 'refund the duplicate charge only' },
{ tool: 'file_receipt', note: 'file the receipt for audit' },
],
});Every move lands on the typed stream (skill.step_advanced / step_skipped / steps_unfinished), and a step whose tool asks a human pauses the run and advances on resume — runnable end to end.
The cursor picks the brain, too (9.19.0). The graph already decides where the run is; defineSkill({ provider, model }) — or skillGraph(g, { providers }) — lets that position decide who answers: triage on the small model, the refund skill on the strong one, llm_start.brain recording which rung won. Add escalation: { provider, model, afterRefusals: N } and N recorded routing refusals in one turn flip the rest of it onto the bigger brain (skill.escalated — evidence, never vibes), and decider: { provider, model } resolves an ambiguous turn-start menu out of band (turn_routed { by: 'decider' } — the sanctioned resolver for rails). And tools steer with data, not string conventions: return { content, status, effects } and a 'denied' refund routes by a declared onToolStatus edge — meaning, not prose — while a propose-transition / require-instruction effect is validated against the graph's own law, every acceptance and teaching refusal a typed tools.effect event. Runnable end to end
Then keep the conversation
run() is one turn. It seeds the conversation from the message you pass and nothing else, so calling it twice gives you two conversations — right for one-shot work, and not what a chat wants. Continuing is something you name:
await agent.run({ message: 'Book me a table for two on Friday.' });
await agent.followUp('Make it three.'); // same conversation
// …or hand the conversation around: plain JSON, any store, any machine.
const conversation = agent.checkpoint();
await agent.run({ message: 'Make it three.', continueFrom: conversation });The conversation carries its own identity, so a continued turn writes its memory where the earlier turns can read it. Two things that look like this and are not: identity.conversationId is a namespace key (it scopes memory, RAG and permissions — it does not join two runs), and .memory() gives you recall in the system prompt rather than the verbatim window. standingAgent({ agent, sessions, host }) does the whole store-and-continue dance per session for you. Worked example, printing the wire each way →
The conversation keeps its place, too (9.17.0). On a skill graph mounted with .skillGraph(graph, { continuity: 'conversation' }), the skill the last turn ended on rides the same checkpoint — followUp('one more thing') starts there instead of re-routing from scratch. A sticky default, never a lock: each new message is still judged by the graph's declared rules and intents (match: { intent, examples } + classify: a scorer), a decisively different topic moves, a genuinely ambiguous one offers the model a menu, and every verdict — winners, losers, the thresholds that judged them — is recorded as agentfootprint.skill.turn_routed. The routing cascade, in five minutes → · Runnable example →
Then compose control flow
One agent is a Runner. So is every composition of agents — four control-flow primitives, and anything that runs composes into anything else:
import { Sequence, Parallel, Conditional } from 'agentfootprint';
const pipeline = Sequence.create()
.step('classify', classifyAgent) // sequence: step → step
.step('review',
Parallel.create() // parallel: fan out, then merge
.branch('legal', legalAgent)
.branch('ethics', ethicsAgent)
.mergeWithLLM({ provider, model, prompt: 'Synthesize:' })
.build())
.step('respond',
Conditional.create() // conditional: one branch runs
.when('urgent', (i) => i.message.startsWith('URGENT'), urgentAgent)
.otherwise('normal', normalAgent)
.build())
.build();
await pipeline.run({ message: 'URGENT: refund dispute on order #4411' });The fourth primitive is Loop — Loop.repeat(agent).until(guard).times(5), with a mandatory budget guard. And the named patterns from the research literature ship pre-composed from the same four: selfConsistency · reflection · debate · mapReduce · tot · swarm · llmSwarm (a swarm whose hand-offs an LLM decides). Because every composition is a flowchart, the structure you wrote is the structure you see in the UI — and the trace spans the whole pipeline, not one agent at a time. Designing systems of agents ↓
How — we abstract context engineering
Skills, steering, RAG, facts, memory, guardrails — every name for context does one thing: it injects into one of three LLM slots. So we abstracted the injection itself.
What tracking buys you
See it in 30 seconds — four questions logs can't answer, each answered by code in this repo from a real run:
Q: Why did the model pick refund_full instead of refund_partial?
A: margin 0.02 — ⚠ NARROW: the two tool descriptions read nearly identical
(toolChoiceRecorder — and the catalog lint flags the pair before you ever run)
Q: Why was this loan declined?
A: decision ← [control: "DTI above the 0.43 affordability ceiling"] ← dti 0.52 ← monthlyDebt / income
(decide() evidence + the causal slice — every hop is a real recorded edge)
Q: Which piece of context made the answer wrong?
A: CAUSAL: ablating fact 'vip-override' flipped the outcome in 3/3 seeded reruns
(localizeContextBug — ranked proxies, counterfactual proof)
Q: Prove nobody edited this run's record.
A: verifyAuditBundle → valid: false, brokenAt: #16 — the tampered record, named
(hash-chained audit export, offline verification)And you don't have to read the trace yourself — the trace toolpack lets a debugger model do it: in one run it found a planted bug while reading 9.5% of the trace (guide).
And the watching costs the run nothing:
One contextual error, walked end to end
The third question above, in full — every value below is the captured output of
examples/observability/05-context-bisect.ts
and 06-backtrack-trace.ts, runnable offline.
The bug. A refunds agent carries a poisoned customer-profile fact. It answers:
"Refund APPROVED: Dana Reyes holds VIP tier override status, so the 47-day-old order qualifies for a refund beyond the 30-day window."
The policy says 30 days. The logs look fine — the model was given bad context, and classical logging has no row for that.
The walk. Because context here is state, the decision backtracks like a variable: who read it, who wrote it, who let it in, where it was born —
ANSWER "Refund APPROVED…" ← the bug
READ call-llm#40 assembled the system prompt ← exactly what the model saw
LANDED context#6 wrote systemPromptInjections ← who mutated state
ALLOWED trigger { kind: 'always' } — active every iteration ← why it was let in
BORN defineFact('vip-override-fact') ← who wrote itThat chain is the provable candidate set — everything that demonstrably reached
the call, nothing else. Influence scoring then ranks inside it, and counterfactual
ablation proves: removing vip-override-fact flips APPROVED → DECLINED in
3/3 seeded reruns (the benign fact and the lookup tool: 0/3). Scores are proxies;
only ablation makes the causal claim.
Three interfaces, one per shape of bug — ship-a-default, bring-your-own:
| interface | finds the culprit when it is… | confirm by |
|---|---|---|
| influence ranking (scoreInfluence + rankingConfidence) | present — orders suspects, says when it can't rank | — |
| ablation (localizeContextBug) | present — remove it, see the outcome flip | removal |
| missing-context finder (findDroppedContext) | absent — available but never reached the model (available − sent) | restoration |
The third closes a gap the first two are blind to: a key instruction truncated out of
the window has nothing to ablate, so you confirm by restoration. And when scoring is
too flat to trust, rankingConfidence returns a shortlist to confirm rather than a
confident, wrong #1. Guides: ranking-confidence ·
missing-context.
The same walk, visual. toBacktrackTrace() serializes the report into
Story Lens's <BacktrackView>
— the "why?" board, triggerable from any decision point (final answer, a mid-loop tool
choice, a deterministic decide() rule):
Pick how influence is scored — then re-run without the noisy source
Influence ranking is a pluggable strategy. Two ship out of the box:
semantic-alignment (embeddings, the default) and lexical-overlap (plain word
overlap — free and deterministic). listInfluenceStrategies() gives a UI everything
it needs to offer a picker, and to grey out what it can't run.
Found the source you suspect? rerunWithoutSources runs the same scenario again without
it. It reports what changed — the answer, how often it flipped across seeded re-runs, and
(with checkBaseline) a causal verdict. Scores suggest; re-runs convict.
const again = await rerunWithoutSources({ report, ignore: ['social-sentiment'], runner, originalAnswer, embedder });
again.answer; // 'HOLD …'
again.whatChanged.summary; // 'Removing sources [social-sentiment] changed the answer in 3/3 seeded re-runs …'Chat sessions that can explain themselves
recordedChat wraps your agent factory into a recorded conversation. Every send()
freezes that turn's evidence. reason(k) shows what influenced reply K. rerunTurn(k,
{ ignore }) re-runs that exact turn — same recorded history, byte for byte — without a
source, and returns the same honest result rerunWithoutSources gives you. fork(k,
{ fromRerun }) continues the conversation from the what-if as a new recorded session.
Branch, never rewrite.
import { recordedChat } from 'agentfootprint/observe';
const chat = recordedChat({ makeAgent }); // your factory, specs applied at construction
await chat.send('Should we BUY or HOLD?'); // recorded turn (frozen evidence)
const rerun = await chat.rerunTurn(1, { // that turn, minus one source
ignore: ['social-sentiment'], embedder, checkBaseline: true,
});
const fork = chat.fork(1, { fromRerun: rerun }); // continue from the what-ifPick your door
| 🔧 Building an agent? | 🐛 Agent misbehaving? | 🏛️ Need audit / compliance? | |---|---|---| | Typed agents with skills, steering, RAG, memory, guardrails — and the trace for free. | Lint your tool catalog in 5 minutes — works on any framework's tool list (plain JSON / MCP / OpenAI / Anthropic shapes). Then causal slices, context bisection, and the debugger-LLM toolpack. | Hash-chained, tamper-evident run records with an offline verifier — record-keeping in the EU-AI-Act shape. | | → Quick start · → Build ↓ | → Debug ↓ · → Tool-catalog lint · → Trace debugging | → Audit ↓ · → Security guide |
The model — what we abstract
You collect domain-specific data and instructions — Skills · Steering · Guardrails · RAG · Tool APIs · Memory, with more on the way. They all do one thing: inject into one of three slots (system, messages, tools). So we abstracted the injection itself.
The abstraction is three rules:
- Three slots are fixed.
system,messages,tools— the LLM API surface. - N flavors are open. You declare what you have. Tomorrow's flavor (few-shot, reflection, persona, A2A handoff…) plugs in the same way.
- Rules decide where and when. You provide the rules. We collect your data, fire the right one, land it in the right slot at the right iteration.
That's the whole model: Injection = slot × trigger × cache.
- Slot — which of the 3 LLM API regions the content lands in (
system/messages/tools). - Trigger — when the content fires (see below).
- Cache — how stable the content is across iterations. The framework places provider cache markers for you — stable content gets 80–90% cheaper prefixes.
The 4 triggers
| Trigger | Flavor | Fires when | Illustration | Default slot |
|---|---|---|---|---|
| always | static | Every iteration | .steering(defineSteering({ id, prompt: 'You are a triage agent…' })) | system |
| rule | runtime — predicate | Your rule returns true | .instruction(defineInstruction({ id, activeWhen: s => /price\|refund/.test(s.userQuery), prompt })) | system |
| on-tool-return | runtime — lifecycle | After a specific tool returns | .instruction(defineInstruction({ id, activeWhen: s => s.lastToolResult?.toolName === 'search', prompt: 'Cite source IDs.' })) | system |
| llm-activated | runtime — agent-driven | LLM calls read_skill('id') | .skill(defineSkill({ id: 'refund-policy', description, body })) | system (body) + tools |
[!NOTE] The "Illustration" column shows the shape of each flavor — the typed builder methods (
.steering/.instruction/.skill/.fact/.rag) take anInjection(orMemoryDefinitionfor.rag) produced by the matchingdefineSteering/defineInstruction/defineSkill/defineFact/defineRAGfactory. ASkilltargets more than one slot at once:tools(the schemas it contributes — registered up front and visible from iteration 1 unless the Skill setsautoActivate: 'currentSkill', which scopes them to the iterations where it is active) andsystem(its body — or, withsurfaceMode: 'tool-only', theread_skillresult instead). Themessagesslot both projects the conversation and accepts delivery:slot: 'messages'(with aroleyou name) appends to the window itself, subject to what the attached provider carries insidemessagesand to a sequence rule that defers rather than reorders.
3 slots × 4 triggers × N flavors = the entire context-engineering surface.
Why we chose this abstraction
The agent space has many credible primary abstractions:
| Framework | What it abstracts | |---|---| | LangChain | Pipelines of composable components | | LangGraph | State machines of nodes and edges | | CrewAI · AutoGen | Crews of role-playing agents | | Mastra · Genkit · Pydantic AI | Typed full-stack bundles | | DSPy | Compiled prompts | | Inngest AgentKit | Durable workflows |
We didn't have to choose between them.
agentfootprint is built on footprintjs — the flowchart pattern for backend code. footprintjs gives us every one of those abstractions out of the box:
| Capability | What footprintjs hands us |
|---|---|
| Composition | Sequence · Parallel · Conditional · Loop |
| State machines | The ReAct loop is a flowchart |
| Multi-agent crews | Compose Agents through control flow — no special class needed |
| Durable workflows | pauseHere() plus JSON-portable resume() |
| Typed observation | 60+ events for free, because the framework owns the loop |
So we used the budget those abstractions would have cost us to invest deeply in something they all leave to the developer: the injection loop.
[!IMPORTANT] We abstract context engineering — and hand back the trace. Live to develop · offline to monitor · detailed to improve.
🔧 Build — design your agent or system of agents
Two scales — same alphabet. Four control flows are the entire vocabulary.
import { Sequence } from 'agentfootprint';
const flow = Sequence.create()
.step('a', stageA)
.step('b', stageB)
.step('c', stageC)
.build();import { Parallel } from 'agentfootprint';
const fan = Parallel.create()
.branch('web', searchWeb)
.branch('docs', searchDocs)
.mergeWithFn(synthesizer)
.build();import { Conditional } from 'agentfootprint';
const router = Conditional.create()
.when('billing', s => /bill|invoice|refund/.test(s.message), billingAgent)
.when('tech', s => /error|bug|crash/.test(s.message), techAgent)
.otherwise('default', defaultAgent)
.build();import { Loop } from 'agentfootprint';
const reflexion = Loop.create()
.repeat(thinkAgent)
.until(({ latestOutput }) => latestOutput.includes('DONE'))
.build();Inside one agent — Dynamic vs Classic ReAct
| Iteration | Classic ReAct | Dynamic ReAct (agentfootprint) |
|---|---|---|
| 1 | 12 tools shown | 1 tool (read_skill) |
| 2 | 12 tools shown | 5 tools (skill activated) |
| 3 | 12 tools shown | 5 tools |
Multi-agent — compose with the alphabet
Pick the flows that match your problem. Chain them. That's your Agentic Application.
const research = Loop.create()
.repeat(Sequence.create().step('plan', plan).step('search', searchAll).build())
.until(({ iteration, latestOutput }) => iteration >= 3 || latestOutput.includes('DONE'))
.build();Named patterns — also compositions of the same 4
The patterns the field knows reduce to the same alphabet:
| Pattern | Composition |
|---|---|
| Swarm | Loop( Parallel( Agent×N ) → merge ) |
| Tree-of-Thoughts | Loop( Parallel( Agent×N ) → Conditional(score) ) |
| Reflexion | Loop( Agent → Conditional(critique) → Agent ) |
| Debate | Parallel( Agent_pro, Agent_con ) → Agent_judge |
| Router | Conditional → Agent_A \| Agent_B \| Agent_C |
| Hierarchical | Agent_planner → Sequence( Agent_worker×N ) → synth |
Same trick as the injection model: instead of N libraries for N patterns, we found the M building blocks all N patterns are made of.
📖 Compare: hand-rolled vs declarative · migration from LangChain / CrewAI / LangGraph
Check in with the receipts — human-in-the-loop consent for consequential actions
OpenWorker-class agents check in; agentfootprint checks in with the receipts. A tool declares checkIn: 'always' (or a (args) => boolean predicate). When it trips, the run pauses before the tool executes and hands back an evidence pack: willDo (plain-words claim), read (context the run consumed), drivers (which context drove the choice, ranked, zero LLM calls), and a compact trail. A human answers checkInApproved({ by }) / checkInDeclined({ by, note }); on approve the tool runs, on decline the model sees the note and adapts.
const refund = defineTool({
name: 'issue_refund', description: 'Issue a refund',
inputSchema: { type: 'object', properties: { amount: { type: 'number' } } },
checkIn: (args) => args.amount > 1000, // ask a human only for big refunds
execute: ({ amount }) => `refunded ${amount}`,
});
const out = await agent.run({ message: 'refund 5000' });
if (isCheckInPause(out)) { // distinct from a plain askHuman pause
showToHuman(out.checkIn.evidence); // the receipts
await agent.resume(out.checkpoint, checkInApproved({ by: 'alice@ops' }));
}It rides the same JSON checkpoint as pause/resume — the ask and the decision can be servers and days apart — and lands checkin.request / checkin.decision as typed events (CheckInRecorder captures the audit trail). A tool without checkIn is byte-identical. Permission (policy) still runs first — check-in is consent, not policy.
Run the flagship demo — an AI coworker that drafts a weekly status doc and checks in before posting it (deterministic $0 mock, watch it adapt on decline):
npm run example examples/features/34-checkin-coworker.ts -- --declineSee the Check-in guide.
Act — everything your agent does about its own loop, in one block
Tools do the work. Act decides about the work. Watch remembers both — and nothing can act without being watched.
const agent = Agent.create({ provider, model })
.act({
input: [scrubSSNs], // the message, before the run commits it
beforeTool: [refundCeiling, fourEyes], // every call, before it is dispatched
afterTool: [stripPII], // every result, before the model reads it
window: slidingWindow({ keepRecentTurns: 12 }), // what the live window keeps
output: [noCodenames], // the answer, before the caller gets it
})
.build();Five keys, one per moment, in the order the loop reaches them — so autocomplete on an empty {} teaches the loop. Every rule answers allow(), allow(value, why) or deny(reason) (and ask({ question }) where a person can still change the outcome), and none of them can answer for the tool: the outcome union has no result arm, so what the model finally reads is the real tool's output or a refusal. Every decision files a ledger row stamped with its moment.
It is pure sugar over the five individual doors, pinned byte-equivalent per key — and the bundle's keys are locked against LoopMoment at compile time, so a sixth moment cannot ship without a key for it.
npm run example examples/features/38-act.tsWatch — who is looking while it does
.act() says what the agent may do. .watch() says who is looking while it does it. Observers handed to the builder are attached before build() returns, so there is no window where the agent has run and nobody was watching:
const routes = routeRecorder();
const choices = toolChoiceRecorder({ embedder: staticEmbedder() });
const agent = Agent.create({ provider, model })
.watch(routes, choices) // build-time — sees the very first run
.act({ beforeTool: [refundCeiling] })
.build();
await agent.run({ message: 'refund order 4471' });
console.log(await choices.getFlagged()); // calls where the tool choice was a near-tieVariadic, because observers come in sets. It returns the builder — the runtime door is still agent.attach(observer), which returns an Unsubscribe you own and call when the observer's life ends. Same mechanism underneath, so mixing them is fine and order is preserved.
There is deliberately no list of "watch moments" to go with .act()'s five. A rule has to be told where it may speak, so that list is closed and compiler-pinned; an observer attends the whole stream, and any list we published would be a vocabulary we then had to keep true.
(.recorder() was the same door under its old, internals-flavoured name. It was removed in 9.0.0 and now throws a sentence pointing here — see 9.0.0.)
🐛 Debug — see what your agent did
Because we own the loop, every decision and execution is captured during traversal — not bolted on. The default capture is the causal trace: every stage, read, write, and decision evidence as a JSON-portable, scrubbable, queryable, exportable artifact — and every LLM call backtracks to four typed answers: what was injected, who triggered it (which rule), when it fired, how it landed (slot · position · cache). Beyond the default, wire custom recorders for cost, latency, or quality scoring — any observation hook fires on the same stream.
The same trace serves three downstream consumers — no extra instrumentation:
| Consumer | What the trace gives you |
|---|---|
| Audit / compliance | Six months later, "why was loan #42 rejected?" answers from the chain (creditScore=580 < 620 ∧ dti=0.6 > 0.43 → riskTier=high → REJECTED) — no LLM call. GDPR Art. 22 / ECOA / EU AI Act adverse-action notices write themselves from the captured decision evidence. |
| Cheap-model triage | A Sonnet trace is good input for Haiku to answer follow-ups: ~200 tokens ($0.25/1M) vs ~2,500 at a reasoning model ($15/1M). Memoized thinking — no agent rerun. |
| Training data | Every successful chain is a labeled trajectory; SFT pairs fall out of the snapshot's history field. (The export wrapper, DPO, and process-RL need extra collection layers — roadmap.) |
Four views, one trace — pick by question:
| View | Shows | When to use | |---|---|---| | Story Lens (the hero up top) | The run replayed as an animated, scrubbable story — the brain, the tools, the reasoning | Show anyone what the agent did | | BacktrackView (the board above) | A decision walked backwards — suspects, influence meters, ablation stamps, custody rewind | Answer why it decided that | | Why Lens | Agent-centric — User/Agent[3 slots]/Tool flowchart with iteration scrubber and round commentary | Live debugging, "what did the agent see at step 5?" | | Explainable Trace | Structural — subflow tree, full flowchart, memory inspector, per-stage execution timeline | Architecture review, root-cause analysis |
And two conversational doors over the same evidence — ask instead of look:
// dedicated: a cheap model debugs an expensive run by id — pays for what it opens
const debuggerAi = traceDebugAgent({ artifacts, provider: anthropic(), model: 'claude-haiku-4-5' });
await debuggerAi.run({ message: 'Why was loan APP-7 approved?' });
// in-conversation: the agent answers "why did you…?" from its OWN previous turn
Agent.create({ provider, model }).tool(lookupOrder)
.selfExplain({ delegate: { provider: anthropic(), model: 'claude-haiku-4-5' } })
.build();.selfExplain() mounts one skill: the catalog stays clean until the LLM activates
it, evidence binds only to completed runs (never in-flight), and delegate
answers at the cheap model's price inside the expensive conversation.
Guide · examples
07 ·
08 · the doors walk the
same evidence the board visualizes ▶.
📖 Powered by footprintjs
causalChain()— backward thin-slicing on the commit log. Causal memory deep dive · Explainability & compliance
One recording. Two lenses. Three consumers. Zero extra instrumentation.
Observers stay off the hot path
By default every agent.on() listener runs synchronously inside the producing
statement. One option moves observation off the hot path:
Agent.create({ provider, model, observerDelivery: 'deferred' }) // default 'inline'
// serverless / shutdown: settle async listener work before the freeze
await agent.drainObservers({ timeoutMs: 5_000 });Events are captured into a bounded queue (≈ microseconds on the hot path) and
delivered one beat behind — same typed events, same order, zero loss, a throwing
listener can't kill the run, and per-listener stats land on
getLastSnapshot()?.observerStats to name the hog. Terminal boundaries (resolve,
crash, pause) drain synchronously first, so checkpoints are always complete.
Measured: −8% wall on a 50-iteration agent with a deliberately slow listener
(example 21).
📖 Full semantics (capture policies, backpressure, overflow): deferred-observers guide
Lint your tool catalog — before the model picks the wrong twin
Tool routing is an LLM decision driven by names + descriptions — so lint the catalog like code and gate it in CI. Zero stack buy-in: works on any OpenAI / Anthropic / MCP / plain tool list, no agentfootprint runtime needed.
npx agentfootprint-lint-tools tools.json --threshold 0.94 --strict✗ CONFUSABLE 0.9445 get_fcns_database <> influx_get_fcns_database
hint: names differ only by 'influx' — make the descriptions say WHEN to choose each
~ warn [enum-in-prose] influx_get_port_ranking.metric
suggest: "enum": ["avg_iops","peak_iops","mbps"]Pairwise confusability over what the model reads (embedder pluggable,
content-hash cached) plus a pluggable structural rule pack
(missing/short descriptions, says-WHAT-not-WHEN, enums hiding in prose,
undocumented optional params). The runtime counterpart, toolChoiceRecorder
(agentfootprint/observe), scores each live LLM call's tool choice against
the same geometry and flags narrow margins and proxy disagreements — lazily,
off the hot path.
agentfootprint owns the detection — the scores, the ties, the recorded margins; you own the policy — the threshold, the embedder, the rewrite. We map; you decide.
📖 Tool-catalog lint guide — 5 minutes from a tools.json to a gated CI check ·
examples/observability/02·03·04
🏛️ Audit — prove what happened
Answering "why was the loan rejected?" from captured evidence is the debug door above. The audit door adds the integrity layer: prove the record itself hasn't been edited since capture. auditExport() hash-chains every typed event — decisions, tool calls, validation rejections, permission verdicts, costs — into an append-only bundle (EU AI Act Art. 12 record-keeping shape); verifyAuditBundle() re-checks it offline — no agent, no LLM — and names the exact record any tamper broke.
import { auditExport, verifyAuditBundle } from 'agentfootprint/observe';
const audit = auditExport({ agent: 'ledger-auditor' });
const stop = agent.enable.observability({ strategy: audit });
await agent.run({ message: 'audit account ACCT-1142' });
stop();
const bundle = audit.bundle(); // plain JSON — store anywhere
verifyAuditBundle(bundle); // { valid: true, recordsChecked: 50 }
// flip one byte anywhere → { valid: false, brokenAt: 13, reason: 'hash mismatch — …' }Payloads are PII-bounded by default (tool args as key names, results as a type, content as [N chars] markers). And it's honest about its limits: tamper-evident, not tamper-proof — for non-repudiation, anchor both chain ends in external storage (WORM store, signed log).
📖 Tamper-evident audit guide ·
examples/features/19-audit-export.ts— capture → verify → tamper → drain ·20-regulated-decisioning.ts— an offline auditor reconstructs a loan decline from persisted files, both chain ends anchored
Mocks first, production second
Build the entire app against in-memory mocks with zero API cost, then swap real infrastructure one boundary at a time.
| Boundary | Dev | Prod |
|---|---|---|
| LLM provider | mock(...) | ollama('<model>') free · anthropic() · openai() · bedrock() |
| Memory store | InMemoryStore | RedisStore · AgentCoreStore |
| MCP | mockMcpClient(...) | mcpClient({ transport }) |
| Cache strategy | NoOpCacheStrategy | auto-selected per provider |
The flowchart, recorders, and tests don't change between dev and prod.
What ships today
Core
- 2 primitives —
LLMCall,Agent(the ReAct loop) - 4 control flows —
Sequence,Parallel,Conditional,Loop(plusworkflow(), the same sequence with every hand-off type-checked by the compiler, andgraph(), a fixed DAG whose independent nodes run concurrently) - 1 Injection primitive —
defineSkill/defineSteering/defineInstruction/defineFact - 1 reliability gate —
.reliability({ preCheck, postDecide, providers, circuitBreaker, fallback }) - 1 tool dispatch primitive —
ToolProvider(sync OR async) —staticTools·gatedTools·skillScopedTools· or a customToolProviderthat discovers over hubs / MCP / per-tenant catalogs
LLM providers (8)
| Factory | Use for |
|---|---|
| anthropic | Claude (Sonnet, Opus, Haiku) via @anthropic-ai/sdk |
| openai | GPT-4o, GPT-4-turbo via openai SDK |
| bedrock | Claude / Titan / Mistral via AWS Bedrock runtime |
| gemini | Gemini via @google/genai — two doors, Vertex (project + ADC) or the Gemini API (one key). Not every door/model pair works: read the door/model matrix first |
| ollama | Local models, over Ollama's native API — no SDK, no key, real token counts, refusals that name ollama serve / ollama pull · openai({ baseURL }) reaches llama.cpp, vLLM, and any other OpenAI-compatible endpoint |
| browserAnthropic | Browser-side Claude calls (no proxy server) |
| browserOpenai | Browser-side OpenAI calls (no proxy server) |
| mock | Deterministic dev/test (zero API cost) |
Memory + adapters
- Memory factory — 4 types (
episodic/semantic/narrative/causal) × 7 strategies (window/budget/summarize/topK/extract/decay/hybrid) - Memory stores —
InMemoryStore,RedisStore(peer-depioredis),AgentCoreStore(peer-dep AWS SDK) - RAG · MCP adapters —
mockMcpClient(...)/mcpClient({ transport })
Operability
- Provider-agnostic prompt caching — declarative per-injection, per-iteration marker recomputation
- Artifacts (the claim check) — tools check data into a governed store (
ctx.artifacts; in-memory / file / SQLite / S3 / Cloud Storage) and the model routes ~26-char tickets instead of hauling payloads:wantsresolves ref arguments at dispatch with teaching refusals, the placement threshold (artifacts: { store, placement }) auto-refs oversized tool results, and the auto-attachedpresenttool hands a ref to the screen with a durable description snapshot — every mint / resolve / refusal / presentation a typed event - Artifact stores everywhere the agent runs — one five-verb port, five adapters:
inMemoryArtifacts·fileArtifacts·sqliteArtifacts·s3Artifacts·gcsArtifacts, all running the same contract suite. Scope-partitioned keys (a..tenant is a name, never a hop), the ticket as one metadata entry, digest verified on read, retention stated at mint. OptionalputStream/getStreamare feature-detected (canStreamArtifacts) — a store that cannot stream leaves them absent rather than faking it. The cloud adapters ship contract-shaped and tested, awaiting field use - Skill artifact vocabularies — a skill or a step declares
produces/consumes(artifact kinds), andgraph.checkup()warnsartifact-kind-unsatisfiedwhen nothing on the agent claims to make what a consumer needs. Honest by construction: a warning, never an error, because it reads declarations only and artifacts outlive the turn that made them - Human-in-the-loop pause / resume — a tool calls
pauseHere(...)(oraskHuman(...));isPaused(result)hands you a JSON-serializable checkpoint, andagent.resume(checkpoint, input)continues hours later on a different server - Resilience primitives —
withRetry,withFallback,withCircuitBreaker,.outputFallback,agent.resumeOnError - Context Integrity — deterministic checks at the seams where a run contradicts ITSELF: a tool parked but still on the wire, a tool offered after the results grounding it were evicted, an answer field that disagrees with the fact it claims to report (
.claims(), requires.outputSchema()). Nothing is blocked or rewritten — each defect is one typed finding, and every run files a disposition ledger so "no findings" and "no check ran" stay different states.integrityPosture: 'dev'adds the liveness proofs (a start-of-run canary;CheckerDeadErrorinstead of a green report from a checker that never ran). Read it back withfind_context_errorsover a recording — Context Integrity - 60+ typed observability events —
agent·composition·context·stream·tools·skill·memory·cache·cost·permission·eval·embedding·pause·error·fallback·resilience·reliability·risk
Debugging & compliance (agentfootprint/observe)
- Tool-catalog lint —
npx agentfootprint-lint-tools(any framework's tool list) + runtimetoolChoiceRecordermargins - Contextual-bug localizer —
localizeContextBug(causal slice → influence ranking → counterfactual ablation) +bisectCulprits toBacktrackTrace— render any decision as the BacktrackView "why?" board- Trace toolpack — 11 bounded, LLM-callable tools so a debugger model walks the trace by id, including
find_context_errors(what the run contradicted itself about, joined to the step that filed it) andinspect_tool_run: the descent INTO a tool call, when the tool kept its own record (flowchartAsTool({ keepRecord: true })) traceDebugAgent(dedicated debugger session) ·.selfExplain()(in-conversation why-questions, skill-gated, with a cheap-modeldelegateswitch)- OTel GenAI span export · hash-chained tamper-evident audit bundles with an offline verifier
Tooling
- Story Lens — animated run player + BacktrackView why-board (separate
agentthinkinguipackage) - Why Lens · Explainable Trace — two visual replays of the causal trace (separate
agentfootprint-lenspackage) - AI-coding-tool support — Claude Code · Cursor · Windsurf · Cline · Kiro · Copilot
Where to next
| If you are... | Go here | |---|---| | New to agents | 5-minute quick start | | Coming from LangChain / CrewAI / LangGraph | Migration guide | | Architecting an enterprise rollout | Production guide | | Doing due diligence | Architecture overview | | Researcher / academic background | Citations & prior art | | Curious about design | Inspiration docs |
Or jump into the examples gallery — every example is also an end-to-end CI test.
Tree-shakeable & ESM-first
Import one thing, ship one thing. agentfootprint is built so your bundle grows only with what you actually use:
- Dual build, true ESM. Ships CommonJS (
require) and real ECMAScript Modules (import) with TypeScript types. The ESM build istype:modulewith explicit.jsimport extensions, so it loads as true ESM under Node, Vite, Next, Deno, and Bun — no shims. - Per-file modules + honest
sideEffects. The dist is emitted file-by-file (never pre-bundled), so bundlers drop every export you don't touch. A smallimport { defineTool }doesn't pull in the Agent runtime, injection engine, memory stores, or LLM providers. - Ten doors, named for what you're doing.
agentfootprint·/providers(plug in a backend) ·/memory(state that outlives a turn) ·/observe(everything that watches) ·/context(how context gets assembled) ·/resilience(when the call fails) ·/security(who may do what) ·/hosting(behind a wire) ·/events(the typed wire vocabulary) ·/cache(prompt caching). 8.0.0 consolidated 26 internals-named subpaths into these; every old path still resolves for all of 8.x. - Lazy peer-deps. Heavyweight integrations load their SDK only when you instantiate them — importing agentfootprint never bundles
@anthropic-ai/sdk,ioredis, the AWS SDKs, or the MCP SDK unless you actually use that adapter.
Proven, not promised. A CI smoke test bundles a minimal import { defineTool } and asserts the Agent runtime, injection engine, memory stores, and providers are pruned; a second test loads the main barrel and every subpath as true ESM and verifies the lazy-adapter loader works under ESM (createRequire, not a bare require). See test/esm-packaging.test.ts.
Built on
footprintjs — the flowchart pattern for backend code. agentfootprint's decision-evidence capture, narrative recording, and time-travel checkpointing are footprintjs primitives at the runtime layer.
You don't need to learn footprintjs to use agentfootprint — but if you want to build your own primitives at this depth, start there.
Citing
agentfootprint is part of a research program on making software systems explain themselves — every run records why it did what it did, as a causal trace. Researching agent transparency, observability, or explainable AI? The ecosystem map is a good starting point.
If you use agentfootprint in academic work, please cite it (or use the "Cite this repository" button on GitHub):
@software{anbalagan_agentfootprint,
author = {Anbalagan, Sanjay Krishna},
title = {agentfootprint: the explainable AI agent framework},
url = {https://github.com/footprintjs/agentfootprint},
license = {MIT},
year = {2025}
}