floe-guard
v0.15.1
Published
A local budget guardrail for AI agents — hard-stops your agent before its next LLM call crosses a USD ceiling. Vercel AI SDK middleware + LiveKit / Vapi / Retell voice adapters.
Downloads
1,996
Maintainers
Readme
floe-guard (Vercel AI SDK)
A local budget guardrail for AI agents — the TypeScript counterpart to the
Python floe-guard. It hard-stops your agent before its next
LLM or paid tool call when it would cross a USD spend ceiling — tokens and
tool calls under one local ceiling. No account, no signup, no network. Runs in
your process.
Works with both AI SDK v4 and v5 (ai@4 / ai@5).
npm i floe-guard ai @ai-sdk/openaiimport { wrapLanguageModel } from "ai";
import { openai } from "@ai-sdk/openai";
import { BudgetGuard, budgetGuardMiddleware } from "floe-guard";
const guard = new BudgetGuard(5.0); // your ceiling, in USD
const model = wrapLanguageModel({
model: openai("gpt-4o"),
middleware: budgetGuardMiddleware(guard),
});
// generateText / streamText with `model` now stop at $5 — the call that would
// cross the ceiling throws `BudgetExceeded` BEFORE it runs.The middleware sits in the call path: it check()s before doGenerate /
doStream (throwing BudgetExceeded to halt the run) and record()s priced
token usage after — for streaming it reads usage from the finish part.
Pricing
Tokens are priced offline from a bundled
LiteLLM cost map. A model that isn't in the map (and has no
manual price) fails closed: record throws UnpriceableModelError rather
than silently treating spend as free — you can't cap spend you can't measure.
const guard = new BudgetGuard(5.0, {
priceOverrides: {
"my-self-hosted-model": { inputCostPerToken: 1e-6, outputCostPerToken: 2e-6 },
},
// or failClosed: false to warn-and-skip for models you accept un-metered.
});Context-aware budgeting
guard.advisory() returns a soft signal you can act on before a call — nearLimit,
usedBps, remainingUsd — so the agent can taper near the cap instead of being
cut off. The hard-stop (check) is still the guarantee. This is the same shape
the Python package exposes and that hosted Floe returns on every proxied call
(X-Floe-Budget-Advisory), so the logic ports unchanged to the hosted path.
const guard = new BudgetGuard(0.1, { nearLimitBps: 7000 }); // flag at 70% used
const adv = guard.advisory();
const model = adv.nearLimit ? openai("gpt-4o-mini") : openai("gpt-4o");Request-sized estimates
To ensure the ceiling is enforced on the first run or for a call much larger than the previous one, you can price the actual incoming request using estimateCall() and pass the estimate to reserve() or check():
const est = guard.estimateCall("gpt-4o", 12_000, 4_096);
const handle = guard.reserve(est); // throws BudgetExceeded NOW if this call alone would cross
try {
const response = await callYourLlm({ model: "gpt-4o", ... });
guard.settle("gpt-4o", response.usage.promptTokens, response.usage.completionTokens, { reserved: handle });
} catch (err) {
guard.release(handle);
throw err;
}If the model is unpriceable, estimateCall() returns undefined and reserve(undefined) / check(undefined) fall back gracefully to the last-cost prediction.
Tool spend under the same ceiling
Paid tool calls (Apollo, Exa, scrapers) burn the same budget as tokens. The full reserve/settle contract applies — and the price is known before the call, so the pre-call hard-stop is exact:
const handle = guard.reserveTool(0.02); // throws BudgetExceeded BEFORE the call
const result = await apollo.peopleLookup(...);
guard.settleTool("apollo.people_lookup", 0.02, { reserved: handle });
guard.recordTool("exa.search", 0.004); // post-hoc, for metered APIs
guard.toolCosts; // { "apollo.people_lookup": 0.42, "exa.search": 0.11 }Token ceilings and per-step budgets
Cap total token usage — every bucket the guard counts: prompt, completion, and cache (a token ceiling) and keep one step of a sequential loop from starving the rest (a per-step cap) — a second dimension on the same reserve/settle machinery, not a second guard:
import { BudgetGuard, TokenBudgetExceeded } from "floe-guard";
// aggregate token ceiling alongside the USD one
const guard = new BudgetGuard(100, { tokenLimit: 20_000 });
guard.check(undefined, { estimatedTokens: 1_200 }); // throws TokenBudgetExceeded if it'd cross
guard.record("gpt-4o", 800, 400); // tokens accrue for free from the counts
// a per-step cap for one step of a sequential loop (callback — g IS guard)
guard.step({ maxTokens: 5_000 }, (g) => {
g.record("gpt-4o", 3_000, 1_500);
g.check(undefined, { estimatedTokens: 1_000 }); // 4_500 + 1_000 > 5_000 → scope "step"
});
const adv = guard.advisory();
adv.tokenUsedBps; // aggregate token utilization (null if no tokenLimit)
adv.remainingTokens; // tokens left before the ceiling (null if no tokenLimit)
adv.stepRemainingTokens; // active step's headroom (null if no step, or its token cap is unset)TokenBudgetExceeded extends BudgetExceeded, so budget-aware retry treats a
token block as terminal automatically. With no tokenLimit and no step(), USD
enforcement is unchanged and reserve() still returns a plain number — a
BudgetReservation handle appears only when tokens are actually reserved or a
step is active. (advisory() gains the token/step fields above; additive and
null when their dimension is unused.)
Per-call spend log
The guard keeps a typed, in-memory ledger of everything it priced: each
record() / settle() appends one SpendEvent, and recordTool() lets paid
non-LLM calls spend the same budget and land in the same log. The events sum to
spentUsd (unless a maxLogEvents ring buffer has evicted old ones).
const guard = new BudgetGuard(1.0); // { maxLogEvents: N } caps memory
guard.record("gpt-4o", 1_200, 350, { label: "researcher" });
guard.recordTool("serpapi.search", 0.01, { label: "researcher" });
guard.spendLog; // [{ timestamp, kind: "llm", modelOrTool: "gpt-4o", … }, …]
process.stdout.write(guard.exportLog()); // JSONL, one event per lineexportLog() emits a stable snake_case schema —
{timestamp, kind: llm|tool, model_or_tool, prompt_tokens, completion_tokens,
cost_usd, label?, reserved?} — identical to the Python package's
export_log(), so every agent produces the same shape regardless of stack.
Compatibility
ai is declared as a peer dependency with the range >=4.0.0 <6.0.0:
ai@4—LanguageModelV1MiddlewareviawrapLanguageModel/experimental_wrapLanguageModel; usage read frompromptTokens/completionTokens.ai@5—LanguageModelV2MiddlewareviawrapLanguageModel; usage read frominputTokens/outputTokens.
The middleware imports nothing from ai at runtime or in its types, so one
build serves both majors. If a provider reports no usable token counts, the
call is rejected (fail-closed) rather than metered as $0.
Development
npm install
npm run build
npm test
npm run typecheckLicense
MIT — see ../LICENSE.
