@actaclad/agentguard
v3.2.2
Published
Agent Guard observability and guardrails SDK for Node.js
Downloads
956
Readme
@actaclad/agentguard
Observability and runtime guardrails for AI agents in Node.js/TypeScript.
One init() auto-instruments your LLM clients — every call is traced (into
the AgentGuard console) and guarded (server-configured guardrails run on the
prompt and the response, and can redact or block). Covers OpenAI,
Anthropic, AWS Bedrock, LangGraph.js, MCP tools, and OpenAI voice.
npm install @actaclad/agentguardQuick start (auto-instrumentation)
# .env
AGENTGUARD_PUBLIC_KEY=pk-lf-...
AGENTGUARD_SECRET_KEY=sk-lf-...
AGENTGUARD_BASE_URL=https://<your-agentguard-host>
AGENTGUARD_PROJECT_ID=<project-id> # enables live guardrail-config pollingimport OpenAI from "openai";
import * as agentguard from "@actaclad/agentguard";
agentguard.init(); // once, at startup — keys/base/project from env
agentguard.instrumentOpenAI(OpenAI); // explicit class = ESM/bundler-safe (recommended)
const openai = new OpenAI();
// This call is now traced AND guarded — no other code changes.
const res = await openai.chat.completions.create({
model: "gpt-4o",
messages: [{ role: "user", content: "summarise the crew roster" }],
});
await agentguard.flush(); // flush telemetry before the process exitsGuardrails are configured in the AgentGuard console and fetched live (ETag polling, fail-safe to last-known-good). The SDK only enforces what the server enables, so you tune policy without redeploying.
Why
instrumentOpenAI(OpenAI)and not justinit()? In ESM/bundled apps,require("openai")and yourimport OpenAIcan resolve to different module copies, so the auto-patch may miss the class you actually use. Passing the exact class you imported removes that ambiguity. The same applies toinstrumentAnthropic(Anthropic).
Providers
OpenAI (chat + voice)
import OpenAI from "openai";
agentguard.init();
agentguard.instrumentOpenAI(OpenAI); // covers chat completions AND audio (STT/TTS)- Chat completions — prompt runs input guardrails; response runs output guardrails; streaming is scanned chunk-by-chunk.
- Voice —
audio.transcriptions.create(STT): the transcript runs the output guardrails.audio.speech.create(TTS): the input text runs the input guardrails (redact/block before synthesis).
Anthropic (incl. Bedrock/Vertex)
import Anthropic from "@anthropic-ai/sdk";
agentguard.init();
agentguard.instrumentAnthropic(Anthropic); // also covers AnthropicBedrock / AnthropicVertexAWS Bedrock (Converse / ConverseStream / InvokeModel)
Two ways to instrument, depending on how you get your client:
import * as Bedrock from "@aws-sdk/client-bedrock-runtime";
import {
instrumentBedrock,
instrumentBedrockClient,
} from "@actaclad/agentguard/bedrock";
agentguard.init();
// (a) You own the client instance — add middleware to it:
const client = new Bedrock.BedrockRuntimeClient({ region: "us-east-1" });
instrumentBedrockClient(client);
// (b) You DON'T hold the client (e.g. LangChain's ChatBedrockConverse builds its
// own internal one) — patch the class so EVERY client is covered:
instrumentBedrock(Bedrock);Bedrock model ids are normalized so cost matches server-side (Claude ids resolve to their canonical names).
Azure, Bedrock and Vertex clients report their real provider (azure,
bedrock, vertex). List those models under the matching provider in the
Allowed Model List guardrail.
If your code runs the tools itself (for example a Bedrock Converse tool loop),
wrap each call with runTool(). See Tool security.
LangGraph.js
createLangGraphHandler() returns a LangChain callback handler that
reconstructs the full graph → node → llm/tool trace, wired to your init()
credentials with identity propagation.
It brings no LangChain packages with it. @langchain/core is an optional peer
used for types only — the handler satisfies LangChain's callback contract
structurally — so installing this SDK never adds a LangChain version to your
tree or a resolution conflict to your install.
import { createLangGraphHandler } from "@actaclad/agentguard/langgraph";
agentguard.init();
const handler = createLangGraphHandler({ userId, sessionId, feature });
await graph.invoke(input, { callbacks: [handler] });One linked trace (recommended). With both tracing (handler) and enforcement
(instrumented model clients) active, pair them via startTrace so everything
lands in a single trace instead of two:
const session = agentguard.startTrace({ name: "crew-triage", userId });
const handler = createLangGraphHandler({ root: session.root });
await session.run(() => graph.invoke(input, { callbacks: [handler] }));
console.log(session.url ?? session.id);Enforcement inside graph nodes comes from the model layer, not the handler (LangChain callbacks are observe-only).
ChatOpenAIis covered via the sharedopenaipackage;ChatBedrockConverseviainstrumentBedrock(...).
LangGraph internals are filtered out. LangGraph emits a run for every channel
write, branch selector and wrapper sequence — in a two-node graph, six of eleven
observations. They are dropped so the tree reads like the graph you wrote; model
calls re-parent onto the node above them. Pass traceInternals: true to keep
them. The console's flow diagram is drawn from node metadata and is unaffected.
Automatic tracing (no callbacks argument). init() registers a LangChain
configure hook so every run is traced without passing a handler. In an ESM
application — Next.js, or any "type": "module" project — you must register it
yourself:
import * as agentguard from "@actaclad/agentguard";
import * as langchainContext from "@langchain/core/context";
agentguard.init();
agentguard.instrumentLangChain(langchainContext);We cannot do this for you, for the same reason instrumentOpenAI exists:
@langchain/core ships separate CommonJS and ESM builds, and this SDK's
CommonJS build reaches the wrong one. Registration must also be synchronous
in your own async context, because LangChain stores configure hooks in
AsyncLocalStorage. Passing a handler explicitly needs none of this and works
everywhere, which is why it stays the recommended path.
Tool security
An agent reaches the outside world through two doors: the model call and the tool call. Instrumenting a provider covers the first. Tools are separate, and an unwrapped tool is traced but not enforced — the callback handler can see it and cannot stop it.
import { guardTools, runTool } from "@actaclad/agentguard/tools";
import { instrumentMcpClient } from "@actaclad/agentguard/mcp";
agentguard.init();
const tools = guardTools([search, refundOrder]); // LangChain / LangGraph tools
instrumentMcpClient(mcpClient); // MCP: after connect()
// Your own tool loop (e.g. Bedrock Converse):
const result = await runTool(use.name, use.input, (args) =>
myTools[use.name](args),
);Per call: (1) a BLOCK / un-approved APPROVAL_REQUIRED tool is refused
before it runs (even if the model asked); (2) arguments are
injection/secret/PII-scanned (block or redact); (3) the result content is
PII/injection-scanned (block or redact) before it reaches the model; (4) the
call is recorded as a tool.<name> span under the graph step that ran it.
Server-side defense-in-depth is available via instrumentMcpServer(server).
For JSON arguments and results, prompt-injection and toxic-content read only the sentences (three or more words). Dates, IDs and codes are not scored, so plain tool data is not flagged.
Narrowing the tool door for latency
Steps 2 and 3 each cost one call to the scan service, so a guarded tool call costs two. Tools that run in parallel scan concurrently; sequential graph nodes pay serially, and that is where a long run adds up.
Both entry points take policy.disable, which drops named guardrails at the tool
boundary only — the model door is unaffected. With no sidecar-backed guardrail
left, the scan request is not made at all:
// 2 scan calls per tool call -> 0
guardTools(tools, {
policy: { disable: ["pii-redaction", "toxic-content", "prompt-injection"] },
});
instrumentMcpClient(mcpClient, {
policy: { disable: ["pii-redaction", "toxic-content", "prompt-injection"] },
});Tool permission still enforces — it is evaluated locally from your config,
not through the scan service, so a BLOCK tool is still refused before it runs.
Before switching both off, note the two directions are not worth the same:
| | Protects against | Worth keeping when | | ------------------------ | -------------------------------------------------- | ------------------------------------------------- | | Arguments (outbound) | Exfiltration — data leaving your process to a tool | The tool is third-party | | Results (inbound) | Injection and PII entering the model's context | Almost always — this is the untrusted content |
Tool results are where indirect prompt injection actually arrives, and where a query drops customer records into the conversation. If you cut one, cut arguments.
Policy (per-request overrides)
withPolicy sets request-scoped identity and overrides (propagates across
await). Anything not passed inherits the parent scope.
import { withPolicy } from "@actaclad/agentguard";
await withPolicy(
{
userId: req.user.id,
sessionId: req.session.id,
feature: "triage",
onBlock: "refuse",
},
async () => {
// all instrumented LLM/tool calls in here are tagged with this identity
return handleRequest();
},
);userId,sessionId,feature,metadata— identity on the trace.onBlock—"raise"(default; throwsAgentGuardBlocked) or"refuse"(returns a synthetic provider-shaped refusal). Applies to both directions: an output-side block always replaces the offending text first — it never reaches you — and then raises or returns according to this setting.disable— guardrail names to skip for this scope.
Guardrails
Guardrails are turned on and configured in the console. init() registers all
of them; there is nothing else to install or call.
- Built in: allowed-model-list, budget-guard, rate-limit, token-limit, tool-permission, execution-limit-guard, pii-redaction, secret-scanner, token-scanner, prompt-injection, toxic-content, schema-validation, hallucination. Plus your own custom policies from the console.
- Where they run:
pii-redaction,prompt-injectionandtoxic-contentare scored by your AgentGuard host in one batched/api/public/guardrails/scancall. If the host can't be reached they let the call through and log a warning. The rest run in your process. - Hallucination is judged through your own model client (OpenAI, Anthropic or Bedrock), so no extra key is needed.
- Passwords are an opt-in secret type: select them under the secrets scanner's secret types in the console.
Observe mode
agentguard.init({ guardrails: "observe" });Every guardrail runs and records its findings, but nothing is blocked; the console shows those findings as WOULD BLOCK. Redaction still applies. Use it to try a policy before you enforce it.
Kill switch
Turn on the kill switch in the console to block every model call in the project
with your message. The SDK checks it every 10 seconds (killSwitchPollInterval),
so it takes effect without a redeploy. If a check fails, the last known state
stays in force.
registerMlGuardrails() (@actaclad/agentguard/guardrails-ml) is kept for
future on-device guardrails. It does nothing today and you don't need to call it.
Prompt management
const prompt = await agentguard.getPrompt("movie-critic");
const text = prompt.compile({ criticLevel: "expert", movie: "Dune 2" });Production is fetched by default; pass { label: "staging" } or { version: 3 }
to select another deployment.
Configuration
Environment variables (an explicit init({ ... }) argument always wins):
| Variable | Purpose |
| ------------------------------------------------- | ---------------------------------------------------------------- |
| AGENTGUARD_PUBLIC_KEY / AGENTGUARD_SECRET_KEY | API keys (AGENT_GUARD_* aliases accepted) |
| AGENTGUARD_BASE_URL | AgentGuard host (legacy AGENTGUARD_HOST accepted) |
| AGENTGUARD_PROJECT_ID | Enables live guardrail-config polling |
| AGENTGUARD_CAPTURE_CONTENT | true to record prompt/response text on traces (off by default) |
| AGENTGUARD_INCLUDE_INFRA_SPANS | Include infra spans in traces |
| AGENTGUARD_KILL_SWITCH_POLLER_ENABLED | false to turn off the 10-second kill-switch check |
Main init() options: projectId, environment (a label on traces),
guardrails (auto · observe · off), onBlock (raise · refuse),
tracing (batch · realtime), pollInterval (seconds, default 900),
killSwitchPollInterval (seconds, default 10).
Subpath exports
| Import | What |
| ------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------- |
| @actaclad/agentguard | init, flush, instrumentOpenAI, instrumentAnthropic, instrumentLangChain, startTrace, withPolicy, getPrompt, errors |
| @actaclad/agentguard/bedrock | instrumentBedrock, instrumentBedrockClient |
| @actaclad/agentguard/langgraph | createLangGraphHandler |
| @actaclad/agentguard/mcp | instrumentMcpClient, instrumentMcpServer |
| @actaclad/agentguard/tools | guardTool, guardTools, runTool |
| @actaclad/agentguard/guardrails-ml | registerMlGuardrails |
| @actaclad/agentguard/langchain | AgentGuardCallbackHandler — deprecated, use /langgraph |
Peer deps (openai, @anthropic-ai/sdk, @aws-sdk/client-bedrock-runtime,
@modelcontextprotocol/sdk, @langchain/core) are
optional — install only the ones you use. Their version ranges are
deliberately open, so this SDK never constrains yours.
Manual tracing API (legacy)
The AgentGuard class is the low-level, explicit tracing API (traces / spans /
generations / scores). Auto-instrumentation above is preferred for new code, but
this remains available:
import { AgentGuard } from "@actaclad/agentguard";
const ag = new AgentGuard({ publicKey, secretKey, baseUrl });
const trace = ag.trace({ name: "chat-request", userId, input: { message } });
const gen = trace.generation({ name: "llm-call", model: "gpt-4o" });
const answer = await callLLM(message);
gen.end({ output: answer });
trace.update({ output: answer });
await ag.flush();