@altheris/sdk
v0.2.0
Published
Runtime security for AI agents — scope enforcement, prompt-injection screening, output redaction, human approvals, cost circuit breakers, honeypots.
Maintainers
Readme
@altheris/sdk
Runtime security for AI agents. Wrap your LLM client in one line and every call is screened for prompt injection, held to the agent's declared scope, filtered for leaked secrets/PII, gated behind human approval where you've required it, cost-tracked against a circuit breaker, and subject to an emergency kill switch — enforced by the Altheris platform, not by prompts.
Quick start
npm install @altheris/sdkPut the agent's alsk_ SDK key in your .env — never commit this file:
ALTHERIS_API_KEY=alsk_your_sdk_keyThen wrap the LLM client you already construct:
import { protect } from "@altheris/sdk";
const client = protect(new OpenAI(), { agentId: "your-agent-uuid" });protect() also takes onBlocked (every block event) and onError (telemetry
and fail-open availability errors). With no onBlocked, each tool call the
scope gate strips is logged once with console.warn, naming the tool and the
reason — pass onBlocked to receive the event instead, or onBlocked: () => {}
to silence it.
protect() auto-detects OpenAI-like, Anthropic-like, LangChain runnable, and Vercel AI model clients, and throws on anything it does not recognise — it never leaves a client silently unprotected.
Every chat.completions.create call now runs the full security pipeline.
Get your alsk_ SDK key from Dashboard → Agent → SDK key.
Advanced: explicit configuration
protect() is the one-call path and reads ALTHERIS_API_KEY from the
environment. Construct the class directly when you need to configure anything
else — a key provider function, baseUrl, the other callbacks, or close():
import { Altheris } from "@altheris/sdk";
const altheris = new Altheris({ apiKey: "alsk_…", agentId: "your-agent-uuid" });
const openai = altheris.protectOpenAI(new OpenAI());What you get on every call
| # | Feature | How | |---|---------|-----| | 1 | Emergency lockdown (kill switch) | polled every 30s; locked agents refuse all calls | | 2 | Conditional access policies | time windows enforced locally; IP/geo server-side | | 3 | Prompt-injection screening | last user message → platform pattern corpus | | 4 | Runtime scope enforcement | every call checked against allow/deny action globs | | 5 | Human-in-the-loop approval | flagged actions wait for an approver (or time out) | | 6 | Output filtering / redaction | credit cards, SSNs, API keys, JWTs, passwords… | | 7 | Cost circuit breaker | usage reported per call; runaway spend auto-pauses | | 8 | Decision provenance | forensic record of every protected call | | 9 | Honeypot canaries | bait URLs detected → agent suspended instantly | | 10–13 | Config versioning, behavioural baselining, threat signals, audit trail | server-side, fed by the SDK |
Adapters
OpenAI
import OpenAI from "openai";
import { Altheris } from "@altheris/sdk";
const altheris = new Altheris({ apiKey: process.env.ALTHERIS_API_KEY!, agentId: AGENT_ID });
const openai = altheris.protectOpenAI(new OpenAI());
const res = await openai.chat.completions.create({
model: "gpt-4o",
messages: [{ role: "user", content: userMessage }],
});
// res.choices[0].message.content is already redacted if it leaked anythingWraps chat.completions.create and legacy completions.create. Streaming
(stream: true) is fully supported — see Streaming.
The legacy completions.create has no tools parameter and returns no tool
calls, so the tool gates are N/A there by API shape — not a coverage gap.
Anthropic
import Anthropic from "@anthropic-ai/sdk";
const anthropic = altheris.protectAnthropic(new Anthropic());
const msg = await anthropic.messages.create({
model: "claude-sonnet-4-6",
max_tokens: 1024,
messages: [{ role: "user", content: userMessage }],
});Wraps messages.create. String and content-block-array messages both
screened; response text blocks filtered individually.
LangChain (JS 0.3.x)
const chain = prompt.pipe(model).pipe(parser);
const safeChain = altheris.protectLangChain(chain);
const answer = await safeChain.invoke({ input: userMessage });Wraps invoke, batch and stream of any Runnable. Inputs may be
strings, { input | question | text | content | query } objects, or message
arrays. Cost is best-effort from usage_metadata /
response_metadata.tokenUsage when the chain surfaces them.
Response tool calls are screened (default on). The tool_calls on a
returned message — and the tool_call_chunks / tool_calls streamed by
stream — are checked against the pinned tool registry and then against the
agent's intent grants before your executor loop can iterate them. A call that
fails either check is removed (streamed calls are buffered and only emitted
once they pass), and each removal fires onBlocked. Disable with
toolRegistry: false / toolCallScopeGate: false.
Because this fails closed, note the LangChain-specific consequence: LangChain
binds tools inside the Runnable, so if Altheris cannot sniff them and you did
not pass tools, nothing was pinned — and registry enforcement will then strip
every tool call in the response. That is on top of the loud onError the
adapter already fires when sniffing fails. Pass the tools explicitly:
const safeChain = altheris.protectLangChain(chain, { tools: [myTool, otherTool] });Vercel AI SDK
import { openai } from "@ai-sdk/openai";
import { generateText } from "ai";
const model = altheris.protectVercelAI(openai("gpt-4o"));
const { text } = await generateText({ model, prompt: userMessage });Wraps the model object (doGenerate/doStream, v1 and v2 specs), so
generateText, streamText, generateObject and streamObject are all
covered with one wrap.
Response tool calls are screened (default on). The toolCalls[] / content[]
tool-call parts on a doGenerate result and the streamed tool-call parts on
doStream are checked against the pinned tool registry and then against the agent's
intent grants before the AI SDK's tool-execution loop ever sees them — so a call
that fails either check never runs. Streamed calls are buffered while they arrive and
emitted only once they pass; a denied call reaches your stream in no form at all, while
text parts keep flowing live. Disable with toolRegistry: false /
toolCallScopeGate: false.
Removals are reported server-side by BOTH layers, and both fire onBlocked, but
they count differently — worth knowing before you count events. A scope-gate
removal fires onBlocked once per removed call, so two calls to the same
ungranted tool give you two events; a registry removal fires once per distinct
blocked name, so duplicate calls to the same blocked tool collapse into a
single event. Server-side, a registry removal is reported to report-tool-block
as it always was, and a scope-gate strip sends one report-tool-block request
per response (checkpoint: "tool_call_scope", carrying each stripped tool with
its intent and the reason it was refused) — fire-and-forget, never awaited, never
on your response path. And a tool call whose name cannot be read at all is
stripped fail-closed by the registry path without an event, because there is
no name to report.
Because this fails closed, note the consequence for the AI SDK: with the gates
on by default, a tool the server has not mapped to an intent — or an agent with
no intent grants at all — has every one of its tool calls stripped, on both
doGenerate and doStream. Absence is denial, not an oversight: a tool is
callable only once its intent is ticked for the agent in the dashboard.
Anything else — generic protect()
const safeAgent = altheris.protect(myAgent, {
inputMethods: ["sendMessage"], // 1st string arg → input screening + scope gate
outputMethods: ["sendMessage", "receiveResponse"], // primary string of result → filtering
});Proxy-based: the original object (and its TypeScript type) is untouched;
named methods get the pipeline. The result's "primary string" is the value
itself if it's a string, else the first of text / content / output /
response / message / answer / result.
toolMethods are gated at invocation (default on). This is the one adapter
where a tool's actual EXECUTION can be stopped — the method is invoked through
Altheris, so a blocked call simply never runs, rather than merely being stripped
from a response. Two checks fire, registry first: the tool's pinned registry
status (a revoked or suspended tool throws before it executes), then the Layer 2
intent scope gate. Each block throws AltherisBlockedError with
context.executed === false and fires onBlocked. Disable with
toolRegistry: false / toolCallScopeGate: false.
The scope gate here has a deliberate limit — read this before relying on it.
Unlike the other adapters, it enforces only when both of these hold: an
intent snapshot has been received from Altheris (i.e. this agent has made at least
one check-action, which means listing the method in inputMethods too), and
you declared tool definitions via the tools: config. Anything less passes
through unenforced. That is intentional: a generic toolMethod is only a
method name, and a toolMethods-only agent never calls check-action, so
applying absence-is-denial unconditionally would block every tool method of every
generic agent that never touched the tool-definition workflow. Registry status
enforcement is unaffected and applies regardless. To get scope enforcement on a
generic agent, declare tools: and list the method in inputMethods as well:
const altheris = new Altheris({ /* … */ tools: [processRefund, sendEmail] });
const safeAgent = altheris.protect(myAgent, {
inputMethods: ["processRefund", "sendEmail"], // supplies the intent snapshot
toolMethods: ["processRefund", "sendEmail"], // …which the invocation gate reads
});Configuration
const altheris = new Altheris({
apiKey: "alsk_…", // required — the agent's SDK key
agentId: "…", // required — the agent's UUID
endUserId: () => currentUser(),// string | () => string — explicit end-user identity,
// for per-end-user escalation. Never inferred; ≤256
// chars; omit for anonymous (escalation is skipped,
// never widened to the whole agent).
baseUrl: "https://altheris.io",// default
failureMode: "fail-safe", // "fail-safe" (default) | "fail-open"
timeout: 5000, // per-request ms
cacheConfigTTL: 60000, // agent-config cache
lockdownPollInterval: 30000, // kill-switch poll
approvalPollInterval: 5000, // poll while waiting on a human
approvalTimeout: undefined, // default: the approval's own expiry
onBlocked: (e) => {}, // { phase, reason, agentId, detail }; default warns once per scope-stripped tool call
onRedacted: (e) => {}, // { agentId, redactions, postStream? }
onError: (err) => {}, // telemetry + fail-open availability errors
debug: false,
});fail-safe vs fail-open. failureMode governs availability failures
only. fail-safe (default): if Altheris is unreachable, the call is blocked
with AltherisNetworkError. fail-open: the call proceeds and onError
fires. Server verdicts (blocked input, out-of-scope action) and the human
approval gate always bind in both modes — an unreachable approval flow
never defaults to "approved".
Error handling
import {
AltherisError, // base — code, retryable, context
AltherisBlockedError, // .phase: "input"|"action"|"policy"|"honeypot", .reason
AltherisLockdownError, // kill switch engaged
AltherisApprovalRejectedError, // .approvalId, .reason
AltherisApprovalExpiredError, // retryable: true — re-request the action
AltherisAuthError, // bad alsk_ key / wrong agent — fix config
AltherisNetworkError, // Altheris unreachable — retryable
} from "@altheris/sdk";
try {
await openai.chat.completions.create(/* … */);
} catch (e) {
if (e instanceof AltherisBlockedError) {
return `Request refused (${e.phase}): ${e.reason}`;
}
if (e instanceof AltherisNetworkError && e.retryable) {
/* back off and retry */
}
throw e;
}Streaming
Streams are passed through live and unaltered — redaction cannot be
applied retroactively to chunks the caller has already received. The SDK
accumulates the streamed text and runs output screening at stream end; if
anything sensitive surfaced, onRedacted fires with postStream: true and
the detection is logged server-side. Input screening and the scope/approval
gates still run before the stream opens. If post-hoc detection isn't
acceptable for your use case, use non-streaming calls.
Cleanup is guaranteed even if you leave early. If the consumer stops a
stream before it ends — break, return, a thrown error, or cancelling the
stream — the SDK still screens what was delivered and reports cost +
provenance. In that case the screened text is only what the consumer actually
received (not the full LLM output), and onRedacted carries partial: true
so you can tell the difference:
const altheris = new Altheris({
apiKey, agentId,
onRedacted: (e) => {
if (e.partial) log.warn("sensitive content in a partially-read stream", e.redactions);
},
});
for await (const chunk of stream) {
if (enough(chunk)) break; // cleanup still runs on the delivered text
}Honeypots
Canary URLs from your agent's config are detected automatically in any text the pipeline screens — taking the bait suspends the agent immediately. Two manual hooks for layers the adapters can't see:
await altheris.scanForHoneypots(urlAboutToBeFetched); // throws on a canary
await altheris.reportHoneypotTrigger(honeypotId, { path }); // from your HTTP layerLatency
Budget: <150ms SDK overhead per call; in practice the warm path is two
parallel POSTs before the LLM call (input + action screening, both edge-
deployed) and one after (output). Agent config is cached 60s, lockdown
state 30s (background poll), cost/provenance reporting never blocks. Call
altheris.close() on shutdown to stop the poller (it's unref'd, so it
won't hold your process open either way).
Notes
- Node ≥ 18.17. Zero runtime dependencies. ESM and CommonJS.
- The SDK never mutates your framework's response objects — redacted
responses are clones (class instances like LangChain's
AIMessageare rewritten in place to preserve their prototype). - Full docs: https://altheris.io/docs (coming soon).
