@atlanai/sdk
v0.4.0
Published
Atlan Agent Gateway management SDK
Readme
@atlanai/sdk
TypeScript/JavaScript SDK for the Atlan Agent Gateway: manage agents, skills, sessions, and workspaces, and optionally trace what your agents do.
Install
npm install @atlanai/sdkQuickstart
import { AtlanClient } from "@atlanai/sdk";
const client = new AtlanClient({
gatewayOrigin: "https://<your-gateway-host>",
bearerToken: "<your-api-token>",
});
const agent = await client.agents.create({
name: "support-triage",
workspaceId: "workspace_01example",
});
const page = await client.agents.list({ limit: 10 });
console.log(`${page.items.length} agents`);Methods read as client.<resource>.<action> — client.agents.get(agentId),
client.sessions.messages.create(sessionId, {...}), and so on. Every public
Agent Gateway operation is reachable this way; nothing requires reaching into
a generated client directly.
Pass a default workspace once instead of repeating it on every call:
const client = new AtlanClient({ gatewayOrigin, bearerToken, workspace: "workspace_01example" });Run an eval
Eval() executes the task and scorers, creates the Registry experiment, emits
one trace per case, uploads trace-linked results, and finalizes the run. Registry
derives the durable score summary during that finalization:
import { Eval } from "@atlanai/sdk";
const result = await Eval("Anthropic Evaluation", {
data: () => [
{ input: "What is 2+2?", expected: "4" },
{ input: "What is the capital of France?", expected: "Paris" },
],
task: async (input) => callModel(input),
scores: [{
name: "accuracy",
scorer: ({ output, expected }) => output === expected ? 1 : 0,
}],
}, { subjectKind: "agent", subjectId: "agent_01example" });
console.log(result.experimentId);subjectKind and subjectId link the run to the agent it tested, which is
how the Evals UI finds it; the agent ID (agent_...) is the id that
client.agents.list() returns. A run with neither is still saved, but lists
under no agent, and Eval() warns when it creates one.
Set ATLAN_API_KEY and ATLAN_WORKSPACE_ID. Replace inline data with
dataset: "dataset_..." or one exact Registry dataset name. The runner refuses
to proceed when tracing is disabled because trace-less eval rows break the
evidence chain. After flushing, it reads every case trace back through the
experiment filter before it uploads results; a missing trace marks the
experiment failed. A Registry dataset case passes the value at the dataset's
inputKey (default question) to the task and unwraps a single-key
{ value: ... } expected value. A run that produced no usable score at all is
finalized and then refused; pass requireScores: false if an empty score set
is expected.
Scorer return values
A scorer returns one of these. The scorer's own name is the metric name unless the value carries one.
() => 0.8 // a number (true counts as 1, false as 0)
() => null // abstain: recorded, but not averaged
() => ({ score: 1, name: "exact", metadata: { why: "match" } })
() => [{ score: 1, name: "a" }, { score: 0, name: "b" }] // a list of { score, ... } objectsA list holds { score, name?, metadata? } objects, not bare numbers. A plain
object needs a score key.
By default a scorer that throws marks its whole case as errored. Pass
{ scorerErrors: "score" } in the options to keep the failure to that scorer:
its metric is recorded as null. The message is stored on the case result (and under
sourceMetadata.atlan_eval.scorer_errors on the persisted row), and the case
stays successful unless no scorer produced a number. Only an exception thrown by
the scorer counts: an unsupported return value, a duplicate or colliding score
name, or a failed registration still errors the case.
Run part of a dataset
recordIds, where and limit run a subset of a Registry dataset. The run
stays linked to the dataset, so each result keeps its datasetRecordId and
the run can be compared with a full run of the same dataset on the records both
ran. They apply in that order to the experiment's frozen snapshot, and a resume
must pass the same selectors. where must be a synchronous function that
returns a bool.
await Eval("Support smoke", {
dataset: "support-daily-tasks",
task,
scores,
}, {
where: (item) => (item.tags ?? []).includes("billing"),
limit: 10,
});Control an eval lifecycle directly
Resolve an existing dataset by artifact ID or exact name, then create the Registry experiment before another runner emits any traces:
import { createContextManifest, startExperiment } from "@atlanai/sdk";
import { propagateAttributes } from "@atlanai/sdk/tracing";
const contextManifest = await createContextManifest([{
kind: "file",
name: "CLAUDE.md",
version: "git:0123456789abcdef0123456789abcdef01234567",
digest: `sha256:${"0".repeat(64)}`,
}]);
const run = await startExperiment(
client,
"conversational-studio-daily", // exact name, or dataset_... ID
{ name: "candidate-run", config: { model: "example-model" } },
{ contextManifest },
);
await propagateAttributes(run.traceOptions, () => existingRunner());run.experiment is the generated create response and run.id is its
experimentId. The helper pins the resolved dataset version and context
manifest in the immutable experiment config. Use it when another harness owns
execution, result upload, and finalization.
Tracing
import { initLogger, wrapTraced } from "@atlanai/sdk/tracing";
const logger = initLogger({ projectName: "support-agent" });
const handleRequest = wrapTraced(async (input: string) => callModel(input));
await logger.traced(async (span) => {
const output = await handleRequest("hello");
span.log({ input: "hello", output });
}, { name: "handle-request", type: "task" });
await logger.flush();Tracing has its own API key and lifecycle. It does not reuse the management
client's transport. It ships as a separate subpath
(@atlanai/sdk/tracing) so importing it doesn't pull OpenTelemetry into a
management-only bundle. npm still installs the OpenTelemetry dependencies with
the package because npm cannot make dependencies conditional on an imported
subpath; splitting tracing into a second package would trade that install size
for a second version and release lifecycle.
Vercel AI SDK v3-v6 uses its native OpenTelemetry path:
import * as ai from "ai";
import { initLogger, wrapAISDK } from "@atlanai/sdk/tracing";
initLogger({ projectName: "support-agent" });
const { generateText, streamText } = wrapAISDK(ai);See evals and tracing for framework setup, context manifests, serverless flushing, and the tested integration matrix.
Errors
Every non-2xx response rejects with AtlanAPIError, with .status, .code,
and (where the gateway includes one) .traceId. A failure before any response
is also normalized as AtlanAPIError with status === 0:
import { AtlanAPIError } from "@atlanai/sdk";
try {
await client.agents.get("agent_does_not_exist");
} catch (error) {
if (error instanceof AtlanAPIError) {
console.log(error.status, error.code);
}
}Runtimes
Node ≥20. The tracing subpath is written to also run in non-Node runtimes
like Cloudflare Workers — pass configuration to init() explicitly there
rather than relying on environment variables, since Workers has no ambient
environment.
