@cubos/agent-sdk
v0.0.1152175
Published
Client for the Cubos Agent conversation API. Runs anywhere fetch does.
Readme
@cubos/agent-sdk
Client for Cubos Agent. Two surfaces: conversation, for your end-user app, and admin, for your backend.
Zero dependencies. Uses only fetch, ReadableStream and AbortController, so the same code runs in the browser, Node ≥ 18, Bun, Deno and Cloudflare Workers.
bun add @cubos/agent-sdk # or npm / pnpm / yarnThe token: minted by your backend, never in the client
A Cubos Agent api_key is tenant-wide and must never reach a browser. Your backend trades it for a short-lived token scoped to one user:
import { createAdminClient } from "@cubos/agent-sdk";
// On YOUR server.
const admin = createAdminClient({ baseUrl, apiKey: process.env.CUBOS_API_KEY! });
const { token, expires_at } = await admin.tenant("acme").userTokens.create({
user_id: session.userId, // from YOUR auth, never from the client
agent_slugs: ["support"],
ttl_seconds: 900, // optional; defaults to 3600
});The api_key needs the user_tokens:write grant. The resulting token reaches only the conversation endpoints, and only that user's own conversations — there is no way to widen it from the client.
agent_slugs lists the agents the token may open a conversation with. ttl_seconds defaults to 3600 and caps at 86400. There is no revocation list: prefer the shortest TTL your app can comfortably refresh. Blocking the user (users.block) invalidates their tokens on the next request.
Conversation
import { createUserClient } from "@cubos/agent-sdk";
const client = createUserClient({
baseUrl: "https://agent.acme.com",
// Called per request — cache it yourself. On forceRefresh, return a freshly
// minted token (the SDK only asks for that after a 401).
getToken: async ({ forceRefresh }) => fetchTokenFromMyBackend({ forceRefresh }),
});
const conversation = await client.createConversation({ agentSlug: "support" });
await client.sendMessage(conversation.id, "My invoice was charged twice");
const subscription = client.subscribe(conversation.id, {
onMessage: (m) => console.log(m.role, m.content),
onActivity: (a) => setTyping(a.isProcessing || a.hasPendingTurn),
onTodos: (todos) => setPlan(todos),
});
subscription.close();If you already hold a token — a script, a test — pass it directly as token: "…" instead of writing a closure. A static token cannot be renewed, so it eventually expires for good; use getToken for anything real.
subscribe backfills the history on connect and resumes from the exact point after any drop, so what it delivers is the whole conversation — no need to pair it with listMessages, which exists for paging long history backwards.
Telling the agent what is on screen
await client.setContext(conversationId, `screen: customers\nfilters: status=active`);Call it whenever the screen moves — it writes no event and starts no turn, so it is safe on every navigation and even mid-turn. The server shows it to the agent the next time a turn reads the conversation, and only when it changed, so repeating an unchanged context costs nothing.
That is the raw route. @cubos/agent-sdk-react's useConversation({ context })
goes one better and holds the value until a turn could read it — browsing eleven
screens before typing anything makes one request, not eleven.
The alternative most apps reach for is prefixing the context onto the user's message. That spends it on every message, puts it in the transcript unless every render path strips it out, and only works for whoever owns the composer.
onActivity rides that same stream, and the ordering is a guarantee you can build on: a frame saying the turn ended is never delivered before the reply that ended it. Turn the typing indicator off on it and the answer is already in hand. (It used to come from the conversation-metadata stream, a second connection with no ordering against the log, and the gap between them — milliseconds on a local socket, seconds through a proxy — showed a finished turn with nothing in it.)
For a long conversation that is the wrong trade: every reconnect re-delivers the whole log. Pass a cursor and the stream starts from there instead, leaving history to pagination:
const page = await client.listMessagesPage(id, { limit: 50 });
setMessages(page.messages);
// Only what is newer than the page. Nothing is lost in between: the server
// catches up everything past the cursor before going live.
client.subscribe(id, { onMessage: append }, { since: page.latestChangeSeq ?? undefined });
// Scrolled to the top — the page before this one.
if (page.hasOlder) {
const older = await client.listMessagesPage(id, { before: page.oldestSeq, limit: 50 });
prepend(older.messages);
}@cubos/agent-sdk-react wires exactly this into loadOlder / hasOlder.
The log itself, when the conversation is not enough
The conversation surface hides the event log on purpose: a chat draws messages,
not llm_call rows. But hiding is not the same as withholding — the same
subscription carries both, and an app that needs the log gets it from the same
place:
const page = await client.listEventsPage(id, { limit: 50 });
client.subscribe(id, {
onEvent: (e) => append(e), // every row, in stream order
onMessage: (m) => draw(m), // …and what it became, if you want both
}, { since: page.latestChangeSeq ?? undefined });One connection, one cursor, one catch-up. onEvent fires before the projections
see the frame, so an app taking both never has to reconcile them.
What differs is the promise, not the access. Message | Todo | Activity is
curated and stays put across internal refactors; ConversationEvent is the
server's own DTO and moves when the server does — the same trade the REST API
already offers, since these rows are what it returns. Reach for it to build an
operator console, an audit view or an activity tree; stay with Message to build
a chat.
@cubos/agent-sdk-react wires this into useConversationEvents.
Reopening is free
Conversations already loaded are kept in a bounded in-memory LRU, so opening one a second time paints with no request:
const start = await client.loadHistory(id); // start.fromCache on a repeat
client.subscribe(id, { onMessage: append, onCursor: (seq) => (cursor = seq) },
{ since: start.latestChangeSeq ?? undefined });
// Hand the current state back whenever it moves; the next open resumes from here.
await client.saveHistory(id, { messages, oldestSeq, latestChangeSeq: cursor, hasOlder });A hit is not stale. The server's event log never rewrites a row's content,
seq or type, and every mutation that does happen — a tentative event being
consolidated or discarded, a delivery receipt — goes through a trigger that
issues a fresh change_seq. So subscribing with the cached cursor replays
exactly what changed while the app was away, inserts and updates alike.
Pass cache: null to switch it off, cacheMessageLimit to change how much is
kept per conversation (300 messages by default; older ones are dropped from the
cache, not the server, and come back through paging). To survive a reload,
implement ConversationCache over IndexedDB, AsyncStorage or SQLite:
interface ConversationCache {
read(key: string): Promise<CachedConversation | null>;
write(key: string, entry: CachedConversation): Promise<void>;
clear(key?: string): Promise<void>;
}Keys arrive already scoped to tenant and user — a persistent store is shared by every session on the device, and two users must not read each other's messages out of it. Treat them as opaque. A store that rejects is treated as a miss, so a corrupt cache degrades to the network rather than breaking the chat.
The tenant is resolved from the token via GET /me on the first call. Pass tenant in the options to skip that round-trip.
Pagination
When you want everything and would rather not manage a cursor:
for await (const conversation of client.iterateConversations()) { … }
// The 20 most recent messages, newest first.
const recent = [];
for await (const message of client.iterateMessages(id)) {
recent.push(message);
if (recent.length === 20) break;
}Breaking out does not fetch the next page.
pageSize counts events, not messages: the log also holds the agent's tool calls and bookkeeping, so a page can hold fewer messages than you asked for. The iterator accounts for that; a page with no messages at all does not end the walk.
Response components
Your app can let the agent answer with your components — a price chart, a date picker, an order card — instead of markdown alone. What the agent may use comes from component libraries, which an operator authors over the admin API; your app enables the ones the open screen can draw:
const conversation = await client.createConversation({
agentSlug: "support",
componentLibraries: ["support-widgets"],
});setComponentLibraries(id, slugs) replaces the list at any time — send the
complete one; an empty list takes the agent back to plain markdown. A change
applies from the agent's next turn.
The split is a trust boundary. A component's summary, prop descriptions and example are shown to the model verbatim, so writing them is writing prompt content and needs an api_key; enabling a library is only wiring, which is why this client — holding a short-lived end-user token — can do it and nothing more.
Both the enable call and its list counterpart answer with tags, every
component the enabled libraries resolve to, in the order the agent sees them
indexed. Check it against what you can actually render: a tag with no renderer is
a block the agent will happily write and nothing will draw.
Every agent message then arrives split into blocks, with props already parsed
and checked against the library's schema:
{message.blocks?.map((block, i) =>
block.type === "markdown" ? (
<Markdown key={i}>{block.text}</Markdown>
) : (
<MyComponent key={i} tag={block.tag} {...block.props} />
),
)}A component is always self-closing — there is no <Tag>…</Tag>. Whatever it
should display goes in a prop, and the markdown around it stays outside as its
own blocks:
// leia **isto**
// <PriceChart symbol="PETR4" points={[1,2]} />
[
{ "type": "markdown", "text": "leia **isto**" },
{ "type": "component", "tag": "PriceChart", "props": { "symbol": "PETR4", "points": [1, 2] } }
]Worth knowing:
blocksis always there on an agent message. A reply with no component is onemarkdownblock, so you write one rendering path, not two.- The server enforces the contract. A component no enabled library offers,
HTML, or props that miss the schema are refused before the message is written —
the agent gets the error and rewrites. Validation is the full JSON Schema, not
a shallow type check, and the refusal names the path that failed
(
points.0.v: "muito" is not of type "number"), so the agent can fix it. contentis still the whole message as the agent wrote it, if you would rather render it yourself.- Blocks are derived when read, never stored. Dropping a library does not rewrite old messages, so keep a renderer for anything the transcript may still contain.
- Libraries are enabled per conversation only where your app is the renderer. Doing it on a channel-backed conversation (WhatsApp, Telegram) is refused — an external channel declares its libraries on the channel instead.
- Interactivity is yours: the agent emits the block, your component handles the
click, and whatever it should mean goes back through
sendMessage.
Live conversation list
const page = await client.listConversations();
client.subscribeToConversations({ onConversation: (c) => upsertAndResort(c) });A conversation is re-emitted whenever its activity advances, which is what keeps a chat list sorted without polling.
Images
await client.sendImage(id, file, { caption: "is this the right charge?" });
// Up to 10 in one message, so the agent reasons over the set rather than one
// turn per picture. A label names an image for the model, which lets it answer
// about "the receipt" instead of "the second image".
await client.sendImages(id, [
{ image: receipt, label: "receipt" },
{ image: statement, label: "statement" },
], { caption: "compare these" });png, jpeg, webp and gif, up to 10 MB each. The agent sees the picture natively when its model has vision, otherwise a description from the agent's fallback vision model — either way it costs a turn and is budgeted as one.
An image arrives as a Message whose attachments describe the media and whose content is the caption. Sending without a caption is normal, so render attachments even when content is empty. The bytes come back as a Blob, not a URL, because the token travels in a header — an <img src> pointing at the route would arrive unauthenticated:
const blob = await client.fetchAttachment(conversationId, message.id, attachment.id);
img.src = URL.createObjectURL(blob); // revoke it when the element goes awayWhen the fallback vision model describes a picture, that description stays out of the transcript: it is the model talking to itself about something already on screen, not a message the user sent.
Voice messages
await client.sendAudio(conversation.id, blobFromMediaRecorder);The clip is transcribed by the agent's STT model before the turn runs. It arrives as a Message carrying one audio attachment, with content holding the transcription — empty until STT finishes, which is the window where a client shows the player with no text under it yet. Bytes for playback come from fetchAttachment too.
Client tools
Tools with no server-side implementation: the agent calls one, its turn suspends, and your app answers.
const session = await client.serveClientTools(id, {
tools: {
get_location: {
description: "Where the user is, as their device reports it.",
inputSchema: { type: "object", properties: {} },
readOnlyHint: true,
handler: async () => ({ city: "Fortaleza" }),
},
},
});That declares the set, watches the conversation for calls, leases each one so two tabs don't run it twice, renews the lease while a slow handler works, and retries the result POST — the answer is worth more than one attempt once the side effect has happened.
It opens its own event stream to do the watching. If you already subscribe to the conversation, don't pay for a second connection:
const session = await client.serveClientTools(id, { tools, watch: false });
client.subscribe(id, {
onOpen: () => session.poke(), // catches up after a (re)connect
onClientToolCall: () => session.poke(), // and on every new call
onMessage: append,
});setClientTools declares without running anything — useful right before the
first message, since a tool the agent was never told about can't be called in the
turn that follows. @cubos/agent-sdk-react does all of this for you.
Conversation surface
| | |
|---|---|
| me() / refreshIdentity() | the token's user, tenant and agents |
| listConversations() / iterateConversations() | one page with a cursor, or all of them |
| createConversation() | opens a conversation with an agent |
| getConversation() / renameConversation() / archiveConversation() | |
| listMessages() / iterateMessages() | history |
| listMessagesPage() | history plus the cursors for paging back and for subscribe({ since }) |
| loadHistory() / saveHistory() / forgetHistory() | the conversation cache |
| sendMessage() / sendAudio() | the user's message, typed or spoken |
| sendImage() / sendImages() | one image, or up to 10 in one message |
| setClientTools() / serveClientTools() | declare your functions, and run the calls |
| setComponentLibraries() / listComponentLibraries() | which component libraries the agent may use |
| fetchAttachment() | an attachment's bytes, as a Blob |
| steer() | injects an instruction mid-turn |
| subscribe() / subscribeToConversations() | live state |
All of them accept an AbortSignal.
The exposed types (Conversation, Message, Attachment, Block,
EnabledComponents, Todo, Activity) are a curated
surface, narrower than the server's DTOs: the event log has dozens of types that
exist for the operator dashboard, and pinning those here would make every
internal refactor a breaking change.
Admin
The whole API, authenticated with an api_key, namespaced by resource and scoped per tenant:
import { createAdminClient } from "@cubos/agent-sdk";
const admin = createAdminClient({ baseUrl: "https://agent.acme.com", apiKey });
await admin.tenant("acme").agents.list();
await admin.tenant("acme").users.block(userId);
await admin.tenants.list();
await admin.apiKeys.rotate(keyId);Types come from the server's OpenAPI document (Schemas["Agent"], …),
regenerated on every route change — so the client tracks the server on its own.
admin.raw is the escape hatch for a route not yet wrapped.
Server-side only. An api_key is tenant-wide: it can read and write every
conversation of every user in the tenant. Use createUserClient in a browser.
Errors
Everything the SDK throws extends AgentError, so one catch covers the two
cases that matter:
try {
await admin.tenant("acme").agents.get("support");
} catch (err) {
if (err instanceof AgentNetworkError) {
// Never reached the server: wrong URL, server down, CORS, or the timeout
// (`err.timedOut`). `err.cause` holds the original failure.
} else if (err instanceof AgentApiError) {
// The server answered, and refused.
if (err.isNotFound) …
if (err.isConflict) … // 409
if (err.isAuthError) … // 401/403
if (err.isRetryable) … // 429 or 5xx
}
}AgentApiError carries status, requestId and the verbatim body
(err.json() to parse it). AgentConfigError comes out of client construction
rather than a call — usually a baseUrl with no http://.
Every request has a 30s timeout; change it with timeoutMs, or 0 to disable.
Streams are exempt, since staying open is the point.
Rate limits
The server budgets end-user tokens per instance: roughly 20 agent turns and 300
other calls per minute, plus a cap on concurrent streams. A refusal is a 429 with
Retry-After, and the SDK handles it — it waits the advertised interval and
replays the request, twice by default. A 429 means refused, not half-applied, so
replaying a POST is safe.
It gives up and throws when the retries run out, when there is no Retry-After
to honour, or when the wait would exceed 20 seconds — blocking a caller for
minutes is worse than telling them now. maxRetries: 0 opts out. An abort during
the wait cancels the retry rather than finishing it.
api_keys are exempt from all of this.
Transient stream failures do not end a subscription: the SDK reconnects with
exponential backoff and reports them through onError, for logging or a
"reconnecting" hint.
Three failures that look like silence get handled rather than reported. A
401 buys one immediate reconnect with a forced token refresh — a short-lived
token expiring under a stream that outlives it is the expected case, not a
misconfiguration — and only a second rejection is final. A 429 — the per-user
cap on open streams, which one tab too many reaches — waits out the server's
Retry-After and reconnects, since the cap clears the moment any tab closes. A
connection that stops sending anything, keep-alive comments included, is
dropped and reopened after idleTimeoutMs (30s, twice the server's keep-alive
interval): a half-open socket otherwise leaves the reader waiting forever with
no error to react to.
And a stream has to prove itself before it is trusted. The server sends a
frame on every connect, so until a connection has delivered one — and again
whenever one ends — subscribe polls the same log every 3s through the after
cursor, and the conversation row with it. On a healthy network the frame lands
inside the first poll's 2s grace and no poll is ever made. On a network whose
proxy holds a text/event-stream response until it ends (corporate TLS
inspection, some antivirus — Chrome negotiating HTTP/1.1 with a host that speaks
h2 is the tell), every connection sits silent until something cuts it, and the
reply used to show up only after a page reload; now it shows up within 3s. The
same rule covers the gap between a drop and the reconnect that follows. One
cursor serves both paths, so nothing is delivered twice.
What remains final — a 4xx that survived the refresh, a deleted conversation —
closes the subscription for good and fires onFatal (onError sees it too).
That is the one a UI has to surface: everything still renders, and nothing will
ever arrive again.
Also in the package
@cubos/agent-sdk/sse— the Server-Sent Events reader on its own (reconnect, cursor replay, backoff), for consumers who talk to the API directly and want just that piece.@cubos/agent-sdk-react— the same conversation surface as React hooks, in a separate package. Hooks only: it renders nothing, so the chat UI stays yours.@cubos/agent-sdk-react-dom— the few components that are generic enough to share, starting with the agent's markdown. Web only, and it styles itself.
