mandala-computer
v0.9.0
Published
TypeScript SDK for Mandala Computer — cloud desktops for AI agents
Maintainers
Readme
mandala-computer
TypeScript SDK for Mandala Computer — cloud desktops for AI agents.
A real Linux desktop your code can see and drive: screenshots come back as bytes, clicks go in as coordinates, and a shell in the guest is one call away.
Status: alpha. The surface is settling; expect breaking changes before 1.0. Tracks the platform's
/api/v1, which is itself still moving.
Zero runtime dependencies. Node 22+, and anywhere else with fetch — Bun, Deno,
workers, the edge. (The mandala CLI is Node-only; the library is not.)
Install
npm install mandala-computerPublished as ES modules, with type declarations alongside. The install also
puts a mandala command on your PATH; see The mandala CLI.
To install only the CLI, for use outside a project, use Homebrew or the installer script:
brew install mandalacomputer/tap/mandala
# or
curl -fsSL https://mandala.computer/install.sh | shBoth install this same package. The formula brings Homebrew's own node; the
script needs Node 22+ and npm already present, never uses sudo, and falls
back to ~/.local when npm's global prefix is not writable. It takes
--version and --prefix after sh -s --.
Sign in once on your local Node installation:
mandala login
# Or request exactly one workspace and save a named profile:
mandala login --profile Work --workspace ResearchCompare the code shown in your terminal with the browser approval page, confirm
the account and scope, and approve. The first command requests access to the
whole account. The CLI always prints the URL and code for manual or headless use.
A successful login saves a normal device-named API key in
~/.mandala/credentials.json; it never prints the key. Revoke it in Settings →
Credentials → API keys when access is no longer needed. Treat the file like a
password and never ship a credential to browser users.
You can also create a key in Settings and use MANDALA_API_KEY or apiKey.
Resolution happens once when constructing a client: a supplied apiKey wins,
then a nonempty MANDALA_API_KEY, then the credentials file. An explicit empty
or wrong-type key fails instead of falling through. Empty environment values
are absent. A key supplied explicitly or through the environment performs no
credential-file or home-directory lookup, even when profile is set.
File selection uses new Client({ profile: 'Work' }), then MANDALA_PROFILE,
then the saved default. Profile names are case-sensitive ASCII names of 1–64
characters, starting with a letter or digit and continuing with letters, digits,
dots, underscores or hyphens; object-prototype names are reserved. A file key
is bound to its saved API base, including its complete path prefix. An explicit
or environment base override must match that canonical base. Invalid stores,
permissions, missing profiles, or mismatched bases fail locally before a request.
Existing clients retain their selected key after the file changes; a new client
reads the latest file. A revoked key raises the normal AuthenticationError
with status 401 and its revoked reason, without fallback, replay, or automatic login.
Local file authentication requires POSIX owner/mode verification: an owned real
0700 ~/.mandala directory and an owned regular 0600 file with one link.
Symlinks, hardlinks, unsafe modes, malformed JSON, files above 64 KiB, and more
than 100 profiles are refused. Windows file protection is unsupported in this
version; explicit/environment keys still work. Browser and default package
conditions support explicit keys and injected fetch; they never import the
local filesystem or use browser storage for profiles.
The same file can be read by the Python SDK and local MCP server. After login, no API-key environment variable is needed for local clients:
mandala computers list --profile Work
claude mcp add mandala -- npx -y mandala-computer-mcp --profile WorkRequests go to https://app.mandala.computer/api/v1; MANDALA_BASE_URL or
new Client({ baseUrl }) points them elsewhere, and apiKey is the same
option for the key. timeoutMs on the client is the per-request budget — 60
seconds unless a call knows it needs longer, 0 to disable — and fetch takes
an implementation of your own if you have proxies or certificates to configure.
Every request carries User-Agent: mandala-computer-ts/<VERSION> node/<version>;
userAgent: 'my-app/1.2' appends your own token to it. In a browser, which does
not let a page set that header, none is sent.
The timeoutMs a wait takes is a different number and is documented with each:
it bounds the whole loop rather than one request, and can be far longer, because
what those wait for outlives any single request. Each method takes
a signal among its options, so any one request can be cancelled.
Use
import { Client } from 'mandala-computer';
const client = new Client(); // environment key or saved default profile
const c = await client.computers.launch({ template: 'base' });
try {
await c.open('https://example.com'); // on the screen, not as root
const png = await c.screenshot();
await c.click(640, 400);
await c.type('hello');
} finally {
await c.delete();
}launch() creates once, waits for the disk, starts the computer if needed, and
returns when its guest agent answers and, on Linux, when its desktop session
exists (waitForDesktop()), so a first exec(..., { desktop: true }) is not
refused as having no desktop session. It accepts every create() option;
start: false is sent unchanged to create, then launch starts the computer after
its disk is ready. An already admitted start is waited on, and a failed start is
reported without being retried in that call (a launch resent under the same key
reads the computer afresh and may start it again; see Operations).
With secrets bound, it also waits until they
have reached the desktop, so the first command on the returned computer sees
them; a delivery that failed throws, naming why. With a browser proxy it waits
until the browsers have it (waitForBrowserProxy()), and with an egress proxy
that names credentials until the host holds them (waitForEgressProxy()):
until then every connection the computer opens is closed.
Pass { timeoutMs: 600_000, signal } as the second argument for a larger build or
cancellation. The default readiness budget is 180,000 milliseconds, beginning
after create returns. Disk, running, guest, secrets, proxy and desktop waits share the remaining budget,
including elapsed start work. Create and start retain their usual transport
deadlines, so this is not a total wall-clock limit on launch. pollMs defaults
to 3,000 for all stages.
The returned computer is persistent. A failed or cancelled launch can leave a
computer behind; SDK errors after creation include its id and keep their type
(including TimeoutError). Cancellation preserves the caller's original reason.
No failure automatically deletes the computer. The finally above cleans up
after a successful return; use the id in a readiness error to inspect or delete
a computer whose launch failed.
For cleanup tied to a callback, use ephemeral():
await client.computers.ephemeral({ template: 'base' }, async (c) => {
await c.waitForGuest();
await c.open('https://example.com');
}); // destroyed here, even if the block threwcreate() returns as soon as provisioning responds and never deletes anything.
Use it when you want to manage each readiness stage yourself.
A create takes a name, and start: false leaves the computer stopped. Finding
one again is computers.get(id) or computers.list(); a handle you already
hold is re-read with c.refresh(), and renamed with c.rename('staging'). Every
field on the handle — status, os, cpu, ramMb, createdAt and the rest —
is what the last payload said, and c.raw is that payload.
On a runtime with explicit resource management, ephemeral also works as a
disposable:
await using c = await client.computers.ephemeral({ template: 'base' });
await c.waitForGuest();If the block throws and the cleanup delete fails too, both errors arrive —
a SuppressedError whose .suppressed is the block's own error, the fault to
read first, and whose .error is a MandalaError naming the machine that is
still billable. Both spellings fill those two fields the same way, so a cleanup
failure is never the thing that goes unmentioned.
Read .error, not .message: the top-level message is the one field the two
spellings cannot agree on, because the runtime writes its own generic text over
it for await using. And on Node 22, where SuppressedError is not a global,
the same three fields arrive on a plain Error named SuppressedError — so
err.name === 'SuppressedError' is the portable test and instanceof is not.
A 404 from the cleanup is not one of these: the block deleted the machine itself, and its own error stands alone.
Every computer is a Linux desktop today. Windows guests are not offered on any plan; where this README mentions Windows it is describing behaviour the client already supports for when they are.
Sizes
size names a template and a CPU/RAM/disk shape together, and these are the
shapes the platform keeps pre-booted — so naming one is the likeliest way to get
a computer in about a second rather than a cold boot.
for (const s of await client.sizes.list()) {
console.log(s.id, s.label, s.cpu, s.ramMb, s.allowed ? '' : `needs ${s.cheapestPlan}`);
}
const c = await client.computers.create({ size: 'large' });allowed is about your plan's per-computer ceilings only — what the account
already holds is not counted, so a create at an allowed size can still be refused
against the plan's pools.
It cannot be combined with template, cpu, ramMb or diskGb. Sending both
throws before any request is made.
Your own templates
A template is a mandala/v1 document — the image family it resolves to, what it
is layered onto, and the shape a computer gets when the create names no numbers.
Publishing one gives it a ref you can launch by name.
const doc = await readFile('devbox.yaml', 'utf8');
// Worth doing while you iterate: this reports EVERY problem at once, and claims
// no ref. It does not throw for an invalid document — that is the answer.
const check = await client.templates.validate(doc);
if (!check.valid) throw new Error(check.problems.join('\n'));
const t = await client.templates.publish(doc);
const c = await client.computers.create({ template: t.ref });If create returns a ConflictError whose body has
code: "template_image_preparing", inspect body.preparation.state and
body.preparation.error before deciding to continue. The server's valid
Retry-After header is exposed as err.retryAfterMs; it is undefined when
absent or malformed. After that delay, repeat the original create body,
including the same nonempty template, and add body.template_transfer as
templateTransfer. Preserve the token exactly; it must be a nonblank string
and cannot be combined with size.
The token preserves the selected image and is not an idempotency key. Stop
after a successful create, and do not automatically replay after an ambiguous
response. isTransient returns false for this preparation conflict because
continuing requires the token, and the SDK never retries the create for you.
A valid answer carries more than the verdict. docDigest identifies the
document: it changes with anything that changes what the document means, a
label included, and not with comments, key order, whitespace or YAML vs JSON.
buildDigest covers only what decides the image, so comparing it against a
previous run tells you whether an edit means a rebuild. Publishing an invalid
document is a 400 that lists every problem too, as err.body.problems.
buildDigest is present only for a document with no spec.from, which builds
nothing (build steps need a parent to layer onto). A document that names a
parent has no buildDigest — it cannot be computed without the parent's —
and gets buildDigestNeeds in its place, a sentence saying what is missing.
The two are never both present, so read the second when the first is absent
rather than treating the absence as unexplained:
if (!check.buildDigest) console.log(check.buildDigestNeeds);canonical is the document as the digests were taken over it, key order and
whitespace normalised — hash it yourself to check docDigest rather than
trusting the platform to have done it honestly. template is the catalogue row
the document describes, in the same shape templates.list() answers — so what a
validated document would look like in a picker is readable before it is
published.
The namespace is your account. metadata.namespace has to be your account
id — anything else is a 403, system included — and this SDK does not rewrite
it, because publishing a ref that is not the one in your file would be worse than
refusing.
A ref is immutable. Publishing the identical document again succeeds and
changes nothing, so a pipeline that republishes on every commit is safe.
Publishing a different document under the same ref is a ConflictError; bump
metadata.version. What counts as different is the digest, so a changed label is
a change. Retrying does not help with that conflict, nor with a retired ref or
either of the account's template ceilings: every ConflictError from
publish() carries a permanent reason (exists where the platform sent none),
so isTransient answers false for it.
Read one back — yours or system, so you can see what you are layering onto:
// The namespace is your account id — the same one your document's
// `metadata.namespace` names. `ref` is where to read it back off a publish.
const namespace = t.ref.split('/')[0] ?? '';
const base = await client.templates.get('system', 'base');
const pinned = await client.templates.get(namespace, 'devbox', { version: '1.0.0' });Without version you get the newest, which is also what a create naming the
unpinned namespace/name resolves to.
templates.list() is the catalogue — the images a computer can be created
from, each with the ref a create names it by — and templates.schema() is
the JSON Schema for a mandala/v1 document, returned as it is so an editor or
a validator can be pointed at it. The URL it is served from needs an API key,
so an editor cannot fetch it by its $id: save what templates.schema() (or
mandala templates schema) returns to a file and point the editor at the
file.
Retiring one
await client.templates.retire(namespace, 'devbox', { version: '1.0.4' }); // one version
await client.templates.retire(namespace, 'devbox'); // every versionOmitting version retires the whole name — deliberately not get's "the
newest", which on a delete would let a loop walk backwards through a history it
never asked about. An empty string is refused before it is sent, for the same
reason.
Computers are not affected. A computer is built from the image the ref resolved to and holds no reference to the document, so anything already running, stopped or suspended is untouched. What a retire breaks is resolution: a new create naming the ref is refused.
The ref stays spoken for, and still counts once. Publishing it again is a
ConflictError, identical bytes included, and refsClaimed on the result does
not go down — it is the count against a much larger, separate ceiling than
templates. A ref you retired is a NotFoundError whose message names the date
it went, rather than claiming the template never existed; read the message before
concluding you mistyped something.
Building one
A document that declares spec.build steps or spec.env has to be compiled
into an image before anything can launch it. That is minutes of work — an agent
image is roughly fifteen — so it never blocks:
const build = await client.builds.start(doc);
const out = await client.builds.wait(build.id);
if (out.status !== 'succeeded') {
const failed = out.steps.find((s) => s.status === 'failed');
console.error(`step ${failed?.n} (${failed?.kind} ${failed?.label}) failed: ${out.error}`);
}wait does not throw for a build that failed. succeeded and failed are
two situations with two remedies — one has an image, the other has a step to fix
— and an exception flattens them into "something went wrong". Read status.
Identical documents share an image, which is what makes a repeated build cheap;
builds.start(doc, { noReuse: true }) builds again regardless. The namespace
and the spec.family both have to be yours, and either one that is not is a
PermissionDeniedError; spec.from has to name a system/... template, or it
is a 400. A ConflictError means the host is busy — one build runs per host at
a time — or that a secret's value could not be read just now, and is worth
retrying. builds.get(id) is the job, builds.progress(id) is what it is doing
and stays readable after it has finished, and builds.list() is every build the
fleet still holds a record of.
A build's steps can read secrets: list them under spec.secrets (on by
default), each by its id and the variable to read it as, never by value. They
are resolved in your key's scope — a workspace's own first, then the account's —
each frozen at the revision current when the build is submitted, and a document
may name at most 32 (templates.validate() does not check that limit). A
malformed reference is a 400 that says what is wrong; one that does not
resolve is a 400 that does not say which; neither is worth retrying. A 503
saying secrets are not available is worth retrying, like the 409 above.
For a terminal, stream it instead of polling:
for await (const p of client.builds.events(build.id)) {
console.log(`${p.phase} ${p.step}/${p.of} ${p.note}`);
}Each event is news — the platform sends one only when something moved — and the
last one is the done, including for a build that failed.
The loop above throws for three reasons, all of them about the stream rather than
the build: an error event, a final event whose payload is malformed, and a
stream that ends without a final event at all. The last two matter because
returning quietly would make a cut stream indistinguishable from a finished
build. All three say the build is probably still running and point at
builds.progress. Breaking out early is not one of them — that closes the stream
and throws nothing. An account may hold eight streams open at once.
After a successful build, launch the published template with
client.computers.create({ template: t.ref }). A pinned template version
selects the latest successful image built from that exact document; an
unpinned ref first selects the newest published template version. Existing
computers keep their original image when a later build becomes available.
A create may need to prepare that image before it can launch. Follow the
template preparation guidance for a
template_image_preparing response: inspect its state and error, then make a
deliberate continuation with the returned delay and token and all original
create arguments. For other image refusals, inspect the response before
deciding whether to rebuild or retry.
Two of those refusals never clear, and both arrive as a 400 — a bare
APIError rather than the UnavailableError a 503 becomes, because a retry
loop reads a 503 as an answer worth waiting for (isTransient says yes to
it) and these are not. A document that builds into a family this account may
not build into is one: no hypervisor will launch it, and what changes the
answer is publishing a new version, not sending the create again. A document
that builds nothing and names an account family is the other: a create is
served only from the families that ship with the product.
Resolution
Create-time, and only create-time: the screen is part of the machine QEMU builds, so changing it needs a new computer.
const c = await client.computers.create({ template: 'base', resolution: '1920x1080' });
const { width, height } = c.screen; // what every coordinate is inUnreachable, deleted and lost records can have no screen to report. Check
c.resolution directly before reading c.screen; when it is empty, c.screen
throws ValidationError:
if (c.resolution) {
const { width, height } = c.screen;
console.log(width, height);
}Read c.screen rather than assuming 1280x800. Computer-use accuracy is
resolution-sensitive, and a model told the wrong numbers clicks proportionally
short of everything it aims at:
const tool = {
type: 'computer_20250124',
name: 'computer',
display_width_px: c.screen.width,
display_height_px: c.screen.height,
};Driving the desktop
The verb set is Anthropic's computer tool, in full — so whatever a computer-use model emits, there is a method for it:
await c.move(100, 200);
await c.click(100, 200);
await c.click(); // where the pointer already is
await c.click(100, 200, ['shift']); // held for the click
await c.rightClick(100, 200);
await c.middleClick(100, 200);
await c.doubleClick(100, 200);
await c.tripleClick(100, 200); // selects a line in most editors
await c.drag(400, 300, { x: 100, y: 200 }); // one gesture, not two clicks
await c.mouseDown(100, 200);
await c.mouseUp(400, 300);
await c.scroll(640, 400, { direction: 'down', amount: 3 });
await c.scroll(640, 400, { direction: 'right', modifiers: ['shift'] });
const { mechanism } = await c.type('hello'); // 'physical' | 'unicode' | 'mixed'
await c.paste('Café — 東京 😀'); // clipboard + Ctrl+V; { shortcut: 'ctrl+shift+v' } for terminals
await c.key('ctrl', 'c'); // X11 keysyms work too: Page_Down, BackSpace
await c.key(['ctrl', 'c'], { signal }); // the same chord, cancellable
await c.holdKey(['shift'], 1.5);
await c.wait(2);
const at = await c.cursorPosition(); // undefined if nothing has placed itNo coordinate means "where the pointer already is", which is a real and different
request from clicking (0, 0). Half a coordinate — click(5) — is refused rather
than completed with a zero: it would succeed, at the wrong place, and nothing
would say so. Modifiers are a positional on the clicks and an option on
scroll, and the wrong spelling of either is refused rather than sent with
nothing held down. A drag with no from starts where the pointer is, and is
refused if nothing has placed it yet.
type takes 1 to 400 characters. Plain ASCII is typed as key events, about
12 ms a character; text with anything else in it is typed whole by a guest
helper using GTK Unicode composition, which works in supported Chromium GTK3 and
Xfce Terminal setups on Linux X11 and is refused, not typed partially,
elsewhere (Firefox, GTK4, Windows, Wayland). Unsupported text is refused before
a key is pressed; bare CR and other control characters are refused too. The
answer's mechanism says how it was typed. It confirms dispatch, not that the
application accepted the text, so check the result — and a failure part-way can
leave partial text, so look before retrying. For long text, paste writes the
clipboard (up to 8192 bytes) and presses the shortcut; it replaces the clipboard
and leaves it replaced, and a success likewise means delivered, not inserted.
Every one of these but cursorPosition takes { context: true } (on key, in
the array form) and then answers the desktop as it stands just after the
action: the windows windows() lists by default and the focused one, saving
a second request. An action that answers nothing resolves to that
InputContext instead of undefined, and type carries it as
result.context:
const after = await c.key(['ctrl', 'l'], { context: true });
after.focused?.title; // what has the keyboard now
const { mechanism, context } = await c.type('hello', { context: true });The windows are read once, straight after the action, with no settle wait, so a
window still opening may not be listed yet. When they cannot be read — a
Windows guest, no desktop session — windows is null and error says why;
the action itself still happened, so do not send it again.
Screenshots, and the one flag a drive loop needs
const png = await c.screenshot(); // full-resolution PNG
const thumb = await c.screenshot(320); // downscaled JPEG
const now = await c.screenshot(undefined, { fresh: true });
const liveThumb = await c.screenshot(320, { fresh: true }); // both, and honouredPass fresh whenever the image is feeding a decision. A bare screenshot may
be served from a cache up to 1.5 seconds old — which is what makes ten watchers
of one desktop cost a single screendump, and what makes a drive loop read the
screen as it was before its own last click. A model handed that frame concludes
the click missed and clicks again, and the second click lands on whatever the
first one revealed. A thumbnail can have the cached frame; a decision cannot.
fresh composes with a width. screenshot(320, { fresh: true }) is a downscaled
JPEG built from a capture taken after the request arrived — the cheap way to watch
a desktop that is actually moving. Earlier versions of this SDK refused that call
on the grounds that a width made the flag a no-op; it is honoured now, so it is
sent. Omit fresh, or pass false, to ask for the last frame already held.
A suspended computer serves only that cached frame: its saved desktop, a JPEG
up to 640 pixels wide whatever the width asked for, and even with none. It is a
stored picture rather than a screen to drive, so its pixels are not screen
coordinates. Asking a suspended computer for a fresh capture is refused — 409, telling you to start it first — with or
without a width, so a poller that suspends its own machine should drop fresh
rather than treat the 409 as the computer having vanished.
A crop, a smaller picture, a cheaper encoding
A full-resolution PNG is the most expensive frame to hand a model. region,
scale, format and quality shape it, applied in that order to the same
capture:
// The top-left quarter of a 1280x800 screen, halved, as a JPEG.
const corner = await c.screenshot(undefined, {
fresh: true,
region: { x: 0, y: 0, width: 640, height: 400 }, // screen pixels, before scaling
scale: 0.5, // 0 < scale <= 1; not with a width
format: 'jpeg', // 'png' by default ('jpeg' with a width)
quality: 60, // 1-100, JPEG only
});A cropped or scaled picture is in its own pixel space. To click on something in
it, divide its position by the scale and add the region's x and y: (100, 50)
in corner is (200, 100) on the screen. A region that reaches past the screen's
edge is refused with a 400 naming the screen size rather than clipped; the value
checks (a scale out of range or beside a width, a quality on a PNG) throw a
ValidationError before anything is sent.
A suspended computer has only its saved JPEG and cannot shape it: a crop, a
scale, format: 'png' or a quality is a ConflictError whose reason is
unavailable, which does not clear by waiting. Start the computer, or drop the
shaping (and fresh) for the saved picture; a width or format: 'jpeg' may stay.
Windows
A screenshot says what the desktop looks like; this says what any of it is, which is how you tell a browser that failed to open from one that has not painted yet. Linux only.
for (const w of await c.windows()) {
console.log(w.id, w.windowClass, w.title, w.focused, w.visible, w.pid);
}
const moved = await c.windowAction('0x2600003', 'move', { x: 100, y: 100 });
console.log(moved.window?.x, moved.window?.y); // 105, 129, probably
const shut = await c.windowAction('0x2600003', 'close');
shut.gone; // true — there is no window leftMatch on windowClass, not title: the class is the application, the title is
whatever page it is showing. visible is false for a minimised window,
which stays on the list — clicking at the coordinates of one puts the click on
whatever is actually there. Panels, the wallpaper and the rest of the desktop's
furniture are left off by default — a stock guest with one terminal open has
five windows, four of which are not applications — and
windows({ includeAll: true }) puts them back.
The actions are focus, raise, minimize, maximize, unmaximize,
close, move and resize; the geometry argument — x, y, width,
height — is what move and resize read.
pid is the process that owns the window, and is undefined where the guest
did not say — never 0, which is a pid a guest may genuinely advertise. It
does not identify the window: an application that keeps one process for
several windows — xfce4-terminal is one — reports the same pid on all of
them, so killing this pid can take windows you never asked about.
x, y, width and height are undefined on the same rule and for a
sharper reason: 0 is a place a window really is — the top-left corner — so a
coordinate this client could not read must not come back as one. The live route
sends all four on every window, so absent means something is already wrong, and
w.x ?? 0 is the wrong repair: there is no fallback for a place. A listing
carrying a window with no id is refused outright rather than handed back,
because every window action takes that id and a row without one names nothing
you can act on.
Prefer focus over raise: raising without focusing gives a window that is
visibly in front and silently not receiving keystrokes.
The reply is the window afterwards, not an acknowledgement — window managers
snap to their own grid, so a move to 300,200 routinely lands at 305,229 —
wrapped in a result rather than returned bare, because two outcomes have no
window to describe and gone is the only thing that separates them: true
after a close, and false when the action happened and the guest could not
describe what it left. The second is an outcome, not a failure, and not a reason
to repeat the action.
Two desktops answer this, and c.desktop says which. os is linux for a
Wayland guest and an X11 one alike, so it is the only field that tells them
apart — 'wayland', 'x11', or undefined from a host too old to have been
asked, which is not the same as 'x11'. Two things change with it:
c.desktop; // 'wayland' | 'x11' | undefined
// The call is the same on both, and on X11 it simply works. What Wayland adds
// is a refusal: a TILED window's geometry belongs to the compositor's layout,
// so this comes back a 400 rather than a move that quietly changed nothing.
// The message names the way past it — float the window (Super+V in a stock
// Omarchy) and the same call takes.
await c.windowAction('0x2600003', 'move', { x: 100, y: 100 });The other change is what a window's id is: a Hyprland client address
rather than an X window id. Both are 0x and hex and both are what
windowAction takes, so nothing on this API changes — but an id handed to
xdotool or xprop through c.exec() finds no window on a Wayland guest.
Clipboard
The desktop's CLIPBOARD selection — what Ctrl-C writes and Ctrl-V pastes — read
and written from outside the guest. Linux only, and it needs nothing of the
hardware: no cold boot, no permission from a browser. What it does need is
xclip in the guest, which every image built since August 2026 carries — so in
practice this is the road that works on every computer, and where it is not, the
refusal says so. (The other road is RFB extended cut text over the desktop
socket, which is live and conditional; see
Showing somebody the desktop.)
await c.setClipboard('https://mandala.computer');
await c.key(['ctrl', 'v']); // into whatever has focus
const onClipboard = await c.clipboard(); // '' is an empty clipboardsetClipboard() takes at most 64 KiB of UTF-8; clipboard() returns at most
128 KiB. They are different bounds on different channels, and the read is
refused rather than truncated past its own — half a password is not less of
an answer, it is a wrong one that looks completely normal. A NUL is refused
here, before the request goes out. setClipboard('') clears the clipboard, and
is the only way to: putting other text there only replaces one value with
another.
The platform confirms the write by reading the selection back before it answers,
so setClipboard() returning means the desktop is holding the text rather
than that a command ran.
Not every ConflictError here is worth retrying, and err.reason is how you
tell. contention is the one that clears by itself — the desktop did not
take the text means something else claimed the selection in that instant, a
clipboard manager settling, usually — and starting clears too, more slowly:
the guest agent has not answered inside its boot window yet. unavailable does
not clear at all, because the computer is not running and start() is the fix
rather than another attempt. Desktop-session and X-server failures carry no
reason, deliberately: the platform cannot tell a guest still coming up from a
logged-out desktop or a crashed window manager, so it offers no retry advice
there. Branch on the word, never on the sentence, which is prose and is
rewritten.
isTransient() reads it, so it no longer says true to the stopped computer —
which is what it used to do, and what a blanket retry loop spun on until its
deadline. An unclassified refusal falls back to the old type answer, so bound a
loop that meets one.
A 400 is the other one to know, because it never clears: a computer built
from an image that predates xclip is refused permanently. Install xclip in
the guest — you have root there — or create a new computer.
The two differ on one thing worth knowing: setClipboard() resumes a
suspended computer, because putting text on a clipboard is the first half of
pasting it and that is somebody working on the machine. clipboard() does not —
what somebody copied is not worth waking a machine for — so reading a suspended
computer is a 409 rather than a start you did not ask for.
A read failure is an exception, not an empty string. That is the distinction the
exec recipe these replace could not make.
Why not exec and xclip yourself
Because exec runs a login shell: the desktop user's profile is sourced,
and anything it prints lands on the same stdout as your command's output, ahead
of it. That is wanted when you asked to run a command the way the user would,
and fatal when you are reading a value — an echo in the guest's .profile
corrupts the answer and a deliberate one forges it, and no framing you add
fixes that, since a profile that prints your frame owns everything after it.
The clipboard endpoints do not share that stream. The write is worse still: an
X selection belongs to a live process, so the holder has to outlive the exec,
the text has to travel quoted, and being granted a selection is asynchronous,
so the result has to be polled for — each poll a billable exec. setClipboard()
does all of it in one call.
Running commands
const res = await c.exec('ls /home/user');
if (!res.ok) console.error(res.stderrText);
if (res.truncated) { /* the guest agent capped output at 16 MiB */ }stdout and stderr are Uint8Array — the bytes the command wrote — and
stdoutText and stderrText are those bytes decoded as UTF-8 for the ordinary
case of reading a line back. The platform sends both streams base64-encoded, and
this SDK hands you the bytes rather than a string, because a string is where the
output used to be quietly damaged: JSON strings are UTF-8 by definition, so a
command that printed a tarball, a PNG or a latin-1 build log came back with every
invalid byte replaced by U+FFFD, with a 200 and no flag saying so. The text
accessors do replace — a build log with one stray byte in it is still a log — so
reach for the bytes when the exact ones matter:
const png = await c.exec('cat /tmp/shot.png');
await writeFile('shot.png', png.stdout); // bytes, unaltered
console.log(png.stderrText); // text, when text is meantA non-zero exit is returned, not thrown. The guest gets timeoutS to finish —
30 seconds unless you say otherwise — and a command that outlives it keeps
running in the guest with its output unreachable. By default the command runs
as root with no display; anything with a window needs the desktop session:
await c.exec('nohup firefox https://example.com >/dev/null 2>&1 &', { desktop: true });env is the right way to hand a build a token, since the alternative is
interpolating one into the command line where the guest's shell history and
process list can both read it:
await c.exec('./deploy.sh', { cwd: '/src', env: { CI: '1', TOKEN: token } });On Linux those variables go on top of the guest's profile — PATH and the rest
survive, because the command runs through bash -lc. On Windows they replace
it: cmd.exe /c sources no profile, so the command sees exactly what you passed
and nothing else, PATH and SystemRoot included. Pass what it needs there.
Or call open() and let the SDK write that line — it picks the browser from
what the image has installed (firefox-esr, then firefox, then chromium),
quotes the URL, and detaches the launch:
await c.open('https://example.com');The images do not share one browser name: the Debian-based ones carry
firefox-esrand the Omarchy one onlychromium, so the line above that namesfirefoxopens nothing there, and still exits 0.open()looks before it detaches, and throws aMandalaErrorwhen the image has none of the three rather than returning a result that reads as success.
Linux only, and the platform is what says so. A desktop-session exec on a
Windows guest is refused before the computer is asked whether it is running,
with reason: "unsupported" — so isTransient reads it as settled and a retry
loop stops rather than starting the computer to ask again. There is no OS check
in the SDK: it knows only what the last payload said, and a computer whose os
never arrived would be refused for a command it could have run.
Long-running commands
Foreground timeoutS must be an integer from 1 through 600 seconds; the
SDK rejects larger or fractional values before sending the command. Use
execBackground for longer commands.
Hosted exec requests can still time out after about two minutes.
The HTTP budget is derived from it and the platform stretches its own deadline
to match, but a proxy in front of the platform abandons a request that has
produced no response for roughly that long and answers 524, which arrives as
GatewayTimeoutError. Measured against app.mandala.computer:
| command | timeoutS | result | wall clock |
|---|---|---|---|
| sleep 110 | 230 | ok | 110.6s |
| sleep 130 | 300 | GatewayTimeoutError | 125.2s |
The server's 600-second maximum does not extend this hosted proxy deadline.
The command also survives the request that abandoned it, so the
call after one of these often raises ConflictError — the guest agent still
busy with it, which is the first failure continuing rather than a second one.
So execBackground is not merely the tidier option past a few seconds; past two
minutes it is the only one that works. Strictly better than backgrounding with
&, which throws away the exit code and the output:
const job = await c.execBackground('apt-get install -y build-essential');
for (;;) {
const s = await c.execPoll(job.pid);
process.stdout.write(s.stdout); // only the NEW bytes, as bytes
process.stderr.write(s.stderr);
if (!s.running && !s.more) break;
if (!s.more) await new Promise((r) => setTimeout(r, 1000));
}
await c.execKill(job.pid); // if you change your mindA command can stop while several chunks of output remain. Keep polling while
more is true, and stop only once it is no longer running and its output is drained.
The output is a cursor, not a buffer: each poll gives you only what has been printed since the last one, so two readers on one pid split the output between them rather than each seeing all of it.
And it is cut at 1 MiB on a byte offset, so a chunk can begin or end part-way
through a multi-byte character. Write the bytes to a stream as above, or join
them and decode once at the end; s.stdoutText decodes each chunk on its own,
which is what you want for a line of output and lossy across a cut.
Independent execution reads
Newer platforms supply job.executionId on an accepted background command.
Older replies may omit it; the SDK never fabricates an ID or falls back to a PID
when you request a stable read. PID polling and killing above keep their existing
shared, consuming behavior and can address a newer command after PID reuse.
if (job.executionId) {
const signal = new AbortController().signal;
const observed = await c.execution(job.executionId, { signal });
console.log(observed.status);
// Each reader owns two independent byte positions; both are always explicit.
const first = await c.executionOutput(job.executionId, {
stdoutOffset: 0, stderrOffset: 0, limit: 65536, signal,
});
const next = await c.executionOutput(job.executionId, {
stdoutOffset: first.stdoutOffset,
stderrOffset: first.stderrOffset,
limit: 65536, signal,
});
for (const chunk of [first, next]) {
process.stdout.write(chunk.stdout); // Uint8Array, including NUL or partial UTF-8
process.stderr.write(chunk.stderr);
}
// Another reader can still read from zero, independently of these calls or execPoll.
}execution() reports the last observed running, exited, or lost state.
Only exited includes endedAt and a signed exitCode; running does not prove
that the computer is awake, and lost establishes no success or failure.
executionOutput() returns separate stdoutMore and stderrMore flags. False
means EOF at this instant, not final completion. Decode text with a streaming
TextDecoder across chunks to preserve split UTF-8. The separate diagnostic
bytes repeat in full on every read, may be truncated (diagnosticTruncated), and
never advance either stream offset.
These methods perform one request by default; opt-in safe GET retries reuse the same offsets. There is no automatic execution, resume, wait, output capture, or fallback. Output reads perform guest I/O and are unsuitable for passive history views. The files are mutable guest content, not retained artifacts. Handles disappear on platform/computer state loss, replacement or cleanup; observed exits expire after ten minutes. Unavailable reads throw the normal API error instead of returning an empty successful result.
Retained results and nominated artifacts
Retained versions are explicit, immutable snapshots with expiry. Background capture reads the guest's volatile output once; later metadata and byte reads use retained storage and do not resume the computer or poll a PID. Each explicit capture creates a new version, even for the same execution.
if (job.executionId) {
const captured = await c.retainExecutionOutput(job.executionId, {
maxBytesPerStream: 1024 * 1024, retentionSeconds: 86400, signal,
});
const metadata = await c.result(captured.resultId, { signal });
const page = await c.resultOutput(metadata.resultId, {
stream: 'stdout', offset: 0, limit: 65536, signal,
});
// page.bytes is authoritative, including invalid UTF-8 and binary data.
// The next independent read can use page.nextOffset. There is no shared cursor.
await c.deleteResult(metadata.resultId, { signal });
}
const finished = await c.exec('make test', { retainOutput: true, signal });
if (finished.resultId) {
const retained = await c.result(finished.resultId, { signal });
console.log(retained.kind); // synchronous-output
}retainOutput is synchronous exec only: absent or false keeps the ordinary
request; true or { maxBytesPerStream, retentionSeconds } requests retention.
Older servers and unconfirmed optional capture leave resultId absent. Missing or
malformed optional retention metadata does not invalidate the executed command.
No helper retries a capture or replays a command after an uncertain response.
RetainedResult distinguishes background-output from synchronous-output.
Background observations can still be running. Page EOF describes the retained
prefix, and ready describes stored bytes; neither means the task succeeded or
that all original output was retained. Synchronous prefixes report
sourceResponseBytes and upstreamTruncated separately from the retained
endReason. Their diagnostic is null; requesting that stream preserves the
server's 409 rather than fabricating empty output. Background diagnostics are
separate wrapper bytes and are independently paged like stdout and stderr.
Capture limits are 1..4 MiB per stream; retention is 1..604800 seconds (default
86400). A page is at most 65536 bytes. Metadata is finite and excludes unknown
response fields.
Artifact publication requires the caller to already know the exact guest path, byte count and SHA-256. It does not stat, read, hash or run a guest command to prepare the nomination. Paths remain byte-for-byte intact; Linux absolute paths, Windows drive-qualified paths and UNC forms are accepted for backend OS validation.
const artifact = await c.publishArtifact('/tmp/report.bin', {
expectedSize: reportSize, expectedSha256: reportSha256,
maxBytes: 8 * 1024 * 1024, retentionSeconds: 86400, signal,
// executionId: job.executionId, // optional caller selection, not provenance
});
const info = await c.artifact(artifact.artifactId, { signal });
const bytes = await c.downloadArtifact(info.artifactId, {
maxBytes: 8 * 1024 * 1024, signal,
});
await c.deleteArtifact(info.artifactId, { signal });downloadArtifact reads metadata and, if its size is within the independent
download cap, downloads the whole content. Each GET is one attempt by default;
opt-in transport retries may repeat an interrupted GET from the beginning.
It returns a Uint8Array only after exact length and SHA-256 verification
against that invocation's metadata.
The default download cap is 8 MiB; the maximum is 64 MiB. Publication's maxBytes
is a separate capture cap with the same default and maximum. Empty artifacts
still require the correct empty digest. Web Crypto SHA-256 support is required
and checked before transfer, including with an injected fetch in a browser or
worker. These helpers never follow a download URL or filename from the response,
write local files, return a partial prefix, use Range, or fall back to guest files.
New manifest and byte transports enforce caps during streaming, including an EOF
check at the exact cap. Invalid framing, incomplete content, hash mismatch,
authority loss and cancellation return no successful bytes. Explicit capture,
publication and the whole-content download request use at least a 90-second
request allowance; a larger configured client timeout remains larger, and
timeoutMs: 0 still disables the SDK deadline. Caller cancellation remains active.
This is a per-request allowance, not a deadline for a multi-request operation.
A lost publication response leaves its commit unconfirmed. Do not automatically
repeat the POST. deleteResult and deleteArtifact make one DELETE each; 204
returns void, while a repeated 404 remains unavailable. Activities and their
result-detail convenience methods remain outside this SDK surface.
Events
A computer says what it is doing. Waiting for something to happen is a socket, not a screenshot every second that mostly reports that nothing has changed:
await c.waitFor('computer.ready'); // the desktop is up
const done = await c.waitFor('process.exited'); // a background command ended
console.log(done.pid, done.exitCode);or the whole stream:
for await (const ev of c.events()) {
if (ev.type === 'window.opened') console.log('opened', ev.window?.title);
if (ev.type === 'process.exited' && ev.pid === job.pid) break; // closes the socket
}Windows opening, closing and taking focus; the clipboard changing hands; a
background command exiting; the desktop becoming ready; every power transition.
ev.data always holds the payload verbatim, and the fields worth reading are
promoted onto the event: window, windowId, pid, exitCode, lost,
selection, watch, path, kind, dir, armed, lostReason, status,
previous, idleSeconds, oldestCursor, detail.
When connecting your own WebSocket client, use the exact returned events_url,
including its desktop capability. A REST Bearer key alone is insufficient; a
REST call to the events path returns guidance to use events_url.
It keeps your place. Every event carries an opaque cursor, and the position
after the last event you actually consumed is what a reconnect resumes from —
so a socket that drops mid-loop does not lose the process.exited you were
waiting for. Each reconnect re-reads the computer for a fresh events_url,
because a restart rotates that credential and a restart is one of the ordinary
reasons the socket dropped. stream.cursor is that position if you want to keep
it across a process restart; pass it back as since.
Where the host can no longer replay that far you get a gap event rather than
silence. It is not an error and it is not swallowed: it is the signal that what
you missed is unrecoverable, and to reconcile against windows() or
execPoll() instead of assuming nothing happened.
Reconnecting is on by default and is most of what events() is for. backoffMs
doubles up to maxBackoffMs between attempts, maxRetries gives up after that
many consecutive failures to deliver an event (0, the default, never does — a
connection that delivers an event resets the count), and
connectTimeoutMs bounds the handshake. maxQueued is how many frames may sit
unread before the socket is closed and reopened from where you had got to —
nothing dropped, nothing sent twice — because a websocket cannot be paused and
something has to bound a consumer that is not keeping up. reconnect: false
ends the iteration when the socket does, for a caller running their own
supervision; signal ends it on demand, without throwing. The defaults are
exported as EVENT_STREAM_DEFAULTS; waitFor takes the same options plus a
timeoutMs, three minutes unless you say otherwise.
computer.ready has a trap in it, and this SDK takes it out. It fires once
per desktop session, so a machine that has been up for an hour will never send
it again — a raw socket waiting for it waits forever. The opening frame carries
the state instead, and a stream that joins an already-ready desktop yields a
computer.ready marked synthesized: true as its first event. That is what
makes waitFor('computer.ready') return at once on a computer somebody else
already brought up.
Only where it could not arrive as an event: a stream opened with since either
already had the readiness or is about to be handed it out of the backlog, so
nothing is made up there. A resume that gapped does get one, because the
backlog it would have been in is what the gap says is gone — including a
second time, if the stream gaps again. A desktop can be replaced inside a
running computer, and a gap is exactly where the event saying so went missing,
so one extra readiness per gap is the price of never suppressing a real one.
A second computer.ready is real news: restarting the display manager inside
a guest destroys the desktop and brings up a new one without the computer ever
leaving running. The new desktop's windows arrive as window.opened before
that second ready, so a client that empties its map when it arrives throws away
the openings it was just handed. Nothing on the wire marks where the
replacement begins — ask windows(), which asks the machine.
Watching a directory
file.changed is the one event that never arrives unasked. Nominate the
trees you want on the way in, and only those are reported:
const stream = c.events({ watch: '/home/user/project' }); // up to four
for await (const ev of stream) {
if (ev.type !== 'file.changed') continue;
if (ev.armed) continue; // this tree is live from here on
if (ev.lostReason) continue; // my picture of this tree is wrong
console.log(ev.kind, ev.path); // created | modified | deleted
}Because it is a nomination rather than a filter, it is an option on the stream
and not a type to watch for: without one, no file.changed can reach the
socket at all. It is fixed for the life of the subscription and re-sent on every
reconnect — a socket that came back without it would be healthy and silent,
which is the one failure you cannot tell from a quiet directory.
Match on what you were given, not on what you sent. The host normalises a
nomination — a trailing slash and a . segment are cleaned away — and the
cleaned form is what every event carries in ev.watch. stream.watching is the
answer, one entry per tree — and onConnect is where to read it before the
first event, since it is the opening frame that carries it:
const stream = c.events({
watch: '/home/user/project/',
onConnect: (hello) => console.log(hello.watching),
}); // [{ path: '/home/user/project', armed: false }]hello.watchingIncomplete and hello.windowsIncomplete are the opening frame's
version of a listing's short answer, one per collection: null when every entry
of that collection was usable, and a count of the entries that were not. An
entry this client cannot use — a row that is not an object, or one that names no
window and no tree — is dropped rather than allowed to end the connection (a
stream is worth more than one malformed row), so the count is the only trace it
was there. hello.watchingIncomplete is what to check before reading
hello.watching.length as what the host accepted; a window the host could not
describe moves the other one and says nothing about your nominations, which is
why there are two counts and not one.
And armed is the half that is easy to get wrong. A tree is not being
watched the moment the opening frame accepts it: the guest has to be asked, and
on a computer nobody has opened a terminal on the host installs the watcher
first — seconds, not milliseconds. inotify reports changes and not state, so
anything that happens in that window is never reported and never will be.
armed: false in stream.watching means wait for that tree's file.changed
carrying armed: true; armed: true there means live now, and no event is
coming to say so, because the guest answers a nomination once and somebody else
got there first. stream.watching is each tree's state rather than the opening
frame's claim about it: an armed moves an entry to live and an unwatchable
moves it back, while flood and budget leave it, because under those the tree
is watched and is merely being reported incompletely. stream.hello.watching
stays what the connection was told when it joined, the same way hello.events
stays the opening vocabulary and stream.eventTypes is the live one — read
stream.watching to decide what silence means.
Same split as ready: state in the opening frame, transitions on the stream.
An armed also comes again after anything that re-arms the watch
— a stop and a start, a guest reboot — and means what the first one did:
reporting starts here, so re-read the tree if the interruption mattered.
The other shape carrying no path is a loss, in ev.lostReason. flood is
transient — the tree changed faster than the cap allows, so re-read it and keep
listening; a build under a watched path costs one of these rather than thousands
of events. budget means the tree is bigger than the directory budget one watch
gets, so part of it is not watched at all: permanent, and the fix is a narrower
path. unwatchable is the only one that means the tree is not being watched —
it is not there yet, is not a directory, cannot be read, or is a symlink,
which is refused rather than followed because inotify pins whatever the link
resolved to. That one recovers on its own where it can: nominating the directory
a job is about to create is supported, and the watch starts by itself when it
appears, announced by an armed and by nothing else.
Renames are a deleted and a created, not a move. Writes are coalesced, so
what you get is the truth about a path when the window closed rather than a
transcript of every write. Nothing is announced about what is already in a
tree when you nominate it — those are not changes.
Nominate the narrowest tree you can. Four distinct trees per stream — counted
the way the platform counts them, after normalising, so ['/a/b', '/a/b/'] is
one — and a computer watches at most 32 across every stream open on it; a
nomination past that one is refused on the upgrade. The replay history is per
computer and shared with every other subscriber, so a broad watch spends the
history a client resuming with a cursor needs.
Nominations are checked before a socket is opened, because the platform's 400
reaches a websocket client as the same empty close a rotated credential gives —
and with reconnect on, that is a stream that reopens forever and never says
why. Absolute paths, at most 256 bytes, no control characters, and not the root
however it is spelled: watching everything would spend the directory budget on
/usr before reaching anything you care about.
file.changed needs only the terminal channel, not the X bindings the window
watcher runs on — so it is advertised on Linux computers that emit no
window.* at all. The guest half is not one capability; read
stream.eventTypes rather than assuming the two travel together.
Three frames are about the stream rather than about the computer, and they
arrive as events too, because a client cannot ignore what it was never handed:
gap, closed (this host ending the socket deliberately, with a sentence
saying why) and capabilities (the vocabulary being revised under an open
socket). Ignore a type you do not recognise — the vocabulary grows.
A closed is reopened like any other drop rather than being sorted by its
wording, and the reconnect's own GET computers/:id is what sorts it: a
computer that moved to another host hands back that host's URL and the stream
carries on, and one that is gone answers 404 and ends it.
ev.source is worth reading. daemon means the platform observed it; guest
means the machine reported it about itself — every window.*,
clipboard.changed, file.changed and computer.ready — and anyone with root
inside that guest can make those say anything.
waitFor refuses rather than waiting out three cases. An event type this
computer cannot emit: a Windows guest, or an image built without the X bindings
the watcher needs, produces no window.* and no computer.ready, the opening
frame says so, and stream.eventTypes is that list. A waitFor('file.changed')
with no watch nominated, which the advertised list alone would call reachable
and which nothing would ever satisfy. And a computer that is suspended or
stopped — the stream is the one part of this API that does not resume a
suspended computer for you.
That last one is a refusal on the upgrade, and no refusal on the upgrade
reaches a websocket client as a status: a 409, a 401 and a TCP reset are the
same 1006 close, so the SDK reads the computer afterwards and says which it was.
A nomination the host will not honour arrives the same silent way, which is why
watch is checked before a socket is opened.
waitFor('file.changed') ends on a change, and not on the arming marker or
a loss. Three shapes share that type and only one of them is a change, so a wait
matched on the name alone came back with the arming on a fresh nomination and
with a real change on a tree somebody else had already armed — the same call
meaning two things depending on who got there first. The markers still arrive on
events(); they simply do not answer that question. A timeout says which
nominated tree never armed, because a watch that did not arm is silent in
exactly the way a tree where nothing happened is.
Windows guests have no event stream at all: there is nowhere in the guest to run the watcher the guest half needs.
Webhooks
The other transport for events. The socket above is for a caller attached to a computer and waiting. A webhook is for one that is not — CI, a queue worker, anything that wants to be woken rather than to wait. The platform POSTs one request per event to a URL you name, and its body is the event object exactly as the socket would frame it, byte for byte, with nothing wrapped around it.
const hook = await client.webhooks.create({
url: 'https://ci.example.com/mandala',
events: ['process.exited', 'computer.ready'], // omit for every type
computers: ['vm-3f9a1c2b7d4e'], // omit for every computer
});
await vault.put('mandala-webhook-secret', hook.secret); // shown ONCEThe secret on that answer is the only time you will see it. It is not on a
get or a list, and rotate() is the only way to get another — which mints a
new one and keeps honouring the old for 24 hours, so a receiver can switch over
at leisure.
Verifying a delivery is one call. Hand it the secret, the request headers in whatever shape your framework holds them, and the raw body — the bytes as they arrived, never the parsed object:
import { verify } from 'mandala-computer';
// Express: express.raw() so req.body is the bytes, not a parsed object.
app.post('/mandala', express.raw({ type: 'application/json' }), async (req, res) => {
if (!(await verify(process.env.MANDALA_WEBHOOK_SECRET!, req.headers, req.body))) {
return res.status(401).end();
}
res.status(200).end(); // acknowledge first, then work
const event = JSON.parse(req.body.toString('utf8'));
if (event.type === 'process.exited') queue.push(event.computer, event.data);
});
// fetch-shaped runtimes (Workers, Deno, Bun, Next route handlers):
export async function POST(request: Request) {
const raw = await request.text();
if (!(await verify(secret, request.headers, raw))) return new Response(null, { status: 401 });
return new Response(null, { status: 200 });
}The scheme is Standard Webhooks v1,
verbatim — HMAC-SHA256 over webhook-id.webhook-timestamp.body — so any of
that specification's libraries verifies a Mandala delivery too; this one holds
the platform's own test vector and needs no dependency. It is async because it
uses WebCrypto, which is what lets it run on the edge, where webhook receivers
tend to live.
Three things the verifier does for you, and one it cannot. A delivery whose
webhook-timestamp is more than five minutes from your clock is refused before
the signature is checked. A header carrying two signatures — every delivery
inside the rotation window — passes under either secret. And a secret pasted
without its whsec_ prefix throws rather than returning false forever, since
that is a configuration error and not a bad delivery. What it cannot do is
remember: record every accepted webhook-id before processing and refuse
repeats. How long for is replayRetentionS() — twice the tolerance from the
moment of acceptance, inclusively: keep the id while the elapsed time is less
than or equal to it, and drop it only once past.
import { replayRetentionS } from 'mandala-computer';
const keepFor = replayRetentionS(); // 600 seconds, with the default tolerance
const custom = replayRetentionS(60); // 120, if you overrode toleranceSA function rather than a sentence because the obvious reading of the window gives half the right answer. A timestamp can be one tolerance ahead of your clock when you first accept it, and the same bytes still verify one tolerance behind it, so the gap a retention has to span is two tolerances. An id remembered for one is forgotten while its signature is still good, and whoever captured the request replays it then.
Expire the id on the same clock verify reads, and on one that never goes
backwards. That clock is the wall clock unless you pass now, so a TTL measured
on a monotonic c
