@iso4/sandbox
v0.6.3
Published
Fast, sandboxed V8 isolate runtime for agent-generated JavaScript. Two-process architecture for crash isolation.
Maintainers
Readme
@iso4/sandbox
Fast, sandboxed V8 isolate runtime for agent-generated JavaScript. Runs user code in a separate Rust process for full crash isolation — an OOM or panic in the sandbox kills only the subprocess, not your host application.
Built for the AI-agent prefix/postfix pattern: prepare host setup (globals,
libraries, tool bindings) once, then run many agent-generated code strings
against it in parallel. Every run gets its own limits, guards, and bridge
bindings; one-off sandbox.run() always gets a fresh isolate, while prefix
runs reuse resident warm instances when one is free (see below).
Status: core execution works end-to-end. Not yet at 1.0 — every API, option, and observability surface (including
stats()) may still change.
Install
npm i @iso4/sandbox
# hardened fetch defaults (recommended):
npm i @iso4/fetchPrebuilt binaries exist for macOS (arm64, x64) and glibc Linux (arm64, x64;
glibc ≥ 2.34 — Debian 12 "bookworm" / Ubuntu 22.04 or newer). musl-based
images such as Alpine (node:*-alpine) are not supported — the binary
won't start there; use a node:*-slim image instead (#186
tracks musl support).
Quick start
import { createSandbox } from '@iso4/sandbox'
import { createSafeFetch } from '@iso4/fetch'
const sandbox = await createSandbox({ memoryMb: 128 }) // per-isolate heap cap
// Validate and prepare host setup once
const prefix = await sandbox.prepare({
code: `
const config = { apiBase: 'https://api.example.com' }
globalThis.config = config
`,
globals: {
fetch: createSafeFetch({ policy: ({ host }) => host === 'api.example.com' }),
},
})
// Run agent-generated code against the prefix — as many times as needed
const result = await prefix.execute({
code: `
const res = await fetch(config.apiBase + '/users')
export default { count: res.length }
`,
limits: { cpuTimeMs: 200, wallTimeMs: 5_000 }, // heap cap: prepare({ memoryMb }) or the sandbox default
})
if (result.ok) {
console.log(result.exports.default) // { count: 42 }
} else {
console.error(result.error.code, result.error.message)
}
await sandbox.dispose()
sandbox.prepare()andprefix.execute()are the current names. The former names —sandbox.precompile()andprefix.run()— remain as deprecated aliases with identical behavior and will be removed in a future major.
How globals work
globals wires any non-reserved name directly into the sandbox's global
object as a bridge stub. The bridge is fully generic — fetch is not
special-cased:
const options = {
globals: {
searchWeb: async (query: string) => {
const res = await fetch(`https://api.example.com/search?q=${encodeURIComponent(query)}`)
return res.json()
},
}
}Functions in bridge return values are currently dropped — return plain data, not class instances with methods.
TypeScript-checked rebinding
prepare() infers the globals shape G from what you pass, and the
returned Prefix<G> only allows rebinding those names at run time:
const prefix = await sandbox.prepare({
globals: { fetch: defaultFetch, myTool: defaultTool },
})
prefix.execute({ globals: { fetch: perUserFetch } }) // ✅ rebind one
prefix.execute({ globals: { unknown: handler } }) // ❌ TS errorResource limits
prefix.execute({
code: agentCode,
limits: {
cpuTimeMs: 200, // active JS execution only (await time excluded)
wallTimeMs: 5_000, // hard cap including async waits
// memoryMb: per prefix via prepare({ memoryMb }) or the sandbox default —
// never per prefix-run, the cap is baked into the shared warm isolates.
// One-off sandbox.run() may set limits.memoryMb (fresh isolate anyway).
maxBridgeCalls: 10, // max host-bridge calls per run (0 = unlimited)
maxBridgeCallBytes: 0, // max bytes per bridge call (0 = 64 MiB framing cap)
maxExportBytes: 0, // max bytes of encoded exports (0 = unlimited)
maxStdoutBytes: 0, // per-run console capture caps (0 = unlimited)
maxStderrBytes: 0,
},
})Warm instances and the memory budget
Runs against a prepared prefix are served by warm instances: resident isolates with the prefix already evaluated, so boot and prefix evaluation are paid once instead of per run. Automatic — no flag, no separate API.
Warmth is a cache, never a guarantee. An instance can vanish between any
two calls (memory pressure, dispose()), and a run interrupted
mid-execution — a CPU or memory limit firing, or an abort landing on
actively running code — costs its instance: the next call cold-starts
clean. Failures that arrive while a run is waiting (a wall timeout during
a host call, an abort of a suspended run, a dropped connection) fail that
run alone and the instance keeps serving, state intact. Don't rely on state
carrying over either way: globalThis writes and patched builtins stay
visible to later runs until eviction silently wipes them. Keep durable
state in a database and do expensive setup lazily in the handler
(conn ??= await connect()).
Carryover is not an isolation boundary. Because a warm instance is reused
across runs of the same prefix, one run can change what a later run sees —
including reassigning the runtime's own globals (Response, fetch, …) or
patching prototypes, which then affects that later run. This is intended and
matches a shared-isolate worker: run one prefix only for callers that trust
each other, and give mutually-distrusting workloads separate prefixes.
Instances of one prefix share no state with each other; today each serves
one call at a time, so concurrency means more instances (the engine can
already interleave several runs on one instance — that switches on with
wire multiplexing). Top-level names never collide across runs
(each run is its own module). A prefix that can't finish evaluating under
the fixed warm-up budget is rejected by prepare() with
ERR_WARMUP_LIMIT. console.* from prefix evaluation arrives on the
cold-starting call's result.
Residency is bounded by memory, not by a count:
const sandbox = await createSandbox({
memoryMb: 128, // default heap cap per isolate (override per prefix at prepare())
memoryBudgetMb: 2048, // RSS mark for the whole runtime process (0 = off)
hostReserveMb: 128, // container memory held back for this host process
})memoryMb means two different things depending on whether the isolate is
reused.
On a prefix, the number is the retirement line: an instance whose heap sits
above it when two runs in a row finish takes no new runs and drops once the
runs it is already serving complete. Nothing fails. The terminating line sits a
headroom band above it (round(sqrt(8 × memoryMb)) MB, so 128 → 160); crossing
that kills the run and fails its co-residents, and only a runaway allocation
gets there. The band is what lets a grown instance be replaced instead of
killing a run at the number — an optimization for reuse.
await sandbox.prepare({ code, memoryMb: { soft: 96, hard: 256 } }) // exact lines
await sandbox.prepare({ code, memoryMb: { hard: 128 } }) // terminate at 128, never retireOn a one-off run(), the number is enforced exactly — memoryMb: 128
dies at 128. A fresh isolate is never reused, so there is nothing to retire
and no band. It takes a plain number for the same reason: a soft line would
be inert.
The runtime watches its own process RSS. At or above memoryBudgetMb it
evicts idle instances (largest heap × longest idle first) and stops adding new
warm ones — prefix runs then execute on cold one-off isolates — until RSS falls
back to 80 % of the mark. maxConcurrentRuns caps concurrent runs; this caps
memory.
The default is derived from the memory available to the process
(container-aware), so most hosts never set it: 80 % of the container limit
minus hostReserveMb, which is headroom for this host process to grow into
(its current usage is already metered) — raise it when the host caches
heavily, 0 to hand the sandbox the whole limit.
Above the budget sits the admission line, and a run that needs a new isolate
whose heap ceiling would cross it fails with ERR_CAPACITY_MEMORY rather than
risking the container. That refusal is honest, not a bug to route around:
nothing ran, the telemetry is zero, and a retry a moment later usually
succeeds. The runtime does not evict warm instances to squeeze the refused
run in — eviction hands the freed pages to V8's pool, which the next
isolate draws from, so the measured usage the line compares against does not
move for seconds either way. Instead the budget mark sheds idle warmth
gradually while the line holds, and V8 returns what it no longer needs to the
operating system on its own. If refusals are routine rather than occasional,
the container is genuinely too small for maxConcurrentRuns × memoryMb —
createSandbox warns at startup when those two contradict each other.
sandbox.stats() reports the live picture — active runs, queue depth, warm and
idle instance counts, summed idle heap, budgetBytes / rssBytes, whether the
runtime is currently underPressure, and per-prefix counts. It answers on a
dedicated connection, so it works even when every run slot is busy.
Concurrency, and why shedding is yours
How many runs execute at once is the runtime's decision, not a constant. It
derives the number from the shape of the runs it is serving —
5 × cores × wall^0.7 ÷ cpu^1.2 over the uncontended minimum wall and CPU
time of each prefix, recomputed every 32 completions — so a prefix that
waits 200 ms on an upstream call is granted thousands of slots while one
that burns a core for 3 ms is granted a handful. Callers beyond the current
number queue FIFO, stats() reports it as slotLimit, and a run that had
to wait carries queueWaitMs on its result.
const sandbox = await createSandbox({
maxConcurrentRuns: 64, // pins the number — derivation is off entirely
maxQueuedRuns: 10_000, // waiters behind it; past this, ERR_QUEUE_FULL
})That number is about the sandbox, not about your process. For wait-heavy work it grows into the thousands, and every one of those runs is also promises, frames and host handlers on your Node event loop — the same loop serving your own requests. Measured on an 8-core pod against runs that wait 250 ms on a host call: ~8k concurrent runs sustain ~28k runs/s at ~43 ms event-loop delay, while ~12k push that delay to ~286 ms and throughput down to ~22k/s. Zero errors at either level — nothing fails, the process just gets slow, your own request handling included.
Neither side refuses on that. The runtime grants without enforcing and never
looks at your event loop, so the ceiling is yours to impose: keep your own
in-flight bound upstream of execute() and reject or defer past it —
perf_hooks.monitorEventLoopDelay() is the signal to set it by — or pin
maxConcurrentRuns to a number you have measured this host at. A pin is a
fixed number, not a cap on the derived one: it opts out of derivation for
every workload the sandbox serves.
Async context (AsyncLocalStorage)
Run/postfix code can import a minimal, Node-compatible AsyncLocalStorage to
carry an ambient value across await points — concurrency-safe, unlike a
module variable:
prefix.execute({
code: `
import { AsyncLocalStorage } from 'node:async_hooks'
const als = new AsyncLocalStorage()
export default await als.run('trace-42', async () => {
await somethingAsync()
return als.getStore() // 'trace-42', even several awaits deep
})
`,
})Only run(store, callback, ...args) and getStore() are provided. Built on
V8's continuation-preserved embedder data; no promise hooks, so it's free
unless used. Not available in prepare() (prefix) code — it's for the
postfix. See DESIGN.md §16.
Calling into the sandbox
Invoke a function the module exports — export default { fetch } or any
named export — with real typed arguments. On a prepared prefix nothing is
compiled per request:
const prefix = await sandbox.prepare({ code: workerBundle })
const result = await prefix.call({
export: 'default.fetch',
args: [new Request('https://example.com/', { method: 'POST', body: 'hi' })],
})
if (result.ok) {
const response = result.value as Response // a real Response instance
}sandbox.run({ code, call }) does the same against a freshly evaluated
module. With call the success result carries the function's return value
instead of exports — never both. The receiver is the exported object
(this works); prototype methods resolve (export default new Worker()),
and a path that does not reach a callable fails with
ERR_CALL_TARGET_NOT_FOUND. prefix.call() accepts the same per-call
globals / imports rebinds as prefix.execute().
Non-serializable exports no longer fail a plain run: they are absent from
exports and reported in skippedExports, so a module that exports handlers
still reads cleanly. sandbox.readExports({ code }) wraps that for the
deploy path — load once, read the declaration exports, and get the skipped
handler names back.
Batch event streams
Arguments cross as one blob per call, so a handler that takes an array of events and returns an array of results pays one crossing per batch instead of one per event. For event-shaped work this is the biggest throughput lever available:
export default {
transform(events) {
return events.map((event) => ({ id: event.eventId, bucket: event.type }))
},
}const result = await prefix.call({
export: 'default.transform',
args: [events], // one array, not one call per event
})Measured on a warm prefix over an 8-slot pool with ~750 B analytics events
(bench/warm.bench.ts, "batched calls"): ~98k events/sec one at a time,
~319k at 8 per call, ~365k at 32. That is 3.6× at the peak, and most of it
is already there at 8.
Two things bound it. One call runs on one slot, so a single batch of 128
comes out slower than four batches of 32 running in parallel — size
batches so several are in flight across the pool. And a batch shares one
run's limits: cpuTimeMs, wallTimeMs and memoryMb cover the whole
array, so an oversized batch turns a per-event cost into a per-run timeout.
Result shape
type RunResult
= | { ok: true, exports: SandboxExports, skippedExports: string[], stdout: string[], stderr: string[], durationMs: number, cpuTimeMs: number, bridgeCalls: BridgeCallEntry[] }
| { ok: false, error: RunError, stdout: string[], stderr: string[], durationMs: number, cpuTimeMs: number, bridgeCalls: BridgeCallEntry[] }
// prefix.call() / run({ code, call }) resolve to a CallResult instead:
// the success arm carries `value` (the function's return value) in place of
// `exports` + `skippedExports`; failure/aborted arms are shared.
// durationMs — wall-clock time of the run; cpuTimeMs — active V8 execution
// time (bridge waits excluded). Both measured in the runtime, µs resolution.
type BridgeCallEntry = { // recorded in the Rust runtime; one per attempt, in order
name: string // 'fetch', 'myTool', or '<specifier>.<path>' for host-module imports
startMs: number // offset from run start (same clock as durationMs)
durationMs: number // round-trip the sandbox waited (handler + IPC)
argBytes: number // serialized call payload size
responseBytes: number // serialized response value size (0 unless ok)
} & (
| { ok: true }
| { ok: false, reason: 'blocked' | 'error' | 'unanswered' | 'dropped' }
)
// blocked: refused runtime-side (limit, function argument, ...), never reached the host
// error: the handler threw / rejected, or the host could not decode or encode a value
// unanswered: no answer when the run ended (handler still ran)
// dropped: the run was already aborted, the handler was never called
interface RunError {
code: RunErrorCode
name: string
message: string
stack?: string
fields?: Record<string, unknown> // all other own-enumerable props of the thrown error
resetCause?: 'cpu' | 'memory' | 'wall' | 'abort' | 'internal' // ERR_INSTANCE_RESET only
culpritRunId?: number // ERR_INSTANCE_RESET only: the run whose interruption reset the instance
}run() never throws for sandboxed failures — only for infrastructure errors
(subprocess crashed, binary not found). ok: false with an error code is the
normal failure path.
One code is about a neighbor, not your own run: ERR_INSTANCE_RESET means a
run sharing the same warm instance had to be terminated mid-execution
(resetCause says why, culpritRunId names it), so this run — in flight on
the now-untrusted instance — was failed with its real partial telemetry. It is
never retried automatically, because it may already have had side effects.
Thrown errors keep their identity across the bridge, in both directions:
- Sandbox → host: an uncaught sandbox throw surfaces as
ERR_USER_CODEwith the error's realname,message,stack, and every other own-enumerable property undererror.fields(namespaced so a customcodeproperty can't collide with the iso4error.code).name/message/stackare reserved and never appear insidefields. - Host → sandbox: a host handler that throws rejects the sandbox call with
a real
Errorcarrying the samename(instanceof TypeErrorworks for built-ins) and its extra properties re-attached directly (e.status,e.reason, …). Sandbox code can catch it and continue; uncaught it fails the run withERR_HOST_BRIDGE. The host stack never crosses into the sandbox.
Every own property you attach to a host-thrown error is visible to sandbox code (the auto-populated stack is the one exception). Sandbox code is untrusted, and this includes properties a third-party SDK attaches to its own error objects — so if an error may carry request context or credentials (an SDK's
config/request/response), throw a cleanErrorfrom your handler instead of re-throwing it wholesale. Sanitising is the handler's job, the same as with any value you return.
Architecture
V8 runs in a separate Rust subprocess communicating over a Unix domain socket.
A slot pool admits as many runs at once as the runtime currently grants (the
rest queue FIFO — an AbortSignal, e.g. AbortSignal.timeout(), bounds the
wait);
connections to the subprocess open on demand and are reused, and stats()
reports the live count as openConnections. Five concurrent
prefix.execute() calls each get their own run slot and execute in
parallel.
License
MIT
