@castle-ai/node-runner
v0.1.0-beta.6
Published
Run Castle agents in Node.js: sessions, persistence, approval, workspaces, MCP and child agents.
Readme
@castle-ai/node-runner
@castle-ai/node-runner is the public Node.js composition root for Castle Agents.
It builds a directly held TypeScript Agent declaration into a Program for each
Session and binds local Adapters. Start with openSession; applications that
manage several Sessions can share one lower-level createNodeRunner.
Model/Skill/Sandbox resolution, Session identity and persistence now belong to
this package. There is no separate @castle/execution-host package to install.
It includes:
- Agent capabilities:
workspacefor workspace tools, commands, OS sandboxing and directory-scoped instructions;skills,mcp,childAgent,goalandinstructionsFile; - root-bound workspace list, search, read, exact-edit, and patch Tools;
createWorkspaceInstructionsPluginfor Run-scoped root and target-directoryAGENTS.mdcontext, with explicit byte budget and a host-owned read-receipt callback;createWorkspaceEnvironmentPluginfor observed working-directory and Git worktree facts;- POSIX command Session plugins for explicit host or OS-sandboxed processes;
- a session-scoped local browser QA plugin with an optional Playwright Adapter;
- a session-scoped local stdio MCP Tool plugin;
- an explicit provider-hosted Web Search declaration boundary;
- a local file Session Store and
SKILL.mdresolver; and - bounded foreground child Agents.
Start with examples/session for a small persistent text
conversation: Agent, Model binding, file store, streaming and resource closure.
It runs again in a new process using the same saved Session. The offline Adapter
echoes supplied context so the example needs no credentials.
Its OpenRouter entry lists models with their
capabilities and starts a persistent conversation with only a key and your
selected model ID. The generic API-key entry remains available for other endpoints.
The Anthropic entry verifies a native API key,
lists the account's catalog and binds the chosen model through the same Session API.
See examples/local-coding-agent for a
deterministic, provider-free read → edit → test → final run with approval,
automatic compaction, a child Agent, and restart history.
Assistant results expose both result.content (ordered text/image/Provider
blocks) and the derived result.text. File stores preserve the content and
opaque continuation across processes; opening a saved Session does not start
a model request. For image replies, read result.content rather than treating
empty text as failure. See the content contract.
Alpha quickstart
The packages currently have private 0.1.0-alpha.1 manifests and are not available
from a registry. From this alpha source checkout, run the exact external-package
flow:
pnpm --dir packages/node-runner test:docsThat command cleans and packs the required packages, extracts the example
from the packed @castle-ai/node-runner tarball, installs only those tarballs in
an unrelated temporary consumer, typechecks every example entry, and runs the
deterministic coding flow. It executes the real-model CLI with mock HTTP to
check tool/search binding, Skill loading, process restart and interrupt cleanup.
It also compiles the minimum Session example and runs
it across independent OS processes, checking continuation and rejection of
invalid or wrong-Agent history without overwriting saved files. The API-key
entry is also run across processes against a local HTTP fixture, without
provider credentials. None of these checks uses workspace package resolution.
The OpenRouter entry additionally verifies public discovery, inference-key
rejection, selected capabilities, unknown-model rejection and interruption during
both discovery and streaming against a local HTTP fixture.
The Anthropic entry checks its native authenticated, paginated catalog, selected
thinking/output limits, cross-process history and discovery/stream cancellation.
To keep a portable local candidate with versioned tarballs, both examples and relative dependency bindings, choose a new output directory:
pnpm pack:sdk /absolute/path/to/new-castle-sdk-candidateThis uses the same packed-example verification, then retains the archives,
example sources and tested lockfiles. It removes installed dependencies and QA
history before delivery. Follow the generated README to install either example
from the kit; the output folder must not already exist. No registry publication
occurs. The seven runtime packages share 0.1.0-alpha.1; optional source-build and
private product packages are outside this candidate.
The complete, compiling quickstart is intentionally split by responsibility:
agent.tsdeclares the parent and child withdefineAgent, ahostToolreference, and theskillsandworkspacecapabilities;run.tsprovides the complete publiccreateNodeRunnercomposition, including the Session Store, permission Hook, approval Adapter, compaction budget, foreground child, cancellation-compatible event stream, and restart; andmock-model.tssupplies the deterministic provider-neutralModelRuntimeused by the smoke.
Declare a local workspace
Add workspace({ root }) to an Agent's capabilities. It contributes the file,
image and command Tools, the Sandbox, and a workspace instructions section with
their usage rules. It opens the root once per Session and loads applicable
AGENTS.md guidance for each Run. The coding example uses this capability
without manually registering file and command tools.
import { defineAgent } from "@castle-ai/harness";
import { workspace } from "@castle-ai/node-runner";
export const Coder = defineAgent("example.coder", {
model,
instructions: "Make the smallest correct change, run the tests, and report.",
capabilities: [workspace({
root: new URL("./project/", import.meta.url),
network: { allowedDomains: ["registry.npmjs.org"] },
})],
});root is a file URL, or a function of the Session's parameters returning one,
such as root: params => new URL(String(params.checkout)). A fixed root is
validated when the capability is created, a derived one when each Session opens.
The defaults are /bin/sh -c, a 120-second ordinary-command timeout, 64 KiB
per output stream, the core environment, and sandbox: "os". An unavailable
OS sandbox fails opening; it never falls back to unsandboxed execution.
sandbox: "none" explicitly runs commands without OS containment. readOnly
exposes only file and image reading tools, with no commands or Sandbox.
network.allowedDomains / deniedDomains set the OS sandbox's rules; omission
permits no external hosts. Publication requires a host publisher that returns
a verified delivery reference. An OS workspace declares the Agent's one Sandbox,
so the Agent and its other capabilities cannot declare another. Declare a Skill
directory separately with skills; installation and choosing source scopes
remain the application's responsibility.
Each OS workspace Session owns an isolated sandbox worker, including when
two Sessions use the same root or a parent delegates to a child. Network
hosts approved through request_network_access last only for that Session. Closing waits
for its commands, resets its sandbox and terminates its worker; another
Session keeps its own configuration and process lifecycle.
Source development must build the worker with
pnpm --filter @castle-ai/node-runner build:worker before opening an OS workspace.
The package's test command includes this step; direct Vitest or source CLI
invocations need it explicitly. Both source and installed use the fixed
compiled worker entry. Published tarballs include it, so consumers need no
TypeScript loader or build step. The sandbox-worker package subpath is for
Node/Desktop host integration, not an additional Agent-authoring API.
Bind your models
Models can come directly from the Agent declaration. For example, with an Anthropic API key and an explicit model ID:
import { defineAgent } from "@castle-ai/harness";
import { openSession } from "@castle-ai/node-runner";
import { anthropic } from "@castle-ai/models/anthropic";
const model = anthropic({
apiKey: process.env.ANTHROPIC_API_KEY!,
model: process.env.ANTHROPIC_MODEL!,
});
const Assistant = defineAgent("app.assistant", {
model,
instructions: `Answer clearly using ${model.specifier}.`,
});
const session = await openSession({ agent: Assistant });
try {
console.log((await session.run("Explain an Agent Session.").result).text);
} finally {
await session.close();
}anthropic defaults to a 4,096-token response limit; use maxOutputTokens to
set another limit. The other native factories are openaiCompatible, openrouter
and codex in their respective runtime packages. Their connection inputs retain
the documented provider-specific shapes; a shorter name does not add catalog
discovery or credential storage. The old create*ModelRuntime factory names and
createTerminalToolApprovalPlugin are removed in this alpha contract reset.
openSession({ agent, id?, params?, store?, approval?, permissions?, signal? })
uses the existing Runner lifecycle internally. store accepts a SessionStore
directly; approval accepts a NodeRunnerPlugin, including its cleanup (for
example, the existing terminalApproval). Opening with the same
ID and Store restores history and parameters without starting a model or Tool.
Closing waits for active work to settle and releases the Session's private Runner.
Failed opening also cleans up; an already-aborted signal rejects before plugins
are applied. Always close a successfully opened Session in finally.
Choose a Tool permission preset
openSession defaults to ask-on-write. Only an executable Tool declared by the
application with readOnly: true qualifies as read-only.
| permissions | Declared read-only | Other executable Tools |
| --- | --- | --- |
| read-only | Allow | Deny |
| ask-on-write | Allow | Ask through the configured approval plugin |
| full-access | Allow | Allow |
Without an approval adapter, a call that requires approval is denied. A preset
answers only calls that every toolPermission hook left undecided (no answer or
defer). A host policy's deny wins, then its ask, then its allow, so a host can
tighten or relax the preset for the calls it recognizes. Hooks and approval
requests carry agent, the identity of the Agent whose model made the call, and
for a call inside a foreground child Agent, parent: the parent Session, Run,
Tool call id and Tool name that started it. Presets do not remove root or OS-sandbox limits,
classify provider-hosted actions, or infer effects from a Tool's name, description,
parallel scheduling or MCP annotations. Declare readOnly: true only when the
whole Tool is read-only; a Tool with mixed effects stays unclassified.
The built-in workspace list, search, file reader and image reader declare this
trait. Edits and commands require approval under the default preset. Existing
hosts using createNodeRunner keep their own policy unless they explicitly pass
permissions; the SDK modes are distinct from Desktop's product permission modes.
Direct store/approval options and plugins register through the same owner;
registering two Stores or two approval adapters is rejected.
Model binding details
Anthropic, OpenAI-compatible and OpenRouter factories return a ModelBinding:
the streaming runtime plus its specifier and known contextWindow. A custom
runtime can provide those same fields; a runtime with dynamic connections, such
as the current Codex Adapter, can be named with { specifier, ...runtime }.
The connection and credentials remain inside the runtime; the Program contains
only its model reference. No model connection is opened when the Agent is
declared or a Session opens.
model: "app/chat" selects a named host binding. A missing registered name
fails on Session open; a dynamic ModelResolver still resolves the admitted
Run's model when that request starts. A directly declared model and a registered
model with the same name are ambiguous and rejected. Explicit Run or compaction
model changes resolve the requested name; they never silently reuse a different
declared model. Plugins are optional when the declaration supplies its own model.
When your application already has a ModelRuntime, register it under the same
name used in the Agent's model:
// Inside a plugin's apply(host):
host.registerModel("app/chat", chatModel);
host.registerModel("app/review", reviewModel);The Agent declares model: "app/chat". These names are opaque
application bindings; the Provider Adapter still owns the real model ID,
endpoint and credentials. No resolver callback is needed for fixed bindings.
Multiple plugins may register distinct names, and the same bindings serve Runs,
foreground children and compaction. An unknown name fails before contacting a
model; there is no default-model fallback. Registration closes at Runner creation,
and later replacement of the passed runtime's stream method has no effect.
Use host.registerModelResolver(resolver) instead when the host must resolve a
connection for each request, such as Desktop's changing authenticated accounts.
One Runner uses either named models or one resolver; mixing them and registering
a name twice are errors. Both feed the same Runner model-resolution path and
preserve the admitted model name and request identity. See the complete
API-key example for fixed bindings and the
existing composition example for a resolver.
A typed Tool value in tools supplies its contract and implementation.
Declaring it does not execute it, and the Program contains only data. For
host-owned catalogs, hostTool("lookup") resolves the implementation from
host.registerTool(lookupTool); missing names fail when a Session opens.
Registered Tools run with the same Tool context as declared ones. Binding the
same name both directly and through the host is an error. Each foreground child
builds its own Program and still obeys toolGrants.
enabled on a Tool, a hostTool reference or a capability controls whether the
model sees it. It is true, false, or a function of the turn context evaluated
at each turn. Hidden Tools still need valid contracts and implementations; they
are absent from model requests and cannot be called by a stale model name.
Skills hidden with their capability cannot be activated explicitly or by the
model, cannot read reference files, and do not project saved activation
instructions. Their durable history remains intact. beforeModel hooks may
transform messages but must preserve the projected Tool list and contracts.
A SkillBinding in skills binds a value directly. Metadata is available at
Session admission; load() and readResource() run only through Skill
activation/resource reading. A string in skills uses the host Skill resolver,
so direct and named Skills can coexist. A direct value owns its reference; the
host resolver is only called for the remaining named references.
sandbox: { specifier, open } contributes an acquisition request.
The Runner calls open({ signal }) once when opening that Session and releases
the returned { sandbox, close } lease on close or admission failure.
sandbox: "name" uses the host Sandbox resolver instead. Declaring never
acquires the Sandbox; both forms use the existing admission and cleanup owner.
A capability's resource(kind, spec) shares this Session acquisition/cleanup
owner. Each typed kind owns acquire(spec, { signal }) → { value, close }; a
frozen handle exposes value after acquisition. See the
Harness capability contract.
BindingKind, BindingHandle, BindingLease and SandboxLease are exported
from @castle-ai/node-runner, which owns acquisition and cleanup. Host approval
plugins also import ToolPermissionContribution, ToolApprovalAdapter,
ToolApprovalRequest, ToolApprovalAnswer and ToolExecutionHostError from this
package. These are host contracts, not Harness authoring exports.
Load application instructions
instructionsFile(fileUrl, { name?, placement? }) is a capability whose one
section is read from a file:
import { defineAgent } from "@castle-ai/harness";
import { instructionsFile } from "@castle-ai/node-runner";
export const ReviewAgent = defineAgent("example.file-review", {
model: "review-model",
instructions: "Review the requested change.",
capabilities: [instructionsFile(new URL("./review-rules.md", import.meta.url))],
});The section is named after the file (review-rules.md becomes review-rules)
unless name is given, and renders with that heading after the Agent's own
sections. The application authorizes that explicit file URL and its resolved
target, including a symlink target. The Runner resolves and reads a regular
UTF-8 file once per Session as a capability resource, closes the file, and
retains immutable text. A changed or deleted file does not alter an open
Session; a newly opened Session reads again. The default placement is system;
{ placement: "context" } sends the same text as current user-side context
each turn. Contents are not written into Session history.
Declaring the Agent does no file I/O. projectAgent returns a pending resource
reference for a system-placed file; it cannot preview a context-placed one,
whose text exists only after acquisition. After acquiring resources the Runner
performs its initial projection, so even a Run stopped before its first model
request has resolved instructions. Missing, non-regular or invalid UTF-8 files,
and blank system-placed files, fail Session opening and release already
acquired resources. This is application prompt loading, separate from workspace
authority and AGENTS.md's per-Run discovery/refresh policy. It does not impose
the later K4 byte budget.
Change instructions and Tools each turn
Instruction sections and enabled options may be functions of the turn
context. The Runner builds the Program once when a Session opens, runs these
functions once after acquiring resources, then again synchronously before each
agent turn. Use them to change instructions and visibility without changing a
Tool's contract:
import { defineAgent, defineTool } from "@castle-ai/harness";
const Agent = defineAgent("example.bounded-search", {
model,
instructions: {
identity: "Answer questions about the project.",
phase: ctx => ctx.turn <= 3 ? "Search for relevant evidence." : "Answer using the evidence already collected.",
},
tools: [defineTool({
name: "search",
description: "Search the project",
inputSchema: { type: "object", properties: { query: { type: "string" } }, required: ["query"] },
readOnly: true,
enabled: ctx => ctx.turn <= 3,
async execute({ query }, ctx) {
return await searchProject(query, { turn: ctx.turn });
},
})],
});model is your bound model and searchProject is application code. turn is
the turn within the current Run. The Tool's context carries the turn that
requested the call; its schema and description remain fixed. Foreground child
projections remain restricted to their original Tool grants.
Use projectAgent from @castle-ai/harness/testing
to assert Tool visibility and instructions for specific state/parameter inputs.
Parameters and acquired resources are Session-local. Projection does not reopen the Sandbox, reload file resources, or rerun settled Tools. Skill activation still uses the normal Tool/explicit-input path; instructions of Skills hidden with their capability are omitted from the current projection while their original history is preserved. Run binding/model selection, manual compaction and restoration keep their existing explicit lifetimes. Resources belong to capabilities; writable Session state follows the contract below.
Inspect a live Session
session.inspect() returns a frozen, read-only view of accepted declarations.
It does not rerun the Agent, acquire resources, or make model requests:
const inspection = session.inspect();
console.log(inspection.current.declaration?.capabilities); // Capability names in declaration order
console.log(inspection.current.declaration?.placements); // Section placement table
console.log(inspection.turns.at(-1)?.changes); // Tools, Skills and sections changed
console.log(inspection.stateWrites); // Committed writes with source / Run / turnThe SessionInspection type is exported from @castle-ai/node-runner. initial
is the projection after resource acquisition; turns records successful bound
projections with Run identity and turn number, and current is the latest one.
Failed projections do not add turn entries. An explicit same-Run resume can
observe a turn again; entries are observations, not unique historical turn records.
Tool visibility includes internal Skill tools and child grant restrictions.
Sections and tools describe the accepted declaration before beforeModel
hooks rewrite a request; this is not the final Provider payload. Each placement
has the section's identity, name, source (agent, or capability with its
name), placement, lifetime and order. order is declaration order, not a
Provider message index. Manually supplied Programs have no declaration and report
declaration: null.
stateWrites includes only settled state facts after the configured Store
acknowledges them. Tentative or aborted Tool writes are excluded. pendingState
is the current host-write queue, not a durability acknowledgement; continue to
await session.setState() when queuing a durable host write.
projectionScope: "current-open" makes the lifetime explicit: reopening restores
saved state facts, but starts a new projection history. Reading after close still
returns cached data. Inspection omits acquired handles, binding specs, model
credentials and inactive Skill bodies; instruction and state content is visible
as authored and is not automatically redacted.
React to settled Tool results
import { defineAgent, defineTool } from "@castle-ai/harness";
const createReport = defineTool({
name: "create_report",
description: "Create the weekly report.",
enabled: ctx => !ctx.toolCalls("create_report").some(call => call.ok),
async execute() { return await createWeeklyReport(); },
});
const Agent = defineAgent("app.report-once", {
model,
instructions: {
task: "Create the report and summarize the result.",
retry: ctx => ctx.toolCalls("create_report").at(-1)?.ok === false
? "The last report attempt failed. Inspect its result before trying again."
: "",
},
tools: [createReport],
});createWeeklyReport is application code. After a successful result, the next
turn hides the Tool; an error leaves it available and adds the retry context
section, which is omitted while empty. ctx.toolCalls(name) lists settled
receipts with input, the complete result, ok, runId and turn. These are
frozen Session facts, retained across new Runs, compaction and restart. They do
not replace an external service's idempotency contract when a process fails
before saving a receipt.
The incident-reporter example reads the same receipts to report completed updates without storing a second counter. Its installed-package test reads those receipts in a fresh OS process after compaction.
SandboxContribution and BuiltinPromptOverrides are host configuration types
exported from @castle-ai/node-runner, alongside SandboxLease and
HarnessContextBudget.
Keep state across Tools and turns
Most progress can be derived from ctx.toolCalls. Use durable state for values
the model never produced or that cannot be derived, such as an external ID:
import { defineAgent, defineTool } from "@castle-ai/harness";
const Reporter = defineAgent("app.reporter", {
model,
state: { updates: 0 },
instructions: {
identity: "Record completed report updates.",
progress: ctx => `Report updates so far: ${ctx.state.updates}.`,
},
tools: [defineTool({
name: "record_update",
description: "Record a completed report update",
enabled: ctx => (ctx.state.updates as number) < 3,
async execute(_input, ctx) {
ctx.setState(state => ({ updates: (state.updates as number) + 1 }));
return `Recorded update ${String(ctx.state.updates)}.`;
},
})],
});Here model is a bound Model as above. state declares keys and their initial
JSON values; keys from the Agent and its capabilities share one flat space, and
a key declared twice fails when a Session opens. Instructions and enabled read
the committed values. Inside the Tool, ctx.state includes earlier writes in the
same batch, so the result above reports the new count. setState takes a patch
of declared keys, or a function of the latest state returning one. Values must
be JSON and are deeply frozen. Updates commit only after the Tool batch settles.
Abort or framework failure discards tentative state, while Tool receipts remain.
Ordinary Tool error results still settle. Queuing a run.join() input does not
itself abort the batch.
The host can call await session.setState("updates", 0) to queue an update for the next
turn. This does not change the current Tool's reads. Writes emit state-changed
on the Run and host.hooks.stateChanged after commitment, with source: "host"
or "tool". State is private unless the Agent puts it into instructions.
The file Store appends committed state facts before publishing state-changed.
await setState confirms that the queued write is saved, including when the
Session is idle. It checkpoints the current history and pending queue through
the same ordered Store path as transcript appends. Reopening waits for an
explicit Run before applying the queue. During a text-only turn, queued host
updates cause a further turn unless a stop condition intervenes. Without a Store
the acknowledgment is in memory. A failed save rejects and stops active work;
close waits for writes and releases resources. Each host write currently requires
a full checkpoint, so this API is not a high-frequency telemetry channel.
Compaction preserves state facts. Each batch records its start and a committed
or discarded settlement. Committed state and that settlement share one append;
notification failure afterward cannot change the decision. Individual receipts
still precede the batch's state decision; interrupted work uses the explicit
same-Run continuation and host reconciliation contracts below. See the executable repository example
stateful-report.ts,
verified with installed packages and two separate processes by test:package.
HarnessContextBudget is a Runner configuration type and is imported from
@castle-ai/node-runner.
Declare a durable Goal
The goal capability declares get_goal and update_goal and a goal context
section with the current Goal observation. With an existing ModelBinding named
model:
import { defineAgent } from "@castle-ai/harness";
import { createFileGoalStore, createFileSessionStore, goal, openSession } from "@castle-ai/node-runner";
const goals = createFileGoalStore({ directory: new URL("./goals/", import.meta.url) });
const agent = defineAgent("release-review", {
model,
instructions: "Review the project against its release ticket.",
capabilities: [goal({ store: goals })],
});
const session = await openSession({
agent, id: "release-review",
store: createFileSessionStore({ directory: new URL("./sessions/", import.meta.url) }),
});
try {
// Open only in response to the user's explicit request.
await goals.control({ sessionId: session.id, action: "open", objective: "Complete the release review." });
const result = await session.run("Review the requirements and current implementation.").result;
console.log(result.text);
} finally {
await session.close();
}The Goal Store owns the objective and lifecycle. Host controls open, edit, pause,
resume or delete it; Agent tools can read and report complete/blocked using the
exact current identity and objective. A paused, deleted or changed Goal rejects
an obsolete settlement. Task completion never completes a Goal. GoalPort lets
a product adapt its existing authoritative Store without a Conversation DTO.
Each new turn reads the Store before projection and durably queues that current
observation through Session state (key castle.goal.observation), superseding
any older pending observation. update_goal is visible only while the latest
observation is an active Goal.
The turnStart lifecycle hook runs before queued host writes are applied and the
turn is checkpointed/projected; replaying an existing checkpoint does not call it.
The context section is not appended as a conversation message. Existing checkpoint
replay preserves the original projection; the next new turn refreshes it. Opening
or reopening a Session does not read the Goal, run the model, resume paused work,
or schedule continuation. Store read/write failures preserve their cause.
Omitting the capability's Store uses one shared file Store at .castle/goals relative
to the current working directory; it does not open a Goal. Use an explicit shared
Store for host controls. File Stores require one writing instance/process per
Session; separate Sessions can share a directory. Goal storage and Session history
storage are distinct.
Continue a Goal across Runs
Call driveGoal explicitly for an open, idle Session whose Agent declares the
goal capability. Pass the same Store to both. For example:
import { driveGoal } from "@castle-ai/node-runner";
const stop = new AbortController();
const result = await driveGoal(session, {
store: goals,
objective: "Complete the release review.",
policy: { idleMs: 1_000, maxCheckins: 5 },
observe: () => ({ queuedInput: false, interactionPending: false }),
signal: stop.signal,
async onRun(run) {
for await (const event of run) {
if (event.kind === "assistant-text-delta") process.stdout.write(event.text);
}
},
});
console.log(result.decision, result.runs);The example has no UI queue; a UI host supplies its actual queued-input and
pending-interaction facts synchronously. onRun exposes the real Run, including
its event stream and join. Normal Session permissions still apply. The driver
does not answer approvals, grant permissions, or change the Agent's toolset.
An objective is an explicit host request to open a Goal. A different unfinished
objective is rejected; a paused or blocked Goal is never implicitly resumed.
Omit objective to work on the existing Goal. The first Run starts immediately;
maxCheckins caps subsequent automatic Runs, separated by idleMs. The defaults
are 60 seconds and five check-ins. Exhausting the allowance leaves the Goal active.
After each Run, the driver reads authoritative Goal state. It checks synchronous
host facts again after onDecision and immediately before the next admission.
Queued input, pending interactions, an existing Run, pause, deletion or completion
return a typed hold to the caller. Only the idle cooldown waits internally; a host
must explicitly call again after resolving another hold. Stop conditions return
run-stopped; failures reject with their cause. An edited objective is reread, and
a replacement Goal identity ends this invocation. Live command handles can be
provided in observe().activeCommandSessions for a check-in; they do not force the
Agent to wait for a healthy service to terminate.
Abort stop to stop the driver and its own active Run without closing the Session.
If a user pauses or deletes a Goal during a Run, apply that Store control and
abort this signal together; Store controls alone do not interrupt executing tools.
session.signal exposes the existing Session lifetime: closing the Session or
Runner also cancels the driver and its cooldown. Opening saved history never
starts a driver. An explicit new call arms it again; it never automatically retries
an interrupted Run or consumes saved unanswered input.
Each admitted input carries an opaque, durable ID of the form
castle.goal/<encoded-goal-id>/<uuid>. This records the host's Goal origin outside
model messages using the existing input identity contract; it is not a new Harness
message type. History, Goal storage and task progress remain independent.
Hosts that already own a durable input queue can use createGoalDriverController,
the same controller used by driveGoal. It owns cooldown, allowance and cancellation
generations; the host keeps its existing queue and Run executor. Supply observe
for current Goal and host facts, and admit(goal, decision, admission) to prepare
a submission. Call admission.accept(freshFacts) synchronously immediately before
durably admitting it, with no asynchronous gap. Return true only for an accepted
submission. Use admission.signal to cancel preparation where supported.
Explicit user activity calls arm; wake rechecks facts without restarting the
idle deadline, and idle starts a fresh cooldown after a Run settles. cancel
disarms the generation; close also waits for any preparation already in flight.
An old generation's failure reaches onError without disarming a newer one.
snapshot exposes armed state and counts for host presentation. A controller
starts disarmed, so restoring a Goal alone cannot start work. By default every
admitted continuation counts toward the cap. driveGoal arms with initialRun
to exclude its explicit initial Run; Desktop retains its cap on all automatic
submissions. Host adapters remain responsible for stopping their active Runs.
Track tasks with durable defaults
createTaskTools() provides create, read, list, revise and progress tools. Its
default file Store is .castle/tasks under the working directory where the tools
are created. Declaring tools performs no I/O. With an existing model:
import { defineAgent } from "@castle-ai/harness";
import { createTaskTools, openSession } from "@castle-ai/node-runner";
const agent = defineAgent("release-work", {
model,
instructions: "Track meaningful release deliverables and their dependencies.",
tools: createTaskTools(),
});
const session = await openSession({ agent, id: "release-work", permissions: "full-access" });
try {
console.log((await session.run("Plan the release tasks.").result).text);
} finally {
await session.close();
}This example grants execution to its Task-only toolset. In an application with other tools, keep the appropriate permission policy and approval adapter; Task writes go through the same authorization path as other tools.
Task lists are scoped to the Session ID. Reusing that ID restores its Task list; persist conversation history separately with a Session Store. Creation records a pending task and does not start execution. Revisions preserve identity and progress; dependencies must exist in the list and cannot form a cycle. Pending tasks have progress 0, completed tasks 100, and in-progress tasks require an integer measurement from 0 to 100. Completing the list never completes a Goal.
To read tasks from the host or choose a location, create one
createFileTaskStore({ directory: new URL("./tasks/", import.meta.url) }), pass it
to createTaskTools(store), and use store.load(session.id). TaskStore extends
the neutral TaskPort; an application can still provide its own TaskPort.
Use one writing Store instance/process per Session. Saved task corruption is
reported with its cause and is never replaced by an empty list.
Give each Session its own parameters
Use immutable JSON parameters for constants such as an issue key or tenant ID:
import { defineAgent } from "@castle-ai/harness";
const IssueAgent = defineAgent("app.issue-agent", {
model: "app/chat",
params: ["issueKey"],
instructions: {
identity: "Resolve the Session's issue.",
issue: ctx => `Work on issue ${String(ctx.params.issueKey)}.`,
},
});
// Configure the model and Session Store through the Runner's plugins.
const session = await runner.createSession({
id: "issue:SDK-42",
params: { issueKey: "SDK-42" },
});Plugin registration happens once when creating the Runner. The Agent's Program
is built when opening each Session, after loading and validating saved state and
before acquiring resources. ctx.params is available to instructions, enabled,
Tools and capability setup (for example a workspace root). Parameters, state
and capability resources are isolated between Sessions; an Agent's own Tool
values are shared, so keep per-Session data in parameters, state or a capability
resource rather than module variables. Parameters are copied and deeply frozen;
keys listed in params that are missing, and non-JSON values, fail admission.
Values are typed as JSON; the SDK is not a runtime schema validator. Parameters
enter the model's context only where the Agent explicitly uses them in its
instructions.
With a Session Store, nonempty parameters are saved before createSession
returns, even before the first Run. Reopen using the same ID and omit params
to restore the saved values. If supplied again, parameters must be structurally
equal to the saved values. Opening or restoring does not call the model or replay
Tools. Existing v5 snapshots without a params field mean empty parameters;
they cannot be reopened with new nonempty parameters. A Session without parameters
is first saved before the initial Run input is sent to the model, before a
compaction checkpoint commits, or on close when it has queued host state updates.
Resources are acquired once per Session. Each turn then projects current state, instructions and visibility without reacquiring them.
After package publication
Only after the package manifests are published at real versions, install them from the registry:
pnpm add @castle-ai/harness @castle-ai/node-runnerDeclare an Agent value and pass that same value to the Runner:
// agent.ts
import { defineAgent, hostTool } from "@castle-ai/harness";
export const CodingAgent = defineAgent("example.coding", {
model: "provider/model",
instructions: "Read the target, edit it exactly once, run its test, and report.",
tools: ["web_search", "read_workspace_file", "edit_workspace_file", "exec_command", "command_session"]
.map(name => hostTool(name)),
skills: ["review-code"],
sandbox: "local-sandbox",
});Use the linked packed example as the complete starting composition instead of copying a partial snippet with application-owned Adapters left undefined.
Register provider-hosted search at the Node composition boundary. It uses the
same hostTool reference but has no local execute binding:
import { CodingAgent } from "./agent.js";
const runner = await createNodeRunner({
agent: CodingAgent,
plugins: [
{
name: "example.hosted-search",
apply(host) {
host.registerHostedTool({
name: "web_search",
hosted: { kind: "web-search" },
});
host.registerModelResolver(modelResolver);
},
},
],
});An Agent that does not declare web_search does not expose it to the Model.
The provider Adapter converts a declared hosted descriptor to its native wire
shape. Search activity is emitted as typed Harness events and saved in Session
state; reopening history reads those facts without replaying the search.
Real OpenAI composition
openai-agent.ts keeps the
Agent declaration provider-neutral: it declares only the opaque
coding/primary specifier. openai.ts
is the executable composition owner. It binds the access token and complete
model description inside codex, then supplies that
runtime to createNodeRunner. The docs smoke installs the packed
@castle-ai/models/codex and executes this exact entry with mock HTTP;
live provider acceptance is a separate check.
In the installed candidate's examples/local-coding-agent directory, supply
CASTLE_OPENAI_CODEX_ACCESS_TOKEN through your environment, then run:
CASTLE_WORKSPACE=/absolute/path/to/your/project \
pnpm openai -- "Inspect this workspace and fix the failing test."Set CASTLE_WORKSPACE to choose another existing directory,
CASTLE_OPENAI_CODEX_MODEL to override the default model ID, and the optional
comma-separated CASTLE_SANDBOX_ALLOWED_DOMAINS to grant command network
destinations. The entry never prints the token. It prompts on the exact final
input before edits, patches, or commands, then runs approved commands through
the declared OS Sandbox. Ctrl+C aborts the Run and waits for command cleanup.
Repeating the command with the same workspace reopens saved history. Use
CASTLE_SESSION_DIRECTORY to change its default <workspace>/.castle-sessions
location; this example uses one fixed Session ID per history directory.
OpenAI-compatible and OpenRouter composition
For an API-key model, use the complete
examples/session/compatible.ts entry. In
its installed and built example directory, set MODEL_BASE_URL, MODEL_API_KEY
and MODEL_ID, then run:
pnpm chat "Remember our project is Cedar."
pnpm chat "What is our project called?"The second process reopens the saved conversation. This Agent declares only a
model, so it needs no coding tools, sandbox or approval setup. See the
example README for the
configuration and the models README
for protocol limits. For coding tools, extend the existing coding composition;
the compatible runtime does not support provider-hosted web_search.
The compatible adapter preserves text tool replies and emits tool images as
labelled user image content after the reply group. The selected model must
support images. The Codex adapter maps ordered text/image blocks into Responses
function_call_output items. In this repository, the opted-in real
packed-package portability check is:
node --env-file=.env.local \
packages/node-runner/scripts/provider-portability-canary.mjs --require-liveThat portability check requires both an OpenAI Codex access token and the OpenRouter connection variables. It builds the same Agent declaration for each Session and binds the same opaque model specifier to each Castle Adapter. Each selected Provider runs two Runs in one Session: the first performs an approved command Tool call and the second reads that result from Session history. Both Runs verify that streamed events and the final result retain the same execution identity and text.
To exercise only one configured Provider without weakening the default two-Adapter portability check, select it explicitly:
node --env-file=.env.local \
packages/node-runner/scripts/provider-portability-canary.mjs \
--require-live --provider=openrouterThe real visual Tool loop is a separate opted-in check:
node --env-file=.env.local \
packages/node-runner/scripts/provider-visual-canary.mjs --require-liveIt installs clean Castle tarballs into an unrelated consumer, starts a local page through an approved command, captures an unlabeled randomized color as text plus PNG, lets the real model choose the exact approved edit from that image, and captures the changed page again. The check also verifies Tool call identity, Responses image projection, Browser cleanup, and that the access token does not enter Tool content, Session state, stdout, or stderr.
Local capability boundaries
Workspace instructions
For repository guidance, construct one createWorkspaceInstructionsPlugin
with rootDirectory, maxBytes and the host's record callback. Include it in
the Runner's plugins and pass that same object as instructions to
createRootBoundWorkspaceTools for the same root. Root guidance is available
on the first request; targeted file/directory operations discover nested
guidance. New guidance must enter a model request before an affected edit can
execute. Each source is frozen for that Run and rediscovered on later Runs.
This does not inspect arbitrary shell-command paths or promise model compliance.
Workspace Tools
createRootBoundWorkspaceTools binds one existing file-URL root and returns,
in order, list_workspace, search_workspace, read_workspace_file,
edit_workspace_file, and apply_workspace_patch. Model paths are relative
POSIX-style paths. The Tools reject absolute, noncanonical, and escaping paths;
directory traversal does not follow symlinks. List, search, and read results are
stable byte-bounded JSON pages with nextOffset.
read_workspace_file accepts either a one-based startLine (including the
line returned by search) or an exact UTF-8 byte offset, never both. Omit
both to read from the beginning. Lines are LF-delimited and original CRLF bytes
are preserved; an empty file has line 1, and a trailing LF starts an empty
last line. Out-of-range lines fail instead of returning another location.
Line lookup scans with a fixed buffer. Subsequent pages use the returned
nextOffset as offset, omitting startLine, so continuation does not rescan
earlier lines. Output shape and byte limits are identical in either mode.
The workspace root is filesystem authority for only these Tools. It is not a
Sandbox, virtual filesystem, command boundary, or permission grant.
edit_workspace_file requires one exact expected substring. The patch Tool
accepts one or more ordered Add File, Update File, or Delete File operations
in the *** Begin Patch format per call. Each canonical path may occur only once.
A Delete File line has no body and removes an existing regular file; a missing
target fails before permission, and its receipt is { path, operation: "deleted" }. An Update can include context-only @@ hunks: they must match unique,
non-overlapping content in the original file, and do not count as edits in the
receipt. At least one hunk must change content. Stale or ambiguous context fails
before writing that file. All syntax and paths are checked before permission;
files commit individually in order, with no whole-patch rollback. A later
failure or cancellation retains each committed file as an ordered JSON receipt
text block, followed by a failure diagnostic with isError: true. Do not retry
already committed changes. New files require existing parent directories.
The final validated patch exposes paths for complete approval scope. Receipt
capacity is checked before effects; diagnostics may be explicitly truncated,
but committed file receipts are never discarded.
Bounded workspace review
createWorkspaceReviewTool({ rootDirectory, instructions, model }) contributes
review_workspace_files with { task, paths }. Reuse the same instruction plugin
registered on the parent Runner. The model({ sessionId, runId }) callback must
return that active Run's effective { specifier, thinkingLevel, runtime } from
the host's model binding owner; do not resolve a new connection or use stale
Program defaults.
The Tool reads at most 12 UTF-8 files / 128 KiB in total, reuses frozen applicable instruction snapshots, and supplies them to an isolated one-turn Session with no Tools or parent history. Unavailable guidance, invalid paths, non-UTF-8 input and over-limit files fail before calling the reviewer. Parent cancellation closes and drains the child. The 16 KiB bounded review result and file/instruction hashes return through the original Tool call; a model failure is not a successful review.
This is advisory and costs an additional model request. It does not make review mandatory, supply an automatic repair loop, execute tests, or prove completion. The calling Agent decides which feedback warrants edits and verification under its existing permissions. Hosts retain responsibility for logging/accounting; child identities in the Tool result are not admitted Desktop product Runs.
Command Sessions
createCommandSessionPlugin registers the exclusive exec_command and
command_session Tools together. exec_command waits up to the configured
initial-yield deadline. A process that finishes inside that window returns its
terminal result; a live process returns one opaque handle. command_session
polls only new output from that handle or stops and awaits it.
Ordinary { command } calls retain the host's executionTimeoutMs hard deadline.
For a preview or service needed across Runs, the model can explicitly request
{ command: "npm run dev", lifetime: "session" }. That invocation has no execution
timer; it stays under the same supervisor, Sandbox, permission and output limits.
Its final immutable input includes the lifetime for host approval. Hosts must not
treat an existing ordinary-command grant as authorization for a longer lifetime.
Reuse the returned handle instead of shell-backgrounding or restarting a live
service. A normal Run finish does not stop it; Run abort, command_session stop,
or Runner Session close terminates and awaits its process group. This is not a
daemon, persisted process, or automatic restart policy.
Each Tool result contains one command-session-report JSON object. Its
commandId is stable across the initial call and later polls, while stdout
and stderr contain only that report's delta plus explicit truncation
metadata. state: "running" means the process remains live even though the
individual exec_command Tool call has returned; terminal reports carry the
exit, timeout, stop, signal, or cleanup outcome.
Hosts that persist command activity can provide onTerminalUpdate. It receives
each spawned command's final outcome with its original Session, Run, Tool call,
and command identities, including cancellation before the first yield. Its
output is the complete bounded stdout/stderr, independent of poll deltas. This
notification can precede Tool completion and requires no additional model Turn.
Session close waits for outstanding observers and reports their failures.
Such hosts can also provide readTerminal({ sessionId, commandId }). Once a
live record has been released, command_session reads its retained receipt
through this port. The host must scope lookup to the exact owning Session and
return undefined for unknown or foreign handles. A retained receipt contains
the complete bounded final output, not a new delta; repeated reads return the
same result. Reading a stopped or completed handle never launches or stops
another process. Desktop supplies this port from its existing command ledger,
so a later Run can verify a command that finished between Runs or before restart.
Hosts that show running services to their user read them from the same owner.
createSandboxRuntimeCommandSessionPlugin(...).sessionCommands(sessionId) lists
the Session's lifetime: "session" commands whose process has not exited: the
command text, handle, owning Run and Tool call, launch time and the bounded
output since launch (not a poll delta, and reading it consumes nothing).
stopSessionCommand(sessionId, commandId) stops one exactly like the
command_session stop action, publishes its stopped terminal through
onTerminalUpdate, and resolves after that; a command that already ended needs
no stop. The onSessionCommandsChanged({ sessionId }) command option is called
when such a command starts and after it ends, so the host can refresh its list.
Before each agent model request, the same command plugin supplies current
observations of yielded handles: state, exit outcome, code, signal and deadline.
Late terminals supersede initial running observations without rewriting Tool
history or requiring a poll. This context contains no command text or output and
does not authorize another launch. A running process is not proof of service
health. Observations survive compaction in the live Runner Session, are excluded
from compaction requests, and are removed on Session close. On Session start,
the plugin reconstructs unresolved handles from the validated original transcript,
including exchanges omitted from the compacted model context. It resolves them
through readTerminal; without live ownership or a receipt they are explicitly
unavailable. This restores observations only, never processes or command effects.
The supervisor is POSIX-only and non-PTY. It owns process groups in memory per
Runner Session, applies the ordinary command execution deadline unless the
invocation explicitly requests Session lifetime, bounds stdout and
stderr independently per report, and closes every owned process when the
Runner Session closes. A process group still live when the host process exits,
for example because the host exited before its Runner finished closing, is
killed on the host's exit event; only a host killed outright (SIGKILL, a
crash) leaves its commands behind. Handles cannot cross Runner Sessions and are never
reattached after process restart. Reading a persisted terminal receipt does not
reattach a process. The working directory is not a Sandbox.
Every command Module also requires one environment policy. inherit is
"all", "core", or "none"; "core" keeps the host's shell, path, home,
temporary-directory, locale, and user variables. Case-insensitive * / ?
patterns in exclude run after the built-in *KEY*, *SECRET*, and *TOKEN*
filter. Explicit set values run next, then includeOnly narrows the result.
Castle/provider identity variables remain unavailable even through set.
The resolved environment is frozen when the Module is created; Session history
never stores or restores it.
The initial exec_command call should pass the host's effect-approval policy.
Polling or stopping its already-admitted handle can be allowed by a separate
stable policy key; it must not manufacture a second approval for the same
process authority.
Sandboxed Command Sessions
createSandboxRuntimeCommandSessionPlugin contributes the same
exec_command / command_session lifecycle plus the SandboxResolver that
binds them to a real OS process boundary. Declare its exact specifier as
sandbox on every Agent that receives exec_command:
const localCommands = createSandboxRuntimeCommandSessionPlugin({
sandbox: {
specifier: "local/workspace",
config: {
network: { allowedDomains: [], deniedDomains: [] },
filesystem: {
denyRead: ["~/.ssh"],
allowWrite: [workspacePath],
denyWrite: [],
},
},
},
command: {
workingDirectory: pathToFileURL(`${workspacePath}/`),
shell: { executable: "/bin/sh", arguments: ["-c"] },
environment: { inherit: "core" },
executionTimeoutMs: 120_000,
initialYieldMs: 10_000,
pollWaitMs: 1_000,
maxOutputBytesPerStream: 64 * 1024,
},
});The Node Adapter uses @anthropic-ai/sandbox-runtime: Seatbelt on macOS and
bubblewrap/seccomp on Linux. Filesystem and network policy are explicit input;
Castle does not infer policy from the working directory or command approval. The
final rewritten and approved command is transformed into sandboxed argv/env
immediately before the existing command supervisor spawns it. Each invocation
uses its command handle as the Sandbox Runtime attribution key, so filesystem,
seccomp, and proxy denials are appended to that command's stderr without a
Castle-owned regex classifier.
When filesystem confinement is enabled, the Adapter first creates the temporary directory selected by Sandbox Runtime inside the same sandboxed process. The requested command runs only if that succeeds. This does not grant a write path or depend on a directory left by another application. Filesystem-disabled mode keeps its existing environment behavior; command attribution retains the original approved command text.
runtimeAdapter is the narrow host integration seam for
initialize / allowNetworkHosts / prepare / finalize / reset. Omitting it uses the pinned
Sandbox Runtime in the current Node process. An injected Adapter may perform
these operations in a workspace-owned local Worker, but it returns only
the transformed argv: the existing command supervisor still uniquely owns
spawn, process groups, output, timeout, abort, and stop. Finalization is awaited
after process-group cleanup and returns that command's annotated stderr before
the Tool result settles.
One plugin profile can serve multiple concurrent Runner Sessions. Each Session has its own frozen binding and command handles; the final lease resets the profile. The upstream manager is process-global, so a Node process may have one active profile at a time. Another profile fails before initialization instead of replacing policy. Windows command Session process ownership, automatic unsandboxed escalation, containers, remote execution, and restart adoption are not part of this Module.
The plugin also contributes request_network_access; declare it with
hostTool("request_network_access") when the Agent should request
additional domains. Its validated input contains 1–16 exact canonical hosts
and a non-empty reason. The host must apply its Tool permission policy before
execution. Command approval alone does not grant network access.
A grant adds those hosts to the current Runner's Sandbox allowlist. Existing denied domains and filesystem policy remain authoritative. It does not execute or retry the command. All Sessions sharing this Runner share the grant, including Sessions opened after the last lease resets; a fresh Runner starts with the original profile, even when reusing the Plugin object. Desktop maps this lifetime to the current workspace until the app closes. Grants are not stored in history or replayed as effects.
Browser QA
createBrowserSessionPlugin contributes browser_navigate,
browser_snapshot, browser_switch_tab, browser_click, browser_hover, browser_drag, browser_select_option,
browser_fill, browser_press_key, and
browser_capture. The
Agent must declare each Tool it needs. The host supplies an Origin authority, a
BrowserSessionFactory, a text observation byte limit, and an image byte
limit. The authority admits
and canonicalizes each requested URL at the host trust boundary; product
permission consumes that validated admission through the Harness policy hook
instead of parsing the URL again.
The plugin opens a Browser only after browser_navigate admits an Origin. That
Browser context is reused for the Session. Every navigation admits the requested
Origin before the Adapter receives it. newTab: true preserves the current page;
browser_switch_tab accepts an ID from the observation's tabs list and
reactivates that page without navigation. Tab identities expire on restart. Capture and other page actions fail explicitly before
the first navigate. The Browser closes through the normal sessionEnd
lifecycle. Browser actions are exclusive ordinary Tools: final input still
passes permission and approval before the Adapter sees it. Saved transcript
history never restores or adopts an old Browser or page. browser_capture
accepts only integer width/height, light | dark colorScheme, and reduce
| no-preference reducedMotion; it returns one bounded text observation
followed by one PNG block under the same Tool call identity.
Each observation from Castle's Playwright and connected-Chrome Adapters includes
the measured layout viewport (window.innerWidth / innerHeight) in CSS pixels,
including scrollbars. This is fresh page state, not the last requested capture
size or the PNG's device-pixel dimensions. It remains visible in ordinary key,
click and snapshot results across Runs in the same live Session. External
Adapters may omit the optional viewport; the text then explicitly says it was
not reported. Old recorded observations are not retroactively assigned a size.
Text observations return at most the configured byte limit or 8 KiB, whichever
is smaller. A long result ends with browser_snapshot continuation arguments
(observationId, offset). Pass both unchanged to read the next slice of the
same immutable observation without refreshing target refs. Calling snapshot
with no arguments observes the page again. Any new browser action invalidates
the previous continuation, including an action that fails; the page may have
changed. Continuations are Session-local and ephemeral, never recovered from
transcript history. To find relevant content without paging unrelated text, pass
query alone: a case-insensitive literal search returns the first matching line
and following content from that same observation, with unchanged refs. It does
not refresh the browser; no match is a successful read, not a Tool failure.
Screenshot text contains metadata rather than the ARIA tree, so refresh the
snapshot to search page text after a capture.
Accessible targets accept optional within: { role, name?, exact? } to limit
matching to an observed ancestor, for example a date's named region or an
unnamed main. Click, fill, select, and both drag endpoints use this same
contract. Use the current snapshot's roles and names; the resulting target must
match exactly one element. Missing scopes never fall back to the whole page,
and ambiguous targets never select the first match.
browser_hover moves the pointer over one observed visible target without
clicking. Hover the visible row/title to reveal hover-only controls, then use
the returned observation for the next action. It uses the same role/name/ref
and optional ancestor contract as click, without bypassing hit testing.
browser_drag takes source and target, each with the same accessible
role, name, and optional exact fields used by click. It runs in the
current Browser Session and returns the updated observation. Both targets must
resolve uniquely; a missing or ambiguous target is an action failure.
browser_select_option selects a native select's option by its visible label.
The control uses the same accessible role, name, and optional exact
fields. Selection dispatches input/change and returns the updated observation;
custom menus use browser_click. This shares the existing Browser Session,
Origin authority, and permission lifecycle.
For an Agent declaration, add the browser capability from the optional
@castle-ai/connectors/browser entry to capabilities (install playwright-core alongside). It declares these same
ten tools and a browser instructions section, owns a lazy browser per Session, and supports
managed browsing or an explicitly host-selected Chrome tab. See that package's
README for defaults and examples; Node Runner alone does not install a browser
dependency. Custom adapters can use createBrowserSessionTools(options) directly;
its tools, closeSession(id) and close() share the plugin's implementation.
Awaiting close() blocks new calls, releases browsers and settles all admitted
operations, including screenshot writes.
The Playwright Adapter uses accessible role/name locators and returns URL,
title, and a bounded ARIA snapshot after every action. For capture it applies
the requested viewport and media preferences, disables screenshot animation,
and captures the visible viewport as PNG. The Adapter limits top-level navigation
and redirects to the Session's admitted Origins. Its default resourceAccess:
"same-origin" keeps subresources within those Origins; hosts can explicitly select
"http" for websites using external HTTP(S) resources and WS(S) connections.
Service Workers remain disabled. See the Adapter README for the resource policy
and limitations; this is not a general network Sandbox.
Local stdio MCP Tools
Declare a server with the mcp capability. It adds no instructions: its Tools
already carry model-visible descriptions.
import { defineAgent } from "@castle-ai/harness";
import { mcp, openSession } from "@castle-ai/node-runner";
const Agent = defineAgent("example.mcp", {
model,
instructions: "Use the project tools to complete the request.",
capabilities: [mcp({
transport: "stdio",
command: "project-mcp-server",
args: ["--stdio"],
cwd: new URL("file:///absolute/project/"),
serverName: "project",
})],
});
const session = await openSession({ agent: Agent, approval });
try { console.log((await session.run("Read the current issue.").result).text); }
finally { await session.close(); }Declaring the Agent does not spawn the server. Opening each Session starts one
process, discovers its full catalog, and uses that same connection for all calls
until close. No separate plugin or hostTool reference is needed. enabled (a
boolean, or a function of the turn context) hides the whole server for a turn;
a disabled catalog is still acquired and
validated but is absent from model requests and cannot be di
