npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@castle-ai/node-runner

v0.1.0-beta.6

Published

Run Castle agents in Node.js: sessions, persistence, approval, workspaces, MCP and child agents.

Readme

@castle-ai/node-runner

@castle-ai/node-runner is the public Node.js composition root for Castle Agents. It builds a directly held TypeScript Agent declaration into a Program for each Session and binds local Adapters. Start with openSession; applications that manage several Sessions can share one lower-level createNodeRunner. Model/Skill/Sandbox resolution, Session identity and persistence now belong to this package. There is no separate @castle/execution-host package to install.

It includes:

  • Agent capabilities: workspace for workspace tools, commands, OS sandboxing and directory-scoped instructions; skills, mcp, childAgent, goal and instructionsFile;
  • root-bound workspace list, search, read, exact-edit, and patch Tools;
  • createWorkspaceInstructionsPlugin for Run-scoped root and target-directory AGENTS.md context, with explicit byte budget and a host-owned read-receipt callback;
  • createWorkspaceEnvironmentPlugin for observed working-directory and Git worktree facts;
  • POSIX command Session plugins for explicit host or OS-sandboxed processes;
  • a session-scoped local browser QA plugin with an optional Playwright Adapter;
  • a session-scoped local stdio MCP Tool plugin;
  • an explicit provider-hosted Web Search declaration boundary;
  • a local file Session Store and SKILL.md resolver; and
  • bounded foreground child Agents.

Start with examples/session for a small persistent text conversation: Agent, Model binding, file store, streaming and resource closure. It runs again in a new process using the same saved Session. The offline Adapter echoes supplied context so the example needs no credentials. Its OpenRouter entry lists models with their capabilities and starts a persistent conversation with only a key and your selected model ID. The generic API-key entry remains available for other endpoints. The Anthropic entry verifies a native API key, lists the account's catalog and binds the chosen model through the same Session API.

See examples/local-coding-agent for a deterministic, provider-free read → edit → test → final run with approval, automatic compaction, a child Agent, and restart history.

Assistant results expose both result.content (ordered text/image/Provider blocks) and the derived result.text. File stores preserve the content and opaque continuation across processes; opening a saved Session does not start a model request. For image replies, read result.content rather than treating empty text as failure. See the content contract.

Alpha quickstart

The packages currently have private 0.1.0-alpha.1 manifests and are not available from a registry. From this alpha source checkout, run the exact external-package flow:

pnpm --dir packages/node-runner test:docs

That command cleans and packs the required packages, extracts the example from the packed @castle-ai/node-runner tarball, installs only those tarballs in an unrelated temporary consumer, typechecks every example entry, and runs the deterministic coding flow. It executes the real-model CLI with mock HTTP to check tool/search binding, Skill loading, process restart and interrupt cleanup. It also compiles the minimum Session example and runs it across independent OS processes, checking continuation and rejection of invalid or wrong-Agent history without overwriting saved files. The API-key entry is also run across processes against a local HTTP fixture, without provider credentials. None of these checks uses workspace package resolution. The OpenRouter entry additionally verifies public discovery, inference-key rejection, selected capabilities, unknown-model rejection and interruption during both discovery and streaming against a local HTTP fixture. The Anthropic entry checks its native authenticated, paginated catalog, selected thinking/output limits, cross-process history and discovery/stream cancellation.

To keep a portable local candidate with versioned tarballs, both examples and relative dependency bindings, choose a new output directory:

pnpm pack:sdk /absolute/path/to/new-castle-sdk-candidate

This uses the same packed-example verification, then retains the archives, example sources and tested lockfiles. It removes installed dependencies and QA history before delivery. Follow the generated README to install either example from the kit; the output folder must not already exist. No registry publication occurs. The seven runtime packages share 0.1.0-alpha.1; optional source-build and private product packages are outside this candidate.

The complete, compiling quickstart is intentionally split by responsibility:

  • agent.ts declares the parent and child with defineAgent, a hostTool reference, and the skills and workspace capabilities;
  • run.ts provides the complete public createNodeRunner composition, including the Session Store, permission Hook, approval Adapter, compaction budget, foreground child, cancellation-compatible event stream, and restart; and
  • mock-model.ts supplies the deterministic provider-neutral ModelRuntime used by the smoke.

Declare a local workspace

Add workspace({ root }) to an Agent's capabilities. It contributes the file, image and command Tools, the Sandbox, and a workspace instructions section with their usage rules. It opens the root once per Session and loads applicable AGENTS.md guidance for each Run. The coding example uses this capability without manually registering file and command tools.

import { defineAgent } from "@castle-ai/harness";
import { workspace } from "@castle-ai/node-runner";

export const Coder = defineAgent("example.coder", {
  model,
  instructions: "Make the smallest correct change, run the tests, and report.",
  capabilities: [workspace({
    root: new URL("./project/", import.meta.url),
    network: { allowedDomains: ["registry.npmjs.org"] },
  })],
});

root is a file URL, or a function of the Session's parameters returning one, such as root: params => new URL(String(params.checkout)). A fixed root is validated when the capability is created, a derived one when each Session opens. The defaults are /bin/sh -c, a 120-second ordinary-command timeout, 64 KiB per output stream, the core environment, and sandbox: "os". An unavailable OS sandbox fails opening; it never falls back to unsandboxed execution. sandbox: "none" explicitly runs commands without OS containment. readOnly exposes only file and image reading tools, with no commands or Sandbox. network.allowedDomains / deniedDomains set the OS sandbox's rules; omission permits no external hosts. Publication requires a host publisher that returns a verified delivery reference. An OS workspace declares the Agent's one Sandbox, so the Agent and its other capabilities cannot declare another. Declare a Skill directory separately with skills; installation and choosing source scopes remain the application's responsibility.

Each OS workspace Session owns an isolated sandbox worker, including when two Sessions use the same root or a parent delegates to a child. Network hosts approved through request_network_access last only for that Session. Closing waits for its commands, resets its sandbox and terminates its worker; another Session keeps its own configuration and process lifecycle.

Source development must build the worker with pnpm --filter @castle-ai/node-runner build:worker before opening an OS workspace. The package's test command includes this step; direct Vitest or source CLI invocations need it explicitly. Both source and installed use the fixed compiled worker entry. Published tarballs include it, so consumers need no TypeScript loader or build step. The sandbox-worker package subpath is for Node/Desktop host integration, not an additional Agent-authoring API.

Bind your models

Models can come directly from the Agent declaration. For example, with an Anthropic API key and an explicit model ID:

import { defineAgent } from "@castle-ai/harness";
import { openSession } from "@castle-ai/node-runner";
import { anthropic } from "@castle-ai/models/anthropic";

const model = anthropic({
  apiKey: process.env.ANTHROPIC_API_KEY!,
  model: process.env.ANTHROPIC_MODEL!,
});
const Assistant = defineAgent("app.assistant", {
  model,
  instructions: `Answer clearly using ${model.specifier}.`,
});
const session = await openSession({ agent: Assistant });
try {
  console.log((await session.run("Explain an Agent Session.").result).text);
} finally {
  await session.close();
}

anthropic defaults to a 4,096-token response limit; use maxOutputTokens to set another limit. The other native factories are openaiCompatible, openrouter and codex in their respective runtime packages. Their connection inputs retain the documented provider-specific shapes; a shorter name does not add catalog discovery or credential storage. The old create*ModelRuntime factory names and createTerminalToolApprovalPlugin are removed in this alpha contract reset.

openSession({ agent, id?, params?, store?, approval?, permissions?, signal? }) uses the existing Runner lifecycle internally. store accepts a SessionStore directly; approval accepts a NodeRunnerPlugin, including its cleanup (for example, the existing terminalApproval). Opening with the same ID and Store restores history and parameters without starting a model or Tool. Closing waits for active work to settle and releases the Session's private Runner. Failed opening also cleans up; an already-aborted signal rejects before plugins are applied. Always close a successfully opened Session in finally.

Choose a Tool permission preset

openSession defaults to ask-on-write. Only an executable Tool declared by the application with readOnly: true qualifies as read-only.

| permissions | Declared read-only | Other executable Tools | | --- | --- | --- | | read-only | Allow | Deny | | ask-on-write | Allow | Ask through the configured approval plugin | | full-access | Allow | Allow |

Without an approval adapter, a call that requires approval is denied. A preset answers only calls that every toolPermission hook left undecided (no answer or defer). A host policy's deny wins, then its ask, then its allow, so a host can tighten or relax the preset for the calls it recognizes. Hooks and approval requests carry agent, the identity of the Agent whose model made the call, and for a call inside a foreground child Agent, parent: the parent Session, Run, Tool call id and Tool name that started it. Presets do not remove root or OS-sandbox limits, classify provider-hosted actions, or infer effects from a Tool's name, description, parallel scheduling or MCP annotations. Declare readOnly: true only when the whole Tool is read-only; a Tool with mixed effects stays unclassified.

The built-in workspace list, search, file reader and image reader declare this trait. Edits and commands require approval under the default preset. Existing hosts using createNodeRunner keep their own policy unless they explicitly pass permissions; the SDK modes are distinct from Desktop's product permission modes. Direct store/approval options and plugins register through the same owner; registering two Stores or two approval adapters is rejected.

Model binding details

Anthropic, OpenAI-compatible and OpenRouter factories return a ModelBinding: the streaming runtime plus its specifier and known contextWindow. A custom runtime can provide those same fields; a runtime with dynamic connections, such as the current Codex Adapter, can be named with { specifier, ...runtime }. The connection and credentials remain inside the runtime; the Program contains only its model reference. No model connection is opened when the Agent is declared or a Session opens.

model: "app/chat" selects a named host binding. A missing registered name fails on Session open; a dynamic ModelResolver still resolves the admitted Run's model when that request starts. A directly declared model and a registered model with the same name are ambiguous and rejected. Explicit Run or compaction model changes resolve the requested name; they never silently reuse a different declared model. Plugins are optional when the declaration supplies its own model.

When your application already has a ModelRuntime, register it under the same name used in the Agent's model:

// Inside a plugin's apply(host):
host.registerModel("app/chat", chatModel);
host.registerModel("app/review", reviewModel);

The Agent declares model: "app/chat". These names are opaque application bindings; the Provider Adapter still owns the real model ID, endpoint and credentials. No resolver callback is needed for fixed bindings. Multiple plugins may register distinct names, and the same bindings serve Runs, foreground children and compaction. An unknown name fails before contacting a model; there is no default-model fallback. Registration closes at Runner creation, and later replacement of the passed runtime's stream method has no effect.

Use host.registerModelResolver(resolver) instead when the host must resolve a connection for each request, such as Desktop's changing authenticated accounts. One Runner uses either named models or one resolver; mixing them and registering a name twice are errors. Both feed the same Runner model-resolution path and preserve the admitted model name and request identity. See the complete API-key example for fixed bindings and the existing composition example for a resolver.

A typed Tool value in tools supplies its contract and implementation. Declaring it does not execute it, and the Program contains only data. For host-owned catalogs, hostTool("lookup") resolves the implementation from host.registerTool(lookupTool); missing names fail when a Session opens. Registered Tools run with the same Tool context as declared ones. Binding the same name both directly and through the host is an error. Each foreground child builds its own Program and still obeys toolGrants.

enabled on a Tool, a hostTool reference or a capability controls whether the model sees it. It is true, false, or a function of the turn context evaluated at each turn. Hidden Tools still need valid contracts and implementations; they are absent from model requests and cannot be called by a stale model name. Skills hidden with their capability cannot be activated explicitly or by the model, cannot read reference files, and do not project saved activation instructions. Their durable history remains intact. beforeModel hooks may transform messages but must preserve the projected Tool list and contracts.

A SkillBinding in skills binds a value directly. Metadata is available at Session admission; load() and readResource() run only through Skill activation/resource reading. A string in skills uses the host Skill resolver, so direct and named Skills can coexist. A direct value owns its reference; the host resolver is only called for the remaining named references.

sandbox: { specifier, open } contributes an acquisition request. The Runner calls open({ signal }) once when opening that Session and releases the returned { sandbox, close } lease on close or admission failure. sandbox: "name" uses the host Sandbox resolver instead. Declaring never acquires the Sandbox; both forms use the existing admission and cleanup owner. A capability's resource(kind, spec) shares this Session acquisition/cleanup owner. Each typed kind owns acquire(spec, { signal }) → { value, close }; a frozen handle exposes value after acquisition. See the Harness capability contract. BindingKind, BindingHandle, BindingLease and SandboxLease are exported from @castle-ai/node-runner, which owns acquisition and cleanup. Host approval plugins also import ToolPermissionContribution, ToolApprovalAdapter, ToolApprovalRequest, ToolApprovalAnswer and ToolExecutionHostError from this package. These are host contracts, not Harness authoring exports.

Load application instructions

instructionsFile(fileUrl, { name?, placement? }) is a capability whose one section is read from a file:

import { defineAgent } from "@castle-ai/harness";
import { instructionsFile } from "@castle-ai/node-runner";

export const ReviewAgent = defineAgent("example.file-review", {
  model: "review-model",
  instructions: "Review the requested change.",
  capabilities: [instructionsFile(new URL("./review-rules.md", import.meta.url))],
});

The section is named after the file (review-rules.md becomes review-rules) unless name is given, and renders with that heading after the Agent's own sections. The application authorizes that explicit file URL and its resolved target, including a symlink target. The Runner resolves and reads a regular UTF-8 file once per Session as a capability resource, closes the file, and retains immutable text. A changed or deleted file does not alter an open Session; a newly opened Session reads again. The default placement is system; { placement: "context" } sends the same text as current user-side context each turn. Contents are not written into Session history.

Declaring the Agent does no file I/O. projectAgent returns a pending resource reference for a system-placed file; it cannot preview a context-placed one, whose text exists only after acquisition. After acquiring resources the Runner performs its initial projection, so even a Run stopped before its first model request has resolved instructions. Missing, non-regular or invalid UTF-8 files, and blank system-placed files, fail Session opening and release already acquired resources. This is application prompt loading, separate from workspace authority and AGENTS.md's per-Run discovery/refresh policy. It does not impose the later K4 byte budget.

Change instructions and Tools each turn

Instruction sections and enabled options may be functions of the turn context. The Runner builds the Program once when a Session opens, runs these functions once after acquiring resources, then again synchronously before each agent turn. Use them to change instructions and visibility without changing a Tool's contract:

import { defineAgent, defineTool } from "@castle-ai/harness";

const Agent = defineAgent("example.bounded-search", {
  model,
  instructions: {
    identity: "Answer questions about the project.",
    phase: ctx => ctx.turn <= 3 ? "Search for relevant evidence." : "Answer using the evidence already collected.",
  },
  tools: [defineTool({
    name: "search",
    description: "Search the project",
    inputSchema: { type: "object", properties: { query: { type: "string" } }, required: ["query"] },
    readOnly: true,
    enabled: ctx => ctx.turn <= 3,
    async execute({ query }, ctx) {
      return await searchProject(query, { turn: ctx.turn });
    },
  })],
});

model is your bound model and searchProject is application code. turn is the turn within the current Run. The Tool's context carries the turn that requested the call; its schema and description remain fixed. Foreground child projections remain restricted to their original Tool grants. Use projectAgent from @castle-ai/harness/testing to assert Tool visibility and instructions for specific state/parameter inputs.

Parameters and acquired resources are Session-local. Projection does not reopen the Sandbox, reload file resources, or rerun settled Tools. Skill activation still uses the normal Tool/explicit-input path; instructions of Skills hidden with their capability are omitted from the current projection while their original history is preserved. Run binding/model selection, manual compaction and restoration keep their existing explicit lifetimes. Resources belong to capabilities; writable Session state follows the contract below.

Inspect a live Session

session.inspect() returns a frozen, read-only view of accepted declarations. It does not rerun the Agent, acquire resources, or make model requests:

const inspection = session.inspect();
console.log(inspection.current.declaration?.capabilities); // Capability names in declaration order
console.log(inspection.current.declaration?.placements); // Section placement table
console.log(inspection.turns.at(-1)?.changes); // Tools, Skills and sections changed
console.log(inspection.stateWrites); // Committed writes with source / Run / turn

The SessionInspection type is exported from @castle-ai/node-runner. initial is the projection after resource acquisition; turns records successful bound projections with Run identity and turn number, and current is the latest one. Failed projections do not add turn entries. An explicit same-Run resume can observe a turn again; entries are observations, not unique historical turn records.

Tool visibility includes internal Skill tools and child grant restrictions. Sections and tools describe the accepted declaration before beforeModel hooks rewrite a request; this is not the final Provider payload. Each placement has the section's identity, name, source (agent, or capability with its name), placement, lifetime and order. order is declaration order, not a Provider message index. Manually supplied Programs have no declaration and report declaration: null.

stateWrites includes only settled state facts after the configured Store acknowledges them. Tentative or aborted Tool writes are excluded. pendingState is the current host-write queue, not a durability acknowledgement; continue to await session.setState() when queuing a durable host write.

projectionScope: "current-open" makes the lifetime explicit: reopening restores saved state facts, but starts a new projection history. Reading after close still returns cached data. Inspection omits acquired handles, binding specs, model credentials and inactive Skill bodies; instruction and state content is visible as authored and is not automatically redacted.

React to settled Tool results

import { defineAgent, defineTool } from "@castle-ai/harness";

const createReport = defineTool({
  name: "create_report",
  description: "Create the weekly report.",
  enabled: ctx => !ctx.toolCalls("create_report").some(call => call.ok),
  async execute() { return await createWeeklyReport(); },
});

const Agent = defineAgent("app.report-once", {
  model,
  instructions: {
    task: "Create the report and summarize the result.",
    retry: ctx => ctx.toolCalls("create_report").at(-1)?.ok === false
      ? "The last report attempt failed. Inspect its result before trying again."
      : "",
  },
  tools: [createReport],
});

createWeeklyReport is application code. After a successful result, the next turn hides the Tool; an error leaves it available and adds the retry context section, which is omitted while empty. ctx.toolCalls(name) lists settled receipts with input, the complete result, ok, runId and turn. These are frozen Session facts, retained across new Runs, compaction and restart. They do not replace an external service's idempotency contract when a process fails before saving a receipt.

The incident-reporter example reads the same receipts to report completed updates without storing a second counter. Its installed-package test reads those receipts in a fresh OS process after compaction.

SandboxContribution and BuiltinPromptOverrides are host configuration types exported from @castle-ai/node-runner, alongside SandboxLease and HarnessContextBudget.

Keep state across Tools and turns

Most progress can be derived from ctx.toolCalls. Use durable state for values the model never produced or that cannot be derived, such as an external ID:

import { defineAgent, defineTool } from "@castle-ai/harness";

const Reporter = defineAgent("app.reporter", {
  model,
  state: { updates: 0 },
  instructions: {
    identity: "Record completed report updates.",
    progress: ctx => `Report updates so far: ${ctx.state.updates}.`,
  },
  tools: [defineTool({
    name: "record_update",
    description: "Record a completed report update",
    enabled: ctx => (ctx.state.updates as number) < 3,
    async execute(_input, ctx) {
      ctx.setState(state => ({ updates: (state.updates as number) + 1 }));
      return `Recorded update ${String(ctx.state.updates)}.`;
    },
  })],
});

Here model is a bound Model as above. state declares keys and their initial JSON values; keys from the Agent and its capabilities share one flat space, and a key declared twice fails when a Session opens. Instructions and enabled read the committed values. Inside the Tool, ctx.state includes earlier writes in the same batch, so the result above reports the new count. setState takes a patch of declared keys, or a function of the latest state returning one. Values must be JSON and are deeply frozen. Updates commit only after the Tool batch settles. Abort or framework failure discards tentative state, while Tool receipts remain. Ordinary Tool error results still settle. Queuing a run.join() input does not itself abort the batch.

The host can call await session.setState("updates", 0) to queue an update for the next turn. This does not change the current Tool's reads. Writes emit state-changed on the Run and host.hooks.stateChanged after commitment, with source: "host" or "tool". State is private unless the Agent puts it into instructions.

The file Store appends committed state facts before publishing state-changed. await setState confirms that the queued write is saved, including when the Session is idle. It checkpoints the current history and pending queue through the same ordered Store path as transcript appends. Reopening waits for an explicit Run before applying the queue. During a text-only turn, queued host updates cause a further turn unless a stop condition intervenes. Without a Store the acknowledgment is in memory. A failed save rejects and stops active work; close waits for writes and releases resources. Each host write currently requires a full checkpoint, so this API is not a high-frequency telemetry channel. Compaction preserves state facts. Each batch records its start and a committed or discarded settlement. Committed state and that settlement share one append; notification failure afterward cannot change the decision. Individual receipts still precede the batch's state decision; interrupted work uses the explicit same-Run continuation and host reconciliation contracts below. See the executable repository example stateful-report.ts, verified with installed packages and two separate processes by test:package.

HarnessContextBudget is a Runner configuration type and is imported from @castle-ai/node-runner.

Declare a durable Goal

The goal capability declares get_goal and update_goal and a goal context section with the current Goal observation. With an existing ModelBinding named model:

import { defineAgent } from "@castle-ai/harness";
import { createFileGoalStore, createFileSessionStore, goal, openSession } from "@castle-ai/node-runner";

const goals = createFileGoalStore({ directory: new URL("./goals/", import.meta.url) });
const agent = defineAgent("release-review", {
  model,
  instructions: "Review the project against its release ticket.",
  capabilities: [goal({ store: goals })],
});
const session = await openSession({
  agent, id: "release-review",
  store: createFileSessionStore({ directory: new URL("./sessions/", import.meta.url) }),
});
try {
  // Open only in response to the user's explicit request.
  await goals.control({ sessionId: session.id, action: "open", objective: "Complete the release review." });
  const result = await session.run("Review the requirements and current implementation.").result;
  console.log(result.text);
} finally {
  await session.close();
}

The Goal Store owns the objective and lifecycle. Host controls open, edit, pause, resume or delete it; Agent tools can read and report complete/blocked using the exact current identity and objective. A paused, deleted or changed Goal rejects an obsolete settlement. Task completion never completes a Goal. GoalPort lets a product adapt its existing authoritative Store without a Conversation DTO.

Each new turn reads the Store before projection and durably queues that current observation through Session state (key castle.goal.observation), superseding any older pending observation. update_goal is visible only while the latest observation is an active Goal. The turnStart lifecycle hook runs before queued host writes are applied and the turn is checkpointed/projected; replaying an existing checkpoint does not call it. The context section is not appended as a conversation message. Existing checkpoint replay preserves the original projection; the next new turn refreshes it. Opening or reopening a Session does not read the Goal, run the model, resume paused work, or schedule continuation. Store read/write failures preserve their cause.

Omitting the capability's Store uses one shared file Store at .castle/goals relative to the current working directory; it does not open a Goal. Use an explicit shared Store for host controls. File Stores require one writing instance/process per Session; separate Sessions can share a directory. Goal storage and Session history storage are distinct.

Continue a Goal across Runs

Call driveGoal explicitly for an open, idle Session whose Agent declares the goal capability. Pass the same Store to both. For example:

import { driveGoal } from "@castle-ai/node-runner";

const stop = new AbortController();
const result = await driveGoal(session, {
  store: goals,
  objective: "Complete the release review.",
  policy: { idleMs: 1_000, maxCheckins: 5 },
  observe: () => ({ queuedInput: false, interactionPending: false }),
  signal: stop.signal,
  async onRun(run) {
    for await (const event of run) {
      if (event.kind === "assistant-text-delta") process.stdout.write(event.text);
    }
  },
});
console.log(result.decision, result.runs);

The example has no UI queue; a UI host supplies its actual queued-input and pending-interaction facts synchronously. onRun exposes the real Run, including its event stream and join. Normal Session permissions still apply. The driver does not answer approvals, grant permissions, or change the Agent's toolset.

An objective is an explicit host request to open a Goal. A different unfinished objective is rejected; a paused or blocked Goal is never implicitly resumed. Omit objective to work on the existing Goal. The first Run starts immediately; maxCheckins caps subsequent automatic Runs, separated by idleMs. The defaults are 60 seconds and five check-ins. Exhausting the allowance leaves the Goal active.

After each Run, the driver reads authoritative Goal state. It checks synchronous host facts again after onDecision and immediately before the next admission. Queued input, pending interactions, an existing Run, pause, deletion or completion return a typed hold to the caller. Only the idle cooldown waits internally; a host must explicitly call again after resolving another hold. Stop conditions return run-stopped; failures reject with their cause. An edited objective is reread, and a replacement Goal identity ends this invocation. Live command handles can be provided in observe().activeCommandSessions for a check-in; they do not force the Agent to wait for a healthy service to terminate.

Abort stop to stop the driver and its own active Run without closing the Session. If a user pauses or deletes a Goal during a Run, apply that Store control and abort this signal together; Store controls alone do not interrupt executing tools. session.signal exposes the existing Session lifetime: closing the Session or Runner also cancels the driver and its cooldown. Opening saved history never starts a driver. An explicit new call arms it again; it never automatically retries an interrupted Run or consumes saved unanswered input.

Each admitted input carries an opaque, durable ID of the form castle.goal/<encoded-goal-id>/<uuid>. This records the host's Goal origin outside model messages using the existing input identity contract; it is not a new Harness message type. History, Goal storage and task progress remain independent.

Hosts that already own a durable input queue can use createGoalDriverController, the same controller used by driveGoal. It owns cooldown, allowance and cancellation generations; the host keeps its existing queue and Run executor. Supply observe for current Goal and host facts, and admit(goal, decision, admission) to prepare a submission. Call admission.accept(freshFacts) synchronously immediately before durably admitting it, with no asynchronous gap. Return true only for an accepted submission. Use admission.signal to cancel preparation where supported.

Explicit user activity calls arm; wake rechecks facts without restarting the idle deadline, and idle starts a fresh cooldown after a Run settles. cancel disarms the generation; close also waits for any preparation already in flight. An old generation's failure reaches onError without disarming a newer one. snapshot exposes armed state and counts for host presentation. A controller starts disarmed, so restoring a Goal alone cannot start work. By default every admitted continuation counts toward the cap. driveGoal arms with initialRun to exclude its explicit initial Run; Desktop retains its cap on all automatic submissions. Host adapters remain responsible for stopping their active Runs.

Track tasks with durable defaults

createTaskTools() provides create, read, list, revise and progress tools. Its default file Store is .castle/tasks under the working directory where the tools are created. Declaring tools performs no I/O. With an existing model:

import { defineAgent } from "@castle-ai/harness";
import { createTaskTools, openSession } from "@castle-ai/node-runner";

const agent = defineAgent("release-work", {
  model,
  instructions: "Track meaningful release deliverables and their dependencies.",
  tools: createTaskTools(),
});
const session = await openSession({ agent, id: "release-work", permissions: "full-access" });
try {
  console.log((await session.run("Plan the release tasks.").result).text);
} finally {
  await session.close();
}

This example grants execution to its Task-only toolset. In an application with other tools, keep the appropriate permission policy and approval adapter; Task writes go through the same authorization path as other tools.

Task lists are scoped to the Session ID. Reusing that ID restores its Task list; persist conversation history separately with a Session Store. Creation records a pending task and does not start execution. Revisions preserve identity and progress; dependencies must exist in the list and cannot form a cycle. Pending tasks have progress 0, completed tasks 100, and in-progress tasks require an integer measurement from 0 to 100. Completing the list never completes a Goal.

To read tasks from the host or choose a location, create one createFileTaskStore({ directory: new URL("./tasks/", import.meta.url) }), pass it to createTaskTools(store), and use store.load(session.id). TaskStore extends the neutral TaskPort; an application can still provide its own TaskPort. Use one writing Store instance/process per Session. Saved task corruption is reported with its cause and is never replaced by an empty list.

Give each Session its own parameters

Use immutable JSON parameters for constants such as an issue key or tenant ID:

import { defineAgent } from "@castle-ai/harness";

const IssueAgent = defineAgent("app.issue-agent", {
  model: "app/chat",
  params: ["issueKey"],
  instructions: {
    identity: "Resolve the Session's issue.",
    issue: ctx => `Work on issue ${String(ctx.params.issueKey)}.`,
  },
});

// Configure the model and Session Store through the Runner's plugins.
const session = await runner.createSession({
  id: "issue:SDK-42",
  params: { issueKey: "SDK-42" },
});

Plugin registration happens once when creating the Runner. The Agent's Program is built when opening each Session, after loading and validating saved state and before acquiring resources. ctx.params is available to instructions, enabled, Tools and capability setup (for example a workspace root). Parameters, state and capability resources are isolated between Sessions; an Agent's own Tool values are shared, so keep per-Session data in parameters, state or a capability resource rather than module variables. Parameters are copied and deeply frozen; keys listed in params that are missing, and non-JSON values, fail admission. Values are typed as JSON; the SDK is not a runtime schema validator. Parameters enter the model's context only where the Agent explicitly uses them in its instructions.

With a Session Store, nonempty parameters are saved before createSession returns, even before the first Run. Reopen using the same ID and omit params to restore the saved values. If supplied again, parameters must be structurally equal to the saved values. Opening or restoring does not call the model or replay Tools. Existing v5 snapshots without a params field mean empty parameters; they cannot be reopened with new nonempty parameters. A Session without parameters is first saved before the initial Run input is sent to the model, before a compaction checkpoint commits, or on close when it has queued host state updates.

Resources are acquired once per Session. Each turn then projects current state, instructions and visibility without reacquiring them.

After package publication

Only after the package manifests are published at real versions, install them from the registry:

pnpm add @castle-ai/harness @castle-ai/node-runner

Declare an Agent value and pass that same value to the Runner:

// agent.ts
import { defineAgent, hostTool } from "@castle-ai/harness";

export const CodingAgent = defineAgent("example.coding", {
  model: "provider/model",
  instructions: "Read the target, edit it exactly once, run its test, and report.",
  tools: ["web_search", "read_workspace_file", "edit_workspace_file", "exec_command", "command_session"]
    .map(name => hostTool(name)),
  skills: ["review-code"],
  sandbox: "local-sandbox",
});

Use the linked packed example as the complete starting composition instead of copying a partial snippet with application-owned Adapters left undefined.

Register provider-hosted search at the Node composition boundary. It uses the same hostTool reference but has no local execute binding:

import { CodingAgent } from "./agent.js";

const runner = await createNodeRunner({
  agent: CodingAgent,
  plugins: [
    {
      name: "example.hosted-search",
      apply(host) {
        host.registerHostedTool({
          name: "web_search",
          hosted: { kind: "web-search" },
        });
        host.registerModelResolver(modelResolver);
      },
    },
  ],
});

An Agent that does not declare web_search does not expose it to the Model. The provider Adapter converts a declared hosted descriptor to its native wire shape. Search activity is emitted as typed Harness events and saved in Session state; reopening history reads those facts without replaying the search.

Real OpenAI composition

openai-agent.ts keeps the Agent declaration provider-neutral: it declares only the opaque coding/primary specifier. openai.ts is the executable composition owner. It binds the access token and complete model description inside codex, then supplies that runtime to createNodeRunner. The docs smoke installs the packed @castle-ai/models/codex and executes this exact entry with mock HTTP; live provider acceptance is a separate check.

In the installed candidate's examples/local-coding-agent directory, supply CASTLE_OPENAI_CODEX_ACCESS_TOKEN through your environment, then run:

CASTLE_WORKSPACE=/absolute/path/to/your/project \
pnpm openai -- "Inspect this workspace and fix the failing test."

Set CASTLE_WORKSPACE to choose another existing directory, CASTLE_OPENAI_CODEX_MODEL to override the default model ID, and the optional comma-separated CASTLE_SANDBOX_ALLOWED_DOMAINS to grant command network destinations. The entry never prints the token. It prompts on the exact final input before edits, patches, or commands, then runs approved commands through the declared OS Sandbox. Ctrl+C aborts the Run and waits for command cleanup. Repeating the command with the same workspace reopens saved history. Use CASTLE_SESSION_DIRECTORY to change its default <workspace>/.castle-sessions location; this example uses one fixed Session ID per history directory.

OpenAI-compatible and OpenRouter composition

For an API-key model, use the complete examples/session/compatible.ts entry. In its installed and built example directory, set MODEL_BASE_URL, MODEL_API_KEY and MODEL_ID, then run:

pnpm chat "Remember our project is Cedar."
pnpm chat "What is our project called?"

The second process reopens the saved conversation. This Agent declares only a model, so it needs no coding tools, sandbox or approval setup. See the example README for the configuration and the models README for protocol limits. For coding tools, extend the existing coding composition; the compatible runtime does not support provider-hosted web_search.

The compatible adapter preserves text tool replies and emits tool images as labelled user image content after the reply group. The selected model must support images. The Codex adapter maps ordered text/image blocks into Responses function_call_output items. In this repository, the opted-in real packed-package portability check is:

node --env-file=.env.local \
  packages/node-runner/scripts/provider-portability-canary.mjs --require-live

That portability check requires both an OpenAI Codex access token and the OpenRouter connection variables. It builds the same Agent declaration for each Session and binds the same opaque model specifier to each Castle Adapter. Each selected Provider runs two Runs in one Session: the first performs an approved command Tool call and the second reads that result from Session history. Both Runs verify that streamed events and the final result retain the same execution identity and text.

To exercise only one configured Provider without weakening the default two-Adapter portability check, select it explicitly:

node --env-file=.env.local \
  packages/node-runner/scripts/provider-portability-canary.mjs \
  --require-live --provider=openrouter

The real visual Tool loop is a separate opted-in check:

node --env-file=.env.local \
  packages/node-runner/scripts/provider-visual-canary.mjs --require-live

It installs clean Castle tarballs into an unrelated consumer, starts a local page through an approved command, captures an unlabeled randomized color as text plus PNG, lets the real model choose the exact approved edit from that image, and captures the changed page again. The check also verifies Tool call identity, Responses image projection, Browser cleanup, and that the access token does not enter Tool content, Session state, stdout, or stderr.

Local capability boundaries

Workspace instructions

For repository guidance, construct one createWorkspaceInstructionsPlugin with rootDirectory, maxBytes and the host's record callback. Include it in the Runner's plugins and pass that same object as instructions to createRootBoundWorkspaceTools for the same root. Root guidance is available on the first request; targeted file/directory operations discover nested guidance. New guidance must enter a model request before an affected edit can execute. Each source is frozen for that Run and rediscovered on later Runs. This does not inspect arbitrary shell-command paths or promise model compliance.

Workspace Tools

createRootBoundWorkspaceTools binds one existing file-URL root and returns, in order, list_workspace, search_workspace, read_workspace_file, edit_workspace_file, and apply_workspace_patch. Model paths are relative POSIX-style paths. The Tools reject absolute, noncanonical, and escaping paths; directory traversal does not follow symlinks. List, search, and read results are stable byte-bounded JSON pages with nextOffset.

read_workspace_file accepts either a one-based startLine (including the line returned by search) or an exact UTF-8 byte offset, never both. Omit both to read from the beginning. Lines are LF-delimited and original CRLF bytes are preserved; an empty file has line 1, and a trailing LF starts an empty last line. Out-of-range lines fail instead of returning another location. Line lookup scans with a fixed buffer. Subsequent pages use the returned nextOffset as offset, omitting startLine, so continuation does not rescan earlier lines. Output shape and byte limits are identical in either mode.

The workspace root is filesystem authority for only these Tools. It is not a Sandbox, virtual filesystem, command boundary, or permission grant. edit_workspace_file requires one exact expected substring. The patch Tool accepts one or more ordered Add File, Update File, or Delete File operations in the *** Begin Patch format per call. Each canonical path may occur only once. A Delete File line has no body and removes an existing regular file; a missing target fails before permission, and its receipt is { path, operation: "deleted" }. An Update can include context-only @@ hunks: they must match unique, non-overlapping content in the original file, and do not count as edits in the receipt. At least one hunk must change content. Stale or ambiguous context fails before writing that file. All syntax and paths are checked before permission; files commit individually in order, with no whole-patch rollback. A later failure or cancellation retains each committed file as an ordered JSON receipt text block, followed by a failure diagnostic with isError: true. Do not retry already committed changes. New files require existing parent directories.

The final validated patch exposes paths for complete approval scope. Receipt capacity is checked before effects; diagnostics may be explicitly truncated, but committed file receipts are never discarded.

Bounded workspace review

createWorkspaceReviewTool({ rootDirectory, instructions, model }) contributes review_workspace_files with { task, paths }. Reuse the same instruction plugin registered on the parent Runner. The model({ sessionId, runId }) callback must return that active Run's effective { specifier, thinkingLevel, runtime } from the host's model binding owner; do not resolve a new connection or use stale Program defaults.

The Tool reads at most 12 UTF-8 files / 128 KiB in total, reuses frozen applicable instruction snapshots, and supplies them to an isolated one-turn Session with no Tools or parent history. Unavailable guidance, invalid paths, non-UTF-8 input and over-limit files fail before calling the reviewer. Parent cancellation closes and drains the child. The 16 KiB bounded review result and file/instruction hashes return through the original Tool call; a model failure is not a successful review.

This is advisory and costs an additional model request. It does not make review mandatory, supply an automatic repair loop, execute tests, or prove completion. The calling Agent decides which feedback warrants edits and verification under its existing permissions. Hosts retain responsibility for logging/accounting; child identities in the Tool result are not admitted Desktop product Runs.

Command Sessions

createCommandSessionPlugin registers the exclusive exec_command and command_session Tools together. exec_command waits up to the configured initial-yield deadline. A process that finishes inside that window returns its terminal result; a live process returns one opaque handle. command_session polls only new output from that handle or stops and awaits it.

Ordinary { command } calls retain the host's executionTimeoutMs hard deadline. For a preview or service needed across Runs, the model can explicitly request { command: "npm run dev", lifetime: "session" }. That invocation has no execution timer; it stays under the same supervisor, Sandbox, permission and output limits. Its final immutable input includes the lifetime for host approval. Hosts must not treat an existing ordinary-command grant as authorization for a longer lifetime. Reuse the returned handle instead of shell-backgrounding or restarting a live service. A normal Run finish does not stop it; Run abort, command_session stop, or Runner Session close terminates and awaits its process group. This is not a daemon, persisted process, or automatic restart policy.

Each Tool result contains one command-session-report JSON object. Its commandId is stable across the initial call and later polls, while stdout and stderr contain only that report's delta plus explicit truncation metadata. state: "running" means the process remains live even though the individual exec_command Tool call has returned; terminal reports carry the exit, timeout, stop, signal, or cleanup outcome.

Hosts that persist command activity can provide onTerminalUpdate. It receives each spawned command's final outcome with its original Session, Run, Tool call, and command identities, including cancellation before the first yield. Its output is the complete bounded stdout/stderr, independent of poll deltas. This notification can precede Tool completion and requires no additional model Turn. Session close waits for outstanding observers and reports their failures.

Such hosts can also provide readTerminal({ sessionId, commandId }). Once a live record has been released, command_session reads its retained receipt through this port. The host must scope lookup to the exact owning Session and return undefined for unknown or foreign handles. A retained receipt contains the complete bounded final output, not a new delta; repeated reads return the same result. Reading a stopped or completed handle never launches or stops another process. Desktop supplies this port from its existing command ledger, so a later Run can verify a command that finished between Runs or before restart.

Hosts that show running services to their user read them from the same owner. createSandboxRuntimeCommandSessionPlugin(...).sessionCommands(sessionId) lists the Session's lifetime: "session" commands whose process has not exited: the command text, handle, owning Run and Tool call, launch time and the bounded output since launch (not a poll delta, and reading it consumes nothing). stopSessionCommand(sessionId, commandId) stops one exactly like the command_session stop action, publishes its stopped terminal through onTerminalUpdate, and resolves after that; a command that already ended needs no stop. The onSessionCommandsChanged({ sessionId }) command option is called when such a command starts and after it ends, so the host can refresh its list.

Before each agent model request, the same command plugin supplies current observations of yielded handles: state, exit outcome, code, signal and deadline. Late terminals supersede initial running observations without rewriting Tool history or requiring a poll. This context contains no command text or output and does not authorize another launch. A running process is not proof of service health. Observations survive compaction in the live Runner Session, are excluded from compaction requests, and are removed on Session close. On Session start, the plugin reconstructs unresolved handles from the validated original transcript, including exchanges omitted from the compacted model context. It resolves them through readTerminal; without live ownership or a receipt they are explicitly unavailable. This restores observations only, never processes or command effects.

The supervisor is POSIX-only and non-PTY. It owns process groups in memory per Runner Session, applies the ordinary command execution deadline unless the invocation explicitly requests Session lifetime, bounds stdout and stderr independently per report, and closes every owned process when the Runner Session closes. A process group still live when the host process exits, for example because the host exited before its Runner finished closing, is killed on the host's exit event; only a host killed outright (SIGKILL, a crash) leaves its commands behind. Handles cannot cross Runner Sessions and are never reattached after process restart. Reading a persisted terminal receipt does not reattach a process. The working directory is not a Sandbox.

Every command Module also requires one environment policy. inherit is "all", "core", or "none"; "core" keeps the host's shell, path, home, temporary-directory, locale, and user variables. Case-insensitive * / ? patterns in exclude run after the built-in *KEY*, *SECRET*, and *TOKEN* filter. Explicit set values run next, then includeOnly narrows the result. Castle/provider identity variables remain unavailable even through set. The resolved environment is frozen when the Module is created; Session history never stores or restores it.

The initial exec_command call should pass the host's effect-approval policy. Polling or stopping its already-admitted handle can be allowed by a separate stable policy key; it must not manufacture a second approval for the same process authority.

Sandboxed Command Sessions

createSandboxRuntimeCommandSessionPlugin contributes the same exec_command / command_session lifecycle plus the SandboxResolver that binds them to a real OS process boundary. Declare its exact specifier as sandbox on every Agent that receives exec_command:

const localCommands = createSandboxRuntimeCommandSessionPlugin({
  sandbox: {
    specifier: "local/workspace",
    config: {
      network: { allowedDomains: [], deniedDomains: [] },
      filesystem: {
        denyRead: ["~/.ssh"],
        allowWrite: [workspacePath],
        denyWrite: [],
      },
    },
  },
  command: {
    workingDirectory: pathToFileURL(`${workspacePath}/`),
    shell: { executable: "/bin/sh", arguments: ["-c"] },
    environment: { inherit: "core" },
    executionTimeoutMs: 120_000,
    initialYieldMs: 10_000,
    pollWaitMs: 1_000,
    maxOutputBytesPerStream: 64 * 1024,
  },
});

The Node Adapter uses @anthropic-ai/sandbox-runtime: Seatbelt on macOS and bubblewrap/seccomp on Linux. Filesystem and network policy are explicit input; Castle does not infer policy from the working directory or command approval. The final rewritten and approved command is transformed into sandboxed argv/env immediately before the existing command supervisor spawns it. Each invocation uses its command handle as the Sandbox Runtime attribution key, so filesystem, seccomp, and proxy denials are appended to that command's stderr without a Castle-owned regex classifier.

When filesystem confinement is enabled, the Adapter first creates the temporary directory selected by Sandbox Runtime inside the same sandboxed process. The requested command runs only if that succeeds. This does not grant a write path or depend on a directory left by another application. Filesystem-disabled mode keeps its existing environment behavior; command attribution retains the original approved command text.

runtimeAdapter is the narrow host integration seam for initialize / allowNetworkHosts / prepare / finalize / reset. Omitting it uses the pinned Sandbox Runtime in the current Node process. An injected Adapter may perform these operations in a workspace-owned local Worker, but it returns only the transformed argv: the existing command supervisor still uniquely owns spawn, process groups, output, timeout, abort, and stop. Finalization is awaited after process-group cleanup and returns that command's annotated stderr before the Tool result settles.

One plugin profile can serve multiple concurrent Runner Sessions. Each Session has its own frozen binding and command handles; the final lease resets the profile. The upstream manager is process-global, so a Node process may have one active profile at a time. Another profile fails before initialization instead of replacing policy. Windows command Session process ownership, automatic unsandboxed escalation, containers, remote execution, and restart adoption are not part of this Module.

The plugin also contributes request_network_access; declare it with hostTool("request_network_access") when the Agent should request additional domains. Its validated input contains 1–16 exact canonical hosts and a non-empty reason. The host must apply its Tool permission policy before execution. Command approval alone does not grant network access.

A grant adds those hosts to the current Runner's Sandbox allowlist. Existing denied domains and filesystem policy remain authoritative. It does not execute or retry the command. All Sessions sharing this Runner share the grant, including Sessions opened after the last lease resets; a fresh Runner starts with the original profile, even when reusing the Plugin object. Desktop maps this lifetime to the current workspace until the app closes. Grants are not stored in history or replayed as effects.

Browser QA

createBrowserSessionPlugin contributes browser_navigate, browser_snapshot, browser_switch_tab, browser_click, browser_hover, browser_drag, browser_select_option, browser_fill, browser_press_key, and browser_capture. The Agent must declare each Tool it needs. The host supplies an Origin authority, a BrowserSessionFactory, a text observation byte limit, and an image byte limit. The authority admits and canonicalizes each requested URL at the host trust boundary; product permission consumes that validated admission through the Harness policy hook instead of parsing the URL again.

The plugin opens a Browser only after browser_navigate admits an Origin. That Browser context is reused for the Session. Every navigation admits the requested Origin before the Adapter receives it. newTab: true preserves the current page; browser_switch_tab accepts an ID from the observation's tabs list and reactivates that page without navigation. Tab identities expire on restart. Capture and other page actions fail explicitly before the first navigate. The Browser closes through the normal sessionEnd lifecycle. Browser actions are exclusive ordinary Tools: final input still passes permission and approval before the Adapter sees it. Saved transcript history never restores or adopts an old Browser or page. browser_capture accepts only integer width/height, light | dark colorScheme, and reduce | no-preference reducedMotion; it returns one bounded text observation followed by one PNG block under the same Tool call identity.

Each observation from Castle's Playwright and connected-Chrome Adapters includes the measured layout viewport (window.innerWidth / innerHeight) in CSS pixels, including scrollbars. This is fresh page state, not the last requested capture size or the PNG's device-pixel dimensions. It remains visible in ordinary key, click and snapshot results across Runs in the same live Session. External Adapters may omit the optional viewport; the text then explicitly says it was not reported. Old recorded observations are not retroactively assigned a size.

Text observations return at most the configured byte limit or 8 KiB, whichever is smaller. A long result ends with browser_snapshot continuation arguments (observationId, offset). Pass both unchanged to read the next slice of the same immutable observation without refreshing target refs. Calling snapshot with no arguments observes the page again. Any new browser action invalidates the previous continuation, including an action that fails; the page may have changed. Continuations are Session-local and ephemeral, never recovered from transcript history. To find relevant content without paging unrelated text, pass query alone: a case-insensitive literal search returns the first matching line and following content from that same observation, with unchanged refs. It does not refresh the browser; no match is a successful read, not a Tool failure. Screenshot text contains metadata rather than the ARIA tree, so refresh the snapshot to search page text after a capture.

Accessible targets accept optional within: { role, name?, exact? } to limit matching to an observed ancestor, for example a date's named region or an unnamed main. Click, fill, select, and both drag endpoints use this same contract. Use the current snapshot's roles and names; the resulting target must match exactly one element. Missing scopes never fall back to the whole page, and ambiguous targets never select the first match.

browser_hover moves the pointer over one observed visible target without clicking. Hover the visible row/title to reveal hover-only controls, then use the returned observation for the next action. It uses the same role/name/ref and optional ancestor contract as click, without bypassing hit testing.

browser_drag takes source and target, each with the same accessible role, name, and optional exact fields used by click. It runs in the current Browser Session and returns the updated observation. Both targets must resolve uniquely; a missing or ambiguous target is an action failure.

browser_select_option selects a native select's option by its visible label. The control uses the same accessible role, name, and optional exact fields. Selection dispatches input/change and returns the updated observation; custom menus use browser_click. This shares the existing Browser Session, Origin authority, and permission lifecycle.

For an Agent declaration, add the browser capability from the optional @castle-ai/connectors/browser entry to capabilities (install playwright-core alongside). It declares these same ten tools and a browser instructions section, owns a lazy browser per Session, and supports managed browsing or an explicitly host-selected Chrome tab. See that package's README for defaults and examples; Node Runner alone does not install a browser dependency. Custom adapters can use createBrowserSessionTools(options) directly; its tools, closeSession(id) and close() share the plugin's implementation. Awaiting close() blocks new calls, releases browsers and settles all admitted operations, including screenshot writes.

The Playwright Adapter uses accessible role/name locators and returns URL, title, and a bounded ARIA snapshot after every action. For capture it applies the requested viewport and media preferences, disables screenshot animation, and captures the visible viewport as PNG. The Adapter limits top-level navigation and redirects to the Session's admitted Origins. Its default resourceAccess: "same-origin" keeps subresources within those Origins; hosts can explicitly select "http" for websites using external HTTP(S) resources and WS(S) connections. Service Workers remain disabled. See the Adapter README for the resource policy and limitations; this is not a general network Sandbox.

Local stdio MCP Tools

Declare a server with the mcp capability. It adds no instructions: its Tools already carry model-visible descriptions.

import { defineAgent } from "@castle-ai/harness";
import { mcp, openSession } from "@castle-ai/node-runner";

const Agent = defineAgent("example.mcp", {
  model,
  instructions: "Use the project tools to complete the request.",
  capabilities: [mcp({
    transport: "stdio",
    command: "project-mcp-server",
    args: ["--stdio"],
    cwd: new URL("file:///absolute/project/"),
    serverName: "project",
  })],
});
const session = await openSession({ agent: Agent, approval });
try { console.log((await session.run("Read the current issue.").result).text); }
finally { await session.close(); }

Declaring the Agent does not spawn the server. Opening each Session starts one process, discovers its full catalog, and uses that same connection for all calls until close. No separate plugin or hostTool reference is needed. enabled (a boolean, or a function of the turn context) hides the whole server for a turn; a disabled catalog is still acquired and validated but is absent from model requests and cannot be di