npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@modelprofile.com/flexharness

v12.1.0

Published

Managed agent sessions with tools, MCP, chat, OCR, documents and browser control through focused subpath imports.

Readme

FlexHarness is a modular toolbox for model inference, agents, managed sessions, chat, OCR, KVM control and provider capabilities. It brings SmartAI, SmartAgent, SmartChat, SmartKVM and SmartOCR's AI helpers into one repository through four focused packages published with tspublish.

Issue Reporting and Security

For reporting bugs, issues, or security vulnerabilities, please visit community.foss.global/. This is the central community hub for all issue reporting. Developers who sign and comply with our contribution agreement and go through identification can also get a code.foss.global/ account to submit Pull Requests directly.

Install and choose an entrypoint

The toolbox publishes four packages under @modelprofile.com, with one shared release version. Integrations use subpath imports instead of additional package names.

| Package | Use it for | | --- | --- | | flexharness | Managed sessions, permissions, tools, chat, OCR, documents and KVM | | flexharness-agent | Standalone runAgent and AgentSession; portable /runner | | flexharness-models | Model contracts, ModelRegistry, AI SDK helpers and caching | | flexharness-providers | All eight model providers, authentication, accounts, media and research |

For a managed agent with model providers:

pnpm add @modelprofile.com/flexharness @modelprofile.com/flexharness-providers

For inference without an agent:

pnpm add @modelprofile.com/flexharness-models @modelprofile.com/flexharness-providers
import { ModelRegistry, generateText } from '@modelprofile.com/flexharness-models';
import { createOpenAiModelProvider } from '@modelprofile.com/flexharness-providers/openai';

const models = new ModelRegistry().register(createOpenAiModelProvider());
const setup = models.getModelSetup({
  provider: 'openai', model: 'gpt-5.5', apiKey: process.env.OPENAI_API_KEY,
});
const result = await generateText({ ...setup, prompt: 'Hello' });

For a standalone agent, install @modelprofile.com/flexharness-agent and pass the same setup to runAgent({ ...setup, prompt: 'Hello' }). Node.js 24 or newer is required for Node entrypoints. Each application owns its model registry; imports do not register providers globally.

Harness subpaths

All entries below belong to @modelprofile.com/flexharness. The root export stays focused on managed sessions. Importing it does not load UI, PDF, browser automation or provider SDKs.

| Import suffix | Capability | Additional install | | --- | --- | --- | | /tools | Host-supplied execution contexts, HTTP and JSON tools | None | | /tools/node | Local filesystem and process tools, file-backed job adapters | None | | /browser | Run-scoped SmartBrowser framed-client adapter for Flex tool providers | @push.rocks/smartbrowser | | /compaction | Model-driven conversation compaction | None | | /media | Vision recipes using an injected model | None | | /ocr | ImageOcr, OCR contracts and Mistral transport | None | | /chat | Portable streaming ChatSession, history and usage | None | | /chat/cli | Ink/React terminal chat | ink ink-text-input react; TypeScript: @types/react | | /chat/web | Lit flexchat-window, flexchat-message, flexchat-input | lit | | /chat/harness | FlexHarnessTranscript and intents for the @design.estate/dees-catalog harness components, browser-safe | @design.estate/dees-catalog ^17.0.0 \|\| ^19.5.0 \|\| ^20.0.0 \|\| ^21.0.0 (types only) | | /kvm | Browser control, terminal framing and OCR observation | puppeteer | | /documents | PDF and image documents, extractTextFromPdf | @push.rocks/smartpdf; for image documents sharp | | /mcp | MCP clients, AI SDK tool conversion and createMcpToolCatalog, the tool catalog of MCP tools | @modelcontextprotocol/sdk | | /migration | Existing versioned harness migrations | None | | /stores/nosqldb | NoSqlFlexHarnessStores: cross-process CAS stores on @lossless.org/client/nosqldb | @lossless.org/client | | /stores/testing | createFlexStoreContractTests: the contract every built-in IFlexHarnessStores passes, to run against a custom store with any test runner | None |

The integrations in the last column use optional peer dependencies. Install them only for the entrypoints your project imports, for example:

pnpm add @modelprofile.com/flexharness lit

For a managed SmartBrowser resource, install its optional browser peer:

pnpm add @modelprofile.com/flexharness @push.rocks/smartbrowser
import { ChatSession } from '@modelprofile.com/flexharness/chat';
import '@modelprofile.com/flexharness/chat/web';

Chat uses the portable flexharness-agent/runner. A browser application can provide a model backed by its own server transport, or supply run in IChatSessionOptions. The web element accepts an IChatSession and maintains its own subscription. Full Node AgentSession remains at the agent package root.

Durable session token metrics

Read await harness.getSessionMetrics(scopeId, sessionId) for bounded, durable accounting independent of transcript visibility. It reads fixed-size projection state rather than summing assistant messages or scanning event/archive/child history. A warm read joins outstanding projection writes and surfaces accounting failures.

| Field | Meaning | | --- | --- | | reportedLifetimeUsedTokens | Sum of durably settled per-call reports. Incomplete accounting makes this a lower bound. | | reportedLifetimeUsedTokensExact | Whether the sum of those reports is exact; invalid reports and safe-integer overflow make it inexact. | | lifetimeUsageComplete | Every owned call is accounted for and no accounting operation remains in flight. | | lifetimeUsedTokens | Present only when the lifetime total is complete. Zero is a valid total. | | inFlight | A generation or actual compactor invocation has not settled its accounting yet. | | lastObservedInputTokens, lastObservedInputTokensAt | Input count and observation time of the last settled generation call with complete input/output totals. | | lastObservedInputTokensFreshness | Always historical. This is not current context size: output, steering, compaction and reversion can change context. | | reportedLifetimeInputTokens | Sum of the settled per-call reports' input tokens, prompt-cache reads included. | | reportedLifetimeCacheReadTokens, reportedLifetimeCacheWriteTokens | Sums of the settled per-call reports' input tokens read from and written to the prompt cache. | | lifetimeTokenBreakdownComplete | The three sums cover every call lifetimeUsedTokens covers: lifetime usage is complete and the session's calls were counted by input and cache since it began. Otherwise they are lower bounds. |

Accounting covers this session's generation calls and reported compaction calls, including failure, cancellation and retries. Child sessions own their own spend; tool-owned model calls are not included. Do not sum cumulative assistant-message usage to reconstruct lifetime spend. Missing/partial reports remain incomplete; unsafe totals saturate at Number.MAX_SAFE_INTEGER without claiming exactness. A compactor that reports no calls cannot prove zero spend.

New sessions start complete at zero. Existing/imported sessions with unknown prior spend start incomplete. Write-ahead operation markers make interrupted recovery conservative: unresolved operations are cleared without certifying their unknown spend. Undo, redo, compaction and transcript archiving never subtract lifetime spend. Importing visible history does not reconstruct historical billing.

Sessions whose metrics were written before input and cache tokens were counted count them from their next settled call on and never report the breakdown complete; their lifetime total keeps its completeness.

Conversational runs that exhaust maxSteps with tool work or unseen input remaining fail with FlexHarnessStepLimitError (FLEX_STEP_LIMIT), rather than being marked completed with an empty answer. Standalone Agent callers retain tool-only results and can inspect terminationReason: 'model' | 'step-limit'.

Provider subpaths

Install @modelprofile.com/flexharness-providers once. Its root exports all provider factories; use a provider subpath to import just that implementation.

| Import suffix | Capability | | --- | --- | | /anthropic, /openai, /google, /groq, /mistral, /xai, /perplexity, /ollama | Model adapters and their request types | | /auth | ChatGPT browser/device authentication, refresh and model connections | | /accounts | Account lifecycle, model catalogs, rate limits and credential envelopes | | /auth/files | Explicit interoperability with external credential files | | /media | OpenAI audio and image generation | | /research | Anthropic research utilities |

Provider SDKs are installed together. Authentication, file access and vendor media are separate entrypoints and are not imported by the provider factory root. For ChatGPT authentication, pass connection: createOpenAiChatGptModelConnection(credentials) from /auth to the OpenAI model options. API-key inference needs no authentication setup.

Migration

Replace the withdrawn 6.x component imports with the subpaths above. Provider imports change from individual provider packages to @modelprofile.com/flexharness-providers/<provider>. Keep direct agent and models imports and upgrade them to the same release version. The removed package names are not republished.

| Previous API | Replacement | | --- | --- | | SmartAI models, caching and AI SDK contracts | flexharness-models and explicit ModelRegistry registrations | | SmartAI provider factories | flexharness-providers/<provider> | | SmartAI authentication, account adapters and external credential files | flexharness-providers/auth, /accounts, /auth/files | | SmartAI vision, document and OCR helpers | flexharness/media, /documents, /ocr | | SmartAI audio, image and research helpers | flexharness-providers/media, /research | | SmartAgent runtime, events and persistence contracts | flexharness-agent | | SmartAgent tools, local tools, compaction and MCP | flexharness/tools, /tools/node, /compaction, /mcp | | SmartChat sessions, CLI and web components | flexharness/chat, /chat/cli, /chat/web | | SmartKVM | flexharness/kvm: BrowserKvm, KvmTerminal, createKvmTools | | SmartOCR image AI helper | ImageOcr.recognizeImageBytes from flexharness/ocr | | SmartOCR PDF AI helper | extractTextFromPdf from flexharness/documents |

For image OCR, rename smartAiOcrEngine to engine and mistralOcrOptions to mistralOptions; pass the API key explicitly. Native searchable-PDF processing through SmartOcr.processPdfBuffer remains in SmartOCR.

Existing ISmartAi* and TSmartAi* contracts, durable event/job schemas, harness projections, credential envelopes and external credential paths remain unchanged. The OpenAI account SmartAiProviderRegistry is separate from inference's ModelRegistry. The package consolidation requires no data migration.

Developing and releasing

Four named tspublish.json descriptors own the published packages. Other folder descriptors specify compilation order only. Each package explicitly declares its owned folders, exports and required or optional peer dependencies. Source and compiled declaration mappings support development without workspace links.

Run pnpm build, pnpm test and pnpm run check:test. Tests install packed artifacts in disposable directories to check minimal dependency graphs, every public subpath and browser imports. Live provider and browser authentication tests remain opt-in in test_integration/.

The configured GitZone release prepares four packages, records their exact artifacts, and publishes models before its dependents to npmjs and Verdaccio. Third-party OpenAI Codex notices accompany the providers package.

Core Setup

import {
  FlexHarness,
  JsonFileFlexHarnessStores,
  type IFlexResolvedModel,
  type TFlexAgentToolSet,
} from '@modelprofile.com/flexharness';

interface IProjectScope {
  projectRoot: string;
}

const stores = new JsonFileFlexHarnessStores({
  directory: '/var/lib/my-app/model-sessions',
});

const harness = new FlexHarness<IProjectScope>({
  scopeResolver: {
    async resolveScope(scopeId) {
      const project = await projectRegistry.get(scopeId);
      return {
        // Aliases that resolve to this same key share sessions and save ordering.
        storageKey: project.accountAndProjectKey,
        scope: { projectRoot: project.root },
      };
    },
  },
  modelResolver: {
    async resolveModel({ scope, modelHint, signal }): Promise<IFlexResolvedModel> {
      const configured = await modelRegistry.resolve({ scope, modelHint, signal });
      return {
        model: configured.model,
        identity: {
          provider: configured.providerId,
          model: configured.modelId,
          displayName: configured.label,
        },
        providerOptions: configured.providerOptions,
      };
    },
  },
  toolProvider: {
    async provideTools(context) {
      const tools: TFlexAgentToolSet = await createProjectTools({
        root: context.scope.projectRoot,
        signal: context.signal,
        requestPermission: context.requestPermission,
      });
      return {
        tools,
        close: async () => closeProjectTools(tools),
      };
    },
  },
  stores,
  builtInTools: {
    renameSession: true,
    projectManagement: {
      // task, goal, and scratchpad default to true when this block exists.
    },
  },
  toolOutputLimits: {
    maxDepth: 12,
    maxBytes: 256 * 1024,
  },
  callbackLimits: {
    maxEvents: 10_000,
    maxOutputBytes: 1024 * 1024,
    maxParts: 2_000,
  },
  promptQueueLimits: {
    maxOutstandingPromptsPerSession: 16,
    maxOutstandingBytesPerSession: 64 * 1024 * 1024,
    maxPendingAdmissions: 64,
    maxPendingAdmissionBytes: 128 * 1024 * 1024,
    maxTerminalEntriesPerSession: 64,
  },
  reversionLimits: {
    maxCompletedTurns: 100,
    maxSegments: 300,
    maxExcludedRunIds: 1000,
    maxPendingReversionReleases: 1000,
  },
  sessionImportLimits: {
    maxOpenImports: 2,
    maxMessagesPerSession: 2048,
    maxSessionBytes: 64 * 1024 * 1024,
    maxPageBytes: 1024 * 1024,
    stagedImportIdleTimeoutMs: 5 * 60 * 1000,
  },
  resultLimits: {
    maxStoredBytes: 8 * 1024 * 1024,
    maxResultsPerSession: 32,
  },
  subagents: [
    {
      name: 'researcher',
      description: 'Research a focused question and return one final answer.',
      modelHint: 'reasoning-model',
      system: 'Investigate the assigned question. Return a concise evidence-based answer.',
      maxSteps: 8,
    },
  ],
  maxSubagentDepth: 6,
  maxSubagentCallsPerRun: 32,
  externalErrorProjector: (_error, context) => ({
    name: 'ModelOperationError',
    message: `The ${context.source} operation failed.`,
    code: 'MODEL_OPERATION_FAILED',
  }),
});

modelRegistry, projectRegistry, createProjectTools, and closeProjectTools in this example are application-owned integrations. FlexHarness passes the same run AbortSignal to the model resolver and tool provider. Every model-resolver, application tool-provider, and resource tool-provider context also carries the required canonical sessionGenerationId and sessionGenerationSequence, allowing host operations to authorize the exact session generation rather than a reusable session ID alone.

Resource Tool Providers

resourceToolProviderResolver composes zero or more resource-owned providers with the existing application toolProvider. The resolver runs fresh for every prompt and returns the current resource attachment descriptors:

resourceToolProviderResolver: {
  async resolveResourceToolProviders({
    scope,
    sessionId,
    sessionGenerationId,
    sessionGenerationSequence,
    runId,
    signal,
  }) {
    const attachments = await resourceRegistry.listAttached({
      scope,
      sessionId,
      sessionGenerationId,
      sessionGenerationSequence,
      runId,
      signal,
    });
    return attachments.map((attachment) => ({
      resourceId: attachment.resourceId,
      attachmentRevision: attachment.attachmentRevision,
      provider: createResourceToolProvider(attachment),
    }));
  },
},

Each descriptor uses the existing IFlexToolProvider<TScope> contract. Its provider receives the normal run context and must return a fresh run-scoped handle. The resource resolver context carries the same required canonical session generation as the tool-provider contexts. The original toolProvider remains optional and its tool names remain unchanged. Resource tool names are deterministic and bounded:

  1. resourceIdentity is the lowercase hexadecimal SHA-256 of JSON.stringify([resourceId, attachmentRevision]).
  2. The namespace is resource_ plus the first 16 digest characters.
  3. The exposed name is <namespace>__<stem>__<toolDigest>. stem replaces characters outside [A-Za-z0-9_-] with _, keeps the first 16 characters, and falls back to tool; toolDigest is the first 12 lowercase hexadecimal characters of SHA-256 over the original tool name.

The resolver accepts at most 128 descriptors per run. resourceId must be non-empty and at most 512 UTF-8 bytes, attachmentRevision must be a non-negative safe integer, and each original resource tool name must be non-empty and at most 512 UTF-8 bytes. FlexHarness rejects duplicate resourceId values even across revisions, duplicate derived namespaces, and duplicate final exposed tool names before model execution. Descriptor identity and namespace validation completes before any application or resource provider is acquired.

Resource permission requests are scoped with the complete 64-character resourceIdentity, not the shortened tool namespace. FlexHarness rewrites kind to resource.<resourceIdentity>.<providerKind> and an optional rememberKey to resource:<resourceIdentity>:<providerRememberKey>. Harness-owned metadata contains resourceId, attachmentRevision, resourceIdentity, and toolNamespace; provider metadata is nested under providerMetadata, so it cannot override attachment identity.

FlexHarness owns every acquired handle. Normal close and partial-failure cleanup run in reverse acquisition order, attempt every handle, aggregate multiple failures, and retain failed cleanup for retirement or disposal retry. Cancellation uses the same path. If model resolution fails while a resource provider is still settling, a late returned handle remains tracked and disposal waits for its closure. Application and resource providers may not define a harness built-in name while that built-in is enabled for the current run. Disabled names are not reserved.

Project Management Tools

Harness-owned project tools are opt-in and session-local:

builtInTools: {
  renameSession: true,
  projectManagement: {
    task: true,
    goal: true,
    scratchpad: true,
  },
},

renameSession enables rename_session. The projectManagement block enables the public project-management APIs and contains the model-tool flags; task, goal, and scratchpad each default to enabled unless explicitly set to false. Without that block, the public project-management APIs reject with FlexHarnessValidationError, while the required stores.projectManagement domain still participates in session cleanup. With no builtInTools configuration, none of these four tools is present. Constructor options are copied and frozen.

Project-management records use FLEX_PROJECT_MANAGEMENT_SCHEMA_VERSION, currently 1, and form a strict live-or-tombstone union:

interface IFlexProjectManagementSnapshot {
  schemaVersion: 1;
  revision: number;
  sessionGenerationId: string;
  sessionGenerationSequence: number;
  goal?: string;
  scratchpad: string;
  tasks: Array<{
    id: string;
    content: string;
    status: 'pending' | 'in_progress' | 'completed' | 'cancelled';
    priority: 'high' | 'medium' | 'low';
    createdAt: string;
    updatedAt: string;
  }>;
}

interface IFlexProjectManagementTombstone {
  schemaVersion: 1;
  revision: number;
  sessionGenerationId: string;
  sessionGenerationSequence: number;
  deletedAt: string;
}

type TFlexProjectManagementRecord =
  | IFlexProjectManagementSnapshot
  | IFlexProjectManagementTombstone;

The tools use strict action-discriminated inputs:

  • task: list, create, update, delete, or clear. Create defaults to pending and medium.
  • goal: get, set, or clear.
  • scratchpad: get, set, append, or clear. Append concatenates the supplied content exactly.
  • rename_session: sets the active session title and returns the authoritative session.

Every project action returns the authoritative revision and state; task mutations also return the affected task, and clear returns the removed tasks. Reads never save. A set, clear, append, update, idempotent create, or empty task clear that makes no state change returns the current revision without writing. Mutations load once, apply once, validate the complete next snapshot, and issue one compare-and-swap save at revision + 1. FlexHarness never retries or merges an external conflict.

Tool task creation accepts an optional id. When omitted, FlexHarness requires the stable AgentSession toolCallId and derives task_ plus the SHA-256 of JSON.stringify(['flexharness-project-task-v1', storageKey, sessionId, runId, toolCallId]). Repeating an explicit or deterministic ID with identical content, status, and priority is idempotent; different creation data conflicts. Application callers must supply an explicit id to createProjectTask() because no tool-call identity exists at that boundary.

The same engine is available to applications:

await harness.getProjectState(scopeId, sessionId);
await harness.listProjectTasks(scopeId, sessionId);
await harness.createProjectTask(scopeId, sessionId, { id, content, status, priority });
await harness.updateProjectTask(scopeId, sessionId, { id, content, status, priority });
await harness.deleteProjectTask(scopeId, sessionId, id);
await harness.clearProjectTasks(scopeId, sessionId);
await harness.getProjectGoal(scopeId, sessionId);
await harness.setProjectGoal(scopeId, sessionId, goal);
await harness.clearProjectGoal(scopeId, sessionId);
await harness.getProjectScratchpad(scopeId, sessionId);
await harness.setProjectScratchpad(scopeId, sessionId, content);
await harness.appendProjectScratchpad(scopeId, sessionId, content);
await harness.clearProjectScratchpad(scopeId, sessionId);

Public writes use { actor: 'application' }. Tool writes use { actor: 'agent', runId, toolCallId, agent? }, allowing custom stores to preserve attribution. Project side effects commit independently of the later model outcome and are intentionally outside transcript undo/redo.

FLEX_PROJECT_MANAGEMENT_LIMITS exports the hard UTF-8 and aggregate limits: goal 8 KiB, scratchpad 128 KiB, task content 8 KiB, task ID 512 bytes, title 2048 bytes, 512 tasks, and a 1 MiB serialized snapshot. The aggregate bound leaves room for worst-case JSON escaping of a controller-valid scratchpad. Loaded snapshots reject extra fields, duplicate IDs, invalid status/priority/timestamps, non-JSON data, wrong schema/revision, and every exceeded bound before use.

IFlexProjectManagementStore is exact per (storageKey, sessionId): load, CAS save, CAS tombstoneSession, and purgeNamespace must not collapse multiple sessions or storage namespaces. load() returns TFlexProjectManagementRecord | undefined and receives an optional IFlexProjectManagementSessionContext as its third argument; tombstoneSession() receives the same optional context as its fifth argument. FlexHarness always supplies both, while existing two-argument loads, four-argument tombstones, and shorter store implementations remain compatible.

IFlexProjectManagementSessionContext contains sessionGenerationId, sessionGenerationSequence, and optional subagent. The atomic IFlexSubagentProvenance block contains parentSessionId, parentSessionGenerationId, parentSessionGenerationSequence, originParentRunId, originParentToolCallId, agent, and actual session depth. Both the context and its separately cloned nested block are frozen.

Same-generation live saves use normal revision CAS, and a same-generation tombstone permanently rejects later live saves. A higher sessionGenerationSequence with a different sessionGenerationId may replace only an older tombstone using expected revision 0; it cannot replace a live record. This resets the PM revision for a recreated core session while stale saves and tombstones from older generations remain fenced. Deleting a recreated session that made no PM writes still replaces the prior-generation tombstone with a revision-1 tombstone for the new generation.

Every newly created core session exposes and persists a sessionGenerationId plus its monotonic sessionGenerationSequence. FlexHarness generates a strong random ID when sessionGenerationId is omitted. Applications may supply the ID to createSession() when they need to persist creation authority before dispatch; a supplied ID must be nonblank, contain no control characters, and fit within FLEX_SESSION_GENERATION_ID_MAX_BYTES (128 UTF-8 bytes). FlexHarness still assigns the sequence atomically. A legacy scope session without those fields is assigned a deterministic bounded ID derived from its immutable storageKey, sessionId, and createdAt; FlexHarness persists the repaired scope snapshot before accepting work. Grouped core deletion tombstones retain both fields after live metadata is removed.

Normal Flex session cleanup always waits in-flight local project operations, then loads and CAS-tombstones stores.projectManagement, regardless of whether PM tools are enabled in that harness. Child cleanup persists one complete IFlexSubagentProvenance block on its core tombstone before live metadata is removed. One cleanup invocation reuses the identical doubly frozen context object across bounded CAS retries; restart or a later cleanup invocation reconstructs a new frozen context from the persisted provenance. If a concurrent same-generation save wins first, cleanup reloads and retries; unresolved conflict or store failure retains the core Flex session cleanup tombstone for a later retry. The durable PM tombstone is not physically removed during normal session cleanup.

purgeNamespace(storageKey) is the explicit destructive reclamation operation and physically removes every live record and tombstone in that exact PM namespace. Applications may call it only after serializing every scope alias, preventing new admission, awaiting retireScope() on every harness owner, and deleting or purging the application-owned core scope namespace. retireScope() itself remains non-destructive and never calls purgeNamespace(). Purging PM first, purging only one alias, or racing a stale harness can remove the fence that makes session-generation reuse safe.

InMemoryFlexProjectManagementStore is the standalone in-memory implementation. InMemoryFlexHarnessStores and JsonFileFlexHarnessStores include projectManagement as a required bundle member. assertFlexProjectManagementSnapshot() validates live records, assertFlexProjectManagementTombstone() validates tombstones, and assertFlexProjectManagementRecord() validates the union. createEmptyFlexProjectManagementSnapshot(sessionGenerationId, sessionGenerationSequence) returns a revision-0 live state for the supplied current generation.

Subagents

subagents enables the harness-owned delegate, delegate_result, and delegate_send tools when at least one definition exists and the session depth is below maxSubagentDepth. A tool provider cannot replace any enabled built-in; global and agent tool allowlists still intersect to control model-facing exposure. Delegation is asynchronous by default: delegate returns { taskId, status: 'running', runId } after durable child admission, so the parent can continue independent work. delegate_result({ taskId, waitMs? }) reads live status and, once settled, bounded final text and model identity; its optional wait is limited to 240 seconds and observes parent cancellation. Set foreground: true to await the final answer in delegate itself. Children remain owned by the parent turn: Stop cancels its exact child runs, and the parent joins them before releasing authority and finalizing. A child result enters the parent's next model context as untrusted task data. When the parent's model ends its turn while background children still run, the turn waits for the next child to settle and takes its result in with another step (while steps remain), so a delegated result is never lost because the parent finished first.

The model calls it with this exact input shape:

interface IDelegateInput {
  description: string;
  prompt: string;
  subagentType: string;
  taskId?: string;
  foreground?: boolean;
}

When upgrading from 9.x, callers that require the previous synchronous delegate result must pass foreground: true. Tool providers must also leave the reserved delegate_result and delegate_send names available when subagents are enabled.

delegate_send({ taskId, prompt }) sends follow-up instructions to a running direct child owned by this exact parent turn. Both fields are required and extra fields are rejected; taskId is non-empty and at most 512 UTF-8 bytes, and prompt is non-empty and at most 64 KiB. It returns { steerId, queueId, runId } on admission, with a stable steer identity derived from the parent run and tool call. The child takes the input at its next inference boundary, without interrupting a running tool or permission wait. A child waiting to finish while its own descendants run can take the follow-up and continue inference; those descendants keep running. Use delegate_result to inspect progress and collect the final answer.

Controllers can use the same admission path explicitly:

await harness.sendSubagentInput(scopeId, parentSessionId, parentRunId, taskId, {
  steerId: 'follow-up-1',
  prompt: 'Prefer primary sources and include the publication date.',
});

Sending requires the exact active parent run, its currently delegated direct child, matching captured configuration and parent/child generations, and live writable ancestors. It does not authorize a root to skip levels, or infer ownership from immutable creation-origin metadata: a resumed child belongs to the current delegate invocation. Foreign, unowned, stale, cancelled, settled, or ending runs reject; sending never starts an idle generation, restarts a child, or resumes one. General public steerPrompt() still forbids child sessions. Child sends share root steering's duplicate rejection, pending count/byte bounds, accepted/applied events, durable user-message application, and terminal unapplied-input accounting.

Before creating or resuming a child, FlexHarness requests permission on the parent run with kind: 'subagent.start', the parent toolCallId, and separately bounded harness-owned agent/task metadata. Its metadata includes childSessionId, the exact ID reserved for this call: FlexHarness derives it deterministically when taskId is omitted and copies the supplied candidate when taskId is present. A resume also retains that candidate as taskId. This block is not truncated by toolOutputLimits. The reserved ID binds permission handling before child creation but does not prove that the child exists or is owned: applications must treat the later delegated admission context as authoritative. The controller answers through the normal permission APIs. This request has no rememberKey, so always is invalid; controllers use once or reject.

Each new invocation creates a durable child IFlexSession with immutable parentSessionId, origin parentRunId, origin parentToolCallId, agent, and depth. New public roots persist depth: 0; legacy schema-1 roots may omit it. These fields are harness-owned; public createSession() accepts sessionId, sessionGenerationId, configurationRef, and title. A child inherits its parent's configurationRef and cannot resume under a different reference. Child sessions reject direct prompt(), startPrompt(), enqueuePrompt(), and schedulePrompt() calls and run only through the delegate tool. The model and tool resolver contexts receive optional immutable parentSessionId and agent values so integrations can apply agent-specific model and tool policy. Child prompts use the definition's modelHint, system, and maxSteps.

The parent tool part receives childSessionId as soon as the child is acquired, so a host can offer an authorized live transcript. A foreground call also adds the resolved child model to its part and returns bounded JSON:

{
  taskId: 'subagent_...',
  status: 'completed',
  text: 'The child final answer, limited to 64 KiB.',
  model: { provider: '...', model: '...', displayName: '...', variant: '...' },
}

Omitting taskId creates a deterministic child for the parent session, run, and tool call. The model-visible tool description and taskId schema state this creation rule directly. Repeating that same invocation does not create another child. If the deterministic child already has messages, FlexHarness reports an uncertain prior execution and never silently reruns it. This preserves AgentSession's durable parent tool intent as crash authority; controllers use listUncertainToolExecutions() and reconcileToolExecution() for uncertain parent calls.

Supplying taskId deliberately resumes an idle, live child from a later run of the same immutable parent session and the same configured agent. The caller must use the exact ID returned by an earlier completed delegate call; an unknown ID fails with safe corrective guidance and never creates a child under the supplied label. Resume starts a new child prompt while retaining the child's original parent run and tool-call origin. A child owned by another parent or agent, a deleted child, an active child, a same-run resume, or a second acquisition of the same child within one later parent run is rejected. Parent cancellation propagates only to the exact child run started by that delegate call.

delegatedRunAdmissionProvider optionally adds an application-owned admission lease around each internally delegated child run. The public contracts are IFlexDelegatedRunAdmissionProvider<TScope>, IFlexDelegatedRunAdmissionContext<TScope>, and IFlexDelegatedRunAdmissionLease. The provider is never called for root prompts or direct public prompt APIs. It receives a frozen context containing the resolved scopeId, exact captured scope, and storageKey; the exact child sessionId, session generation, queue, and run; the exact current parent session generation, queue, run, and delegate toolCallId; immutable originParentRunId and originParentToolCallId; the child agent and depth; and the child run AbortSignal. For taskId resume, the current parent queue/run/tool-call fields identify this delegate invocation, while the origin fields remain fixed to the invocation that created the durable child.

delegatedRunAdmissionProvider: {
  async acquireDelegatedRunAdmission(context) {
    const admission = await controller.acquireDelegatedRun({
      child: {
        sessionId: context.sessionId,
        sessionGenerationId: context.sessionGenerationId,
        sessionGenerationSequence: context.sessionGenerationSequence,
        queueId: context.queueId,
        runId: context.runId,
      },
      parent: {
        sessionId: context.parentSessionId,
        sessionGenerationId: context.parentSessionGenerationId,
        sessionGenerationSequence: context.parentSessionGenerationSequence,
        queueId: context.parentQueueId,
        runId: context.parentRunId,
        toolCallId: context.parentToolCallId,
        originRunId: context.originParentRunId,
        originToolCallId: context.originParentToolCallId,
      },
      signal: context.signal,
    });
    return {
      close: () => admission.close(),
    };
  },
},

Acquisition completes before generation-side branch reversion, context compaction, model resolution, application or resource tool-provider callbacks, and model execution. Child session store and runtime initialization may already have occurred before acquisition. Providers must honor the supplied AbortSignal; an abort can settle the active child and parent without waiting for an acquisition that ignores cancellation, while FlexHarness retains ownership and closes any lease returned later.

For a normally acquired lease, FlexHarness gives close() an awaited attempt after AgentSession generation and before canonical accepted, rejected, or interrupted finalization and the terminal prompt.finished event. If an abort detaches an acquisition that ignores its signal, the run may settle before acquisition returns; FlexHarness retains that owner and closes any late lease. close() must be idempotent and safe to retry after rejection. A close failure prevents successful child acceptance and remains owned by the exact child session generation for retry by later exact-session deletion, scope retirement, or disposal. Retirement and disposal truthfully wait for late acquisition and lease cleanup.

Limits are validated and frozen at construction: at most 32 unique definitions; names are non-empty and at most 128 UTF-8 bytes; descriptions 2048 bytes; optional model hints 512 bytes; optional system prompts 64 KiB; and optional maxSteps a positive safe integer. maxSubagentDepth defaults to 6 and must be a positive safe integer at most 8. The root is depth 0, so the default permits six child levels (depths 1 through 6); delegation tools are absent at the configured maximum. maxSubagentCallsPerRun defaults to 32 and must be a positive safe integer at most 128. A call slot is consumed synchronously at the start of every schema-valid delegate execution, before semantic bounds, subagent type/depth validation, permission, or child work. Inputs rejected by the tool schema never start delegate execution and do not consume a slot. After successful semantic validation, the child ID candidate is reserved for the rest of the parent run, including after permission rejection or later failure. Permission rejection creates no child session. Omitting taskId reserves a deterministic new child ID; supplying taskId reserves that unverified candidate and attempts resume after permission only if it identifies a resumable child. Delegate descriptions are non-empty and at most 256 UTF-8 bytes, prompts non-empty and at most 64 KiB, subagent types at most 128 bytes, and task IDs at most 512 bytes.

Scope snapshots remain schema 1 and legacy child tombstones without subagent provenance continue to load. New child tombstones write the atomic provenance block and reject partial, parent-generation-mismatched, depth-mismatched, or duplicate-origin records. FlexHarness versions before this provenance addition reject that new optional key under their strict reader, so downgrading or mixing old readers with newly written scope snapshots is unsupported.

Sessions And Prompts

const session = await harness.createSession('project:billing', {
  title: 'Invoice import',
});

const result = await harness.prompt(
  'project:billing',
  session.sessionId,
  [
    { type: 'text', text: 'Extract the invoice totals.' },
    {
      type: 'file',
      data: invoicePdfBase64,
      mediaType: 'application/pdf',
      name: 'invoice.pdf',
    },
  ],
  { modelHint: 'document-model', maxSteps: 12 },
);

console.log(result.assistantMessage.parts);
console.log(result.usage);

TFlexPrompt is deliberately JSON-safe. It accepts a string or an ordered array of:

  • { type: 'text', text }
  • { type: 'image', data, mediaType?, name? }
  • { type: 'file', data, mediaType, name? }

Attachment data is a string containing base64, a data URL, or a remote URL. Public input never requires Buffer or URL objects. Remote URL strings are converted only at the private AgentSession invocation boundary.

Attachment payloads are never copied into public audit messages or events. Public attachment parts contain metadata only:

{
  type: 'attachment',
  partId: '...',
  attachmentType: 'file',
  source: 'inline-base64', // or data-url / remote-url
  sizeBytes: 48231,        // omitted when it cannot be determined
  mediaType: 'application/pdf',
  name: 'invoice.pdf',
}

The original string remains only in canonical private Agent events, so a later model turn can receive the attachment again. A turn that succeeded stays in future context, and so does a turn stopped through abort() or cancelPrompt() once its model produced output: its prompt, the steers it took in and what the model produced before the stop, with every tool call that has no recorded result closed by an error result saying its effect is unknown. A turn stopped before any model output leaves no trace, so its prompt can be sent again without appearing twice; its run.finished event and finished queue entry carry keptInContext: false. Failed, resolver-failed, cleanup-failed, and persistence-failed turns, and turns a process restart interrupted, do not add anything to future context.

The main session methods are:

await harness.listSessions(scopeId);
await harness.createSession(scopeId, { sessionId, sessionGenerationId, title });
await harness.getSession(scopeId, sessionId);
await harness.getMessages(scopeId, sessionId);
await harness.listMessagePage(scopeId, sessionId, { limit: 50, before: cursor });
await harness.getMessage(scopeId, sessionId, messageId);
await harness.listSlashCommands(scopeId, sessionId);
const command = await harness.executeSlashCommand(scopeId, sessionId, '/init focus on tests');
if (command.type === 'prompt-admission') {
  console.log(command.admission.queueId, command.admission.runId);
  await command.admission.completion;
}
await harness.updateSession(scopeId, sessionId, { title: 'Renamed', archived: true });
await harness.updateSession(scopeId, sessionId, { title: null, archived: false });
await harness.getProjectState(scopeId, sessionId);
await harness.createProjectTask(scopeId, sessionId, { id: 'tests', content: 'Add tests' });
await harness.setProjectGoal(scopeId, sessionId, 'Ship the next release');
await harness.appendProjectScratchpad(scopeId, sessionId, 'One durable note.');
await harness.deleteSession(scopeId, sessionId);
await harness.deleteSessionGenerationCohort(scopeId, {
  root: { sessionId, sessionGenerationId, sessionGenerationSequence },
  authorizedCohort,
});
await harness.prompt(scopeId, sessionId, prompt, options);
const queued = await harness.enqueuePrompt(scopeId, sessionId, prompt, options);
console.log(queued.queueId);
await queued.completion;
const admission = await harness.startPrompt(scopeId, sessionId, prompt, options);
console.log(admission.queueId);
console.log(admission.runId);
await admission.completion;
const scheduled = await harness.schedulePrompt(
  scopeId,
  sessionId,
  'refresh-index',
  prompt,
  { debounceMs: 250 },
);
await harness.cancelScheduledPrompt(scopeId, sessionId, scheduled.scheduleKey);
await harness.getPromptQueueEntry(scopeId, sessionId, queued.queueId);
await harness.listPromptQueueEntries(scopeId, sessionId);
await harness.cancelPrompt(scopeId, sessionId, queued.queueId);
await harness.steerPrompt(scopeId, sessionId, admission.queueId, { steerId, prompt });
await harness.abort(scopeId, sessionId);
await harness.listPendingPermissions(scopeId, sessionId);
await harness.respondToPermission(scopeId, sessionId, permissionId, 'once');
await harness.pushRuntimeEvent(scopeId, sessionId, { type: 'workspace.changed', path: 'src/' });
await harness.listUncertainToolExecutions(scopeId, sessionId);
await harness.reconcileToolExecution(scopeId, sessionId, intentId, {
  resolution: 'executed',
  output: { committed: true },
});
const reversion = await harness.getSessionReversionInfo(scopeId, sessionId);
console.log(reversion.undoAvailable, reversion.redoAvailable, reversion.groups);
const undone = await harness.undoSession(scopeId, sessionId);
console.log(undone.revertedRunId);
const redone = await harness.redoSession(scopeId, sessionId);
console.log(redone.restoredRunId);
await harness.compactSession(scopeId, sessionId, { modelHint: 'summary-model' });
await harness.archiveSessionEvents(scopeId, sessionId, compactionEventId);
await harness.listBackgroundExecutions(scopeId, sessionId);
await harness.getBackgroundExecution(scopeId, sessionId, executionId);
await harness.abortBackgroundExecution(scopeId, sessionId, executionId);
await harness.retireScope(scopeId);
await harness.dispose();

Only one run may be active in a session. Additional prompts enter a bounded FIFO owned by FlexHarness, while different sessions can run concurrently. enqueuePrompt() resolves with { queueId, completion } after the immutable prompt and options have been accepted into that runtime queue and prompt.queued has been emitted. startPrompt() keeps its durable-admission behavior: it waits for its FIFO turn and resolves with { queueId, runId, completion } only after the canonical generation claim, run ID, and initial public audit messages have been durably reserved and the corresponding start events have been emitted. prompt() preserves the simpler behavior by awaiting completion internally.

Queue entries expose queued, starting, scheduled, running, completed, failed, and cancelled status through getPromptQueueEntry() and listPromptQueueEntries(). The list is ordered by process-local queueSequence. getPromptQueueEntry() throws FlexHarnessNotFoundError for an unknown or evicted ID. cancelPrompt() cancels one exact queue ID: a waiting entry leaves the FIFO and releases capacity immediately but remains queryable as cancelled until terminal retention evicts it; a promoted entry uses the canonical run cancellation path. Cancelling a terminal entry returns false, while an unknown ID throws. abort() remains scoped to the currently active run.

The displayed queue limits are the defaults. Outstanding count and byte limits apply per session and include every non-terminal queued or active prompt until it settles. Pending-admission limits apply to the complete harness while scope aliases are unresolved. Terminal retention applies per session. Exceeding an admission limit throws FlexHarnessQueueFullError.

Queue payloads, status records, and prompt.* queue events are process-local. The existing stores do not have a private generic queue domain: projections are deliberately redacted, Agent events are canonical conversation transactions, and jobs are AgentSession background executions. FlexHarness therefore never writes a never-started prompt into those unrelated domains. A process restart drops never-started entries; a prompt that reached durable run admission continues to use the existing canonical recovery policy and is repaired to a safe terminal state instead of being replayed.

schedulePrompt() waits for its FIFO turn, performs the same durable admission, exposes session status scheduled, and starts model preparation after its bounded debounceMs delay. Schedule keys remain unique across waiting and active prompts. cancelScheduledPrompt() returns true only while the matching schedule key can still be cancelled. Cancelling while it is still waiting rejects the schedulePrompt() call itself; cancelling after durable admission rejects the returned completion and marks its reserved audit messages cancelled.

The reservation save is the admission point. A save failure produces no start events or active audit. If disposal begins while that save is in flight and the save commits, admission still resolves and its completion settles as cancelled; disposal waits for terminal finalization.

listMessagePage() returns the newest contiguous page in chronological order. limit must be an integer from 1 through 50 and defaults to 50. nextCursor is opaque, limited to 4096 UTF-8 bytes, bound to the resolved storage namespace and session, and remains stable when newer messages are appended. Mismatched and stale cursors fail validation. Each page carries startIndex, the zero-based index of its first message in the session's visible history (for an empty page, the index the page ends at), so a consumer can order and page by absolute position. Appending messages leaves every index unchanged; a branch after an undo reuses the positions of the history it replaces, so a consumer re-reads its pages after session.history.changed. getMessage() performs an exact lookup. Transfer identifiers are limited to 512 bytes, text and reasoning parts to 96 KiB, complete messages to 480 KiB, and complete page envelopes to 512 KiB. A page may therefore contain fewer messages than requested. Oversized text is truncated and an otherwise oversized parts collection is replaced with an explicit elision marker; metadata that still cannot fit fails validation. Canonical private Agent events are unchanged.

updateSession() supports title replacement, explicit title clearing with null, and archive state through archived. Title-only updates remain available while prompts are queued or running, while permission is pending, and after archival. Requests containing archived are rejected while the session has any outstanding prompt or pending permission; a mixed title-and-archive request is rejected atomically without changing the title. Archived sessions expose archivedAt. Deleting a session cascades through its complete descendant subtree. One durable root-keyed tombstone group hides every newly affected live session, and the delete also joins any already-separate descendant cleanup groups without rewriting their roots. FlexHarness then cancels queued and active subtree work, emits terminal queue events, waits for admitted initialization, and purges runtime queue status while cleaning runtime and persisted domains child-first. The requested root tombstone is removed last after every domain confirms cleanup; project-management cleanup confirmation is a retained durable project tombstone rather than physical removal. An imported session can hold no project-management state, so its deletion writes no project tombstone. A successful live deleteSession() call emits session.deleted for each session it newly tombstoned; retries of an existing tombstone and automatic load, retirement, or disposal cleanup emit no deletion events. Direct deletion of a descendant cascades only through that descendant's subtree. Cleanup authority follows the resolved storage namespace, so scope aliases share the same groups. A partial failure retains durable ownership for retry by a later deleteSession(), namespace load, retireScope(), or dispose() call.

deleteSessionGenerationCohort() is the generation-fenced destructive form for controllers. Its root and every authorizedCohort entry use the exact { sessionId, sessionGenerationId, sessionGenerationSequence } shape. The cohort contains unique entries in strictly ascending sessionId order and is limited by FLEX_SESSION_GENERATION_COHORT_MAX_ENTRIES (2048). FlexHarness atomically validates the matching root, its complete live subtree, and every retained descendant tombstone group before reserving deletion. A missing or different root generation returns { matched: false } without touching a newer generation; a missing or mismatched cascade entry throws before new tombstoning or cleanup. Extra valid entries do not expand the cascade. A matched partial cleanup remains retryable with the same authority and resolves to { matched: true } when cleanup completes. The existing deleteSession() method remains available for callers that intentionally own the complete reusable-ID namespace.

abort() returns true only while cancellation is still accepted. Terminal persistence is the run's commit point; once it starts, abort() returns false and the already-fixed terminal outcome completes while the session remains busy.

Steering a Running Turn

steerPrompt() hands a message to the running run of a prompt instead of queueing it behind that run:

const admission = await harness.startPrompt(scopeId, sessionId, 'Plan the release.');
await harness.steerPrompt(scopeId, sessionId, admission.queueId, {
  steerId: 'steer-1',
  prompt: 'Keep the database migration out of it.',
});
const result = await admission.completion;
console.log(result.unappliedSteerIds); // [] once the run took the steer in

The run takes a steer in at its next step boundary: after the tool results of its current model step are recorded and before its next model call. A running tool call or a pending permission is never interrupted; the steer waits for the boundary that follows it. A steer taken in is a user message of the run (steerId set) that the model sees in its next call and every later turn. If the run had already answered, the assistant message that answered is completed and the run answers on in a new assistant message after the steer; a steer taken in before the run answered anything is placed before its still empty assistant message. IFlexPromptResult.userMessage stays the prompt and assistantMessage is the run's last assistant message. Undo and redo treat the run's messages as one turn.

steerPrompt() resolves with { steerId, queueId, runId } once the run accepted the steer and prompt.steer.accepted was emitted; prompt.steer.applied carries the messageId of the user message the steer became. Every accepted steer ends in exactly one of two ways: it is in the model context of later turns, or it is listed in unappliedSteerIds, in arrival order, on IFlexPromptResult and run.finished, and on the finished queue entry when that list is not empty. A steer is listed when the run did not take it in before it ended, because its last model step was already running or because it was stopped or failed, and also when the run itself does not stay in context. keptInContext on run.finished and on the finished queue entry says whether it does: it is false for a failed run and for a run stopped before any model output, and such a run lists every steer it accepted, including the ones it took in and reported with prompt.steer.applied. Send its prompt and its unapplied steers again to have them answered. The run stops taking steers atomically when its generation ends or abort()/cancelPrompt() stops it. A steer for a prompt without a running run, or whose run is ending, is refused with FlexHarnessSteerRejectedError reason not-running; the caller sends it as the next turn instead. Reusing a steerId within a run is refused with reason duplicate, an unknown queueId throws FlexHarnessNotFoundError, and child sessions reject this general API (their exact parent run uses sendSubagentInput() or delegate_send instead). A run holds at most maxOutstandingPromptsPerSession pending steers, and their bytes count against the session's outstanding prompt byte limit until the run takes them in or ends, so startPrompt() and enqueuePrompt() are refused with FlexHarnessQueueFullError while pending steers fill it. They are process-local like queue entries: a process restart ends the run and drops its pending steers, while steers it took in are durable parts of the turn.

Slash Commands

parseSlashCommand() is the public strict pure parser. It accepts at most 768 KiB, returns not-command for input not starting with /, malformed for invalid slash syntax, and parsed with the exact input, lowercase command name, separator-stripped raw argument text, and OpenCode-compatible quoted tokenization. Command names match [a-z][a-z0-9_-]{0,63}. The public isValidSlashCommandName(name) and isReservedSlashCommandName(name) helpers let applications validate registration names against the same rules. Single and double quotes group tokens and are stripped; escapes are not interpreted.

Applications register immutable custom commands at construction:

const harness = new FlexHarness({
  // scopeResolver, modelResolver, and other options...
  slashCommands: [
    {
      name: 'review-area',
      description: 'Review one area of the workspace.',
      template: 'Review $1 with these additional constraints: $ARGUMENTS',
    },
    {
      name: 'refresh-index',
      description: 'Refresh the application-owned workspace index.',
      async handler({
        scopeId,
        scope,
        storageKey,
        sessionId,
        sessionGenerationId,
        sessionGenerationSequence,
        rawArguments,
        arguments,
        signal,
      }) {
        return indexer.refresh({
          scopeId,
          scope,
          storageKey,
          sessionId,
          sessionGenerationId,
          sessionGenerationSequence,
          rawArguments,
          arguments,
          signal,
        });
      },
    },
  ],
});

compact, init, undo, and redo are reserved. listSlashCommands() verifies the scope and session and returns immutable data-only descriptors with kind, placeholder hints, current immediate availability, and workspaceReversion. compact, undo, redo, and custom handlers require an otherwise idle command session. Prompt templates and init use the normal bounded FIFO and may wait behind an active prompt. Their descriptors report availability from the same queue-admission conditions used by execution: lifecycle, slash ownership, pending reversion, root-session eligibility, outstanding count, and estimated prompt bytes. A dynamic capacity race may still produce FlexHarnessQueueFullError during admission. Only one slash-command execution may own a session at a time.

At most 128 custom commands may be registered. Every registration must be a plain object with exactly one of template or handler; names must match [a-z][a-z0-9_-]{0,63}, be unique, and not use a reserved name. Optional descriptions must be non-empty and at most 2048 UTF-8 bytes. Templates must be non-empty and at most 768 KiB, and the expanded prompt must also fit 768 KiB. Registrations are copied and frozen during construction.

/undo and /redo, plus undoSession() and redoSession(), move a durable history cursor. The direct methods return { revertedRunId } and { restoredRunId }; the slash forms return { type: 'operation', name: 'undo' | 'redo' }. Each committed cursor move emits one session.history.changed event with direction, runId, and the selected session identity. A committed branch emits the same event with direction: 'branch' and no runId, so controllers should refresh the complete selected session. Capture finalization, cleanup, and metadata-only changes do not emit this event.

A completed root-session turn defines an operation-group boundary. Failed and cancelled turns after it belong to that group; leading failed or cancelled turns belong to the first completed group. No completed boundary means there is nothing to undo. Undo applies selected segments in reverse order and redo applies them in forward order. Without turnReversionProvider, only transcript and future model context move. Hidden messages disappear from getMessages(), message pages, exact message lookup, and future model context. Starting a new prompt, template, handler, or compaction from an undone position commits a branch: hidden messages and segments are removed durably and cannot be redone. Successful event archival also commits hidden redo history; a missing compaction or failed archive leaves it intact. Automatic budget compaction archives covered events after generation finalization and advances the same undo horizon; set contextCompactionBytes: false to disable that automatic policy.

Two horizons bound undo. Retention pruning removes the oldest complete visible units when reversionLimits is exceeded. Explicit event archival marks covered turns context-unavailable and prunes complete prefixes that can no longer be rebuilt; FlexHarness never crosses that archive horizon. Manual compaction without archival retains the original events and remains undoable. Schema-1 projection history and sessions migrated from 2.x have no reversion segments, so historical turns are not retroactively undoable; newly written turns are tracked normally.

Session metadata archival through updateSession(..., { archived: true }) only sets archivedAt. It does not archive Agent events, retire captures, or remove undo history.

executeSlashCommand() is the authoritative parser and lookup boundary. Its result distinguishes not-command, malformed, unknown, completed operation, bounded handler-result, and prompt-admission. A prompt admission contains the normal { queueId, runId, completion }; await admission.completion for the model result. Unknown commands are never admitted as literal prompts. Known unavailable commands and invalid arguments throw typed FlexHarness errors. Options accept modelHint, system, maxSteps, and signal; commands do not accept attachments. /compact takes modelHint and signal only and refuses system and maxSteps with FlexHarnessValidationError. Aborting a template or init execution cancels its exact queued or started prompt without affecting another queue entry.

Templates replace every $ARGUMENTS with untouched raw argument text. $1 through the highest referenced positional placeholder use tokenized arguments, with the highest position receiving all remaining tokens joined by spaces. Missing positions become empty. A template with no placeholders appends non-empty raw arguments after a blank line. /init uses the OpenCode 1.18.15 AGENTS.md initialization prompt with provider-neutral active-workspace wording.

Handler context is frozen and contains only the resolved scope identity, session identity, required canonical sessionGenerationId and sessionGenerationSequence, raw and tokenized arguments, and an AbortSignal. Handler results are converted with the configured toolOutputLimits; void becomes JSON null. Handler failures use externalErrorProjector with source slashCommand. Same-session command overlap is rejected, including reentry from a handler. Scope retirement and disposal abort and await active handlers; prompt-admission commands transfer immediately to the normal prompt lifecycle.

Workspace Reversion Provider

reversionPolicy defaults to transcript-optional, preserving the V1 behavior described above. Set it to workspace-required when transcript and workspace traversal must move together. This policy requires an IFlexTurnReversionProviderV2 at construction.

Applications using the original protocol can continue to provide all six unchanged IFlexTurnReversionProvider operations:

import type { IFlexTurnReversionProvider } from '@modelprofile.com/flexharness';

const turnReversionProvider: IFlexTurnReversionProvider<IProjectScope> = {
  prepare: (context) => workspaceSnapshots.prepare(context),
  inspectCapture: (context) => workspaceSnapshots.inspectCapture(context),
  finalize: (context) => workspaceSnapshots.finalize(context),
  inspectApply: (context) => workspaceSnapshots.inspectApply(context),
  apply: (context) => workspaceSnapshots.apply(context),
  release: (context) => workspaceSnapshots.release(context),
};

Protocol 2 adds the protocolVersion discriminant and a tagged finalized outcome. The prepare, apply, apply-inspection, and release contexts remain the V1 shapes:

import type {
  IFlexTurnReversionProviderV2,
} from '@modelprofile.com/flexharness';

const turnReversionProvider: IFlexTurnReversionProviderV2<IProjectScope> = {
  protocolVersion: 2,
  prepare: (context) => workspaceHistory.prepare(context),
  inspectCapture: (context) => workspaceHistory.inspectCapture(context),
  async finalize(context) {
    const capture = await workspaceHistory.finalize(context);
    if (capture.changedPaths.length === 0) {
      return {
        disposition: 'no-change',
        reference: capture.cleanupReference,
      };
    }
    if (!capture.revertible) {
      return {
        disposition: 'nonrevertible',
        reference: capture.cleanupReference,
        reasonCode: 'git.unmerged',
        affectedWorkspaces: [{ id: capture.workspaceId, label: capture.workspaceLabel }],
      };
    }
    return {
      disposition: 'revertible',
      reference: capture.reference,
      affectedWorkspaces: [{ id: capture.workspaceId, label: capture.workspaceLabel }],
    };
  },
  inspectApply: (context) => workspaceHistory.inspectApply(context),
  apply: (context) => workspaceHistory.apply(context),
  release: (context) => workspaceHistory.release(context),
};

const harness = new FlexHarness<IProjectScope>({
  // scopeResolver, modelResolver, stores, and other options...
  turn