@termwright/mcp
v0.7.7
Published
thin MCP server over the public driver API
Downloads
1,506
Readme
@termwright/mcp
An MCP server that lets an agent drive terminal programs the way a person would: launch a real pseudo-terminal, read a compact accessibility-style snapshot, click a button by its ref, wait on a condition, and ask what changed.
It is deliberately thin. Every tool validates its arguments with zod, calls the
public @termwright/driver API, and renders the result. There is no locator
engine, no wait loop and no matching heuristic here — a behaviour that differed
between this server and the Native Host would be a bug in this package.
Install
pnpm add -D @termwright/mcpNode >= 22, ESM only.
Usage
Register the binary with your MCP host (stdio is what hosts spawn):
{
"mcpServers": {
"termwright": { "command": "termwright-mcp" }
}
}A typical agent loop:
import { serveStdio } from '@termwright/mcp';
const running = await serveStdio(); // stdio, one implicit session
process.on('SIGINT', () => void running.close());// terminal.launch -> { terminal: "t1", semanticTree: "available", compact: … }
// envMode defaults to "replace": the child gets a minimal environment, not the
// operator's secrets. Pass "inherit" (or name variables in env) when it needs more.
{ "command": ["node", "app.js"], "columns": 100, "rows": 30 }
// terminal.snapshot -> the compact format, plus refs / cursor / modes / scroll
Terminal t1 100x30 revision 42
semanticTree: available
dialog "Permission" ref=semantic:n7@42 bounds=(8,20,40,9) modal
button "Approve" ref=semantic:n8@42 bounds=(14,23,11,1) focused
button "Reject" ref=semantic:n9@42 bounds=(14,36,10,1)
visible text:
…
// terminal.click { "terminal": "t1", "ref": "semantic:n9@42" }
// terminal.wait_for { "terminal": "t1", "wait": "text", "text": "Rejected" }
// terminal.wait_for { "terminal": "t1", "wait": "focused", "testId": "reject" }
// terminal.capture_since{ "terminal": "t1", "cursor": 42 } -> changed rows + subtrees
// terminal.close { "terminal": "t1" }OpenTUI applications can request explicit probe injection without resolving a preload path themselves:
{ "command": ["bun", "src/app.ts"], "probe": "opentui" }The text wait proves the PTY output. When the next operation requires a semantic
result, wait for that explicit state (focused, checked, selected, …)
before capturing. The condition must represent a change from the baseline; an
already-satisfied condition proves no new commit. Screen-only output and the
prefix of a future semantic frame are intentionally indistinguishable until the
probe publishes a causal signal.
For a condition that may outlive the MCP transport's per-request timeout, start
a session-owned watcher. It uses the same canonical conditions and revision
notifications as terminal.wait_for; it does not poll snapshots. Cancelling a
watch.wait request only detaches that receiver. The watcher continues and
buffers its result until another watch.wait receives it or watch.cancel,
terminal.close, or session shutdown removes it.
// watch.start -> { "watchId": "w1", "startRevision": 42 }
{ "terminal": "t1", "wait": "text", "text": "Import complete" }
// watch.wait { "watchId": "w1" }
// watch.cancel{ "watchId": "w1" }Local MCP Monitor
Add --monitor to the stdio MCP command to open a small, read-only browser
view when the agent launches its first terminal:
{
"mcpServers": {
"termwright": {
"command": "termwright-mcp",
"args": ["--monitor"]
}
}
}MCP startup itself does not open an empty window. The monitor follows the
terminal lifecycle: it starts with the first terminal.launch, supports tabs
while several terminals are live, and closes after the last terminal.close.
The monitor shows every live terminal, its styled cell screen, cursor, process and recording status, durable watchers, and the semantic tree. Hovering or selecting a semantic node outlines its known bounds on the terminal and shows its role, state, value, actions, test id, and geometry. Updates stream over a local SSE connection. This is an observer for following an agent's work; the full UI remains the place for test runs and trace replay.
The terminal grid preserves ANSI palette and RGB colours, text attributes, and cursor position. It scales the complete grid to the available panel instead of introducing a second terminal scrollbar, but never enlarges it beyond its real cell size automatically. The scale indicator explains when fitting is active; the controls allow manual zoom, reset to automatic sizing, and fullscreen of the terminal alone. Termwright starts a dedicated local Chromium or Firefox app process with an isolated temporary profile, then closes that process and removes the profile after the last terminal closes.
The listener binds to 127.0.0.1 and uses a fresh random capability path.
--monitor --no-open prints the URL to stderr when the first terminal launches
without launching a browser; --monitor-port N selects a fixed local port.
Closing the MCP session also closes the monitor and its live connections.
Streamable HTTP, for hosts that connect over a socket:
termwright-mcp --http --port 7333 --show-auth-token
# Explicit opt-in prints the endpoint and its fresh per-launch bearer token.Every MCP-bearing HTTP request, including initialize and DELETE, must send
Authorization: Bearer <token>; an allowlisted CORS preflight is the only
bearer-free request and cannot reach MCP routing. Library callers receive the
token as handle.authToken; the SDK transport accepts it through
requestInit.headers:
import { StreamableHTTPClientTransport } from '@modelcontextprotocol/sdk/client/streamableHttp.js';
import { serveHttp } from '@termwright/mcp';
const handle = await serveHttp();
const transport = new StreamableHTTPClientTransport(
new URL(`http://127.0.0.1:${handle.port}/mcp`),
{ requestInit: { headers: { authorization: `Bearer ${handle.authToken}` } } },
);Sessions are keyed by Mcp-Session-Id in this package's own SessionRegistry,
not inside transport objects. A session id is routing metadata, never a
credential: authentication happens before session lookup, refresh, initialize
or request-body buffering. Each session owns its terminals, DELETE disposes
them, and the ceiling (16 sessions, 16 terminals each) is enforced before a
transport exists.
The listener binds to loopback by default. A non-loopback host is refused
unless allowNonLoopback: true is set (CLI: --allow-non-loopback). That opt-in
does not add TLS: put a remotely reachable listener behind a private,
authenticated TLS boundary because the bearer otherwise crosses the network in
cleartext. Browser requests carrying Origin are rejected by default; an
embedding may allow exact HTTP(S) origins with allowedOrigins. A bounded
per-peer rate limiter protects authenticated work and accepted preflights
without letting invalid credentials or preflight traffic consume a legitimate
client's bucket; configure its window/request/client ceilings with rateLimit
when the deployment has a known proxy or concurrency envelope.
Streamable HTTP gives no disconnect signal, so a session also expires after
idleTtlMs without a request (10 minutes by default, 0 to disable). Every
request naming a session refreshes it; expiry runs the full teardown — terminals
closed, children gone, traces released, slot returned — and writes a line to
stderr. stdio has no TTL: there, EOF on the pipe is the signal.
Tools
Live terminal tools
| Tool | Purpose |
| --- | --- |
| terminal.launch | Starts a program in a real pseudo-terminal and returns a terminal handle plus the first snapshot. The child gets a minimal environment unless envMode is "inherit"; values passed in env are never echoed back. |
| terminal.capabilities | What this session supports: whether a semantic tree is published, which adapter publishes it, and the terminal geometry. Call it before relying on role-based targeting. |
| terminal.snapshot | One typed view of the terminal: compact semantic refs, visible text, cursor, terminal modes and scroll position. variant "full" writes the complete dump (text, ANSI, HTML, semantic tree) to disk and returns only refs plus the file path. The returned revision is the cursor for terminal.capture_since. |
| terminal.capture_since | Incremental view: the screen rows that differ and the semantic subtrees that were added, removed or updated in the latest committed semantic tree since the given cursor. A screen change alone does not imply a future semantic commit; wait for an explicit semantic state when the caller requires one. The cursor must be a revision this server handed out earlier (snapshot or capture_since); older cursors fail with history-truncated. |
| terminal.query | Queries the current state without acting or waiting for a match. Zero matches return immediately. Use terminal.wait_for to await appearance, then query to inspect matches. |
| terminal.checkpoint | Returns the atomic session/contract/screen/semantic identity used by revision-safe actions and waits. |
| terminal.actionability | Runs the same ActionPlanner used by execution, but sends no input. Reports every authoritative requirement and the chosen strategy or typed rejection. |
| terminal.click | Sends a real click mouse report through the pseudo-terminal. Success confirms delivery of input, not completion of the application operation; use terminal.wait_for for its resulting state. Fails closed with input-mode-disabled when required tracking or encoding is disabled or unobservable. |
| terminal.double_click | Sends a real double-click mouse report through the pseudo-terminal. Success confirms delivery of input, not completion of the application operation; use terminal.wait_for for its resulting state. Fails closed with input-mode-disabled when required tracking or encoding is disabled or unobservable. |
| terminal.hover | Sends a real motion mouse report through the pseudo-terminal. Success confirms delivery of input, not completion of the application operation; use terminal.wait_for for its resulting state. Fails closed with input-mode-disabled when required tracking or encoding is disabled or unobservable. |
| terminal.click_at | Sends a physical mouse click at a viewport cell without claiming that a semantic control receives it. Useful for composite controls whose visible child owns the hit cell; a child may stop event propagation. Success confirms input delivery only, so verify the application result with terminal.wait_for. |
| terminal.press | Sends key chords as real bytes, honouring the modes the program enabled (application cursor keys, keypad). Examples: "Enter", "Escape", "Control+K Control+U". With a target, the node must already be focused. |
| terminal.type | Types text as individual keystrokes (not a paste). With a target, the node must already be focused; use terminal.fill for focus + replacement. |
| terminal.fill | Ensures the semantic control receives focus through the real input path, selects its current value, and types the replacement. |
| terminal.check | Uses the central action planner and real terminal input to check a checkbox or radio, then verifies semantic state. |
| terminal.uncheck | Uses the central action planner and real terminal input to uncheck a checkbox or radio, then verifies semantic state. |
| terminal.paste | Pastes text, wrapped in bracketed-paste markers when the program enabled that mode. Use it for multi-line input instead of terminal.type. |
| terminal.write_raw | Writes bytes to the pseudo-terminal verbatim — no newline, no key encoding. The escape hatch for sequences the key encoder does not model. |
| terminal.drag | Drags with real mouse reports: either from one target to another (toTarget), or between two cell positions inside the source target (from/to). |
| terminal.wheel | Sends wheel reports over a target. Positive deltaY scrolls down. |
| terminal.resize | Resizes the pseudo-terminal; the child sees a real SIGWINCH. |
| terminal.signal | Sends INT, TERM, KILL or HUP to the child. Destructive by design: terminal.close cleans up without signalling. |
| terminal.scrollback | Emulator-side history: read a line range, search it, or move the viewport. The child sees nothing — no input is sent. |
| terminal.select_cells | Selects cells in Termwright's emulator only. No mouse input reaches the application, so this cannot test application selection or Ctrl+C. Use terminal.drag followed by terminal.press and terminal.wait_for to test that flow. |
| terminal.copy_selection | Returns text selected in Termwright's emulator by terminal.select_cells; this does not invoke the application's copy behavior. |
| terminal.wait_for | Revision-driven waits — never a sleep. "text"/"title" wait for content, locator states use the driver's canonical Conditions, "quiet" explicitly waits for heuristic silence, "render" for a render after a given revision, "exit" for the child to exit. |
| terminal.close | Bounded physical cleanup: hangs up the pseudo-terminal, finalizes an optional recording, and forgets the handle. Send signals explicitly with terminal.signal if the child must be killed first. |
Durable watcher tools
| Tool | Purpose |
| --- | --- |
| watch.start | Starts a revision-driven wait owned by the MCP session. It keeps running when the initiating request ends; use watch.wait to receive its buffered result. |
| watch.wait | Waits for the next buffered result. Cancelling or timing out this MCP request leaves the watcher alive; call watch.wait again with the same id. |
| watch.cancel | Cancels and forgets a session-owned watcher. |
Trace tools
| Tool | Purpose |
| --- | --- |
| trace.open | Validates a .twtrace directory or zip and returns a handle plus its metadata: the recorded command, viewport, duration, exit status and whether the session published a semantic tree. Start every replay investigation here. |
| trace.overview | The shape of a recording: every step with its status and timing, the cast markers, the exit status, and which step failed. Use it to pick the moment worth reconstructing before calling trace.frame_at. |
| trace.frame_at | Rebuilds the screen at a moment — named by timeMs, stepIndex or marker — by replaying the recording into a headless emulator, and pairs it with the semantic tree of the nearest revision at or before that moment. Reads exactly like a live terminal.snapshot. |
| trace.diff | Reconstructs two moments of a recording and reports what moved: changed screen rows and changed semantic subtrees, in the same shape as terminal.capture_since on a live session. |
Targeting
Targeting precedence is ref, selector, testId, role (+name), label, text, screenText.
semanticTree: unavailable means the program ships no integration — target physical output with screenText, never semantic text or role.
Names and text accept /pattern/flags. Locators are strict: more than one match returns
ambiguous-locator unless nth is explicit.
Replaying a recorded failure
A failing run leaves a .twtrace archive; the trace.* tools read it with the
same vocabulary as a live session.
// trace.open { "path": "out/login.twtrace" } -> handle tr1 + what was recorded
// trace.overview { "traceId": "tr1" } -> steps, markers, exit, which step failed
// trace.frame_at { "traceId": "tr1", "stepIndex": 1 }
Terminal tr1 40x6 revision 2
semanticTree: available
dialog "Permission" ref=semantic:n1@2 modal
button "Approve" ref=semantic:n2@2 disabled
visible text:
…
// trace.diff { "traceId": "tr1", "fromMs": 0, "toMs": 3000 } -> changed rows + subtreesReconstruction is @termwright/trace's: stateAt() returns the cast prefix and
the nearest semantic snapshot, and the prefix is replayed through the same
headless emulator the HTML report uses. A moment is named by timeMs,
stepIndex or marker — exactly one of them.
Archives are per session, capped at 8 open and 128 MB each; at the ceiling the coldest handle is closed and named in the result, and re-opening a path always works.
Screenshots
terminal.snapshot and trace.frame_at take screenshot: true and attach a PNG
as ImageContent, rendered by @termwright/screenshot — a cell grid becomes an
SVG with the glyph outlines embedded and resvg rasterises it, so there is no
browser in the loop and no dependency on the agent's machine having the right
font.
{ "terminal": "t1", "screenshot": true, "screenshotScale": 2, "screenshotTheme": "light" }The image is always additional: the compact tree and the screen text are in the
same result, so an agent that cannot see pictures loses nothing.
structuredContent.screenshot carries the size and selfContained — false when
a character had no embedded outline and fell back to a font the viewer may not
have. PNGs above 3 MB are refused with capacity rather than blowing a context
window; lower screenshotScale or resize the terminal.
Refs and revisions
A ref is semantic:n8@42: node id at semantic revision 42 (screen matches get
screen:1,2,9,1@7). Refs go straight to harness.locatorForRef(), so they resolve
by node identity — two buttons with the same name stay distinct. A producer
which promises stable identity can resolve that node again in later revisions.
Frame-local identities and grid refs remain revision-bound; take a fresh
snapshot when either becomes stale.
terminal.snapshot also returns a screen revision; pass it back as the
cursor of terminal.capture_since to get only the rows and semantic subtrees
that changed in the latest committed tree. A changed screen does not promise a
future semantic commit; wait for an explicit semantic condition when one is
required. Cursors the server never handed out fail with history-truncated (the
last 16 captures per terminal are retained).
Programs without a framework probe or custom semantic producer report
semanticTree: unavailable. There are no invented roles: target physical
output with screenText.
Application logs
A terminal shows what a program drew; its log says what it decided. Follow one at launch and read it alongside the screen:
// terminal.launch
{ "command": ["node", "app.js"], "logs": [{ "path": "out/app.log", "label": "app" }] }
// terminal.capture_since -> changed rows, changed subtrees, and:
logs: 2
1840ms [app] ERROR upstream refused the token
1841ms [app] WARN falling back to cached profileAn existing file is followed from its end, so a session never replays the
previous run. Entries are buffered per terminal (1000 deep) and returned since
your cursor, with logsOmitted counting anything that fell out in between —
computed when you read, so a program that went quiet still reports its last
drops. Files are polled, so a line written moments ago may arrive on the next
call; re-asking with the same cursor is lossless.
Structured records from an instrumented adapter keep their level, logger and attributes; a followed file yields the raw line.
The same view exists for a recording: trace.frame_at returns the entries
leading up to that moment (maxLogs, default 20) and trace.diff the ones
between the two, so "what was it saying when the screen looked like this" reads
the same live and in replay.
Crashes
When a child dies on its own, the driver records what the session knew and this
server surfaces it three ways: attached to whatever call failed next,
in terminal.capabilities and terminal.snapshot instead of a bare closed
session, and in trace.overview for a recording whose meta.json carries one.
crash: the program exited on its own — code=7 signal=null at 812ms
last input: key "\r"
screen tail:
Error: boom
at thing (app.js:3:9)That matters because a locator which never resolved because the program is gone
otherwise reports a plain timeout, and an agent reading a timeout waits longer.
In a recording, crash.timeMs is the cast offset, so
trace.frame_at { traceId, timeMs } jumps to the moment of death with the screen
and the semantic tree of that revision.
The screen tail is unredacted — it is what the terminal displayed, secrets included. It is bounded (40 lines, 500 characters each) and never logged, but treat it like a screenshot when storing or forwarding a result. Paste contents are the one thing never recorded: the driver keeps their size only.
Errors
Failures come back as tool results with isError set. The text content reads
error stale-snapshot: ref semantic:n8@42 no longer exists at semantic revision 43
suggestion: re-resolve the locator; the node identity is no longer present
semanticTree: trueand the same payload — kind, message, suggestion, bounded candidates,
screenExcerpt — travels structured in _meta["io.termwright/error"]. Stack
traces never leave the server, and neither the child's environment nor the
session token appears in any result or log.
CLI
termwright-mcp # serve over stdio
termwright-mcp --monitor # open the live monitor with the first terminal
termwright-mcp --monitor --no-open # print its URL when the first terminal launches
termwright-mcp --http --port N # serve Streamable HTTP on /mcp
termwright-mcp agent-context # versioned JSON: every tool, param, enum, exit code
termwright-mcp usage # one-screen cheat sheet
termwright-mcp skill --out DIR # emit an agent-skill package (SKILL.md + reference)
termwright-mcp --json … # machine-readable errors carrying `kind`Exit codes: 0 ok / 1 assertion / 2 usage / 3 no-session / 4 ipc / 5 internal.
agent-context and the skill package are generated from the live zod schemas,
so neither can drift from the tools. skill writes SKILL.md (what an agent
reads), reference.md (every tool and parameter) and agent-context.json; with
no --out it prints them instead. The umbrella termwright CLI imports
buildAgentContext(), buildUsage() and buildAgentSkill() rather than
spawning this binary.
Testing this package
pnpm build && pnpm typecheck && pnpm testThe end-to-end suite runs a real MCP client over InMemoryTransport against the
real driver and the fixtures in packages/driver/test-fixtures. It skips itself
where no pseudo-terminal can be opened, or with TERMWRIGHT_SKIP_PTY=1.
