@cubicecho/agent-mcp-pool
v3.7.0
Published
A pool of long-lived Model Context Protocol clients, exposing every connected server's tools to an OpenAI-compatible agent loop as `<slug>__<tool>`.
Maintainers
Readme
@cubicecho/agent-mcp-pool
A pool of long-lived Model Context Protocol clients, exposing every connected server's tools to
an OpenAI-compatible agent loop as <slug>__<tool name>.
Connections are long-lived and shared across runs: a stdio server is a child process, and spawning one per run would cost more than the run.
Install
npm install @cubicecho/agent-mcp-pool @modelcontextprotocol/sdk@modelcontextprotocol/sdk (>=1.30) is the one peer dependency, because client() hands back
the SDK's own Client and an instanceof against a second copy in the tree means nothing.
openai is not one: tools() returns ToolDefinition, which is declared here and assignable
to OpenAI's ChatCompletionTool because TypeScript is structural. It was a required peer for two
type positions that are erased at compile time, so a consumer proxying MCP and never calling a
model installed 24 MB to satisfy them. ESM only, Node >=22.
The seam
The two servers this came from both did import { db } and read an mcp_servers table. The
pool now asks for its rows instead:
import { McpPool } from "@cubicecho/agent-mcp-pool";
export const mcp = new McpPool({
load: () => db.select().from(mcpServers),
clientName: "task-server",
clientVersion: SERVER_VERSION,
});clientName and clientVersion are the clientInfo of the MCP handshake — the whole of what a
dialled server learns about who is calling it, and so what it logs, gates a behaviour on, or
quotes back in a support channel. The version defaults to this package's own, read from its
manifest; set it beside the name, since a name that is yours next to a version that is the
pool's tells the server something untrue.
A row is one of two shapes, discriminated on transport, and carries only its own arm's fields:
{ id: "git", label: "Git", enabled: true, transport: "stdio", command: "uvx", args: ["mcp-server-git"] }
{ id: "docs", label: "Docs", enabled: true, transport: "http", url: "https://mcp.example.com/mcp" }StdioServerConfig and HttpServerConfig are both exported; McpServerConfig is their union.
Flat, an http row still had to write command: "", args: null, env: null — three fields
nothing would ever read, and a command that reads as configured rather than absent — while a
stdio row with no command at all compiled, and createTransport could only refuse it at runtime,
one connect too late. Migrating: a row written as an object literal that supplies the other
arm's fields now fails excess-property checking; one that arrives from a typed variable — a Drizzle
select, a parsed config — is unaffected. The fix is deleting the fields that row's transport never
used.
A stdio server may also name a cwd; absent, it inherits this process's. Several servers resolve
relative paths — a filesystem root, a sqlite file — against their working directory rather than
against an argument, and it counts as part of the connection: editing it restarts the child.
sync(configs?) reconciles against what it is given, or against load when given nothing.
syncSoon() debounces a reconcile past a transaction commit, and flush() pays one off early
for a reader that would otherwise be shown the pool as it stood before its own write.
load is optional: a consumer that owns its own configuration passes the rows every time instead.
On such a pool, a sync() or reconnect(id) with no configs has nothing to reconcile against and
is refused with a no-configs McpPoolError rather than treated as an empty set — reconciling
against nothing closes and forgets every server, and doing that because an argument was left off
is the most destructive thing this API could do by accident. sync([]) still closes everything,
from a caller who said so.
State
state() reports every configured server in the order it was configured, and hands back the
row it was configured from:
for (const { config, status, error, tools, pid, startedAt, instructions } of mcp.state()) {
// config is a copy of the row you passed in, minus its credentials;
// everything beside it is what the pool made of it, or what the server said for itself
}Both halves exist so a consumer does not have to keep its own copy of the rows beside the pool's.
A UI that draws the edit form and the connection state as one line needs the row, and a shadow
map of the same rows goes stale the moment anything reconciles without going through it — which
syncSoon() and a load-driven sync() both do. The order is the caller's array rather than
Map insertion order for the same reason: entries are created by parallel connects, so without
this the operator's list reorders itself according to which child started quickest.
id, slug and label stay alongside config — those are the effective values the pool
actually used.
Each tool in state().tools carries both of its names: name is the server's own, qualified is
<slug>__<name> — what the model is offered and what call() takes. Both, because an operator
reads the first and debugs with the second, and because catalog()'s identically shaped list
carries the qualified one under name. Two { name, description } lists meaning different
things is a transposition waiting to happen, and the fix is to stop making the reader remember
which is which.
env and headers are left out of it. That UI is a browser, sending state() to it is the
shortest way to draw that line, and for a real server those two fields are an API key and an
Authorization: Bearer — so the default is the safe one, and the caller genuinely rendering the
edit form server-side is the one that asks:
mcp.state({ secrets: true }); // config carries env and headers againconfig is a copy rather than the row itself. A caller that holds its rows and edits one in
place used to get a pool that never reconnected — sameConnection was being asked whether a row
differed from itself — while state() reported the edit as though the child had been restarted
for it.
pid and startedAt describe the connection rather than the configuration, so both are absent
unless one is up, and pid over http, which has no child. They are what make ready mean
something concrete to an operator: a pid finds a wedged child in ps, and a start time is how a
server that is quietly crash-looping is spotted, since status reads ready either side of a
restart.
instructions and capabilities are the rest of what initialize returned, and are absent for
the same reason — a server that is not connected has no handshake to report:
const guidance = mcp
.state()
.flatMap(({ label, instructions }) => (instructions ? [`## ${label}\n${instructions}`] : []))
.join("\n\n"); // straight into a system prompt, synchronously, spawning nothinginstructions is what a server says about itself for a model to read — the things a tool
description has no room for, like "resolve the library id before querying docs" — so a system
prompt is where it belongs. It is reported here rather than left to client() because that door
dials: it connects an idle server by design, so under lazy, or with idleTimeoutMs set and a
server just reaped, building a prompt would spawn children. capabilities is what says whether
resources/list or prompts/list on a client() is worth attempting at all; without it the
choice is an error round trip per server per surface, or not offering the surface. Both are
reported under indexTools: false as well, unlike tools — they were already in hand, and the
consumer that turned indexing off is the one proxying the protocol.
toolsFingerprint is a hash of what the server last listed — every tool's name, description,
input schema and annotations — and is absent while it has no list. A cold server reports the one
its last-known list hashes to, beside stale: true. A server whose descriptions
change after an operator approved them is the "rug pull" of the MCP security write-ups, and it can
only be noticed against something stored:
const { toolsFingerprint } = await mcp.probe(row); // what the operator approved
await save({ ...row, approved: toolsFingerprint });
for (const { id, toolsFingerprint } of mcp.state()) {
if (toolsFingerprint && toolsFingerprint !== approvedFor(id)) warn(id);
}It moves when the list does — a reconnect to an upgraded server, or a tools/list_changed — and
not when the row is renamed or the same tools arrive in another order. Annotations are in it
because a tool that keeps its description and flips destructiveHint is the same attack on a host
that auto-approves by the hint. toolsFingerprint(tools) is exported for a consumer hashing a
tools/list it fetched itself.
serversWith(state, capability) is that check written once — the connected servers whose
handshake offered it, which every host had spelled out for itself:
import { serversWith } from "@cubicecho/agent-mcp-pool";
for (const { id } of serversWith(mcp.state(), "prompts")) {
// worth a prompts/list
}Scope
tools, catalog and call all take an optional set of server ids. Absent means every
connected server; empty means none of them — the two must not collapse, because "this agent
has no servers linked" is a real and correct state. call re-checks the scope rather than
trusting the definitions the caller was given: a model that has seen a tool name once will call
it again from memory.
A scoped call does not start a server the run cannot reach. A name the index cannot answer
wakes the servers whose slug could have produced it — a cold one under lazy, a crashed one past
its backoff — and the run's scope gates that too, not only the refusal after it. Otherwise a run
scoped to one server spawned another's child, drained its tool list, and then said no server
offers that tool: a process started for a run that may not reach it, and a latency difference that
answers the question the shared message exists to leave unanswered.
tools names both of its own:
pool.tools({ names: ["echo__add"], servers: [agent.serverId] });tools(names, servers) was two collections of strings in an order nothing could check, so a
transposition was not a type error — and its answer is an empty array, which is also the right
answer for a run scoped to servers that offer nothing. A consumer adopting the pool swapped them,
and the migration compiled, connected and offered its model no tools at all. catalog(servers)
stays positional: it has no two arguments that could be confused for each other. call takes
its scope as the third argument as before, or an object when it needs more:
pool.call("echo__add", input, { servers, signal, timeoutMs: 5000 });Hooks
A row can carry hooks: tool calls of its own that the host makes at points in a session, the way
Claude Code's hooks run at SessionStart or Stop — except a hook here is an MCP tool call and never
a command. Rows are edited from UIs, and a command hook would make "can edit the server list" the
same permission as "can run anything on the host".
The case it was written for is memory: a server that recalls before every turn and remembers after it, without the model having to decide to call either.
const memory = {
id: "zeromem",
label: "Memory",
// ...transport fields...
// The model is not offered these; hooks may still call them.
hiddenTools: ["zeromem_remember", "zeromem_forget_session"],
hooks: [
{
id: "recall", on: "beforeTurn", tool: "zeromem_recall", inject: true, maxTokens: 800,
args: { query: "{{prompt}}", exclude_session: "app:{{session.id}}", format: "text" },
},
{
id: "remember", on: "afterTurn", tool: "zeromem_remember",
args: { session_id: "app:{{session.id}}", turns: "{{turn.messages}}" },
},
],
};
// Before the request:
const outcomes = await pool.runHooks("beforeTurn", { session: { id }, prompt }, { signal, onNotice });
const { text } = contextBlocks(outcomes); // <context source="Memory">…</context>, or ""
// After the reply — not awaited, and not on the turn's signal:
void pool.runHooks("afterTurn", { session: { id }, prompt, reply, turn: { index, messages } });The pool never fires a hook itself — only the host knows when a turn starts. It supplies the runner, so every host runs the same rows the same way.
| Event | Runs | Context beyond session.id, host, now, vars.* | inject |
|---|---|---|---|
| sessionStart | before a session's first turn | prompt | yes |
| beforeTurn | before each turn's request | prompt, turn.index | yes |
| afterTurn | once a turn has its reply | prompt, reply, turn.index, turn.messages | no |
| beforeCompact | before old messages are summarised away | compacting, range.from, range.through | no, but see veto |
| sessionEnd | when a run that ends, ends | status, reply | no |
| sessionDelete | when the host deletes a session | — | no |
- Templates. A string that is exactly
"{{path}}"becomes the value itself, so an array or a number goes through as one.{{path}}inside a longer string is interpolated as text. A path the context has no value for skips the hook: asession_idsent as"app:"would file a turn under the wrong session. - Validation.
validateHooks(row.hooks)reports an unknown event, a placeholder the event does not offer,injecton an event that runs too late,vetoon an event that announces nothing stoppable, and duplicate ids. It takesunknownand reports a hook's shape too — a missingon, aninjectthat is not a boolean — since what a form holds is not yet aToolHook. Run it when the row is saved. - Never rejects. A failed call, a timeout, an abort and a skipped hook each come back as an
outcome with
ok: false, and are passed toonNotice. A memory server that is down costs the turn its recall, not the turn. - At once, in order. An event's hooks run in parallel, and their outcomes come back in configuration order.
- Bounded. A hook gets its
timeoutMs. Otherwise it gets 3s onsessionStartandbeforeTurn, because the user is waiting on those, and on a hook that can veto, because the compaction is — and the SDK's timeout everywhere else. The bound covers waking a server that is down as well as the request itself. - Read and add only. A hook cannot stop a turn or rewrite it. What it returns reaches the model
only through
contextBlocks. That function caps each block at its hook'smaxTokens(1000 by default) and the total at 2000. Itsinjectedlists each hook that made it in, with itstokensand itstextexactly as it went into the block — trimmed, cut with…if over a cap, without the wrapper — for a host that shows the user what each hook added.
Declining a compaction
The one thing a hook can stop is a compaction. A beforeCompact hook whose row says veto: true
may ask for the summary not to be written — a memory server mid-write, a session the operator has
pinned. Its outcome then carries veto: true, and agent-core's consult (and
compactTranscript with hooks.honourVeto) is what acts on it. Nothing in the pool acts on it:
the pool reports what the hook said.
hooks: [{ id: "hold", on: "beforeCompact", tool: "zeromem_status", veto: true }],- The row grants it, the tool asks for it. Like
inject,vetois set on the row, so which servers may stall a compaction is the operator's choice rather than the choice of whoever wrote the tool. Without the flag, a tool answering{"veto": true}has simply answered that. - How a tool asks. Its output parses as a JSON object whose
vetoistrue— MCP gives a tool no channel but its result. Areasonstring becomes the outcome'stext, so a host's note can say why. Prose, a list,{"veto": "true"}: none of them is a veto, so a tool that has never heard of any of this cannot stop a compaction by accident.readVetois that rule, and is exported for a host that wants to hold a server to it. - A failure is no opinion.
vetois only ever set besideok: true. A call that failed, timed out or was skipped leaves it unset, because agent-core readsok: falseas "no opinion" — a memory server that is down must cost a compaction nothing, not stall every one of them. - Only
beforeCompact, validated or not.validateHooksrefusesvetoon any other event, andrunHooksholds a row to the same rule, since a row need not have been through the validator. - A forced compaction ignores it. That is agent-core's rule: a run already past its window has
no better option, and a veto there trades the summary for a
ContextOverflow.
hiddenTools is the other half. A hidden tool is left out of tools() and catalog(), and
call() refuses it as one that does not exist, unless the caller passes { hidden: true }. Hooks
pass it. state() still reports hidden tools, marked hidden, so a form can offer to unhide them.
Both fields are read at call time, so an edit applies without a reconnect.
Editing rows in a browser
The root entry spawns children, so importing it pulls in node:child_process and the SDK's stdio
transport — which is why each UI that edits rows kept its own copy of the hook events and its own
config importer. Two subpath entries import nothing from Node or the SDK at runtime, and are what
a browser bundle takes instead:
@cubicecho/agent-mcp-pool/hooks—HOOK_EVENTS,INJECT_EVENTS,VETO_EVENTS,hookVars,validateHooksand the rest ofsrc/hooks.ts, with the hook types.@cubicecho/agent-mcp-pool/servers—fromMcpServersJson,validateServerConfig,sameConnectionandserversWith, with the row types.
Both are also exported from the root, for a server that already imports it. A test compiles them
with no Node types and walks their import graph, so a node: import there fails CI rather than a
consumer's build.
import { fromMcpServersJson, validateServerConfig } from "@cubicecho/agent-mcp-pool/servers";
// A paste of Claude Desktop's, Cursor's or a README's config. `{ "mcpServers": { … } }`, VS Code's
// `servers`, `{ "fs": { … } }` and a bare body (with `{ name }`) are all read.
const rows = fromMcpServersJson(text, { env: {} });
const problems = rows.flatMap((row) => validateServerConfig(row).map((p) => `${row.id}: ${p}`));- Each key is the row's
id,slugandlabel, anddisabled: trueisenabled: false. Aurlwith notypeis http.sseis refused with a message rather than imported as a row that could never connect: the pool speaks streamable HTTP. ${VAR}and${VAR:-default}are filled fromenv, which defaults toprocess.envwhere there is one and to nothing in a browser. A variable with no value is left as written, so the form shows what still needs filling in rather than an empty string.validateServerConfig(row)reports every problem with a row of any shape — a missing command, a url that is not http, a timeout that is not a whole number, an id that cannot namespace tool names — and its hooks' problems throughvalidateHooks. Empty meanssync()can have it. Whether the server answers isprobe()'s question.sameConnection(a, b)is the pool's own answer to whether an edit reconnects: only the fields a child is made of, with absent and empty the same. A host that decides before handing the row over should ask this rather than compare rows itself, or the two disagree about which edits restart a server.
Lifecycle
Eager and long-lived by default: sync() connects every enabled server and holds the connection,
because for an agent loop spawning one child per run costs more than the run.
A gateway has the opposite pressure — dozens of installed servers, most idle most of the time — so the lifecycle is a policy rather than a fixed behaviour:
new McpPool({
load,
lazy: true, // sync() registers entries; the child waits for a use
idleTimeoutMs: 300_000, // close a server after five minutes without one
});A registered-but-unconnected server sits at idle, which is neither disabled (switched off) nor
error (tried, failed, waiting out a backoff). call() and client() start it; tools() and
catalog() do not — they answer from the tools it was last known to have.
What a cold server offers
A server that was reaped or stopped keeps its list. Nothing went wrong, so what it offered a moment
ago is the best account of what it will offer when a call brings it back; dropping it made a
reaped server vanish from the catalogue until something happened to call it by name, which nothing
could, since the catalogue is where a model gets the names. state() and catalog() mark such a
server stale: true, and the mark goes when it connects.
A call to a last-known tool starts its server — that one alone, and only after the scope and
hiddenTools refusals, so a run that may not reach a tool does not spawn a child to be told so.
If the server comes back offering something else, the list is replaced, tools-changed fires, and
a tool that is gone is refused as one that does not exist. A crash is the other case: a server at
error keeps nothing, so state() never reads as one that is down and still has tools to offer.
That covers a server the pool has seen. For one it has not — a lazy pool's first boot, or any
restart — toolsCache is where a list outlives the process:
new McpPool({
load,
lazy: true,
toolsCache: {
load: (id) => db.tools.get(id), // { connection, tools } | undefined
save: (id, cached) => db.tools.put(id, cached),
},
});save is handed every list a server gives that the store does not already hold — on a connect and
on a tools/list_changed — and is not awaited: a slow store must not hold a connect open, and a
failing one is logged and costs only the next cold start's catalogue. load is asked once per row
that sync() registers without dialling. connection is a hash of the command or url the list was
fetched over, env and headers included, so a list from before a row was pointed somewhere else
is not offered as its tools and the cache never holds a secret. Delete a row's entry when you
delete the row.
With that, a lazy pool sends a complete catalogue on its first turn having spawned nothing.
idleTimeoutMs resets on every use, and McpServerConfig.idleTimeoutMs overrides it per server —
0 opts one out entirely. A reap is a success path, not a crash: the server goes back to
idle with no error and no backoff, so the next call reconnects immediately rather than waiting
out a penalty for something that did not go wrong. Both options absent is exactly today's
behaviour.
A use that lands after the clock has already fired keeps the server: the reap runs on the
reconcile queue, and it checks that the timer it was armed with is still the one the server is
waiting on. Cancelling is not enough on its own — clearTimeout on a timer that has already fired
does nothing — so without that check a call arriving in the window between the fire and the queued
close had its client closed underneath it, and the model was handed a transport error from a server
that was in use.
Stopping and restarting one server
shutdown() is every server and forgets them, and sync() only closes what the configs dropped.
One server on its own is stop() and reconnect():
await pool.stop(id); // close the child, keep the row: `idle`, no error, no backoff
await pool.reconnect(id); // close it and dial again, changed or notstop() is the reap path reached by a person instead of a clock, which is why it lands on idle
rather than a status of its own: idle already means registered, no child, nothing wrong. Nothing
stands between a stopped server and the next call() or client(), including the backoff a crash
would have armed — so "restart this wedged server" is a stop() and then a use, and an operator
uninstalling a server can close its child before the row leaves disk rather than racing a live
process holding files open. It stays stopped in the meantime: reconcile steps over an idle
entry whether or not the pool is lazy, so the next write to the server table will not dial it.
reconnect() is the unconditional one, lazy included. It used to drop the entry and let the
reconcile rebuild it, and a lazy reconcile registers an entry and waits for a use — so on a lazy
pool it stopped the server it was asked to restart, and the caller found out on whichever later
call spawned a child. Which of the two things the method did depended on a constructor flag set
somewhere else entirely; now the two pools mean the same by it. A disabled row is still not
started: it is off for a reason a reconnect does not overrule.
Not indexing at all
The drain that fills the index is the pool doing its job for an agent loop, and pure cost for a
gateway that proxies tools/list through from the client that asked. indexTools: false skips it:
new McpPool({ load, lazy: true, indexTools: false });What that buys is a round trip per page off the first request that spawns a server — a lazy
gateway pays the spawn on a user-facing request, and against a server that pages ten at a time the
walk is several more before the request it actually made is even sent. It also stops holding a
second copy of every tool for nobody, and stops a server that answers initialize and then wedges
on tools/list from failing to connect at all: that one is still good for a resources/read.
The index is then empty for ever, so tools(), catalog() and state().tools are empty and
call() refuses every name — without waking anything, since connecting a server could not index
it either. client() is the surface that remains, and probe() is unaffected: a probe exists to
report what a config offers.
Loading tools on demand
Thirty MCP tools with real schemas are past a small model's context before the system prompt, so an agent loop sends a catalogue and a handful of schemas and lets the model ask for the rest. The pool's half of that is four things:
const servers = agent.serverIds;
mcp.catalog(servers); // every name and description, no schemas
mcp.tools({ names: mcp.alwaysLoaded(servers), servers }); // the schemas sent before it asks
mcp.search("read the config file", { servers, limit: 8 }); // the catalogue, cut to this turn
mcp.tools({ names: asked, servers }); // what `load_tools` came back withNone of them connects a server; a cold one answers from its last-known list.
alwaysLoad is on the row, the mirror of hiddenTools: the tools an operator wants in front
of the model from the first turn, by the server's own names. The row is where it is said so each
consumer does not keep that list somewhere else. alwaysLoaded() answers with qualified names and
only ones tools() would hand back — a name the server does not offer, or a hidden one, is left
out rather than sent on to be skipped and logged.
search() returns catalog()'s shape holding only what matched, servers ordered by their best
tool and tools best first. The ranking is word overlap — a query word scores once, by the best
place it turned up: the tool's name, then its title, then its description, then its server's slug
and label. Names are split on punctuation and camelCase, and reads, reading and read are one
word. No embedding and no dependency: it is a shortlist, where roughly right about which thirty of
a hundred and forty is the whole job. rankTools is exported for ranking a list of your own.
tokens is on every tool in catalog(), search() and state(): an estimate of what its
whole definition costs to send, at four characters a token over the JSON tools() hands out —
label prefix and schema included, and the same rough count agent-core budgets with. And tools()
says so in log.info when the set it was asked for comes to more than toolsTokenWarning (default
3000, 0 to silence). Ollama's default num_ctx is 4096 and a runtime that is overrun truncates
without a word, so this is the only place the problem is ever stated. It is said once, and again
only for a larger set than the last one reported — an agent loop asks every iteration.
Past the agent surface
tools() returns OpenAI definitions and call() returns a string, because a string is what goes
back into a message array. A consumer that is proxying MCP rather than driving a model wants
neither, and wants resources, prompts, subscriptions and logging that a string was never going to
carry. client(id) hands back the connected client:
const { resources } = await (await pool.client(id)).listResources();A remote server can lose the session that client was opened with — it restarted, or it reaped a
session it took for abandoned — and go on answering. Nothing closes, so the pool is not told, the
server stays ready, and every request on that client is refused: 404 for a session id the
server no longer knows, or 400 for a missing one where a server that was stateless has come back
stateful. use(id, run) is client(id) with that handled:
const { resources } = await pool.use(id, (client) => client.listResources());The first such refusal redials the server and calls run once more with the new client; any
other failure is passed through. Both refusals come before the server dispatches the request, so
a tools/call sent again has not run twice. Requests that share the dead client share one redial,
and it waits for the ones still in flight. call() does the same on its own, and sessionLost
is the test both use, exported for a consumer that keeps its own clients.
listAllTools(client, options?) is the other half a raw client needs: tools/list is paginated,
the page size is the server's choice rather than the caller's, and a tool left on page two is
not merely unlisted — it is absent from the index, so call() refuses it as one that does not
exist. The pool and probe() both drain the cursor; a consumer driving the client itself wants
the same walk rather than one listTools. A timeout in options bounds that walk rather than
each page of it, for the reason above; everything else in options is passed to every page.
resultText is the flattening call() does, exported separately: MCP answers with a list of
content blocks and a message array holds one string. A consumer driving the client itself and
still putting the answer in front of a model wants the same rule rather than its own. A block that
came with text arrives as that text, wherever that block keeps it, and what has none is named
rather than dropped — so a model that asked for a screenshot is told it got one instead of being
handed an empty string and left to conclude the call failed:
| Block | Flattened to |
| --- | --- |
| text | its text |
| resource, text arm | the resource's own text — a file the server read is an answer, not a placeholder |
| resource, blob arm | [resource <uri> <mime>, <size> omitted], keeping the uri a follow-up call needs |
| resource_link | [resource_link <uri> — <name>: <description>], since the uri is what makes a link followable |
| image, audio | [<type> <mime>, <size> omitted], such as [image image/png, 42 KB omitted] |
| anything else | [<type> content] |
A result with no content at all but a structuredContent — what a server with an outputSchema
tends to answer with — is flattened to that structure as JSON, rather than reaching the model as
call()'s "(no output)". Text blocks win where there are any: they are what the server wrote
for a reader. The exception is text that only mirrors the structure, which the spec asks servers
to send and most do pretty-printed. That arrives as the compact JSON, since it says the same thing
in fewer tokens.
The definitions tools() hands back are the pool's own objects rather than copies — the agent
loop rebuilds its tool array every iteration, and the schema behind one cannot change without the
connection being torn down and remade — so they are frozen. An edit that would otherwise have
silently rewritten what every later run is offered fails at the edit instead. Shallow: parameters
is the server's own schema, passed through untouched.
Everything else the pool does applies unchanged — reconcile, the queue, crash detection with the
stderr tail, backoff, retry-on-use. A server that is merely down is retried first, the same as
call() does; a disabled one is refused, because off is not the same as out of scope. It
bypasses the scope check by construction: that guard defends against a model calling a name it
remembers, and a caller holding a server id is not a model.
What a definition has no room for
An OpenAI definition is a name, a description and a schema. A server says more than that about a
tool, and a host deciding whether a call needs a person's approval wants the rest.
describe(qualified) returns it:
const info = pool.describe("files__delete", { servers });
if (info?.annotations?.destructiveHint !== false && !info?.annotations?.readOnlyHint) {
await askUser(info);
}annotations are readOnlyHint, destructiveHint, idempotentHint and openWorldHint, as the
server sent them. title, outputSchema, the untouched inputSchema and the tool's _meta (as
meta) come with them. state().tools, catalog() and probe() carry title and annotations
too, so a UI can badge a destructive tool before anything calls it.
Annotations are claims, not facts. The spec calls them untrusted, and a server that says
readOnlyHint: true has only said so. Trust them as far as you trust the server.
describe() never connects, like tools(), and refuses what call() would: a tool outside the
scope, or a hidden one unless hidden: true is passed, answers undefined.
Results too big for the window
A tool that reads a file or fetches a page can return more than a local model's whole context
window. maxResultChars caps what call() returns, at the same three levels as the timeouts:
new McpPool({ load, maxResultChars: 16_000 }); // the pool's; unset is no cap
{ id: "fetch", url: "...", maxResultChars: 4_000 } // this server's; 0 is no cap
await pool.call(name, input, { maxResultChars: 0 }); // this call'sOver the cap, the first two thirds of the budget and the last third are kept, with a marker between them on its own line:
[truncated: kept 15958 of 204800 chars]The head is where most answers start and the tail is where a log ends. The marker tells the model
it did not see everything, so it can ask for less. A tool-error message is capped the same way.
truncateText(text, maxChars) is the same cut, exported.
A consumer that wants the blocks themselves, an image to render or structuredContent to read,
passes raw: true and gets the server's CallToolResult back:
const result = await pool.call("charts__render", input, { servers, raw: true });Unlike client(), that keeps the scope and hiding checks, coercion and the timeout. Nothing is
truncated, and an isError result is returned rather than thrown.
Calls a turn makes twice
A local model often makes the same read twice in one turn, and each is a round trip and the same
tokens of answer. resultCache answers the second from memory:
const mcp = new McpPool({ load, resultCache: { ttlMs: 30_000, maxEntries: 256 } }); // the defaults
await mcp.call("notes__read", { path: "a" }); // asks the server
await mcp.call("notes__read", { path: "a" }); // does not
await mcp.call("notes__read", { path: "a" }, { cache: false }); // asks, and keeps the new answerOff unless the option is given, and then only where two things hold: the tool declares
readOnlyHint or idempotentHint, and its row sets trustAnnotations: true. Annotations are the
server's claims about itself, and this is a place a false one costs something — a tool with side
effects marked read-only has its second call swallowed. So the operator says which servers to
believe, the same judgement as auto-approving on destructiveHint.
- Keyed on the tool and its arguments after coercion, key order aside, so
{ limit: "5" }and{ limit: 5 }are one entry where the schema says integer. - Looked up after the refusals. A cached answer is never given to a run outside the scope, or for a hidden tool, because the lookup is not reached.
- Never a failure. A
tool-erroris thrown as ever and asked again next time. - Kept whole, cut on the way out, so a hit obeys the
maxResultCharsof the call that asked. Arawcall is neither answered from the cache nor stored: what is kept is text. - Cleared per server when its connection closes for any reason, when its tool list changes, and when any call that is not read-only reaches it — what was read before a write is not what would be read after it. An idempotent write is cached and clears the reads all the same.
A hit is reported as cached: true on the call event. ttlMs is the bound on what the pool
cannot see — a file edited by something else — so keep it at the length of a turn, not a session.
Arguments
Local models get argument types wrong far more often than names. They send "5" for a number,
"true" for a boolean, an object serialised into a string, and "" for a parameter they meant to
leave out. A server with a strict validator refuses each one, and the model reads a stack trace.
So call() checks the arguments against the tool's own inputSchema first, and repairs what has
only one reading:
| The model sent | The schema says | Sent as |
| --- | --- | --- |
| "5" | integer or number | 5 |
| "true", "FALSE" | boolean | true, false |
| '{"a": 1}', '[1, 2]' | object, array | the parsed value |
| 5, true | string | "5", "true" |
| "" or null, for an optional property | a type that is neither | left out |
| the whole input as a JSON string | | the parsed object |
It recurses through properties and items. Nothing is added, a property the schema does not
name passes through, and anyOf, $ref and formats are left to the server. What still does not
fit is refused as invalid-arguments, with a message written for the model to correct from:
invalid arguments for "files__read": `limit` must be an integer, got "abc"; missing required `path`That refusal comes after the scope and hiding checks, so a tool this run may not reach is still answered as one that does not exist.
On by default, and configured at the same three levels as the timeouts, nearest first:
new McpPool({ load, coerceArguments: false }); // off for the pool
{ id: "legacy", command: "...", coerceArguments: false } // off for one server with a wrong schema
await pool.call(name, input, { coerce: false }); // off for one callOff sends the input exactly as given. coerceArguments(input, schema) is exported for a consumer
driving client() itself.
Failure
A stdio server is a child process, and child processes die. The pool watches for it: an
unexpected close moves the server to error, drops its tools from the index so the model is not
offered tools whose process is gone, and records what the child last wrote to stderr — which
for a server that failed to start is usually the only useful explanation ("no module named
mcp_server_git" rather than "MCP error -32000: Connection closed").
Every refusal from client() and call() is an McpPoolError with a code, because the
message alone cannot separate the two that matter most:
try {
await pool.client(id);
} catch (error) {
if (error instanceof McpPoolError) respond(status[error.code], error.detail);
}unknown-server, disabled, backoff (a failure recent enough that nothing was dialled — with
retryAt for when one will be) and connect-failed (dialled just now, and could not — with the
child's stderr in detail). call() adds unknown-tool and out-of-scope, which deliberately
share their message: a run must not learn that a server it was not scoped to exists. The
messages are unchanged from the plain Errors these replaced. sync() and reconnect() have one
of their own, no-configs — see the seam.
tool-error is the odd one in the list: not a refusal from the pool at all, but the pool reporting
that the server ran the tool and the tool failed — MCP's isError result, whose text becomes the
message. It was a plain Error for exactly that reason, and it carries a code anyway because it is
the one a caller most needs to tell from the others: retrying a backoff makes sense and retrying
a tool that rejected its arguments does not. The message is unchanged, so anything reading
.message is unaffected.
invalid-arguments is the pool refusing a call before the server sees it, because the arguments
still failed the tool's schema after coercion. See arguments. timeout is a call
that ran out of time. See timeouts.
A gateway answering over HTTP wants a status for each code, and every one that sat in front of
this pool wrote the same table. httpStatusFor(code) is that table:
| Code | Status |
| --- | --- |
| unknown-server, disabled, unknown-tool, out-of-scope | 404 |
| invalid-arguments | 400 |
| no-configs | 500 |
| connect-failed, tool-error | 502 |
| backoff | 503 |
| timeout | 504 |
A failed server is then retried, which is the other half: sync leaves a healthy unchanged
server alone but treats a failed one as work to do, and call brings back a server that is
merely down rather than telling the model its tool does not exist. Both are held off by
crashBackoffMs (5s), or a server that cannot start would be respawned on every write.
Timeouts
connectTimeoutMs caps how long a server gets to answer initialize and tools/list. The SDK
already applies its own 60s, so this is not about an unbounded hang — it is about how long a boot
is willing to stall. sync connects in parallel, but one wedged server still holds the whole
reconcile open for the full timeout, so the number to pick is the one your startup can afford,
not the one a healthy server needs.
Unset is the SDK's 60s, which is a ceiling rather than a budget.
It is one budget for the whole connect, not a fresh allowance per request: initialize and
every page of tools/list spend the same clock, and what is left when the handshake finishes is
what the walk gets. Page size is the server's choice — a hundred-tool server answering ten at a
time is eleven requests — so a per-request bound would really have been connectTimeoutMs × (1 +
pages), which is not a number a startup budget can be picked from without knowing a server's page
count in advance. A connect that runs out fails as a timeout rather than keeping the pages it
managed: a tool missing from the index is one call() refuses as a tool that does not exist, and
a short list is a wrong answer that looks right.
McpServerConfig.connectTimeoutMs overrides it per server, the way idleTimeoutMs does, because
connect cost is a property of the server rather than of the pool:
{ id: "git", command: "node", args: ["./node_modules/.bin/git-mcp"] } // up in milliseconds
{ id: "docs", command: "uvx", args: ["some-mcp-server@latest"], connectTimeoutMs: 120_000 }uvx on a cold cache resolves and downloads a package before it says anything. One pool-wide
number has to be the maximum of those, which leaves the fast server with no useful bound — the
wedged node child this option exists to catch still hangs for the two minutes the slow one
legitimately needs. null or absent is the pool's number, which is what every row that predates
the field says; unlike idleTimeoutMs there is no special 0, which is simply a server given no
time at all.
The row is re-read on every reconcile, so a consumer whose configuration is hand-editable does not need a restart to change it. An edited timeout does not bounce a running child — it is read at connect time, so it applies to the next one.
callTimeoutMs is the same idea for one call(), and for an agent loop it is the number most
worth setting: unset, a tool call takes the SDK's 60s, which is most of a turn. It reads the same
three levels — McpServerConfig.callTimeoutMs, then the pool's, then the SDK's — for the same
reason the connect one does: a filesystem read and a deep-research server that thinks for ninety
seconds cannot share a number, and the number that accommodates both leaves the fast server
effectively unbounded. Read at call time, so an edit applies to the next call without a reconnect.
new McpPool({ load, callTimeoutMs: 30_000 }); // the pool's
{ id: "research", url: "...", callTimeoutMs: 180_000 } // this server'sDeliberately not the SDK's resetTimeoutOnProgress: a long call that reports progress is still
cut off at this number, because a bound a server can hold open indefinitely by talking is not a
bound. A consumer that wants the other reading has client().
A call that runs out is refused as McpPoolError code timeout, carrying the timeoutMs that ran
out. It used to surface as the SDK's "MCP error -32001: Request timed out", which a caller could
only recognise by its wording.
Probing
A config is easy to get subtly wrong, and finding out at 3am when the task runs is too late.
probe() connects a config that may not be saved yet, lists its tools, and hangs up:
const { ok, error, tools, toolsFingerprint, instructions } = await mcp.probe(row); // { ok: false, error: "no module named …" }instructions is there for the same reason tools is: "Test connection" is where an operator
finds out what a row actually offers, and what its tools are for is the half a tool list does
not show. Empty where the server sent none, or where the dial never got that far.
It reports rather than throws, and error is the child's stderr where there is one — the same
tail that makes a failed server diagnosable above, which is the whole difference between "no
module named mcp_server_git" and "MCP error -32000: Connection closed".
Going through the pool is what binds the client identity, the childEnv policy and the timeout to
whatever this pool uses, so the probe dials the way the pool will. Every consumer that called the
free probe() wrote that wrapper itself, and one that bound them differently showed up only in a
remote server's logs. The free function stays exported for a caller with no pool, and takes that
identity as one argument — probe(row, "my-gateway") for the name alone, or
probe(row, { name: "my-gateway", version: "1.4.0" }) for both. A -probe suffix is appended to
whichever name arrives, so a server's log tells a test connection apart from a real one.
probeTimeoutMs overrides connectTimeoutMs for probes alone, defaulting to it. The two have
different audiences: a reconcile of thirty servers at boot can afford to be patient, and a person
who has just pressed "Test connection" cannot. A row's own connectTimeoutMs outranks both: a
server that needs two minutes to start needs them behind the button too, or the button reports a
failure for a server that works.
The free probe() reads that row too, having no pool to ask — probe(row) waits as long as the
row says. Its timeoutMs option outranks the row, since a number passed at the call site is a
decision about that one probe, and neither set leaves the SDK's 60s:
await probe(row); // the row's connectTimeoutMs, else the SDK's 60s
await probe(row, "my-gateway", { timeoutMs: 5_000 }); // this probe, whatever the row saysLogging
The pool logs — a server's tool count on connect, what it wrote on the way out, a name nothing
offers — and by default it logs to console. A consumer wondering why MCP chatter is in its
stdout wants log:
import { McpPool, type PoolLog } from "@cubicecho/agent-mcp-pool";
new McpPool({ load, log: { info: logger.debug, error: logger.warn } });Both halves are optional, so {} is silence and { error: logger.warn } keeps the failures and
drops the rest. PoolLog is exported so a consumer can declare one rather than infer it.
Notifications
Anything a server sends that the SDK does not handle itself — tools/list_changed,
resources/list_changed, prompts/list_changed, resources/updated, logging/message — is
dropped unless someone is listening:
const stop = pool.onNotification((id, notification) => {
if (notification.method === "notifications/tools/list_changed") reload(id);
});The server id comes first because a listener hears from every server at once and the notification
does not say where it came from. The handler is installed before each connect and reinstalled on
a respawn, so a logging/message sent during a server's own startup is not missed. An agent loop
can ignore all of this, but a consumer relaying the protocol onward cannot.
tools/list_changed is the one the pool acts on as well as forwards. It lists that server's tools
again — every page, on the same budget as a connect — and reindexes, so tools(), catalog() and
call() follow the server without a reconnect() restarting the child. A walk that fails leaves
the list as it was, and a pool with indexTools: false does not list at all.
The order of tools() is a guarantee: configuration order, then each server's own order, and
a re-list keeps the server's slot. Prompt caching on every local runtime — llama.cpp's prefix
cache, Ollama, vLLM's automatic prefix caching — depends on the tool array being byte-identical
from turn to turn, and a changed server appended to the end would move every definition after it.
Events
PoolLog is for a person reading a console. A tracer, a metrics exporter or a UI's activity feed
wants the same moments as data, and used to wrap every method to get them. onEvent hands them
over:
const stop = pool.onEvent((event) => {
if (event.type === "call") span(event.qualified, event.ms, event.ok, event.code);
});| Event | Carries |
| --- | --- |
| connect | serverId, ms, and the number of tools it listed |
| connect-failed | serverId, ms, and the error the row now reports |
| close | serverId and a reason, plus the error for a crash |
| tools-changed | serverId, the before and after fingerprints, the number of tools now, and the names added and removed |
| call | qualified, serverId and toolName once resolved, ms, ok, the refusal code and error, chars before any cap, truncated, hidden, raw, and cached where resultCache answered |
A close's reason is one of idle, stop, reconnect, removed, changed, redial,
shutdown or crash. It fires only for a connection that was open, and after state() already
shows where the close left the server.
tools-changed follows a tools/list_changed that moved the fingerprint; one that changed
nothing is not reported. So does a cold server that comes back offering other tools than its
last-known list. A tool whose description or schema changed is in neither added nor
removed, so the fingerprints are what say a change happened.
Listeners run synchronously as each thing happens. One that throws is logged and skipped, since a tracer's bug must not fail a tool call. With no listener subscribed, a call does no extra work.
Elicitation
A server can stop in the middle of a tool and ask the user something: a missing field, a
confirmation, a link to open. MCP calls that elicitation/create, and a client has to declare it
at the handshake. Without onElicit the pool declares nothing, so the server refuses to ask and
its tool fails. With it, every connection declares the capability and each request lands in the
handler with the id of the server that sent it:
const pool = new McpPool({
clientName: "my-agent",
onElicit: async (serverId, params, { signal }) => {
const answer = await askTheUser(serverId, params.message, params, signal);
return answer ? { action: "accept", content: answer } : { action: "decline" };
},
});params carries requestedSchema for a form, or url when the server wants a link opened.
Only form is declared by default; pass elicitationModes: ["form", "url"] when the host can open
a link. A handler that throws is logged and answered cancel, so the server gets an answer it
can handle rather than a protocol error.
The call that caused the request is still on its clock while a person reads the question, so give
that server a callTimeoutMs sized for the wait.
Roots and sampling, the other two client capabilities, are deprecated in the 2026-07-28 specification release candidate, so the pool does not offer them.
The child's environment
By default a stdio child inherits all of process.env, which is how the two servers this
came from did it. That hands third-party code every secret this process was started with, so
childEnv narrows it:
import { McpPool, MINIMAL_CHILD_ENV } from "@cubicecho/agent-mcp-pool";
new McpPool({ load, childEnv: MINIMAL_CHILD_ENV });The permissive default is kept deliberately: narrowing breaks any server that quietly depends on a variable the allowlist does not name, and the stderr tail above is what makes that diagnosable when it happens.
Connections to http servers
An http server is reached with fetch, and Node's keeps an idle connection for 4s. An agent
thinks for longer than that between two tool calls, so each call after a pause opened a new
connection first: a TCP handshake, and against a remote server a TLS one, which measured 300ms
where a kept connection took 30. The pool dials with keepAliveFetch() instead, which keeps an
idle connection for 30s.
import { keepAliveFetch, McpPool } from "@cubicecho/agent-mcp-pool";
new McpPool({ load, fetch: keepAliveFetch(120_000) }); // another idle time
new McpPool({ load, fetch: viaProxy }); // or a fetch of your ownThe number is what the pool does where a server is silent. One that answers with
Keep-Alive: timeout=N is taken at its word, and Node's own http server says timeout=5 unless
its keepAliveTimeout is raised, so a server you run needs that raised for the pool to have
anything to keep. Thirty seconds rather than more because the far end decides too: a connection
held past what a load balancer allows is closed under the client, and the next request finds out.
fetch is also on createTransport and the free probe(), and pool.probe() uses the pool's.
A transport built without one shares a single keepAliveFetch(), so a probe of a server the pool
already holds reuses its connection. This is why undici is a dependency: Node exposes no way to
set the idle time of its own fetch, and a global dispatcher would change every request the
consumer makes rather than the pool's.
Naming
Tools are <slug>__<tool name>, capped at 64 characters for OpenAI's function-name limit, and
resolved by whole-string lookup rather than by splitting on __ — the split of a shortened name
is a tool its server never had.
slug is optional and defaults to id. A consumer whose ids are already namespace-shaped has
nothing else to put in a slug column, and a second name beside such an id only gives an operator
a way to make the two disagree. state() reports the effective value, so a consumer that never
set one still sees what its tools are called.
Every character outside [A-Za-z0-9_-] becomes _ first. MCP allows a dot in a tool name and
OpenAI does not, so a server calling its tool fs.read sends a function.name the API refuses —
and it refuses the request, so one such tool costs the model every other server's tools too.
A name that then does not fit keeps its first 57 characters and spends the rest on _ plus six
hex digits of a SHA-256 of the whole name, as the server gave it. Truncating alone made two
tools sharing a 64-character prefix collapse onto one key, so the second silently replaced the
first and the model was offered a name that dispatched to the wrong tool. Hashing the name before
the substitution keeps a.b and a_b apart when they are cut. Names that already fit and already
qualify are returned byte-for-byte, so nothing that was unambiguous before changes on the wire.
When two tools want the same name
Two rows with the same effective slug, or one server's fs.read and fs_read, want a single slot
in the index. Nothing is silently overwritten: a namespace goes to the lower of the two ids, the
other server's tools are left out of tools(), catalog() and call() entirely, and the clash is
logged at error once — not again on each reindex, which runs on every connect, close and reap.
The tie is broken by the rows rather than by which child was up, or the name would move to the
other server after a reap and move back on the next call.
Catch it at save time instead. validateServers(rows) checks a whole list the way
validateServerConfig checks one, and only a list can see a duplicate id or a shared slug:
const errors = validateServers(rows);
if (errors.length > 0) return { ok: false, errors };sync() deliberately does not refuse a clashing list. Rows are edited from a UI, and a pool that
threw on one bad row would take every other server down with it.
Descriptions too long for the API
OpenAI also caps a function description at 1024 characters, and a server is free to send a usage
guide several times that. maxDescriptionChars caps what tools() sends:
new McpPool({ load, maxDescriptionChars: 1024 }); // unset is no capCounted over the whole of what the model is sent, [Label] prefix included, since that is what
the API measures, and cut like a result — head, marker, tail. state(), catalog() and
describe() still report the server's own text in full: this is a wire limit, not an opinion
about what a tool should say for itself.
mcp-router has a namespacing scheme that looks identical and is not: it splits names a foreign
MCP client invented, longest-prefix-first, and applies the same scheme to resource URIs and prompt
names. The two cannot be shared — unify on this truncation and its resource URIs corrupt.
Testing
Every suite that exercised the pool spawned a child it was not about, and every consumer declared
its own makePool() beside its own stdio fixture. @cubicecho/agent-mcp-pool/testing is that
helper, published:
import { echoServer, makePool, memoryRow } from "@cubicecho/agent-mcp-pool/testing";
const pool = makePool({ servers: { echo: () => echoServer() } });
await pool.sync([memoryRow("echo")]);
await pool.call("echo__add", { a: 1, b: 2 }); // 'add({"a":1,"b":2})'No process starts. makePool passes memoryTransport(servers) as the pool's createTransport,
which links the SDK's in-memory pair to a server built for that connect, found by the host of the
row's memory:// url. A builder rather than a server, because an SDK server connects once and a
reconnect needs another. Any McpServer or low-level Server fits there, so a consumer tests
against its own server as easily as against the echo one.
createTransport is also the seam for a transport the pool does not know, such as a socket or a
worker. It receives the row and the pool's childEnv policy, and returns an unconnected
Transport; a factory that throws fails the connect the way a child that will not start does.
probe() takes the same option. When the SDK ships the stateless HTTP transport of the
2026-07-28 draft, this is where it plugs in before the pool learns it natively.
A test that does want a real child has echoServerPath, a stdio script serving the same tools.
Where the merged behaviour came from
- The reconcile queue (
running,queue,syncSoon,flush) istask_server's. Without it, two syncs interleaving both spawn a child for the same edited server and the second orphans the first — a live process with nothing holding a handle to close it.kanban_serverstill has that race. - The per-run server scoping is
kanban_server's idea withtask_server's signature.
