@nexusbloom/mcp-server
v2.2.0
Published
MCP server for NexusBloom — agents discover tools by intent, read exact schemas, and execute them. Built on @nexusbloom/core.
Maintainers
Readme
@nexusbloom/mcp-server
MCP (Model Context Protocol) server for NexusBloom.
An agent connects once and can then discover tools by what it wants to do, read their exact parameters, and run them. No local runtime, no credentials beyond the API, no state on disk.
Built on @nexusbloom/core —
the same engine the CLI, the VS Code extension and the embed use, so a type rule
or validation change lands once and every surface inherits it.
Install
Add it to any MCP client that speaks stdio:
{
"mcpServers": {
"nexusbloom": {
"command": "npx",
"args": ["-y", "@nexusbloom/mcp-server"]
}
}
}Or run it directly:
npx @nexusbloom/mcp-serverConfiguration
All variables are optional.
| Variable | Default | Purpose |
|---|---|---|
| NEXUSBLOOM_API_KEY | (none) | API key. Without it the server runs anonymously under the free-tier rate limit. |
| NEXUSBLOOM_API_URL | https://nexusbloom.dev/api | API base. Accepts the bare host or the full /api form. |
| NEXUSBLOOM_MCP_TIMEOUT_MS | 15000 | Per-request timeout. |
| NEXUSBLOOM_MCP_CACHE_TTL_MS | 60000 | Tool-list cache lifetime. 0 disables caching. |
| NEXUSBLOOM_MCP_DEBUG | (off) | 1 logs requests to stderr. |
| NEXUSBLOOM_MCP_EXECUTION | remote | local or auto to execute tool source on this machine. Off by default on purpose — see below. |
Diagnostics always go to stderr; stdout is the MCP transport and writing anything else there corrupts the protocol.
What an agent sees
Every published tool is advertised as an MCP tool named by its slug, with its real JSON Schema — so it can be called directly:
One caveat on slugs. A slug has to match MCP's tool-name grammar (
[A-Za-z0-9_-]{1,64}) to be advertised. A catalogue slug that does not — a space, a slash, non-ASCII, or more than 64 characters — is withheld from the tool list and reported once on stderr, because a strict host validates every name in that array and rejects the whole response over one bad entry. That would cost an agent every tool rather than one. A withheld tool is still reachable through{"command":"run"}andnexusbloom://catalogue; it is only not directly callable.npm run catalogue-checkverifies a catalogue against the grammar.
{ "name": "env-validator", "arguments": { "env_content": "DEBUG=true" } }Alongside it is a nexusbloom meta-tool for when the agent does not yet know
which tool it wants:
| Command | What it does |
|---|---|
| {"command":"search","query":"validate env file"} | Rank tools by intent. Works on prose, not just slug fragments. |
| {"command":"list"} | Every published tool, one line each. |
| {"command":"schema","slug":"…"} | Exact parameters, plus a ready-to-send example invocation. |
| {"command":"run","slug":"…","params":{…}} | Execute a tool. |
| {"command":"batch","runs":[{"slug":"…","params":{…}},…]} | Execute up to 10 tools in one call. |
| {"command":"history"} | What this server process has run, newest first. |
| {"command":"history","show":"<id>"} | One run in full: input, result, timing. |
| {"command":"diff","from":"<id>","to":"<id>"} | Compare two runs field by field. |
Batching
Several tools belonging to one task go in a single turn:
{ "command": "batch", "runs": [
{ "slug": "env-validator", "params": { "env_content": "DEBUG=true" } },
{ "slug": "data-generator", "params": { "count": 2 } }
] }Two behaviours make it safe rather than merely fast:
- Validate everything first. Every slug is resolved and every schema checked before the first run. A typo in item four costs zero quota and rejects the whole batch — a batch is one intent, so half of it is not a useful answer.
- Then isolate failures. Past that point each run is independent. One tool erroring does not discard the seven that worked; the response reports the tally, lists failures first, and the structured block keeps every payload.
Capped at 10 runs, because anonymous execution allows 30 requests/minute and one turn should not spend the budget the next ten calls need.
Progress notifications
If the host sends a progressToken in _meta, the server emits
notifications/progress while a batch runs — one tick per entry, plus an opening
tick so the host is never silent. Progress is monotonic, and a failed
notification never fails the call. A host that sends no token receives none:
unsolicited progress is a protocol violation, not a nicety.
Resources — read the catalogue without spending a call
Hosts that support MCP resources can pull the catalogue directly, which is how an agent holds the whole surface in context instead of discovering it one slug at a time:
| URI | Contents |
|---|---|
| nexusbloom://guide | This usage guide as markdown: meta commands, recovery paths, limits. |
| nexusbloom://catalogue | Every tool as JSON — slug, description, category, tags, parameter counts, has_schema, and the URI of its full manifest. |
| nexusbloom://tools/{slug} | One tool's full manifest: input and output JSON Schema, plus call: { tool, arguments } — a ready-to-send example. |
The catalogue deliberately omits full schemas; that is what the per-tool URI is for. A 31-manifest inline blob would exhaust the context budget before the agent had chosen anything.
Slugs in resource URIs resolve by exact match or unambiguous abbreviation only, identically to tool calls, and an unknown slug produces a protocol error naming the close matches rather than an empty body an agent might quote as fact.
Prompts — one-click starts a host can offer
Four prompts, each assembled against the live catalogue rather than canned, so the text a model receives names tools that actually exist and carries calls that already validate:
| Prompt | Arguments | Gives the model |
|---|---|---|
| find-tool | goal | Tools matching the goal, ranked, with each one's parameter cost and the manifest URI. |
| use-tool | slug | One tool's exact parameters, its output_schema, and a valid call to copy. |
| plan-batch | task | Likely participants, the batch call shape, the 10-run cap, and the whole-batch validation rule. |
| recover | error, slug | What each error code means, what to do about it, and the failed tool's real call shape. |
recover exists because the common failure is not "the tool broke" — it is an
agent retrying a validation failure unchanged, or guessing slugs. The prompt
states the codes, marks which are retryable, and embeds the required fields of the
tool that just failed.
use-tool resolves slugs exactly, unlike tool calls. A prompt naming a tool
is a deliberate user choice; silently running a different tool because a prefix
happened to match would be the wrong answer to a request the user made explicitly.
Local execution (opt-in, and read this first)
NEXUSBLOOM_MCP_EXECUTION chooses where tool code runs:
| Mode | Behaviour |
|---|---|
| remote (default) | Always call the API. No tool source is ever fetched. |
| local | Always execute the published source in a child process. A tool with no source is an error. |
| auto | Execute locally when source exists, otherwise fall back to the API. |
An unrecognised value falls back to remote, so a typo cannot silently switch on
code execution.
The threat model. Executing a tool locally means executing code that arrived
over the network. The sandbox from @nexusbloom/core gives the tool its own
process, a SIGKILL deadline, a heap cap, and no access to this process's stdout.
It is isolation, not a security sandbox: that code still runs with your user
permissions and can read files, open sockets, and spawn processes.
That is why it is off by default. An MCP server is driven by whatever the model
decides to ask for, so the default posture has to be the one where a surprising
tool call cannot become arbitrary code execution on a developer's machine. Turn it
on when you trust the catalogue — a self-hosted deployment publishing only your
own tools — or to keep working while the API is unreachable (auto).
Enabling it prints the same warning to stderr at startup, once, because the threat model belongs on the record at the moment it is switched on.
Credentials are stripped from the child. The sandbox forks the tool with the
parent's environment minus anything whose name looks like a credential —
NEXUSBLOOM_API_KEY, STRIPE_SECRET_KEY, AWS_SESSION_TOKEN, *_PASSWORD,
*_AUTH, and so on, matched by substring. Tool code has no legitimate reason to
read your API key, and without this a published tool could simply read it and
exfiltrate it — which turns "runs with your permissions" into "runs with your
account". Ordinary variables (PATH, project config, anything you have set) pass
through untouched, so tools that read their own configuration still work.
Run history and diff
Every run this server performs is recorded, so an agent can see what already happened instead of guessing:
{"command":"history"}— an index: run id, tool, status, duration, newest first.{"command":"history","show":"r3"}— one run with its input and its result.{"command":"diff","from":"-2","to":"-1"}— two runs compared field by field.-1is the most recent run,-2the one before it, so comparing consecutive runs needs no lookup.
A diff leads with the verdict and flags changed inputs, because a result that moved because its input moved is not a regression — conflating the two is how a diff gets dismissed as noise and then ignored when it mattered.
Three deliberate constraints:
- In memory, never on disk. Run payloads are user data and frequently secrets. A log on disk would be an unencrypted store nobody asked for, and losing it on restart costs nothing the tool call did not.
- Scoped to the process, not the conversation. Most hosts keep the stdio process alive across turns, so history spans every conversation turn the process has served — it is not reset between them. Every history response says so, and carries the process start time, because an agent that finds a run it cannot account for needs somewhere to stop guessing. Observed in a real session: an agent correctly reported that an unexplained earlier run had different input, then invented a story about who had changed it and when, and escalated it into a security incident. The facts were right; the cause was fiction. Naming the boundary is what makes "this predates me" a checkable statement.
- Bounded. 50 runs, each payload capped at 8,000 characters. Truncation is recorded, and a diff says so — a truncated payload reporting "no differences" would be the worst possible failure mode here.
- Inputs redacted. Any key matching
key,secret,token,password,auth,cookieorsession— pluspasswd,passphraseandcredential— is masked before storage, at any depth up to 6 levels, without preserving length. Beyond that depth the value is replaced with[deep]rather than inspected, so a deeply nested payload cannot smuggle a credential past the walk.
The index is also readable as nexusbloom://history, for hosts that would rather
pull it than call for it.
Errors are actionable
Failures come back as an isError result with a stable code and a
retryable flag, so an agent can decide whether to retry rather than parsing
prose:
{
"success": false,
"error": "No tool named \"env-validatr\" is published. Closest matches: env-validator. …",
"code": "TOOL_NOT_FOUND",
"retryable": false
}A wrong slug suggests the right one. A missing argument names the field. A timeout is marked retryable; a validation failure is not.
Validation before the request
Input is checked against the tool's own schema before anything is sent, so a malformed call costs no API quota and returns an error naming the exact field. The API remains authoritative — this only rejects what is provably wrong.
Developer scripts
| Command | What it does |
|---|---|
| npm run shell | Interactive MCP client over real stdio. shell:mock starts a fake API for you. |
| npm run agent-demo | Walks the path an LLM takes — discover, search, schema, run, batch, recover — with no model needed. agent-demo:mock starts its own mock. |
| npm run catalogue-check | Every tool in the seed catalogue walked through the real server: advertise, describe, validate, run, batch, resolve as a resource. Exits non-zero on any failure, so it works as a CI gate. |
| npm run mock-api [port] | The fake API on its own. Reuses a running one or explains what holds the port. |
agent-demo checks the API is reachable before it starts. That is deliberate:
when the mock is not running, every check otherwise fails with its own
ECONNREFUSED, and ten identical failures read as a broken server rather than a
missing dependency.
Design notes
Execution never guesses. Slugs resolve by exact match or unambiguous
abbreviation only. Asking for env-validatr is an error with a suggestion, not
a silent run of env-validator — a wrong-but-successful result is worse than a
wrong-but-obvious failure.
Search is lexical, not embedded. Ranking runs with no network, no model and
no vector store, is deterministic, and returns identical output for identical
input. That makes it testable and fast enough to run on every call. It matches
across slug, name, description, tags and input field names, with light
morphological handling so "environment" reaches env-validator.
Failures degrade, they do not crash. An unreachable API produces a working server that explains itself on each request, not a process that exits or a host that concludes the server is broken. A failed catalogue refresh keeps serving the last good list rather than emptying it.
