@opencraw/mcp
v0.1.21
Published
An MCP server exposing OpenCraw's crawler as agent tools: probe a page, validate recipes, run a crawl, list recipes.
Maintainers
Readme
@opencraw/mcp
An MCP (Model Context Protocol) server exposing @opencraw/core as tools an agent can call
directly: probe a page, validate recipes, run a crawl, list what's already authored, compare two runs. Local transport only
(stdio) — the host launches it as a subprocess, the same shape as the opencraw cli.
For a shared, remote endpoint that authors recipes where they will run (and publishes them to a crawl host), see
@opencraw/azure-durable's /mcp.
This does not auto-author recipes from a sentence. The five tools are, list_recipes aside, the same
primitives the opencraw cli gives a terminal; the calling agent still writes the JSON recipes, using probe and
validate to iterate, the same way this project's own example recipes were built by hand. An agent that
can write files passes their paths; one that can't (a chat-only host) passes the recipes inline.
Register it
{
"mcpServers": {
"opencraw": {
"command": "node",
"args": ["/absolute/path/to/opencraw/packages/mcp/bin/opencraw-mcp.mjs"],
"env": { "OPENCRAW_CHROMIUM": "/path/to/chrome" }
}
}
}OPENCRAW_CHROMIUM and OPENCRAW_INSECURE_TLS=1 in the server's environment are the defaults for every
probe/run call's browserPath/insecureTls, so a sandbox without Playwright's own bundled browser
does not need every tool call to repeat them.
OPENCRAW_PLUGINS (or OPENCRAW_HOOKS) points at a plugins module: named exports hooks (the recipes'
hook steps and transforms), accessPlugins (for { kind: "plugin" } access profiles) and captchaSolvers;
a module whose default export is { name: function } is read as hooks alone. Only the server's environment
names it, never a tool call, so an agent can run recipes that use your plugins but can't make the server load
a module of its choosing.
OPENCRAW_ACCESS points at an access config (access.md): proxy profiles, with
credentials as {{env.NAME}} read from the server's environment. The probe and run tools then take an
access argument naming a profile. The file and the credentials never pass through a tool call.
OPENCRAW_PROFILES is the directory of the persistent browser profiles recipes name in
session.browserProfile (default .opencraw/profiles in the server's working directory).
Tools
probe
Fetches a page and reports where its data lives: JSON-LD blocks, inline JSON objects above a size with
their keys, .json URLs referenced in the markup, script hosts, and links that look like an API. With
browser: true it also renders the page and lists the JSON responses observed while it settles, for
endpoints only a script fetches after load.
| Input | Meaning |
|---|---|
| url | The page to fetch. |
| browser? | Also render it in a browser (slower, a few seconds). |
| browserPath?, insecureTls?, userAgent? | As in the cli. |
| access? | An access profile from the server's OPENCRAW_ACCESS config. |
Returns { url, status, findings: { jsonLd, inlineJson, jsonUrls, scriptHosts, apiLinks }, observed }. For a
PDF (served as application/pdf, or a local path) it adds pdf: { pages, rows, headers }: the first rows and
every likely table header with the selector a table extract needs. For a spreadsheet or a CSV (served as a
spreadsheet type or text/csv, or a local .xlsx, .csv or .tsv path) it adds
workbook: { csv?: { encoding, delimiter }, sheets, rows, headers }. For a presentation it adds
deck: { width, height, slides, headers, charts, grids }, and for JSON, JSON Lines or YAML
json: { format, type, tree, lists }: the structure, and every list of records with its path. An HTML page
adds html: { tables, outline?, frontMatter? }: its tables' header rows, and for Markdown or a Word document its
sections and front matter keys. XML (a feed, a sitemap) adds xml: { root, namespaces, lists, tree, sitemap? }.
validate
Loads and binds recipes (an output recipe plus its input recipes), from files or passed inline: parses each against its schema, checks every mapping resolves, reports every problem with its JSON path.
| Input | Meaning |
|---|---|
| paths | Recipe files (.json, .jsonl) or directories of them. |
| recipes | Instead of paths: the recipes themselves, as an array of recipe objects or as JSON / JSON Lines text. For a host that can't write files. An inline recipe's issues name it by position: recipes[1], or recipes:2 for a line. |
Give exactly one of paths and recipes.
Returns { ok, output?, inputs: string[], issues: [{ path, message, source }] }.
run
Crawls the recipes at the given paths, or passed inline.
| Input | Meaning |
|---|---|
| paths or recipes | As in validate: exactly one output recipe, any number of input recipes. |
| out? | Write records to this JSON Lines file instead of returning them inline. Use this for anything beyond a handful of records — the tool result is not the place for a large crawl's output. |
| append?, resume? | As in the cli (resume needs append). |
| only? | Run only these input recipe ids. |
| dryRun? | One record per input recipe instead of a full crawl — for checking a recipe under construction. |
| headed?, browserPath?, insecureTls?, userAgent? | As in the cli. |
| access? | The access profile for recipes that name none, from the server's OPENCRAW_ACCESS config. |
| parallel? | How many input recipes run at once (1 to 16, default 1). |
Returns { report: CrawlReport, records?, truncated? }. records/truncated are present only when out
was not given, and records is capped at 50 even then.
list_recipes
Lists the recipe files in a directory, split by kind, with their ids — cheaper than a full validate
call, for checking what already exists before writing a new recipe.
| Input | Meaning |
|---|---|
| dir | A directory of recipe files. |
Returns { outputs: [{ path, id }], inputs: [{ path, id, output, mode }], others: string[] }.
diff
Compares two runs' JSON Lines files by record key: what was added, what was removed, and which fields changed, before and after. For "what changed in this price list since last week" without re-reading it all.
| Input | Meaning |
|---|---|
| previous, current | The two runs' files (a run's out). |
| key? | The fields that identify a record (the output recipe's key fields). Default: each line's _key, which a run with append writes. |
| ignore? | Fields not compared, such as a scrapedAt the crawl stamps. |
Returns { added, removed, changed, unchanged, shrunk?, summary, changes }: summary is the report a person
reads, changes the first 200 changes (removed, changed with fields: [{ field, before, after }], added),
and shrunk is set when the current run holds under half the records the previous one did.
Building
nx build @opencraw/mcp, nx test @opencraw/mcp (unit tests per tool), nx run @opencraw/mcp:e2e
(spawns the built server and drives it with the MCP SDK's own Client/StdioClientTransport, against the
fixture shop @opencraw/core ships; needs a browser — npm run playwright:install once, or
OPENCRAW_CHROMIUM=/path/to/chrome).
