@opencraw/cli
v0.1.21
Published
Command-line tools for OpenCraw: probe a page for scrapable data, validate recipes, run a crawl.
Readme
@opencraw/cli
Command-line tools for @opencraw/core: probe a page for scrapable data, validate recipes,
and run a crawl from the terminal.
Install
npm install -g @opencraw/cli
# or, inside this workspace, after `npm run cli:build`: node packages/cli/bin/opencraw.mjs <command>Commands
opencraw validate <recipe files or directories...>
opencraw run <recipe files or directories...> [options]
opencraw probe <url> [options]
opencraw diff <previous.jsonl> <current.jsonl> [--key a,b] [--ignore a,b] [--changes <file>]validate
Loads every recipe file (or every .json and .jsonl file in a directory; a .jsonl file holds one
recipe per line), parses it against its schema and binds
the input recipes to the output recipe, printing every problem with its JSON path. Exit code 1 if anything
fails.
opencraw validate recipes/run
Crawls the recipes at the given paths (files or directories; exactly one output recipe, any number of input recipes) and reports what happened.
| Flag | Meaning |
|---|---|
| --out <file> | Write records to this JSON Lines file. Without it, records print to stdout as JSON Lines. |
| --append | Keep what --out holds and add to it; each line carries _key. |
| --resume | Skip records --out already has (needs --append). |
| --only <id> | Run only this input recipe; repeat for more than one. |
| --dry-run | One record per input recipe, printed with the scope it was mapped from — for checking a recipe under construction without a full run. |
| --trace | Print the crawl trace to stderr. |
| --headed | Show the browser instead of running headless. |
| --parallel <n> | Run this many input recipes at once (default 1). Iterations inside a recipe follow its limits.concurrency. See §3.9 of the authoring guide. |
| --retries <n> | Tries per request that fails in passing (a dropped connection, a timeout, a 429, a 5xx), for recipes whose limits.retry says nothing. Default 3; 1 turns retrying off. See §7 of the authoring guide. |
| --profiles <dir> | Where the persistent browser profiles of session.browserProfile live (or OPENCRAW_PROFILES; default .opencraw/profiles). See access.md. |
| --host-delay <ms>, --host-concurrency <n> | Per-site politeness across every recipe: at least ms between two requests to one site, at most n in flight. They override the access config's throttle defaults. See access.md. |
| --plugins <file> (or --hooks) | A plugins module: named exports hooks (the recipes' hook steps and transforms), accessPlugins (for { kind: "plugin" } access profiles) and captchaSolvers. A module whose default export is { name: function } is read as hooks alone. OPENCRAW_PLUGINS / OPENCRAW_HOOKS when not given; probe reads it too. It runs as your code; load only files you trust. See §6 of the authoring guide. |
opencraw run recipes/ --out out/products.jsonl --trace
opencraw run recipes/movie.output.json recipes/tmdb.input.json --dry-rundiff
opencraw diff last-week.jsonl today.jsonl --key model,trim --ignore scrapedAt
opencraw run recipes/ --out today.jsonl --diff last-week.jsonl --changes changes.jsonlCompares two runs' records by key, not line by line, so a reordered file is no change:
1 added, 1 removed, 2 changed, 38 unchanged (41 records before, 41 now)
- 600e · La Prima
~ Avenger · Summit price.amount: 24950 → 23950; inStock: true → false
~ Pandina · Hybrid title: "Pandina" → "Pandina Hybrid"
+ Grande Panda · Icondiff <previous> <current>:--keynames the fields that identify a record (default: each line's_key, which--appendwrites);--ignoreleaves fields out;--changes <file>also writes every change as a JSON line ({ change, key, before?, after?, fields? }).run --diff <previous>compares the records this run emits with a previous file, taking the key and the fields to ignore (generated: now/uuid) from the output recipe. The file is read before the crawl, so it may be--outitself:run … --out prices.jsonl --diff prices.jsonlkeeps one file and reports each change. Not with--resume, which skips records.- A run that holds under half the records the previous one did gets a warning: when the source did not shrink, the site changed and the recipe quietly stopped finding everything.
probe
Fetches a page and reports where its data lives: JSON-LD blocks, inline JSON objects, .json URLs
referenced in the markup, script hosts, and links that look like an API. With --browser, it also renders
the page in a browser and lists the JSON responses it observes while the page settles — useful for
endpoints only a script fetches after load.
opencraw probe https://example.com/product/1
opencraw probe https://example.com/configurator --browserA PDF, a spreadsheet, a CSV or a presentation (a URL served with its content type, or a local .pdf,
.xlsx, .csv, .tsv or .pptx path) is read instead: probe lists its pages, sheets or slides, its first
rows, and every row that looks like a table header, with the selector a table extract needs. For a CSV it
also names the encoding and the delimiter it detected; for a presentation, its charts and the slides whose
text boxes look like a table. JSON, JSON Lines and YAML show their structure instead, and every list of
records with the jsonpath that walks it. An HTML page also lists its tables' header rows, and Markdown
(a .md URL is read as Markdown even when served as text/plain) its sections and front matter keys. XML
(a feed, a sitemap, an export) shows its namespaces with the namespaces line an xpath extract needs, the
elements that repeat (its records), and its structure; a sitemap says how many pages it lists. A Word document
(.docx) lists its sections and tables, like Markdown.
opencraw probe https://example.com/price-list.pdf
opencraw probe ./sheets/september.pdf
opencraw probe ./exports/listino.csv
opencraw probe ./exports/incentivi.xlsx
opencraw probe ./decks/incentivi.pptxUse it before writing an input recipe, to find the shape a site's data actually takes (§9 of the authoring guide).
Options common to run and probe
| Flag | Meaning |
|---|---|
| --browser-path <path> | A browser binary other than the one Playwright installed (or OPENCRAW_CHROMIUM). |
| --insecure-tls | Accept an intercepting proxy's certificate (or OPENCRAW_INSECURE_TLS=1). |
| --access <file> | An access config: proxy profiles, with credentials as {{env.NAME}} (or OPENCRAW_ACCESS). See access.md. |
| --access-profile <name> | The profile used by recipes that name none, overriding the config's default. probe uses it for its fetch. |
| --user-agent <ua> | The user agent to send. |
Building
Run nx build @opencraw/cli to build this project, nx test @opencraw/cli for its unit tests, and
nx run @opencraw/cli:e2e for the spawn-driven suite against the fixture shop @opencraw/core ships
(needs a browser: npm run playwright:install once, or OPENCRAW_CHROMIUM=/path/to/chrome).
