@zorinik/scraper
v0.1.0
Published
Recipe-driven web scraping engine: JSON recipes in, plain objects out.
Maintainers
Readme
@zorinik/scraper
A recipe-driven web scraping engine: each source is a JSON recipe, not code. The engine reads a recipe and some params, fetches the listing pages (with pagination) and each item's detail page, and returns plain objects whose keys come from the recipe.
- Plain HTTP (
fetch) or a real browser (Playwright, optional), per stage - A recursive, config-driven parser over HTML (Cheerio) or JSON: text, markdown, dates per locale, embedded JSON, regex, templates
- URL and in-page pagination, templated detail URLs, list-only sources
- Item validation, progress events, cancellation, a
limitfor dry runs - Field-level LLM extraction through an adapter you supply, immediate or deferred for batch APIs, with an optional OpenAI adapter
- A CLI to run recipes, save fixtures and snapshot-test recipes offline
The full recipe format is in docs/recipe.md.
Install
npm install @zorinik/scraperRequires Node 22 or later. Two peer dependencies are optional:
playwright, for recipes that use thebrowserstrategy:npm install playwright && npx playwright install chromiumopenai, for the@zorinik/scraper/openaiadapter
Without them, fetch recipes work as usual: they are only imported when used.
Quick start
A recipe, jobs.json:
{
"name": "example_jobs",
"version": 1,
"delay": 1,
"params": {
"query": { "type": "string", "required": true }
},
"strategies": {
"list": {
"strategy": { "name": "fetch", "config": { "url": "https://jobs.example.com/search?q={query}&page={page}" }, "output": "html" },
"pagination": { "start": 1, "max_pages": 5 },
"parser": {
"list": {
"type": "array",
"selector": ".result",
"items": {
"type": "object",
"properties": {
"url": { "type": "string", "selector": "a", "attribute": "href" },
"title": { "type": "string", "selector": "h2" }
}
}
}
}
},
"detail": {
"strategy": { "name": "fetch", "config": {}, "output": "html" },
"parser": {
"deadline": { "type": "date", "selector": ".deadline" },
"description": { "type": "string", "selector": "main", "format": "markdown" }
}
}
}
}From the shell:
npx recipe-scraper validate jobs.json
npx recipe-scraper run jobs.json --param query=physics --limit 5 --out items.jsonFrom code:
import {loadRecipe, runRecipe} from '@zorinik/scraper';
const recipe = await loadRecipe('jobs.json');
const {items, skipped, failed} = await runRecipe(recipe, {params: {query: 'physics'}, limit: 5});Programmatic API
Everything is exported from the package root, except the test helpers
(@zorinik/scraper/testing) and the OpenAI adapter (@zorinik/scraper/openai).
| Area | Main exports |
|---|---|
| Recipes | loadRecipe(path \| object), validateRecipe(data), RecipeValidationError (every issue, with its JSON path), resolveParams, resolveUrl, the Recipe types |
| Pipeline | runRecipe(recipe, options) → {items, skipped, failed, deferredLlm}; runDetail(recipe, item, options) to resume the detail stage of a stored item; validateItem |
| Fetchers | HttpFetcher, BrowserFetcher, RoutingFetcher (the default), FetchError, Throttle |
| Parser | parseHtml(parser, html), parseJson(parser, data), parse, extract, parseDate, toMarkdown |
| LLM | LlmAdapter, buildLlmRequest, applyLlmResults, enumSourceNames |
runRecipe options: params, fetcher, delay, signal, onProgress, limit,
concurrency, listOnly, now, dateNormalizers, llm, enums,
llmInstructions. See Pipeline for each one.
A host that stores listings first and fetches details later:
import {RoutingFetcher, Throttle, runDetail, runRecipe} from '@zorinik/scraper';
const {items} = await runRecipe(recipe, {params, listOnly: true});
// … store the items, then later, for each one:
const fetcher = new RoutingFetcher();
const throttle = new Throttle(recipe.delay);
const outcome = await runDetail(recipe, item, {fetcher, throttle});
await fetcher.close();LLM-filled fields, deferred for a batch API:
import {applyLlmResults, runRecipe, validateItem} from '@zorinik/scraper';
import {parseResponse, responsesBody} from '@zorinik/scraper/openai';
const {items, deferredLlm} = await runRecipe(recipe, {llm: 'defer'});
// send responsesBody(request, model) for each deferred request, then read the answers back:
const answers = batchLines.map(line => ({itemId: line.custom_id, output: parseResponse(line.response.body)}));
const kept = applyLlmResults(items, answers).filter(item => validateItem(recipe, item, 'post_llm') === null);Or immediately, per item: runRecipe(recipe, {llm: new OpenAiAdapter({model: 'gpt-4.1-mini'})}).
See Field-level LLM extraction.
CLI
The package installs a recipe-scraper binary. Relative paths are resolved from the
working directory. Exit codes: 0 success, 1 failure, 2 usage error.
| Command | Does |
|---|---|
| validate <recipe.json>... | Validates each recipe and prints every issue with its JSON path |
| run <recipe.json> [--param name=value]... [--list-only] [--limit N] [--delay S] [--concurrency N] [--out file.json] [--quiet] | Runs the recipe live. The result JSON goes to --out or stdout; progress and a summary go to stderr (silenced by --quiet). Ctrl+C cancels |
| fixture <recipe.json> --out <dir> [--param name=value]... [--limit N] [--list-only] [--delay S] [--force] [--quiet] | Runs the recipe live (by default with --limit 3, so the listing and 3 detail pages) and saves every response into <dir> |
| test <recipe.json> --fixtures <dir> [--update] | Runs the recipe offline against the saved responses and compares the result with <dir>/snapshot.json; --update writes it |
--param is repeatable. Values are strings, converted by the recipe's param
declarations (number, boolean, enum). recipe-scraper <command> --help prints
the options of a command.
Testing recipes
The fixture workflow keeps recipe tests offline and deterministic:
npx recipe-scraper fixture recipes/jobs.json --param query=physics --out recipes/fixtures/jobs
npx recipe-scraper test recipes/jobs.json --fixtures recipes/fixtures/jobs --update # writes snapshot.json
npx recipe-scraper test recipes/jobs.json --fixtures recipes/fixtures/jobs # comparesA fixture directory holds the saved responses (001.html, 002.json, …), a
fixtures.json manifest (the params, limit and moment of the recording, and each
request with its response file or its error), and snapshot.json. The replay uses
the recorded params and moment, so date tokens and relative dates come out the same,
and it runs with no delay. Review snapshot.json like code, and commit the whole
directory. When the site changes, record again with --force.
The same functions are available to a test runner, from @zorinik/scraper/testing:
import {expect, it} from 'vitest';
import {loadRecipe} from '@zorinik/scraper';
import {testFixtures} from '@zorinik/scraper/testing';
it('jobs recipe', async () => {
const recipe = await loadRecipe('recipes/jobs.json');
const outcome = await testFixtures(recipe, 'recipes/fixtures/jobs');
expect(outcome.diff).toEqual([]);
expect(outcome.status).toBe('pass');
});recordFixtures and replayFixtures are exported too. For hand-written tests,
FakeFetcher serves canned responses by URL: see
Testing with FakeFetcher.
Porting a Python data-extractor config
The format keeps the Python configs' shape, so a config converts mechanically:
- Unwrap
{"sources": [...]}: one recipe per file. - Add
"version": 1, and alocale(e.g."it") for itsdateproperties. - Rename
validation.eventtovalidation.item, andhas_future_or_today_datetodate_not_past. - Drop the strategy-level
llmblock (field-levelllmblocks are kept). - Declare a
paramsentry for each{token}in the URLs that is not reserved ({page},{date_from},{date_to},{date_future}). - Replace
\1backreferences inregex.replacewith$1.
Then run recipe-scraper validate: it lists every remaining problem with its JSON
path. The behavioral differences (dates per locale, stricter validation, merge rules)
are listed in Differences from the Python format.
License
MIT
