@volter/twin-firecrawl
v0.1.35
Published
Local Firecrawl twin (scrape, crawl, map — the v1 web-scraping API) built on @volter/world-core.
Readme
@volter/twin-firecrawl
A local Firecrawl twin — offline, deterministic replicas of the Firecrawl
v1 web-scraping API (api.firecrawl.dev) in vendor-faithful response shapes,
with the real Authorization: Bearer fc-… gate and
400/401/402/404/422/429 error semantics. The unmodified Firecrawl SDK
(@mendable/firecrawl-js@^1) pointed at it gets Firecrawl-correct responses.
world-firecrawl serve [--port N] [--root DIR] [--read-only]
world-firecrawl conformance [--root DIR]import FirecrawlApp from '@mendable/firecrawl-js';
const app = new FirecrawlApp({ apiKey: 'fc-test-key', apiUrl: process.env.FIRECRAWL_TWIN_URL });
const doc = await app.scrapeUrl('https://example.com/docs', { formats: ['markdown', 'links'] });Coverage
The capability manifest (src/firecrawl-capabilities.ts) is an honest
partial denominator across the whole Firecrawl v1 surface — scrape, crawl,
map, batch scrape, search, extract, LLMs.txt, deep research, webhooks and the
team/usage endpoints. The scrape/crawl/map core is done (each with a failable,
deterministic verify asserting response values plus the vendor's negative
4xx); the long tail is honest todo. Total surface is >= 50 entries; this
is the live denominator, not a self-portrait, and it grows as more of the API is
modeled.
How honest is the coverage number? The headline counts every done, and
about a third of them are the twin checking itself — the state.* and
honesty.* families plus scrape.appends_state and the connector caps. Roughly
19 are vendor features (scrape shapes, map, the crawl lifecycle) and 16 are the
vendor's negative 4xx/auth paths. Read the headline as "how much of the modeled
contract is proven", not "how much of Firecrawl exists here" — counted purely as
vendor features against the in-scope denominator it is closer to 19%, the
expected shape for a new pack.
Grounding. No single downloadable first-party OpenAPI document was found for
v1, so spec-sources.json records kind: none. Two first-party sources
ground the surface anyway: Firecrawl's own official v1 client
(@mendable/[email protected], fetched read-only via npm pack and read
directly — every /v1/* path it builds, the FirecrawlDocument /
CrawlStatusResponse / MapResponse / ErrorResponse types, the Bearer header
form, the handleError branches), and the live OpenAPI-driven v1 reference pages
at docs.firecrawl.dev/api-reference/v1-endpoint/*, which settle the wire details
the SDK does not: metadata carries sourceURL and no url (that is a v2
field), a missing job answers {"error":"Crawl job not found."}, cancel answers
{"status":"cancelled"}, and the 402/429 bodies are quoted verbatim. Rate-limit
numbers come from docs.firecrawl.dev/rate-limits. The one point the vendor
genuinely does not publish — the exact 401 error string — is annotated
doc-UNVERIFIED in firecrawl-auth.ts and is asserted only as a status plus an
envelope shape, never as a vendor fact.
Two error envelopes, both real. Where the vendor publishes an example body
the twin serves it exactly: the auth/credit/rate family and the missing-crawl-job
404 carry { error } with no success field, and the 429 also carries
Retry-After. The undocumented responses (validation 400s, the twin's own
unmodeled-endpoint 404) use { success: false, error, details? }, the shape the
v1 SDK's ErrorResponse type declares.
Known hole in "the unmodified SDK works". The v1 SDK's checkCrawlErrors
issues a DELETE to /v1/crawl/{id}/errors, while the vendor documents that
route as GET. The twin follows the docs, so that one SDK method 404s against
it. Aliasing DELETE would mean reproducing an SDK bug as if it were vendor
behavior, so it is recorded here instead. Every other v1 SDK method used by the
fidelity test works unmodified.
SDK version pin. The fidelity test (src/firecrawl-sdk.integration.test.ts)
drives the unmodified SDK, pinned to @mendable/firecrawl-js@^1.29.3,
through its own public apiUrl option. The pin is the SDK's public surface, not
a patch: the 1.x line is the major that speaks /v1, while 4.x moved to
/v2. This is the same "un-patched, not un-configured" allowance the notion pack
records for Notion-Version.
Deterministic page content, NOT the real web. A local twin cannot fetch the
live internet. It returns faithful document shapes with stable, hash-derived
VALUES per URL: a given URL always maps to the same markdown/HTML/links/metadata,
in every checkout, offline. Every scrape is folded into the @volter/world-core
append-only log as a document, so a re-scrape, a crawl and a seeded override all
read the same projection.
Modeled (done):
- Auth gate. Every v1 endpoint gates on
Authorization: Bearer fc-…. Missing/blank/fc--less token → 401 in Firecrawl's{ success:false, error }envelope; sentinel keysfc-out-of-credits→ 402 andfc-rate-limited→ 429. POST /v2/scrape— Firecrawl's current API (docs.firecrawl.dev/api-reference/endpoint/scrape), which@mendable/[email protected]and LibreChat's web search call by default: the same pages and modelled formats as v1, a format given as a string or an object{ type }, andmetadata.urlandcontentTypebesidesourceURL. v2's other formats (summary,screenshot,json, …) answer 422 not modelled; an unknown one 400, as is a bare string for the formats the docs give only as objects (json,changeTracking,question,highlights). Its refusals reuse v1's error body (v2 documents a stringcodebesideerror, whose values are not published; not probed). v2's other endpoints (map, crawl, batch scrape, search, extract) aretodo.POST /v1/scrape— the real document shape ({ success, data:{ markdown, html?, rawHtml?, links?, metadata } }) with a fullmetadatablock (title,description,language,og*,sourceURL,statusCode). Honorsformats(markdowndefault;html,rawHtml,links), canonicalizes the URL (trailing slash / fragment), and is deterministic per URL. Missing url → 400 with a zod-shapeddetails; a format outside the v1 enum → 400; a real-but-unmodeled format (screenshot,json,extract,changeTracking) → 422, never a document silently missing the field.POST /v1/crawl+GET/DELETE /v1/crawl/{id}+GET /v1/crawl/{id}/errors— start returns{ success, id, url }with a UUID-shaped id; status returns{ status, total, completed, creditsUsed, expiresAt, data[] }and advances one page per poll, deterministically, fromscrapingtocompletedwith no wall-clock dependency. HonorslimitandscrapeOptions.formats; two crawls of the same URL get distinct ids and advance independently; cancel flips the job tocancelled; an unknown id 404s on all three routes.POST /v1/map—{ success, links[] }over the host's page set unioned with every page already observed locally; honorssearchandlimit.- State + honesty. Seeded pages (including a non-200
statusCode) are served verbatim by a later scrape and by a crawl — one projection, not two; a re-seed after a scrape wins (the stored copy is refreshed, not stuck).--read-onlyforbids crawl starts, cancels and seeds (405), serves reads, appends nothing, and never advances a job. An unmodeled endpoint 404s in the vendor envelope. - Connector. Injected-client pull of scraped documents
(
pullFirecrawlDocuments/syncFirecrawlFromReal), idempotent on the canonical URL, and storing nothing on a failed, unsuccessful or empty response — an empty document would otherwise be stored and permanently shadow the deterministic content for that URL.
A note on --read-only. For most twins read-only means "serve exactly what was
observed". Here a scrape of a URL the twin has never seen still returns 200 with
freshly synthesized content — it just refuses to record it. That is deliberate (a
scrape is a read, and refusing every unseen URL would make read-only mode useless
for a scraper), but it means read-only is not a pure mirror of prior observations.
No connector push. Firecrawl v1 has no write endpoint at all — every verb
submits work, none stores a document — so there is no vendor path to push local
state back to. pushPendingFirecrawlActions exists for local convergence against
the twin's own seed route, and liveFirecrawlExecute refuses to send that route to
the vendor rather than faking a round-trip.
- Rate budget.
liveFirecrawlExecuteis the pack's one construction site for a real network-calling client, with the kernel's fail-closed budget charged before every request. Sized to Firecrawl's documented free plan: 10 units/60s at 1 per call = 10 scrapes/minute, withPOST /v1/crawlpriced at 5 = 2 crawls/minute.
Not yet modeled (honest todo): most scrape options (onlyMainContent,
includeTags/excludeTags, headers, waitFor, timeout, mobile,
location, blockAds, parsePDF, caching via maxAge/storeInCache,
changeTracking, zeroDataRetention, full OpenGraph/dcterms metadata
passthrough); crawl path/depth/link-following rules, sitemap options, pacing,
status pagination (next/skip/limit), GET /v1/crawl/active, idempotency
keys, the failed terminal status, populated crawl errors and the WebSocket
watcher; map includeSubdomains/sitemapOnly/useIndex; the entire batch
scrape family; the /v1/search, /v1/extract, /v1/llmstxt and
/v1/deep-research job envelopes; crawl webhooks; and the /v1/team/* +
/v1/concurrency-check usage endpoints.
No UI mirror
Firecrawl is an API-first vendor: someone doing its core job writes code
against POST /v1/scrape, they do not open a browser. The firecrawl.dev
dashboard is incidental tooling for API keys, credit usage and job logs — a
state/billing console, not where the work happens — so this pack ships no
mirror and has no UI capabilities. Coverage is API + connector.
Architecture
State lives in the @volter/world-core event/action log (document, crawl, audit
subjects); there is no ad-hoc store. The serve path makes no real network calls;
connector functions accept injected executors for real Firecrawl I/O and are not
used by the local handler. POST /v1/_twin/page is a clearly-namespaced
twin-only seed route (Firecrawl's real v1 API has no endpoint that stores a
page), deliberately kept out of the capability manifest as scaffolding rather than
vendor surface — and liveFirecrawlExecute refuses to send a /v1/_twin/* path to
the real vendor.
