npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@volter/twin-azureformrecognizer

v0.1.37

Published

Local Azure AI Document Intelligence (formerly Form Recognizer) twin — async analyze → Operation-Location poll lifecycle, prebuilt document models, layout/key-value/table extraction — built on @volter/world-core.

Readme

@volter/twin-azureformrecognizer

Local Azure AI Document Intelligence twin (the service formerly called Azure Form Recognizer) — the async analyze lifecycle (POST /documentintelligence/documentModels/{modelId}:analyze → **202

  • Operation-Location** → poll GET .../analyzeResults/{resultId}, notStarted → running → succeeded/failed), the prebuilt model catalog + detail, resource info, and analyze-result deletion, for api-version 2024-11-30 (v4.0 GA).

The real @azure-rest/ai-document-intelligence npm SDK works unmodified — including its getLongRunningPoller, which reads the twin's Operation-Location header and follows it home (proven in src/azureformrecognizer-sdk.integration.test.ts).

bun packages/twin/azureformrecognizer/src/cli.ts                        # serve
bun packages/twin/azureformrecognizer/src/cli.ts serve --scenario s.json
bun packages/twin/azureformrecognizer/src/cli.ts serve --api-key <key>  # pin the resource key
import DocumentIntelligence, { getLongRunningPoller } from '@azure-rest/ai-document-intelligence';
const client = DocumentIntelligence(twinUrl, { key: 'twin-local-key' }, { allowInsecureConnection: true });
const initial = await client.path('/documentModels/{modelId}:analyze', 'prebuilt-invoice')
  .post({ contentType: 'application/json', body: { urlSource } });
const { body } = await getLongRunningPoller(client, initial).pollUntilDone();
body.analyzeResult.documents[0].fields.InvoiceTotal.valueCurrency; // { amount, currencyCode, currencySymbol }

Injection. The endpoint is per-resource (https://<your-resource>.cognitiveservices.azure.com), so the injector matches the domain by suffix + path: this pack claims ^/(documentintelligence|formrecognizer)/ on it, and the azure pack's Azure OpenAI area claims ^/openai/. That domain is shared by every Azure AI service, so neither pack may claim it bare — a bare claim here would swallow Azure OpenAI traffic into the document-analysis twin and answer it with a plausible-but-wrong 404. The two patterns are disjoint, so routing does not depend on table order, and a Cognitive Services path belonging to neither (Speech, Language, Vision) resolves to no vendor and gets a loud refusal rather than a wrong answer. Set AZUREFORMRECOGNIZER_TWIN_URL and an unmodified app keeps constructing its client exactly as in production. (scripts/vendor-hosts.test.ts pins both directions; the claim lives on this pack's pack descriptor, not in inject.cjs's legacy hand table.)

The store door: GET /twin/store/analyses

Document Intelligence has no endpoint that lists analyses. The only vendor read of a submitted analysis is the per-operation poll (GET /documentModels/{modelId}/analyzeResults/{id}), and that poll deliberately advances the modeled lifecycle one step per GET — notStarted on submit, running on the first poll, succeeded on the next — so an unmodified SDK poller terminates in two polls with no wall-clock dependence. Faithful, but it means nothing in the vendor surface can observe an analysis without moving it.

GET /twin/store/analyses is the twin's own named, read-only, deterministic projection over stored state (the twin programming model's store door, GET /twin names it under stores). It folds the kernel log into the analyses the resource holds — ids, model, source, status, timestamps — with internal bookkeeping and purged rows omitted. It is a twin door, not vendor surface, so it needs no subscription key and is deliberately absent from the capability manifest: counting scaffolding Azure does not have would pad the coverage denominator. It is also what makes this pack's determinism checkable at the resource level — the gate's R9 replay reads it back on two fresh roots.

The extraction stub (honest by design)

OCR and the document models are real hosted compute — the twin has no OCR engine and no trained models. Extracted content is a clearly-labeled deterministic stub ([twin-stub:document] … (seed <hash>), seeded from the document URL or, for inline/binary documents, the bytes' sha256 — same document, same result, always) while the AnalyzeResult envelope is faithful:

  • the 202 + Operation-Location submit and the notStarted → running → succeeded/failed poll lifecycle, with createdDateTime / lastUpdatedDateTime;
  • the content/span/offset algebra — for every element in the result (word, line, paragraph, table cell, key-value key/value, document field, selection mark), content.substr(span.offset, span.length) === element.content exactly, which is what a real consumer relies on to highlight a value back on the page;
  • page geometry (1-based pageNumber, angle, width/height, unit: "inch", 4-point 8-number polygons, per-word confidence), paragraph roles (layout/custom only), rowIndex/columnIndex-addressed table cells with a columnHeader row, selection marks rendered inline as :selected:;
  • the typed field union on documents[].fields — valueString / valueDate / valueNumber / valueCurrency ({ amount, currencySymbol, currencyCode }) / valueArray of valueObject line items — with confidences and bounding regions;
  • the vendor's { error: { code, message, innererror } } envelope on every 4xx.

Model-shape fidelity follows the vendor's own "Model analysis features" matrix: prebuilt-read returns text and unroled paragraphs only (the matrix ticks Paragraphs for read but leaves Paragraph roles blank); prebuilt-layout adds paragraph roles, selection marks and tables; prebuilt-invoice and prebuilt-receipt return typed documents[].fields and no paragraphs. Key-value pairs are an optional add-on on both models that support them — absent by default from prebuilt-layout and prebuilt-invoice, emitted only with ?features=keyValuePairs. ?features= accepts exactly the seven v4.0 DocumentAnalysisFeature values (ocrHighResolution, languages, barcodes, formulas, keyValuePairs, styleFont, queryFields) and 400s anything else — including searchablePdf, which is an ?output= option, not a feature.

prebuilt-document is deliberately NOT served. Microsoft deprecated the general document model for v4.0 — "use the Layout model with the optional query string parameter features=keyValuePairs" — so under api-version=2024-11-30 the real service answers 404 NotFound / ModelNotFound, and so does this twin (the inner message names the replacement). Serving it would be a fabricated success. Key-value extraction is available exactly where the vendor puts it: prebuilt-layout?features=keyValuePairs. The retired v3.1 /formrecognizer/... surface (where prebuilt-document does live) is enumerated as a todo, not silently emulated.

Scenario scripting (src/azureformrecognizer-scenario.ts, same convention as the anthropic/assemblyai/deepgram packs): a JSON file { rules: [{ match: { urlSourceIncludes | urlSourceEquals | modelId | nthAnalysis }, analyze: { fields | lines | error } }] } scripts the exact extracted field values — or a vendor-style failed operation (HTTP 200 with status: "failed" and an error object, which is how Document Intelligence surfaces an unfetchable/corrupt document) — per document. Rules are strictly validated at load (unknown keys / mistyped values fail startup loudly). Wire with createAzureFormRecognizerTwinServer({ scenarioPath }), TWIN_AZUREFORMRECOGNIZER_SCENARIO, or --scenario. Scenario support is twin-only scaffolding and is deliberately not counted in the capability manifest; it is gated by src/azureformrecognizer-scenario.test.ts.

Error-string fidelity note: HTTP status codes, the top-level error.code values (InvalidRequest, InvalidArgument, NotFound, MethodNotAllowed, UnsupportedMediaType) and the innererror.code values (ModelNotFound, OperationNotFound, ParameterMissing, InvalidContentSourceFormat, InvalidParameter, InvalidContent, ModelReadOnly, NotSupportedApiVersion) are grounded in Microsoft's published error reference, which is also where the pairings come from (InvalidParameter / ParameterMissing sit under InvalidArgument). The exact prose of inner messages — and the pairing chosen for an unparseable request body, which the reference does not enumerate — is an approximation (not verifiable offline), tracked as azureformrecognizer.errors.exact_message_parity.

Modeled vendor semantics worth knowing:

  • Auth is the APIM gateway's, not the service's: a missing/wrong Ocp-Apim-Subscription-Key returns 401 with { error: { code: "401", message } } (the code is the string status, which is what Azure API Management emits). Start the server with apiKey / --api-key to pin a resource key; without it any non-empty key is accepted and only a missing key 401s.
  • api-version is required on every operation — omitting it is 400 InvalidArgument / ParameterMissing (a missing required parameter); asking for a version this resource doesn't serve is 400 InvalidRequest / NotSupportedApiVersion. Only 2024-11-30 is served.
  • The lifecycle advances one step per poll (deterministic, no wall clock): the first GET observes running, the second reaches the terminal state. So GET .../analyzeResults/{id} is a state-advancing read (a deliberate trade-off for determinism) — and in --read-only mode reads never mutate, so a pending operation stays pending forever there.
  • Each :analyze call mints a new resultId; DELETE .../analyzeResults/{id} marks it for deletion (204, then 404 on read) and the id is never re-issued, so re-analyzing the same document after a delete can't resurrect the old result.
  • Prebuilts are service-owned: DELETE /documentModels/prebuilt-* is 400 InvalidRequest / ModelReadOnly, and the connector never folds prebuilts into the observed log (only custom models are twin state).
  • Prebuilt createdDateTime is reported as 2024-11-30T00:00:00Z (a twin-chosen stable value — the real service reports its own per-region build date). For a model the connector pulled, any field the pull did not observe is omitted, never defaulted — an abbreviated pull stays visibly abbreviated.
  • The twin's stub document is one physical page. ?pages= is honoured for page 1; a selector naming a page the document doesn't have is 400 InvalidArgument / InvalidParameter rather than a synthesised empty page. Real multi-page extraction is a todo (azureformrecognizer.extraction.multi_page_documents).
  • ?locale= is a recognition hint, not a detection. features=languages always reports the detected locale (en for the English stub content); passing ?locale=fr does not change it.
  • A locally deleted custom model stays deleted across a re-pull that still observes it upstream in the local-delete conformance case. Branch ancestry, fetch and rebase follow the shared model. The case is azureformrecognizer.models.delete_survives_repull.

Coverage

The capability manifest (src/azureformrecognizer-capabilities.ts) is an honest partial denominator for the Document Intelligence v4.0 API, authored top-down from the Microsoft Learn REST reference (Document Models / Document Classifiers / Miscellaneous operation groups), the model-overview feature matrix and add-on list, and the published error reference — cross-checked against the official @azure-rest/ai-document-intelligence v1.1.0 SDK's routes and generated types. Coverage is partial and will read low as the denominator grows; see the manifest's todo entries for what is enumerated but not yet modeled.

Modeled: the analyze → poll lifecycle (URL / base64 / raw-binary document sources), the Operation-Location contract, deterministic per-document results, analyze-result deletion, request validation (missing/invalid source, malformed JSON, unsupported media type, bad features / pages / stringIndexType), the AnalyzeResult envelope + span algebra, pages/words/lines/ selection marks, paragraph roles, tables, features=keyValuePairs and features=languages, typed invoice + receipt fields (including line-item arrays and internally consistent totals), the model catalog + detail + field schemas, custom-model delete + analyze (and delete surviving a re-pull), /info, subscription-key and api-version parity, read-only mode, unmodeled-route 404s, and the connector's pull of custom models.

Not yet modeled (a sample of the todo denominator): custom model build / compose / copy and their build operations, document classifiers (build/list/get/classify/split/copy), batch analysis (:analyzeBatch), the remaining ~26 prebuilt field schemas (idDocument, tax.us., mortgage.us., contract, healthInsuranceCard.us, bankStatement, check.us, payStub.us, creditCard, marriageCertificate.us), the other add-on features (queryFields, barcodes, formulas, styleFont, ocrHighResolution), outputContentFormat=markdown, figures/sections, genuinely multi-page documents, Entra ID (AAD) auth, list pagination (nextLink), rate limits (429), the 24-hour result expiry, real model-build /operations (the twin serves only the empty-resource answer), and the v3.1 /formrecognizer surface.

No UI mirror

The API is the product. A Document Intelligence integration lives in backend code: you POST a document to :analyze and poll Operation-Location (exactly what the official SDK's getLongRunningPoller does), then read analyzeResult. Document Intelligence Studio is a labeling/inspection console used to build and try custom models — the analyze traffic an integrator actually cares about never passes through it — so this pack ships no mirror and has no UI capabilities.

Connector (D7)

syncAzureFormRecognizerFromReal(client, { root }) pulls a real resource's custom document models over an injected client (a fake in tests; a thin wrapper over the real REST client in prod — the pack imports no SDK and holds no key) and folds them into the event log via the kernel observation path, so a re-pull of identical state appends zero deltas. Pulled models then appear in the twin's own GET /documentModels and can be analyzed against locally. Prebuilts are filtered out (service-owned constants, not observed state). Analyze results are not pullable — Document Intelligence exposes no list endpoint for them — which is filed as azureformrecognizer.connector.pull_analyze_results and recorded in pull-audit.json.