@volter/twin-azureformrecognizer
v0.1.37
Published
Local Azure AI Document Intelligence (formerly Form Recognizer) twin — async analyze → Operation-Location poll lifecycle, prebuilt document models, layout/key-value/table extraction — built on @volter/world-core.
Readme
@volter/twin-azureformrecognizer
Local Azure AI Document Intelligence twin (the service formerly called Azure Form Recognizer) —
the async analyze lifecycle (POST /documentintelligence/documentModels/{modelId}:analyze → **202
Operation-Location** → pollGET .../analyzeResults/{resultId},notStarted → running → succeeded/failed), the prebuilt model catalog + detail, resource info, and analyze-result deletion, for api-version2024-11-30(v4.0 GA).
The real @azure-rest/ai-document-intelligence npm SDK works unmodified — including its
getLongRunningPoller, which reads the twin's Operation-Location header and follows it home
(proven in src/azureformrecognizer-sdk.integration.test.ts).
bun packages/twin/azureformrecognizer/src/cli.ts # serve
bun packages/twin/azureformrecognizer/src/cli.ts serve --scenario s.json
bun packages/twin/azureformrecognizer/src/cli.ts serve --api-key <key> # pin the resource keyimport DocumentIntelligence, { getLongRunningPoller } from '@azure-rest/ai-document-intelligence';
const client = DocumentIntelligence(twinUrl, { key: 'twin-local-key' }, { allowInsecureConnection: true });
const initial = await client.path('/documentModels/{modelId}:analyze', 'prebuilt-invoice')
.post({ contentType: 'application/json', body: { urlSource } });
const { body } = await getLongRunningPoller(client, initial).pollUntilDone();
body.analyzeResult.documents[0].fields.InvoiceTotal.valueCurrency; // { amount, currencyCode, currencySymbol }Injection. The endpoint is per-resource
(https://<your-resource>.cognitiveservices.azure.com), so the injector matches the domain by
suffix + path: this pack claims ^/(documentintelligence|formrecognizer)/ on it, and the
azure pack's Azure OpenAI area claims ^/openai/. That domain is shared by every Azure AI
service, so neither pack may claim it bare — a bare claim here would swallow Azure OpenAI traffic
into the document-analysis twin and answer it with a plausible-but-wrong 404. The two patterns
are disjoint, so routing does not depend on table order, and a Cognitive Services path belonging
to neither (Speech, Language, Vision) resolves to no vendor and gets a loud refusal rather
than a wrong answer. Set AZUREFORMRECOGNIZER_TWIN_URL and an unmodified app keeps constructing
its client exactly as in production. (scripts/vendor-hosts.test.ts pins both directions; the
claim lives on this pack's pack descriptor, not in inject.cjs's legacy hand table.)
The store door: GET /twin/store/analyses
Document Intelligence has no endpoint that lists analyses. The only vendor read of a
submitted analysis is the per-operation poll (GET /documentModels/{modelId}/analyzeResults/{id}),
and that poll deliberately advances the modeled lifecycle one step per GET — notStarted on
submit, running on the first poll, succeeded on the next — so an unmodified SDK poller
terminates in two polls with no wall-clock dependence. Faithful, but it means nothing in the
vendor surface can observe an analysis without moving it.
GET /twin/store/analyses is the twin's own named, read-only, deterministic projection over
stored state (the twin programming model's store door, GET /twin names it under stores). It
folds the kernel log into the analyses the resource holds — ids, model, source, status,
timestamps — with internal bookkeeping and purged rows omitted. It is a twin door, not vendor
surface, so it needs no subscription key and is deliberately absent from the capability
manifest: counting scaffolding Azure does not have would pad the coverage denominator. It is
also what makes this pack's determinism checkable at the resource level — the gate's R9 replay
reads it back on two fresh roots.
The extraction stub (honest by design)
OCR and the document models are real hosted compute — the twin has no OCR engine and no trained
models. Extracted content is a clearly-labeled deterministic stub ([twin-stub:document] …
(seed <hash>), seeded from the document URL or, for inline/binary documents, the bytes' sha256 —
same document, same result, always) while the AnalyzeResult envelope is faithful:
- the 202 +
Operation-Locationsubmit and thenotStarted → running → succeeded/failedpoll lifecycle, withcreatedDateTime/lastUpdatedDateTime; - the content/span/offset algebra — for every element in the result (word, line, paragraph,
table cell, key-value key/value, document field, selection mark),
content.substr(span.offset, span.length) === element.contentexactly, which is what a real consumer relies on to highlight a value back on the page; - page geometry (1-based
pageNumber,angle,width/height,unit: "inch", 4-point 8-numberpolygons, per-word confidence), paragraphroles (layout/custom only),rowIndex/columnIndex-addressed table cells with acolumnHeaderrow, selection marks rendered inline as:selected:; - the typed field union on
documents[].fields—valueString/valueDate/valueNumber/valueCurrency({ amount, currencySymbol, currencyCode }) /valueArrayofvalueObjectline items — with confidences and bounding regions; - the vendor's
{ error: { code, message, innererror } }envelope on every 4xx.
Model-shape fidelity follows the vendor's own "Model analysis features" matrix:
prebuilt-read returns text and unroled paragraphs only (the matrix ticks Paragraphs for read
but leaves Paragraph roles blank); prebuilt-layout adds paragraph roles, selection marks and
tables; prebuilt-invoice and prebuilt-receipt return typed documents[].fields and no
paragraphs. Key-value pairs are an optional add-on on both models that support them — absent by
default from prebuilt-layout and prebuilt-invoice, emitted only with
?features=keyValuePairs. ?features= accepts exactly the seven v4.0 DocumentAnalysisFeature
values (ocrHighResolution, languages, barcodes, formulas, keyValuePairs, styleFont,
queryFields) and 400s anything else — including searchablePdf, which is an ?output= option,
not a feature.
prebuilt-document is deliberately NOT served. Microsoft deprecated the general document model
for v4.0 — "use the Layout model with the optional query string parameter features=keyValuePairs"
— so under api-version=2024-11-30 the real service answers 404 NotFound / ModelNotFound,
and so does this twin (the inner message names the replacement). Serving it would be a fabricated
success. Key-value extraction is available exactly where the vendor puts it:
prebuilt-layout?features=keyValuePairs. The retired v3.1 /formrecognizer/... surface (where
prebuilt-document does live) is enumerated as a todo, not silently emulated.
Scenario scripting (src/azureformrecognizer-scenario.ts, same convention as the
anthropic/assemblyai/deepgram packs): a JSON file
{ rules: [{ match: { urlSourceIncludes | urlSourceEquals | modelId | nthAnalysis }, analyze: { fields | lines | error } }] }
scripts the exact extracted field values — or a vendor-style failed operation (HTTP 200 with
status: "failed" and an error object, which is how Document Intelligence surfaces an
unfetchable/corrupt document) — per document. Rules are strictly validated at load (unknown keys /
mistyped values fail startup loudly). Wire with
createAzureFormRecognizerTwinServer({ scenarioPath }), TWIN_AZUREFORMRECOGNIZER_SCENARIO, or
--scenario. Scenario support is twin-only scaffolding and is deliberately not counted in the
capability manifest; it is gated by src/azureformrecognizer-scenario.test.ts.
Error-string fidelity note: HTTP status codes, the top-level error.code values
(InvalidRequest, InvalidArgument, NotFound, MethodNotAllowed, UnsupportedMediaType) and
the innererror.code values (ModelNotFound, OperationNotFound, ParameterMissing,
InvalidContentSourceFormat, InvalidParameter, InvalidContent, ModelReadOnly,
NotSupportedApiVersion) are grounded in Microsoft's published error reference, which is also
where the pairings come from (InvalidParameter / ParameterMissing sit under
InvalidArgument). The exact prose of inner messages — and the pairing chosen for an
unparseable request body, which the reference does not enumerate — is an approximation (not
verifiable offline), tracked as azureformrecognizer.errors.exact_message_parity.
Modeled vendor semantics worth knowing:
- Auth is the APIM gateway's, not the service's: a missing/wrong
Ocp-Apim-Subscription-Keyreturns401with{ error: { code: "401", message } }(the code is the string status, which is what Azure API Management emits). Start the server withapiKey/--api-keyto pin a resource key; without it any non-empty key is accepted and only a missing key 401s. api-versionis required on every operation — omitting it is400 InvalidArgument / ParameterMissing(a missing required parameter); asking for a version this resource doesn't serve is400 InvalidRequest / NotSupportedApiVersion. Only2024-11-30is served.- The lifecycle advances one step per poll (deterministic, no wall clock): the first GET
observes
running, the second reaches the terminal state. SoGET .../analyzeResults/{id}is a state-advancing read (a deliberate trade-off for determinism) — and in--read-onlymode reads never mutate, so a pending operation stays pending forever there. - Each
:analyzecall mints a newresultId;DELETE .../analyzeResults/{id}marks it for deletion (204, then 404 on read) and the id is never re-issued, so re-analyzing the same document after a delete can't resurrect the old result. - Prebuilts are service-owned:
DELETE /documentModels/prebuilt-*is400 InvalidRequest / ModelReadOnly, and the connector never folds prebuilts into the observed log (only custom models are twin state). - Prebuilt
createdDateTimeis reported as2024-11-30T00:00:00Z(a twin-chosen stable value — the real service reports its own per-region build date). For a model the connector pulled, any field the pull did not observe is omitted, never defaulted — an abbreviated pull stays visibly abbreviated. - The twin's stub document is one physical page.
?pages=is honoured for page 1; a selector naming a page the document doesn't have is400 InvalidArgument / InvalidParameterrather than a synthesised empty page. Real multi-page extraction is atodo(azureformrecognizer.extraction.multi_page_documents). ?locale=is a recognition hint, not a detection.features=languagesalways reports the detected locale (enfor the English stub content); passing?locale=frdoes not change it.- A locally deleted custom model stays deleted across a re-pull that still observes it
upstream in the local-delete conformance case. Branch ancestry, fetch and rebase follow
the shared model. The case is
azureformrecognizer.models.delete_survives_repull.
Coverage
The capability manifest (src/azureformrecognizer-capabilities.ts) is an honest partial
denominator for the Document Intelligence v4.0 API, authored top-down from the Microsoft Learn
REST reference (Document Models / Document Classifiers / Miscellaneous operation groups), the
model-overview feature matrix and add-on list, and the published error reference — cross-checked
against the official @azure-rest/ai-document-intelligence v1.1.0 SDK's routes and generated
types. Coverage is partial and will read low as the denominator grows; see the manifest's todo
entries for what is enumerated but not yet modeled.
Modeled: the analyze → poll lifecycle (URL / base64 / raw-binary document sources), the
Operation-Location contract, deterministic per-document results, analyze-result deletion, request
validation (missing/invalid source, malformed JSON, unsupported media type, bad features /
pages / stringIndexType), the AnalyzeResult envelope + span algebra, pages/words/lines/
selection marks, paragraph roles, tables, features=keyValuePairs and features=languages,
typed invoice + receipt fields (including line-item arrays and internally consistent totals), the
model catalog + detail + field schemas, custom-model delete + analyze (and delete surviving a
re-pull), /info, subscription-key and api-version parity, read-only mode, unmodeled-route 404s,
and the connector's pull of custom models.
Not yet modeled (a sample of the todo denominator): custom model build / compose / copy
and their build operations, document classifiers (build/list/get/classify/split/copy), batch
analysis (:analyzeBatch), the remaining ~26 prebuilt field schemas (idDocument, tax.us.,
mortgage.us., contract, healthInsuranceCard.us, bankStatement, check.us, payStub.us, creditCard,
marriageCertificate.us), the other add-on features (queryFields, barcodes, formulas,
styleFont, ocrHighResolution), outputContentFormat=markdown, figures/sections, genuinely
multi-page documents, Entra ID (AAD) auth, list pagination (nextLink), rate limits (429), the
24-hour result expiry, real model-build /operations (the twin serves only the empty-resource
answer), and the v3.1 /formrecognizer surface.
No UI mirror
The API is the product. A Document Intelligence integration lives in backend code: you POST a
document to :analyze and poll Operation-Location (exactly what the official SDK's
getLongRunningPoller does), then read analyzeResult. Document Intelligence Studio is a
labeling/inspection console used to build and try custom models — the analyze traffic an
integrator actually cares about never passes through it — so this pack ships no mirror and has no
UI capabilities.
Connector (D7)
syncAzureFormRecognizerFromReal(client, { root }) pulls a real resource's custom document
models over an injected client (a fake in tests; a thin wrapper over the real REST client in
prod — the pack imports no SDK and holds no key) and folds them into the event log via
the kernel observation path, so a re-pull of identical state appends zero deltas. Pulled models then appear in the
twin's own GET /documentModels and can be analyzed against locally. Prebuilts are filtered out
(service-owned constants, not observed state). Analyze results are not pullable — Document
Intelligence exposes no list endpoint for them — which is filed as
azureformrecognizer.connector.pull_analyze_results and recorded in pull-audit.json.
