@volter/twin-assemblyai
v0.1.35
Published
Local AssemblyAI twin — async speech-to-text transcription API (submit → poll lifecycle, uploads, derived subtitles/search) built on @volter/world-core.
Readme
@volter/twin-assemblyai
Local AssemblyAI twin — the async speech-to-text transcription API (POST /v2/transcript →
poll GET /v2/transcript/:id, queued → processing → completed/error), binary uploads
(POST /v2/upload), the transcript list envelope, and the derived-content endpoints
(word-search, sentences, paragraphs, srt/vtt subtitles). The real assemblyai npm SDK works
unmodified via its public baseUrl option (proven in src/assemblyai-sdk.integration.test.ts).
bun packages/twin/assemblyai/src/cli.ts # serve
bun packages/twin/assemblyai/src/cli.ts serve --scenario s.jsonThe generative stub (honest by design)
Transcription is real hosted compute — the twin has no acoustic model. Transcript CONTENT is a
clearly-labeled deterministic stub ([twin-stub:transcript] … (seed <hash>), hashed from the
audio URL or, for twin uploads, the uploaded bytes' sha256 — same audio, same transcript, always)
while the protocol envelope is faithful: the submit → poll lifecycle, word timings in
milliseconds with per-word confidence, capital-letter speaker labels + utterances when
speaker_labels is on, language_code/language_confidence when language_detection is on,
config echo (speech_models → speech_model_used), error transcripts (status: "error" with a
message — never an HTTP error, exactly how the vendor surfaces unfetchable/no-speech audio), and
the { error } 4xx envelope the SDK converts into thrown Errors.
Scenario scripting (src/assemblyai-scenario.ts, same convention as the anthropic pack):
a JSON file { rules: [{ match: { audioUrlIncludes | audioUrlEquals | nthTranscript }, transcript:
{ text | words | error, audio_duration?, language_code? } }] } scripts the exact transcript — or a
vendor-style error such as language_detection cannot be performed on files with no spoken audio.
— per audio file, so eval worlds can stage silent b-roll, specific dialogue, etc. Rules are
strictly validated at load (unknown keys/mistyped values fail startup loudly). Wire with
createAssemblyAITwinServer({ scenarioPath }), TWIN_ASSEMBLYAI_SCENARIO, or --scenario.
Error-string fidelity note: HTTP status codes and the { error } shape are vendor-faithful;
the exact prose of most error messages is an approximation (not verifiable offline). Where the
string itself is load-bearing for consumers (the no-spoken-audio family), script it verbatim via
a scenario.
Modeled vendor semantics worth knowing:
DELETE /v2/transcript/:idmodels AssemblyAI's data-deletion semantics — the transcript's data is redacted (audio_url→http://deleted_by_user, text/words removed) but the id stays retrievable. Deletion is terminal: a transcript deleted while still pending never advances or regenerates content.- The lifecycle advances one step per poll (deterministic, no wall clock): first GET observes
processing, the second reaches the terminal state. This meansGET /v2/transcript/:idis a state-advancing read (a deliberate trade-off for determinism) — and in--read-onlymode reads never mutate, so a still-pending transcript stays pending forever there (an SDKtranscribe()polling a read-only mirror will run until its ownpollingTimeout). Pulled real transcripts are already terminal, so a pure mirror serves them fine. language_detection: truealways resolves tolanguage_code: "en"with a deterministic seed-derivedlanguage_confidence(0.95–0.99) — no detection actually runs (part of the deterministic stub); scriptlanguage_codevia a scenario when a test needs another language.word-searchomits searched terms with zero hits frommatches(a modeling assumption — not verified against the live API offline).
Coverage
The capability manifest (src/assemblyai-capabilities.ts) is an honest partial denominator for
the AssemblyAI API, authored top-down from the vendor API reference and the official SDK v4
paths/types. Coverage is partial and will read low as the denominator grows — see the manifest's
todo entries (LeMUR envelopes, realtime/streaming WebSockets, webhooks, PII redaction, the
audio-intelligence models, cursor pagination, auth/rate-limit parity) for what is enumerated but
not yet modeled.
What grounds those todos (2026-09-02). AssemblyAI's published OpenAPI document
(www.assemblyai.com/docs/openapi.yaml 1.3.4 — this pack's census denominator) is 8 paths / 10
operations: POST /v2/upload and the /v2/transcript family. It names neither LeMUR nor
POST /v2/realtime/token, and the AssemblyAI/assemblyai-api-spec repository that used to name
them no longer exists. Those todos therefore rest on the installed official SDK
(assemblyai 4.36.4), whose lemur, realtime and streaming services still call
/lemur/v3/..., POST /v2/realtime/token, wss://api.assemblyai.com/v2/realtime/ws and
wss://streaming.assemblyai.com/v3/ws — the surface exists; only its documentation moved.
Modeled: transcript submit → poll lifecycle (queued/processing/completed/error), uploads
(content-derived URLs), word/utterance shapes (ms timings, confidence, speaker labels), language
detection fields, config echo, list envelope (page_details, limit, status filter), delete
(redaction semantics), word-search, sentences/paragraphs, srt/vtt subtitles
(chars_per_caption), scenario scripting, read-only mode, and the connector's pull.
No UI mirror
AssemblyAI is a vendor whose product is the API: when someone does AssemblyAI's core job they
write code — integrators call POST /v2/transcript from a backend (Ponder's transcription
pipeline is exactly this), and the web dashboard is incidental tooling for API keys, usage, and
billing rather than where the work happens. Per
../../../docs/contributing/adding-a-twin.md ("Does this vendor get a mirror?" — "An
integrator calls the API; the vendor's console is incidental tooling … → the API is the product
→ omit the mirror") this pack ships no React mirror and no UI capabilities — omit rather than
fabricate a dashboard. Coverage is API + connector.
