@nationsunite/subtitles
v0.1.4
Published
Speaker-aware transcription and subtitle parsing, formatting, and timeline utilities.
Readme
@nationsunite/subtitles
Speaker-aware transcription and subtitle infrastructure shared by browser, training, scraping, and AI workflows.
SubtitleCollection owns two related representations:
- a canonical provider-neutral
SubtitleDocumentcontaining source utterances, separate timed events, and optional words nested beneath their owning utterance. A word addresses an exact substring throughtextStart/textEnd; punctuation and spacing stay in the utterance text instead of becoming synthetic transcript items; and - a display cue plan containing one voice per cue, one or two readable lines, and exact timed text activations.
The canonical structure does not declare transcript, timing, or speaker authority and does not store a timing-granularity flag. Consumers inspect the utterances and their nested words. An ElevenLabs adapter maps its flat transport items at the package boundary. VTT, SRT, and EBU-TT-D parsing preserve all evidence their formats contain. VTT and SRT serialization share the same cue plan; SRT only drops voice and inline word timing because the format cannot encode them.
Wire formats are explicit package subpaths. A format module exports parse, format, or both,
according to what the workspace actually needs. There is no required format interface and callers
do not ask the collection to detect an input format.
Parsers decode source syntax and preserve its evidence. Temporal validity belongs to the canonical
collection: validate() reports typed issues, including whether each issue is repairable, and
fix() returns a repaired copy, a repair ledger, and any remaining issues. Both methods use the
same repair planner, so repairable cannot diverge from what fix() can actually do. Formatting
is lazy, which allows malformed-but-decodable timing to reach this shared validation layer instead
of failing while a parser is still decoding the source.
import { SubtitleCollection } from '@nationsunite/subtitles';
import { parse as parseElevenLabs } from '@nationsunite/subtitles/elevenlabs';
import { format as formatSrt } from '@nationsunite/subtitles/srt';
import { format as formatVtt } from '@nationsunite/subtitles/vtt';
const subtitles: SubtitleCollection = parseElevenLabs(transcription);
const validation = subtitles.validate();
const fixed = subtitles.fix();
const trainingTargetAndDisplayTrack = formatVtt(fixed.collection);
const plainFallback = formatSrt(subtitles);Available modules are ./vtt (parse, format), ./srt (parse, format),
./ebu-tt-d (parse) and ./elevenlabs (parse, format). Documented corpus
annotation formats are also public parser-only subpaths:
./chat, ./eaf, ./exb, ./nxt, ./textgrid, and ./trs.
Those corpus parsers receive source content (or, for linked NXT documents, an explicit document
bundle) and return the same SubtitleCollection. Every parsed utterance, nested word, and event
retains its standard-format ID, ordinal, original source text, timing, speaker, and optional
provider path. The package does
not detect formats, discover files, or interpret provider-specific notation. Each parser has a
public-boundary conformance case derived from one checked-in canonical SubtitleDocument loaded
through SubtitleCollection.fromDocument. The test compares the semantic projection each wire
format can actually preserve; format-specific malformed syntax and exact-evidence behavior remain
focused tests. XML formats use a real XML parser; exact raw element spans are retained separately
when lossless provenance requires them.
Generated WebVTT uses standard voice spans and inline timestamps. When silence follows a word, one timestamp records its end and a later timestamp records the next word's start. This retains provider word boundaries without custom control tokens:
WEBVTT
00:00:00.000 --> 00:00:00.800
<v speaker_1>Good<00:00:00.240><00:00:00.280> morning.<00:00:00.640>The formatter plans each source-scoped speaker lane independently. Genuine crosstalk remains as overlapping voiced cues; minimum display padding is limited by free timeline space and never creates a new speaker overlap. Gap discovery and chunk splitting operate on connected occupied-time components, so they cannot split through overlapping dialogue.
The compact formatting-world E2E covers English, German, Russian, Chinese, Persian, and Hebrew. Each artificial fixture combines continuous speech, semantic wrapping, proper names, quotes/brackets, a long indivisible token, mixed scripts/numbers, equal and zero-duration timestamps, audio events, silence, rapid replies, two- and three-speaker crosstalk, speaker resumption through an interruption, interleaved punctuation, and reused raw speaker IDs across provider responses. Separate focused tests cover the ElevenLabs, VTT, SRT, EBU-TT-D, and automatic format adapters. Non-standard model-training grammars are deliberately not exported from this package.
