@qvac/asr-ggml
v0.7.0
Published
Multi-engine ASR (Whisper + NVIDIA Parakeet) inference addon for qvac on the Bare runtime
Readme
@qvac/asr-ggml
Multi-engine automatic speech recognition for QVAC runtime applications on the
Bare runtime. One npm package and one native prebuild serve two
ggml-based ASR engines behind a single class, ASRGgml:
| Engine | Native library | Good for |
| --- | --- | --- |
| Whisper | whisper.cpp | Multilingual offline transcription, translation, Silero-VAD-segmented live capture |
| Parakeet | parakeet-cpp through the speech-cpp umbrella port (NVIDIA Parakeet / Sortformer) | Low-latency streaming ASR, native end-of-turn detection, 4-speaker diarization |
This package replaces @qvac/transcription-whispercpp and
@qvac/transcription-parakeet. See CHANGELOG.md for the
breaking changes the merge introduced.
Table of Contents
- Supported Engines and Models
- Choosing a model
- Supported Platforms
- Installation
- Quickstart
- Engine Selection
- API Surface
- Assessing fit
- Configuration Reference
- Audio Input
- Backends and GPU Acceleration
- Staging Models
- Error Codes
- Development
- Benchmarking
- Examples
- Documentation
- Glossary
- License
Supported Engines and Models
Whisper (engine: 'whisper')
Legacy single-file GGML .bin checkpoints from
ggerganov/whisper.cpp:
| Model | Size | Description |
|-------|------|-------------|
| ggml-tiny.bin | 78 MB | Smallest, fastest |
| ggml-base.bin | 148 MB | Balanced size/accuracy |
| ggml-small.bin | 488 MB | Better accuracy |
| ggml-medium.bin | 1.5 GB | High accuracy |
| ggml-large-v3.bin | 3.1 GB | Best accuracy |
| ggml-large-v3-turbo.bin | 1.6 GB | Best accuracy, faster |
Quantized variants (q8_0, q5_1, q5_0) exist for all sizes. Whisper
covers ~99 languages plus translation-to-English, and the fine-tuned
per-language checkpoints listed in NOTICE also load.
VAD model (required for runStreaming()), from
ggml-org/whisper-vad:
| Model | Size | Description |
|-------|------|-------------|
| ggml-silero-v5.1.2.bin | 885 KB | Silero VAD for voice-activity detection |
Parakeet (engine: 'parakeet')
Single-file .gguf checkpoints. The model type is auto-detected from the
GGUF metadata — there is no modelType to pass.
| Variant | Languages | Decoder | ~Size (q8_0) | Notes |
|---------|-----------|---------|-------------:|-------|
| CTC (parakeet-ctc-0.6b) | English | argmax CTC | ~700 MiB | Fast, no punctuation/capitalization |
| TDT (parakeet-tdt-0.6b-v3) | ~25 | RNN-T greedy + duration | ~715 MiB | Recommended default; PnC + language auto-detect |
| Unified (parakeet-unified-en-0.6b) | English | RNN-T | ~715 MiB | One checkpoint for batch and cache-aware streaming at 80/160/560/1040 ms; PnC |
| EOU (parakeet-eou-120m-v1) | English | RNN-T greedy + <EOU> | ~132 MiB | Streaming-trained; native end-of-turn token |
| Indic Conformer CTC (indic-conformer-ctc) | Indic aggregate | argmax CTC + language mask | ~701 MiB | Multilingual Indic; set parakeetConfig.language (e.g. "hi") |
| Sortformer v1 (sortformer-4spk-v1) | n/a | Diarization head (sliding history) | ~141 MiB | 4-speaker. Default for offline diarization |
| Sortformer v2.1 + AOSC (diar_streaming_sortformer_4spk-v2.1) | n/a | Diarization head + speaker cache | ~141 MiB | 4-speaker. Default for streaming diarization; AOSC anchors speaker slots across silence, auto-detected from GGUF metadata |
On macOS and iOS, TDT, Unified, EOU, and Sortformer v2.1 can also run their encoder on an optional Core ML sidecar; CTC and Indic Conformer CTC cannot with the pinned engine, and Sortformer v1 has no sidecar. See Core ML encoder sidecars.
Upstream .nemo checkpoints are NVIDIA's; see the
Parakeet model cards
for the per-checkpoint NVIDIA Open Model License terms.
Choosing a model
Pick a specific checkpoint, not only an engine. Whisper and Parakeet overlap on English batch transcription; they diverge on streaming semantics, language coverage, translation, and diarization.
Decision guide
| If you need… | Use this model | Notes |
| --- | --- | --- |
| Default multilingual / English ASR (batch or duplex stream) | parakeet-tdt-0.6b-v3 (q8_0 GGUF) | Recommended Parakeet default: ~25 languages, punctuation/capitalization, language auto-detect, low-latency streaming. |
| English batch and low-latency streaming with one checkpoint | parakeet-unified-en-0.6b | Standard RNN-T with punctuation and capitalization; use when multilingual TDT or native EOU tokens are not required. Streaming uses the native cache-aware encoder: streamingChunkMs accepts 80, 160, 560, or 1040 and streamingRightLookaheadMs 0, 80, 160, 240, 320, 560, or 1040, both snapped down to the nearest trained value. Defaults to 560 ms. |
| Native end-of-turn for conversational / duplex English | parakeet-eou-120m-v1 | Emits <EOU>; smallest Parakeet (~132 MiB). Pair with TDT when you need broader language coverage and EOU. |
| Fast English-only, no punctuation | parakeet-ctc-0.6b | Lowest decode cost in the Parakeet family; no PnC. |
| Indic-language ASR (Hindi and other Indic ids) | indic-conformer-ctc | Pass parakeetConfig.language (e.g. "hi"). Same Parakeet engine; GGUF lives under indic_conformer/ in the registry. |
| Offline 4-speaker diarization | sortformer-4spk-v1 | Default offline diarization head. |
| Streaming 4-speaker diarization | diar_streaming_sortformer_4spk-v2.1 | AOSC keeps speaker slots across silence; prefer over v1 for live streams. |
| Broadest language set + translate-to-English | ggml-large-v3-turbo.bin (or ggml-small.bin on edge) | Whisper: ~99 languages, translation, Silero-VAD live capture. Turbo is the accuracy/speed sweet spot; use tiny/base only when size dominates. |
| Live capture with VAD segmentation (Whisper path) | Whisper ASR model + ggml-silero-v5.1.2.bin | Silero VAD is required for Whisper runStreaming(). |
Engine vs model
| Engine | Prefer when… | Prefer the other when… | | --- | --- | --- | | Parakeet | Low-latency streaming, native EOU, diarization, English / ~25-lang product ASR | You need Whisper’s language breadth or translate-to-English | | Whisper | Multilingual offline, translation, VAD-segmented live capture on the whisper.cpp path | You need Parakeet EOU / Sortformer diarization or tighter streaming RTF |
Always pass config.engine (or top-level engine) explicitly in library and
SDK code — see Engine Selection.
Supported Platforms
| Platform | Architecture | Min Version | Status | GPU Support |
|----------|-------------|-------------|--------|-------------|
| macOS | arm64, x64 | 14.0+ | ✅ Tier 1 | Metal |
| iOS | arm64 | 17.0+ | ✅ Tier 1 | Metal |
| Linux | arm64, x64 | Ubuntu-22+ | ✅ Tier 1 | Vulkan; CUDA (x64 prebuild; arm64 via build:cuda / ASR_CUDA=ON) |
| Android | arm64 | 12+ | ✅ Tier 1 | Vulkan, OpenCL (Adreno) |
| Windows | x64 | 10+ | ✅ Tier 1 | Vulkan; CUDA via build:cuda / ASR_CUDA=ON |
Dependencies:
qvac-lib-inference-addon-cpp— C++ addon framework (vcpkg port; version pinned invcpkg.json)speech-cpp— umbrella vcpkg port enabling thewhisperandparakeetengine features plus each platform's GPU features- Bare runtime — see
engines.bareinpackage.json - Linux requires Clang/LLVM 22 with libc++
Installation
Make sure the Bare runtime is installed:
npm install -g bare bare-make
bare -v # must satisfy package.json engines.bareThen:
npm install @qvac/asr-ggmlPlatform packages
@qvac/asr-ggml is a meta package that ships the JavaScript wrapper only.
The native prebuild for each desktop host lives in a version-locked platform
package selected at install time through os/cpu filtered
optionalDependencies:
| Host | Package |
| --- | --- |
| linux-x64 (glibc) | @qvac/asr-ggml-linux-x64 |
| linux-arm64 (glibc) | @qvac/asr-ggml-linux-arm64 |
| darwin-arm64 | @qvac/asr-ggml-darwin-arm64 |
| darwin-x64 | @qvac/asr-ggml-darwin-x64 |
| win32-x64 | @qvac/asr-ggml-win32-x64 |
Do not depend on desktop platform packages directly. Supported installers are
npm 7+, pnpm, bun, and Yarn Berry. Yarn v1 and --omit=optional installs skip
the platform package and fail at require time with an error naming the missing
package; a locally built prebuilds/ directory in the package root always
takes precedence. Use require('@qvac/asr-ggml').resolveBackendsDir() to
locate the directory holding the host's prebuilt binaries and dynamically
loaded ggml backends.
Mobile targets are cross-built, so no install host ever matches their os,
and optionalDependencies filtering can never select them. Mobile
applications must declare the target's platform package as a direct
dependency, pinned to the exact @qvac/asr-ggml version:
| Target | Package |
| --- | --- |
| android-arm64 | @qvac/asr-ggml-android-arm64 |
| ios (device + simulators) | @qvac/asr-ggml-ios |
{
"dependencies": {
"@qvac/asr-ggml": "x.y.z",
"@qvac/asr-ggml-android-arm64": "x.y.z"
}
}Quickstart
All four snippets assume:
const fs = require('bare-fs')
const ASRGgml = require('@qvac/asr-ggml')Audio must be 16 kHz mono. See Audio Input for the accepted shapes.
Whisper — batch run()
const model = new ASRGgml({
files: { model: './models/ggml-tiny.bin' },
config: {
engine: 'whisper',
whisperConfig: {
language: 'en',
audio_format: 's16le' // how raw Uint8Array bytes are decoded
}
}
})
await model.load()
const audioStream = fs.createReadStream('./audio.raw', { highWaterMark: 16000 })
const response = await model.run(audioStream)
// Push-based: segments arrive as whisper.cpp emits them.
await response
.onUpdate((out) => {
for (const segment of Array.isArray(out) ? out : [out]) {
console.log(segment.start, '→', segment.end, segment.text)
}
})
.await()
await model.destroy()Or pull-based, with iterate() instead of onUpdate():
const response = await model.run(audioStream)
for await (const out of response.iterate()) {
for (const segment of Array.isArray(out) ? out : [out]) {
console.log(segment.text)
}
}Whisper — VAD streaming runStreaming()
Silero VAD splits the incoming audio into utterances and each one is
transcribed as it completes. A VAD model is required — pass it as
files.vadModel, config.vadModelPath, or
config.whisperConfig.vad_model_path; a missing file throws
VAD_MODEL_NOT_FOUND (6018) from the constructor, and a missing path throws
VAD_MODEL_REQUIRED (6009) from runStreaming().
const model = new ASRGgml({
files: {
model: './models/ggml-tiny.bin',
vadModel: './models/ggml-silero-v5.1.2.bin'
},
config: { engine: 'whisper', whisperConfig: { language: 'en' } }
})
await model.load()
const response = await model.runStreaming(micStream, {
emitVadEvents: true, // adds { type: 'vad', speaking, score, source }
endOfTurnSilenceMs: 800 // adds { type: 'endOfTurn', source: 'vad-silence' }
})
for await (const out of response.iterate()) {
if (out.type === 'vad') console.log('speaking:', out.speaking)
else if (out.type === 'endOfTurn') console.log('--- turn ended ---')
else for (const s of Array.isArray(out) ? out : [out]) console.log(s.text)
}Parakeet — batch run()
const model = new ASRGgml({
files: { model: './models/parakeet-tdt-0.6b-v3.q8_0.gguf' },
config: {
engine: 'parakeet',
parakeetConfig: { useGPU: true, maxThreads: 4 }
}
})
await model.load()
const response = await model.run(float32Samples) // Float32Array, 16 kHz mono
const segments = []
await response
.onUpdate((out) => {
for (const s of Array.isArray(out) ? out : [out]) {
// `toAppend` marks a segment that continues the previous one rather
// than replacing it.
if (s.text && s.toAppend) segments.push(s)
}
})
.await()
console.log(segments.map((s) => s.text).join(''))
await model.unload()Buffer cap (
run()only): every chunk of onerun()call is normalized to Float32 and batched into a single nativeprocess()call. The cap is 500 MiB of caller-supplied audio — ≈4.55 h of 16 kHz monos16le(the default byte format) or ≈2.27 h of 16 kHz monof32le, which needs no expansion on the way to native. Exceeding it raisesBUFFER_LIMIT_EXCEEDED(6015), whose message names the source format the budget is denominated in. For longer captures userunStreaming(), which feeds the engine as audio arrives, or split into sequentialrun()calls.
Parakeet — duplex streaming runStreaming()
runStreaming() opens one long-lived native session for the lifetime of the
call and forwards each chunk as it arrives — no batching, no per-chunk session
recreation. Speaker IDs stay stable across appends, and an EOU model's <EOU>
boundaries surface both as segment.isEndOfTurn and as a synthesized
{ type: 'endOfTurn', source: 'model-eou' } event.
const model = new ASRGgml({
engine: 'parakeet', // no config → alias form
files: { model: './models/parakeet-eou-120m-v1.q8_0.gguf' }
})
await model.load()
const response = await model.runStreaming(micStream, { chunkMs: 480 })
await response
.onUpdate((out) => {
if (out.type === 'endOfTurn') return console.log('--- turn ended ---')
for (const s of Array.isArray(out) ? out : [out]) process.stdout.write(s.text)
})
.await()Only one streaming session may be open per instance; a concurrent run() or
runStreaming() throws STREAMING_SESSION_ACTIVE (6020).
Engine Selection
The engine is resolved once, in the constructor, from three sources in strict precedence order:
config.engine— authoritative, and the recommended form. Ifconfigis supplied at all, it must carryengine; a missing or unrecognized value throwsINVALID_ENGINE(6021).engine— a top-level convenience alias, used only whenconfigis omitted entirely. Unrecognized values throwINVALID_ENGINE.- Model-file sniffing — last resort when neither is given. The first four
bytes of
files.modelare read:GGUF⇒ parakeet, anything else ⇒ whisper.
Sniffing is a convenience for scripts, not an integration path. It opens
the model file synchronously inside the constructor, it cannot distinguish a
GGUF whisper build from a GGUF parakeet build, and it throws INVALID_ENGINE
if the file exists but cannot be read (a missing file is reported first, as
MODEL_NOT_FOUND (24009)). Library and SDK callers should always pass
config.engine explicitly.
Validation and sniffing both target the file the driver actually opens: for
whisper that is config.path when set, otherwise files.model; parakeet only
ever loads files.model.
getEngineType() reports the resolved engine; ASRGgml.ENGINE_WHISPER and
ASRGgml.ENGINE_PARAKEET are available as statics.
API Surface
Every verb has one signature and one meaning regardless of engine.
| Member | Description |
| --- | --- |
| new ASRGgml({ files, config?, engine?, enableStats?, logger?, exclusiveRun? }) | Resolves the engine, validates model files and the engine config vocabulary. Throws on any problem — nothing is deferred to load(). |
| load() | Creates the native instance and activates the model. Calling it on a loaded instance unloads first. Throws INSTANCE_DESTROYED after destroy(). |
| run(audio) | Batch transcription. Returns a QvacResponse; drain it with onUpdate(cb) (push) or iterate() (pull). |
| runStreaming(audio, opts?) | Duplex/VAD-segmented streaming. Resolves once the native session is open; opts is the engine's streaming vocabulary. |
| reload(newConfig?) | Applies an engine-scoped partial config in place where possible. Rejects with NOT_SUPPORTED (6019) on an engine whose driver has no native reload. |
| cancel(jobId?) | Cancels the active job and fails it, so a draining iterate() throws. The native verb takes no id; jobId is accepted for source compatibility only. |
| status() | Native state string. Before load(), rejects with FAILED_TO_GET_STATUS: 6004 for Whisper and 24004 for Parakeet. |
| addon | The native interface, or undefined before load() (not cleared by unload(), as in both pre-merge packages). Escape hatch for a native hard cancel that stops the decode without failing the job (what the SDK's model-wide cancel uses). Not otherwise part of the supported surface. |
| unload() / destroy() | Release the model / retire the instance. |
| getState() | { configLoaded, weightsLoaded, destroyed }. |
| getEngineType() | 'whisper' | 'parakeet'. |
| getBackendInfo() | BackendInfo or null before load(). |
| pause() / unpause() | Always reject with NOT_SUPPORTED (6019). Neither engine implements a correct pause/resume. |
Constructor options:
| Option | Default | Description |
| --- | --- | --- |
| files.model | — | Required. Path to the .bin (whisper) or .gguf (parakeet) checkpoint. Must exist. |
| files.vadModel | — | Whisper streaming only: path to the Silero VAD model. Must exist if given. |
| config | — | Engine-scoped config; must carry engine when present. |
| engine | — | Alias for config.engine, used when config is omitted. |
| enableStats | true | Attach RuntimeStats to the job-end payload. |
| logger | null | A @qvac/logging LoggerInterface. |
| exclusiveRun | true | Serialize run() calls to completion on the inference lane; runStreaming() holds that lane for session setup only. reload/unload/destroy serialize on a separate lifecycle lane, so teardown pre-empts an in-flight run() instead of queueing behind it. |
run() / runStreaming() output payloads are:
- bare
TranscriptionSegment[](or a single segment) for transcripts — there is no{ type: 'segment' }wrapper; { type: 'vad', speaking, score, source, timestamp?, speakerId? }for voice-activity events, emitted on each speech/silence change.sourceis'silero'(Whisper streaming),'energy'(Parakeet ASR energy detector;scoreis the window RMS), or'sortformer'(Parakeet speaker activity;speakerIdnames the dominant speaker when speech starts). Parakeet events carrytimestamp, in seconds from the start of the session;{ type: 'endOfTurn', source, silenceDurationMs? }for turn boundaries.
Engine-specific segment fields:
- Whisper segments carry
noSpeechProb, andlanguagewhenever whisper reports one (the decoded language, the detected one underlanguage: 'auto'). Withtoken_timestamps: truethey addtokens: [{ text, start, end, probability }](special tokens left out); withtdrz_enable: trueon a tinydiarize model they addspeakerTurnNext. - Parakeet Sortformer keeps the
"Speaker N: start - end"text and adds the same data in structured form: a streamed diarization segment carriesspeakerId, and the offline transcript carriesspeakerSegments: [{ speakerId, start, end }], one entry per text line.
Assessing fit
assessFit projects a load against the memory free right now. It reads model metadata and never weight data. Parakeet is GGUF, so the registry's weightless copy of a parakeet model answers the same as the model itself and the projection can run before it is downloaded. Whisper ships as .bin, which the registry has no weightless form for, so a whisper projection needs the file. It is a module export, not an instance method — nothing is loaded to call it.
const ASRGgml = require('@qvac/asr-ggml')
const fit = ASRGgml.assessFit({
engine: 'whisper',
modelPath: '/models/whisper.bin',
vadModelPath: '/models/silero-vad.bin',
audioSeconds: 300,
decoders: 5
})
fit.status // 'fits' | 'does-not-fit' | 'error'
fit.reason // the engine's own wording, e.g. 'model-unreadable', 'workload-too-large'
fit.modelType // whisper: 'tiny' … 'large v3'; parakeet: 'ctc' | 'rnnt' | 'tdt' | 'eou' | 'nemotron' | 'sortformer'
fit.deviceName
fit.deviceBytes
fit.weightsBytes
fit.hostBytes
fit.reportengine picks the fitter and defaults to parakeet. Each engine fills its own breakdown on the result: whisper reports kvBytes, computeBytes, vadBytes and hostOverflowBytes; parakeet reports encoderComputeBytes, decoderStateBytes and decoderComputeBytes.
| Option | Description |
| --- | --- |
| modelPath | Required. Absolute path to the model, or to the registry's weightless copy where one exists. |
| audioSeconds | Longest single transcribe the projection must cover. Defaults to 300. |
| gpuLayers | Greater than 0 requests the GPU stack, with the fallbacks a real load applies. Omitted, parakeet projects on the CPU and whisper on the GPU, matching what each load does. |
| marginBytes | Free memory that must remain for the projection to count as fitting. Defaults to the engine's own headroom, which is 256 MiB for parakeet. |
| backendsDir | The prebuilds root. The backends are read from the per-target subdir under it, the same path a load reads. |
| vadModelPath | Whisper: projected alongside the model. Omitted means no VAD. |
| decoders | Whisper: worst-case resident decoders, the best_of or beam_size the run will use. The KV cache and decode graph grow with it. |
| flashAttn, gpuDevice | Whisper: as the load takes them. |
| threads, longFormWindowFrames, longFormContextFrames | Parakeet: as the load takes them. |
| nemotronChunkMs | Nemotron: the streaming operating point the projection must also cover. 0 projects the largest allowed one. |
A model the fitter cannot read is status: "error" with the engine's reason; only a broken request throws.
Configuration Reference
Configuration vocabularies are engine-scoped — there is no merged config namespace. A key belonging to one engine is rejected, not ignored, by the other, and unknown keys throw at construction time.
Whisper: config.whisperConfig / contextParams / miscConfig
const config = {
engine: 'whisper',
contextParams: {
model: './models/ggml-tiny.bin',
use_gpu: true, // opt-in; defaults to false
gpu_device: 0
},
whisperConfig: {
language: 'en', // 'auto' for language detection
duration_ms: 0,
temperature: 0.0,
suppress_nst: true,
n_threads: 0,
audio_format: 's16le',
vad_model_path: './models/ggml-silero-v5.1.2.bin',
vad_params: {
threshold: 0.6,
min_speech_duration_ms: 250,
min_silence_duration_ms: 200
}
},
miscConfig: { caption_enabled: false }
}The public Whisper keys, including vad_params, are declared by
WhisperConfig. The constructor accepts a
curated subset of whisper_full_params covering decoder strategy, thresholds,
timestamps, VAD, and prompts.
WhisperDriver maps vad_params to the native vadParams object; callers
must not pass vadParams. Any key outside the list throws from the constructor. contextParams accepts only
model, use_gpu, flash_attn, gpu_device; miscConfig only
caption_enabled, seed. For what each flag means see the upstream
whisper_full_params reference.
Notes:
- GPU is opt-in.
use_gpudefaults tofalse; set it incontextParams. - Four context keys force a full reload —
model,use_gpu,flash_attn,gpu_device,main-gpu,main_gpu. Changing any of them destroys and rebuilds the whisper context (seconds, depending on model size). Everything inwhisperConfigis applied in place. backendsDir(inwhisperConfig) overrides where dynamically-loaded ggml backend libraries are found. See Backends and GPU Acceleration.max_secondsis a convenience that derivesduration_ms.carry_initial_prompt: trueprependsinitial_promptto every decode window instead of only the first.
Whisper runStreaming(audio, opts) options:
| Option | Description |
|--------|-------------|
| emitVadEvents / conversationMode | Emit { type: 'vad' } as speech starts and stops. |
| endOfTurnSilenceMs | Emit { type: 'endOfTurn' } after this much trailing silence. |
| vadRunIntervalMs | How often the VAD is evaluated. |
Parakeet: config.parakeetConfig
const config = {
engine: 'parakeet',
parakeetConfig: {
useGPU: true,
maxThreads: 4,
timestampsEnabled: true,
streaming: false // true opens a session at load time
}
}The authoritative vocabulary is ParakeetConfig in
published driver.d.ts — every
key is documented inline there, and any key outside it throws
INVALID_CONFIG (24015) from the constructor. Groups:
| Group | Keys |
| --- | --- |
| Compute | maxThreads, useGPU, seed |
| Audio | sampleRate (16000), channels (1) |
| Output | captionEnabled, timestampsEnabled |
| Language | language — multilingual CTC id (e.g. "hi"); required for Indic Conformer GGUFs that advertise parakeet.ctc.lang_* ranges; ignored on monolingual CTC |
| Streaming (ASR) | streaming, streamingChunkMs, streamingEmitPartials, streamingEnergyVad, streamingEnergyVadThresholdDb, streamingEnergyVadWindowMs, streamingEnergyVadHangoverMs, streamingLeftContextMs, streamingRightLookaheadMs |
| Streaming (Sortformer) | streamingHistoryMs, streamingSpeakerVad, streamingSpkCacheEnable, streamingSpkCacheLen, streamingFifoLen, streamingChunkLeftContextMs, streamingChunkRightContextMs, streamingSpkCacheUpdatePeriod |
| Diarization (Sortformer, offline and streaming) | diarizationThreshold (0.641), diarizationMinSegmentMs (510) |
| Engine | prewarm, prewarmAudioSeconds, longFormWindowFrames, longFormContextFrames |
| Backends | backendsDir, openclCacheDir |
The streamingSpkCache* / streamingFifoLen /
streamingChunk{Left,Right}ContextMs defaults are the NeMo-port tuning
parakeet-cpp ships — keep them unless you are A/B comparing AOSC against the
v1 sliding-window path. There is no modelType: CTC / TDT / EOU / Sortformer,
and Sortformer v1 vs v2.1+AOSC, are all detected from the GGUF metadata.
Parakeet runStreaming(audio, opts) options are per-call overrides of the same
knobs without the streaming prefix: chunkMs, historyMs, leftContextMs,
rightLookaheadMs, emitPartials, emitEnergyVad, energyVadThresholdDb,
energyVadWindowMs, energyVadHangoverMs, emitSpeakerVad,
diarizationThreshold, diarizationMinSegmentMs, spkCacheEnable,
spkCacheLen, fifoLen, chunkLeftContextMs, chunkRightContextMs,
spkCacheUpdatePeriod.
Voice-activity events are opt-in. streamingEnergyVad / emitEnergyVad runs
the engine's RMS energy detector on CTC, TDT, RNN-T, and Nemotron sessions
(EOU models signal turns through <EOU> instead); the threshold is in dBFS
(default -35) and the hangover is how long the audio must stay below it
before the state returns to silence (default 200 ms).
streamingSpeakerVad / emitSpeakerVad reports Sortformer speaker activity.
prewarm runs one synthetic encoder pass while loading so the first request
does not pay the GPU shader or kernel compile. longFormWindowFrames bounds
the offline encoder window (0 picks it from the model, a negative value
always runs one pass), which keeps memory flat on long run() inputs;
cancel() also takes effect between those windows.
For streaming diarization, use the Sortformer v2.1 GGUF. Its metadata enables
AOSC automatically; keep the speaker-cache defaults unless you are comparing
against the v1 sliding-window path. Sortformer v1 remains the offline
diarization default. The conversion scripts support both variants and read
NVIDIA .nemo archives directly.
Audio Input
Both engines take 16 kHz mono audio. run() and runStreaming() accept a
stream, an iterable, a single chunk, or an array of chunks. A chunk's class
decides how it is interpreted:
| Chunk type | Interpretation |
| --- | --- |
| Float32Array | f32 samples in [-1, 1] |
| Int16Array | s16 samples |
| Uint8Array | raw bytes, decoded as s16le by default; whisper's whisperConfig.audio_format: 'f32le' switches the byte interpretation |
audio_format accepts 's16le', 'f32le' and 'decoded' (an alias for
'f32le'); anything else throws INVALID_AUDIO_FORMAT (24010) rather than
being decoded as little-endian s16. It only ever describes raw Uint8Array
bytes — it never reinterprets a typed array.
A byte length that is not a whole number of samples raises
INVALID_AUDIO_INPUT (6011). In a stream of byte chunks the check is
applied in aggregate, so a sample split across a chunk boundary (arbitrary
socket/pipe read sizes) is carried over into the next chunk; only a stream
that ends mid-sample is rejected.
Backends and GPU Acceleration
GPU backends are selected per platform via vcpkg.json features; no
bare-make generate flag is needed:
- Linux / Windows — Vulkan (needs the Vulkan SDK on the build host); the linux-x64 prebuild additionally bundles CUDA, see below
- Android — Vulkan + OpenCL (Adreno) as dynamically-loaded
.sobackends shipped beside the prebuild - macOS / iOS — Metal, statically linked, plus the optional Parakeet Core ML encoder sidecars
CUDA (Linux / Windows on NVIDIA) needs nvcc on the build host, so it is
gated behind the ASR_CUDA CMake option — supported on linux-x64,
linux-arm64 and win32-x64. The published linux-x64 prebuild turns it on (the
prebuild workflow installs the CUDA toolkit); elsewhere build it yourself
with npm run build:cuda (or bare-make generate -D ASR_CUDA=ON). The
option adds the cuda feature to the speech-cpp dependency and turns on
GGML_CUDA. Every linux-x64 and linux-arm64 build, and win32-x64 with the
cuda feature, uses ggml's hybrid dynamically-loaded backend mode: the
per-arch CPU-variant and Vulkan backends ship as runtime-loaded modules
(.so on Linux, .dll on Windows) next to the addon, the cuda builds add
the CUDA module, and only that module depends on the CUDA runtime. Engaging
CUDA requires the NVIDIA driver plus the CUDA 13 runtime libraries (cudart
and cuBLAS) resolvable at load time; hosts that cannot resolve them —
including CPU-only and non-NVIDIA machines — skip the module and fall back
to Vulkan or CPU instead of failing to load the addon. CUDA is compiled
alongside Vulkan rather than replacing it; ggml registers CUDA ahead of
Vulkan, so a use_gpu / useGPU request lands on CUDA when a supported
device is present and falls back to Vulkan otherwise. Both engines report
the winner through getBackendInfo() as backendId: 2 (BackendId.CUDA).
On x64 a CUDA build's module targets compute capability 7.5 and newer, with native code for Turing (7.5 — RTX 20xx, GTX 16xx, T4), Ampere (8.0, 8.6), Ada (8.9), Hopper (9.0) and Blackwell (12.0, 12.1). Anything newer JIT-compiles from the bundled 8.0 PTX on first use, a one-off compile the driver caches. On linux-arm64 the native set is Jetson Orin (8.7), Grace-Hopper (9.0) and GB10 / DGX Spark (12.1), with discrete Ampere+ cards and newer parts covered through the bundled 8.0 PTX. Volta and Pascal fall outside CUDA 13's support entirely, so they have no code path here: the backend skips such devices at registration and the addon falls back to Vulkan or CPU.
The addon takes no direct CUDA linkage — the CUDA module carries its own CUDA
DT_NEEDED entries, which is what makes the graceful fallback possible — and
nvcc's clang host-compiler setup lives in
vcpkg-overlays/toolchains/linux-clang.cmake, shared by every addon that
compiles the CUDA backend.
Both engines default to CPU: whisper needs contextParams.use_gpu: true,
parakeet needs parakeetConfig.useGPU: true.
For Whisper GPU selection, set contextParams['main-gpu'] (or the alias
contextParams.main_gpu) to a raw ggml registry index, an integer string,
'dedicated', or 'integrated' (class names are case-insensitive). With GPU
enabled and no explicit selector, dedicated GPUs are preferred. A class
selector is strict: if that class is unavailable, execution falls back to CPU.
An in-range numeric selector preserves its registry identity before backend
filtering; a CPU, excluded backend, or refused Adreno Vulkan slot falls back to
CPU without selecting another GPU. An out-of-range index logs a warning and
uses normal selection. The supported local families are Metal, CUDA, Vulkan,
and OpenCL; the existing Adreno OpenCL guard still applies.
main-gpu does not enable GPU execution by itself. use_gpu: false always
selects CPU. The legacy contextParams.gpu_device retains its existing
Whisper GPU/IGPU-ordinal meaning and Adreno guard. Combining it with either
new selector spelling, or supplying both new spellings, is rejected.
This selector currently applies to the Whisper engine.
getBackendInfo() reports what actually ran — backendName, backendId
(see the BackendId enum), string backendDevice, backendDescription,
encoderBackend, encoderOnCoreml (Apple: whether a Parakeet Core ML
encoder sidecar loaded; see
Core ML encoder sidecars), and
modelType ('whisper', or the Parakeet family detected from the GGUF:
'ctc', 'tdt', 'rnnt', 'eou', 'nemotron', 'sortformer'). Whisper
additionally reports
gpuMemTotalMb / gpuMemFreeMb. This differs from
RuntimeStats.backendDevice, which is the native numeric device-class code.
Two paths matter on Android and Linux:
backendsDir(inwhisperConfig/parakeetConfig) — root directory holding dynamically-loaded ggml backend libraries (CUDA, Vulkan, OpenCL, per-arch CPU variants). Defaults toresolveBackendsDir(): the package's ownprebuilds/when present (local builds, mobile flatten), otherwise the installed platform package (see Platform packages); the native addon appends<bare-target>/<module-name>before scanning. Pass an explicit path when backend libraries ship elsewhere — e.g. Android'sApplicationInfo.nativeLibraryDirwhen they are packaged inside the APK. No-op on Apple, where backends are statically linked.openclCacheDir(parakeet) — persistent directory for ggml-opencl's compiled program-binary cache. Android-only; pass the host app's cache directory to avoid a coldclBuildProgramon every process start.
Core ML encoder sidecars (Apple)
The macOS and iOS prebuilds are built with the speech-cpp coreml feature,
so a Parakeet model can run its FastConformer encoder on Apple Core ML (the
Neural Engine) while mel preprocessing and the decoder or speaker head stay on
the ggml backend. It is opt-in by presence: at load() the engine looks for a
compiled <stem>-encoder.mlmodelc next to the GGUF, where <stem> is the GGUF
name with its quantization suffix stripped, so one sidecar serves every tier
(parakeet-tdt-0.6b-v3.q8_0.gguf and .f16.gguf both resolve to
parakeet-tdt-0.6b-v3-encoder.mlmodelc). The published models ship without
sidecars, so a model directory behaves as before until you stage one. A
missing sidecar, an input shape it does not take, or a failed prediction falls
back to ggml, and PARAKEET_COREML_DISABLE=1 in the process environment forces
ggml.
| Model | Sidecar | Inputs it takes | Runs on ggml instead |
| --- | --- | --- | --- |
| TDT (parakeet-tdt-0.6b-v3) | <stem>-encoder.mlmodelc | any length: shorter inputs are zero-padded to the compiled shape, longer offline inputs are split into overlapping windows | only on fallback |
| Unified (parakeet-unified-en-0.6b) | <stem>-encoder.mlmodelc | run(), padded or windowed like TDT | runStreaming(), which uses the cache-aware encoder |
| EOU (parakeet-eou-120m-v1) | <stem>-encoder.mlmodelc | inputs of exactly the compiled mel length | every other length, including streaming windows of a different length |
| Sortformer v2.1 + AOSC | <stem>-encoder.mlmodelc (batch), <stem>-encoder-bypass-pre-encode.mlmodelc (AOSC) | batch: exactly the compiled length; AOSC: slabs up to the masked capacity (410 encoder frames for the default geometry) | other batch lengths, larger AOSC slabs; the speaker head always |
| CTC (parakeet-ctc-0.6b), Indic Conformer CTC | none in the pinned speech-cpp | — | always (the engine adds CTC sidecars from speech-cpp 2026-09-24) |
| Sortformer v1 | none | — | always |
| Whisper | none: the whisper feature builds without WHISPER_COREML | — | always |
getBackendInfo().encoderOnCoreml (with encoderBackend: 'coreml') and
RuntimeStats.encoderOnCoreml report that a sidecar loaded at load(), not
that a given call ran on it: an EOU input of another length, a Unified
runStreaming() session, or a failed prediction still runs the encoder on
ggml with the flag set. For what a job actually did, read
RuntimeStats.encoderUsedCoreml: 1 when every offline ASR transcription in
the job ran its encoder on Core ML, 0 when any ran on ggml. The engine
reports per-call routing only for offline ASR, so the field is absent after
Sortformer diarization and streaming jobs. Export sidecars with
engines/parakeet/scripts/export-encoder-coreml.py, preferably from an f16
or f32 GGUF (a quantized source works, but its rounding is baked into the
sidecar every tier shares), from the
qvac-fabric-speech.cpp
tree at the ref speech-cpp pins; its
Parakeet backends guide
has the per-model export commands. The
Core ML RTF lanes record what the
TDT sidecar gains over Metal.
Staging Models
The addon loads weights from local paths only — it never downloads. Stage files with the bundled scripts, from the QVAC model registry, or by hand.
Whisper (HuggingFace):
npm run download-models # interactive picker into ./models/Parakeet, prebuilt GGUFs from the QVAC registry (fastest path):
npm run download-models:parakeet:registry # all types
npm run download-models:parakeet:registry -- -t tdt # just TDT
npm run download-models:parakeet:registry -- -t unifiedParakeet, converting NVIDIA .nemo yourself:
npm run setup-models:parakeet # venv + download + convert (all types, q8_0)
npm run setup-models:parakeet -- -t tdt # just TDT
npm run setup-models:parakeet -- -t unified # Unified RNN-T
npm run setup-models:parakeet -- -t eou -q f16 # full-precision EOUsetup-models:parakeet chains setup:venv → download-models:parakeet →
convert-models:parakeet and is idempotent. Each step is also flag-driven on
its own (scripts/setup-venv.sh, scripts/parakeet-download-models.sh,
scripts/convert-nemo.sh — all accept --help). The converter reads the
.nemo archive directly and does not need the heavy nemo_toolkit
package, but it does need sentencepiece to decode the tokenizer (without it
transcripts come out as raw token IDs). Full requirements:
scripts/requirements.txt.
Error Codes
Thrown errors are QvacErrorAddonASRGgml instances (extending
QvacErrorBase) carrying a numeric .code, so callers can match
programmatically. ERR_CODES is exposed as ASRGgml.ERR_CODES and the class
as ASRGgml.Error.
The map spans two ranges, because no historical code was allowed to move:
- Shared verbs are canonical in the whisper
6001–7000range —FAILED_TO_LOAD_WEIGHTS6001 …VAD_MODEL_NOT_FOUND6018, plus the codes the merge added:NOT_SUPPORTED6019,STREAMING_SESSION_ACTIVE6020,INVALID_ENGINE6021. - Parakeet-only names keep their historical
24001–25000numbers —MODEL_NOT_FOUND24009,INVALID_AUDIO_FORMAT24010,INVALID_CONFIG24015,INSTANCE_DESTROYED24018,JOB_CANCELLED24019. Shared verbs raised by the parakeet engine also stay in24xxx(a parakeet append failure is 24003, not 6003).
Every number from both historical tables is registered, so pre-merge codes
still resolve to a name and message. See
published error.d.ts for the full table and
the engine architecture document for the
rationale.
Development
Prerequisites
- CMake ≥ 3.25, Git with submodule support, a C++20 compiler
- Linux: Clang/LLVM 22 with libc++ (
clang libc++-dev libc++abi-dev) - macOS: Xcode command-line tools
- Windows: Visual Studio 2019+ or Build Tools
- Linux: Clang/LLVM 22 with libc++ (
- vcpkg — clone it, run
bootstrap-vcpkg.sh/.bat, and exportVCPKG_ROOT=/path/to/vcpkg - Optional GPU SDKs: Vulkan SDK on
Linux/Windows (
vulkan-tools libvulkan-dev vulkan-utility-libraries-dev spirv-toolson Ubuntu/Debian); Metal needs nothing on macOS/iOS; the opt-in CUDA build additionally needs the CUDA Toolkit (nvcc) on Linux/Windows
Build
git clone https://github.com/tetherto/qvac.git
cd qvac/packages/asr-ggml
git submodule update --init --recursive
npm install
npm run build # build:ts (TypeScript) + build:native (bare-make)
npm run build:cuda # same, with the CUDA backend compiled in (needs nvcc)build:native runs bare-make generate → bare-make build →
bare-make install. build:native:cuda is the same chain with
-D ASR_CUDA=ON on the generate step.
Test
npm test # complete standard gate
npm run test:all # same aggregate, named explicitly
npm run test:unit
npm run test:package # packed tarball and consumer contract
npm run test:integration # standard suites for both engines
npm run test:integration:whisper
npm run test:integration:parakeet
npm run test:cpp # native gtest suite
npm run lintThe standard integration gate covers representative transcription
correctness, streaming, validation, and lifecycle behavior. Accuracy,
long-audio, cold-start, GPU, and C++ suites stay explicit because they need
specialized models, hardware, timing conditions, or toolchains:
test:integration:accuracy, test:integration:long,
test:integration:cold-start, test:integration:gpu (Whisper),
test:integration:parakeet:gpu, and test:cpp.
The Parakeet GPU command is manual; ASR CI keeps
test:integration:gpu Whisper-only.
On CPU the standard suites run only brief inference, one short transcription
per test. Multi-run, long-audio and paced streaming tests run on the GPU and
are skipped when NO_GPU=true, as on the CPU-only CI rows.
test:integration:live-stream-simulation runs only the long-lived Whisper
stream test; the misspelled test:integration:live-stream-simultion remains
as a temporary alias.
Typical loop: npm install && npm run build && npm run test:integration.
Benchmarking
Accuracy (WER / CER / AraDiaWER) and RTF benchmarks live under
benchmarks/ and are engine-keyed throughout:
benchmarks/server/— one bare HTTP server;POST /runtakes anengine: "whisper" | "parakeet"discriminator.benchmarks/client/— one Python client;src/main.pydispatches on the config's required top-levelengine:key oversrc/whisper/andsrc/parakeet/.benchmarks/client/config/config-whisper*.yaml(incl. three Common Voice Arabic variants) andconfig-parakeet{,-unified,-ctc,-eou,-sortformer,-indic-conformer}.yaml.benchmarks/manual-results/{whisper,parakeet}/— drop RTF artifacts for backends CI cannot host.benchmarks/ci/— the HuggingFace → GGML conversion step the accuracy workflow uses for whisper'scustom_model_repoinput.
# accuracy
cd benchmarks/client && poetry install
poetry run python -u -m src.main --config config/config-whisper.yaml
poetry run python -u -m src.main --config config/config-parakeet.yaml
# RTF matrix (engine-keyed entries via QVAC_ASR_GGML_BENCHMARK_MATRIX_JSON)
npm run test:benchmark:rtf:matrixscripts/trigger-benchmark.sh -e whisper|parakeet dispatches the CI accuracy
workflow. Aggregated historical results:
benchmarks/results/results_summary.md.
Core ML (Apple Neural Engine) RTF lanes
On darwin the parakeet TDT matrix also has coreml lanes, which run the
FastConformer encoder on the Neural Engine via an exported
<stem>-encoder.mlmodelc sidecar while the TDT decoder stays on Metal. Add
"coreml": true to a parakeet matrix entry:
{ "engine": "parakeet", "modelType": "tdt", "quant": "f16", "useGPU": true, "coreml": true }Three things are worth knowing before touching these lanes:
- The sidecar is presence-driven. parakeet.cpp derives the sidecar path
from the GGUF path and strips a trailing quant tag, so one
parakeet-tdt-0.6b-v3-encoder.mlmodelcserves the f16/q8_0/q4_0 GGUFs. It therefore cannot live inmodels/— every CPU and Metal lane would silently start measuring the ANE. Core ML entries run against an isolatedmodels/coreml/copy instead, staged by the matrix runner. - The export is traced at one mel length. That length is the TDT
sidecar's fixed capacity: shorter utterances are zero-padded up to it and
longer ones are split into overlapping windows, so every length runs on the
ANE, but only a matching one runs as a single unpadded pass. The lanes'
sidecar is therefore sized to the benchmark's own sample —
examples/samples/sample.raw(20.13 s ⇒1 + 322137/160= 2014 mel frames). Changing that sample means re-exporting, or the lane measures padding or windowing. A variable-length (--flexible) export exists but places zero ops on the ANE, so it is for numerical checks only, never for benchmarking. - A lane can never publish a mislabelled number.
activeBackendis derived from each measured run'sencoderUsedCoremlstat, which reports where that run's encoder actually ran. The benchmark refuses to write an artifact when a Core ML lane has any run whose encoder fell back to ggml, or when a non-Core ML lane has any run on Core ML. The check runs before the artifact is written, because the artifact is written before the test's own assertions run.
Sidecars are pinned in
test/integration/parakeet-coreml.manifest.json
and staged by scripts/stage-integration-models.mjs. Until a bundle is
published there the Core ML lanes skip themselves loudly and the rest of the
matrix is unaffected.
To produce a sidecar, use export-encoder-coreml.py from the speech-cpp
source tree at the ref pinned in vcpkg.json, seeded with an f16 or f32
GGUF so the sidecar every quant tier shares carries unrounded weights:
python scripts/export-encoder-coreml.py \
--gguf models/parakeet-tdt-0.6b-v3.f16.gguf \
--n-mel-frames 2014 \
--out parakeet-tdt-0.6b-v3-encoder.mlpackage \
--compile-dir models/coremlIt prints the op placement; a good export is overwhelmingly NeuralEngine
(the reference export is NeuralEngine=1334, GPU=14, CPU=1).
Measured on an Apple M1 Pro (macOS 15.1.1, addon 0.4.2, sample.raw, 5 runs
per lane) — full artifacts in
benchmarks/manual-results/parakeet/:
| Quant | CPU | Metal | Core ML | ANE vs Metal | |-------|-----|-------|---------|--------------| | f16 | 0.09614 | 0.00751 | 0.00625 | 1.20x | | q8_0 | 0.04246 | 0.00847 | 0.00706 | 1.20x | | q4_0 | 0.04269 | 0.00718 | 0.00566 | 1.27x |
Mean RTF, lower is better. The ANE encoder costs peak RSS (~+120-180 MB over the Metal lane) and additional load time to initialise the sidecar, so it pays off across many utterances in one process rather than for a single short one.
Examples
Whisper:
examples/quickstart.js— basic transcription (npm run example:whisper -- [audioPath] [modelPath] [vadModelPath])examples/example.streaming-vad.js— VAD-segmentedrunStreaming()examples/example.mic-conversation.js— mic capture with VAD state and end-of-turn eventsexamples/example.live-transcription.js— small chunks into one long-lived jobexamples/example.audio-ctx-chunking.js— long recordings via per-chunkreload()examples/example.reload.js— reloading with a different language/temperatureexamples/example.decoder.js— the FFmpeg decoder standalone
Parakeet:
examples/parakeet-transcribe.js— CTC, TDT, Unified, EOU, or Sortformer transcriptionexamples/parakeet-unified-transcribe.js— batch transcription withparakeet-unified-en-0.6bexamples/parakeet-indic-conformer-transcribe.js— Indic Conformer transcription with the required--language <id>optionexamples/parakeet-diarized-transcribe.js— Sortformer + ASR, "who said what"examples/parakeet-live-mic.js— live mic via the duplex streaming sessionexamples/parakeet-live-mic-diarized.js— live mic with speaker tagsexamples/parakeet-live-mic-diarized-aosc.js— same, with the AOSC tuning knobs as CLI flagsexamples/parakeet-decode-audio.js— decode + transcribe any FFmpeg-supported container
The npm tarball includes the dependency-clean Whisper quickstart. The other examples are repository examples. Run their matching commands from a source checkout:
npm run example:whisper:streaming-vad
npm run example:whisper:mic
npm run example:whisper:live-transcription
npm run example:whisper:audio-ctx-chunking
npm run example:whisper:reload
npm run example:whisper:decoder
npm run example:parakeet
npm run example:parakeet:unified
npm run example:parakeet:indic-conformer
npm run example:parakeet:diarize
npm run example:parakeet:mic
npm run example:parakeet:mic-diarize
npm run example:parakeet:mic-diarize-aosc
npm run example:parakeet:decode-audioThe published quickstart uses the Bare global for arguments and exit handling
so it does not require the repository-only bare-process development
dependency.
The live-mic examples capture the default input device via sox -d
(brew install sox / apt install sox / choco install sox). With
npm run example:whisper -- ..., keep the -- separator or npm eats the flags.
Documentation
docs/engines.md— the orchestrator + driver layout, the native verb table, engine resolution, and how to add a third enginedocs/architecture.md— full architecture write-up (heritage: whisper engine, pre-merge naming)docs/data-flows-detailed.md— sequence diagrams for load / run / streaming / reload (heritage: whisper engine)docs/whisper-addon-help.md— whisper.cpp parameter referencedocs/PARAKEET-README.md— heritage@qvac/transcription-parakeetREADME; still the deepest reference for Sortformer/AOSC behaviour and the.nemo→.ggufpipelinedocs/WHISPER-CHANGELOG.md/docs/PARAKEET-CHANGELOG.md— the two pre-merge histories, preserved verbatimCHANGELOG.md— the merged package's history, starting at0.1.0
Glossary
- Bare — small, modular JavaScript runtime for desktop and mobile. Learn more.
- GGUF — single-file model format used by ggml-based runtimes; carries weights, tokenizer, and hyperparameters together.
- QVAC — Tether's open-source AI SDK for building decentralized AI applications.
- RTF — real-time factor: processing time divided by audio duration. Lower is better; below 1.0 is faster than real time.
- AOSC — Audio-Online Speaker Cache, the NeMo-derived mechanism that anchors Sortformer v2.1 speaker slots across silence.
- EOU — end of utterance; Parakeet's EOU model emits a native
<EOU>token at turn boundaries.
Resources
- NVIDIA Parakeet model cards — upstream
.nemocheckpoints - whisper.cpp GGML models — upstream whisper checkpoints
License
This project is licensed under the Apache-2.0 License — see LICENSE for details, and NOTICE for third-party components. Parakeet model files are distributed under the NVIDIA Open Model License; see the upstream HuggingFace model cards for the per-checkpoint terms.
For questions or issues, please open an issue on the GitHub repository.
