npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@qvac/asr-ggml

v0.7.0

Published

Multi-engine ASR (Whisper + NVIDIA Parakeet) inference addon for qvac on the Bare runtime

Readme

@qvac/asr-ggml

Multi-engine automatic speech recognition for QVAC runtime applications on the Bare runtime. One npm package and one native prebuild serve two ggml-based ASR engines behind a single class, ASRGgml:

| Engine | Native library | Good for | | --- | --- | --- | | Whisper | whisper.cpp | Multilingual offline transcription, translation, Silero-VAD-segmented live capture | | Parakeet | parakeet-cpp through the speech-cpp umbrella port (NVIDIA Parakeet / Sortformer) | Low-latency streaming ASR, native end-of-turn detection, 4-speaker diarization |

This package replaces @qvac/transcription-whispercpp and @qvac/transcription-parakeet. See CHANGELOG.md for the breaking changes the merge introduced.

Table of Contents

Supported Engines and Models

Whisper (engine: 'whisper')

Legacy single-file GGML .bin checkpoints from ggerganov/whisper.cpp:

| Model | Size | Description | |-------|------|-------------| | ggml-tiny.bin | 78 MB | Smallest, fastest | | ggml-base.bin | 148 MB | Balanced size/accuracy | | ggml-small.bin | 488 MB | Better accuracy | | ggml-medium.bin | 1.5 GB | High accuracy | | ggml-large-v3.bin | 3.1 GB | Best accuracy | | ggml-large-v3-turbo.bin | 1.6 GB | Best accuracy, faster |

Quantized variants (q8_0, q5_1, q5_0) exist for all sizes. Whisper covers ~99 languages plus translation-to-English, and the fine-tuned per-language checkpoints listed in NOTICE also load.

VAD model (required for runStreaming()), from ggml-org/whisper-vad:

| Model | Size | Description | |-------|------|-------------| | ggml-silero-v5.1.2.bin | 885 KB | Silero VAD for voice-activity detection |

Parakeet (engine: 'parakeet')

Single-file .gguf checkpoints. The model type is auto-detected from the GGUF metadata — there is no modelType to pass.

| Variant | Languages | Decoder | ~Size (q8_0) | Notes | |---------|-----------|---------|-------------:|-------| | CTC (parakeet-ctc-0.6b) | English | argmax CTC | ~700 MiB | Fast, no punctuation/capitalization | | TDT (parakeet-tdt-0.6b-v3) | ~25 | RNN-T greedy + duration | ~715 MiB | Recommended default; PnC + language auto-detect | | Unified (parakeet-unified-en-0.6b) | English | RNN-T | ~715 MiB | One checkpoint for batch and cache-aware streaming at 80/160/560/1040 ms; PnC | | EOU (parakeet-eou-120m-v1) | English | RNN-T greedy + <EOU> | ~132 MiB | Streaming-trained; native end-of-turn token | | Indic Conformer CTC (indic-conformer-ctc) | Indic aggregate | argmax CTC + language mask | ~701 MiB | Multilingual Indic; set parakeetConfig.language (e.g. "hi") | | Sortformer v1 (sortformer-4spk-v1) | n/a | Diarization head (sliding history) | ~141 MiB | 4-speaker. Default for offline diarization | | Sortformer v2.1 + AOSC (diar_streaming_sortformer_4spk-v2.1) | n/a | Diarization head + speaker cache | ~141 MiB | 4-speaker. Default for streaming diarization; AOSC anchors speaker slots across silence, auto-detected from GGUF metadata |

On macOS and iOS, TDT, Unified, EOU, and Sortformer v2.1 can also run their encoder on an optional Core ML sidecar; CTC and Indic Conformer CTC cannot with the pinned engine, and Sortformer v1 has no sidecar. See Core ML encoder sidecars.

Upstream .nemo checkpoints are NVIDIA's; see the Parakeet model cards for the per-checkpoint NVIDIA Open Model License terms.

Choosing a model

Pick a specific checkpoint, not only an engine. Whisper and Parakeet overlap on English batch transcription; they diverge on streaming semantics, language coverage, translation, and diarization.

Decision guide

| If you need… | Use this model | Notes | | --- | --- | --- | | Default multilingual / English ASR (batch or duplex stream) | parakeet-tdt-0.6b-v3 (q8_0 GGUF) | Recommended Parakeet default: ~25 languages, punctuation/capitalization, language auto-detect, low-latency streaming. | | English batch and low-latency streaming with one checkpoint | parakeet-unified-en-0.6b | Standard RNN-T with punctuation and capitalization; use when multilingual TDT or native EOU tokens are not required. Streaming uses the native cache-aware encoder: streamingChunkMs accepts 80, 160, 560, or 1040 and streamingRightLookaheadMs 0, 80, 160, 240, 320, 560, or 1040, both snapped down to the nearest trained value. Defaults to 560 ms. | | Native end-of-turn for conversational / duplex English | parakeet-eou-120m-v1 | Emits <EOU>; smallest Parakeet (~132 MiB). Pair with TDT when you need broader language coverage and EOU. | | Fast English-only, no punctuation | parakeet-ctc-0.6b | Lowest decode cost in the Parakeet family; no PnC. | | Indic-language ASR (Hindi and other Indic ids) | indic-conformer-ctc | Pass parakeetConfig.language (e.g. "hi"). Same Parakeet engine; GGUF lives under indic_conformer/ in the registry. | | Offline 4-speaker diarization | sortformer-4spk-v1 | Default offline diarization head. | | Streaming 4-speaker diarization | diar_streaming_sortformer_4spk-v2.1 | AOSC keeps speaker slots across silence; prefer over v1 for live streams. | | Broadest language set + translate-to-English | ggml-large-v3-turbo.bin (or ggml-small.bin on edge) | Whisper: ~99 languages, translation, Silero-VAD live capture. Turbo is the accuracy/speed sweet spot; use tiny/base only when size dominates. | | Live capture with VAD segmentation (Whisper path) | Whisper ASR model + ggml-silero-v5.1.2.bin | Silero VAD is required for Whisper runStreaming(). |

Engine vs model

| Engine | Prefer when… | Prefer the other when… | | --- | --- | --- | | Parakeet | Low-latency streaming, native EOU, diarization, English / ~25-lang product ASR | You need Whisper’s language breadth or translate-to-English | | Whisper | Multilingual offline, translation, VAD-segmented live capture on the whisper.cpp path | You need Parakeet EOU / Sortformer diarization or tighter streaming RTF |

Always pass config.engine (or top-level engine) explicitly in library and SDK code — see Engine Selection.

Supported Platforms

| Platform | Architecture | Min Version | Status | GPU Support | |----------|-------------|-------------|--------|-------------| | macOS | arm64, x64 | 14.0+ | ✅ Tier 1 | Metal | | iOS | arm64 | 17.0+ | ✅ Tier 1 | Metal | | Linux | arm64, x64 | Ubuntu-22+ | ✅ Tier 1 | Vulkan; CUDA (x64 prebuild; arm64 via build:cuda / ASR_CUDA=ON) | | Android | arm64 | 12+ | ✅ Tier 1 | Vulkan, OpenCL (Adreno) | | Windows | x64 | 10+ | ✅ Tier 1 | Vulkan; CUDA via build:cuda / ASR_CUDA=ON |

Dependencies:

  • qvac-lib-inference-addon-cpp — C++ addon framework (vcpkg port; version pinned in vcpkg.json)
  • speech-cpp — umbrella vcpkg port enabling the whisper and parakeet engine features plus each platform's GPU features
  • Bare runtime — see engines.bare in package.json
  • Linux requires Clang/LLVM 22 with libc++

Installation

Make sure the Bare runtime is installed:

npm install -g bare bare-make
bare -v   # must satisfy package.json engines.bare

Then:

npm install @qvac/asr-ggml

Platform packages

@qvac/asr-ggml is a meta package that ships the JavaScript wrapper only. The native prebuild for each desktop host lives in a version-locked platform package selected at install time through os/cpu filtered optionalDependencies:

| Host | Package | | --- | --- | | linux-x64 (glibc) | @qvac/asr-ggml-linux-x64 | | linux-arm64 (glibc) | @qvac/asr-ggml-linux-arm64 | | darwin-arm64 | @qvac/asr-ggml-darwin-arm64 | | darwin-x64 | @qvac/asr-ggml-darwin-x64 | | win32-x64 | @qvac/asr-ggml-win32-x64 |

Do not depend on desktop platform packages directly. Supported installers are npm 7+, pnpm, bun, and Yarn Berry. Yarn v1 and --omit=optional installs skip the platform package and fail at require time with an error naming the missing package; a locally built prebuilds/ directory in the package root always takes precedence. Use require('@qvac/asr-ggml').resolveBackendsDir() to locate the directory holding the host's prebuilt binaries and dynamically loaded ggml backends.

Mobile targets are cross-built, so no install host ever matches their os, and optionalDependencies filtering can never select them. Mobile applications must declare the target's platform package as a direct dependency, pinned to the exact @qvac/asr-ggml version:

| Target | Package | | --- | --- | | android-arm64 | @qvac/asr-ggml-android-arm64 | | ios (device + simulators) | @qvac/asr-ggml-ios |

{
  "dependencies": {
    "@qvac/asr-ggml": "x.y.z",
    "@qvac/asr-ggml-android-arm64": "x.y.z"
  }
}

Quickstart

All four snippets assume:

const fs = require('bare-fs')
const ASRGgml = require('@qvac/asr-ggml')

Audio must be 16 kHz mono. See Audio Input for the accepted shapes.

Whisper — batch run()

const model = new ASRGgml({
  files: { model: './models/ggml-tiny.bin' },
  config: {
    engine: 'whisper',
    whisperConfig: {
      language: 'en',
      audio_format: 's16le'   // how raw Uint8Array bytes are decoded
    }
  }
})

await model.load()

const audioStream = fs.createReadStream('./audio.raw', { highWaterMark: 16000 })
const response = await model.run(audioStream)

// Push-based: segments arrive as whisper.cpp emits them.
await response
  .onUpdate((out) => {
    for (const segment of Array.isArray(out) ? out : [out]) {
      console.log(segment.start, '→', segment.end, segment.text)
    }
  })
  .await()

await model.destroy()

Or pull-based, with iterate() instead of onUpdate():

const response = await model.run(audioStream)
for await (const out of response.iterate()) {
  for (const segment of Array.isArray(out) ? out : [out]) {
    console.log(segment.text)
  }
}

Whisper — VAD streaming runStreaming()

Silero VAD splits the incoming audio into utterances and each one is transcribed as it completes. A VAD model is required — pass it as files.vadModel, config.vadModelPath, or config.whisperConfig.vad_model_path; a missing file throws VAD_MODEL_NOT_FOUND (6018) from the constructor, and a missing path throws VAD_MODEL_REQUIRED (6009) from runStreaming().

const model = new ASRGgml({
  files: {
    model: './models/ggml-tiny.bin',
    vadModel: './models/ggml-silero-v5.1.2.bin'
  },
  config: { engine: 'whisper', whisperConfig: { language: 'en' } }
})

await model.load()

const response = await model.runStreaming(micStream, {
  emitVadEvents: true,      // adds { type: 'vad', speaking, score, source }
  endOfTurnSilenceMs: 800   // adds { type: 'endOfTurn', source: 'vad-silence' }
})

for await (const out of response.iterate()) {
  if (out.type === 'vad') console.log('speaking:', out.speaking)
  else if (out.type === 'endOfTurn') console.log('--- turn ended ---')
  else for (const s of Array.isArray(out) ? out : [out]) console.log(s.text)
}

Parakeet — batch run()

const model = new ASRGgml({
  files: { model: './models/parakeet-tdt-0.6b-v3.q8_0.gguf' },
  config: {
    engine: 'parakeet',
    parakeetConfig: { useGPU: true, maxThreads: 4 }
  }
})

await model.load()

const response = await model.run(float32Samples)   // Float32Array, 16 kHz mono
const segments = []
await response
  .onUpdate((out) => {
    for (const s of Array.isArray(out) ? out : [out]) {
      // `toAppend` marks a segment that continues the previous one rather
      // than replacing it.
      if (s.text && s.toAppend) segments.push(s)
    }
  })
  .await()

console.log(segments.map((s) => s.text).join(''))
await model.unload()

Buffer cap (run() only): every chunk of one run() call is normalized to Float32 and batched into a single native process() call. The cap is 500 MiB of caller-supplied audio — ≈4.55 h of 16 kHz mono s16le (the default byte format) or ≈2.27 h of 16 kHz mono f32le, which needs no expansion on the way to native. Exceeding it raises BUFFER_LIMIT_EXCEEDED (6015), whose message names the source format the budget is denominated in. For longer captures use runStreaming(), which feeds the engine as audio arrives, or split into sequential run() calls.

Parakeet — duplex streaming runStreaming()

runStreaming() opens one long-lived native session for the lifetime of the call and forwards each chunk as it arrives — no batching, no per-chunk session recreation. Speaker IDs stay stable across appends, and an EOU model's <EOU> boundaries surface both as segment.isEndOfTurn and as a synthesized { type: 'endOfTurn', source: 'model-eou' } event.

const model = new ASRGgml({
  engine: 'parakeet',                                  // no config → alias form
  files: { model: './models/parakeet-eou-120m-v1.q8_0.gguf' }
})

await model.load()

const response = await model.runStreaming(micStream, { chunkMs: 480 })
await response
  .onUpdate((out) => {
    if (out.type === 'endOfTurn') return console.log('--- turn ended ---')
    for (const s of Array.isArray(out) ? out : [out]) process.stdout.write(s.text)
  })
  .await()

Only one streaming session may be open per instance; a concurrent run() or runStreaming() throws STREAMING_SESSION_ACTIVE (6020).

Engine Selection

The engine is resolved once, in the constructor, from three sources in strict precedence order:

  1. config.engine — authoritative, and the recommended form. If config is supplied at all, it must carry engine; a missing or unrecognized value throws INVALID_ENGINE (6021).
  2. engine — a top-level convenience alias, used only when config is omitted entirely. Unrecognized values throw INVALID_ENGINE.
  3. Model-file sniffing — last resort when neither is given. The first four bytes of files.model are read: GGUF ⇒ parakeet, anything else ⇒ whisper.

Sniffing is a convenience for scripts, not an integration path. It opens the model file synchronously inside the constructor, it cannot distinguish a GGUF whisper build from a GGUF parakeet build, and it throws INVALID_ENGINE if the file exists but cannot be read (a missing file is reported first, as MODEL_NOT_FOUND (24009)). Library and SDK callers should always pass config.engine explicitly.

Validation and sniffing both target the file the driver actually opens: for whisper that is config.path when set, otherwise files.model; parakeet only ever loads files.model.

getEngineType() reports the resolved engine; ASRGgml.ENGINE_WHISPER and ASRGgml.ENGINE_PARAKEET are available as statics.

API Surface

Every verb has one signature and one meaning regardless of engine.

| Member | Description | | --- | --- | | new ASRGgml({ files, config?, engine?, enableStats?, logger?, exclusiveRun? }) | Resolves the engine, validates model files and the engine config vocabulary. Throws on any problem — nothing is deferred to load(). | | load() | Creates the native instance and activates the model. Calling it on a loaded instance unloads first. Throws INSTANCE_DESTROYED after destroy(). | | run(audio) | Batch transcription. Returns a QvacResponse; drain it with onUpdate(cb) (push) or iterate() (pull). | | runStreaming(audio, opts?) | Duplex/VAD-segmented streaming. Resolves once the native session is open; opts is the engine's streaming vocabulary. | | reload(newConfig?) | Applies an engine-scoped partial config in place where possible. Rejects with NOT_SUPPORTED (6019) on an engine whose driver has no native reload. | | cancel(jobId?) | Cancels the active job and fails it, so a draining iterate() throws. The native verb takes no id; jobId is accepted for source compatibility only. | | status() | Native state string. Before load(), rejects with FAILED_TO_GET_STATUS: 6004 for Whisper and 24004 for Parakeet. | | addon | The native interface, or undefined before load() (not cleared by unload(), as in both pre-merge packages). Escape hatch for a native hard cancel that stops the decode without failing the job (what the SDK's model-wide cancel uses). Not otherwise part of the supported surface. | | unload() / destroy() | Release the model / retire the instance. | | getState() | { configLoaded, weightsLoaded, destroyed }. | | getEngineType() | 'whisper' | 'parakeet'. | | getBackendInfo() | BackendInfo or null before load(). | | pause() / unpause() | Always reject with NOT_SUPPORTED (6019). Neither engine implements a correct pause/resume. |

Constructor options:

| Option | Default | Description | | --- | --- | --- | | files.model | — | Required. Path to the .bin (whisper) or .gguf (parakeet) checkpoint. Must exist. | | files.vadModel | — | Whisper streaming only: path to the Silero VAD model. Must exist if given. | | config | — | Engine-scoped config; must carry engine when present. | | engine | — | Alias for config.engine, used when config is omitted. | | enableStats | true | Attach RuntimeStats to the job-end payload. | | logger | null | A @qvac/logging LoggerInterface. | | exclusiveRun | true | Serialize run() calls to completion on the inference lane; runStreaming() holds that lane for session setup only. reload/unload/destroy serialize on a separate lifecycle lane, so teardown pre-empts an in-flight run() instead of queueing behind it. |

run() / runStreaming() output payloads are:

  • bare TranscriptionSegment[] (or a single segment) for transcripts — there is no { type: 'segment' } wrapper;
  • { type: 'vad', speaking, score, source, timestamp?, speakerId? } for voice-activity events, emitted on each speech/silence change. source is 'silero' (Whisper streaming), 'energy' (Parakeet ASR energy detector; score is the window RMS), or 'sortformer' (Parakeet speaker activity; speakerId names the dominant speaker when speech starts). Parakeet events carry timestamp, in seconds from the start of the session;
  • { type: 'endOfTurn', source, silenceDurationMs? } for turn boundaries.

Engine-specific segment fields:

  • Whisper segments carry noSpeechProb, and language whenever whisper reports one (the decoded language, the detected one under language: 'auto'). With token_timestamps: true they add tokens: [{ text, start, end, probability }] (special tokens left out); with tdrz_enable: true on a tinydiarize model they add speakerTurnNext.
  • Parakeet Sortformer keeps the "Speaker N: start - end" text and adds the same data in structured form: a streamed diarization segment carries speakerId, and the offline transcript carries speakerSegments: [{ speakerId, start, end }], one entry per text line.

Assessing fit

assessFit projects a load against the memory free right now. It reads model metadata and never weight data. Parakeet is GGUF, so the registry's weightless copy of a parakeet model answers the same as the model itself and the projection can run before it is downloaded. Whisper ships as .bin, which the registry has no weightless form for, so a whisper projection needs the file. It is a module export, not an instance method — nothing is loaded to call it.

const ASRGgml = require('@qvac/asr-ggml')

const fit = ASRGgml.assessFit({
  engine: 'whisper',
  modelPath: '/models/whisper.bin',
  vadModelPath: '/models/silero-vad.bin',
  audioSeconds: 300,
  decoders: 5
})

fit.status // 'fits' | 'does-not-fit' | 'error'
fit.reason // the engine's own wording, e.g. 'model-unreadable', 'workload-too-large'
fit.modelType // whisper: 'tiny' … 'large v3'; parakeet: 'ctc' | 'rnnt' | 'tdt' | 'eou' | 'nemotron' | 'sortformer'
fit.deviceName
fit.deviceBytes
fit.weightsBytes
fit.hostBytes
fit.report

engine picks the fitter and defaults to parakeet. Each engine fills its own breakdown on the result: whisper reports kvBytes, computeBytes, vadBytes and hostOverflowBytes; parakeet reports encoderComputeBytes, decoderStateBytes and decoderComputeBytes.

| Option | Description | | --- | --- | | modelPath | Required. Absolute path to the model, or to the registry's weightless copy where one exists. | | audioSeconds | Longest single transcribe the projection must cover. Defaults to 300. | | gpuLayers | Greater than 0 requests the GPU stack, with the fallbacks a real load applies. Omitted, parakeet projects on the CPU and whisper on the GPU, matching what each load does. | | marginBytes | Free memory that must remain for the projection to count as fitting. Defaults to the engine's own headroom, which is 256 MiB for parakeet. | | backendsDir | The prebuilds root. The backends are read from the per-target subdir under it, the same path a load reads. | | vadModelPath | Whisper: projected alongside the model. Omitted means no VAD. | | decoders | Whisper: worst-case resident decoders, the best_of or beam_size the run will use. The KV cache and decode graph grow with it. | | flashAttn, gpuDevice | Whisper: as the load takes them. | | threads, longFormWindowFrames, longFormContextFrames | Parakeet: as the load takes them. | | nemotronChunkMs | Nemotron: the streaming operating point the projection must also cover. 0 projects the largest allowed one. |

A model the fitter cannot read is status: "error" with the engine's reason; only a broken request throws.

Configuration Reference

Configuration vocabularies are engine-scoped — there is no merged config namespace. A key belonging to one engine is rejected, not ignored, by the other, and unknown keys throw at construction time.

Whisper: config.whisperConfig / contextParams / miscConfig

const config = {
  engine: 'whisper',
  contextParams: {
    model: './models/ggml-tiny.bin',
    use_gpu: true,      // opt-in; defaults to false
    gpu_device: 0
  },
  whisperConfig: {
    language: 'en',     // 'auto' for language detection
    duration_ms: 0,
    temperature: 0.0,
    suppress_nst: true,
    n_threads: 0,
    audio_format: 's16le',
    vad_model_path: './models/ggml-silero-v5.1.2.bin',
    vad_params: {
      threshold: 0.6,
      min_speech_duration_ms: 250,
      min_silence_duration_ms: 200
    }
  },
  miscConfig: { caption_enabled: false }
}

The public Whisper keys, including vad_params, are declared by WhisperConfig. The constructor accepts a curated subset of whisper_full_params covering decoder strategy, thresholds, timestamps, VAD, and prompts. WhisperDriver maps vad_params to the native vadParams object; callers must not pass vadParams. Any key outside the list throws from the constructor. contextParams accepts only model, use_gpu, flash_attn, gpu_device; miscConfig only caption_enabled, seed. For what each flag means see the upstream whisper_full_params reference.

Notes:

  • GPU is opt-in. use_gpu defaults to false; set it in contextParams.
  • Four context keys force a full reload — model, use_gpu, flash_attn, gpu_device, main-gpu, main_gpu. Changing any of them destroys and rebuilds the whisper context (seconds, depending on model size). Everything in whisperConfig is applied in place.
  • backendsDir (in whisperConfig) overrides where dynamically-loaded ggml backend libraries are found. See Backends and GPU Acceleration.
  • max_seconds is a convenience that derives duration_ms.
  • carry_initial_prompt: true prepends initial_prompt to every decode window instead of only the first.

Whisper runStreaming(audio, opts) options:

| Option | Description | |--------|-------------| | emitVadEvents / conversationMode | Emit { type: 'vad' } as speech starts and stops. | | endOfTurnSilenceMs | Emit { type: 'endOfTurn' } after this much trailing silence. | | vadRunIntervalMs | How often the VAD is evaluated. |

Parakeet: config.parakeetConfig

const config = {
  engine: 'parakeet',
  parakeetConfig: {
    useGPU: true,
    maxThreads: 4,
    timestampsEnabled: true,
    streaming: false     // true opens a session at load time
  }
}

The authoritative vocabulary is ParakeetConfig in published driver.d.ts — every key is documented inline there, and any key outside it throws INVALID_CONFIG (24015) from the constructor. Groups:

| Group | Keys | | --- | --- | | Compute | maxThreads, useGPU, seed | | Audio | sampleRate (16000), channels (1) | | Output | captionEnabled, timestampsEnabled | | Language | language — multilingual CTC id (e.g. "hi"); required for Indic Conformer GGUFs that advertise parakeet.ctc.lang_* ranges; ignored on monolingual CTC | | Streaming (ASR) | streaming, streamingChunkMs, streamingEmitPartials, streamingEnergyVad, streamingEnergyVadThresholdDb, streamingEnergyVadWindowMs, streamingEnergyVadHangoverMs, streamingLeftContextMs, streamingRightLookaheadMs | | Streaming (Sortformer) | streamingHistoryMs, streamingSpeakerVad, streamingSpkCacheEnable, streamingSpkCacheLen, streamingFifoLen, streamingChunkLeftContextMs, streamingChunkRightContextMs, streamingSpkCacheUpdatePeriod | | Diarization (Sortformer, offline and streaming) | diarizationThreshold (0.641), diarizationMinSegmentMs (510) | | Engine | prewarm, prewarmAudioSeconds, longFormWindowFrames, longFormContextFrames | | Backends | backendsDir, openclCacheDir |

The streamingSpkCache* / streamingFifoLen / streamingChunk{Left,Right}ContextMs defaults are the NeMo-port tuning parakeet-cpp ships — keep them unless you are A/B comparing AOSC against the v1 sliding-window path. There is no modelType: CTC / TDT / EOU / Sortformer, and Sortformer v1 vs v2.1+AOSC, are all detected from the GGUF metadata.

Parakeet runStreaming(audio, opts) options are per-call overrides of the same knobs without the streaming prefix: chunkMs, historyMs, leftContextMs, rightLookaheadMs, emitPartials, emitEnergyVad, energyVadThresholdDb, energyVadWindowMs, energyVadHangoverMs, emitSpeakerVad, diarizationThreshold, diarizationMinSegmentMs, spkCacheEnable, spkCacheLen, fifoLen, chunkLeftContextMs, chunkRightContextMs, spkCacheUpdatePeriod.

Voice-activity events are opt-in. streamingEnergyVad / emitEnergyVad runs the engine's RMS energy detector on CTC, TDT, RNN-T, and Nemotron sessions (EOU models signal turns through <EOU> instead); the threshold is in dBFS (default -35) and the hangover is how long the audio must stay below it before the state returns to silence (default 200 ms). streamingSpeakerVad / emitSpeakerVad reports Sortformer speaker activity.

prewarm runs one synthetic encoder pass while loading so the first request does not pay the GPU shader or kernel compile. longFormWindowFrames bounds the offline encoder window (0 picks it from the model, a negative value always runs one pass), which keeps memory flat on long run() inputs; cancel() also takes effect between those windows.

For streaming diarization, use the Sortformer v2.1 GGUF. Its metadata enables AOSC automatically; keep the speaker-cache defaults unless you are comparing against the v1 sliding-window path. Sortformer v1 remains the offline diarization default. The conversion scripts support both variants and read NVIDIA .nemo archives directly.

Audio Input

Both engines take 16 kHz mono audio. run() and runStreaming() accept a stream, an iterable, a single chunk, or an array of chunks. A chunk's class decides how it is interpreted:

| Chunk type | Interpretation | | --- | --- | | Float32Array | f32 samples in [-1, 1] | | Int16Array | s16 samples | | Uint8Array | raw bytes, decoded as s16le by default; whisper's whisperConfig.audio_format: 'f32le' switches the byte interpretation |

audio_format accepts 's16le', 'f32le' and 'decoded' (an alias for 'f32le'); anything else throws INVALID_AUDIO_FORMAT (24010) rather than being decoded as little-endian s16. It only ever describes raw Uint8Array bytes — it never reinterprets a typed array.

A byte length that is not a whole number of samples raises INVALID_AUDIO_INPUT (6011). In a stream of byte chunks the check is applied in aggregate, so a sample split across a chunk boundary (arbitrary socket/pipe read sizes) is carried over into the next chunk; only a stream that ends mid-sample is rejected.

Backends and GPU Acceleration

GPU backends are selected per platform via vcpkg.json features; no bare-make generate flag is needed:

  • Linux / Windows — Vulkan (needs the Vulkan SDK on the build host); the linux-x64 prebuild additionally bundles CUDA, see below
  • Android — Vulkan + OpenCL (Adreno) as dynamically-loaded .so backends shipped beside the prebuild
  • macOS / iOS — Metal, statically linked, plus the optional Parakeet Core ML encoder sidecars

CUDA (Linux / Windows on NVIDIA) needs nvcc on the build host, so it is gated behind the ASR_CUDA CMake option — supported on linux-x64, linux-arm64 and win32-x64. The published linux-x64 prebuild turns it on (the prebuild workflow installs the CUDA toolkit); elsewhere build it yourself with npm run build:cuda (or bare-make generate -D ASR_CUDA=ON). The option adds the cuda feature to the speech-cpp dependency and turns on GGML_CUDA. Every linux-x64 and linux-arm64 build, and win32-x64 with the cuda feature, uses ggml's hybrid dynamically-loaded backend mode: the per-arch CPU-variant and Vulkan backends ship as runtime-loaded modules (.so on Linux, .dll on Windows) next to the addon, the cuda builds add the CUDA module, and only that module depends on the CUDA runtime. Engaging CUDA requires the NVIDIA driver plus the CUDA 13 runtime libraries (cudart and cuBLAS) resolvable at load time; hosts that cannot resolve them — including CPU-only and non-NVIDIA machines — skip the module and fall back to Vulkan or CPU instead of failing to load the addon. CUDA is compiled alongside Vulkan rather than replacing it; ggml registers CUDA ahead of Vulkan, so a use_gpu / useGPU request lands on CUDA when a supported device is present and falls back to Vulkan otherwise. Both engines report the winner through getBackendInfo() as backendId: 2 (BackendId.CUDA).

On x64 a CUDA build's module targets compute capability 7.5 and newer, with native code for Turing (7.5 — RTX 20xx, GTX 16xx, T4), Ampere (8.0, 8.6), Ada (8.9), Hopper (9.0) and Blackwell (12.0, 12.1). Anything newer JIT-compiles from the bundled 8.0 PTX on first use, a one-off compile the driver caches. On linux-arm64 the native set is Jetson Orin (8.7), Grace-Hopper (9.0) and GB10 / DGX Spark (12.1), with discrete Ampere+ cards and newer parts covered through the bundled 8.0 PTX. Volta and Pascal fall outside CUDA 13's support entirely, so they have no code path here: the backend skips such devices at registration and the addon falls back to Vulkan or CPU.

The addon takes no direct CUDA linkage — the CUDA module carries its own CUDA DT_NEEDED entries, which is what makes the graceful fallback possible — and nvcc's clang host-compiler setup lives in vcpkg-overlays/toolchains/linux-clang.cmake, shared by every addon that compiles the CUDA backend.

Both engines default to CPU: whisper needs contextParams.use_gpu: true, parakeet needs parakeetConfig.useGPU: true.

For Whisper GPU selection, set contextParams['main-gpu'] (or the alias contextParams.main_gpu) to a raw ggml registry index, an integer string, 'dedicated', or 'integrated' (class names are case-insensitive). With GPU enabled and no explicit selector, dedicated GPUs are preferred. A class selector is strict: if that class is unavailable, execution falls back to CPU. An in-range numeric selector preserves its registry identity before backend filtering; a CPU, excluded backend, or refused Adreno Vulkan slot falls back to CPU without selecting another GPU. An out-of-range index logs a warning and uses normal selection. The supported local families are Metal, CUDA, Vulkan, and OpenCL; the existing Adreno OpenCL guard still applies.

main-gpu does not enable GPU execution by itself. use_gpu: false always selects CPU. The legacy contextParams.gpu_device retains its existing Whisper GPU/IGPU-ordinal meaning and Adreno guard. Combining it with either new selector spelling, or supplying both new spellings, is rejected.

This selector currently applies to the Whisper engine.

getBackendInfo() reports what actually ran — backendName, backendId (see the BackendId enum), string backendDevice, backendDescription, encoderBackend, encoderOnCoreml (Apple: whether a Parakeet Core ML encoder sidecar loaded; see Core ML encoder sidecars), and modelType ('whisper', or the Parakeet family detected from the GGUF: 'ctc', 'tdt', 'rnnt', 'eou', 'nemotron', 'sortformer'). Whisper additionally reports gpuMemTotalMb / gpuMemFreeMb. This differs from RuntimeStats.backendDevice, which is the native numeric device-class code.

Two paths matter on Android and Linux:

  • backendsDir (in whisperConfig / parakeetConfig) — root directory holding dynamically-loaded ggml backend libraries (CUDA, Vulkan, OpenCL, per-arch CPU variants). Defaults to resolveBackendsDir(): the package's own prebuilds/ when present (local builds, mobile flatten), otherwise the installed platform package (see Platform packages); the native addon appends <bare-target>/<module-name> before scanning. Pass an explicit path when backend libraries ship elsewhere — e.g. Android's ApplicationInfo.nativeLibraryDir when they are packaged inside the APK. No-op on Apple, where backends are statically linked.
  • openclCacheDir (parakeet) — persistent directory for ggml-opencl's compiled program-binary cache. Android-only; pass the host app's cache directory to avoid a cold clBuildProgram on every process start.

Core ML encoder sidecars (Apple)

The macOS and iOS prebuilds are built with the speech-cpp coreml feature, so a Parakeet model can run its FastConformer encoder on Apple Core ML (the Neural Engine) while mel preprocessing and the decoder or speaker head stay on the ggml backend. It is opt-in by presence: at load() the engine looks for a compiled <stem>-encoder.mlmodelc next to the GGUF, where <stem> is the GGUF name with its quantization suffix stripped, so one sidecar serves every tier (parakeet-tdt-0.6b-v3.q8_0.gguf and .f16.gguf both resolve to parakeet-tdt-0.6b-v3-encoder.mlmodelc). The published models ship without sidecars, so a model directory behaves as before until you stage one. A missing sidecar, an input shape it does not take, or a failed prediction falls back to ggml, and PARAKEET_COREML_DISABLE=1 in the process environment forces ggml.

| Model | Sidecar | Inputs it takes | Runs on ggml instead | | --- | --- | --- | --- | | TDT (parakeet-tdt-0.6b-v3) | <stem>-encoder.mlmodelc | any length: shorter inputs are zero-padded to the compiled shape, longer offline inputs are split into overlapping windows | only on fallback | | Unified (parakeet-unified-en-0.6b) | <stem>-encoder.mlmodelc | run(), padded or windowed like TDT | runStreaming(), which uses the cache-aware encoder | | EOU (parakeet-eou-120m-v1) | <stem>-encoder.mlmodelc | inputs of exactly the compiled mel length | every other length, including streaming windows of a different length | | Sortformer v2.1 + AOSC | <stem>-encoder.mlmodelc (batch), <stem>-encoder-bypass-pre-encode.mlmodelc (AOSC) | batch: exactly the compiled length; AOSC: slabs up to the masked capacity (410 encoder frames for the default geometry) | other batch lengths, larger AOSC slabs; the speaker head always | | CTC (parakeet-ctc-0.6b), Indic Conformer CTC | none in the pinned speech-cpp | — | always (the engine adds CTC sidecars from speech-cpp 2026-09-24) | | Sortformer v1 | none | — | always | | Whisper | none: the whisper feature builds without WHISPER_COREML | — | always |

getBackendInfo().encoderOnCoreml (with encoderBackend: 'coreml') and RuntimeStats.encoderOnCoreml report that a sidecar loaded at load(), not that a given call ran on it: an EOU input of another length, a Unified runStreaming() session, or a failed prediction still runs the encoder on ggml with the flag set. For what a job actually did, read RuntimeStats.encoderUsedCoreml: 1 when every offline ASR transcription in the job ran its encoder on Core ML, 0 when any ran on ggml. The engine reports per-call routing only for offline ASR, so the field is absent after Sortformer diarization and streaming jobs. Export sidecars with engines/parakeet/scripts/export-encoder-coreml.py, preferably from an f16 or f32 GGUF (a quantized source works, but its rounding is baked into the sidecar every tier shares), from the qvac-fabric-speech.cpp tree at the ref speech-cpp pins; its Parakeet backends guide has the per-model export commands. The Core ML RTF lanes record what the TDT sidecar gains over Metal.

Staging Models

The addon loads weights from local paths only — it never downloads. Stage files with the bundled scripts, from the QVAC model registry, or by hand.

Whisper (HuggingFace):

npm run download-models                 # interactive picker into ./models/

Parakeet, prebuilt GGUFs from the QVAC registry (fastest path):

npm run download-models:parakeet:registry              # all types
npm run download-models:parakeet:registry -- -t tdt    # just TDT
npm run download-models:parakeet:registry -- -t unified

Parakeet, converting NVIDIA .nemo yourself:

npm run setup-models:parakeet                  # venv + download + convert (all types, q8_0)
npm run setup-models:parakeet -- -t tdt        # just TDT
npm run setup-models:parakeet -- -t unified    # Unified RNN-T
npm run setup-models:parakeet -- -t eou -q f16 # full-precision EOU

setup-models:parakeet chains setup:venv → download-models:parakeet → convert-models:parakeet and is idempotent. Each step is also flag-driven on its own (scripts/setup-venv.sh, scripts/parakeet-download-models.sh, scripts/convert-nemo.sh — all accept --help). The converter reads the .nemo archive directly and does not need the heavy nemo_toolkit package, but it does need sentencepiece to decode the tokenizer (without it transcripts come out as raw token IDs). Full requirements: scripts/requirements.txt.

Error Codes

Thrown errors are QvacErrorAddonASRGgml instances (extending QvacErrorBase) carrying a numeric .code, so callers can match programmatically. ERR_CODES is exposed as ASRGgml.ERR_CODES and the class as ASRGgml.Error.

The map spans two ranges, because no historical code was allowed to move:

  • Shared verbs are canonical in the whisper 6001–7000 range — FAILED_TO_LOAD_WEIGHTS 6001 … VAD_MODEL_NOT_FOUND 6018, plus the codes the merge added: NOT_SUPPORTED 6019, STREAMING_SESSION_ACTIVE 6020, INVALID_ENGINE 6021.
  • Parakeet-only names keep their historical 24001–25000 numbers — MODEL_NOT_FOUND 24009, INVALID_AUDIO_FORMAT 24010, INVALID_CONFIG 24015, INSTANCE_DESTROYED 24018, JOB_CANCELLED 24019. Shared verbs raised by the parakeet engine also stay in 24xxx (a parakeet append failure is 24003, not 6003).

Every number from both historical tables is registered, so pre-merge codes still resolve to a name and message. See published error.d.ts for the full table and the engine architecture document for the rationale.

Development

Prerequisites

  • CMake ≥ 3.25, Git with submodule support, a C++20 compiler
    • Linux: Clang/LLVM 22 with libc++ (clang libc++-dev libc++abi-dev)
    • macOS: Xcode command-line tools
    • Windows: Visual Studio 2019+ or Build Tools
  • vcpkg — clone it, run bootstrap-vcpkg.sh/.bat, and export VCPKG_ROOT=/path/to/vcpkg
  • Optional GPU SDKs: Vulkan SDK on Linux/Windows (vulkan-tools libvulkan-dev vulkan-utility-libraries-dev spirv-tools on Ubuntu/Debian); Metal needs nothing on macOS/iOS; the opt-in CUDA build additionally needs the CUDA Toolkit (nvcc) on Linux/Windows

Build

git clone https://github.com/tetherto/qvac.git
cd qvac/packages/asr-ggml
git submodule update --init --recursive
npm install
npm run build          # build:ts (TypeScript) + build:native (bare-make)
npm run build:cuda     # same, with the CUDA backend compiled in (needs nvcc)

build:native runs bare-make generate → bare-make build → bare-make install. build:native:cuda is the same chain with -D ASR_CUDA=ON on the generate step.

Test

npm test                              # complete standard gate
npm run test:all                      # same aggregate, named explicitly
npm run test:unit
npm run test:package                  # packed tarball and consumer contract
npm run test:integration              # standard suites for both engines
npm run test:integration:whisper
npm run test:integration:parakeet
npm run test:cpp                      # native gtest suite
npm run lint

The standard integration gate covers representative transcription correctness, streaming, validation, and lifecycle behavior. Accuracy, long-audio, cold-start, GPU, and C++ suites stay explicit because they need specialized models, hardware, timing conditions, or toolchains: test:integration:accuracy, test:integration:long, test:integration:cold-start, test:integration:gpu (Whisper), test:integration:parakeet:gpu, and test:cpp. The Parakeet GPU command is manual; ASR CI keeps test:integration:gpu Whisper-only. On CPU the standard suites run only brief inference, one short transcription per test. Multi-run, long-audio and paced streaming tests run on the GPU and are skipped when NO_GPU=true, as on the CPU-only CI rows. test:integration:live-stream-simulation runs only the long-lived Whisper stream test; the misspelled test:integration:live-stream-simultion remains as a temporary alias.

Typical loop: npm install && npm run build && npm run test:integration.

Benchmarking

Accuracy (WER / CER / AraDiaWER) and RTF benchmarks live under benchmarks/ and are engine-keyed throughout:

  • benchmarks/server/ — one bare HTTP server; POST /run takes an engine: "whisper" | "parakeet" discriminator.
  • benchmarks/client/ — one Python client; src/main.py dispatches on the config's required top-level engine: key over src/whisper/ and src/parakeet/.
  • benchmarks/client/config/config-whisper*.yaml (incl. three Common Voice Arabic variants) and config-parakeet{,-unified,-ctc,-eou,-sortformer,-indic-conformer}.yaml.
  • benchmarks/manual-results/{whisper,parakeet}/ — drop RTF artifacts for backends CI cannot host.
  • benchmarks/ci/ — the HuggingFace → GGML conversion step the accuracy workflow uses for whisper's custom_model_repo input.
# accuracy
cd benchmarks/client && poetry install
poetry run python -u -m src.main --config config/config-whisper.yaml
poetry run python -u -m src.main --config config/config-parakeet.yaml

# RTF matrix (engine-keyed entries via QVAC_ASR_GGML_BENCHMARK_MATRIX_JSON)
npm run test:benchmark:rtf:matrix

scripts/trigger-benchmark.sh -e whisper|parakeet dispatches the CI accuracy workflow. Aggregated historical results: benchmarks/results/results_summary.md.

Core ML (Apple Neural Engine) RTF lanes

On darwin the parakeet TDT matrix also has coreml lanes, which run the FastConformer encoder on the Neural Engine via an exported <stem>-encoder.mlmodelc sidecar while the TDT decoder stays on Metal. Add "coreml": true to a parakeet matrix entry:

{ "engine": "parakeet", "modelType": "tdt", "quant": "f16", "useGPU": true, "coreml": true }

Three things are worth knowing before touching these lanes:

  • The sidecar is presence-driven. parakeet.cpp derives the sidecar path from the GGUF path and strips a trailing quant tag, so one parakeet-tdt-0.6b-v3-encoder.mlmodelc serves the f16/q8_0/q4_0 GGUFs. It therefore cannot live in models/ — every CPU and Metal lane would silently start measuring the ANE. Core ML entries run against an isolated models/coreml/ copy instead, staged by the matrix runner.
  • The export is traced at one mel length. That length is the TDT sidecar's fixed capacity: shorter utterances are zero-padded up to it and longer ones are split into overlapping windows, so every length runs on the ANE, but only a matching one runs as a single unpadded pass. The lanes' sidecar is therefore sized to the benchmark's own sample — examples/samples/sample.raw (20.13 s ⇒ 1 + 322137/160 = 2014 mel frames). Changing that sample means re-exporting, or the lane measures padding or windowing. A variable-length (--flexible) export exists but places zero ops on the ANE, so it is for numerical checks only, never for benchmarking.
  • A lane can never publish a mislabelled number. activeBackend is derived from each measured run's encoderUsedCoreml stat, which reports where that run's encoder actually ran. The benchmark refuses to write an artifact when a Core ML lane has any run whose encoder fell back to ggml, or when a non-Core ML lane has any run on Core ML. The check runs before the artifact is written, because the artifact is written before the test's own assertions run.

Sidecars are pinned in test/integration/parakeet-coreml.manifest.json and staged by scripts/stage-integration-models.mjs. Until a bundle is published there the Core ML lanes skip themselves loudly and the rest of the matrix is unaffected.

To produce a sidecar, use export-encoder-coreml.py from the speech-cpp source tree at the ref pinned in vcpkg.json, seeded with an f16 or f32 GGUF so the sidecar every quant tier shares carries unrounded weights:

python scripts/export-encoder-coreml.py \
  --gguf models/parakeet-tdt-0.6b-v3.f16.gguf \
  --n-mel-frames 2014 \
  --out parakeet-tdt-0.6b-v3-encoder.mlpackage \
  --compile-dir models/coreml

It prints the op placement; a good export is overwhelmingly NeuralEngine (the reference export is NeuralEngine=1334, GPU=14, CPU=1).

Measured on an Apple M1 Pro (macOS 15.1.1, addon 0.4.2, sample.raw, 5 runs per lane) — full artifacts in benchmarks/manual-results/parakeet/:

| Quant | CPU | Metal | Core ML | ANE vs Metal | |-------|-----|-------|---------|--------------| | f16 | 0.09614 | 0.00751 | 0.00625 | 1.20x | | q8_0 | 0.04246 | 0.00847 | 0.00706 | 1.20x | | q4_0 | 0.04269 | 0.00718 | 0.00566 | 1.27x |

Mean RTF, lower is better. The ANE encoder costs peak RSS (~+120-180 MB over the Metal lane) and additional load time to initialise the sidecar, so it pays off across many utterances in one process rather than for a single short one.

Examples

Whisper:

Parakeet:

The npm tarball includes the dependency-clean Whisper quickstart. The other examples are repository examples. Run their matching commands from a source checkout:

npm run example:whisper:streaming-vad
npm run example:whisper:mic
npm run example:whisper:live-transcription
npm run example:whisper:audio-ctx-chunking
npm run example:whisper:reload
npm run example:whisper:decoder
npm run example:parakeet
npm run example:parakeet:unified
npm run example:parakeet:indic-conformer
npm run example:parakeet:diarize
npm run example:parakeet:mic
npm run example:parakeet:mic-diarize
npm run example:parakeet:mic-diarize-aosc
npm run example:parakeet:decode-audio

The published quickstart uses the Bare global for arguments and exit handling so it does not require the repository-only bare-process development dependency.

The live-mic examples capture the default input device via sox -d (brew install sox / apt install sox / choco install sox). With npm run example:whisper -- ..., keep the -- separator or npm eats the flags.

Documentation

Glossary

  • Bare — small, modular JavaScript runtime for desktop and mobile. Learn more.
  • GGUF — single-file model format used by ggml-based runtimes; carries weights, tokenizer, and hyperparameters together.
  • QVAC — Tether's open-source AI SDK for building decentralized AI applications.
  • RTF — real-time factor: processing time divided by audio duration. Lower is better; below 1.0 is faster than real time.
  • AOSC — Audio-Online Speaker Cache, the NeMo-derived mechanism that anchors Sortformer v2.1 speaker slots across silence.
  • EOU — end of utterance; Parakeet's EOU model emits a native <EOU> token at turn boundaries.

Resources

License

This project is licensed under the Apache-2.0 License — see LICENSE for details, and NOTICE for third-party components. Parakeet model files are distributed under the NVIDIA Open Model License; see the upstream HuggingFace model cards for the per-checkpoint terms.

For questions or issues, please open an issue on the GitHub repository.