@fluidinference/fluidaudio-web
v0.1.0
Published
Local speech AI for the browser — ASR (Parakeet, Whisper, Nemotron), TTS (Kokoro), VAD (Silero), speaker diarization (Sortformer) on hand-written WebGPU + WASM-SIMD kernels. No onnxruntime; model weights stream from Hugging Face and cache locally.
Readme
@fluidinference/fluidaudio-web
Local speech AI in the browser: ASR, TTS, VAD, and speaker diarization on hand-written WebGPU + WASM-SIMD kernels (no onnxruntime). Model weights stream from Hugging Face on first use and cache locally.
import { ParakeetV3Engine } from "@fluidinference/fluidaudio-web/asr-parakeet";
import { decodeToMono16k } from "@fluidinference/fluidaudio-web";
const asr = new ParakeetV3Engine();
await asr.load((p) => console.log(p.file, p.fraction));
asr.setVocabulary(["NVIDIA", "Newrez"]); // optional fuzzy correction
asr.setItn(true); // optional "twenty one" → "21"
const audio = await decodeToMono16k(fileArrayBuffer);
const { text } = await asr.transcribe(audio);
await asr.dispose();Engines (one subpath each, tree-shakeable): /asr-parakeet, /asr-whisper, /asr-nemotron, /tts-kokoro (new KokoroTtsEngine({ lang: "en" | "zh" })), /vad-silero, /diarization-sortformer, /eou-parakeet. To enumerate dynamically, use /registry and instantiate via each entry's make() — registry ids are NOT all valid subpaths (the two Kokoro ids share one subpath).
Requirements: a bundler that supports new URL(..., import.meta.url) assets, module workers, and JSON imports (Vite and webpack 5 out of the box; Rollup needs @rollup/plugin-json + an import-meta-assets plugin). WebGPU strongly recommended (WASM-SIMD fallback runs everywhere). Weights download from Hugging Face at runtime — no build-time model assets.
Demo/playground (same code): https://fluidaudio-web.hanweng9.workers.dev
