npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

react-ai-voice-avatar

v0.7.1

Published

Open-source alternative to real-time avatar APIs like HeyGen Interactive Avatar and Tavus. A talking, lip-synced 3D avatar that runs on your user's GPU, driven by your own LLM. No video stream, no per-minute billing. Or voice mode alone, ChatGPT-style, as

Readme

React AI Voice Avatar (react-ai-voice-avatar) 🚀🗣️🧬

NPM Version TypeScript Tested with Playwright Live Demo License: MIT

The open-source alternative to real-time avatar APIs, as a React component.

HeyGen Interactive Avatar, Tavus and Soul Machines sell a talking avatar that holds a conversation, rendered on their servers and billed per streaming minute. This does the same job as an npm install: the avatar renders on your user's GPU, so there is no video stream, no per-minute cost, and no third party in the middle of your conversations.

You bring the model. Point it at OpenAI, Anthropic, your own fine-tune or your existing chat endpoint, and keep your keys on your own backend. The package owns the parts that are tedious to build and easy to get wrong: microphone capture, knowing when someone has finished speaking, interrupting the avatar mid-sentence when they talk over it, streaming speech synthesis, and driving 52 ARKit facial blendshapes at 60 FPS so the mouth matches the words.

It can also run with no backend at all. Speech recognition, generation and voice all have in-browser implementations, which makes for a convincing demo and a genuinely offline kiosk. Most production apps will use their own model and keep only speech and lip-sync on the device.

Don't need a face? The same engine is a headless React hook for voice mode, the way ChatGPT and Gemini do it: react-ai-voice-avatar/headless, with no three.js in your bundle. See it ➔ · How ➔

🌐 Try the live demo ➔ · Voice only ➔

Speech recognition, voice synthesis and lip-sync, all running in your browser tab.

The first visit downloads about 590 MB of models, roughly a minute and a half on a 50 Mbit/s connection. Your browser keeps them, so a return visit is talking in about 2 seconds with nothing downloaded. Replies on the demo come from a hosted model on Groq's free tier, shared by everyone trying it, and start about 3 seconds after you stop talking; on a busy day that tier can run out, and examples/groq-voice runs the same thing on your own free key.

Talking to the avatar: it answers with captions and gestures, is interrupted mid-answer, stops and answers the new question

🧪 Status: beta

The engine works and is tested on every push, but it has been checked on far fewer browsers and devices than it is meant for. Here is exactly where things stand, so you can decide what to verify yourself before shipping.

Tested

  • Desktop Chrome on macOS (Apple M4, WebGPU), by hand: the live demo (local speech, hosted replies) and the fully local setup.
  • Headless Chromium on Linux, in CI on every push: whole turns through your own speech-to-text, model and voice, with stand-ins for the providers; real speech through a fake microphone for push-to-talk; empty replies; sentences kept in order; the in-browser model speaking in your voice; and that no model downloads when you supply your own.

Not tested yet

  • Safari on Mac, iPhone and iPad, and Firefox.
  • Real Android phones and iPhones: speed, memory, and the microphone.
  • Cloud speech against real providers, beyond the stand-ins in CI.
  • Long sessions. Nothing has yet held a conversation for 30 minutes and checked memory and turn-taking along the way.

The iPhone behaviour (the smaller MMS voice, and the guard that stops a tab reloading into the same crash) is built around Safari's known memory limits and has not been run on a device. The compatibility and hardware tables mark what is tested and what is expected.

Versions. It is 0.x: a minor version can change behaviour. Since 0.6.0 each CHANGELOG entry says at the top whether it breaks anything. Pin an exact version in production. If it fails on your device, an issue with the browser, the device and the onError output is the most useful thing you can send.

🧩 Built with it

Interview Room, a spoken mock interviewer I built on this package for the DEV Hacktoberfest challenge. Everything runs in the browser: Whisper, Gemma, Kokoro and the avatar. It uses the headless hook for its phone screen, the avatar for its video interview, onSubmit with its own Gemma worker, and speechMs for speaking pace. Live demo ➔

Interview Room: the home page, choosing a video interview with the avatar or a phone screen


🌟 Two Entry Points (How to use it)

react-ai-voice-avatar provides two distinct ways to integrate into your app depending on your design needs. Both share the exact same underlying conversational state machine, Voice Activity Detection (VAD), and turn-taking logic.

🎧 Entry 1: The Headless Hook ("Voice Mode for your App")

Voice mode for your app: talk to it the way you talk to ChatGPT or Gemini, and interrupt it mid-sentence. The hook owns the microphone, knowing when someone has finished speaking, interruption, and streaming the reply into speech. You own the UI, and every frame it hands you a loudness level to animate.

Voice mode with no avatar: an orb that pulses with the voice, live captions, and a reply interrupted mid-answer by a new question

Try it ➔: a page built on this hook alone. examples/voice-only is the same page as an app to copy.

Import it from react-ai-voice-avatar/headless and Three.js never enters your module graph. Measured on the same Next.js App Router build, one route rendering the 3D avatar and one rendering only the hook:

| Route | First Load JS | | --- | --- | | 3D avatar | 389 kB | | Headless hook | 117 kB |

npm install react-ai-voice-avatar
import { useRef, useState } from 'react';
// The /headless subpath is what keeps Three.js out of your bundle.
// Importing the hook from the package root pulls the 3D stack in with it.
import { useAiVoiceAvatar } from 'react-ai-voice-avatar/headless';

export function VoiceMode() {
  const orb = useRef<HTMLDivElement>(null);
  const [micOn, setMicOn] = useState(false);

  const voice = useAiVoiceAvatar({
    // Your model. Return a string, or a stream so speech starts sooner.
    onSubmit: text => fetch('/api/chat', { method: 'POST', body: text }).then(r => r.body),

    // 0 to 1 every frame, from whoever is talking. Written straight to the
    // DOM: routing it through state would re-render sixty times a second.
    onAudioLevelChange: level => {
      if (orb.current) orb.current.style.transform = `scale(${1 + level * 0.3})`;
    },
  });

  // The mic stays open across turns, so it is tracked apart from `status`,
  // which moves through listening, thinking and speaking.
  const toggle = () => {
    if (micOn) {
      voice.stopListening();
      voice.interrupt();
    } else {
      voice.startListening(); // From a click: this is where the mic prompt appears.
    }
    setMicOn(!micOn);
  };

  return (
    <>
      <div ref={orb} className="orb" data-status={voice.status} />
      <button onClick={toggle} disabled={!voice.isReady}>
        {voice.isReady ? (micOn ? 'Stop' : 'Talk') : 'Loading…'}
      </button>
    </>
  );
}

For captions, onTranscriptUpdate(text, 'user') gives what the user said and onSpeechStart(text) each sentence of the reply as it starts playing. speak(text) says something without a model turn, such as a greeting, and sendText(text) takes typed input. Everything it returns is in the hook reference.

🚀 Voice mode in production

Each stage runs in the browser or on your backend, and the choice is mostly about the download:

| Setup | Options | Downloaded by each visitor | Audio leaves the device | | :--- | :--- | :--- | :--- | | Cloud speech | onTranscribe + onSubmit + onSynthesize | Nothing until the mic opens, then ~4 MB for the voice detector | Yes, to your providers | | Local speech (the demo) | onSubmit | ~590 MB once, kept for later visits (~320 MB on iPhone) | No | | Fully local | none | ~1.3 GB once | Nothing leaves at all |

Cloud speech is the setup to reach for on phones and on pages people visit once: it is ready as soon as the page is. Local speech keeps what people say on their device and costs nothing per minute, and suits an app they come back to, since the models are stored after the first visit. Fully local is for kiosks and offline use. Use loadModels to hold any download until someone actually engages.

☁️ Cloud Adapters

onTranscribe and onSynthesize replace the local speech models with your own providers. Each replaces its model entirely: supply one and that model is never downloaded. Drop it later and the local model loads, so a provider outage can fall back to the browser mid-session.

To have that fallback ready rather than downloading it at the moment you need it, pass preloadLocalSpeech: the local hearing and voice download behind the conversation, and onLocalSpeechReady fires when they could take over. examples/groq-voice starts on Groq and offers the switch then.

const voice = useAiVoiceAvatar({
  // Instead of Whisper: 16 kHz mono samples of one utterance, to any
  // speech-to-text service. Encode them as WAV if the service wants a file.
  onTranscribe: async (samples: Float32Array) => {
    return await mySpeechToText(samples);
  },

  // Instead of Kokoro: return an encoded MP3/WAV (ArrayBuffer), which the
  // hook decodes, or raw 24 kHz PCM as a Float32Array. Called a sentence at
  // a time, so the first plays while the rest are fetched.
  onSynthesize: async (text: string) => {
    const res = await fetch('/api/tts', { method: 'POST', body: text });
    return await res.arrayBuffer();
  },

  onSubmit: async (text) => fetch('/api/chat', { method: 'POST', body: text }).then(r => r.body),
});

🧑‍💼 Entry 2: The Full 3D Avatar

If you want the full visual presence with 60FPS ARKit lip-syncing, use the drop-in 3D component. Under the hood, this is just a wrapper around the useAiVoiceAvatar hook that procedurally maps the audio to a 3D model!

📦 View Package on the Official NPM Registry ➔

# If using the 3D Avatar, you must also install the Three.js ecosystem
npm install react-ai-voice-avatar three @react-three/fiber @react-three/drei

[!NOTE] React 18 Users: Installing the latest @react-three/drei defaults to version 10, which demands React 19. If your project runs on React 18, install compatible Three.js React bindings explicitly:

npm install @react-three/drei@^9 @react-three/fiber@^8

[!IMPORTANT] React 19.3 and ERESOLVE: @react-three/[email protected] still declares its React peer as >=19 <19.3, so a default npm install alongside React 19.3 or newer fails with ERESOLVE unable to resolve dependency tree. This is a Three.js binding constraint, not a limit of this package: our own peer range accepts React 19.3.

That range is over-cautious. We build and run a Next.js App Router app against React 19.3.0 with @react-three/[email protected] and the avatar renders, loads its models and speaks with no errors. So install past it rather than downgrading:

npm install react-ai-voice-avatar three @react-three/fiber @react-three/drei --legacy-peer-deps

If you would rather keep strict peer resolution, pinning React works too:

npm install react@~19.2.0 react-dom@~19.2.0

The headless entry point (react-ai-voice-avatar/headless) pulls in no Three.js at all, so it never hits this and works on any React 18 or 19 version.

⚡ DX & Performance (Lazy Code-Splitting)

To prevent the massive ML assets (WebGPU workers, 3D engines) from bloating your initial page load, use the built-in lazy wrapper. It will automatically code-split the 3D dependencies and render a sleek holographic Skeleton UI while the assets download in the background!

import { AiVoiceAvatarLazy } from 'react-ai-voice-avatar';

// Use it exactly like the normal component!
<AiVoiceAvatarLazy avatarPreset="ananya" />

Show the avatar before anyone commits to a download

Lazy loading defers the 3D engine. The speech and language models are the larger cost, and by default they start downloading as soon as the avatar mounts. On a landing page most visitors only look, so render the avatar idle and load the models when someone actually engages:

const [engaged, setEngaged] = useState(false);

<AiVoiceAvatar loadModels={engaged} hideStatusPill={!engaged} />
<button onClick={() => setEngaged(true)}>Talk to it</button>

The live demo's landing page works this way.

📐 Architectural Best Practices

  • Standard Import (AiVoiceAvatar): Recommended for full-screen applications where the avatar is the primary product (e.g., Kiosks, Digital Tutors). The browser starts downloading the 3D canvas and ML models straight away, so the avatar is ready as early as it can be.
  • Lazy Import (AiVoiceAvatarLazy): Recommended for widgets, modals, or sub-routes (e.g., a "Support Desk" chat bubble in a SaaS dashboard). Defers downloading the 1.5MB 3D engine and WebWorkers until the user actually opens the widget.

⚙️ Server Configuration (Optional Performance Boost)

The react-ai-voice-avatar engine is truly zero-config. You do not need to configure Vite optimizeDeps, Next.js Webpack overrides, or manually host Web Worker files—everything is dynamically bundled and executed automatically!

However, because our ONNX WebGPU engine leverages modern multi-threaded SharedArrayBuffer memory pipelines for maximum inference speed, your hosting server can optionally emit standard Cross-Origin Isolation HTTP headers (COOP/COEP) to unlock peak performance. If these headers are not present, the engine automatically falls back to single-threaded WebAssembly without crashing.

🌐 Enabling Multi-threading on Production (Vercel, Netlify & Cloudflare)

To unlock multi-threaded performance, specify these isolation headers in your routing manifests:

  • Vercel (vercel.json): Add "headers": [{ "source": "/(.*)", "headers": [{ "key": "Cross-Origin-Opener-Policy", "value": "same-origin" }, { "key": "Cross-Origin-Embedder-Policy", "value": "require-corp" }] }].
  • Netlify / Cloudflare Pages (_headers or netlify.toml): Add /*\n Cross-Origin-Opener-Policy: same-origin\n Cross-Origin-Embedder-Policy: require-corp to public/_headers.

⚡ Enabling Multi-threading in Local Dev (Vite & Next.js)

Vite (vite.config.ts):

export default defineConfig({
  plugins: [react()],
  server: {
    headers: {
      'Cross-Origin-Opener-Policy': 'same-origin',
      'Cross-Origin-Embedder-Policy': 'require-corp',
    },
  },
});

Next.js (next.config.mjs):

export default {
  async headers() {
    return [
      {
        source: '/(.*)',
        headers: [
          { key: 'Cross-Origin-Opener-Policy', value: 'same-origin' },
          { key: 'Cross-Origin-Embedder-Policy', value: 'require-corp' },
        ],
      },
    ];
  },
};

[!CAUTION] Strict CSP Policies: If your enterprise enforces strict Content Security Policies that block blob: workers (worker-src 'self'), you can bypass our zero-config Blob loaders by passing the workerBaseUrl prop to the avatar and hosting the pre-compiled .worker.js files from our dist/assets/ directory yourself.

⚡ 3D Avatar Quickstart

import React, { useRef, useState } from 'react';
import { Canvas } from '@react-three/fiber';
import { OrbitControls } from '@react-three/drei';
import { AiVoiceAvatar, type AiVoiceAvatarHandle } from 'react-ai-voice-avatar';

export function App() {
  const avatarRef = useRef<AiVoiceAvatarHandle>(null);
  const [text, setText] = useState('');

  return (
    <div style={{ width: '100vw', height: '100vh', position: 'relative' }}>
      <Canvas camera={{ position: [0, 0.15, 2.2], fov: 32 }}>
        <color attach="background" args={['#101116']} />
        
        {/* Subtle studio lighting */}
        <pointLight position={[-3, 2, -2]} intensity={25} color="#E67E22" distance={6} />
        <pointLight position={[3, 1, -2]} intensity={20} color="#2980B9" distance={6} />
        
        <OrbitControls target={[0, 0.05, 0]} />
        
        {/* Connected Brain: your model answers, so no language model is downloaded */}
        <AiVoiceAvatar
          ref={avatarRef}
          avatarPreset="ananya"
          lightingPreset="studio"
          ttsEngine="kokoro"
          ttsVoice="af_heart"
          // Connect your backend here (receives user speech transcript):
          onSubmit={async (text) => {
            const res = await fetch('/api/chat', { 
              method: 'POST', 
              body: JSON.stringify({ prompt: text }) 
            });
            return res.body; // Avatar natively reads streams!
          }}
        />
      </Canvas>

      {/* Fallback Text Input for noisy environments */}
      <form
        onSubmit={(e) => {
          e.preventDefault();
          if (text.trim() && avatarRef.current) {
            avatarRef.current.sendText(text);
            setText('');
          }
        }}
        style={{ position: 'absolute', bottom: '20px', left: '50%', transform: 'translateX(-50%)', display: 'flex', gap: '8px', zIndex: 100 }}
      >
        <input 
          value={text} 
          onChange={e => setText(e.target.value)} 
          placeholder="Type a message..." 
          style={{ padding: '8px 16px', borderRadius: '20px', border: 'none', background: 'rgba(255,255,255,0.9)', width: '300px' }}
        />
        <button type="submit" style={{ padding: '8px 16px', borderRadius: '20px', border: 'none', background: '#3b82f6', color: 'white', cursor: 'pointer' }}>
          Send
        </button>
      </form>
    </div>
  );
}

🔌 Connect Your Backend (Recipes)

The onSubmit prop natively accepts a string, an AsyncIterable<string>, or a ReadableStream. To connect your actual backend, simply drop in one of these copy-paste recipes to parse your streaming format!

[!CAUTION] API Keys Belong on the Server! Never put your OpenAI or Anthropic API keys directly in the frontend browser code. Always route through your own backend endpoint (/api/chat).

Recipe 0: Plain Text Stream (Fastest & Simplest)

If your backend uses Vercel AI SDK's streamText(...).toTextStreamResponse() or otherwise streams plain raw text, you can pass the stream natively without any parsing!

onSubmit={async (text) => {
  const res = await fetch('/api/chat', { method: 'POST', body: JSON.stringify({ prompt: text }) });
  return res.body; // Natively supported!
}}

Recipe 1: Vercel AI SDK (≤v4 Data Stream)

Older versions of the Vercel AI SDK stream data using a specific protocol (e.g., 0:"Hello"). This recipe parses those chunks into clean text with a carry-over buffer for safe network boundaries.

onSubmit={async function* (text) {
  const res = await fetch('/api/chat', { method: 'POST', body: JSON.stringify({ prompt: text }) });
  if (!res.body) return;
  const reader = res.body.getReader();
  const decoder = new TextDecoder();
  let buffer = '';
  
  while (true) {
    const { done, value } = await reader.read();
    if (done) break;
    buffer += decoder.decode(value, { stream: true });
    const lines = buffer.split('\n');
    buffer = lines.pop() ?? ''; // keep the trailing partial chunk
    for (const line of lines) {
      if (line.startsWith('0:')) {
        try { yield JSON.parse(line.substring(2)); } catch { /* ignore keep-alive / non-JSON frames */ }
      }
    }
  }
}}

Recipe 2: OpenAI-Compatible SSE Endpoint (and AI SDK v5)

Standard Server-Sent Events (SSE) stream data: {...} blocks. This handles safe parsing across broken network chunk boundaries.

onSubmit={async function* (text) {
  const res = await fetch('/api/chat', { method: 'POST', body: JSON.stringify({ prompt: text }) });
  if (!res.body) return;
  const reader = res.body.getReader();
  const decoder = new TextDecoder();
  let buffer = '';
  
  while (true) {
    const { done, value } = await reader.read();
    if (done) break;
    buffer += decoder.decode(value, { stream: true });
    const lines = buffer.split('\n');
    buffer = lines.pop() ?? ''; // keep the trailing partial chunk
    for (const line of lines) {
      if (line.startsWith('data: ') && line !== 'data: [DONE]') {
        try {
          const parsed = JSON.parse(line.substring(6));
          // AI SDK v5 emits {type:'text-delta', delta:'...'}; OpenAI emits choices[0].delta.content
          if (parsed.type === 'text-delta' && parsed.delta) {
            yield parsed.delta;
          } else if (parsed.choices?.[0]?.delta?.content) {
            yield parsed.choices[0].delta.content;
          }
        } catch { /* ignore keep-alive / non-JSON frames */ }
      }
    }
  }
}}

Staying quiet, and how long they spoke

Return an empty string (or a stream that ends without text) and the avatar says nothing: the turn ends and it keeps listening. That lets you gather one answer across several pauses, replying only once the user has finished.

The second argument says how long the user spoke, from their first word to their last. The padding kept from before speech and the silence waited out after it are left out, so it is the figure for a speaking rate. It is undefined for text sent with sendText.

onSubmit={async (text, { speechMs }) => {
  answer += ' ' + text;
  if (!answerIsComplete(answer)) return ''; // say nothing, keep listening
  const wpm = speechMs ? text.split(/\s+/).length / (speechMs / 60000) : undefined;
  return askMyBackend(answer, { wpm });
}}

🗣️ Languages

English and Hindi both run entirely on the device. Set ttsLanguage and the engine picks a matching voice, a matching speech-recognition model, and a matching phoneme path.

<AiVoiceAvatar ttsLanguage="hi-IN" />

Hindi needed real work rather than a config flag. Kokoro ships four Hindi voices inside the same checkpoint as the English ones, but the JavaScript wrapper does not list them, and the bundled eSpeak build carries English data only and rejects hi outright. So this package includes its own Devanagari-to-phoneme converter. Devanagari is close to phonemic, which makes that tractable; the hard part is schwa deletion, the rule that makes कमल read as "kamal" rather than "kamala", and getting it wrong produces speech that still sounds like speech while sounding like someone spelling Hindi out.

Because the converter emits real phonemes, Hindi gets phoneme-driven lip-sync with distinct mouth shapes for the retroflex consonants, not the amplitude-only mouth flapping most engines fall back to outside English. Code-switching works too: English words inside a Hindi sentence are routed to the English phonemiser, so "मुझे coffee चाहिए" is pronounced correctly throughout.

[!NOTE] Local Hindi is a demo, not a product. Models small enough to run in a browser are far weaker in Hindi than in English. It is good enough to show the pipeline working end to end and not good enough to ship. For production Hindi, route hearing and thinking to an API through onTranscribe and onSubmit; the voice and lip-sync stay local and are genuinely good.

A session is one language at a time. Someone who switches language mid-conversation will be mistranscribed, because the recognition model is told which language to expect and browser Whisper cannot detect it.


🎤 Taking turns

The user can talk over the avatar and cut it off mid-sentence. This is on by default, because waiting for a reply to finish is the thing that makes a voice agent feel like a walkie-talkie.

<AiVoiceAvatar
  allowInterruption={false}   // default: true
  onUserInterrupt={() => analytics.track('barge_in')}
/>

Turn it off for a kiosk or a noisy room, where the avatar hearing its own voice through the speakers and stopping itself is worse than waiting. It is ignored in push-to-talk, which owns the floor explicitly.

Push-to-talk

With listenMode="push-to-talk" the button decides when a turn ends, not a silence. Call startListening() when it goes down and stopListening() when it comes up: what was said in between is handed over as one turn, however long the pauses in it. Pressing during a reply stops the reply. interrupt() abandons a hold without sending it, and a hold with no speech in it sends nothing.

<button
  onPointerDown={() => voice.startListening()}
  onPointerUp={() => voice.stopListening()}
>
  Hold to talk
</button>

How a turn is decided

Voice detection scores every 96ms frame for how much it sounds like speech. A cough, a door and a chair all clear a loudness bar as easily as a word does, so loudness alone cannot separate them — sustain can. A sound has to keep scoring as speech for minSpeechMs before the engine treats it as a turn.

Below that bar, a sound only ducks the avatar's voice, reversibly. If it turns out to be a cough the reply resumes from where it paused, at the right place in the sentence and with the mouth still in sync. Nothing about the conversation changed, because nothing was decided on a noise.

Tuning for your room

The defaults suit a quiet room and headphones. A shop floor is a different problem, and only you know which you have.

<AiVoiceAvatar
  speechDetection={{ positiveSpeechThreshold: 0.6, minSpeechMs: 700 }}
/>

| Field | Default | Raise it when | Lower it when | | :--- | :--- | :--- | :--- | | positiveSpeechThreshold | 0.5 | Passing noise is mistaken for talking | Quiet speakers go unheard | | negativeSpeechThreshold | 0.35 | Turns end too slowly in a noisy room | Turns end while someone is still talking | | minSpeechMs | 500 | Short noises still start turns | Single-word answers are ignored | | redemptionMs | 1400 | People are cut off while thinking mid-sentence | Replies feel slow to start | | preSpeechPadMs | 800 | The first word is still being clipped | Rarely — this is the audio kept from before the trigger, and it is what stops the first syllable going missing |

Two costs worth knowing before you change anything. A speaker who sits below positiveSpeechThreshold produces no events at all — raising it trades quiet voices for quiet rooms. And a filler like "hmm" held long enough to pass minSpeechMs still reaches transcription; a denylist catches the common ones after the fact, but sustain cannot tell a long "hmm" from a short word.


🚨 Handling failures

Pass onError and your app learns when something breaks, rather than finding out from a console message it cannot see.

<AiVoiceAvatar
  onError={(e) => {
    if (e.severity === 'fatal') showFallbackUI(e.stage);
    logToSentry(e);
  }}
/>

Check severity before reacting. Most failures here are survivable because the engine falls back: WebGPU to WASM, Kokoro to a smaller voice model. Those arrive as degraded and the avatar still works, so treating them as fatal would hide a working experience behind an error screen. A refused microphone is fatal for listening while typed input still works, which is a judgement only your app can make.

stage is one of microphone, speech-recognition, language-model, speech-synthesis, audio-output, worker, model-storage or conversation. There is also a detail string carrying the engine's internal stage name for bug reports; it is not stable across versions, so do not branch on it.

model-storage is always degraded: a model loaded and works, but could not be kept, so the next visit downloads it again. The message says why — usually not enough free space, which on a phone is the common case.

Models are kept between visits

Downloaded models are stored in the browser's Origin Private File System, so a returning visitor skips the download: on an Apple M4 with WebGPU, a return visit is ready to talk in about 2 seconds, with nothing fetched. The browser's Cache API — what the model loader uses by default — refuses any single file of 256 MiB or more, and both the voice (~310 MB) and the local language model (~750 MB) are larger than that, so before this they were downloaded again on every visit. Where OPFS is unavailable the Cache API is used as before.

Browsers can clear this storage when disk space runs low. If your app is one people return to, call navigator.storage.persist() from a user gesture to ask the browser to keep it; the engine does not, because Firefox can answer that call with a permission prompt, and a library should not put one in front of your visitors unasked.

Keeping your own models

A model of your own that is over 256 MiB has the same problem. The store the engine uses is its own entry point, with no React and nothing else in it, so a worker that loads your model can import it:

import { env } from '@huggingface/transformers';
import { createModelCache, isModelCacheSupported } from 'react-ai-voice-avatar/model-cache';

if (isModelCacheSupported()) {
  env.useCustomCache = true;
  env.customCache = createModelCache();
}

It is an ordinary match / put / delete cache, so anything that loads through one can use it. onStoreFailed(url, reason) reports a file that could not be kept, such as when the disk is full; the model still loads, and is downloaded again next visit.


🎨 Bring Your Own 3D Avatar (Custom GLB)

You are not locked into our built-in avatars (ananya and aarav)! You can use any custom .glb humanoid model by passing its URL or local path to the modelSrc prop:

<AiVoiceAvatar
  modelSrc="/models/my-custom-avatar.glb"
  // ...
/>

📋 Custom Avatar Requirements

To ensure the lip-sync and procedural facial dynamics engines work correctly, your custom model must meet the following standard requirements:

  1. Format: .glb (GLTF Binary).
  2. Facial Blendshapes (Morph Targets): The model's head/face mesh must contain the standard 52 Apple ARKit blendshapes (e.g., jawOpen, eyeBlinkLeft, mouthSmileRight). Our engine automatically traverses your model to find these targets.
  3. Bone Naming: For the interactive mouse-tracking and head-tilting physics to function, the armature should use standard bone names (e.g., a neck/head bone named Head, head, Neck, or neck).

Where the built-in avatars come from

ananya and aarav are converted from the Microsoft Rocketbox library, which is MIT licensed like this project, so you can redistribute them without restriction. The library has 115 avatars and any of them can be converted with scripts/convert-rocketbox.py. See assets/avatars/LICENSE.md.

Models exported from Ready Player Me also work, since they carry the same ARKit blendshapes and bone names. Note that Ready Player Me shut down in January 2026, so you can no longer create new avatars there, and existing exports are licensed CC BY-NC-SA rather than MIT.


🏗️ What runs where

Four stages, and you choose where each one happens. The defaults are all local, which is why the demo needs no keys, but the interesting production setups are mixed.

flowchart TD
  mic(["🎙️ You speak"]) --> vad["Voice detector<br/><i>in the browser</i>"]
  vad --> hear["Hearing<br/><i>Whisper in the browser</i><br/>or your onTranscribe"]
  hear --> think["Thinking<br/><i>a model in the browser</i><br/>or your onSubmit"]
  think --> speak["Speaking<br/><i>Kokoro in the browser</i><br/>or your onSynthesize"]
  speak --> face["🔊 Voice, with a lip-synced<br/>3D face or your own UI"]
  vad -. "talk over it and it stops" .-> speak

Each stage runs in the browser unless you pass the adapter beside it, which hands that stage to your own service.

| Stage | On the device | Your backend instead | | :--- | :--- | :--- | | Hearing — speech to text | Whisper via ONNX | onTranscribe | | Thinking — the reply | Qwen or Gemma via WebGPU | onSubmit | | Speaking — text to audio | Kokoro-82M | onSynthesize | | Face — lip-sync and animation | Always here | Not applicable |

The face never leaves the device, which is the whole point: that is what avatar APIs charge per minute for, and it is the one stage that cannot be outsourced without a video stream.

The common production shape is onSubmit alone. Hearing and speaking stay local, so no audio ever leaves the browser, while generation goes to whatever model you already run. That keeps the download to about 600 MB (Whisper base and the Kokoro voice), keeps your keys on your server, and still gives you a conversation nobody else can read.

Add onTranscribe and onSynthesize as well and nothing downloads at all, apart from the ~4 MB voice detector once the microphone opens. Each adapter replaces its model entirely, so the page is ready as soon as it loads. That is the shape for phones and for pages people visit once.

Fully local is real, not a demo trick, and it is the right answer for a kiosk, a regulated environment, or anywhere without reliable connectivity. Be aware of the cost: in English the first visit downloads roughly 1.3 GB before anyone can speak (the language model alone is 750 MB), about 2.1 GB in Hindi, and the quality ceiling is whatever a model that size can do.

⏱️ How fast it answers

Measured with scripts/measure-latency.mjs on an Apple M4 with WebGPU in Chromium: a spoken question through a fake microphone, the median of five turns after a warm-up turn, timed by the hook's own callbacks.

| Stage | Time | | :--- | :--- | | Waiting for silence, to be sure you've finished | 1,400 ms (redemptionMs, adjustable) | | Hearing: Whisper base transcribes the utterance | 330 ms | | Thinking, hosted: GPT-OSS 20B on Groq, the demo's route, one round trip | 670 ms | | Thinking and the first sentence of voice, all in the browser (Qwen 0.5B, Kokoro) | 740 ms | | Speaking: Kokoro's first sentence ready and playing, with an instant reply | 460 ms | | Starting up on a return visit, models already stored | about 2 s, nothing downloaded |

So from the moment you stop talking to the first word of the reply:

| Setup | First word after you stop | | :--- | :--- | | All in the browser | about 2.5 s | | Hosted replies, local hearing and voice (the demo) | about 2.9 s | | With an instant reply, the floor for speech alone | about 2.2 s |

The hosted figure adds the route's round trip to the measured speech stages, since that route answers only the live site. Most of the wait is the pause for silence, which is what stops the avatar cutting people off while they think. Lower redemptionMs in speechDetection for a snappier feel, at the cost of answering half-finished sentences. Without WebGPU, recognition and the voice run on the CPU and are several times slower.

🔀 Hosted now, local when warm

You do not have to choose once. preloadLocalLlm fetches the in-browser model behind the conversation while your backend answers, and onLocalLlmReady tells you when it can take over:

const [localReady, setLocalReady] = useState(false);

<AiVoiceAvatar
  // Dropping onSubmit is what hands the conversation over.
  onSubmit={localReady ? undefined : askMyBackend}
  preloadLocalLlm
  onLocalLlmReady={() => setLocalReady(true)}
/>

Nobody waits for a gigabyte to say the first word, and the turns your backend answered are handed to the local model, so it does not restart the conversation from nothing. Useful for a kiosk that must keep working when the wifi drops, a demo on someone else's quota, or a phone, which will never accept the download but can hold a conversation the moment it loads.

The live demo runs exactly this: replies come from a hosted model, and on a desktop with WebGPU the in-browser model downloads during the conversation and takes the floor when it lands. The page says which one is answering at any moment.

Explore the canonical patterns in the examples/ directory:

| Example Pattern | Folder | Highlights & Architecture | | :--- | :--- | :--- | | Live Interactive Demo | sandbox | Deploy on Vercel ➔ — Our full-featured interactive testbed featuring live character switching (ananya, aarav), voice persona switching (af_heart, am_michael), real-time diagnostic probe metrics, and Leva 3D lighting controls. | | Quickstart | examples/quickstart | Minimal, zero-configuration plug-and-play AI voice avatar deployment with built-in studio lighting & sizing. | | Local Kiosk | examples/local-kiosk | Retail and restaurant ordering kiosk with the menu reasoning on the device. Demonstrates the On-Device Brain; works without internet once the models are stored. | | Connected App | examples/hybrid-cloud | Illustrates the Connected Brain (onSubmit). Bypasses gigabyte-scale local LLM downloads by routing reasoning to OpenAI, Claude, or corporate APIs while speech recognition, the voice and lip-sync stay on the device. | | Voice Only | examples/voice-only | Voice mode with no avatar, like ChatGPT or Gemini voice: the react-ai-voice-avatar/headless hook, an audio-reactive orb, live captions and interruption. No three.js in the bundle, and no backend needed to try it. | | Groq Voice | examples/groq-voice | Voice mode on Groq's free tier with your own key: hearing, replies and voice start on Groq with nothing to download, the in-browser hearing and voice download behind the conversation (preloadLocalSpeech), and the page offers to switch when they are ready. Shows Groq's live limits. | | Headless Custom UI| examples/headless-custom-ui| Still renders the 3D avatar; for no avatar at all, see Voice Only above. Demonstrates hiding built-in DOM overlays (hideStatusPill={true}, showCaptions={false}), streaming transcripts into a custom enterprise UI, and controlling voice outputs imperatively via ref.current?.speak(text). |


📖 API Reference

<AiVoiceAvatar /> Props

| Prop | Type | Default | Description | | :--- | :--- | :--- | :--- | | avatarPreset | 'ananya' \| 'aarav' \| 'default' \| 'kiosk' | 'ananya' | Built-in 3D character models featuring both female ('ananya') and male ('aarav') voice concierges out of the box with full ARKit facial blendshapes! | | avatarSize | 'sm' \| 'md' \| 'lg' \| number | 'md' (0.48) | Intuitive model sizing presets or custom decimal scaling multiplier applied directly to the 3D humanoid mesh. | | modelSrc | string | undefined | Absolute local path or remote URL to a custom GLTF/GLB humanoid armature avatar model. | | lightingPreset | 'studio' \| 'cyberpunk_violet' \| 'cool_azure' \| 'warm_amber' \| 'clean_white' \| 'none' | 'studio' | Pre-built cinematic studio lighting atmospheres directly applied to your 3D viewport without manual Three.js configuration! | | systemPrompt | string | "You are Ananya..."| Conversational persona directives and context injected into active LLMs. | | llmModel | string | by language | Hugging Face id for the local WebGPU reasoning model, used only when onSubmit is absent. Defaults to Qwen2.5-0.5B for English and Gemma 3 1B for Hindi, which Qwen that size cannot speak. Set it to pin one model for every language. | | asrModel | string | by language | Hugging Face id for the local Whisper model. Defaults to Whisper base for English and Whisper small for Hindi, which base transcribes badly. Pass "Xenova/whisper-tiny" for a faster download and worse accuracy. | | ttsEngine | 'kokoro' \| 'mms' | 'kokoro' | High-fidelity neural voice synthesis engine executing inside dedicated Web Workers. | | ttsVoice | string | by language | Kokoro voice id. af_* and am_* American, bf_* and bm_* British, hf_* and hm_* Hindi (af_heart, am_michael, bf_emma, hf_alpha, hm_omega). Defaults to one matching ttsLanguage. A voice whose language disagrees is corrected with a warning. | | ttsLanguage| 'en-US' \| 'en-GB' \| 'hi-IN' | 'en-US' | Conversation language. Selects the voice, the recognition model and the phoneme path. Also sets asrLanguage unless you set that yourself. For any other language, pass onSynthesize and use a cloud voice provider. | | asrLanguage| string | ttsLanguage | Language to transcribe. Follows ttsLanguage by default, since a conversation is almost always held in one language. | | showCaptions | boolean | true | Renders a sleek glassmorphic subtitle overlay displaying spoken interaction dialog. | | hideStatusPill| boolean | false | When true, suppresses the default bottom-left microphone interactive control pill. | | listenMode | 'continuous' \| 'push-to-talk' | 'continuous' | continuous keeps the mic hot after the avatar finishes speaking naturally, but explicitly clicking Stop forces it off until tapped again. push-to-talk listens only while held: startListening() on press, stopListening() on release sends the turn. See Push-to-talk. | | loadModels | boolean | true | Set false to render the avatar without downloading any models, then flip it true when the visitor engages. For landing pages and widgets most visitors never talk to. Status stays 'loading' until it is true and the models are up, so show your own call to action meanwhile. | | onModelLoaded | () => void | undefined | Fires once the 3D mesh is parsed and in the scene. Parsing a multi-megabyte GLB leaves the canvas empty for a few seconds; use this to hold a placeholder over it. | | allowInterruption | boolean | true | Lets the user talk over the avatar and cut it off mid-sentence. Turn off for a kiosk or noisy room, where the avatar hearing itself through the speakers is worse than waiting. Ignored in push-to-talk. See Taking turns. | | speechDetection | { positiveSpeechThreshold?, negativeSpeechThreshold?, minSpeechMs?, redemptionMs?, preSpeechPadMs? } | see Taking turns | Tunes how the microphone decides someone is talking. Every field optional. The defaults suit a quiet room; a shop floor needs a higher threshold and a longer minSpeechMs. | | onUserInterrupt | () => void | undefined | Fires when the user talks over the avatar and takes the floor. Only fires if the avatar actually had audio playing. | | gestures | boolean \| number | true | Hand and arm gestures while the avatar speaks: one hand or both brought up in front of the chest for each phrase, with small beats on stressed syllables, and the arms back at rest when it stops. false keeps the arms still; a number sets the size, 0 to 1.5. Needs a skeleton with LeftArm, LeftForeArm and LeftHand and the right-hand equivalents; custom avatars without them simply do not gesture. | | onAudioLevelChange | (level: number, source: 'mic' \| 'tts' \| 'idle') => void | undefined | Real-time audio amplitude (0-1) callbacks for the active stream. Essential for building highly responsive, audio-reactive 3D Visualizers and HUDs! Fires with 0 and 'idle' between turns, so a meter falls to rest rather than freezing. | | onSubmit | (text: string, details: { speechMs?: number }) => Promise<string \| AsyncIterable<string> \| ReadableStream> | undefined | Connected Brain API: Bypasses local LLMs; routes transcribed user microphone strings to your cloud or custom LLM API endpoint. Return '' to stay quiet and keep listening. speechMs is how long the user spoke. See Staying quiet. | | preloadLocalLlm | boolean | false | Only meaningful alongside onSubmit, which otherwise skips the local language model download entirely. Set it to fetch that model in the background while your hosted one answers, so the conversation survives a rate limit, an expired quota or a lost network. See Hosted now, local when warm. | | onLocalLlmReady | () => void | undefined | Fires once the model requested by preloadLocalLlm has loaded. Drop onSubmit here to hand the conversation over; the turns your backend answered are carried across, so the local model knows what was already said. | | onTranscribe | (audio: Float32Array) => Promise<string> | undefined | Replaces local speech recognition with your own service, and Whisper is then not downloaded. Receives one utterance as 16 kHz mono samples. | | onSynthesize | (text: string) => Promise<Float32Array \| ArrayBuffer> | undefined | Replaces local voice synthesis with your own service, and Kokoro is then not downloaded. Return an encoded MP3/WAV buffer, or raw 24 kHz PCM. Called a sentence at a time. | | onError | (e: AiVoiceAvatarError) => void | undefined | Fires when a stage fails. Carries stage, message and a severity of degraded or fatal. See Handling failures. | | onTranscriptUpdate | (text: string, speaker: 'user' \| 'avatar') => void | undefined | Callback delivering real-time microphone transcriptions and assistant spoken utterance strings. | | onStatusChange| (status: string) => void | undefined | Emits live state transitions (loading, idle, listening, thinking, speaking). | | debug | boolean | false | When true, renders an interactive floating GUI (Leva) to inspect and tune individual 3D blendshapes. | | vadAssetPath | string | undefined | Optional URL or local path override for self-hosting @ricky0123/vad-web ONNX asset binaries in airgapped deployments. | | onnxWasmPath | string | undefined | Optional URL override for self-hosting onnxruntime-web WASM distribution files. | | workerBaseUrl| string | undefined | CSP Escape Hatch: if blob: workers are blocked by your server, fetch pre-compiled Web Workers from this URL directory. | | enableLocalAssetProbe | boolean | false | When true, HEAD-checks /ananya.glb in your own public directory before falling back to the CDN. Off by default: with no local copy the probe 404s, and that 404 lands in every visitor's console. | | statusPillStyle | React.CSSProperties | undefined | Optional custom CSS styling & absolute positioning overrides for the interactive Status Pill overlay. | | accentColor | string | undefined | Custom CSS color string (e.g., #38BDF8) for the active status indicator rings and highlights. |


Imperative Ref API (AiVoiceAvatarHandle)

Attach a React ref (useRef<AiVoiceAvatarHandle>(null)) to access imperative real-time controls:

interface AiVoiceAvatarHandle {
  /** Command the 3D avatar to speak an arbitrary string with synchronized acoustic lip blending */
  speak: (text: string) => void;
  /** Manually engage microphone recording and Voice Activity Detection (VAD) */
  startListening: () => void;
  /** Pause active microphone listening */
  stopListening: () => void;
  /** Instantly interrupt and halt active voice speech synthesis and clear the audio queue */
  interrupt: () => void;
  /** Manually submit text to the onSubmit handler, simulating a spoken utterance (useful for text-only fallback) */
  sendText: (text: string) => void;
  /** Wipe multi-turn conversation memory history and caption overlay states */
  clearHistory: () => void;
  /** Retrieve live Web Audio API AnalyserNode powering real-time spectral lip sync */
  getAnalyser: () => AnalyserNode | undefined;
}

💬 Text-Only Input (sendText)

If your users cannot use a microphone (e.g., noisy environments, privacy concerns, or lack of permissions), you can easily wire up a standard text input field to bypass the speech recognition pipeline entirely!

Simply attach a ref and call sendText() to pass a string directly to your onSubmit handler (or local LLM):

const avatarRef = useRef<AiVoiceAvatarHandle>(null);

// In your UI, attach this to a standard <form> submission:
const handleTextSubmit = (userInput: string) => {
  avatarRef.current?.sendText(userInput);
}

When you use sendText, the avatar immediately enters the thinking state and processes the interaction exactly as if the user had spoken it aloud.

useAiVoiceAvatar() (headless)

import { useAiVoiceAvatar } from 'react-ai-voice-avatar/headless';

Takes the same options as the component's props, except the ones about the 3D scene and its overlays (avatarPreset, avatarSize, modelSrc, lightingPreset, gestures, showCaptions, hideStatusPill, onModelLoaded, debug and the styling props). onStatusChange is replaced by the returned status. A few options matter mostly without an avatar:

| Option | Type | Description | | :--- | :--- | :--- | | onAudioLevelChange | (level: number, source: 'mic' \| 'tts' \| 'idle') => void | Loudness from 0 to 1, every frame, from whichever side is talking. What an orb or waveform animates from. Write it to the DOM through a ref rather than into state. | | onSpeechStart | (text: string) => void | Each sentence of the reply as it starts playing. Captions that keep pace with the voice. | | onTranscriptUpdate | (text: string, speaker: 'user' \| 'avatar') => void | What the user said, once transcribed, and the reply in full. | | onInferenceStart / onInferenceEnd | () => void | Around each turn, from the moment the user stops talking to the end of the reply. | | onTtsEngineChange | (engine: 'kokoro' \| 'mms' \| 'custom') => void | Fires whenever activeTtsEngine changes. Also a prop on the component. | | preloadLocalSpeech | boolean | Download the in-browser hearing and voice behind onTranscribe and onSynthesize, without waiting for them. Drop the adapters once onLocalSpeechReady fires and the local models take over. Also a prop on the component. | | onLocalSpeechReady | () => void | Fires once, when the models requested by preloadLocalSpeech are loaded. | | loadingProgress | (pct: number, label: string) => void | Download progress per model: 'asr', 'kokoro' (or 'tts' for MMS) and 'llm'. They download in parallel, so keep one figure per label. |

It returns:

| Value | Type | Description | | :--- | :--- | :--- | | status | 'loading' \| 'idle' \| 'listening' \| 'thinking' \| 'speaking' | Where the conversation is. 'loading' until the models are up, and until loadModels is true. | | isLoading, isIdle, isListening, isThinking, isSpeaking | boolean | Shorthands for status. | | isReady | boolean | The models are loaded. startListening does nothing before this. | | isLocalSpeechReady | boolean | The in-browser hearing and voice are loaded, whether or not they are in use. | | activeTtsEngine | 'kokoro' \| 'mms' \| 'custom' | The voice actually speaking. iPhones and iPads get 'mms', a single plainer voice that ignores ttsVoice, because Kokoro runs Safari out of memory; so does any device where Kokoro fails to load. 'custom' is your onSynthesize. Worth showing if your users will compare devices. | | startListening | () => Promise<void> | Opens the microphone and starts listening. Call it from a click: the first call is where the browser asks for permission. In 'continuous' mode the microphone then stays open across turns. | | stopListening | () => void | Closes the microphone. | | interrupt | () => void | Stops the reply mid-sentence and clears what was queued. | | speak | (text: string) => void | Says the text without a model turn: a greeting, a notification. Plays a sentence at a time. | | sendText | (text: string) => void | Typed input, handled as if it had been spoken. | | clearHistory | () => void | Forgets the conversation so far. | | micError | string \| null | Why the microphone could not open, such as a denied permission. null otherwise. | | analyser | AnalyserNode \| undefined | The reply's audio, for a frequency visualiser. onAudioLevelChange is simpler if you only need loudness. |

It also returns several refs (currentSpeechTextRef, audioContextRef and others) that the 3D component uses for lip-sync. A voice UI can ignore them.


🌐 Performance & Asset Caching

  1. Native WebGPU & WASM Degradation:
    • Modern Chromium browsers (Chrome, Edge, Opera, Arc) on desktop and mobile platforms benefit from hardware-accelerated WebGPU neural execution.
    • On systems without WebGPU, inference falls back to WebAssembly (WASM), which is slower.
  2. Persistent Local Caching:
    • AI models (Whisper, Kokoro, the local language model) are downloaded once and kept in the browser's origin private file system, so later visits skip the download. See Models are kept between visits.

🤝 Contributing & Open Issues Roadmap

We actively welcome community contributions. CONTRIBUTING.md has the local development guide, and ROADMAP.md has what is worth doing next, why, and what is already known about each item — including the measurements behind the open questions.

The largest pieces currently open:

  1. 🖐️ Open-palm gestures. Hands now gesture while the avatar speaks, but always with the palms facing inward. Turning them up — the open, offering gesture people use when explaining — needs the forearm's roll, which has to be worked out from the finger bones rather than assumed, for the same reason as everything else in armRig.ts.
  2. 🎤 Turn-taking in real rooms. Turns are tuned against one speaker in a quiet room. A short "yes" can be dropped, and a quiet speaker can go unheard. Recordings from more voices and noisier rooms would let the defaults be set against something other than one person.
  3. 🎭 Expanding regional 3D avatar personas. Ananya and Aarav ship out of the box. Royalty-free character GLBs (~3MB) rigged with the standard 52 Apple ARKit facial blendshapes are welcome. Avatar meshes are served from a CDN rather than bundled, so adding one adds nothing to the npm install (2.3 MB tarball, 6.0 MB unpacked, asserted in CI by scripts/verify-pack.mjs).
  4. 🙌 Gestures. The avatar stands still while it speaks. Hand and arm movement tied to speech is the most visible thing still missing.
  5. 📱 React Native / Expo support. Exploring bindings to run ONNX inference and Three.js on mobile runtimes.

Hindi speech and VAD sensitivity tuning (speechDetection) were previously listed here and have both shipped.


🧭 Browser Compatibility Matrix

The library relies on WebGL, Web Audio, Web Workers and, where available, WebGPU. Where WebGPU is missing the models run on WebAssembly instead, which is slower. Only desktop Chrome has been tested. The rest of this table is what the code is built to do, not what has been seen on that browser.

| Browser | Tested | What to expect | |---|---|---| | Chrome, desktop | ✅ macOS (Apple M4) by hand; Linux in CI | WebGPU for the models. The reference setup. | | Edge, desktop | Not yet | The same engine as Chrome, so expected to behave the same. | | Chrome, Android | Not yet | WebGPU on devices that have it, WebAssembly otherwise. Speed and memory not measured. | | Safari, iPhone and iPad | Not yet | Uses the smaller MMS voice to stay inside Safari's memory limit. Not run on a device. | | Safari, macOS | Not yet | Expected to work. Not checked. | | Firefox | Not yet | Expected to work, more slowly where WebGPU is unavailable. Not checked. |

Tried it on one of these? An issue saying what worked and what didn't fills in this table.

[!NOTE]

  • WebGPU is currently enabled by default in Chrome/Edge. On browsers without WebGPU, the library automatically falls back to WASM execution.
  • iPhone and iPad: Safari gives a tab far less memory than desktop browsers, and the Kokoro voice (~310 MB) is expected to exceed it, so the engine uses the smaller MMS voice on iOS. Elsewhere, if two page loads in a row start building Kokoro and never finish, that browser switches to the MMS voice rather than crash a third time. Neither has been seen on a real device yet.
  • Strict CSP Environments: Safari and Firefox may block blob: worker execution depending on your Content-Security-Policy headers. If this occurs, host the .worker.js files statically and pass their base path via the workerBaseUrl prop.

💻 Hardware Requirements

Running neural networks in the browser needs capable hardware. The figures below are estimates: only an Apple M4 MacBook has been measured, and no phone has been tried yet.

| Deployment Mode | Estimated RAM | GPU | Tested on | |---|---|---|---| | Cloud speech (your STT, model and voice) | Little beyond the page | None | Headless Chromium in CI, with stand-in providers | | Connected Brain (local speech, your model) | 4 GB | None; WebAssembly works, more slowly | Apple M4 MacBook, Chrome | | Full Local AI (local speech + 500M model) | 8 GB | WebGPU preferred | Apple M4 MacBook, Chrome |

[!TIP] Mobile Memory Limits: Mobile browsers rigidly enforce memory limits per tab (often terminating tabs exceeding ~1GB). If your mobile app crashes "after some time", ensure you are utilizing the Connected Brain mode (onSubmit API) which moves the language model's memory to your server while lip-sync and the voice stay local. On a phone, cloud speech uses the least memory of all.


📜 License

MIT © React AI Voice Avatar Contributors.