@aituber-onair/transcription
v0.0.3
Published
Provider-neutral realtime transcription for AITuber OnAir
Maintainers
Readme
@aituber-onair/transcription

Provider-neutral realtime microphone transcription for AITuber OnAir.
This package is an alpha release. Its public API may change before a stable release.
The package supports Web Speech, OpenAI Realtime transcription over browser WebRTC, Gemini Live transcription over browser WebSocket, and local Whisper Tiny, Base, and Small inference through WebGPU. All providers emit the same per-utterance snapshot events. File transcription, server WebSocket input, automatic chat submission, and provider fallback are intentionally out of scope.
Browser example
The package includes a framework-free browser example that exercises all four providers without depending on AITuber OnAir Core:
cd packages/transcription/examples/browser-basic
npm run devAfter installing dependencies from the repository root, you can also start the example with a workspace command:
npm -w @aituber-onair/transcription run example:devOpen the displayed localhost URL and grant microphone permission when starting a session. Web Speech and Local Whisper need no key. For OpenAI or Gemini, enter an end-user-owned API key in the page. The sample connects to the selected service directly from the browser, so avoid using it on a shared device. Local Whisper requires WebGPU; its first start downloads model/runtime assets and caches them in the browser. The interface supports English and Japanese and selects the initial display from the browser language. See the example README for details.
Build the example without starting a server:
npm -w @aituber-onair/transcription run example:buildUsage
import { createRealtimeTranscriptionSession } from '@aituber-onair/transcription';
const session = createRealtimeTranscriptionSession({
provider: 'web-speech',
language: 'ja-JP',
});
session.onTranscript(({ utteranceId, text, isFinal }) => {
console.log({ utteranceId, text, isFinal });
});
await session.start();
// Later:
await session.stop();
await session.dispose();Only providers with potentially long initialization emit onProgress events.
Currently, only local-whisper emits them; Web Speech, OpenAI Realtime, and
Gemini Live do not.
Local Whisper
Local Whisper runs the selected Whisper model in a module worker and emits final transcripts only:
import { createRealtimeTranscriptionSession } from '@aituber-onair/transcription';
const session = createRealtimeTranscriptionSession({
provider: 'local-whisper',
model: 'tiny',
language: 'ja-JP',
silenceDurationMs: 500,
});
session.onTranscript(({ text, isFinal }) => {
if (isFinal) {
console.log(text);
}
});
session.onProgress(({ phase, progress }) => {
updateLoadingIndicator(phase, progress);
});
session.onError((error) => {
console.error(error.code, error.message);
});
await session.start();
// Later:
await session.stop();
await session.dispose();Local Whisper is less accurate than Web Speech or OpenAI Realtime. It is
intended for use cases that prioritize requiring no API key and not sending
microphone audio to a remote service. Choose small when recognition quality
is important.
| Model | First download reported by progress | Quality guide | Inference (Japanese / English) |
| --- | ---: | --- | ---: |
| tiny (default) | About 122 MB | Lower | 237.3 ms / 203.0 ms |
| base | About 209 MB | Middle | 255.9 ms / 311.2 ms |
| small | About 589 MB | Practical | 574.7 ms / 551.6 ms |
These measurements were taken in Chrome with WebGPU using the same short
Japanese and English microphone clips. Inference excludes capture/VAD time and
varies by GPU. First-use download time depends on network speed and can take
several minutes for hundreds of MB. After caching, initialization measured
about 0.9 s for Tiny, 1.2 s for Base, and 2.5 s for Small. Download sizes are
the sum of the latest totalBytes reported for each model file and do not
include assets that do not report progress.
Requirements and behavior:
- A secure browser context (HTTPS or localhost), microphone access, Web Audio, AudioWorklet, module workers, and WebGPU are required.
- No API key is required. There is no automatic fallback to a remote provider or WASM inference when WebGPU initialization fails.
- On first use, the selected model assets are downloaded from the Hugging Face Hub and ONNX Runtime WebAssembly files are downloaded from jsDelivr. These assets are cached by the browser. Larger models take longer to download and infer.
- Microphone audio is processed in the browser and never leaves the browser. The package does not persist audio or transcripts.
languageaccepts a BCP 47-style hint and is optional. The default 500 mssilenceDurationMscan be reduced to a minimum of 150 ms for faster turn completion.modelacceptstiny,base, orsmalland defaults totiny. Model dtype is fixed to an fp32 encoder and q4 merged decoder for every size.- Download progress can include
file,loadedBytes,totalBytes, and a normalizedprogressvalue from 0 to 1. Initialization and ready phases do not require byte totals.
The package normally resolves dist/local-whisper.worker.js relative to its
ESM entry. If a bundler pre-bundles the package and cannot resolve that asset,
set the advanced workerUrl option to the same module worker asset or an
equivalent build. This browser example uses Vite's ?worker&url import for that
reason.
OpenAI Realtime
OpenAI Realtime uses gpt-live-transcribe and browser WebRTC. Because this
transcription model does not accept server turn detection, the package detects
sustained silence through the browser Web Audio API and explicitly commits each
audio turn. The recommended authentication mode obtains a short-lived client
secret from an application backend:
const session = createRealtimeTranscriptionSession({
provider: 'openai-realtime',
auth: {
type: 'client-secret',
getClientSecret: async () => {
const response = await fetch('/api/openai/realtime/client-secret', {
method: 'POST',
});
const data = await response.json();
return data.value;
},
},
languages: ['ja', 'en'],
keywords: ['AITuber OnAir'],
prompt: 'An AITuber livestream.',
delay: 'low',
});Frontend-only, self-hosted applications may explicitly use an end-user-owned standard API key to mint a client secret in the browser:
const session = createRealtimeTranscriptionSession({
provider: 'openai-realtime',
auth: {
type: 'browser-api-key',
getApiKey: async () => readEndUserKeyAtRuntime(),
acknowledgeBrowserKeyRisk: true,
},
languages: ['ja'],
});Gemini Live
Gemini Live uses gemini-3.5-transcribe-live and streams raw 16-bit PCM audio
over a browser WebSocket. It emits low-latency interim snapshots and a final
snapshot for each detected utterance. The model supports automatic language
detection, language hints, custom vocabulary, and verbatim or smart
transcription modes.
The recommended browser authentication mode obtains a short-lived ephemeral token from an application backend:
const session = createRealtimeTranscriptionSession({
provider: 'gemini-live',
auth: {
type: 'ephemeral-token',
getEphemeralToken: async () => {
const response = await fetch('/api/gemini/ephemeral-token', {
method: 'POST',
});
const data = await response.json();
return data.name;
},
},
languages: ['ja-JP', 'en-US'],
keywords: ['AITuber OnAir'],
mode: 'smart',
});Frontend-only, self-hosted applications may explicitly connect with an end-user-owned API key:
const session = createRealtimeTranscriptionSession({
provider: 'gemini-live',
auth: {
type: 'browser-api-key',
getApiKey: async () => readEndUserKeyAtRuntime(),
acknowledgeBrowserKeyRisk: true,
},
languages: [], // Automatic language detection
mode: 'verbatim',
});Gemini Live Transcribe currently has a 10-minute connection limit. This provider reports an unexpected server disconnect as a typed connection error; applications that need longer listening periods should start a new session. Live streaming does not provide speaker diarization or word-level timestamps. See the Gemini Live transcription documentation for current preview status, supported languages, limits, and pricing links.
Security
OpenAI and Google recommend keeping standard API keys on a server and minting short-lived browser credentials. The browser-BYOK modes exist for trusted frontend-only or self-hosted use. They must use a key owned and supplied by the end user; never bundle an application-owner key in source code or built assets.
The package requests credentials through getApiKey(), getClientSecret(), or
getEphemeralToken() for each start() and does not persist, cache, return, or
log them. A consuming application still controls its own storage. Browser
persistence can expose a key to XSS, extensions, local device access, or
compromised dependencies. Credential failures are returned as typed errors and
never trigger an authentication fallback.
Provider differences
| Capability | Web Speech | OpenAI Realtime | Gemini Live | Local Whisper | | --- | --- | --- | --- | --- | | Interim snapshots | Yes | Yes | Yes | No | | Multiple expected languages | No | Yes | Yes | No | | Keywords / custom vocabulary | No | Yes | Yes | No | | Smart transcription | No | No | Yes | No | | Configurable delay | No | Yes | No | Yes | | Utterance boundary | Browser implementation | Browser audio-level detection | Gemini server VAD | Browser PCM/VAD |
All providers require a supported browser and microphone permission. OpenAI WebRTC, Gemini Live, and Local Whisper also require the Web Audio API and HTTPS or localhost; Gemini Live additionally requires WebSocket and Local Whisper requires WebGPU. Web Speech availability and behavior vary by browser. Remote providers can incur usage charges while listening, so applications should expose state clearly and stop sessions when unused.
