npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@aituber-onair/voice

v0.26.0

Published

Voice synthesis library for AITuber OnAir

Readme

AITuber OnAir Voice

AITuber OnAir Voice - logo

@aituber-onair/voice is an independent voice synthesis library that supports multiple TTS (Text-to-Speech) engines. While originally developed for the AITuber OnAir project, it can be used standalone for any voice synthesis needs.

日本語版はこちら

This project is published as open-source software and is available as an npm package under the MIT License.

Table of Contents

Overview

@aituber-onair/voice is a comprehensive voice synthesis library that provides a unified interface for multiple TTS engines. It specializes in emotion-aware speech synthesis, making it ideal for creating expressive virtual characters, AI assistants, and interactive applications.

Key design principles:

  • Engine Independence: Switch between TTS engines without changing your code
  • Emotion Support: Built-in emotion detection and synthesis
  • Browser Ready: Full support for web audio playback
  • TypeScript First: Complete type safety and excellent IDE support
  • Zero Dependencies: Minimal external dependencies for maximum compatibility

Installation

Install using npm:

npm install @aituber-onair/voice

Or using yarn:

yarn add @aituber-onair/voice

Or using pnpm:

pnpm install @aituber-onair/voice

Main Features

  • Multiple TTS Engine Support
    Compatible with VOICEVOX, VoicePeak, OpenAI TTS, OpenRouter TTS (preview opt-in), xAI TTS, Unreal Speech, ElevenLabs, Fish Audio, Cartesia, Deepgram Flux, Inworld, Gradium, Gemini TTS, MiniMax, AivisSpeech, Aivis Cloud, Web Speech API, and more
  • Unified Interface
    Single API for all supported TTS engines
  • Emotion-Aware Synthesis
    Automatically detects and applies emotions from text tags like [happy], [sad], etc.
  • Screenplay Conversion
    Transforms text with emotion tags into structured screenplay format
  • Browser Audio Support
    Direct playback in web browsers using HTMLAudioElement
  • Custom Endpoints
    Support for self-hosted TTS servers
  • Language Detection
    Automatic language recognition for multi-language engines
  • Flexible Configuration
    Runtime engine switching and parameter updates

Basic Usage

Simple Text-to-Speech

import { VoiceService, VoiceServiceOptions } from '@aituber-onair/voice';

// Configure the voice service
const options: VoiceServiceOptions = {
  engineType: 'voicevox',
  speaker: '1',
  // Optional: specify custom endpoint
  voicevoxApiUrl: 'http://localhost:50021'
};

// Create voice service instance
const voiceService = new VoiceService(options);

// Speak text
await voiceService.speak({ text: 'Hello, world!' });

Using VoiceEngineAdapter (Recommended)

import { VoiceEngineAdapter, VoiceServiceOptions } from '@aituber-onair/voice';

const options: VoiceServiceOptions = {
  engineType: 'openai',
  speaker: 'alloy',
  apiKey: 'your-openai-api-key',
  onPlay: async (audioBuffer) => {
    // Custom audio playback handler
    console.log('Playing audio...');
  }
};

const voiceAdapter = new VoiceEngineAdapter(options);

// Speak with emotion
await voiceAdapter.speak({ 
  text: '[happy] I am so excited to talk with you!' 
});

Supported TTS Engines

VOICEVOX

High-quality Japanese speech synthesis engine with multiple character voices.

const voiceService = new VoiceService({
  engineType: 'voicevox',
  speaker: '1', // Character ID
  voicevoxApiUrl: 'http://localhost:50021' // Optional custom endpoint
});

VoicePeak

Professional speech synthesis with rich emotional expression.

const voiceService = new VoiceService({
  engineType: 'voicepeak',
  speaker: 'f1',
  voicepeakApiUrl: 'http://localhost:20202',
  voicepeakEmotion: 'happy',
  voicepeakSpeed: 140,
  voicepeakPitch: 20
});

Single-tag voicepeakEmotion remains backward compatible with existing VoicePeak setups. Weighted emotion maps require vpeakserver >= v0.2.0.

const weightedVoiceService = new VoiceService({
  engineType: 'voicepeak',
  speaker: 'f1',
  voicepeakApiUrl: 'http://localhost:20202',
  voicepeakEmotion: { happy: 40, fun: 60 },
});
  • neutral is ignored when sending weighted emotions.
  • Weight 0 is ignored.
  • {} means "do not send emotion" and does not fall back to Talk.style.
  • undefined means no override, so Talk.style still maps to a single tag.

OpenAI TTS

OpenAI's text-to-speech API with multiple voice options.

const voiceService = new VoiceService({
  engineType: 'openai',
  speaker: 'alloy',
  apiKey: 'your-openai-api-key'
});

OpenRouter TTS (public-preview opt-in)

Preview models require an explicit choice and are not defaults. This dedicated engine supports microsoft/mai-voice-2.1 and microsoft/mai-voice-2.1-flash through OpenRouter's POST /api/v1/audio/speech. The OpenRouter TTS guide documents both exact model IDs, full model-specific voice IDs, and 24 kHz mono PCM output. Microsoft classifies both as public preview, without an SLA and not recommended for production. The model catalogs checked on October 1, 2026 contained 97 voices across 28 locales for each model, with no Japanese voices. Voice availability may change; fetch the selected model's catalog instead of assuming a fixed list.

import {
  VoiceEngineAdapter,
  getVoiceEngineVoiceList,
  type OpenRouterTtsModel,
} from '@aituber-onair/voice';

// Deliberate preview opt-in; no model is chosen when this option is omitted.
const model: OpenRouterTtsModel = 'microsoft/mai-voice-2.1';
const voices = await getVoiceEngineVoiceList('openRouter', {
  openRouterModel: model,
  // Optional for the public models catalog; synthesis requires your key.
  apiKey: 'your-openrouter-api-key',
});
const speaker = voices.find((voice) =>
  voice.id === 'en-US-Harper:MAI-Voice-2.1',
)?.id;
if (!speaker) throw new Error('Select an available voice for this model');

const voiceService = new VoiceEngineAdapter({
  engineType: 'openRouter',
  openRouterModel: model,
  speaker,
  apiKey: 'your-openrouter-api-key',
  // Optional full speech endpoint (not a base URL):
  // openRouterApiUrl: 'https://openrouter.ai/api/v1/audio/speech',
});
await voiceService.speakText('Hello! Welcome to the show.');

The engine sends Bearer authentication and JSON { model, input, voice, response_format: 'pcm' }. The speech API defines audio/pcm as 16-bit little-endian; the MAI route specifies 24 kHz mono. The engine wraps these bytes in a WAV header using the existing PCM helper, so the browser/Node audio player and onPlay callback receive a complete WAV buffer, not compressed MP3 or headerless PCM. Node playback still needs an optional playback dependency as described below. Non-PCM, empty, or incomplete 16-bit sample responses are rejected before playback. This is one-shot synthesis, not realtime audio streaming. Emotion tags do not add voice style, speed, or cloning parameters; these controls are intentionally omitted. The voice locale determines the synthesis language.

getVoiceEngineVoiceList('openRouter', { openRouterModel }) reads GET https://openrouter.ai/api/v1/models?output_modalities=speech and returns only that exact model's supported_voices. Use openRouterModelsApiUrl (or voiceListApiUrl) on the lookup options for a separate models endpoint; a custom openRouterApiUrl does not redirect catalog requests. Relative endpoint URLs work in browsers; Node requires absolute HTTP(S) URLs. Configure only trusted endpoints because requests with an API key send it to the configured destination.

When changing models at runtime, choose a voice from the new model's list and update openRouterModel and speaker together. Updating only the model clears the prior speaker. Full IDs end with :MAI-Voice-2.1 or :MAI-Voice-2.1-Flash; a stale suffix is rejected, never rewritten automatically. The React example starts with no preview model or voice selected and provides model-scoped voice selection, API key and separate endpoint controls.

The example uses the direct OpenRouter route. Browser-side keys are visible to the page, so deploy shared apps with your own credential-protecting backend rather than embedding a shared secret in client code. Direct access depends on the provider's CORS policy and your deployment; custom endpoints must permit your application's origin or run as same-origin backend routes.

xAI TTS

xAI's cloud TTS API with selectable voice IDs, language control, and output format tuning.

const voiceService = new VoiceService({
  engineType: 'xai',
  speaker: 'eve',
  apiKey: 'your-xai-api-key',
  xaiLanguage: 'ja',
  xaiCodec: 'mp3',
  xaiSampleRate: 24000,
  xaiBitRate: 128000,
});

Unreal Speech

Unreal Speech v8 cloud TTS via the /stream endpoint. It returns audio bytes directly, so it works with VoiceEngineAdapter playback without an extra download step.

const voiceService = new VoiceService({
  engineType: 'unrealSpeech',
  speaker: 'af_bella',
  apiKey: 'your-unreal-speech-api-key',
  unrealSpeechBitrate: '192k',
  unrealSpeechSpeed: 0,
  unrealSpeechPitch: 1,
  unrealSpeechCodec: 'libmp3lame',
  unrealSpeechTemperature: 0.25,
});

Use unrealSpeechApiUrl to override the default https://api.v8.unrealspeech.com/stream endpoint.

ElevenLabs

ElevenLabs Text to Speech API support using direct fetch calls. No SDK is required.

const voiceService = new VoiceService({
  engineType: 'elevenLabs',
  speaker: 'JBFqnCBsd6RMkjVDRZzb',
  apiKey: 'your-elevenlabs-api-key',
  elevenLabsModel: 'eleven_flash_v2_5',
  elevenLabsOutputFormat: 'mp3_44100_128',
  elevenLabsStability: 0.5,
  elevenLabsSimilarityBoost: 0.75,
  elevenLabsUseSpeakerBoost: true,
});

Use elevenLabsApiUrl to override the default https://api.elevenlabs.io/v1/text-to-speech endpoint. The speaker value is sent as the ElevenLabs voice_id.

The curated model choices include eleven_v4 for highest-quality speech, eleven_v3 for previous-generation expressiveness, eleven_multilingual_v2 for high-quality multilingual output, and eleven_flash_v2_5 (the default) for low latency. The deprecated eleven_turbo_v2_5 remains accepted as a custom string for backward compatibility, but is no longer presented as a recommended model.

Model-specific settings are filtered when building requests. For eleven_v4, only Stability and Similarity are sent; configured Style, Speed, and Speaker Boost values are retained for other models but omitted for v4. SSML is not supported by v4; use plain text or its documented audio tags. This engine uses one-shot Create speech and waits for the complete audio response. The official TTS guide confirms v4 support for that endpoint. Models outside the verified integration scope are not offered as supported choices: eleven_v4_turbo is documented for the Text to Dialogue WebSocket, which this package does not implement. Turbo has not been verified through this package's HTTP speech path; this is not a claim that the provider cannot support it over HTTP.

Fish Audio

Fish Audio one-shot TTS uses POST /v1/tts and returns audio bytes directly.

const voiceService = new VoiceService({
  engineType: 'fishAudio',
  speaker: 'your-reference-id',
  apiKey: 'your-fish-audio-api-key',
  fishAudioModel: 's2-pro',
  fishAudioFormat: 'mp3',
  fishAudioLatency: 'normal',
});

speaker is sent as reference_id. Use getVoiceEngineVoiceList('fishAudio', { apiKey }) to list usable models/voices. s2-pro is the stable default for this integration; s2.1-pro-free must be selected explicitly and should not be treated as an SLA-backed production tier.

Voice-list lookups return at most 100 entries by default so they do not traverse the full public model catalog. Pass limit and optionally pageSize to request a larger bounded result set.

The official POST /v1/tts endpoint does not currently complete browser CORS preflight requests. Call it directly from Node.js or a server, or set fishAudioApiUrl to a same-origin backend route in browser applications. Keep the Fish Audio API key on the server in production. The React example includes a Vite-only development/preview proxy for this purpose.

Cartesia

Cartesia synchronous TTS uses POST /tts/bytes and returns audio bytes directly.

const voiceService = new VoiceService({
  engineType: 'cartesia',
  speaker: 'your-cartesia-voice-id',
  apiKey: 'your-cartesia-api-key',
  cartesiaModel: 'sonic-3.5',
  cartesiaLanguage: 'ja',
  cartesiaOutputContainer: 'wav',
  cartesiaSampleRate: 44100,
});

Use getVoiceEngineVoiceList('cartesia', { apiKey, language: 'ja' }) to list voices. The engine retains the Cartesia-Version: 2026-03-01 header and voice: { id } request shape. Set cartesiaModel: 'sonic-3.6' explicitly for the newer model, or 'sonic-3.6-2026-08-27' to pin its stable snapshot. The existing sonic-3.5 default is unchanged. Cartesia documents Sonic 3.6 on POST /tts/bytes and backward compatibility for pinned integrations. WAV/MP3 output and existing voice IDs continue through the same path.

Deepgram Flux

Deepgram Flux uses the one-shot POST /v2/speak endpoint with Authorization: Token ... and returns MP3 audio bytes. This integration supports English Flux voices only. It does not implement Aura (/v1/speak), WebSocket streaming, interruption handling, callbacks, or beta expressivity.

const voiceService = new VoiceService({
  engineType: 'deepgram',
  speaker: 'flux-haley-en',
  apiKey: process.env.DEEPGRAM_API_KEY,
  deepgramSpeed: 1.0,
});

The required full voice ID in speaker is sent as the model query parameter. The example defaults to flux-haley-en. deepgramSpeed is optional (0.5–1.5, in 0.05 steps). Use getVoiceEngineVoiceList('deepgram') to fetch the public v2 model catalog (GET /v2/models, no API key) and keep its English Flux entries. The v1 catalog (/v1/models) lists Aura voices only. The React example always offers the documented Haley preset independently of the catalog. Its optional catalog refresh preserves the selected voice, and empty results or errors also preserve the existing list. See the voice catalog.

Call from Node.js/backend, or use deepgramApiUrl for a same-origin speech proxy and voiceListApiUrl for the catalog proxy. Keep the API key on the server in production. Direct browser CORS behavior has not been live-verified; the React example supplies Vite development/preview proxies under /api/deepgram. Production deployments must provide equivalent backend routes.

Inworld

Inworld TTS non-streaming speech synthesis using direct fetch calls. This engine uses the REST endpoint only; WebSocket and HTTP streaming are not implemented.

const voiceService = new VoiceService({
  engineType: 'inworld',
  speaker: 'Ashley',
  apiKey: process.env.INWORLD_API_KEY,
  inworldModel: 'inworld-tts-2',
  inworldAudioEncoding: 'MP3',
  inworldSampleRateHertz: 48000,
});

Use inworldApiUrl to override the default https://api.inworld.ai/tts/v1/voice endpoint. The apiKey value should be the Inworld Basic Base64 authorization value. Do not expose Basic credentials in browser-side code; use a backend proxy or Inworld JWT authentication for browser apps.

The current TTS-2 models are inworld-tts-2 (the default, for quality) and inworld-tts-2-flash (an explicit option for lower latency and cost). Set inworldModel: 'inworld-tts-2-flash' to use Flash, or switch at runtime with voiceService.updateOptions({ inworldModel: 'inworld-tts-2-flash' }). Both use the same REST endpoint and voice-list configuration. This engine waits for the complete audio response, so provider streaming latency figures do not describe its playback latency.

inworldDeliveryMode is supported only by TTS-2 and is omitted for Flash. Earlier model IDs remain accepted as strings for compatibility, but deprecated TTS 1.5 models are no longer offered in the React model selector. See the model catalog and speech API reference.

Gradium

Gradium one-shot REST TTS support using direct fetch calls. This engine uses the raw-audio response mode (only_audio: true) and does not add the Gradium SDK.

Production remains the default. Experimental models require explicit opt-in: set gradiumModel: 'gradium-tts-beta' to try the public beta, or use gradiumModel: 'default' (or omit it) for the production model. The beta is a supported explicit option, not the recommended default. The React example exposes both choices. updateOptions({ gradiumModel: 'default' }) switches an existing adapter back to production; undefined clears the selection.

gradiumModel is sent as top-level model_name in the REST JSON body, separately from json_config. This follows the official REST guide and model selection guide. The existing one-shot raw-audio response and output format handling are unchanged; this does not add streaming synthesis or new emotion controls.

const voiceService = new VoiceService({
  engineType: 'gradium',
  speaker: 'YTpq7expH9539ERJ',
  apiKey: process.env.GRADIUM_API_KEY,
  gradiumOutputFormat: 'wav',
  gradiumTemperature: 0.7,
  gradiumVoiceSimilarity: 2,
  gradiumPaddingBonus: 0,
  gradiumRewriteRules: 'en',
});

Use gradiumApiUrl to override the default https://api.gradium.ai/api/post/speech/tts endpoint. The speaker value is sent as Gradium voice_id. The React example uses Gradium flagship voice presets as a fallback and can fetch the Gradium voice list through getVoiceEngineVoiceList() when an API key is provided. Browser-side voice list requests may fail if the Gradium API does not allow direct CORS access; use a backend proxy for production browser UIs that need dynamic Gradium voice selection.

OpenAI-Compatible TTS

OpenAI-compatible speech endpoints for self-hosted servers such as Kokoro FastAPI.

const voiceService = new VoiceService({
  engineType: 'openaiCompatible',
  openAiCompatibleApiUrl: 'http://localhost:8880/v1/audio/speech',
  openAiCompatibleModel: 'your-model-id'
});

speaker is optional for compatible endpoints. When omitted, the request body does not include a voice field. openAiCompatibleModel should be set explicitly to a model name accepted by your endpoint. openAiCompatibleApiUrl is used as-is, so pass the full /audio/speech URL. openAiCompatibleInstructions and openAiCompatibleResponseFormat are optional and sent as instructions / response_format only when non-empty. Servers that support instructions typically use it as a voice style prompt. Failed requests throw a VoiceEngineError with kind: 'api', the HTTP status in statusCode, and the response body in the message.

Endpoint setup helpers

These helpers make settings UIs easier to build. They accept an origin, an API base URL, or a full /audio/speech URL.

import {
  resolveOpenAICompatibleSpeechEndpoint,
  listOpenAICompatibleSpeechModels,
  listOpenAICompatibleSpeechVoices,
  getOpenAICompatibleSpeechServerInfo,
  testOpenAICompatibleSpeech,
} from '@aituber-onair/voice';

const endpoint = 'http://localhost:8880/v1';
const { speechUrl } = resolveOpenAICompatibleSpeechEndpoint(endpoint);
// -> http://localhost:8880/v1/audio/speech

const models = await listOpenAICompatibleSpeechModels({ endpoint });
const voices = await listOpenAICompatibleSpeechVoices({ endpoint }); // or null
const info = await getOpenAICompatibleSpeechServerInfo({ endpoint }); // or null

const result = await testOpenAICompatibleSpeech({
  endpoint,
  model: models[0],
  voice: voices?.defaultVoice,
});
if (result.ok) {
  // result.audio is an ArrayBuffer you can play back
} else {
  console.error(result.error.code, result.error.status, result.error.detail);
}
  • listOpenAICompatibleSpeechModels uses the standard GET /models.
  • OpenAI's API has no voice listing endpoint, so listOpenAICompatibleSpeechVoices is best-effort. It tries GET /audio/voices (for example Kokoro-FastAPI) and then GET /voices, and returns null when neither answers with a recognizable list.
  • getOpenAICompatibleSpeechServerInfo is also best-effort. It reads the server root and returns { engine, model?, defaultVoice? } only when the root answers with JSON that has an engine field. Otherwise it returns null.
  • testOpenAICompatibleSpeech synthesizes one short sentence and returns a result object instead of throwing. Errors carry a code (invalid-url, network, aborted, timeout, http, invalid-response), plus status and detail for HTTP errors.
  • Browser requests also need CORS permission from the server. A network error in the browser usually means CORS blocked the request or the server is not running.

MiniMax

Multi-language TTS supporting 24 languages with HD quality.

const voiceService = new VoiceService({
  engineType: 'minimax',
  speaker: 'Japanese_IntellectualSenior',
  apiKey: 'your-minimax-api-key',
  minimaxModel: 'speech-2.8-turbo',
  groupId: 'legacy-group-id', // Optional for older accounts
  endpoint: 'global' // or 'china'
});

Note: Current MiniMax T2A v2 endpoints require Bearer authentication but do not require GroupId. The optional groupId field is kept for older accounts or endpoints that still expect the query parameter. Speech 2.8 Turbo is the default; Speech 2.8 HD is available when output quality is preferred.

Use MiniMax system voice IDs for speaker, such as Japanese_IntellectualSenior. MiniMax documents these IDs in its System Voice ID List. The linked dynamic Get Voice API is not currently available, so getVoiceEngineVoiceList() does not expose MiniMax voice-list fetching.

AivisSpeech

AI-powered speech synthesis with natural voice quality.

const voiceService = new VoiceService({
  engineType: 'aivisSpeech',
  speaker: '888753760',
  aivisSpeechApiUrl: 'http://localhost:10101'
});

Aivis Cloud

High-quality cloud-based TTS service with advanced SSML support and streaming capabilities.

const voiceService = new VoiceService({
  engineType: 'aivisCloud',
  speaker: 'unused', // Not used when model UUID is specified
  apiKey: 'your-aivis-cloud-api-key',
  aivisCloudModelUuid: 'a59cb814-0083-4369-8542-f51a29e72af7', // Required
  
  // Optional advanced settings
  aivisCloudSpeakerUuid: 'speaker-uuid', // For multi-speaker models
  aivisCloudStyleId: 0, // Or use aivisCloudStyleName: 'ノーマル'
  aivisCloudUseSSML: true, // Enable SSML tags
  aivisCloudSpeakingRate: 1.0, // 0.5-2.0
  aivisCloudEmotionalIntensity: 1.0, // 0.0-2.0
  aivisCloudOutputFormat: 'mp3', // wav, flac, mp3, aac, opus
  aivisCloudOutputSamplingRate: 44100, // Hz
});

Key Features:

  • SSML Support: Rich markup for prosody, breaks, aliases, and emotions
  • Streaming Audio: Real-time audio generation and delivery
  • Multiple Formats: WAV, FLAC, MP3, AAC, Opus output
  • Emotion Control: Fine-grained emotional intensity settings
  • High Quality: Professional-grade voice synthesis

getVoiceEngineVoiceList('aivisCloud') uses the Aivis Cloud model search API (GET https://api.aivis-project.com/v1/aivm-models/search) and returns model UUID choices that can be passed as speaker or aivisCloudModelUuid. Direct browser requests to model/list endpoints can fail CORS checks, so browser apps should call this helper from a Node.js backend or relay/proxy when they need dynamic model, speaker, or style selection. The React example intentionally keeps manual model UUID input instead of calling the model search endpoint from the browser.

Gemini TTS

Gemini API text-to-speech supports gemini-3.8-flash-lite-tts for fast, cost-efficient speech and gemini-3.8-flash-tts for more expressive speech. Earlier preview models remain available. Both 3.8 models use the Interactions API and return WAV audio; preview models continue to use generateContent.

const voiceService = new VoiceService({
  engineType: 'geminiTts',
  speaker: 'Zephyr',
  apiKey: 'your-google-api-key',
  geminiTtsModel: 'gemini-3.8-flash-lite-tts',
  geminiTtsPrompt: 'cheerful and friendly', // Optional speech_metadata.style for 3.8
  geminiTtsApiUrl:
    'https://generativelanguage.googleapis.com/v1beta', // Optional Gemini API base URL
});

Note: Use a standard Google API key. apiKey is sent as x-goog-api-key to the Gemini API. speaker accepts a prebuilt voice name such as Zephyr or Kore; for 3.8 it also accepts a voice_... ID created in Google AI Studio. Gemini 3.8 detects the input language automatically, so geminiTtsLanguageCode applies only to the preview models. For 3.8, geminiTtsPrompt is sent as a style annotation instead of being spoken as part of the transcript. This engine returns complete audio rather than streaming chunks. Voice creation and replication are handled in Google AI Studio or the Gemini Voices API.

Web Speech API

Browser-native speech synthesis through window.speechSynthesis. This engine does not return audio bytes; the browser plays speech directly.

const voiceService = new VoiceEngineAdapter({
  engineType: 'webSpeech',
  speaker: '', // Optional: SpeechSynthesisVoice name or voiceURI
  webSpeechLanguage: 'ja-JP',
  webSpeechRate: 1.1,
  webSpeechPitch: 1.0,
  webSpeechVolume: 1.0,
});

await voiceService.speak({ text: 'こんにちは' });

Use getVoiceEngineVoiceList('webSpeech') in a browser to list available SpeechSynthesisVoice entries. Some browsers populate voices asynchronously, so the helper waits briefly for voiceschanged. Because no ArrayBuffer is available, onPlay(audioBuffer) is skipped for this engine; onComplete still runs when the utterance ends. Runtime support is browser-only.

None (Silent Mode)

No audio output - useful for testing or text-only scenarios.

const voiceService = new VoiceService({
  engineType: 'none'
});

Emotion-Aware Speech

The library supports emotion tags in text for more expressive speech:

// Emotion tags are automatically detected and processed
await voiceService.speak({ 
  text: '[happy] Great to see you today!' 
});

await voiceService.speak({ 
  text: '[sad] I will miss you...' 
});

await voiceService.speak({ 
  text: '[angry] This is unacceptable!' 
});

// Supported emotions vary by engine
// Common emotions: happy, sad, angry, surprised, neutral

The emotion system works by:

  1. Extracting emotion tags from the text
  2. Converting text to screenplay format with emotion metadata
  3. Passing emotion information to engines that support it
  4. Falling back gracefully for engines without emotion support

Browser Compatibility

The library includes built-in browser audio playback support:

// Option 1: Default browser playback
const voiceService = new VoiceService({
  engineType: 'openai',
  speaker: 'alloy',
  apiKey: 'your-api-key'
  // Audio will play automatically in the browser
});

// Option 2: Custom audio handling
const voiceService = new VoiceService({
  engineType: 'voicevox',
  speaker: '1',
  onPlay: async (audioBuffer: ArrayBuffer) => {
    // Custom audio playback logic
    const audioContext = new AudioContext();
    const audioBufferSource = audioContext.createBufferSource();
    // ... handle audio playback
  }
});

// Option 3: Specify HTML audio element
const voiceService = new VoiceService({
  engineType: 'voicevox',
  speaker: '1',
  voicevoxApiUrl: 'http://localhost:50021',
  audioElementId: 'my-audio-player' // ID of <audio> element
});

Advanced Configuration

Dynamic Engine Switching

const voiceAdapter = new VoiceEngineAdapter({
  engineType: 'voicevox',
  speaker: '1'
});

// Update options within the same engine
voiceAdapter.updateOptions({
  speaker: '3',
  voicevoxSpeedScale: 1.1,
});

// Switch to a different engine at runtime
voiceAdapter.switchEngine({
  engineType: 'openai',
  speaker: 'nova',
  apiKey: 'your-openai-api-key'
});

// Backward compatibility:
// updateOptions with engineType is still accepted.
voiceAdapter.updateOptions({
  engineType: 'openai',
  speaker: 'nova',
  apiKey: 'your-openai-api-key'
});

Custom Endpoints

// For self-hosted or custom TTS servers
const voiceService = new VoiceService({
  engineType: 'voicevox',
  speaker: '1',
  voicevoxApiUrl: 'https://my-custom-voicevox-server.com'
});

Engine Parameter Overrides

VoiceServiceOptions (see API Reference) now covers a consistent set of overrides for each engine. Below is a field-by-field summary to help you discover the right property without scanning the entire interface.

const voiceService = new VoiceService({
  engineType: 'voicevox',
  speaker: '1',
  openAiSpeed: 1.15,
  openAiCompatibleModel: 'your-model-id',
  openAiCompatibleSpeed: 1.1,
  openAiCompatibleTimeoutMs: 300_000, // Optional; default: 30_000 ms; 0 disables the timeout
  unrealSpeechBitrate: '192k',
  unrealSpeechSpeed: 0,
  unrealSpeechPitch: 1,
  elevenLabsModel: 'eleven_flash_v2_5',
  elevenLabsStability: 0.5,
  elevenLabsSimilarityBoost: 0.75,
  inworldModel: 'inworld-tts-2',
  inworldAudioEncoding: 'MP3',
  inworldSampleRateHertz: 48000,
  gradiumOutputFormat: 'wav',
  gradiumTemperature: 0.7,
  gradiumVoiceSimilarity: 2,
  voicevoxSpeedScale: 1.1,
  voicevoxPitchScale: 0.05,
  voicevoxIntonationScale: 1.2,
  voicevoxQueryParameters: { pauseLength: 0.3, outputSamplingRate: 44100 },
  minimaxVoiceSettings: { speed: 1.05, vol: 1.1, pitch: 2 },
  minimaxAudioSettings: { sampleRate: 44100, format: 'mp3' },
  aivisSpeechSpeedScale: 1.05,
  aivisCloudSpeakingRate: 1.1,
  aivisCloudVolume: 1.05,
});

Tip: the React example in packages/voice/examples/react-basic exposes the same controls with collapsible cards + sliders, making it easy to try values before applying them in code.

Engine parameter reference

  • OpenAI TTS

    • openAiModel
    • openAiSpeed
  • OpenAI-Compatible TTS

    • Endpoint: openAiCompatibleApiUrl
    • Optional voice: speaker
    • openAiCompatibleModel
    • openAiCompatibleSpeed
    • openAiCompatibleTimeoutMs
    • openAiCompatibleInstructions
    • openAiCompatibleResponseFormat
  • OpenRouter TTS (public preview)

    • Required model: openRouterModel (no default)
    • Voice: speaker, including the matching model suffix
    • Speech endpoint: openRouterApiUrl; PCM response wrapped as 24 kHz mono WAV
    • Voice-list options: openRouterModel, optional openRouterModelsApiUrl
  • xAI TTS

    • xaiLanguage
    • xaiCodec
    • xaiSampleRate
    • xaiBitRate
  • Unreal Speech

    • Endpoint: unrealSpeechApiUrl
    • Output: unrealSpeechBitrate, unrealSpeechCodec
    • Voice controls: unrealSpeechSpeed, unrealSpeechPitch, unrealSpeechTemperature
  • ElevenLabs

    • Endpoint: elevenLabsApiUrl
    • Identity/output: speaker, elevenLabsModel, elevenLabsOutputFormat, elevenLabsLanguageCode
    • Voice settings: elevenLabsVoiceSettings, elevenLabsStability, elevenLabsSimilarityBoost, elevenLabsStyle, elevenLabsUseSpeakerBoost, elevenLabsSpeed
    • Context/normalization: elevenLabsSeed, elevenLabsPreviousText, elevenLabsNextText, elevenLabsApplyTextNormalization, elevenLabsApplyLanguageTextNormalization, elevenLabsEnableLogging
  • Fish Audio

    • Endpoint: fishAudioApiUrl
    • Identity/output: speaker, fishAudioModel, fishAudioFormat, fishAudioSampleRate, fishAudioMp3Bitrate
    • Voice controls: fishAudioLatency, fishAudioSpeed
  • Deepgram Flux

    • Endpoint: deepgramApiUrl
    • Voice/model: speaker (Flux English ID), optional deepgramSpeed
    • Output: MP3
  • Cartesia

    • Endpoint: cartesiaApiUrl
    • Identity/output: speaker, cartesiaModel, cartesiaLanguage, cartesiaOutputContainer, cartesiaSampleRate, cartesiaMp3Bitrate
  • Inworld

    • Endpoint: inworldApiUrl
    • Identity/output: speaker, inworldModel, inworldAudioEncoding, inworldSampleRateHertz, inworldBitRate
    • Voice controls: inworldSpeakingRate, inworldLanguage, inworldDeliveryMode, inworldTemperature
  • Gradium

    • Endpoint: gradiumApiUrl
    • Identity/output: speaker, gradiumModel, gradiumOutputFormat
    • Voice controls: gradiumTemperature, gradiumVoiceSimilarity, gradiumPaddingBonus, gradiumRewriteRules
  • VOICEVOX

    • Endpoint: voicevoxApiUrl
    • Scalars: voicevoxSpeedScale, voicevoxPitchScale, voicevoxIntonationScale, voicevoxVolumeScale
    • Timing: voicevoxPrePhonemeLength, voicevoxPostPhonemeLength, voicevoxPauseLength, voicevoxPauseLengthScale
    • Output: voicevoxOutputSamplingRate, voicevoxOutputStereo
    • Flags: voicevoxEnableKatakanaEnglish, voicevoxEnableInterrogativeUpspeak
    • Version: voicevoxCoreVersion
    • Low-level overrides: voicevoxQueryParameters
  • AivisSpeech

    • Endpoint: aivisSpeechApiUrl
    • Scalars: aivisSpeechSpeedScale, aivisSpeechPitchScale, aivisSpeechIntonationScale, aivisSpeechTempoDynamicsScale, aivisSpeechVolumeScale
    • Timing: aivisSpeechPrePhonemeLength, aivisSpeechPostPhonemeLength, aivisSpeechPauseLength, aivisSpeechPauseLengthScale
    • Output: aivisSpeechOutputSamplingRate, aivisSpeechOutputStereo
    • Low-level overrides: aivisSpeechQueryParameters
  • Aivis Cloud

    • Identity: aivisCloudModelUuid, aivisCloudSpeakerUuid, aivisCloudStyleId, aivisCloudStyleName, aivisCloudUserDictionaryUuid
    • Behaviour: aivisCloudUseSSML, aivisCloudLanguage, aivisCloudSpeakingRate, aivisCloudEmotionalIntensity, aivisCloudTempoDynamics, aivisCloudPitch, aivisCloudVolume
    • Silence: aivisCloudLeadingSilence, aivisCloudTrailingSilence, aivisCloudLineBreakSilence
    • Output: aivisCloudOutputFormat, aivisCloudOutputBitrate, aivisCloudOutputSamplingRate, aivisCloudOutputChannels
    • Logging: aivisCloudEnableBillingLogs
  • VoicePeak

    • Endpoint: voicepeakApiUrl
    • Emotion: voicepeakEmotion (single tag or weighted map)
    • Scalars: voicepeakSpeed, voicepeakPitch
  • MiniMax

    • Identity: groupId, endpoint, minimaxModel, minimaxLanguageBoost
    • Voice overrides: minimaxVoiceSettings or individual minimaxSpeed, minimaxVolume, minimaxPitch
    • Audio overrides: minimaxAudioSettings or individual minimaxSampleRate, minimaxBitrate, minimaxAudioFormat, minimaxAudioChannel

Error Handling

try {
  await voiceService.speak({ text: 'Hello!' });
} catch (error) {
  if (error.message.includes('API key')) {
    console.error('Invalid API key');
  } else if (error.message.includes('network')) {
    console.error('Network error - check your connection');
  } else {
    console.error('TTS error:', error);
  }
}

Engine-Specific Features

VOICEVOX Features

  • Multiple character voices with unique personalities
  • Adjustable speech parameters (speed, pitch, intonation)
  • Local server support for privacy

OpenAI TTS Features

  • High-quality multilingual support
  • Multiple voice personalities
  • Optimized for conversational AI

xAI TTS Features

  • Cloud TTS endpoint with Bearer token authentication
  • Passes speaker through to voice_id as provided
  • Configurable codec, sample rate, and MP3 bitrate

Unreal Speech Features

  • Cloud TTS endpoint with Bearer token authentication
  • Passes speaker through to VoiceId as provided
  • Configurable bitrate, codec, speed, pitch, and temperature
  • Uses the v8 /stream API, which returns audio bytes directly

ElevenLabs Features

  • Cloud TTS endpoint with xi-api-key authentication
  • Passes speaker through to voice_id as provided
  • Configurable model, output format, language code, and voice settings
  • Supports optional text context, seed, text normalization, and logging flags

Fish Audio Features

  • Bearer-authenticated one-shot TTS with direct audio-byte responses
  • Configurable S2/S1 model, format, sample rate, MP3 bitrate, latency, and speed
  • Paginated model/voice-list lookup through the normalized helper

Cartesia Features

  • Bearer-authenticated synchronous /tts/bytes requests
  • Sonic 3.6 alias/snapshot as explicit options; Sonic 3.5 remains the default
  • Japanese and other documented language codes
  • Paginated voice-list lookup and WAV/MP3 output controls

Deepgram Flux Features

  • Token-authenticated one-shot /v2/speak requests with binary MP3 output
  • English-only Flux voices from the public v2 model catalog
  • Optional speech speed and custom endpoint; no emotion/style mapping

Inworld Features

  • Cloud TTS endpoint with Basic authentication
  • Passes speaker through to voiceId as provided
  • Configurable model, audio encoding, sample rate, bit rate, speaking rate, language, delivery mode, and temperature
  • Uses the non-streaming REST API and decodes the returned audioContent

Gradium Features

  • Cloud TTS endpoint with x-api-key authentication
  • Passes speaker through to voice_id as provided
  • Configurable output format and json_config controls for temperature, voice similarity, speed, and rewrite rules
  • Flagship voice presets provide readable names for browser speaker selectors
  • Dynamic voice-list lookups may require a backend proxy in browser apps if the provider blocks direct CORS access

MiniMax Features

  • 24 language support with automatic detection
  • HD quality audio output
  • Dual-region endpoints (global/china)
  • Advanced emotion synthesis
  • Uses documented system voice IDs instead of dynamic voice-list fetching

Gemini TTS Features

  • Gemini API-based high-quality voice synthesis
  • 30+ voice options (star/moon themed names)
  • Prompt-based style/tone control
  • Simple API key authentication with x-goog-api-key
  • Configurable Gemini API base URL
  • 24+ language support including Japanese

Web Speech API Features

  • Browser-native speechSynthesis playback with no API key
  • Browser-only runtime; no Node.js or server audio byte output
  • Supports voice selection by SpeechSynthesisVoice.name or voiceURI
  • Supports rate, pitch, volume, and language options where the browser voice honors them

Integration with AITuber OnAir Core

While this package can be used independently, it integrates seamlessly with @aituber-onair/core:

import { AITuberOnAirCore } from '@aituber-onair/core';

const core = new AITuberOnAirCore({
  apiKey: 'your-openai-key',
  voiceOptions: {
    engineType: 'voicevox',
    speaker: '1',
    voicevoxApiUrl: 'http://localhost:50021'
  }
});

// Voice synthesis is handled automatically
await core.processChat('Hello!');

API Reference

VoiceServiceOptions

type VoiceServiceOptions =
  | VoiceVoxVoiceServiceOptions
  | VoicePeakVoiceServiceOptions
  | OpenAiVoiceServiceOptions
  | XaiVoiceServiceOptions
  | UnrealSpeechVoiceServiceOptions
  | ElevenLabsVoiceServiceOptions
  | FishAudioVoiceServiceOptions
  | DeepgramVoiceServiceOptions
  | CartesiaVoiceServiceOptions
  | InworldVoiceServiceOptions
  | GradiumVoiceServiceOptions
  | GeminiTtsVoiceServiceOptions
  | OpenAiCompatibleVoiceServiceOptions
  | AivisSpeechVoiceServiceOptions
  | AivisCloudVoiceServiceOptions
  | MinimaxVoiceServiceOptions
  | PiperPlusVoiceServiceOptions
  | WebSpeechVoiceServiceOptions
  | NoneVoiceServiceOptions;

VoiceServiceOptions is a discriminated union keyed by engineType. Use updateOptions(...) for same-engine updates and switchEngine(...) for cross-engine changes. For backward compatibility, cross-engine fields in updateOptions(...) are still accepted.

Engine Capabilities

import {
  getAllVoiceEngineCapabilities,
  getVoiceEngineCapabilities,
} from '@aituber-onair/voice';

const gradium = getVoiceEngineCapabilities('gradium');
console.log(gradium.supportsVoiceList); // true

const allEngines = getAllVoiceEngineCapabilities();

Capabilities are static metadata only. They do not include API keys, endpoints, user configuration, or other sensitive values.

Voice Lists

import { getVoiceEngineVoiceList } from '@aituber-onair/voice';

const voices = await getVoiceEngineVoiceList('elevenLabs', {
  apiKey: process.env.ELEVENLABS_API_KEY,
});

// [{ id: '...', label: 'Rachel (premade)' }, ...]

getVoiceEngineVoiceList() returns normalized { id, label } items for engines that expose list APIs: VOICEVOX, AivisSpeech, Aivis Cloud, xAI, ElevenLabs, Fish Audio, Cartesia, Deepgram Flux, Inworld, Gradium, and Web Speech API. Pass local apiUrl for VOICEVOX-compatible servers, apiKey for cloud engines that require it, and language for Fish Audio, Cartesia, or Inworld filtering.

For browser apps, cloud provider voice-list endpoints must allow CORS. If a provider blocks direct browser requests, call getVoiceEngineVoiceList() from your backend or expose a small backend relay/proxy for the list endpoint.

Aivis Cloud voice-list support is package-level support for the model search endpoint. Browser requests to Aivis Cloud model/list endpoints can be blocked by CORS; call the helper from Node.js/backend or expose a backend proxy before wiring dynamic Aivis Cloud model selection into a production UI.

MiniMax is also excluded from this helper. Use the documented system voice IDs directly because the linked dynamic Get Voice API is currently unavailable.

VoiceService Methods

interface VoiceService {
  speak(screenplay: ChatScreenplay, options?: AudioPlayOptions): Promise<void>;
  speakText(text: string, options?: AudioPlayOptions): Promise<void>;
  isPlaying(): boolean;
  stop(): void;
  updateOptions(
    options: VoiceServiceOptionsUpdate | Partial<VoiceServiceOptions>
  ): void;
  switchEngine?(options: VoiceServiceOptions): void;
}

Screenplay Format

interface Screenplay {
  emotion?: string;
  text: string;
  speechText?: string;
}

Examples

React Integration

See the React example for a complete implementation:

import { useState } from 'react';
import { VoiceService } from '@aituber-onair/voice';

function VoiceDemo() {
  const [voiceService] = useState(
    () => new VoiceService({
      engineType: 'openai',
      speaker: 'alloy',
      apiKey: 'your-api-key'
    })
  );

  const handleSpeak = async (text: string) => {
    await voiceService.speak({ text });
  };

  return (
    <button onClick={() => handleSpeak('[happy] Hello!')}>
      Speak with emotion
    </button>
  );
}

Node.js Usage

The voice package now fully supports Node.js environments with automatic environment detection:

import { VoiceEngineAdapter } from '@aituber-onair/voice';

const voiceService = new VoiceEngineAdapter({
  engineType: 'openai',
  speaker: 'nova',
  apiKey: process.env.OPENAI_API_KEY
});

// Audio will be played using available Node.js audio libraries
await voiceService.speak({ text: 'Hello from Node.js!' });

Audio Playback in Node.js

For audio playback in Node.js, install one of these optional dependencies:

# Option 1: speaker (native bindings, better quality)
npm install speaker

# Option 2: play-sound (uses system audio player, easier to install)
npm install play-sound

If neither is installed, the package will still work but won't play audio. You can still use the onPlay callback to handle audio data:

const voiceService = new VoiceEngineAdapter({
  engineType: 'voicevox',
  speaker: '1',
  voicevoxApiUrl: 'http://localhost:50021',
  onPlay: async (audioBuffer) => {
    // Save to file or process audio data
    writeFileSync('output.wav', Buffer.from(audioBuffer));
  }
});

The package automatically detects the environment and uses the appropriate audio player:

  • Browser: Uses HTMLAudioElement
  • Node.js: Uses speaker or play-sound if available, otherwise silent

Testing

Run the test suite:

# Run all tests
npm test

# Run tests in watch mode
npm run test:watch

# Generate coverage report
npm run test:coverage

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

  1. Fork the repository
  2. Create your feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add some amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

License

This project is licensed under the MIT License - see the LICENSE file for details.