@aituber-onair/voice
v0.26.0
Published
Voice synthesis library for AITuber OnAir
Maintainers
Readme
AITuber OnAir Voice

@aituber-onair/voice is an independent voice synthesis library that supports multiple TTS (Text-to-Speech) engines. While originally developed for the AITuber OnAir project, it can be used standalone for any voice synthesis needs.
This project is published as open-source software and is available as an npm package under the MIT License.
Table of Contents
- Overview
- Installation
- Main Features
- Basic Usage
- Supported TTS Engines
- Emotion-Aware Speech
- Browser Compatibility
- Advanced Configuration
- Engine-Specific Features
- Integration with AITuber OnAir Core
- API Reference
- Examples
- Testing
- Contributing
Overview
@aituber-onair/voice is a comprehensive voice synthesis library that provides a unified interface for multiple TTS engines. It specializes in emotion-aware speech synthesis, making it ideal for creating expressive virtual characters, AI assistants, and interactive applications.
Key design principles:
- Engine Independence: Switch between TTS engines without changing your code
- Emotion Support: Built-in emotion detection and synthesis
- Browser Ready: Full support for web audio playback
- TypeScript First: Complete type safety and excellent IDE support
- Zero Dependencies: Minimal external dependencies for maximum compatibility
Installation
Install using npm:
npm install @aituber-onair/voiceOr using yarn:
yarn add @aituber-onair/voiceOr using pnpm:
pnpm install @aituber-onair/voiceMain Features
- Multiple TTS Engine Support
Compatible with VOICEVOX, VoicePeak, OpenAI TTS, OpenRouter TTS (preview opt-in), xAI TTS, Unreal Speech, ElevenLabs, Fish Audio, Cartesia, Deepgram Flux, Inworld, Gradium, Gemini TTS, MiniMax, AivisSpeech, Aivis Cloud, Web Speech API, and more - Unified Interface
Single API for all supported TTS engines - Emotion-Aware Synthesis
Automatically detects and applies emotions from text tags like[happy],[sad], etc. - Screenplay Conversion
Transforms text with emotion tags into structured screenplay format - Browser Audio Support
Direct playback in web browsers using HTMLAudioElement - Custom Endpoints
Support for self-hosted TTS servers - Language Detection
Automatic language recognition for multi-language engines - Flexible Configuration
Runtime engine switching and parameter updates
Basic Usage
Simple Text-to-Speech
import { VoiceService, VoiceServiceOptions } from '@aituber-onair/voice';
// Configure the voice service
const options: VoiceServiceOptions = {
engineType: 'voicevox',
speaker: '1',
// Optional: specify custom endpoint
voicevoxApiUrl: 'http://localhost:50021'
};
// Create voice service instance
const voiceService = new VoiceService(options);
// Speak text
await voiceService.speak({ text: 'Hello, world!' });Using VoiceEngineAdapter (Recommended)
import { VoiceEngineAdapter, VoiceServiceOptions } from '@aituber-onair/voice';
const options: VoiceServiceOptions = {
engineType: 'openai',
speaker: 'alloy',
apiKey: 'your-openai-api-key',
onPlay: async (audioBuffer) => {
// Custom audio playback handler
console.log('Playing audio...');
}
};
const voiceAdapter = new VoiceEngineAdapter(options);
// Speak with emotion
await voiceAdapter.speak({
text: '[happy] I am so excited to talk with you!'
});Supported TTS Engines
VOICEVOX
High-quality Japanese speech synthesis engine with multiple character voices.
const voiceService = new VoiceService({
engineType: 'voicevox',
speaker: '1', // Character ID
voicevoxApiUrl: 'http://localhost:50021' // Optional custom endpoint
});VoicePeak
Professional speech synthesis with rich emotional expression.
const voiceService = new VoiceService({
engineType: 'voicepeak',
speaker: 'f1',
voicepeakApiUrl: 'http://localhost:20202',
voicepeakEmotion: 'happy',
voicepeakSpeed: 140,
voicepeakPitch: 20
});Single-tag voicepeakEmotion remains backward compatible with existing
VoicePeak setups. Weighted emotion maps require vpeakserver >= v0.2.0.
const weightedVoiceService = new VoiceService({
engineType: 'voicepeak',
speaker: 'f1',
voicepeakApiUrl: 'http://localhost:20202',
voicepeakEmotion: { happy: 40, fun: 60 },
});neutralis ignored when sending weighted emotions.- Weight
0is ignored. {}means "do not send emotion" and does not fall back toTalk.style.undefinedmeans no override, soTalk.stylestill maps to a single tag.
OpenAI TTS
OpenAI's text-to-speech API with multiple voice options.
const voiceService = new VoiceService({
engineType: 'openai',
speaker: 'alloy',
apiKey: 'your-openai-api-key'
});OpenRouter TTS (public-preview opt-in)
Preview models require an explicit choice and are not defaults. This dedicated
engine supports microsoft/mai-voice-2.1 and
microsoft/mai-voice-2.1-flash through OpenRouter's
POST /api/v1/audio/speech.
The OpenRouter TTS guide
documents both exact model IDs, full model-specific voice IDs, and 24 kHz mono PCM output.
Microsoft classifies both as public preview, without an SLA and not recommended
for production.
The model catalogs checked on October 1, 2026 contained 97 voices across 28 locales
for each model, with no Japanese voices. Voice availability may change; fetch the
selected model's catalog instead of assuming a fixed list.
import {
VoiceEngineAdapter,
getVoiceEngineVoiceList,
type OpenRouterTtsModel,
} from '@aituber-onair/voice';
// Deliberate preview opt-in; no model is chosen when this option is omitted.
const model: OpenRouterTtsModel = 'microsoft/mai-voice-2.1';
const voices = await getVoiceEngineVoiceList('openRouter', {
openRouterModel: model,
// Optional for the public models catalog; synthesis requires your key.
apiKey: 'your-openrouter-api-key',
});
const speaker = voices.find((voice) =>
voice.id === 'en-US-Harper:MAI-Voice-2.1',
)?.id;
if (!speaker) throw new Error('Select an available voice for this model');
const voiceService = new VoiceEngineAdapter({
engineType: 'openRouter',
openRouterModel: model,
speaker,
apiKey: 'your-openrouter-api-key',
// Optional full speech endpoint (not a base URL):
// openRouterApiUrl: 'https://openrouter.ai/api/v1/audio/speech',
});
await voiceService.speakText('Hello! Welcome to the show.');The engine sends Bearer authentication and JSON
{ model, input, voice, response_format: 'pcm' }. The speech API defines
audio/pcm as 16-bit little-endian; the MAI route specifies 24 kHz mono. The
engine wraps these bytes in a WAV header using the existing PCM helper, so the
browser/Node audio player and onPlay callback receive a complete WAV buffer,
not compressed MP3 or headerless PCM. Node playback still needs an optional
playback dependency as described below. Non-PCM, empty, or incomplete 16-bit
sample responses are rejected before playback. This is
one-shot synthesis, not realtime audio streaming. Emotion tags do not add voice
style, speed, or cloning parameters; these controls are intentionally omitted.
The voice locale determines the synthesis language.
getVoiceEngineVoiceList('openRouter', { openRouterModel }) reads
GET https://openrouter.ai/api/v1/models?output_modalities=speech and returns
only that exact model's supported_voices. Use openRouterModelsApiUrl (or
voiceListApiUrl) on the lookup options for a separate models endpoint; a custom
openRouterApiUrl does not redirect catalog requests. Relative endpoint URLs work
in browsers; Node requires absolute HTTP(S) URLs. Configure only trusted endpoints
because requests with an API key send it to the configured destination.
When changing models at runtime, choose a voice from the new model's list and
update openRouterModel and speaker together. Updating only the model clears
the prior speaker. Full IDs end with :MAI-Voice-2.1 or
:MAI-Voice-2.1-Flash; a stale suffix is rejected, never rewritten automatically.
The React example starts with no preview model or voice selected and provides
model-scoped voice selection, API key and separate endpoint controls.
The example uses the direct OpenRouter route. Browser-side keys are visible to the page, so deploy shared apps with your own credential-protecting backend rather than embedding a shared secret in client code. Direct access depends on the provider's CORS policy and your deployment; custom endpoints must permit your application's origin or run as same-origin backend routes.
xAI TTS
xAI's cloud TTS API with selectable voice IDs, language control, and output format tuning.
const voiceService = new VoiceService({
engineType: 'xai',
speaker: 'eve',
apiKey: 'your-xai-api-key',
xaiLanguage: 'ja',
xaiCodec: 'mp3',
xaiSampleRate: 24000,
xaiBitRate: 128000,
});Unreal Speech
Unreal Speech v8 cloud TTS via the /stream endpoint. It returns audio bytes
directly, so it works with VoiceEngineAdapter playback without an extra
download step.
const voiceService = new VoiceService({
engineType: 'unrealSpeech',
speaker: 'af_bella',
apiKey: 'your-unreal-speech-api-key',
unrealSpeechBitrate: '192k',
unrealSpeechSpeed: 0,
unrealSpeechPitch: 1,
unrealSpeechCodec: 'libmp3lame',
unrealSpeechTemperature: 0.25,
});Use unrealSpeechApiUrl to override the default
https://api.v8.unrealspeech.com/stream endpoint.
ElevenLabs
ElevenLabs Text to Speech API support using direct fetch calls. No SDK is
required.
const voiceService = new VoiceService({
engineType: 'elevenLabs',
speaker: 'JBFqnCBsd6RMkjVDRZzb',
apiKey: 'your-elevenlabs-api-key',
elevenLabsModel: 'eleven_flash_v2_5',
elevenLabsOutputFormat: 'mp3_44100_128',
elevenLabsStability: 0.5,
elevenLabsSimilarityBoost: 0.75,
elevenLabsUseSpeakerBoost: true,
});Use elevenLabsApiUrl to override the default
https://api.elevenlabs.io/v1/text-to-speech endpoint. The speaker value is
sent as the ElevenLabs voice_id.
The curated model choices include eleven_v4 for highest-quality speech,
eleven_v3 for previous-generation expressiveness,
eleven_multilingual_v2 for high-quality multilingual output, and
eleven_flash_v2_5 (the default) for low latency. The deprecated
eleven_turbo_v2_5 remains accepted as a custom string for backward
compatibility, but is no longer presented as a recommended model.
Model-specific settings are filtered when building requests. For eleven_v4,
only Stability and Similarity are sent; configured Style, Speed, and Speaker
Boost values are retained for other models but omitted for v4. SSML is not
supported by v4; use plain text or its documented audio tags. This engine uses
one-shot Create speech
and waits for the complete audio response. The official TTS guide
confirms v4 support for that endpoint. Models outside the verified integration
scope are not offered as supported choices: eleven_v4_turbo is documented
for the Text to Dialogue WebSocket,
which this package does not implement. Turbo has not been verified through
this package's HTTP speech path; this is not a claim that the provider cannot
support it over HTTP.
Fish Audio
Fish Audio one-shot TTS uses POST /v1/tts and returns audio bytes directly.
const voiceService = new VoiceService({
engineType: 'fishAudio',
speaker: 'your-reference-id',
apiKey: 'your-fish-audio-api-key',
fishAudioModel: 's2-pro',
fishAudioFormat: 'mp3',
fishAudioLatency: 'normal',
});speaker is sent as reference_id. Use getVoiceEngineVoiceList('fishAudio',
{ apiKey }) to list usable models/voices. s2-pro is the stable default for
this integration; s2.1-pro-free must be selected explicitly and should not be
treated as an SLA-backed production tier.
Voice-list lookups return at most 100 entries by default so they do not traverse
the full public model catalog. Pass limit and optionally pageSize to request
a larger bounded result set.
The official POST /v1/tts endpoint does not currently complete browser CORS
preflight requests. Call it directly from Node.js or a server, or set
fishAudioApiUrl to a same-origin backend route in browser applications. Keep
the Fish Audio API key on the server in production. The React example includes
a Vite-only development/preview proxy for this purpose.
Cartesia
Cartesia synchronous TTS uses POST /tts/bytes and returns audio bytes
directly.
const voiceService = new VoiceService({
engineType: 'cartesia',
speaker: 'your-cartesia-voice-id',
apiKey: 'your-cartesia-api-key',
cartesiaModel: 'sonic-3.5',
cartesiaLanguage: 'ja',
cartesiaOutputContainer: 'wav',
cartesiaSampleRate: 44100,
});Use getVoiceEngineVoiceList('cartesia', { apiKey, language: 'ja' }) to list
voices. The engine retains the Cartesia-Version: 2026-03-01 header and
voice: { id } request shape. Set cartesiaModel: 'sonic-3.6' explicitly for
the newer model, or 'sonic-3.6-2026-08-27' to pin its stable snapshot. The
existing sonic-3.5 default is unchanged. Cartesia documents Sonic 3.6 on
POST /tts/bytes and
backward compatibility for pinned integrations.
WAV/MP3 output and existing voice IDs continue through the same path.
Deepgram Flux
Deepgram Flux uses the one-shot POST /v2/speak endpoint
with Authorization: Token ... and returns MP3 audio bytes. This integration
supports English Flux voices only. It does not implement Aura (/v1/speak),
WebSocket streaming, interruption handling, callbacks, or beta expressivity.
const voiceService = new VoiceService({
engineType: 'deepgram',
speaker: 'flux-haley-en',
apiKey: process.env.DEEPGRAM_API_KEY,
deepgramSpeed: 1.0,
});The required full voice ID in speaker is sent as the model query parameter.
The example defaults to flux-haley-en. deepgramSpeed is optional (0.5–1.5,
in 0.05 steps).
Use getVoiceEngineVoiceList('deepgram') to fetch the public v2 model catalog
(GET /v2/models, no API key) and keep its English Flux entries. The v1
catalog (/v1/models) lists Aura voices only. The React example always offers the documented Haley preset independently of
the catalog. Its optional catalog refresh preserves the selected voice, and
empty results or errors also preserve the existing list. See the
voice catalog.
Call from Node.js/backend, or use deepgramApiUrl for a same-origin speech
proxy and voiceListApiUrl for the catalog proxy. Keep the API key on the server
in production. Direct browser CORS behavior has not been live-verified; the
React example supplies Vite development/preview proxies under /api/deepgram.
Production deployments must provide equivalent backend routes.
Inworld
Inworld TTS non-streaming speech synthesis using direct fetch calls. This
engine uses the REST endpoint only; WebSocket and HTTP streaming are not
implemented.
const voiceService = new VoiceService({
engineType: 'inworld',
speaker: 'Ashley',
apiKey: process.env.INWORLD_API_KEY,
inworldModel: 'inworld-tts-2',
inworldAudioEncoding: 'MP3',
inworldSampleRateHertz: 48000,
});Use inworldApiUrl to override the default
https://api.inworld.ai/tts/v1/voice endpoint. The apiKey value should be
the Inworld Basic Base64 authorization value. Do not expose Basic credentials
in browser-side code; use a backend proxy or Inworld JWT authentication for
browser apps.
The current TTS-2 models are inworld-tts-2 (the default, for quality)
and inworld-tts-2-flash (an explicit option for lower latency and cost).
Set inworldModel: 'inworld-tts-2-flash' to use Flash, or switch at runtime
with voiceService.updateOptions({ inworldModel: 'inworld-tts-2-flash' }).
Both use the same REST endpoint and voice-list configuration. This engine
waits for the complete audio response, so provider streaming latency figures
do not describe its playback latency.
inworldDeliveryMode is supported only by TTS-2 and is omitted for Flash.
Earlier model IDs remain accepted as strings for compatibility, but deprecated
TTS 1.5 models are no longer offered in the React model selector.
See the model catalog and
speech API reference.
Gradium
Gradium one-shot REST TTS support using direct fetch calls. This engine uses
the raw-audio response mode (only_audio: true) and does not add the Gradium
SDK.
Production remains the default. Experimental models require explicit opt-in:
set gradiumModel: 'gradium-tts-beta' to try the public beta, or use
gradiumModel: 'default' (or omit it) for the production model. The beta is a
supported explicit option, not the recommended default. The React example
exposes both choices. updateOptions({ gradiumModel: 'default' }) switches
an existing adapter back to production; undefined clears the selection.
gradiumModel is sent as top-level model_name in the REST JSON body,
separately from json_config. This follows the official
REST guide and
model selection guide.
The existing one-shot raw-audio response and output format handling are unchanged;
this does not add streaming synthesis or new emotion controls.
const voiceService = new VoiceService({
engineType: 'gradium',
speaker: 'YTpq7expH9539ERJ',
apiKey: process.env.GRADIUM_API_KEY,
gradiumOutputFormat: 'wav',
gradiumTemperature: 0.7,
gradiumVoiceSimilarity: 2,
gradiumPaddingBonus: 0,
gradiumRewriteRules: 'en',
});Use gradiumApiUrl to override the default
https://api.gradium.ai/api/post/speech/tts endpoint. The speaker value is
sent as Gradium voice_id. The React example uses Gradium flagship voice
presets as a fallback and can fetch the Gradium voice list through
getVoiceEngineVoiceList() when an API key is provided. Browser-side voice
list requests may fail if the Gradium API does not allow direct CORS access;
use a backend proxy for production browser UIs that need dynamic Gradium voice
selection.
OpenAI-Compatible TTS
OpenAI-compatible speech endpoints for self-hosted servers such as Kokoro FastAPI.
const voiceService = new VoiceService({
engineType: 'openaiCompatible',
openAiCompatibleApiUrl: 'http://localhost:8880/v1/audio/speech',
openAiCompatibleModel: 'your-model-id'
});speaker is optional for compatible endpoints. When omitted, the request body
does not include a voice field.
openAiCompatibleModel should be set explicitly to a model name accepted by
your endpoint.
openAiCompatibleApiUrl is used as-is, so pass the full /audio/speech URL.
openAiCompatibleInstructions and openAiCompatibleResponseFormat are
optional and sent as instructions / response_format only when non-empty.
Servers that support instructions typically use it as a voice style prompt.
Failed requests throw a VoiceEngineError with kind: 'api', the HTTP
status in statusCode, and the response body in the message.
Endpoint setup helpers
These helpers make settings UIs easier to build. They accept an origin, an API
base URL, or a full /audio/speech URL.
import {
resolveOpenAICompatibleSpeechEndpoint,
listOpenAICompatibleSpeechModels,
listOpenAICompatibleSpeechVoices,
getOpenAICompatibleSpeechServerInfo,
testOpenAICompatibleSpeech,
} from '@aituber-onair/voice';
const endpoint = 'http://localhost:8880/v1';
const { speechUrl } = resolveOpenAICompatibleSpeechEndpoint(endpoint);
// -> http://localhost:8880/v1/audio/speech
const models = await listOpenAICompatibleSpeechModels({ endpoint });
const voices = await listOpenAICompatibleSpeechVoices({ endpoint }); // or null
const info = await getOpenAICompatibleSpeechServerInfo({ endpoint }); // or null
const result = await testOpenAICompatibleSpeech({
endpoint,
model: models[0],
voice: voices?.defaultVoice,
});
if (result.ok) {
// result.audio is an ArrayBuffer you can play back
} else {
console.error(result.error.code, result.error.status, result.error.detail);
}listOpenAICompatibleSpeechModelsuses the standardGET /models.- OpenAI's API has no voice listing endpoint, so
listOpenAICompatibleSpeechVoicesis best-effort. It triesGET /audio/voices(for example Kokoro-FastAPI) and thenGET /voices, and returnsnullwhen neither answers with a recognizable list. getOpenAICompatibleSpeechServerInfois also best-effort. It reads the server root and returns{ engine, model?, defaultVoice? }only when the root answers with JSON that has anenginefield. Otherwise it returnsnull.testOpenAICompatibleSpeechsynthesizes one short sentence and returns a result object instead of throwing. Errors carry acode(invalid-url,network,aborted,timeout,http,invalid-response), plusstatusanddetailfor HTTP errors.- Browser requests also need CORS permission from the server. A
networkerror in the browser usually means CORS blocked the request or the server is not running.
MiniMax
Multi-language TTS supporting 24 languages with HD quality.
const voiceService = new VoiceService({
engineType: 'minimax',
speaker: 'Japanese_IntellectualSenior',
apiKey: 'your-minimax-api-key',
minimaxModel: 'speech-2.8-turbo',
groupId: 'legacy-group-id', // Optional for older accounts
endpoint: 'global' // or 'china'
});Note: Current MiniMax T2A v2 endpoints require Bearer authentication but
do not require GroupId. The optional groupId field is kept for older
accounts or endpoints that still expect the query parameter. Speech 2.8 Turbo
is the default; Speech 2.8 HD is available when output quality is preferred.
Use MiniMax system voice IDs for speaker, such as
Japanese_IntellectualSenior. MiniMax documents these IDs in its
System Voice ID List.
The linked dynamic Get Voice API is not currently available, so
getVoiceEngineVoiceList() does not expose MiniMax voice-list fetching.
AivisSpeech
AI-powered speech synthesis with natural voice quality.
const voiceService = new VoiceService({
engineType: 'aivisSpeech',
speaker: '888753760',
aivisSpeechApiUrl: 'http://localhost:10101'
});Aivis Cloud
High-quality cloud-based TTS service with advanced SSML support and streaming capabilities.
const voiceService = new VoiceService({
engineType: 'aivisCloud',
speaker: 'unused', // Not used when model UUID is specified
apiKey: 'your-aivis-cloud-api-key',
aivisCloudModelUuid: 'a59cb814-0083-4369-8542-f51a29e72af7', // Required
// Optional advanced settings
aivisCloudSpeakerUuid: 'speaker-uuid', // For multi-speaker models
aivisCloudStyleId: 0, // Or use aivisCloudStyleName: 'ノーマル'
aivisCloudUseSSML: true, // Enable SSML tags
aivisCloudSpeakingRate: 1.0, // 0.5-2.0
aivisCloudEmotionalIntensity: 1.0, // 0.0-2.0
aivisCloudOutputFormat: 'mp3', // wav, flac, mp3, aac, opus
aivisCloudOutputSamplingRate: 44100, // Hz
});Key Features:
- SSML Support: Rich markup for prosody, breaks, aliases, and emotions
- Streaming Audio: Real-time audio generation and delivery
- Multiple Formats: WAV, FLAC, MP3, AAC, Opus output
- Emotion Control: Fine-grained emotional intensity settings
- High Quality: Professional-grade voice synthesis
getVoiceEngineVoiceList('aivisCloud') uses the Aivis Cloud model search API
(GET https://api.aivis-project.com/v1/aivm-models/search) and returns model
UUID choices that can be passed as speaker or aivisCloudModelUuid. Direct
browser requests to model/list endpoints can fail CORS checks, so browser apps
should call this helper from a Node.js backend or relay/proxy when they need
dynamic model, speaker, or style selection. The React example intentionally
keeps manual model UUID input instead of calling the model search endpoint from
the browser.
Gemini TTS
Gemini API text-to-speech supports gemini-3.8-flash-lite-tts for fast,
cost-efficient speech and gemini-3.8-flash-tts for more expressive speech.
Earlier preview models remain available. Both 3.8 models use the Interactions
API and return WAV audio; preview models continue to use generateContent.
const voiceService = new VoiceService({
engineType: 'geminiTts',
speaker: 'Zephyr',
apiKey: 'your-google-api-key',
geminiTtsModel: 'gemini-3.8-flash-lite-tts',
geminiTtsPrompt: 'cheerful and friendly', // Optional speech_metadata.style for 3.8
geminiTtsApiUrl:
'https://generativelanguage.googleapis.com/v1beta', // Optional Gemini API base URL
});Note: Use a standard Google API key. apiKey is sent as
x-goog-api-key to the Gemini API. speaker accepts a prebuilt voice name
such as Zephyr or Kore; for 3.8 it also accepts a voice_... ID created in
Google AI Studio. Gemini 3.8 detects the input language automatically, so
geminiTtsLanguageCode applies only to the preview models. For 3.8,
geminiTtsPrompt is sent as a style annotation instead of being spoken as
part of the transcript. This engine returns complete audio rather than
streaming chunks. Voice creation and replication are handled in Google AI
Studio or the Gemini Voices API.
Web Speech API
Browser-native speech synthesis through window.speechSynthesis. This engine
does not return audio bytes; the browser plays speech directly.
const voiceService = new VoiceEngineAdapter({
engineType: 'webSpeech',
speaker: '', // Optional: SpeechSynthesisVoice name or voiceURI
webSpeechLanguage: 'ja-JP',
webSpeechRate: 1.1,
webSpeechPitch: 1.0,
webSpeechVolume: 1.0,
});
await voiceService.speak({ text: 'こんにちは' });Use getVoiceEngineVoiceList('webSpeech') in a browser to list available
SpeechSynthesisVoice entries. Some browsers populate voices asynchronously,
so the helper waits briefly for voiceschanged. Because no ArrayBuffer is
available, onPlay(audioBuffer) is skipped for this engine; onComplete still
runs when the utterance ends. Runtime support is browser-only.
None (Silent Mode)
No audio output - useful for testing or text-only scenarios.
const voiceService = new VoiceService({
engineType: 'none'
});Emotion-Aware Speech
The library supports emotion tags in text for more expressive speech:
// Emotion tags are automatically detected and processed
await voiceService.speak({
text: '[happy] Great to see you today!'
});
await voiceService.speak({
text: '[sad] I will miss you...'
});
await voiceService.speak({
text: '[angry] This is unacceptable!'
});
// Supported emotions vary by engine
// Common emotions: happy, sad, angry, surprised, neutralThe emotion system works by:
- Extracting emotion tags from the text
- Converting text to screenplay format with emotion metadata
- Passing emotion information to engines that support it
- Falling back gracefully for engines without emotion support
Browser Compatibility
The library includes built-in browser audio playback support:
// Option 1: Default browser playback
const voiceService = new VoiceService({
engineType: 'openai',
speaker: 'alloy',
apiKey: 'your-api-key'
// Audio will play automatically in the browser
});
// Option 2: Custom audio handling
const voiceService = new VoiceService({
engineType: 'voicevox',
speaker: '1',
onPlay: async (audioBuffer: ArrayBuffer) => {
// Custom audio playback logic
const audioContext = new AudioContext();
const audioBufferSource = audioContext.createBufferSource();
// ... handle audio playback
}
});
// Option 3: Specify HTML audio element
const voiceService = new VoiceService({
engineType: 'voicevox',
speaker: '1',
voicevoxApiUrl: 'http://localhost:50021',
audioElementId: 'my-audio-player' // ID of <audio> element
});Advanced Configuration
Dynamic Engine Switching
const voiceAdapter = new VoiceEngineAdapter({
engineType: 'voicevox',
speaker: '1'
});
// Update options within the same engine
voiceAdapter.updateOptions({
speaker: '3',
voicevoxSpeedScale: 1.1,
});
// Switch to a different engine at runtime
voiceAdapter.switchEngine({
engineType: 'openai',
speaker: 'nova',
apiKey: 'your-openai-api-key'
});
// Backward compatibility:
// updateOptions with engineType is still accepted.
voiceAdapter.updateOptions({
engineType: 'openai',
speaker: 'nova',
apiKey: 'your-openai-api-key'
});Custom Endpoints
// For self-hosted or custom TTS servers
const voiceService = new VoiceService({
engineType: 'voicevox',
speaker: '1',
voicevoxApiUrl: 'https://my-custom-voicevox-server.com'
});Engine Parameter Overrides
VoiceServiceOptions (see API Reference) now covers a consistent set of overrides for each engine. Below is a field-by-field summary to help you discover the right property without scanning the entire interface.
const voiceService = new VoiceService({
engineType: 'voicevox',
speaker: '1',
openAiSpeed: 1.15,
openAiCompatibleModel: 'your-model-id',
openAiCompatibleSpeed: 1.1,
openAiCompatibleTimeoutMs: 300_000, // Optional; default: 30_000 ms; 0 disables the timeout
unrealSpeechBitrate: '192k',
unrealSpeechSpeed: 0,
unrealSpeechPitch: 1,
elevenLabsModel: 'eleven_flash_v2_5',
elevenLabsStability: 0.5,
elevenLabsSimilarityBoost: 0.75,
inworldModel: 'inworld-tts-2',
inworldAudioEncoding: 'MP3',
inworldSampleRateHertz: 48000,
gradiumOutputFormat: 'wav',
gradiumTemperature: 0.7,
gradiumVoiceSimilarity: 2,
voicevoxSpeedScale: 1.1,
voicevoxPitchScale: 0.05,
voicevoxIntonationScale: 1.2,
voicevoxQueryParameters: { pauseLength: 0.3, outputSamplingRate: 44100 },
minimaxVoiceSettings: { speed: 1.05, vol: 1.1, pitch: 2 },
minimaxAudioSettings: { sampleRate: 44100, format: 'mp3' },
aivisSpeechSpeedScale: 1.05,
aivisCloudSpeakingRate: 1.1,
aivisCloudVolume: 1.05,
});Tip: the React example in
packages/voice/examples/react-basicexposes the same controls with collapsible cards + sliders, making it easy to try values before applying them in code.
Engine parameter reference
OpenAI TTS
openAiModelopenAiSpeed
OpenAI-Compatible TTS
- Endpoint:
openAiCompatibleApiUrl - Optional voice:
speaker openAiCompatibleModelopenAiCompatibleSpeedopenAiCompatibleTimeoutMsopenAiCompatibleInstructionsopenAiCompatibleResponseFormat
- Endpoint:
OpenRouter TTS (public preview)
- Required model:
openRouterModel(no default) - Voice:
speaker, including the matching model suffix - Speech endpoint:
openRouterApiUrl; PCM response wrapped as 24 kHz mono WAV - Voice-list options:
openRouterModel, optionalopenRouterModelsApiUrl
- Required model:
xAI TTS
xaiLanguagexaiCodecxaiSampleRatexaiBitRate
Unreal Speech
- Endpoint:
unrealSpeechApiUrl - Output:
unrealSpeechBitrate,unrealSpeechCodec - Voice controls:
unrealSpeechSpeed,unrealSpeechPitch,unrealSpeechTemperature
- Endpoint:
ElevenLabs
- Endpoint:
elevenLabsApiUrl - Identity/output:
speaker,elevenLabsModel,elevenLabsOutputFormat,elevenLabsLanguageCode - Voice settings:
elevenLabsVoiceSettings,elevenLabsStability,elevenLabsSimilarityBoost,elevenLabsStyle,elevenLabsUseSpeakerBoost,elevenLabsSpeed - Context/normalization:
elevenLabsSeed,elevenLabsPreviousText,elevenLabsNextText,elevenLabsApplyTextNormalization,elevenLabsApplyLanguageTextNormalization,elevenLabsEnableLogging
- Endpoint:
Fish Audio
- Endpoint:
fishAudioApiUrl - Identity/output:
speaker,fishAudioModel,fishAudioFormat,fishAudioSampleRate,fishAudioMp3Bitrate - Voice controls:
fishAudioLatency,fishAudioSpeed
- Endpoint:
Deepgram Flux
- Endpoint:
deepgramApiUrl - Voice/model:
speaker(Flux English ID), optionaldeepgramSpeed - Output: MP3
- Endpoint:
Cartesia
- Endpoint:
cartesiaApiUrl - Identity/output:
speaker,cartesiaModel,cartesiaLanguage,cartesiaOutputContainer,cartesiaSampleRate,cartesiaMp3Bitrate
- Endpoint:
Inworld
- Endpoint:
inworldApiUrl - Identity/output:
speaker,inworldModel,inworldAudioEncoding,inworldSampleRateHertz,inworldBitRate - Voice controls:
inworldSpeakingRate,inworldLanguage,inworldDeliveryMode,inworldTemperature
- Endpoint:
Gradium
- Endpoint:
gradiumApiUrl - Identity/output:
speaker,gradiumModel,gradiumOutputFormat - Voice controls:
gradiumTemperature,gradiumVoiceSimilarity,gradiumPaddingBonus,gradiumRewriteRules
- Endpoint:
VOICEVOX
- Endpoint:
voicevoxApiUrl - Scalars:
voicevoxSpeedScale,voicevoxPitchScale,voicevoxIntonationScale,voicevoxVolumeScale - Timing:
voicevoxPrePhonemeLength,voicevoxPostPhonemeLength,voicevoxPauseLength,voicevoxPauseLengthScale - Output:
voicevoxOutputSamplingRate,voicevoxOutputStereo - Flags:
voicevoxEnableKatakanaEnglish,voicevoxEnableInterrogativeUpspeak - Version:
voicevoxCoreVersion - Low-level overrides:
voicevoxQueryParameters
- Endpoint:
AivisSpeech
- Endpoint:
aivisSpeechApiUrl - Scalars:
aivisSpeechSpeedScale,aivisSpeechPitchScale,aivisSpeechIntonationScale,aivisSpeechTempoDynamicsScale,aivisSpeechVolumeScale - Timing:
aivisSpeechPrePhonemeLength,aivisSpeechPostPhonemeLength,aivisSpeechPauseLength,aivisSpeechPauseLengthScale - Output:
aivisSpeechOutputSamplingRate,aivisSpeechOutputStereo - Low-level overrides:
aivisSpeechQueryParameters
- Endpoint:
Aivis Cloud
- Identity:
aivisCloudModelUuid,aivisCloudSpeakerUuid,aivisCloudStyleId,aivisCloudStyleName,aivisCloudUserDictionaryUuid - Behaviour:
aivisCloudUseSSML,aivisCloudLanguage,aivisCloudSpeakingRate,aivisCloudEmotionalIntensity,aivisCloudTempoDynamics,aivisCloudPitch,aivisCloudVolume - Silence:
aivisCloudLeadingSilence,aivisCloudTrailingSilence,aivisCloudLineBreakSilence - Output:
aivisCloudOutputFormat,aivisCloudOutputBitrate,aivisCloudOutputSamplingRate,aivisCloudOutputChannels - Logging:
aivisCloudEnableBillingLogs
- Identity:
VoicePeak
- Endpoint:
voicepeakApiUrl - Emotion:
voicepeakEmotion(single tag or weighted map) - Scalars:
voicepeakSpeed,voicepeakPitch
- Endpoint:
MiniMax
- Identity:
groupId,endpoint,minimaxModel,minimaxLanguageBoost - Voice overrides:
minimaxVoiceSettingsor individualminimaxSpeed,minimaxVolume,minimaxPitch - Audio overrides:
minimaxAudioSettingsor individualminimaxSampleRate,minimaxBitrate,minimaxAudioFormat,minimaxAudioChannel
- Identity:
Error Handling
try {
await voiceService.speak({ text: 'Hello!' });
} catch (error) {
if (error.message.includes('API key')) {
console.error('Invalid API key');
} else if (error.message.includes('network')) {
console.error('Network error - check your connection');
} else {
console.error('TTS error:', error);
}
}Engine-Specific Features
VOICEVOX Features
- Multiple character voices with unique personalities
- Adjustable speech parameters (speed, pitch, intonation)
- Local server support for privacy
OpenAI TTS Features
- High-quality multilingual support
- Multiple voice personalities
- Optimized for conversational AI
xAI TTS Features
- Cloud TTS endpoint with Bearer token authentication
- Passes
speakerthrough tovoice_idas provided - Configurable codec, sample rate, and MP3 bitrate
Unreal Speech Features
- Cloud TTS endpoint with Bearer token authentication
- Passes
speakerthrough toVoiceIdas provided - Configurable bitrate, codec, speed, pitch, and temperature
- Uses the v8
/streamAPI, which returns audio bytes directly
ElevenLabs Features
- Cloud TTS endpoint with
xi-api-keyauthentication - Passes
speakerthrough tovoice_idas provided - Configurable model, output format, language code, and voice settings
- Supports optional text context, seed, text normalization, and logging flags
Fish Audio Features
- Bearer-authenticated one-shot TTS with direct audio-byte responses
- Configurable S2/S1 model, format, sample rate, MP3 bitrate, latency, and speed
- Paginated model/voice-list lookup through the normalized helper
Cartesia Features
- Bearer-authenticated synchronous
/tts/bytesrequests - Sonic 3.6 alias/snapshot as explicit options; Sonic 3.5 remains the default
- Japanese and other documented language codes
- Paginated voice-list lookup and WAV/MP3 output controls
Deepgram Flux Features
- Token-authenticated one-shot
/v2/speakrequests with binary MP3 output - English-only Flux voices from the public v2 model catalog
- Optional speech speed and custom endpoint; no emotion/style mapping
Inworld Features
- Cloud TTS endpoint with Basic authentication
- Passes
speakerthrough tovoiceIdas provided - Configurable model, audio encoding, sample rate, bit rate, speaking rate, language, delivery mode, and temperature
- Uses the non-streaming REST API and decodes the returned
audioContent
Gradium Features
- Cloud TTS endpoint with
x-api-keyauthentication - Passes
speakerthrough tovoice_idas provided - Configurable output format and
json_configcontrols for temperature, voice similarity, speed, and rewrite rules - Flagship voice presets provide readable names for browser speaker selectors
- Dynamic voice-list lookups may require a backend proxy in browser apps if the provider blocks direct CORS access
MiniMax Features
- 24 language support with automatic detection
- HD quality audio output
- Dual-region endpoints (global/china)
- Advanced emotion synthesis
- Uses documented system voice IDs instead of dynamic voice-list fetching
Gemini TTS Features
- Gemini API-based high-quality voice synthesis
- 30+ voice options (star/moon themed names)
- Prompt-based style/tone control
- Simple API key authentication with
x-goog-api-key - Configurable Gemini API base URL
- 24+ language support including Japanese
Web Speech API Features
- Browser-native
speechSynthesisplayback with no API key - Browser-only runtime; no Node.js or server audio byte output
- Supports voice selection by
SpeechSynthesisVoice.nameorvoiceURI - Supports rate, pitch, volume, and language options where the browser voice honors them
Integration with AITuber OnAir Core
While this package can be used independently, it integrates seamlessly with @aituber-onair/core:
import { AITuberOnAirCore } from '@aituber-onair/core';
const core = new AITuberOnAirCore({
apiKey: 'your-openai-key',
voiceOptions: {
engineType: 'voicevox',
speaker: '1',
voicevoxApiUrl: 'http://localhost:50021'
}
});
// Voice synthesis is handled automatically
await core.processChat('Hello!');API Reference
VoiceServiceOptions
type VoiceServiceOptions =
| VoiceVoxVoiceServiceOptions
| VoicePeakVoiceServiceOptions
| OpenAiVoiceServiceOptions
| XaiVoiceServiceOptions
| UnrealSpeechVoiceServiceOptions
| ElevenLabsVoiceServiceOptions
| FishAudioVoiceServiceOptions
| DeepgramVoiceServiceOptions
| CartesiaVoiceServiceOptions
| InworldVoiceServiceOptions
| GradiumVoiceServiceOptions
| GeminiTtsVoiceServiceOptions
| OpenAiCompatibleVoiceServiceOptions
| AivisSpeechVoiceServiceOptions
| AivisCloudVoiceServiceOptions
| MinimaxVoiceServiceOptions
| PiperPlusVoiceServiceOptions
| WebSpeechVoiceServiceOptions
| NoneVoiceServiceOptions;VoiceServiceOptions is a discriminated union keyed by engineType.
Use updateOptions(...) for same-engine updates and switchEngine(...)
for cross-engine changes.
For backward compatibility, cross-engine fields in updateOptions(...)
are still accepted.
Engine Capabilities
import {
getAllVoiceEngineCapabilities,
getVoiceEngineCapabilities,
} from '@aituber-onair/voice';
const gradium = getVoiceEngineCapabilities('gradium');
console.log(gradium.supportsVoiceList); // true
const allEngines = getAllVoiceEngineCapabilities();Capabilities are static metadata only. They do not include API keys, endpoints, user configuration, or other sensitive values.
Voice Lists
import { getVoiceEngineVoiceList } from '@aituber-onair/voice';
const voices = await getVoiceEngineVoiceList('elevenLabs', {
apiKey: process.env.ELEVENLABS_API_KEY,
});
// [{ id: '...', label: 'Rachel (premade)' }, ...]getVoiceEngineVoiceList() returns normalized { id, label } items for
engines that expose list APIs: VOICEVOX, AivisSpeech, Aivis Cloud, xAI,
ElevenLabs, Fish Audio, Cartesia, Deepgram Flux, Inworld, Gradium, and Web Speech API. Pass
local apiUrl for VOICEVOX-compatible servers, apiKey for cloud engines that
require it, and language for Fish Audio, Cartesia, or Inworld filtering.
For browser apps, cloud provider voice-list endpoints must allow CORS. If a
provider blocks direct browser requests, call getVoiceEngineVoiceList() from
your backend or expose a small backend relay/proxy for the list endpoint.
Aivis Cloud voice-list support is package-level support for the model search endpoint. Browser requests to Aivis Cloud model/list endpoints can be blocked by CORS; call the helper from Node.js/backend or expose a backend proxy before wiring dynamic Aivis Cloud model selection into a production UI.
MiniMax is also excluded from this helper. Use the documented system voice IDs directly because the linked dynamic Get Voice API is currently unavailable.
VoiceService Methods
interface VoiceService {
speak(screenplay: ChatScreenplay, options?: AudioPlayOptions): Promise<void>;
speakText(text: string, options?: AudioPlayOptions): Promise<void>;
isPlaying(): boolean;
stop(): void;
updateOptions(
options: VoiceServiceOptionsUpdate | Partial<VoiceServiceOptions>
): void;
switchEngine?(options: VoiceServiceOptions): void;
}Screenplay Format
interface Screenplay {
emotion?: string;
text: string;
speechText?: string;
}Examples
React Integration
See the React example for a complete implementation:
import { useState } from 'react';
import { VoiceService } from '@aituber-onair/voice';
function VoiceDemo() {
const [voiceService] = useState(
() => new VoiceService({
engineType: 'openai',
speaker: 'alloy',
apiKey: 'your-api-key'
})
);
const handleSpeak = async (text: string) => {
await voiceService.speak({ text });
};
return (
<button onClick={() => handleSpeak('[happy] Hello!')}>
Speak with emotion
</button>
);
}Node.js Usage
The voice package now fully supports Node.js environments with automatic environment detection:
import { VoiceEngineAdapter } from '@aituber-onair/voice';
const voiceService = new VoiceEngineAdapter({
engineType: 'openai',
speaker: 'nova',
apiKey: process.env.OPENAI_API_KEY
});
// Audio will be played using available Node.js audio libraries
await voiceService.speak({ text: 'Hello from Node.js!' });Audio Playback in Node.js
For audio playback in Node.js, install one of these optional dependencies:
# Option 1: speaker (native bindings, better quality)
npm install speaker
# Option 2: play-sound (uses system audio player, easier to install)
npm install play-soundIf neither is installed, the package will still work but won't play audio. You can still use the onPlay callback to handle audio data:
const voiceService = new VoiceEngineAdapter({
engineType: 'voicevox',
speaker: '1',
voicevoxApiUrl: 'http://localhost:50021',
onPlay: async (audioBuffer) => {
// Save to file or process audio data
writeFileSync('output.wav', Buffer.from(audioBuffer));
}
});The package automatically detects the environment and uses the appropriate audio player:
- Browser: Uses HTMLAudioElement
- Node.js: Uses speaker or play-sound if available, otherwise silent
Testing
Run the test suite:
# Run all tests
npm test
# Run tests in watch mode
npm run test:watch
# Generate coverage report
npm run test:coverageContributing
Contributions are welcome! Please feel free to submit a Pull Request.
- Fork the repository
- Create your feature branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add some amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
License
This project is licensed under the MIT License - see the LICENSE file for details.
