@breeze.blue/sdk
v0.19.0
Published
ESM-first TypeScript SDK for the Breeze Blue Developer API.
Maintainers
Readme
@breeze.blue/sdk
ESM-first TypeScript SDK for the Breeze Blue Developer API. Covers text-to-speech, voice management, voice preview generation, history audio, models, and account usage.
The root entrypoint is safe to import in Node, browser, and edge runtimes. Local
audio playback helpers live in the Node-only @breeze.blue/sdk/node subpath.
Install
pnpm add @breeze.blue/sdkor:
npm install @breeze.blue/sdkAPI Key
Create an API key in the Breeze Blue Developer Console, then export it:
export BREEZE_API_KEY=brz_...The SDK sends the key with the xi-api-key header. It also sends
x-breeze-sdk for server-side observability.
An API key is required for every REST method. The one exception is
textToSpeech.realtime.connect(...) with a clientSecret, which lets a
browser client run without a key (see the realtime section below).
Quickstart
import { BreezeBlueClient } from "@breeze.blue/sdk";
import { save } from "@breeze.blue/sdk/node";
const client = new BreezeBlueClient();
const audio = await client.textToSpeech.convert(
"voc_...",
{ text: "Hello from Breeze Blue." },
{ outputFormat: "mp3" },
);
await save(audio, "hello.mp3");
console.log(audio.contentType);
console.log(audio.historyItemId);new BreezeBlueClient() reads BREEZE_API_KEY by default and sends requests to
https://api.breeze.blue. To point at another environment, pass baseUrl or
set BREEZE_BASE_URL.
const client = new BreezeBlueClient({
apiKey: "brz_...",
baseUrl: "https://api.breeze.blue",
timeout: 120_000,
});Per-request options accept timeout, signal, and extra headers.
Text to Speech
const audio = await client.textToSpeech.convert("voc_...", {
text: "Render this line.",
});
const audioStream = await client.textToSpeech.stream("voc_...", {
text: "Stream this line.",
});
const enhanced = await client.textToSpeech.enhance({
instruction: "Calm, warm, bedtime narration.",
languageCode: "en",
});Use async text-to-speech for long text, reference-heavy voices, or batch production where the caller should not hold an HTTP connection open:
const job = await client.textToSpeech.createJob(
"voc_...",
{ text: "Render this longer script." },
{ outputFormat: "mp3" },
);
const status = await client.generationJobs.get(job.generationJobId);
if (status.status === "ready") {
const audio = await client.generationJobs.downloadAudio(job.generationJobId);
await save(audio, "async.mp3");
}If the job is still active, downloadAudio(...) rejects with
BreezeBlueGenerationNotReadyError; read error.retryAfter before retrying.
Use realtime text-to-speech when one WebSocket connection should handle multiple
conversation turns. Realtime audio is fixed to raw pcm_s16le, 24000 Hz, mono,
16-bit frames.
Realtime needs a WebSocket implementation. Browsers, edge runtimes, and
Node 22+ provide a global one; on Node 20, pass a constructor (for example the
ws package's WebSocket) via new BreezeBlueClient({ webSocket }).
Start consuming before appending text. Audio can arrive after flush() and may
continue after endTurn(); keep the single consumer running until turn.done.
const connection = await client.textToSpeech.realtime.connect("voc_...", {
modelId: "breeze-tts-2",
});
const pcmChunks: Uint8Array[] = [];
const consumer = (async () => {
for await (const message of connection) {
if (message.type === "audio") {
// Forward this chunk to your playback or transport layer immediately.
pcmChunks.push(message.audio);
} else if (message.type === "error") {
throw new Error(`Realtime TTS failed: ${message.code}: ${message.message}`);
} else if (message.type === "turn.done") {
return;
}
}
throw new Error("Realtime session closed before turn.done");
})();
connection.startTurn("turn_1");
connection.appendText("Hello from Breeze.");
connection.flush();
connection.endTurn();
await consumer;
connection.close();Always consume the connection (or its audio() / events() iterators) while a
turn is active, and call close() when you are done. When the SDK detects a
connection failure — an invalid server frame, or more than 16 MiB of messages
piling up unconsumed — it closes the connection and the iterator rejects with
a BreezeBlueRealtimeError. The audio() helper also throws
BreezeBlueRealtimeError for server error events and abnormal WebSocket
closes instead of silently ending.
When BreezeBlueRealtimeError.recoverable is true, the same connection
remains usable. Most recoverable turn-command errors cancel the active turn.
An active-turn session.update rejection is the exception: the turn keeps
running. If connection.audio() throws that error, start a fresh audio()
iterator immediately on the same connection and continue collecting the
current turn; do not start a replacement turn.
Events are a typed discriminated union (session.ready, session.updated, turn.started,
audio.started, turn.done, turn.cancelled, usage.committed,
session.expiring, session.closed, error, pong), with camelCase fields
such as turnId, historyItemId, expiresAt, and ttfaMs. Treat the union as
non-exhaustive: the server may add event types, and the SDK delivers unknown
JSON events unchanged — ignore event types you do not recognize instead of
switching exhaustively.
Between turns, update the synthesis instructions without replacing the
WebSocket. updateInstructions(...) only sends the request; the matching
session.updated acknowledgement remains on the same ordered iterator. Wait
for it before starting the next turn:
connection.updateInstructions("Speak faster and with more energy.");
for await (const message of connection) {
if (message.type === "session.updated") {
if (message.instructions !== "Speak faster and with more energy.") {
throw new Error("Unexpected instructions acknowledgement");
}
break;
}
}
connection.startTurn("turn_2");Instructions must be a non-empty string of at most 1,000 characters. Only one
update may await acknowledgement, and updates are accepted only while no turn
is active. An update attempted during a turn is rejected without interrupting
that turn. Keep one stream consumer: if a long-lived consumer owns the
iterator, have it signal your turn producer when session.updated arrives
instead of starting a second iterator.
Send a keepalive ping inside the session's inactivityTimeoutSeconds window;
the server answers with a pong event. Keep it running during active turns too:
a long TTFA or upstream stall with no audio does not pause the idle deadline.
const keepalive = setInterval(() => connection.ping(), 10_000);
// ... run turns ...
clearInterval(keepalive);
connection.close();For a long logical conversation, use the opt-in managed connection. It derives
a safe keepalive interval from session.ready and requires a server response
after each ping. A missing acknowledgement replaces an idle half-open socket;
if a turn is active, it reports TURN_INTERRUPTED and never replays commands.
The manager also marks a physical WebSocket for turn-boundary rotation after
10 minutes by default (or earlier when the server deadline requires it). An
active turn can delay the switch beyond that threshold; the server-advertised
hard lifetime still applies. The manager
performs up to three unexpected idle reconnect attempts per interruption with
fresh sessions and jittered exponential backoff:
const connection = await client.textToSpeech.realtime.connectManaged("voc_...", {
modelId: "breeze-tts-2",
});
const consumer = (async () => {
for await (const message of connection) {
if (message.type === "audio") {
// Play or forward this PCM chunk immediately.
} else if (message.type === "session.ready") {
// A new physical epoch is ready; the logical connection remains the same.
} else if (message.type === "turn.done") {
return;
}
}
})();
await connection.startTurnWhenReady("turn_1");
connection.appendText("Hello from a long-running conversation.");
connection.endTurn();
await consumer;
connection.close();Use heartbeatTimeoutMs to tune the acknowledgement deadline and
maxPhysicalSessionMs to tune the client-side epoch cap. Set
maxPhysicalSessionMs: 0 only when you intentionally want to rely solely on
the server-advertised deadline.
When an idle reconnect is handled successfully, the logical iterator stays
open, suppresses the replaced physical socket's terminal event, and emits the
new epoch's session.ready. The replacement sends its first heartbeat
immediately and resumes the derived cadence after an inbound acknowledgement.
Reconnect attempts are bounded per interruption. startTurnWhenReady(...)
waits only when one of these idle replacements is already in progress and
turn.start has not been sent; await it before sending any other turn command.
After a planned rotation, an old epoch with outstanding usage.committed
events drains in parallel for up to five seconds and closes as soon as all
known completed turns settle. This event is best-effort; use history and usage
APIs as the durable source of truth.
The managed connection never buffers turn content or replays a command. The
bounded startTurnWhenReady(...) wait happens before its first WebSocket
write. If a WebSocket is interrupted while a turn is active, its iterator rejects with
BreezeBlueRealtimeError and error.code === "TURN_INTERRUPTED"; decide from
your application conversation state whether and how to start a new turn.
Treat the manager as an active realtime call, not a presence channel. Always
call connection.close() when the call ends, the user leaves, or the page
enters a long-lived background state; otherwise its heartbeat intentionally
keeps a server WebSocket slot occupied.
To cut time to first audio, create the session ahead of time (for example while
your app is still preparing the turn) and connect with its clientSecret when
the first text is ready — only the WebSocket handshake remains:
const session = await client.textToSpeech.realtime.createSession("voc_...", {
modelId: "breeze-tts-2",
});
// Later, when the first text is ready:
const connection = await client.textToSpeech.realtime.connect("voc_...", {
clientSecret: session.clientSecret,
websocketUrl: session.websocketUrl,
directWebsocketUrl: session.directWebsocketUrl,
});When the session response includes directWebsocketUrl, the SDK prefers its
query-free origin and authenticates with the WebSocket subprotocol. It falls
back to websocketUrl for services that have not enabled the direct transport.
The same clientSecret handoff lets a browser connect without ever seeing your
API key: create the session on your server, hand session.clientSecret to the
page, and build a key-less client there:
// Browser — no API key required for clientSecret connections.
const browserClient = new BreezeBlueClient();
const connection = await browserClient.textToSpeech.realtime.connect("voc_...", {
clientSecret, // received from your server
});For a managed browser connection, provide a sessionFactory. The SDK invokes it
for every physical epoch, so the callback must fetch a newly minted secret from
your backend rather than cache the first response. Session configuration belongs
in that backend request when a custom factory is used:
import { BreezeBlueRealtimeError } from "@breeze.blue/sdk";
const connection = await browserClient.textToSpeech.realtime.connectManaged("voc_...", {
sessionFactory: async ({ signal }) => {
const response = await fetch("/api/breeze-realtime-session", {
method: "POST",
signal,
});
if (!response.ok) {
const reconnect =
[408, 425, 429].includes(response.status) || response.status >= 500;
throw new BreezeBlueRealtimeError("Could not create realtime session", {
code: "SESSION_FACTORY_ERROR",
reconnect,
});
}
return response.json(); // { clientSecret, websocketUrl?, directWebsocketUrl? }
},
});Honor the factory signal: the SDK aborts it when the logical connection closes
or the physical-epoch timeout expires. A custom callback that ignores the
signal must still enforce its own bounded request timeout.
After a dynamic instruction update receives session.updated, the managed SDK
carries the confirmed value to later physical epochs. SDK-created replacements
include it in their session request; factory-created replacements receive an
automatic session.update before their next session.ready is exposed to the
logical iterator.
modelId, languageCode, voiceSettings, inactivityTimeoutSeconds, and
enableLogging are fixed when the session is created. The initial
instructions value is also supplied at creation, but it can later be
replaced between turns with updateInstructions(...). connect ignores
initial session options when clientSecret or websocketUrl is provided and
logs a warning.
If a realtime WebSocket is interrupted by a network change, service deployment,
or upstream realtime worker restart, connectManaged(...) handles bounded
reconnects only while there is no active turn. With the lower-level
connect(...), handle error.meta.reconnect === true or a session.closed
event with reconnect === true by creating a new session and starting a new
turn from your own conversation state. Active turns are never resumed in place.
During a managed reconnect, transient session-creation responses (408, 425, 429,
and 5xx) share the same bounded reconnect budget; other 4xx responses fail the
logical connection immediately. A custom sessionFactory should preserve that
distinction with new BreezeBlueRealtimeError(message, { reconnect }); a plain
Error cannot communicate a terminal policy response and is treated as
transient inside the same bounded budget. GENERATION_CAPACITY_EXCEEDED is
turn-scoped: wait for the following turn.cancelled, back off using
meta.retryAfterSeconds, and start a new turn on the same managed logical
connection without replaying text.
The API uses the default text-to-speech model when modelId is omitted. If
you need to select a model explicitly, call client.models.list() and pass one
of the returned modelId values.
All field names follow TypeScript conventions (modelId, voiceSettings,
outputFormat, historyItemId, ...). The SDK translates them to the
snake_case HTTP wire format on send and translates JSON responses back to
camelCase on receive.
Audio responses expose convenient helpers:
await audio.arrayBuffer();
await audio.bytes();
await audio.blob();
audio.contentType;
audio.historyItemId;Node playback helpers:
import { play, save, stream } from "@breeze.blue/sdk/node";
await play(audio); // ffplay, with macOS afplay fallback
await save(audio, "x.mp3");
await stream(audioStream); // mpvStreaming text-to-speech defaults to pcm to reduce time to first audio. Pass
{ outputFormat: "wav" } or { outputFormat: "mp3" } when you need that wire
format explicitly.
Sync and async text-to-speech also take a sample rate and, for the lossy encodings, a bitrate, such as
{ outputFormat: "wav_48000" }. Breeze resamples its native 24000 Hz output for
downstream compatibility; profiles carrying a sample rate are not available for streaming.
Voices
Search every voice available to the account and inspect a single voice. Public catalog results combine semantic similarity with exact, prefix, and substring name matching; saved personal voices remain searchable by name, description, or voice ID.
const voices = await client.voices.search({ search: "calm documentary narrator" });
const newestFirst = await client.voices.search();
const dailyTrend = await client.voices.search({
voiceType: "default",
sort: "trend",
languageCode: "zh",
gender: ["female", "neutral"],
tone: ["warm"],
});
const taggedFavorites = await client.voices.search({
favoritesOnly: true,
tags: ["narration", "novel=三体"],
});
const firstVoiceId = voices.voices[0].voiceId;
const voice = await client.voices.get(firstVoiceId);
const settings = await client.voices.getSettings(firstVoiceId);
const randomVoice = await client.voices.random();
console.log(randomVoice.voiceId, randomVoice.name);
// Narrow the random pool by language and/or voice source.
const randomCatalogVoice = await client.voices.random({ languageCode: "zh", voiceType: "default" });
// Discover the current code-only Voice Metadata contract before building a form.
const metadataOptions = await client.voices.metadataOptions();
console.log(metadataOptions.languageCodes);
console.log(metadataOptions.accentCodesByLanguage.en);Voice listings keep the published default of stable creation-time order from
newest to oldest. Use { sort: "trend", voiceType: "default" } for the shared
Daily Trend order. Eligible voices not yet ranked follow the snapshot members;
returned page tokens pin the initial supplemental candidates. Search queries remain relevance-ranked; Trend is only a
weak prior within the same relevance tier.
Breeze voice creation is always two steps: produce a preview, let the user accept it, then save the preview as a real voice. Three ways to produce a preview:
import { readFile } from "node:fs/promises";
// Option A — clone preview from an audio sample
const clonePreview = await client.voices.createClonePreview({
name: "Demo voice",
file: {
data: await readFile("sample.wav"),
filename: "sample.wav",
contentType: "audio/wav",
},
});
let generatedVoiceId = clonePreview.generatedVoiceId;
// Option B — design preview from a text description (no audio)
const design = await client.voices.createDesignPreview({
voiceDescription: "Warm documentary narrator with clear articulation.",
});
generatedVoiceId = design.previews[0].generatedVoiceId;Option C keeps an existing voice and makes it speak another language.
Localization runs as a background job: start it, then poll until the status is
ready. name defaults to the source voice name.
const job = await client.voices.createLocalizePreview({
voiceId: firstVoiceId,
languageCode: "es",
name: "Documentary narrator (Spanish)",
});
let localized = await client.voices.getLocalizePreview(job.generationJobId);
while (localized.status !== "ready") {
if (localized.status === "failed" || localized.status === "cancelled") {
throw new Error(localized.error?.detail ?? localized.status);
}
await new Promise((resolve) => setTimeout(resolve, 2000));
localized = await client.voices.getLocalizePreview(job.generationJobId);
}
generatedVoiceId = localized.generatedVoiceId!;Voice Remix creates a same-language candidate from an owned or official voice and a custom description or preset. Each successful candidate uses the Voice Design candidate price. The source remains unchanged.
const remixJob = await client.voices.createRemixPreview({
voiceId: firstVoiceId,
prompt: "Make the voice warmer and more expressive.",
guidanceScale: 4,
});
let remix = await client.voices.getRemixPreview(remixJob.generationJobId);
while (remix.status !== "ready") {
if (remix.status === "failed" || remix.status === "cancelled") {
throw new Error(remix.error?.detail ?? remix.status);
}
await new Promise((resolve) => setTimeout(resolve, 2000));
remix = await client.voices.getRemixPreview(remixJob.generationJobId);
}
generatedVoiceId = remix.generatedVoiceId!;Use streamPreview and savePreview with the generated ID to audition and save
a private, independent voice. A client timeout leaves the server job running.
Custom input is enhanced into an instruction and matching script in the source
language. List templates with client.voices.listRemixPresets() and pass
presetId instead of prompt to use one. Edited templates are enhanced;
unchanged templates are translated only when the source language differs.
To compare strengths, call client.voices.prepareRemixPreview(...) first and
pass its preparationToken to each create call with the same input. The token
lasts 24 hours. Values 2 / 4 / 6 / 8 mean Stable / Expressive / Creative / Intense;
4 is the default. Web Auto compares separate candidates at 4, 6, and 8.
Use options.headers["Idempotency-Key"] to retry uncertain submissions safely,
with a distinct key per candidate. Status includes comparisonAudioUrl,
comparisonText, and sourceVoice. Saving without voiceName or
voiceDescription uses the source values.
See the Voice Remix guide.
For the fastest first audio, generate one design preview as a live 24 kHz mono
PCM response. The Node stream() helper consumes the response body directly;
a clean end means generatedVoiceId is ready to save:
import { stream } from "@breeze.blue/sdk/node";
const livePreview = await client.voices.streamDesignPreview({
voiceDescription: "Warm documentary narrator with clear articulation.",
text: "This is a short preview script.",
});
await stream(livePreview);
generatedVoiceId = livePreview.generatedVoiceId!;files is also accepted with exactly one item.
Download a completed preview so the user can audition it, then save the one they pick:
const audio = await client.voices.streamPreview(generatedVoiceId);
await save(audio, "preview.mp3");
const saved = await client.voices.savePreview({
generatedVoiceId,
voiceName: "Documentary narrator",
languageCode: "en",
gender: "neutral",
age: "middle_aged",
tone: ["warm", "articulate"],
accent: "american",
tags: ["narration", "novel=三体"],
});Edit, tune settings, or delete a saved voice:
await client.voices.edit(saved.voiceId, {
name: "Renamed narrator",
languageCode: "en",
tone: ["calm", "measured"],
});
// null clears nullable metadata; [] clears tone. Omitted fields remain unchanged.
await client.voices.edit(saved.voiceId, { accent: null, tone: [] });
await client.voices.editSettings(saved.voiceId, { guidanceScale: 1.2 });
await client.voices.favorite(firstVoiceId);
await client.voices.updateTags(firstVoiceId, ["narration", "novel=三体"]);
await client.voices.updateTags(firstVoiceId, []); // Resume owner tag inheritance.
await client.voices.unfavorite(firstVoiceId);
await client.voices.delete(saved.voiceId);Voice Metadata uses stable code values. gender, age, and accent may be
null; tone contains up to three distinct codes. English and Chinese use
different accent code sets, and other languages require accent: null.
Call client.voices.metadataOptions() to discover every accepted code and the
language-to-accent cascade at runtime.
History, Models, and Account
const models = await client.models.list();
const balance = await client.account.balance();
const currentKey = (await client.account.currentApiKey()).apiKey;
console.log(currentKey.status, currentKey.creditRemaining);
const usage = await client.account.usage({ days: 7 });
const keyUsage = await client.account.usage({
apiKeyId: "key_01hprod",
clientType: "sdk",
});
const history = await client.history.list({ pageSize: 10 });
const item = await client.history.get(history.history[0].historyItemId);
const audio = await client.history.downloadAudio(item.historyItemId);Errors
API failures throw BreezeBlueAPIError subclasses:
BreezeBlueAuthenticationErrorBreezeBlueBadRequestErrorBreezeBlueConflictErrorBreezeBlueForbiddenErrorBreezeBlueNotFoundErrorBreezeBlueValidationErrorBreezeBlueRateLimitErrorBreezeBlueInsufficientCreditsErrorBreezeBlueUpstreamErrorfor upstream generation/storage/model failures, commonly 502/504BreezeBlueServiceUnavailableErrorfor capacity or service-configuration failures, commonly 503
import { BreezeBlueClient, BreezeBlueRateLimitError } from "@breeze.blue/sdk";
try {
await new BreezeBlueClient().models.list();
} catch (error) {
if (error instanceof BreezeBlueRateLimitError) {
console.log(error.headers.get("retry-after"));
}
}Each API error exposes status, code, detail, meta, and headers.
Clone previews automatically generate a short script in the detected reference audio language. They accept a reference sample, name, and optional description. Save the preview as a voice, then use Text to Speech for custom scripts, instructions, or target languages.
HTTP TTS speed
Pass voiceSettings: { speed: 1.25 } to control speech speed without changing pitch, from 0.5 to 2.0. Omission always means 1.0; saved voice speed is not inherited. Sync, async and HTTP streaming support this parameter; Realtime does not.
Pass voiceSettings: { volume: 1.5 } to apply a request-level linear amplitude multiplier. Values range from 0.01 to 2.0; omission means 1.0. Values above 1.0 may clip peaks. Sync, async and HTTP streaming support this parameter; it does not change saved voice settings or apply to Realtime.
Word timing
Stream audio and word/token timestamps with client.textToSpeech.streamWithTimestamps(...). Use for await to consume chunks containing audioBase64 and wordTimestamps; breaking iteration closes the stream. Decode the audio separately and merge repeated word indices. See Speech timing for complete examples and timing semantics.
For a complete response, use client.textToSpeech.convertWithTimestamps(...). It returns base64 audio, its content type, and the full word/token timestamp list. Decode the audio before saving or playback. See Speech timing and Convert with timestamps.
For background generation with timing, use client.textToSpeech.createJobWithTimestamps(...). Poll with client.generationJobs.get(jobId); a ready job includes wordTimestamps and the audio download URL.
Clone preview creation accepts an omitted name. The server uses the filename stem or Cloned Voice; saving a Clone preview without a name keeps its preview name. Other preview types still require a save name. Explicit names longer than 80 Unicode characters are rejected.
