@youka/sdk
v0.2.0
Published
Official Node.js SDK for the Youka API.
Readme
@youka/sdk
Official Node.js SDK for the Youka API.
Use this package to create karaoke videos from Node.js: upload audio or video, sync or transcribe lyrics, create a karaoke project, render the final video, and download the MP4.
Install
npm install @youka/sdkAPI key
Create an API key at online.youka.io/account under API keys.
Keep the key outside your source code:
export YOUKA_API_KEY="yk_..."Create a karaoke video
import { AlignmentModel, YoukaClient } from "@youka/sdk";
const client = new YoukaClient({
apiKey: process.env.YOUKA_API_KEY!,
});
const operation = await client.projects.create({
source: {
type: "path",
path: "./song.mp3",
},
lyricsSource: {
type: "align",
lyrics: "Line one\nLine two\nLine three",
},
title: "Artist - Song",
});
const { project } = await client.projects.wait(operation);
const exportOperation = await client.exports.create(project.id, {
resolution: "1080p",
quality: "high",
});
const finalized = await client.exports.wait(exportOperation);
await client.exports.download(finalized, {
output: "./karaoke.mp4",
});Transcribe lyrics automatically
Use transcription when you do not already have lyrics text:
const operation = await client.projects.create({
source: { type: "path", path: "./song.mp3" },
lyricsSource: { type: "transcribe" },
title: "Artist - Song",
});When syncModel is omitted, Youka uses ElevenLabs Scribe. ElevenLabs
transcription is available to every user; AudioShake, MusicAI, and Whisper
transcription require premium access.
Use ElevenLabs Scribe
Choose ElevenLabs Scribe for word-timed transcription without supplied lyrics:
import { AlignmentModel } from "@youka/sdk";
const operation = await client.projects.create({
source: { type: "path", path: "./song.mp3" },
lyricsSource: {
type: "transcribe",
syncModel: AlignmentModel.ElevenLabsTranscription,
},
title: "Artist - Song",
});If you already have lyrics, pass them as optional transcription guidance:
const operation = await client.projects.create({
source: { type: "path", path: "./song.mp3" },
lyricsSource: {
type: "transcribe",
syncModel: AlignmentModel.ElevenLabsTranscription,
lyrics: "Line one\nLine two\nLine three",
},
title: "Artist - Song",
});Scribe always returns its own word text and timing. Supplied lyrics improve vocabulary recognition; they do not force the output to match the supplied text.
Use a hosted media URL
You can also create a karaoke video from a hosted media URL:
const operation = await client.projects.create({
source: {
type: "url",
url: "https://example.com/song.mp4",
},
lyricsSource: {
type: "align",
lyrics: "Line one\nLine two",
},
});URL helpers automatically ensure the local download binaries on first use. Supported hosted sites depend on yt-dlp: https://github.com/yt-dlp/yt-dlp/blob/master/supportedsites.md
Handle errors
import { YoukaRequestError, YoukaTaskError } from "@youka/sdk";
try {
const operation = await client.projects.create({
source: { type: "path", path: "./song.mp3" },
lyricsSource: { type: "transcribe" },
});
await client.projects.wait(operation);
} catch (error) {
if (error instanceof YoukaRequestError) {
console.error(error.code, error.status, error.message);
} else if (error instanceof YoukaTaskError) {
console.error(error.code, error.status, error.message);
} else {
throw error;
}
}Docs
- Node.js SDK guide: https://docs.youka.io/en/sdk
- Raw HTTP quickstart: https://docs.youka.io/en/api/quickstart
- CLI guide: https://docs.youka.io/en/cli
Model discovery and lyric videos
Use await client.capabilities.get() to discover supported models, workflows,
input requirements, language policies, and account eligibility. Every supported
model is represented; isolated-vocal models are restricted to compatible workflows.
Quotes and creation enforce the same rules.
const operation = await client.projects.create({
kind: "lyric-video",
source: { type: "path", path: "./reference.wav" },
title: "Lyric overlay",
lyricsSource: {
type: "transcribe",
syncModel: AlignmentModel.ElevenLabsTranscription,
lyrics: "Optional recognition hints",
languageHintMode: "auto",
},
});
const { project } = await client.projects.wait(operation);Lyric videos skip separation and reject splitModel. Omitted kind preserves
karaoke behavior. Use lyricsSource: { type: "align", lyrics: "...", syncModel }
for supplied-text alignment or lyricsSource: null to skip lyric processing.
projects.quote() accepts the same source/workflow configuration.
Correct timings and reuse versions
const list = await client.projects.alignments.list(project.id);
const current = await client.projects.alignments.get(
project.id,
list.alignments[0]!.id,
);
const corrected = structuredClone(current.alignment);
corrected.items[0]!.start = 1.25;
corrected.items[0]!.end = 1.75;
await client.projects.alignments.update(project.id, current.id, {
alignment: corrected,
expectedRevision: current.revision,
select: true,
expectedSelectionRevision: current.selectionRevision,
});
const versions = await client.projects.versions.list(project.id);
const versionId = versions[0]!.id;
const settings = await client.projects.getSettings(project.id, { versionId });
await client.projects.updateSettings(project.id, {
versionId,
settings: settings.settings,
});
const fresh = await client.projects.get(project.id);
await client.exports.create(project.id, {
target: "local",
versionId,
transparent: true,
stemVolumes: Object.fromEntries(fresh.stems.map((stem) => [stem.id, 0])),
outputPath: "./overlay.mov",
});Timing replacement uses complete, absolute decimal-second timings. Preserve IDs,
indexes, singer, and translation metadata. Invalid ranges are rejected and stale
revisions return conflicts; refetch and reconcile before retrying. Selection alone
uses projects.alignments.select(projectId, alignmentId, { expectedRevision,
expectedSelectionRevision }). Neither operation starts an AI task.
Local transparent exports use ProRes 4444. Muting stems means silent output; it does not promise removal of the audio stream. Local export starts no paid render or realignment job and retains existing feature eligibility. Line anticipation is layout-dependent; it is not a fixed per-line reveal guarantee.
These methods require the matching server endpoints to be deployed before the new SDK version is published. See the repository's capability-parity guide for rollout and acceptance requirements.
