@babylonjs-toolkit/kie
v1.0.1
Published
Babylon Toolkit MCP servers for kie.ai image + video + sound generation (Nano Banana, Kling, Seedance, Grok, Veo 3.1).
Maintainers
Readme
Babylon Toolkit MCP Servers
| Subcommand | Server | Tool | Models |
|---|---|---|---|
| image (default) | kie-image | generate_image | Nano Banana 2, Imagen 4, Seedream, Flux-2, Qwen, … |
| video | kie-video | generate_video | Kling, Bytedance Seedance, Grok Imagine |
| google | kie-google | generate_google_video | Google Veo 3.1 (veo3 / veo3_fast / veo3_lite) |
| sound | kie-sound | generate_sound | Suno effects/ambience/music, ElevenLabs speech |
Because MCP is an open protocol, the same package works from Claude Code, GitHub
Copilot Chat (VS Code), Cursor, Windsurf, Zed, and any other MCP client — one package,
every AI. Built with only Node built-ins (fetch, fs, readline), so it pulls no
transitive packages. Requires Node 18+.
The command takes one subcommand to pick the server (defaults to image):
kie-image-mcp image # or: video | google | sound (no arg = image)Get an API key
Set your kie.ai key as KIE_KEY (or KIE_AI_API_KEY), or put KIE_KEY=... in a .env
file in your working directory (or its parent). See .env.example. All four servers share the one key.
Install (local, default)
npm install --save-dev @babylonjs-toolkit/kieRegister the servers you need in your project .mcp.json:
{
"mcpServers": {
"kie-image": { "command": "node_modules/.bin/kie-image-mcp", "args": ["image"] },
"kie-video": { "command": "node_modules/.bin/kie-image-mcp", "args": ["video"] },
"kie-google": { "command": "node_modules/.bin/kie-image-mcp", "args": ["google"] },
"kie-sound": { "command": "node_modules/.bin/kie-image-mcp", "args": ["sound"] }
}
}For a source checkout, npm install builds dist/; an MCP entry can instead use
"command": "node", "args": ["/absolute/path/to/kie-image-mcp/dist/index.js", "sound"].
Restart your MCP client after adding the sound server. Use one configuration location
for your client; there is no need to duplicate the same servers in multiple files.
Install (global, when explicitly wanted)
The npm package is published as @babylonjs-toolkit/kie; it installs a command named
kie-image-mcp (that command name is what you reference everywhere else). Install once
so the command is on your PATH:
npm install -g @babylonjs-toolkit/kie # from npm once published
# or, from a local clone (the prepare script builds it for you):
cd kie-image-mcp && npm install -g .Verify and update:
which kie-image-mcp # confirm the command is on PATH
npm install -g @babylonjs-toolkit/kie@latest # update to a newer versionnvm users: the global bin is tied to the active Node version (e.g.
~/.nvm/versions/node/vXX/bin/kie-image-mcp). If you switch Node versions, reinstall for that version. If your MCP client is a GUI app that doesn't inherit your shell PATH, use the absolute bin path (fromwhich kie-image-mcp) in the config.
Configure (any MCP client)
Register one entry per server you want; the subcommand goes in args. KIE_KEY is read
from your environment or a .env file (add "env": { "KIE_KEY": "..." } to an entry to
set it inline).
Claude Code
The following examples use a global install. For a local install, use
node_modules/.bin/kie-image-mcp as shown above.
Add to your project .mcp.json (or user ~/.claude.json):
{
"mcpServers": {
"kie-image": { "command": "kie-image-mcp", "args": ["image"] },
"kie-video": { "command": "kie-image-mcp", "args": ["video"] },
"kie-google": { "command": "kie-image-mcp", "args": ["google"] },
"kie-sound": { "command": "kie-image-mcp", "args": ["sound"] }
}
}GUI-PATH fallback — use the absolute bin path from which kie-image-mcp:
{ "mcpServers": { "kie-image": { "command": "/Users/you/.nvm/versions/node/v24.11.0/bin/kie-image-mcp", "args": ["image"] } } }GitHub Copilot Chat (VS Code)
.vscode/mcp.json:
{
"servers": {
"kie-image": { "command": "kie-image-mcp", "args": ["image"] },
"kie-video": { "command": "kie-image-mcp", "args": ["video"] },
"kie-google": { "command": "kie-image-mcp", "args": ["google"] },
"kie-sound": { "command": "kie-image-mcp", "args": ["sound"] }
}
}Cursor / Windsurf / Zed / generic
Same shape under the client's mcpServers key, using the kie-image-mcp command.
Generate sound
Ask your agent, for example:
Use generate_sound to make loopable nighttime forest ambience, and save it to assets/audio/forest.mp3.
The tool uses the existing KIE authentication and MCP text-result conventions: submit
a task, poll until it completes, download one audio file into out_path, and report
the saved path, model, task ID and source URL. Parent directories are created. The
remote URLs are temporary, so use the saved local file in your game or application.
| kind | Default model | API | Controls |
|---|---|---|---|
| sound_effect (default) | V5 | Suno /api/v1/generate/sounds | Effects, ambience, loops, optional BPM and musical key |
| speech | elevenlabs/text-to-speech-multilingual-v2 | Market /api/v1/jobs/createTask | Text, voice and speech settings |
| music | V5 | Suno /api/v1/generate | Instrumentals, songs, custom lyrics and style |
Common arguments
| Parameter | Default | Meaning |
|---|---|---|
| prompt | Required | Effect description, spoken text, or music description/lyrics |
| out_path | Required | Local .mp3 filename, including relative paths and ~/…; replaced on successful download |
| kind | sound_effect | sound_effect, speech, or music |
| model | Per kind | Exact KIE model name from the lists below |
| track_index | 0 | Zero-based variation to save; other returned track URLs are included in the result |
| callback_url | Unset | HTTP(S) webhook you control; required for music, optional for effects/speech |
Each call submits a new generation; track_index selects an output of that new task.
To download another variation from an existing result, use its returned URL rather than
submitting the same prompt again. A polling/download error includes the task ID so you
can check that paid job before retrying. The tool never automatically resubmits a task.
Sound effects and ambience
Models: V5, V5_5. The prompt is limited to 500 characters. Optional loop is a
boolean (default false); tempo is an integer from 1 to 300 BPM; key accepts major
or minor keys such as C, F#, Am and D#m. Omit key to leave it unspecified.
{
"prompt": "Nighttime forest ambience with gentle wind, distant owls and soft crickets. No music or speech.",
"out_path": "assets/audio/forest.mp3",
"kind": "sound_effect",
"loop": true
}KIE's Sounds API also supports musical loops and background audio. It has no documented
exact duration parameter; this tool rejects duration for effects. Looping is a
generation request, not a local edit or a guarantee of a sample-perfect seam.
KIE Sounds API
Speech
Models: elevenlabs/text-to-speech-multilingual-v2 and
elevenlabs/text-to-speech-turbo-2-5. prompt is the exact spoken text (up to 5000
characters). Optional voice takes a supported KIE voice name or ID; the default is
James (EkK5I93UQWFDigLMpZcX). stability, similarity_boost and speech_style range
from 0 to 1; speed ranges from 0.7 to 1.2. Omitted controls use the model's defaults.
speech_style maps to ElevenLabs' numeric style, keeping it separate from music's
textual style. Only Turbo 2.5 accepts language_code (a two-letter ISO 639-1 code).
{
"prompt": "Welcome back, pilot. Your ship is ready for launch.",
"out_path": "assets/audio/welcome.mp3",
"kind": "speech",
"voice": "Rachel",
"stability": 0.5,
"speed": 1
}Supported voices and controls: Multilingual v2, Turbo 2.5.
Music
Models: V4, V4_5, V4_5PLUS, V4_5ALL, V5, V5_5.
KIE's music schema requires callBackUrl, even when results are polled. Supply
callback_url per call, or set KIE_CALLBACK_URL in your environment or .env.
Use a webhook endpoint you control that accepts KIE's POST callbacks; the stdio MCP
server does not run a webhook listener. The environment default is applied only to
music. Sound effects/ambience and speech need no callback URL.
{
"prompt": "An orchestral exploration theme with warm strings, subtle percussion and a hopeful melody.",
"out_path": "assets/audio/exploration.mp3",
"kind": "music",
"instrumental": true
}This example assumes KIE_CALLBACK_URL is configured. instrumental defaults to
true. Set it to false for singing. In the default custom_mode: false, the prompt
describes the track. With custom_mode: true, style and title become required,
and a vocal track sings the supplied prompt as lyrics. Custom mode also accepts
negative_tags and, for vocal tracks, vocal_gender: "m" | "f". Only V5_5 custom
mode accepts duration (10–360 seconds).
Prompt limits are 3000 characters in simple mode; custom mode allows 3000 for V4 and 5000 for the other models. Custom style limits are 200 for V4 and 1000 otherwise; titles are limited to 80. The tool rejects fields that would be ignored in the chosen mode. KIE Music API
Output and limits
- Saves the provider's MP3 bytes; no local transcoding, audio mixing or normalization.
- Waits for full completion (not streaming/first-track success). Speech polls every 6 seconds for up to 5 minutes; Suno polls every 30 seconds for up to 15 minutes. Individual HTTP requests/downloads have a 60-second timeout.
- HTTP errors, task failures, missing results and invalid audio downloads become MCP
isErrorresponses. A failed HTTP download does not overwrite the destination. - Local reference audio, voice cloning, audio isolation, music extension and WAV conversion use other APIs and are not parameters of this generation tool.
Development and verification
npm test # TypeScript build + offline request/poll/download/MCP tests
npm run build # emits all four servers into dist/Tests use mocked KIE responses and temporary output directories; they spend no KIE
credits. They also verify discovery of the existing image, video, Google and veo
subcommands. Live provider output quality/account availability requires a real API
generation. See research and implementation notes for the
source review, API differences and remaining live-verification limits.
