claude-talkback-mcp
v0.4.0
Published
MCP server that gives Claude Code a conversational voice — speaks short, peer-style summaries aloud, ducking your music so it's audible. Two engines (Windows SAPI + ElevenLabs), works from WSL and native Windows.
Downloads
267
Maintainers
Readme
claude-talkback-mcp
Gives Claude Code a voice. It speaks short, conversational, peer-style summaries of what it's doing out your speakers — while the full detail still scrolls in the terminal as usual. Dual output: the screen has the transcript, the voice has the gist.
- Two engines, toggle live:
sapi— your OS's built-in voice, free and offline with zero setup: Windows SAPI, macOSsay, or Linuxespeak. (macOSsaysounds surprisingly good — no API key needed.)elevenlabs— natural cloud voice. Bring your own API key.
- A free built-in voice on every OS — no API key required anywhere:
- Windows / WSL — Windows SAPI. From WSL it reaches the Windows audio stack via
powershell.exe(SAPI speaks directly; ElevenLabs MP3 plays through Windows MediaPlayer), so there's no audio-passthrough setup. - macOS — the built-in
saycommand. Nothing to install, and its voices (Samantha, Alex, the downloadable "enhanced" ones) sound noticeably better than Windows SAPI. - Linux —
espeakfor the built-in voice (plusffplayfor ElevenLabs playback).
- Windows / WSL — Windows SAPI. From WSL it reaches the Windows audio stack via
- Pick any voice —
list_voices/set_voice, an env default, or a per-line override. - Non-blocking & interruptible — speaking never stalls Claude; a new line can cut off the current one (barge-in) so it pivots to you mid-sentence.
- Graceful fallback — if an ElevenLabs call fails (quota, tier-locked voice, network), it automatically falls back to SAPI so you're never left in silence.
- Agents take turns — multiple Claude Code windows each run their own talkback process, so they coordinate through a shared lockfile and never talk over each other: one voice plays at a time, machine-wide. Combined with per-repo voices, you always know which agent is speaking.
- Ducks your music (Windows/WSL) — if something is actually playing when Claude speaks (iTunes, Pandora, Spotify, a browser tab), it's briefly turned down so the voice is audible, then returned to exactly the level you had it at. Only apps genuinely emitting sound are touched, and volumes are journalled to disk first, so even a mid-sentence interruption restores your level rather than guessing.
Demo
Hear the difference — click to listen/download (GitHub doesn't autoplay audio inline):
- ▶️ Natural voice (ElevenLabs):
assets/demo-elevenlabs.mp3 - ▶️ Free built-in voice (Windows SAPI):
assets/demo-sapi.wav
ElevenLabs sample generated with ElevenLabs.
Tools
| Tool | What it does |
|------|--------------|
| speak | Speak one or two short sentences. interrupt:true cuts off current speech; voice overrides for that line. |
| stop_speaking | Immediately silence and clear the queue. |
| list_voices | List every voice available for the active engine (★ = active). |
| set_voice | Set + remember the voice for this repo. Takes a name/id, a partial ("jessica"), a gender ("female"/"male" → random of that gender), or "random". |
| set_engine | Toggle between sapi and elevenlabs at runtime. |
| voice_status | Report the active engine + current voice, including whether it's male or female. |
| list_repo_voices | Show the central registry — which voice each repo/project has been assigned. |
| check_setup | Report platform, engine, whether the required audio tool + key are installed (with fixes), and whether media ducking is active. |
Requirements
This is a Node/TypeScript MCP server (no Python). Beyond Node, the only dependencies are the system audio tools — and Windows/WSL already ship them:
- Node.js ≥ 18.
- Windows or WSL (primary target): nothing extra. PowerShell + the built-in Windows speech engine handle SAPI and ElevenLabs MP3 playback, and media ducking works here.
- macOS (experimental): the built-in
saygives a good free voice out of the box — no ElevenLabs needed.brew install ffmpegadds theffplayused only for ElevenLabs playback. Ducking is unavailable (no per-app audio API). - Linux (experimental):
sudo apt install espeakfor the free local voice, andsudo apt install ffmpegfor ElevenLabs playback. Ducking is unavailable. - ElevenLabs engine (optional): an API key in
ELEVENLABS_API_KEY+ network access.
If speech doesn't play, ask Claude to run check_setup — the server reports the exact tool
that's missing and how to install it, and every speak result warns you when a backend is absent.
Platform support: built and tested for WSL + Windows. The macOS/Linux paths are best-effort and less battle-tested;
check_setupwill tell you what's needed there.
Install
npm install -g claude-talkback-mcpLatest from GitHub (unreleased changes):
npm install -g github:jarvann/claude-talkback-mcpFrom source (for development):
git clone https://github.com/jarvann/claude-talkback-mcp
cd claude-talkback-mcp
npm install && npm run buildRegister with Claude Code
--scope user makes it available in every project (use --scope project to scope to one repo).
The same commands work on native Windows/PowerShell — the server auto-detects the environment.
Built-in voice (free, offline):
claude mcp add talkback --scope user -- claude-talkback-mcpWith ElevenLabs (natural voice):
claude mcp add talkback --scope user \
--env ELEVENLABS_API_KEY="<your-elevenlabs-key>" \
--env TALKBACK_ENGINE="elevenlabs" \
-- claude-talkback-mcpNot installed globally? Replace
claude-talkback-mcpwith eithernpx -y github:jarvann/claude-talkback-mcp(no install), or — from a source checkout —node /path/to/claude-talkback-mcp/dist/index.js.
Claude Code loads MCP servers at startup — after
mcp add/ changing env, restart Claude Code (or/mcp→ reconnect) for changes to take effect. Confirm with/mcp.
Configuration (env vars)
General
| Var | Default | Meaning |
|-----|---------|---------|
| TALKBACK_ENGINE | elevenlabs if a key is set, else sapi | Which engine to start on. |
| TALKBACK_MAX_CHARS | 600 | Hard cap on spoken length — safety net so nothing long is read aloud. |
| TALKBACK_NO_AUDIO_LOCK | (off) | Set to 1 to disable cross-agent coordination and let windows play audio at the same time. |
| TALKBACK_STATE_DIR | ~/.claude-talkback | Where per-repo voice choices, the audio lock, and ducking journals are stored. |
| TALKBACK_NO_DUCKING | (off) | Set to 1 to leave other apps' volume alone while speaking. |
| TALKBACK_DUCK_LEVEL | 0.15 | How far playing apps are turned down, as a fraction of their current volume (0–1). 0.15 = 15% of wherever you had it. |
SAPI
| Var | Default | Meaning |
|-----|---------|---------|
| TALKBACK_VOICE | system default | SAPI voice name (e.g. Microsoft Zira Desktop). |
| TALKBACK_RATE | 1 | Speech speed, SAPI scale -10..10. |
ElevenLabs
| Var | Default | Meaning |
|-----|---------|---------|
| ELEVENLABS_API_KEY | — | Your key (required for the elevenlabs engine). |
| ELEVENLABS_VOICE_ID | Jessica (cgSgspJ2msm6clMCkdW9) | Default voice id. |
| ELEVENLABS_MODEL | eleven_flash_v2_5 | TTS model. Flash = fastest (~75ms) and half-price; use eleven_multilingual_v2 or eleven_v3 for richer, more expressive quality (slower, full price). |
| ELEVENLABS_FORMAT | mp3_44100_128 | Output format (PCM needs a paid tier). |
| ELEVENLABS_SPEED | 0.9 | Pacing, 0.7–1.2. <1 = slower/less rushed. |
| ELEVENLABS_STABILITY | 0.5 | 0–1. Higher = more consistent delivery. |
| ELEVENLABS_SIMILARITY | 0.75 | 0–1. Similarity boost. |
A note on free-tier ElevenLabs
On a free ElevenLabs plan, only the ~21 premade voices work through the API. "Professional"
/ Voice-Library voices you've added (e.g. Ava) return 402 paid_plan_required via the API even
though they work in the ElevenLabs app — using them here needs a paid plan (Starter and up). When
that happens the server logs it and falls back to SAPI for that line. Once you upgrade,
set_voice ava works with no code changes.
Per-repo voices (a different voice per project)
Every repo automatically gets its own voice, so when you have multiple Claude Code windows open you can tell them apart by ear — like different teammates.
Auto-assigned on first load. The first time the server starts in a repo, it picks a voice at random and remembers it — preferring one no other repo is already using, so projects stay distinct until you run out of voices.
Persistent. Choices are saved to
~/.claude-talkback/repos.json, keyed by repo path, so a repo keeps its voice across restarts. (Override the location withTALKBACK_STATE_DIR.)Change it anytime. Ask Claude to switch —
set_voicetakes a name ("brian"), a gender ("female"/"male"→ a random voice of that gender), or"random", and saves the new pick for that repo.voice_statussays which voice/gender is active right now.See the whole map.
list_repo_voicesprints the central registry (the→marks the current repo):Voices by repo: /work/repoA — Callum - Husky Trickster, male [elevenlabs] → /work/repoB — Alice - Clear, Engaging Educator, female [elevenlabs] /work/repoC — Liam - Energetic, Social Media Creator, male [elevenlabs]
Auto-pick only draws from free-tier premade ElevenLabs voices, so it never lands on a
tier-locked one. To pin a repo to a specific voice, set ELEVENLABS_VOICE_ID in that repo's
registration — an explicit pin always wins and is never overwritten.
Ducking your music (Windows / WSL)
When Claude speaks, anything already playing is briefly turned down so the voice isn't buried, then put back. This is per-application — it moves the app's slider in the Windows Volume Mixer, not your master volume and not the app's own internal volume.
- Only apps actually making noise are touched. Talkback checks each audio session's live peak meter, so an app that merely has the audio device open (a silent browser tab, an idle Zoom) is left alone. Sampling runs over a short window because a single reading often lands between waveform peaks and reads as silence.
- Your level is preserved exactly. Ducking is proportional — if you had Spotify at 55%, it drops to 15% of that and returns to 55%, not to 100%.
- Interruptions are safe. Barge-in kills the speaking process outright, which would otherwise
strand your music quiet forever. Volumes are journalled to
~/.claude-talkback/ducked-<pid>.jsonbefore anything changes, so an interrupted line — or a talkback crash, recovered on next startup — still restores your exact levels. The journal is per-process, and an instance only reclaims journals whose owner has exited, so several Claude Code windows never trip over each other's recovery data.
Set TALKBACK_DUCK_LEVEL to change how far things drop, or TALKBACK_NO_DUCKING=1 to turn it
off. check_setup reports whether ducking is active.
Windows/WSL only. macOS has no public API for ducking another app's audio — the workable options there are lowering system volume (which would quiet Claude's own voice too) or an AppleScript allowlist that only covers a couple of apps and misses browser audio entirely. Rather than ship something that half-works, macOS and Linux simply speak over the music, exactly as before.
How the conversational behavior works
The server ships instructions (surfaced to Claude Code on connect) telling Claude to:
- End each top-level response with a short spoken summary — never read code, logs, or long lists aloud, just the gist.
- Don't go silent during long multi-step work — speak a one-line update at each milestone (not every tool call), so you hear progress as it happens instead of only at the end.
- Give a spoken heads-up before slow work, then run it in the background so its turn ends and you can keep chatting — then speak the result when it finishes.
- Keep turns short so you can interrupt between them; use
interrupt:trueto pivot when you talk. - Talk like a collaborator — warm, first person, no filler.
Note on "talk to it while it works": Claude Code is turn-based — while it's blocked on a single foreground operation it isn't running, so there's no true simultaneous chat within one call. The conversational feel comes from Claude backgrounding slow work so turns stay short and control returns to you frequently. That behavior lives in the instructions above.
License
MIT © Cory Loriot
