dsh-live-voice
v0.3.0
Published
Local-first hands-free AI voice assistant plugin for DeepSeek Harness (DSH), with speech-to-text, text-to-speech, voice commands, and local speech engines.
Maintainers
Keywords
Readme
DSH Live Voice
A local-first, hands-free voice assistant plugin for DeepSeek Harness (DSH). Built in Brazil 🇧🇷 and tested daily with Brazilian Portuguese on macOS.
Speak, listen, answer prompts, and code without touching your keyboard. DSH Live Voice coordinates speech-to-text (STT) and text-to-speech (TTS) into a single, conflict-free conversational flow.
⚡ Quick Start
Install with npx dsh into your DSH web profile:
npx dsh plugin add --profile web dsh-live-voiceOpen DSH Settings → Live Voice after your next DSH startup.
🌍 Interface Languages
DSH Live Voice’s plugin interface is translated into the following languages. This refers to the visible plugin UI—not speech-recognition or text-to-speech language support.
- 🇺🇸 English
- 🇧🇷 Portuguese (Brazil)
- 🇪🇸 Spanish
- 🇫🇷 French
- 🇮🇳 Hindi
- 🇨🇳 Chinese
📚 Documentation
Detailed guides for deep-diving into engines and configurations:
- ⚙️ Configuration & Conversation Flow Guide — Settings overview, speaker vs. headphone modes, sequence diagrams, silence delays, and external engine setup.
- 🧠 Choosing a Speech Recognition Engine — Comparison between Browser STT, Qwen3 ASR, and Whisper, with RAM footprints and OS compatibility.
- 📖 The Story Behind the Project — Why this project was built and the human story behind coordinating voice.
🎯 Which Speech Engine Should I Use?
| Scenario | Recommendation | RAM | Why | | --- | --- | --- | --- | | 🇧🇷 Portuguese on macOS | Qwen3 ASR (HTTP API) | ~3 GB | Best accuracy in daily maintainer use. Whisper is second choice. | | 🇺🇸 English on macOS | Browser SpeechRecognition | ~0 GB | Built-in macOS/browser API. Fast, zero extra RAM. | | 🪟 Windows | Qwen3 ASR or Whisper HTTP | ~2–3 GB | Recommended starting point; Windows browser STT varies. | | 🌐 Multilingual / Other | Whisper HTTP (auto) | ~2 GB | Automatic language detection across dozens of languages. |
👉 For model requirements and server setup, see Choosing a Speech Engine.
✨ Features at a Glance
- 🎙️ Voice Typing: Speak directly into the DSH composer with live interim transcription.
- 👐 Hands-Free Conversation: Continuous dialogue that stays active across chat sessions.
- ❓ Spoken Structured Questions: Narrates DSH prompt questions and submits your spoken answer.
- ⌨️ Hold-to-Talk (Push-to-Talk): Hold
Controlanywhere on the page to speak; release to send. - 🎧 Acoustic Mode Isolation: Gated listening for speakers (no echo) and open-mic interruption for headphones.
- 🗣️ Spoken Commands: Control the chat using phrases like "send", "mute", "clear", and "stop speaking".
- 🧹 Smart Code Filtering: Automatically skips or summarizes large code blocks instead of reading syntax out loud.
- 🏠 Local-First & Private: Audio runs locally on your machine (via Browser APIs, Apple MLX, or whisper.cpp); no external voice telemetry.
- 🌐 Remote-Ready Host Audio: Qwen and macOS Say synthesize on the DSH host, then DSH delivers compact audio to your browser—so playback works over remote and LAN connections.
🧑💻 Coming Soon: Meeting Mode
Meeting Mode is a planned differentiator for collaborative coding conversations. It will keep two independent live transcription streams in the DSH composer: your microphone as “Me:”, and meeting participants from an explicitly shared screen/system-audio stream as “Them:”. This creates an editable, real-time record of a code review or technical discussion, so you can manually ask DSH a question with the meeting context already in the composer.
It will never automatically send the transcript or use meeting audio for voice commands. Sharing system audio will always require explicit browser permission and depends on browser and operating-system support.
🤝 Acknowledgments & Community
Listening and speaking should work together. A heartfelt thank you to GooDAnDReaDY for dsh-voice and Alan2Z for dsh-speak, which inspired this unified coordinator. Read the full story.
I use DSH Live Voice for at least 8 hours every day. Feedback, ideas, and contributions are welcome:
