McSpeak
Multi-engine text-to-speech server exposed as MCP tools via FastMCP 3.0 Streamable HTTP. Plays audio through the host's speakers using PipeWire. Built for giving AI agents a voice -- literally.
Quick Start
McSpeak runs a Streamable HTTP MCP server on :8371. Start the server, then point any MCP client at it.
1. Get the Kokoro model files (~340 MB, one-time):
mkdir -p models/kokoro && cd models/kokoro
curl -LO https://github.com/thewh1teagle/kokoro-onnx/releases/download/model-files-v1.0/kokoro-v1.0.onnx
curl -LO https://github.com/thewh1teagle/kokoro-onnx/releases/download/model-files-v1.0/voices-v1.0.bin
cd -
The defaults look for models/kokoro/kokoro-v1.0.onnx and models/kokoro/voices-v1.0.bin relative to the working directory. Override with TTS_KOKORO_MODEL / TTS_KOKORO_VOICES if you put them elsewhere.
2. Start the server:
# Run directly from PyPI (no install needed)
uvx mcspeak
# Or install and run
pip install mcspeak
mcspeak
Wait for McSpeak ready on 0.0.0.0:8371 in the logs. Piper and Orpheus will report unhealthy unless their backends are running -- that's expected; Kokoro alone is enough to start.
3. Wire it into Claude Code (HTTP transport, server must be running):
claude mcp add --transport http mcspeak http://127.0.0.1:8371/mcp
mcspeak ships as a Streamable HTTP server, not a stdio one, so --transport http is required. The older form claude mcp add mcspeak -- uvx mcspeak registers stdio and will fail to connect.
For multi-engine Docker deployment with Piper and GPU-accelerated Orpheus, see Docker Setup below.
MCP Tools
| Tool | What it does |
|---|---|
speak |
Synthesize text and play through speakers. Blocks with live progress. |
generate_audio |
Synthesize to WAV file without playing. Bypasses the queue. |
list_voices |
List available voices for an engine. |
list_engines |
Show all engines and their health status. |
speech_status |
Check status of a queued/playing speech item. |
cancel_speech |
Kill playback immediately. Plays a cancel tone. |
There's also an audio://recent resource listing the last 20 generated WAV files.
Engines
| Engine | Backend | Speed | Voices | Notes |
|---|---|---|---|---|
| kokoro | ONNX (local CPU) | ~4x realtime | ~50 | Default. Fast, good quality. |
| piper | Wyoming protocol (Docker) | ~2x realtime | Depends on model | Connects to any Wyoming-compatible TTS server. |
| orpheus | llama-server + SNAC | ~1x realtime | 6 | LLM-based. Supports <laugh>, <sigh>, emotion tags. Needs GPU. |
Features
Radio-Style Alert Tones
Speech is bookended by short tones inspired by two-way radio systems. An entry tone fires immediately when speak() is called (before synthesis starts), acting as an audible "I heard you" that covers the 1-3s synthesis latency. An exit tone plays when the queue empties.
Available tones: chirp (Nextel iDEN), apollo (NASA quindar key-up), roger (CB radio descending beep), quindar-out (NASA unkey), scratch (vinyl record scratch for cancellation), and more. All generated programmatically at startup -- no bundled audio files.
Voice Identity
When multiple projects connect simultaneously, each one automatically gets a distinct voice via round-robin assignment from a curated pool. Voices are interleaved across gender and accent (American, British, European) for maximum contrast. Assignments persist across restarts.
Pass project="my-project" in speak() calls, or let it auto-detect from MCP Roots.
Media Ducking
External audio (YouTube, Spotify, etc.) automatically fades down before speech and fades back in after the exit tone, like a radio broadcast interruption. Uses pactl under the hood. Fails silently if PulseAudio isn't available.
Chunked Synthesis Pipeline
Long texts (20+ words with sentence boundaries) are split into chunks and pipelined -- synthesis of chunk N+1 overlaps playback of chunk N. First audio plays within ~1-2s instead of waiting for the full text to synthesize.
Graceful Shutdown
docker compose down lets the current speech finish instead of cutting mid-sentence. Two-phase: 3s handler drain, then up to 30s for the consumer to finish. Configurable via TTS_SHUTDOWN_TIMEOUT.
Configuration
All settings use the TTS_ prefix and can be set via environment variables or .env file.
| Variable | Default | Description |
|---|---|---|
TTS_HOST |
0.0.0.0 |
Server bind address |
TTS_PORT |
8371 |
Server port |
TTS_ENTRY_TONE |
chirp |
Tone before speech (chirp, apollo, none, or WAV path) |
TTS_EXIT_TONE |
roger |
Tone after speech (roger, quindar-out, none, or WAV path) |
TTS_CANCEL_TONE |
scratch |
Tone on cancel (scratch, reverse-roger, none, or WAV path) |
TTS_DUCK_MEDIA |
true |
Fade external audio during speech |
TTS_DUCK_FADE_OUT_MS |
500 |
Media fade-out duration (ms) |
TTS_DUCK_FADE_IN_MS |
1000 |
Media fade-in duration (ms) |
TTS_VOICE_IDENTITY |
true |
Auto-assign distinct voices per project |
TTS_VOICE_IDENTITY_FILE |
(empty) | JSON file to persist voice assignments |
TTS_VOICE_IDENTITY_EXCLUDE |
af_nicole |
Voices excluded from auto-assign pool |
TTS_ANNOUNCE_PROJECT |
false |
Prefix speech with project name |
TTS_VOICE_BLACKLIST |
am_adam,af_jessica |
Voices hidden from listings |
TTS_SHUTDOWN_TIMEOUT |
30 |
Max seconds to wait for speech on shutdown |
TTS_KOKORO_MODEL |
models/kokoro/kokoro-v1.0.onnx |
Path to Kokoro ONNX model |
TTS_KOKORO_VOICES |
models/kokoro/voices-v1.0.bin |
Path to Kokoro voice pack |
TTS_PIPER_HOST |
172.26.0.3 |
Piper Wyoming server host |
TTS_PIPER_PORT |
10200 |
Piper Wyoming server port |
TTS_ORPHEUS_URL |
http://127.0.0.1:8081 |
Orpheus llama-server URL |
Docker Setup
The compose file defaults to a kokoro-only stack — no GPU required. Orpheus (GPU-accelerated, llama-server backed) is opt-in via the with-orpheus profile. Piper integrates via its host/port settings; you supply your own Wyoming server.
Prerequisites (base, kokoro-only):
- PipeWire on the host (the container plays audio via the host's
pw-playsocket) - An external Docker network named
caddy(create withdocker network create caddyif it doesn't exist). The privatemcspeak-internalnetwork is auto-created per-stack. - Kokoro ONNX model files in
./models/kokoro/(see Quick Start for the download)
Additional prerequisites for with-orpheus:
- NVIDIA GPU + nvidia-container-toolkit
- Orpheus GGUF model on disk — path goes in
.envasORPHEUS_GGUF_PATH
Bring it up (kokoro-only):
make up # build + start kokoro-only
make logs # follow logs
make status # health check
Bring it up with Orpheus:
cp .env.example .env
# Edit .env: set ORPHEUS_GGUF_PATH to your Orpheus GGUF file
make up-with-orpheus
Once make status shows the container as healthy, register it with Claude Code. The compose file applies a caddy-docker-proxy label exposing the server at mctalkbox.l.supported.systems -- if you're not using that hostname, point at the container's mapped port directly:
# Direct connection (replace with your hostname/port if different)
claude mcp add --transport http mcspeak http://127.0.0.1:8371/mcp
The container mounts the host PipeWire socket (/run/user/1000/pipewire-0) for audio playback and the PulseAudio compat socket (/run/user/1000/pulse) for media ducking.
stop_grace_period in the compose file is set to 38s (3s drain + 30s shutdown timeout + 5s safety). Adjust if you change TTS_SHUTDOWN_TIMEOUT.
Architecture
src/tts_mcp/
server.py FastMCP lifespan, tool definitions, chunked synthesis
queue.py Producer-consumer speech queue with priority + outcome tracking
tones.py Procedural tone WAV generation (numpy, 48kHz 16-bit PCM)
audio.py WAV writing and async pw-play wrapper
media_duck.py Async pactl volume control for media ducking
voice_identity.py Project-to-voice mapping with round-robin assignment
settings.py Pydantic settings from env vars
engines/
base.py TTSEngine abstract base
kokoro.py Kokoro ONNX engine
piper.py Wyoming protocol client
orpheus.py llama-server + SNAC decoder
See CLAUDE.md for the full technical reference -- design decisions, progress lifecycle, cancellation semantics, and the Python 3.13 asyncio.wait_for pitfall that cost a few hours of debugging.
License
MIT