Stream-restore mute: pw-play streams were being silently restored to 0% volume by PulseAudio's stream-restore module matching the music role key, producing audible-but-clean-exit playback. audio.py now bumps each new pw-play sink-input to 100% as a background task; the new value re-stamps stream-restore on stream end. Entry tone overlap: when speak() arrived during prior playback, its entry tone played immediately from the handler and overlapped the ongoing audio. queue.is_idle() now gates this — tone plays now when idle (preserves latency-hiding intent), else defers to the consumer via _WorkItem.entry_tone, which plays it right before the item's audio. Chunked path defers only on chunk 0. Docker: llama-server gated behind a `with-orpheus` compose profile; mcspeak.depends_on uses required:false. `make up` defaults to kokoro-only with no GPU dependency; `make up-with-orpheus` runs the full stack. Rename dootie-internal network → mcspeak-internal (auto-created per-stack via internal:true). Add 127.0.0.1:8371 port mapping for local MCP clients. README: fix claude mcp add command (was stdio, must be --transport http); document kokoro-only default and the with-orpheus path; add Kokoro model download step. .env.example added so `cp .env.example .env` matches the documented setup.
183 lines
8.8 KiB
Markdown
183 lines
8.8 KiB
Markdown
# McSpeak
|
|
|
|
Multi-engine text-to-speech server exposed as MCP tools via [FastMCP 3.0](https://gofastmcp.com/) Streamable HTTP. Plays audio through the host's speakers using PipeWire. Built for giving AI agents a voice -- literally.
|
|
|
|
[](https://pypi.org/project/mcspeak/)
|
|
[](https://python.org)
|
|
[](LICENSE)
|
|
|
|
## Quick Start
|
|
|
|
McSpeak runs a Streamable HTTP MCP server on `:8371`. Start the server, then point any MCP client at it.
|
|
|
|
**1. Get the Kokoro model files** (~340 MB, one-time):
|
|
|
|
```bash
|
|
mkdir -p models/kokoro && cd models/kokoro
|
|
curl -LO https://github.com/thewh1teagle/kokoro-onnx/releases/download/model-files-v1.0/kokoro-v1.0.onnx
|
|
curl -LO https://github.com/thewh1teagle/kokoro-onnx/releases/download/model-files-v1.0/voices-v1.0.bin
|
|
cd -
|
|
```
|
|
|
|
The defaults look for `models/kokoro/kokoro-v1.0.onnx` and `models/kokoro/voices-v1.0.bin` relative to the working directory. Override with `TTS_KOKORO_MODEL` / `TTS_KOKORO_VOICES` if you put them elsewhere.
|
|
|
|
**2. Start the server:**
|
|
|
|
```bash
|
|
# Run directly from PyPI (no install needed)
|
|
uvx mcspeak
|
|
|
|
# Or install and run
|
|
pip install mcspeak
|
|
mcspeak
|
|
```
|
|
|
|
Wait for `McSpeak ready on 0.0.0.0:8371` in the logs. Piper and Orpheus will report `unhealthy` unless their backends are running -- that's expected; Kokoro alone is enough to start.
|
|
|
|
**3. Wire it into Claude Code** (HTTP transport, server must be running):
|
|
|
|
```bash
|
|
claude mcp add --transport http mcspeak http://127.0.0.1:8371/mcp
|
|
```
|
|
|
|
`mcspeak` ships as a Streamable HTTP server, not a stdio one, so `--transport http` is required. The older form `claude mcp add mcspeak -- uvx mcspeak` registers stdio and will fail to connect.
|
|
|
|
For multi-engine Docker deployment with Piper and GPU-accelerated Orpheus, see [Docker Setup](#docker-setup) below.
|
|
|
|
## MCP Tools
|
|
|
|
| Tool | What it does |
|
|
|------|-------------|
|
|
| `speak` | Synthesize text and play through speakers. Blocks with live progress. |
|
|
| `generate_audio` | Synthesize to WAV file without playing. Bypasses the queue. |
|
|
| `list_voices` | List available voices for an engine. |
|
|
| `list_engines` | Show all engines and their health status. |
|
|
| `speech_status` | Check status of a queued/playing speech item. |
|
|
| `cancel_speech` | Kill playback immediately. Plays a cancel tone. |
|
|
|
|
There's also an `audio://recent` resource listing the last 20 generated WAV files.
|
|
|
|
## Engines
|
|
|
|
| Engine | Backend | Speed | Voices | Notes |
|
|
|--------|---------|-------|--------|-------|
|
|
| **kokoro** | ONNX (local CPU) | ~4x realtime | ~50 | Default. Fast, good quality. |
|
|
| **piper** | Wyoming protocol (Docker) | ~2x realtime | Depends on model | Connects to any Wyoming-compatible TTS server. |
|
|
| **orpheus** | llama-server + SNAC | ~1x realtime | 6 | LLM-based. Supports `<laugh>`, `<sigh>`, emotion tags. Needs GPU. |
|
|
|
|
## Features
|
|
|
|
### Radio-Style Alert Tones
|
|
|
|
Speech is bookended by short tones inspired by two-way radio systems. An entry tone fires immediately when `speak()` is called (before synthesis starts), acting as an audible "I heard you" that covers the 1-3s synthesis latency. An exit tone plays when the queue empties.
|
|
|
|
Available tones: `chirp` (Nextel iDEN), `apollo` (NASA quindar key-up), `roger` (CB radio descending beep), `quindar-out` (NASA unkey), `scratch` (vinyl record scratch for cancellation), and more. All generated programmatically at startup -- no bundled audio files.
|
|
|
|
### Voice Identity
|
|
|
|
When multiple projects connect simultaneously, each one automatically gets a distinct voice via round-robin assignment from a curated pool. Voices are interleaved across gender and accent (American, British, European) for maximum contrast. Assignments persist across restarts.
|
|
|
|
Pass `project="my-project"` in `speak()` calls, or let it auto-detect from MCP Roots.
|
|
|
|
### Media Ducking
|
|
|
|
External audio (YouTube, Spotify, etc.) automatically fades down before speech and fades back in after the exit tone, like a radio broadcast interruption. Uses `pactl` under the hood. Fails silently if PulseAudio isn't available.
|
|
|
|
### Chunked Synthesis Pipeline
|
|
|
|
Long texts (20+ words with sentence boundaries) are split into chunks and pipelined -- synthesis of chunk N+1 overlaps playback of chunk N. First audio plays within ~1-2s instead of waiting for the full text to synthesize.
|
|
|
|
### Graceful Shutdown
|
|
|
|
`docker compose down` lets the current speech finish instead of cutting mid-sentence. Two-phase: 3s handler drain, then up to 30s for the consumer to finish. Configurable via `TTS_SHUTDOWN_TIMEOUT`.
|
|
|
|
## Configuration
|
|
|
|
All settings use the `TTS_` prefix and can be set via environment variables or `.env` file.
|
|
|
|
| Variable | Default | Description |
|
|
|----------|---------|-------------|
|
|
| `TTS_HOST` | `0.0.0.0` | Server bind address |
|
|
| `TTS_PORT` | `8371` | Server port |
|
|
| `TTS_ENTRY_TONE` | `chirp` | Tone before speech (`chirp`, `apollo`, `none`, or WAV path) |
|
|
| `TTS_EXIT_TONE` | `roger` | Tone after speech (`roger`, `quindar-out`, `none`, or WAV path) |
|
|
| `TTS_CANCEL_TONE` | `scratch` | Tone on cancel (`scratch`, `reverse-roger`, `none`, or WAV path) |
|
|
| `TTS_DUCK_MEDIA` | `true` | Fade external audio during speech |
|
|
| `TTS_DUCK_FADE_OUT_MS` | `500` | Media fade-out duration (ms) |
|
|
| `TTS_DUCK_FADE_IN_MS` | `1000` | Media fade-in duration (ms) |
|
|
| `TTS_VOICE_IDENTITY` | `true` | Auto-assign distinct voices per project |
|
|
| `TTS_VOICE_IDENTITY_FILE` | (empty) | JSON file to persist voice assignments |
|
|
| `TTS_VOICE_IDENTITY_EXCLUDE` | `af_nicole` | Voices excluded from auto-assign pool |
|
|
| `TTS_ANNOUNCE_PROJECT` | `false` | Prefix speech with project name |
|
|
| `TTS_VOICE_BLACKLIST` | `am_adam,af_jessica` | Voices hidden from listings |
|
|
| `TTS_SHUTDOWN_TIMEOUT` | `30` | Max seconds to wait for speech on shutdown |
|
|
| `TTS_KOKORO_MODEL` | `models/kokoro/kokoro-v1.0.onnx` | Path to Kokoro ONNX model |
|
|
| `TTS_KOKORO_VOICES` | `models/kokoro/voices-v1.0.bin` | Path to Kokoro voice pack |
|
|
| `TTS_PIPER_HOST` | `172.26.0.3` | Piper Wyoming server host |
|
|
| `TTS_PIPER_PORT` | `10200` | Piper Wyoming server port |
|
|
| `TTS_ORPHEUS_URL` | `http://127.0.0.1:8081` | Orpheus llama-server URL |
|
|
|
|
## Docker Setup
|
|
|
|
The compose file defaults to a **kokoro-only stack** — no GPU required. Orpheus (GPU-accelerated, llama-server backed) is opt-in via the `with-orpheus` profile. Piper integrates via its host/port settings; you supply your own Wyoming server.
|
|
|
|
**Prerequisites (base, kokoro-only):**
|
|
- PipeWire on the host (the container plays audio via the host's `pw-play` socket)
|
|
- An external Docker network named `caddy` (create with `docker network create caddy` if it doesn't exist). The private `mcspeak-internal` network is auto-created per-stack.
|
|
- Kokoro ONNX model files in `./models/kokoro/` (see [Quick Start](#quick-start) for the download)
|
|
|
|
**Additional prerequisites for `with-orpheus`:**
|
|
- NVIDIA GPU + nvidia-container-toolkit
|
|
- Orpheus GGUF model on disk — path goes in `.env` as `ORPHEUS_GGUF_PATH`
|
|
|
|
**Bring it up (kokoro-only):**
|
|
|
|
```bash
|
|
make up # build + start kokoro-only
|
|
make logs # follow logs
|
|
make status # health check
|
|
```
|
|
|
|
**Bring it up with Orpheus:**
|
|
|
|
```bash
|
|
cp .env.example .env
|
|
# Edit .env: set ORPHEUS_GGUF_PATH to your Orpheus GGUF file
|
|
make up-with-orpheus
|
|
```
|
|
|
|
Once `make status` shows the container as healthy, register it with Claude Code. The compose file applies a `caddy-docker-proxy` label exposing the server at `mctalkbox.l.supported.systems` -- if you're not using that hostname, point at the container's mapped port directly:
|
|
|
|
```bash
|
|
# Direct connection (replace with your hostname/port if different)
|
|
claude mcp add --transport http mcspeak http://127.0.0.1:8371/mcp
|
|
```
|
|
|
|
The container mounts the host PipeWire socket (`/run/user/1000/pipewire-0`) for audio playback and the PulseAudio compat socket (`/run/user/1000/pulse`) for media ducking.
|
|
|
|
`stop_grace_period` in the compose file is set to 38s (3s drain + 30s shutdown timeout + 5s safety). Adjust if you change `TTS_SHUTDOWN_TIMEOUT`.
|
|
|
|
## Architecture
|
|
|
|
```
|
|
src/tts_mcp/
|
|
server.py FastMCP lifespan, tool definitions, chunked synthesis
|
|
queue.py Producer-consumer speech queue with priority + outcome tracking
|
|
tones.py Procedural tone WAV generation (numpy, 48kHz 16-bit PCM)
|
|
audio.py WAV writing and async pw-play wrapper
|
|
media_duck.py Async pactl volume control for media ducking
|
|
voice_identity.py Project-to-voice mapping with round-robin assignment
|
|
settings.py Pydantic settings from env vars
|
|
engines/
|
|
base.py TTSEngine abstract base
|
|
kokoro.py Kokoro ONNX engine
|
|
piper.py Wyoming protocol client
|
|
orpheus.py llama-server + SNAC decoder
|
|
```
|
|
|
|
See [CLAUDE.md](CLAUDE.md) for the full technical reference -- design decisions, progress lifecycle, cancellation semantics, and the Python 3.13 `asyncio.wait_for` pitfall that cost a few hours of debugging.
|
|
|
|
## License
|
|
|
|
MIT
|