speak() now blocks until playback finishes, reporting progress via SSE (5% entry tone → 30% synthesis → 35-99% playing → 100% done). Entry tone fires immediately on call to cover synthesis latency. Queue shutdown waits for current speech to finish (configurable timeout, default 30s) before draining pending items — no more mid-sentence cutoffs on container restart. Cancellation via cancel_speech() tool or MCP notifications/cancelled kills pw-play and plays a vinyl scratch tone. Consumer continues to next item after cancel. Progress tracking uses a background ticker task instead of asyncio.wait_for polling — the latter causes stale CancelledError propagation to the consumer under Python 3.13.
8.0 KiB
TTS MCP Server
Multi-engine text-to-speech server exposed via FastMCP 3.0 Streamable HTTP. Engines: Kokoro (ONNX), Piper (Wyoming/Docker), Orpheus (llama-server + SNAC).
Build & Run
make up # build + start (docker compose)
make logs # follow logs
make restart # restart containers
make status # show running containers + health
Entry & Exit Tones (Beep System)
Queued speech playback (speak()) is bookended by short alert tones. generate_audio() is unaffected (file-only, no playback).
Tone Positions
| Position | When | Purpose |
|---|---|---|
| Entry tone | Immediately when speak() is called (before synthesis) |
"I heard you" acknowledgement — covers synthesis latency |
| Exit tone | After speech, queue empty | "Over and out" — channel clear |
| Standby tone | After speech, more queued | "Standby" — more messages coming |
| Cancel tone | After cancel_speech() or MCP cancellation kills playback |
"Nevermind" — speech was aborted |
Available Tones
| Name | Frequency | Duration | Inspired by |
|---|---|---|---|
chirp |
1800 Hz | ~144 ms | Nextel iDEN Talk Permit Tone (TPT) — the 24/24/24/24/48 ms on/off pattern |
apollo |
2525 Hz | 250 ms | NASA quindar intro (key-up) tone used during Apollo missions |
roger |
1400-1000 Hz | ~100 ms | Classic CB radio descending two-tone roger beep |
quindar-out |
2475 Hz | 250 ms | NASA quindar unkey tone (distinct frequency from intro) |
standby |
1000-1400 Hz | ~60 ms | Ascending blip — inverse of roger, signals "more coming" |
scratch |
2000-300 Hz sweep + noise | ~120 ms | Vinyl record scratch — needle yanked off the platter |
reverse-roger |
1000-1400 Hz | ~100 ms | Ascending two-tone — mathematical inverse of roger beep |
Configuration
TTS_ENTRY_TONE=chirp # before speech (chirp, apollo, none, or /path/to/custom.wav)
TTS_EXIT_TONE=roger # after speech, queue empty (roger, quindar-out, none, or path)
TTS_CANCEL_TONE=scratch # on cancel (scratch, reverse-roger, none, or /path/to/custom.wav)
TTS_SHUTDOWN_TIMEOUT=30 # max seconds to wait for current speech on container stop
The standby tone is always the built-in ascending blip. It plays instead of the exit tone when more items are queued.
Tones are generated programmatically at startup (48kHz, 16-bit PCM, -3 dB headroom) in tones.py using numpy. No bundled audio assets.
Voice Identity (Project-Aware Voices)
When multiple Claude Code sessions connect simultaneously, voice identity gives each project a distinct voice via round-robin assignment from a curated English voice pool.
How It Works
- Client calls
speak()orgenerate_audio()without specifyingvoice= - Project is identified via the
projecttool parameter, or falls back to MCP Roots (list_roots()with 2s timeout) - The next unused voice is assigned from the interleaved pool (alternating gender and accent for maximum contrast)
- Assignment is persisted to
/data/voice-assignments.json— survives server restarts
Explicit voice= parameter always overrides auto-assignment. Voice pools are cached for 5 minutes (picks up blacklist/engine changes).
Note: MCP Roots require stateful Streamable HTTP. With stateless_http=True (current default), roots will timeout — the project parameter is the primary identification method.
Configuration
TTS_VOICE_IDENTITY=true # Enable project-aware voice assignment
TTS_VOICE_IDENTITY_PREFIXES=af_,am_,bf_,bm_,ef_,em_ # English voice prefixes
TTS_VOICE_IDENTITY_EXCLUDE=af_nicole # Available explicitly, excluded from auto-assign (whispery)
TTS_VOICE_IDENTITY_FILE=/data/voice-assignments.json # Persist across restarts
TTS_ANNOUNCE_PROJECT=false # Prefix speech with project name
Pool Interleave Order
Voices are interleaved for perceptual diversity: American female → British male → European female → American male → British female → European male. First 6 projects get maximally distinct voices.
Files
voice_identity.py— Pool filtering, interleaving, round-robin assignment, JSON persistence
Architecture
server.py— FastMCP lifespan, tool definitions, engine setupqueue.py— Producer-consumer speech queue with priority tiers and outcome trackingtones.py— Tone WAV generator (entry/exit/standby)audio.py— WAV writing andpw-playasync wrappersettings.py— Pydantic settings from env vars (prefix:TTS_)engines/— TTSEngine implementations (kokoro, piper, orpheus)
speak() Progress Lifecycle
speak() blocks until playback finishes, reporting progress throughout:
| Progress | Phase |
|---|---|
| 5% | Entry tone played — audible "I heard you" |
| 30% | Synthesis complete |
| 35% | Enqueued for playback |
| 35-99% | Playing — progress tracks elapsed time vs expected duration |
| 100% | Playback finished |
MCP-aware clients see a live progress bar. The entry tone fires before synthesis, covering the 1-3s latency gap.
speech_status(speech_id) is still available for checking outcomes after the fact. Possible statuses: completed, playing, queued, unknown.
Outcomes are stored in a bounded ring (last 100 items) — old entries are evicted automatically.
Cancellation
Speech can be cancelled two ways:
- MCP cancellation — the client sends
notifications/cancelledfor an in-flightspeak()call. FastMCP throwsCancelledErrorinto the tool, which triggersqueue.cancel(speech_id). - Explicit
cancel_speech(speech_id)— a separate tool that cancels any queued or playing item.
When a currently-playing item is cancelled, pw-play is killed immediately and the cancel tone plays. When a queued item is cancelled, it's removed from the queue without ever playing. The consumer continues to the next item in both cases.
Graceful Shutdown
On docker compose down or make restart, the server lets the currently-playing speech finish before stopping — no more mid-sentence cutoffs.
How it works: queue.stop() sets a flag and waits up to shutdown_timeout seconds for the consumer to finish the current item. If the timeout expires, it falls back to hard cancel. Docker's stop_grace_period: 35s gives the 30s shutdown timeout room to complete before SIGKILL.
Entry tone timing: The Nextel chirp fires immediately when speak() is called (before synthesis), acting as an audible "I heard you" that covers the 1-3s synthesis latency. Exit/standby tones still play from the consumer (they depend on queue state after playback).
Key Design Decisions
- Speech queue is serialized (one playback at a time) but synthesis is parallel
speak()blocks until playback finishes with live progress (5% → 30% → 35-99% → 100%)- Progress uses a background ticker task, NOT
asyncio.wait_forpolling (see below) - Entry tone is awaited in
speak()before synthesis — covers latency gap - Cancellation via
cancel_speech()or MCPnotifications/cancelledkills pw-play + plays cancel tone - Consumer directly awaits
play_audio(); cancel() targets the consumer task with_item_cancelledflag - Graceful shutdown waits for current speech, then drains pending items with error outcomes
- Tones are non-fatal: if
pw-playfails on a tone, speech still plays - Orpheus uses llama-server (not Ollama) for 15x throughput via continuous batching
- SNAC decoder is lazy-loaded on first Orpheus call to reduce idle memory
Python 3.13 asyncio.wait_for pitfall
Do NOT use asyncio.wait_for(future, timeout) in a polling loop to track progress. In Python 3.13, wait_for cancels its inner task on timeout. When the inner task is awaiting the same asyncio.Future that the consumer will resolve, repeated cancel/re-await cycles cause a stale CancelledError to propagate to the consumer task — killing pw-play mid-playback. Instead, use a background asyncio.create_task ticker for progress and directly await the future for completion.