# McSpeak Multi-engine text-to-speech MCP server exposed via FastMCP 3.0 Streamable HTTP. Engines: Kokoro (ONNX), Piper (Wyoming/Docker), Orpheus (llama-server + SNAC). ## Build & Run ```bash make up # build + start (docker compose) make logs # follow logs make restart # restart containers make status # show running containers + health ``` ## Entry & Exit Tones (Beep System) Queued speech playback (`speak()`) is bookended by short alert tones. `generate_audio()` is unaffected (file-only, no playback). ### Tone Positions | Position | When | Purpose | |----------|------|---------| | **Entry tone** | Immediately when `speak()` is called (before synthesis) | "I heard you" acknowledgement — covers synthesis latency | | **Exit tone** | After speech, queue empty | "Over and out" — channel clear | | **Standby tone** | After speech, more queued | "Standby" — more messages coming | | **Cancel tone** | After `cancel_speech()` or MCP cancellation kills playback | "Nevermind" — speech was aborted | ### Available Tones | Name | Frequency | Duration | Inspired by | |------|-----------|----------|-------------| | `chirp` | 1800 Hz | ~144 ms | Nextel iDEN Talk Permit Tone (TPT) — the 24/24/24/24/48 ms on/off pattern | | `apollo` | 2525 Hz | 250 ms | NASA quindar intro (key-up) tone used during Apollo missions | | `roger` | 1400-1000 Hz | ~100 ms | Classic CB radio descending two-tone roger beep | | `quindar-out` | 2475 Hz | 250 ms | NASA quindar unkey tone (distinct frequency from intro) | | `standby` | 1000-1400 Hz | ~60 ms | Ascending blip — inverse of roger, signals "more coming" | | `scratch` | 2000-300 Hz sweep + noise | ~120 ms | Vinyl record scratch — needle yanked off the platter | | `reverse-roger` | 1000-1400 Hz | ~100 ms | Ascending two-tone — mathematical inverse of roger beep | **Pink Floyd voice kit (current defaults).** The bookends now follow a telephone theme: | Name | Used as | Inspired by | |------|---------|-------------| | `heartbeat` | speak entry | "Speak to Me" heartbeat (DSotM opener) woven with warm MF telephone tones | | `soft-chord` | speak exit | Mellow resolving A3+E4 dyad, pure sines, ear-friendly on repetition | | `mf-dial` | listen start | The **genuine Young Lust R1 MF operator dial** — KP, 0-4-4-1-8-3-1, ST (the 44 is the UK country code), real Bell-System MF pairs | | `machine` | listen end | "Welcome to the Machine" pulsing VCS3 throb — a warm "connected / got it" | | `mf-listen` / `mf-done` | listen (alt) | Gentler stacked-fifths swell / resolving D-major (previous listen defaults) | | `call-waiting` | over ongoing speech | Brief C6 double-blip mixed over the current message when another project queues | There are also softer "natural" tones (`bell-soft`, `chime-tube`, `water-drop`, `soft-pulse`, `hmm-up`, …) in `tones.py` for custom use. ### Configuration ```env TTS_ENTRY_TONE=heartbeat # before speech (heartbeat, chirp, apollo, none, or /path/to/custom.wav) TTS_EXIT_TONE=soft-chord # after speech, queue empty (soft-chord, roger, quindar-out, none, or path) TTS_CANCEL_TONE=scratch # on cancel (scratch, reverse-roger, none, or /path/to/custom.wav) TTS_CALL_WAITING_TONE=call-waiting # mixed over current speech when a DIFFERENT project queues (once/turn) TTS_LISTEN_START_TONE=mf-dial # played once the mic is live (see listen()) TTS_LISTEN_END_TONE=machine # played after recording stops TTS_SHUTDOWN_TIMEOUT=30 # max seconds to wait for current speech on container stop ``` The standby tone is always the built-in ascending blip. It plays instead of the exit tone when more items are queued. ## Secretary Announcements & Call-Waiting When several projects speak concurrently, the queue behaves like a **secretary**: each message plays whole and in order (never interleaved — see the streaming-utterance model below), and a project is announced by name when the speaker changes or returns after a lull. - The announcement is a short `"."` preamble synthesized in a **reserved secretary voice** (`TTS_SECRETARY_VOICE`, default `bf_emma`) — always via Kokoro so it sounds identical regardless of the speaking engine. That voice is excluded from the project auto-assignment pool so no project ever sounds like the secretary. - The play-or-skip decision is made at **play time** in the consumer (`queue.py:_should_announce`), not enqueue time, because urgent reordering means the real speaker order isn't final until then. - **Call-waiting**: when a *different* project's message joins the queue while one is playing, a brief `call-waiting` blip is mixed over the current audio (a second `pw-play` stream — PipeWire mixes it). Fires **at most once per playing turn** (`_call_waiting_fired`, reset when a new utterance starts) so a burst of queued messages never spams the listener. ```env TTS_ANNOUNCE_MODE=secretary # secretary | always | off (legacy TTS_ANNOUNCE_PROJECT=true → always) TTS_REINTRODUCE_AFTER_SECONDS=120 # same project after this much silence gets re-introduced TTS_SECRETARY_VOICE=bf_emma # reserved; excluded from project auto-assignment ``` ## listen() — Voice Conversations `listen()` captures the host mic (`pw-record`), transcribes via Parakeet on the gpu.supported.systems gateway, and returns the text. Pair it with `speak()` for turn-taking: speak a question (let it finish), then `listen()` for the reply — **sequentially, never in parallel**, or the mic records the TTS. - **Defaults are conversation-first**: `wait_for_silence=True` (stop when the person stops), `duration_seconds=30` cap, `vad_aggressiveness=3`, `silence_threshold_ms=2200`. The aggressiveness/threshold defaults were tuned live to stop brief background transients from ending the turn before the real reply. - **No first-word clip**: the "go" tone is played *after* the mic is live (a `warmup_ms` lead, default 150ms). The beep bleeds harmlessly into the head of the recording — VAD treats a pure tone as non-speech and Parakeet ignores it. See `audio.py:record_audio_until_silence`. - Empty transcription (`text == ""`) or a gateway timeout means re-prompt rather than proceed; the recording is saved under `/tmp/mcspeak/` and can be retried with `transcribe()`. Tones are generated programmatically at startup (48kHz, 16-bit PCM, -3 dB headroom) in `tones.py` using numpy. No bundled audio assets. ## Voice Identity (Project-Aware Voices) When multiple Claude Code sessions connect simultaneously, voice identity gives each project a distinct voice via round-robin assignment from a curated English voice pool. ### How It Works 1. Client calls `speak()` or `generate_audio()` without specifying `voice=` 2. Project is identified via the `project` tool parameter, or falls back to MCP Roots (`list_roots()` with 2s timeout) 3. The next unused voice is assigned from the interleaved pool (alternating gender and accent for maximum contrast) 4. Assignment is persisted to `/data/voice-assignments.json` — survives server restarts Explicit `voice=` parameter always overrides auto-assignment. Voice pools are cached for 5 minutes (picks up blacklist/engine changes). **Note:** MCP Roots require stateful Streamable HTTP. With `stateless_http=True` (current default), roots will timeout — the `project` parameter is the primary identification method. ### Configuration ```env TTS_VOICE_IDENTITY=true # Enable project-aware voice assignment TTS_VOICE_IDENTITY_PREFIXES=af_,am_,bf_,bm_,ef_,em_ # English voice prefixes TTS_VOICE_IDENTITY_EXCLUDE=af_nicole # Available explicitly, excluded from auto-assign (whispery) TTS_VOICE_IDENTITY_FILE=/data/voice-assignments.json # Persist across restarts TTS_ANNOUNCE_PROJECT=false # Prefix speech with project name ``` ### Pool Interleave Order Voices are interleaved for perceptual diversity: American female → British male → European female → American male → British female → European male. First 6 projects get maximally distinct voices. ### Files - `voice_identity.py` — Pool filtering, interleaving, round-robin assignment, JSON persistence ## Media Ducking When `speak()` is called, external audio streams (Firefox, Spotify, etc.) are automatically faded down via PulseAudio before speech begins, then faded back up after the exit tone. This creates a radio-broadcast-interruption effect where the entry tone crossfades over the fading media. ### Audio Timeline ``` 0ms ─ Media starts fading (vol 100%) 150ms ─ Entry tone starts! Media at ~70% ← crossfade overlap 294ms ─ Entry tone ends, media at ~40% 500ms ─ Media at 0% (ducked) ─ [synthesis + playback — media stays silent] ─ Exit tone plays (roger/standby) ─ Media fades back in over 1000ms ``` The duck fires in `speak()` (before synthesis), the unduck fires in the consumer (after exit/cancel tone). Media stays ducked across multiple queued items — only unducks when the queue is empty or on cancel/shutdown. ### Configuration ```env TTS_DUCK_MEDIA=true # Enable/disable (on by default) TTS_DUCK_FADE_OUT_MS=500 # Fade-out duration (ms) TTS_DUCK_FADE_IN_MS=1000 # Fade-in duration (slower = natural) ``` ### Docker Requirements The PulseAudio compatibility socket must be mounted in the container: ```yaml volumes: - /run/user/1000/pulse:/run/user/1000/pulse ``` If the socket is missing or pactl fails, ducking silently no-ops — TTS still works normally. **Why pactl, not pulsectl-asyncio?** PipeWire's PulseAudio compat layer silently drops `sink_input_volume_set` operations from pulsectl's native protocol connection, while `pactl` (using libpulse C library) works reliably. The `pulseaudio-utils` package is installed in the container for this reason. ### Non-fatal Design All pactl errors are caught and logged. If PulseAudio is unavailable (no socket, wrong permissions, pactl not installed), duck/unduck become no-ops. This ensures TTS never breaks due to ducking failures. ### Files - `media_duck.py` — `MediaDucker` class (async pactl subprocess volume control with stepped fades) ## Architecture - `server.py` — FastMCP lifespan, tool definitions, engine setup - `queue.py` — Producer-consumer queue of streaming **utterances** (one per `speak()`): priority tiers, secretary announcements + reserved voice, call-waiting, outcome tracking. Synthesized WAVs are reaped after playback (no /tmp leak). - `tones.py` — Tone WAV generator (speak + listen bookends, call-waiting, natural set) - `media_duck.py` — Async PulseAudio volume control for media ducking - `audio.py` — WAV writing and `pw-play` async wrapper - `settings.py` — Pydantic settings from env vars (prefix: `TTS_`) - `engines/` — TTSEngine implementations (kokoro, piper, orpheus) ## `speak()` Progress Lifecycle `speak()` blocks until playback finishes, reporting progress throughout. **Clients should call it in parallel with other tools** rather than blocking on it alone — the return value is informational (speech_id, duration) and never needed for subsequent reasoning. ### Progress notifications A background ticker emits MCP progress notifications every ~0.5 seconds during playback. Each notification includes `progress` (0-100), `total` (100), and a `message` string. MCP-aware clients render this as a live progress bar. | Progress | Phase | Message | |----------|-------|---------| | 5% | Entry tone played — audible "I heard you" | `Entry tone` | | 30% | Synthesis complete | `Synthesizing...` | | 35% | Enqueued for playback | `Playing audio — you can continue working` | | 35-99% | Playing — progress tracks elapsed time vs expected duration | `Playing audio — you can continue working` | | 100% | Playback finished | `Playback complete` | The entry tone fires before synthesis, covering the 1-3s latency gap. The progress message intentionally says "you can continue working" to signal to LLM clients that they don't need to wait. ### Concurrency guidance `speak()` blocks for 10-60s for typical text, but audio plays through physical speakers — the return value doesn't feed into subsequent reasoning. The server instructions and tool docstring explicitly tell clients to **call speak in parallel with other tools** in the same message. Tool annotations provide structured signals: - `openWorldHint=True` — this tool interacts with the physical world (speakers) - `idempotentHint=True` — safe to retry without side effects beyond replaying audio `speech_status(speech_id)` is available for checking outcomes after the fact. Possible statuses: `completed`, `playing`, `queued`, `unknown`. Outcomes are stored in a bounded ring (last 100 items) — old entries are evicted automatically. ## Chunked Synthesis (Pipelined Playback) Long texts (>= 20 words with sentence boundaries) are automatically split into sentence-sized chunks and pipelined — synthesis of chunk N+1 overlaps playback of chunk N. ``` Single-shot: [---12s synthesis---][----------48s playback----------] First audio at T+12s Chunked: [~1s synth][play 1][~1s synth][play 2][~1s synth][play 3]... First audio at T+1s ``` ### Why no gaps Kokoro synthesizes ~4x realtime on CPU, so synthesis always outpaces playback. `pw-play` adds ~10ms startup per chunk — imperceptible. Slower engines (Orpheus) may have brief gaps but still beat the full-synthesis-first approach. ### Short text bypass Texts under 20 words or without sentence boundaries (`.!?` followed by whitespace) take the single-shot path — zero overhead, identical to pre-chunking behavior. ### Tone behavior & per-message coherence Each `speak()` call is **one queue utterance** (not one queue item per chunk). The utterance reserves its ordering slot up front and streams its synthesized chunks in through an internal channel closed by an `_END` sentinel; the consumer stays locked to it from entry tone to exit tone. So when several projects speak at once, a message plays **whole and in order** — chunks never interleave with another project's audio (the old per-chunk `suppress_exit_tone`/`_WorkItem` model is gone). - Entry tone plays once at the start (before first chunk synthesis) - Exit/standby tone plays once, at the very end of the utterance - Pipelining is preserved: synthesis of chunk N+1 overlaps playback of chunk N, but nothing else can be pulled until this utterance finishes ### Cancellation in chunked mode - **Explicit `cancel_speech(speech_id)`** — one speech_id now covers the whole utterance; cancelling it flags an abort (synchronously), reaps un-played chunk WAVs, and plays the cancel tone. - **MCP disconnect** — the `speak()` handler is cancelled; its `finally` closes the utterance channel so the consumer drains the chunks it already has and finishes cleanly. ### Progress lifecycle (chunked) | Progress | Phase | Message | |----------|-------|---------| | 5% | Entry tone played | `Entry tone` | | 5-30% | Synthesis progress across all chunks | `Synthesizing chunk N/M...` | | 35% | All chunks enqueued | `Playing audio — you can continue working` | | 35-99% | Playing — tracks elapsed time vs total duration | `Playing audio — you can continue working` | | 100% | Final chunk playback finished | `Playback complete` | ### Files - `server.py` — `split_text()`, unified `_speak()` (short + chunked share one path), status-aware `_await_with_progress()` - `queue.py` — `_Utterance` (streaming chunk channel + `_END` sentinel), `create_utterance()`, `_play_utterance()`, `_should_announce()`, call-waiting ## Cancellation Speech can be cancelled two ways: 1. **Explicit `cancel_speech(speech_id)`** — kills pw-play immediately, plays the cancel tone, and the consumer moves to the next item. 2. **MCP cancellation** — if a client disconnects or sends `notifications/cancelled`, the `speak()` handler is cancelled but playback continues. The consumer finishes the current audio naturally. This is intentional — audio is already synthesized and playing, so cutting it mid-sentence would be jarring. When a currently-playing item is explicitly cancelled, pw-play is killed immediately and the cancel tone plays. When a queued item is cancelled, it's removed from the queue without ever playing. The consumer continues to the next item in both cases. ## Graceful Shutdown On `docker compose down` or `make restart`, the server lets the currently-playing speech finish before stopping — no more mid-sentence cutoffs. **How it works:** SIGTERM triggers a two-phase shutdown that shares Docker's wall-clock budget: 1. **Phase 1 — Handler drain (3s fixed):** Uvicorn cancels in-flight `speak()` handlers. The handlers re-raise `CancelledError` without touching the consumer — audio keeps playing. 2. **Phase 2 — Queue drain (`shutdown_timeout`, default 30s):** Lifespan finalizer calls `queue.stop()`, which waits for the consumer to finish the current audio. If the timeout expires, it force-cancels. Docker's `stop_grace_period` must exceed the total: `3s + shutdown_timeout + 5s safety margin` (default 38s). If the max expected audio length exceeds `shutdown_timeout`, increase both `TTS_SHUTDOWN_TIMEOUT` and `stop_grace_period`. **Entry tone timing:** The Nextel chirp fires immediately when `speak()` is called (before synthesis), acting as an audible "I heard you" that covers the 1-3s synthesis latency. Exit/standby tones still play from the consumer (they depend on queue state after playback). ## Key Design Decisions - Speech queue is serialized (one playback at a time) but synthesis is parallel; each `speak()` is one **streaming utterance** so concurrent projects never interleave - Secretary behavior: a project is announced by name on speaker-change / after a lull, in a **reserved voice** excluded from the project pool; a different project queuing mid-playback fires a **once-per-turn** call-waiting blip mixed over the current audio - Synthesized WAVs are reaped after playback (and on cancel/shutdown) so `/tmp/mcspeak` doesn't grow unbounded - `speak()` blocks until playback finishes with **status-aware** progress — a queued item reports "waiting in line", not a false "playing" percentage - Progress uses a background ticker task, NOT `asyncio.wait_for` polling (see below) - Entry tone is awaited in `speak()` before synthesis — covers latency gap - Explicit `cancel_speech()` kills pw-play + plays cancel tone; MCP disconnect lets playback finish - Consumer directly awaits `play_audio()`; `cancel()` targets the consumer task with `_item_cancelled` flag - `speak()`'s CancelledError handler does NOT cancel the consumer — only explicit cancel does - Graceful shutdown is two-phase: 3s handler drain (fixed) + `shutdown_timeout` queue drain, both sequential within Docker's `stop_grace_period` - Tones are non-fatal: if `pw-play` fails on a tone, speech still plays - Orpheus uses llama-server (not Ollama) for 15x throughput via continuous batching - SNAC decoder is lazy-loaded on first Orpheus call to reduce idle memory ### Python 3.13 asyncio.wait_for pitfall **Do NOT use `asyncio.wait_for(future, timeout)` in a polling loop to track progress.** In Python 3.13, `wait_for` cancels its inner task on timeout. When the inner task is awaiting the same `asyncio.Future` that the consumer will resolve, repeated cancel/re-await cycles cause a stale `CancelledError` to propagate to the consumer task — killing pw-play mid-playback. Instead, use a background `asyncio.create_task` ticker for progress and directly `await` the future for completion.