Queue-aware exit tones: descending roger beep when queue empties
("over"), ascending standby blip when more items are queued ("standby,
more coming"). Includes quindar-out (2475 Hz) as Apollo-themed
alternative. Configurable via TTS_EXIT_TONE env var.
Also adds CLAUDE.md documenting the full tone system.
2.6 KiB
2.6 KiB
TTS MCP Server
Multi-engine text-to-speech server exposed via FastMCP 3.0 Streamable HTTP. Engines: Kokoro (ONNX), Piper (Wyoming/Docker), Orpheus (llama-server + SNAC).
Build & Run
make up # build + start (docker compose)
make logs # follow logs
make restart # restart containers
make status # show running containers + health
Entry & Exit Tones (Beep System)
Queued speech playback (speak()) is bookended by short alert tones. generate_audio() is unaffected (file-only, no playback).
Tone Positions
| Position | When | Purpose |
|---|---|---|
| Entry tone | Before speech starts | "Incoming transmission" alert |
| Exit tone | After speech, queue empty | "Over and out" — channel clear |
| Standby tone | After speech, more queued | "Standby" — more messages coming |
Available Tones
| Name | Frequency | Duration | Inspired by |
|---|---|---|---|
chirp |
1800 Hz | ~144 ms | Nextel iDEN Talk Permit Tone (TPT) — the 24/24/24/24/48 ms on/off pattern |
apollo |
2525 Hz | 250 ms | NASA quindar intro (key-up) tone used during Apollo missions |
roger |
1400-1000 Hz | ~100 ms | Classic CB radio descending two-tone roger beep |
quindar-out |
2475 Hz | 250 ms | NASA quindar unkey tone (distinct frequency from intro) |
standby |
1000-1400 Hz | ~60 ms | Ascending blip — inverse of roger, signals "more coming" |
Configuration
TTS_ENTRY_TONE=chirp # before speech (chirp, apollo, none, or /path/to/custom.wav)
TTS_EXIT_TONE=roger # after speech, queue empty (roger, quindar-out, none, or path)
The standby tone is always the built-in ascending blip. It plays instead of the exit tone when more items are queued.
Tones are generated programmatically at startup (48kHz, 16-bit PCM, -3 dB headroom) in tones.py using numpy. No bundled audio assets.
Architecture
server.py— FastMCP lifespan, tool definitions, engine setupqueue.py— Producer-consumer speech queue with priority tierstones.py— Tone WAV generator (entry/exit/standby)audio.py— WAV writing andpw-playasync wrappersettings.py— Pydantic settings from env vars (prefix:TTS_)engines/— TTSEngine implementations (kokoro, piper, orpheus)
Key Design Decisions
- Speech queue is serialized (one playback at a time) but synthesis is parallel
- Tones are non-fatal: if
pw-playfails on a tone, speech still plays - Orpheus uses llama-server (not Ollama) for 15x throughput via continuous batching
- SNAC decoder is lazy-loaded on first Orpheus call to reduce idle memory