mcspeak/CLAUDE.md
Ryan Malloy 6a81e0760a Add exit tone (roger beep) and standby tone after speech
Queue-aware exit tones: descending roger beep when queue empties
("over"), ascending standby blip when more items are queued ("standby,
more coming"). Includes quindar-out (2475 Hz) as Apollo-themed
alternative. Configurable via TTS_EXIT_TONE env var.

Also adds CLAUDE.md documenting the full tone system.
2026-02-23 14:44:27 -07:00

2.6 KiB

TTS MCP Server

Multi-engine text-to-speech server exposed via FastMCP 3.0 Streamable HTTP. Engines: Kokoro (ONNX), Piper (Wyoming/Docker), Orpheus (llama-server + SNAC).

Build & Run

make up        # build + start (docker compose)
make logs      # follow logs
make restart   # restart containers
make status    # show running containers + health

Entry & Exit Tones (Beep System)

Queued speech playback (speak()) is bookended by short alert tones. generate_audio() is unaffected (file-only, no playback).

Tone Positions

Position When Purpose
Entry tone Before speech starts "Incoming transmission" alert
Exit tone After speech, queue empty "Over and out" — channel clear
Standby tone After speech, more queued "Standby" — more messages coming

Available Tones

Name Frequency Duration Inspired by
chirp 1800 Hz ~144 ms Nextel iDEN Talk Permit Tone (TPT) — the 24/24/24/24/48 ms on/off pattern
apollo 2525 Hz 250 ms NASA quindar intro (key-up) tone used during Apollo missions
roger 1400-1000 Hz ~100 ms Classic CB radio descending two-tone roger beep
quindar-out 2475 Hz 250 ms NASA quindar unkey tone (distinct frequency from intro)
standby 1000-1400 Hz ~60 ms Ascending blip — inverse of roger, signals "more coming"

Configuration

TTS_ENTRY_TONE=chirp        # before speech (chirp, apollo, none, or /path/to/custom.wav)
TTS_EXIT_TONE=roger          # after speech, queue empty (roger, quindar-out, none, or path)

The standby tone is always the built-in ascending blip. It plays instead of the exit tone when more items are queued.

Tones are generated programmatically at startup (48kHz, 16-bit PCM, -3 dB headroom) in tones.py using numpy. No bundled audio assets.

Architecture

  • server.py — FastMCP lifespan, tool definitions, engine setup
  • queue.py — Producer-consumer speech queue with priority tiers
  • tones.py — Tone WAV generator (entry/exit/standby)
  • audio.py — WAV writing and pw-play async wrapper
  • settings.py — Pydantic settings from env vars (prefix: TTS_)
  • engines/ — TTSEngine implementations (kokoro, piper, orpheus)

Key Design Decisions

  • Speech queue is serialized (one playback at a time) but synthesis is parallel
  • Tones are non-fatal: if pw-play fails on a tone, speech still plays
  • Orpheus uses llama-server (not Ollama) for 15x throughput via continuous batching
  • SNAC decoder is lazy-loaded on first Orpheus call to reduce idle memory