diff --git a/CLAUDE.md b/CLAUDE.md index 1aca469..af2df3a 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -36,17 +36,55 @@ Queued speech playback (`speak()`) is bookended by short alert tones. `generate_ | `scratch` | 2000-300 Hz sweep + noise | ~120 ms | Vinyl record scratch — needle yanked off the platter | | `reverse-roger` | 1000-1400 Hz | ~100 ms | Ascending two-tone — mathematical inverse of roger beep | +**Pink Floyd voice kit (current defaults).** The bookends now follow a telephone theme: + +| Name | Used as | Inspired by | +|------|---------|-------------| +| `heartbeat` | speak entry | "Speak to Me" heartbeat (DSotM opener) woven with warm MF telephone tones | +| `soft-chord` | speak exit | Mellow resolving A3+E4 dyad, pure sines, ear-friendly on repetition | +| `mf-dial` | listen start | The **genuine Young Lust R1 MF operator dial** — KP, 0-4-4-1-8-3-1, ST (the 44 is the UK country code), real Bell-System MF pairs | +| `machine` | listen end | "Welcome to the Machine" pulsing VCS3 throb — a warm "connected / got it" | +| `mf-listen` / `mf-done` | listen (alt) | Gentler stacked-fifths swell / resolving D-major (previous listen defaults) | +| `call-waiting` | over ongoing speech | Brief C6 double-blip mixed over the current message when another project queues | + +There are also softer "natural" tones (`bell-soft`, `chime-tube`, `water-drop`, `soft-pulse`, `hmm-up`, …) in `tones.py` for custom use. + ### Configuration ```env -TTS_ENTRY_TONE=chirp # before speech (chirp, apollo, none, or /path/to/custom.wav) -TTS_EXIT_TONE=roger # after speech, queue empty (roger, quindar-out, none, or path) -TTS_CANCEL_TONE=scratch # on cancel (scratch, reverse-roger, none, or /path/to/custom.wav) -TTS_SHUTDOWN_TIMEOUT=30 # max seconds to wait for current speech on container stop +TTS_ENTRY_TONE=heartbeat # before speech (heartbeat, chirp, apollo, none, or /path/to/custom.wav) +TTS_EXIT_TONE=soft-chord # after speech, queue empty (soft-chord, roger, quindar-out, none, or path) +TTS_CANCEL_TONE=scratch # on cancel (scratch, reverse-roger, none, or /path/to/custom.wav) +TTS_CALL_WAITING_TONE=call-waiting # mixed over current speech when a DIFFERENT project queues (once/turn) +TTS_LISTEN_START_TONE=mf-dial # played once the mic is live (see listen()) +TTS_LISTEN_END_TONE=machine # played after recording stops +TTS_SHUTDOWN_TIMEOUT=30 # max seconds to wait for current speech on container stop ``` The standby tone is always the built-in ascending blip. It plays instead of the exit tone when more items are queued. +## Secretary Announcements & Call-Waiting + +When several projects speak concurrently, the queue behaves like a **secretary**: each message plays whole and in order (never interleaved — see the streaming-utterance model below), and a project is announced by name when the speaker changes or returns after a lull. + +- The announcement is a short `"."` preamble synthesized in a **reserved secretary voice** (`TTS_SECRETARY_VOICE`, default `bf_emma`) — always via Kokoro so it sounds identical regardless of the speaking engine. That voice is excluded from the project auto-assignment pool so no project ever sounds like the secretary. +- The play-or-skip decision is made at **play time** in the consumer (`queue.py:_should_announce`), not enqueue time, because urgent reordering means the real speaker order isn't final until then. +- **Call-waiting**: when a *different* project's message joins the queue while one is playing, a brief `call-waiting` blip is mixed over the current audio (a second `pw-play` stream — PipeWire mixes it). Fires **at most once per playing turn** (`_call_waiting_fired`, reset when a new utterance starts) so a burst of queued messages never spams the listener. + +```env +TTS_ANNOUNCE_MODE=secretary # secretary | always | off (legacy TTS_ANNOUNCE_PROJECT=true → always) +TTS_REINTRODUCE_AFTER_SECONDS=120 # same project after this much silence gets re-introduced +TTS_SECRETARY_VOICE=bf_emma # reserved; excluded from project auto-assignment +``` + +## listen() — Voice Conversations + +`listen()` captures the host mic (`pw-record`), transcribes via Parakeet on the gpu.supported.systems gateway, and returns the text. Pair it with `speak()` for turn-taking: speak a question (let it finish), then `listen()` for the reply — **sequentially, never in parallel**, or the mic records the TTS. + +- **Defaults are conversation-first**: `wait_for_silence=True` (stop when the person stops), `duration_seconds=30` cap, `vad_aggressiveness=3`, `silence_threshold_ms=2200`. The aggressiveness/threshold defaults were tuned live to stop brief background transients from ending the turn before the real reply. +- **No first-word clip**: the "go" tone is played *after* the mic is live (a `warmup_ms` lead, default 150ms). The beep bleeds harmlessly into the head of the recording — VAD treats a pure tone as non-speech and Parakeet ignores it. See `audio.py:record_audio_until_silence`. +- Empty transcription (`text == ""`) or a gateway timeout means re-prompt rather than proceed; the recording is saved under `/tmp/mcspeak/` and can be retried with `transcribe()`. + Tones are generated programmatically at startup (48kHz, 16-bit PCM, -3 dB headroom) in `tones.py` using numpy. No bundled audio assets. ## Voice Identity (Project-Aware Voices) @@ -132,8 +170,8 @@ All pactl errors are caught and logged. If PulseAudio is unavailable (no socket, ## Architecture - `server.py` — FastMCP lifespan, tool definitions, engine setup -- `queue.py` — Producer-consumer speech queue with priority tiers and outcome tracking -- `tones.py` — Tone WAV generator (entry/exit/standby) +- `queue.py` — Producer-consumer queue of streaming **utterances** (one per `speak()`): priority tiers, secretary announcements + reserved voice, call-waiting, outcome tracking. Synthesized WAVs are reaped after playback (no /tmp leak). +- `tones.py` — Tone WAV generator (speak + listen bookends, call-waiting, natural set) - `media_duck.py` — Async PulseAudio volume control for media ducking - `audio.py` — WAV writing and `pw-play` async wrapper - `settings.py` — Pydantic settings from env vars (prefix: `TTS_`) @@ -185,16 +223,18 @@ Kokoro synthesizes ~4x realtime on CPU, so synthesis always outpaces playback. ` Texts under 20 words or without sentence boundaries (`.!?` followed by whitespace) take the single-shot path — zero overhead, identical to pre-chunking behavior. -### Tone behavior +### Tone behavior & per-message coherence + +Each `speak()` call is **one queue utterance** (not one queue item per chunk). The utterance reserves its ordering slot up front and streams its synthesized chunks in through an internal channel closed by an `_END` sentinel; the consumer stays locked to it from entry tone to exit tone. So when several projects speak at once, a message plays **whole and in order** — chunks never interleave with another project's audio (the old per-chunk `suppress_exit_tone`/`_WorkItem` model is gone). - Entry tone plays once at the start (before first chunk synthesis) -- Exit/standby tones are suppressed between chunks (`suppress_exit_tone` flag on `_WorkItem`) -- Final chunk plays the normal exit tone (roger) or standby tone +- Exit/standby tone plays once, at the very end of the utterance +- Pipelining is preserved: synthesis of chunk N+1 overlaps playback of chunk N, but nothing else can be pulled until this utterance finishes ### Cancellation in chunked mode -- **Explicit `cancel_speech(speech_id)`** — cancels the returned speech_id (the final chunk). Already-playing earlier chunks finish naturally. -- **MCP disconnect** — the synthesis loop stops (remaining chunks aren't synthesized). Already-enqueued chunks play through. +- **Explicit `cancel_speech(speech_id)`** — one speech_id now covers the whole utterance; cancelling it flags an abort (synchronously), reaps un-played chunk WAVs, and plays the cancel tone. +- **MCP disconnect** — the `speak()` handler is cancelled; its `finally` closes the utterance channel so the consumer drains the chunks it already has and finishes cleanly. ### Progress lifecycle (chunked) @@ -208,8 +248,8 @@ Texts under 20 words or without sentence boundaries (`.!?` followed by whitespac ### Files -- `server.py` — `split_text()`, `_speak_single()`, `_speak_chunked()`, `_await_with_progress()` -- `queue.py` — `suppress_exit_tone` field on `_WorkItem` +- `server.py` — `split_text()`, unified `_speak()` (short + chunked share one path), status-aware `_await_with_progress()` +- `queue.py` — `_Utterance` (streaming chunk channel + `_END` sentinel), `create_utterance()`, `_play_utterance()`, `_should_announce()`, call-waiting ## Cancellation @@ -235,8 +275,10 @@ Docker's `stop_grace_period` must exceed the total: `3s + shutdown_timeout + 5s ## Key Design Decisions -- Speech queue is serialized (one playback at a time) but synthesis is parallel -- `speak()` blocks until playback finishes with live progress (5% → 30% → 35-99% → 100%) +- Speech queue is serialized (one playback at a time) but synthesis is parallel; each `speak()` is one **streaming utterance** so concurrent projects never interleave +- Secretary behavior: a project is announced by name on speaker-change / after a lull, in a **reserved voice** excluded from the project pool; a different project queuing mid-playback fires a **once-per-turn** call-waiting blip mixed over the current audio +- Synthesized WAVs are reaped after playback (and on cancel/shutdown) so `/tmp/mcspeak` doesn't grow unbounded +- `speak()` blocks until playback finishes with **status-aware** progress — a queued item reports "waiting in line", not a false "playing" percentage - Progress uses a background ticker task, NOT `asyncio.wait_for` polling (see below) - Entry tone is awaited in `speak()` before synthesis — covers latency gap - Explicit `cancel_speech()` kills pw-play + plays cancel tone; MCP disconnect lets playback finish