docs: utterance queue, secretary, tone kit, listen
Update CLAUDE.md for the streaming-utterance model (drop the stale _WorkItem/suppress_exit_tone references), the secretary announcements + reserved voice, call-waiting, the telephone tone kit, and the conversation-first listen() defaults.
This commit is contained in:
parent
3c4e06aa64
commit
fc5057366c
72
CLAUDE.md
72
CLAUDE.md
@ -36,17 +36,55 @@ Queued speech playback (`speak()`) is bookended by short alert tones. `generate_
|
||||
| `scratch` | 2000-300 Hz sweep + noise | ~120 ms | Vinyl record scratch — needle yanked off the platter |
|
||||
| `reverse-roger` | 1000-1400 Hz | ~100 ms | Ascending two-tone — mathematical inverse of roger beep |
|
||||
|
||||
**Pink Floyd voice kit (current defaults).** The bookends now follow a telephone theme:
|
||||
|
||||
| Name | Used as | Inspired by |
|
||||
|------|---------|-------------|
|
||||
| `heartbeat` | speak entry | "Speak to Me" heartbeat (DSotM opener) woven with warm MF telephone tones |
|
||||
| `soft-chord` | speak exit | Mellow resolving A3+E4 dyad, pure sines, ear-friendly on repetition |
|
||||
| `mf-dial` | listen start | The **genuine Young Lust R1 MF operator dial** — KP, 0-4-4-1-8-3-1, ST (the 44 is the UK country code), real Bell-System MF pairs |
|
||||
| `machine` | listen end | "Welcome to the Machine" pulsing VCS3 throb — a warm "connected / got it" |
|
||||
| `mf-listen` / `mf-done` | listen (alt) | Gentler stacked-fifths swell / resolving D-major (previous listen defaults) |
|
||||
| `call-waiting` | over ongoing speech | Brief C6 double-blip mixed over the current message when another project queues |
|
||||
|
||||
There are also softer "natural" tones (`bell-soft`, `chime-tube`, `water-drop`, `soft-pulse`, `hmm-up`, …) in `tones.py` for custom use.
|
||||
|
||||
### Configuration
|
||||
|
||||
```env
|
||||
TTS_ENTRY_TONE=chirp # before speech (chirp, apollo, none, or /path/to/custom.wav)
|
||||
TTS_EXIT_TONE=roger # after speech, queue empty (roger, quindar-out, none, or path)
|
||||
TTS_CANCEL_TONE=scratch # on cancel (scratch, reverse-roger, none, or /path/to/custom.wav)
|
||||
TTS_SHUTDOWN_TIMEOUT=30 # max seconds to wait for current speech on container stop
|
||||
TTS_ENTRY_TONE=heartbeat # before speech (heartbeat, chirp, apollo, none, or /path/to/custom.wav)
|
||||
TTS_EXIT_TONE=soft-chord # after speech, queue empty (soft-chord, roger, quindar-out, none, or path)
|
||||
TTS_CANCEL_TONE=scratch # on cancel (scratch, reverse-roger, none, or /path/to/custom.wav)
|
||||
TTS_CALL_WAITING_TONE=call-waiting # mixed over current speech when a DIFFERENT project queues (once/turn)
|
||||
TTS_LISTEN_START_TONE=mf-dial # played once the mic is live (see listen())
|
||||
TTS_LISTEN_END_TONE=machine # played after recording stops
|
||||
TTS_SHUTDOWN_TIMEOUT=30 # max seconds to wait for current speech on container stop
|
||||
```
|
||||
|
||||
The standby tone is always the built-in ascending blip. It plays instead of the exit tone when more items are queued.
|
||||
|
||||
## Secretary Announcements & Call-Waiting
|
||||
|
||||
When several projects speak concurrently, the queue behaves like a **secretary**: each message plays whole and in order (never interleaved — see the streaming-utterance model below), and a project is announced by name when the speaker changes or returns after a lull.
|
||||
|
||||
- The announcement is a short `"<project>."` preamble synthesized in a **reserved secretary voice** (`TTS_SECRETARY_VOICE`, default `bf_emma`) — always via Kokoro so it sounds identical regardless of the speaking engine. That voice is excluded from the project auto-assignment pool so no project ever sounds like the secretary.
|
||||
- The play-or-skip decision is made at **play time** in the consumer (`queue.py:_should_announce`), not enqueue time, because urgent reordering means the real speaker order isn't final until then.
|
||||
- **Call-waiting**: when a *different* project's message joins the queue while one is playing, a brief `call-waiting` blip is mixed over the current audio (a second `pw-play` stream — PipeWire mixes it). Fires **at most once per playing turn** (`_call_waiting_fired`, reset when a new utterance starts) so a burst of queued messages never spams the listener.
|
||||
|
||||
```env
|
||||
TTS_ANNOUNCE_MODE=secretary # secretary | always | off (legacy TTS_ANNOUNCE_PROJECT=true → always)
|
||||
TTS_REINTRODUCE_AFTER_SECONDS=120 # same project after this much silence gets re-introduced
|
||||
TTS_SECRETARY_VOICE=bf_emma # reserved; excluded from project auto-assignment
|
||||
```
|
||||
|
||||
## listen() — Voice Conversations
|
||||
|
||||
`listen()` captures the host mic (`pw-record`), transcribes via Parakeet on the gpu.supported.systems gateway, and returns the text. Pair it with `speak()` for turn-taking: speak a question (let it finish), then `listen()` for the reply — **sequentially, never in parallel**, or the mic records the TTS.
|
||||
|
||||
- **Defaults are conversation-first**: `wait_for_silence=True` (stop when the person stops), `duration_seconds=30` cap, `vad_aggressiveness=3`, `silence_threshold_ms=2200`. The aggressiveness/threshold defaults were tuned live to stop brief background transients from ending the turn before the real reply.
|
||||
- **No first-word clip**: the "go" tone is played *after* the mic is live (a `warmup_ms` lead, default 150ms). The beep bleeds harmlessly into the head of the recording — VAD treats a pure tone as non-speech and Parakeet ignores it. See `audio.py:record_audio_until_silence`.
|
||||
- Empty transcription (`text == ""`) or a gateway timeout means re-prompt rather than proceed; the recording is saved under `/tmp/mcspeak/` and can be retried with `transcribe()`.
|
||||
|
||||
Tones are generated programmatically at startup (48kHz, 16-bit PCM, -3 dB headroom) in `tones.py` using numpy. No bundled audio assets.
|
||||
|
||||
## Voice Identity (Project-Aware Voices)
|
||||
@ -132,8 +170,8 @@ All pactl errors are caught and logged. If PulseAudio is unavailable (no socket,
|
||||
## Architecture
|
||||
|
||||
- `server.py` — FastMCP lifespan, tool definitions, engine setup
|
||||
- `queue.py` — Producer-consumer speech queue with priority tiers and outcome tracking
|
||||
- `tones.py` — Tone WAV generator (entry/exit/standby)
|
||||
- `queue.py` — Producer-consumer queue of streaming **utterances** (one per `speak()`): priority tiers, secretary announcements + reserved voice, call-waiting, outcome tracking. Synthesized WAVs are reaped after playback (no /tmp leak).
|
||||
- `tones.py` — Tone WAV generator (speak + listen bookends, call-waiting, natural set)
|
||||
- `media_duck.py` — Async PulseAudio volume control for media ducking
|
||||
- `audio.py` — WAV writing and `pw-play` async wrapper
|
||||
- `settings.py` — Pydantic settings from env vars (prefix: `TTS_`)
|
||||
@ -185,16 +223,18 @@ Kokoro synthesizes ~4x realtime on CPU, so synthesis always outpaces playback. `
|
||||
|
||||
Texts under 20 words or without sentence boundaries (`.!?` followed by whitespace) take the single-shot path — zero overhead, identical to pre-chunking behavior.
|
||||
|
||||
### Tone behavior
|
||||
### Tone behavior & per-message coherence
|
||||
|
||||
Each `speak()` call is **one queue utterance** (not one queue item per chunk). The utterance reserves its ordering slot up front and streams its synthesized chunks in through an internal channel closed by an `_END` sentinel; the consumer stays locked to it from entry tone to exit tone. So when several projects speak at once, a message plays **whole and in order** — chunks never interleave with another project's audio (the old per-chunk `suppress_exit_tone`/`_WorkItem` model is gone).
|
||||
|
||||
- Entry tone plays once at the start (before first chunk synthesis)
|
||||
- Exit/standby tones are suppressed between chunks (`suppress_exit_tone` flag on `_WorkItem`)
|
||||
- Final chunk plays the normal exit tone (roger) or standby tone
|
||||
- Exit/standby tone plays once, at the very end of the utterance
|
||||
- Pipelining is preserved: synthesis of chunk N+1 overlaps playback of chunk N, but nothing else can be pulled until this utterance finishes
|
||||
|
||||
### Cancellation in chunked mode
|
||||
|
||||
- **Explicit `cancel_speech(speech_id)`** — cancels the returned speech_id (the final chunk). Already-playing earlier chunks finish naturally.
|
||||
- **MCP disconnect** — the synthesis loop stops (remaining chunks aren't synthesized). Already-enqueued chunks play through.
|
||||
- **Explicit `cancel_speech(speech_id)`** — one speech_id now covers the whole utterance; cancelling it flags an abort (synchronously), reaps un-played chunk WAVs, and plays the cancel tone.
|
||||
- **MCP disconnect** — the `speak()` handler is cancelled; its `finally` closes the utterance channel so the consumer drains the chunks it already has and finishes cleanly.
|
||||
|
||||
### Progress lifecycle (chunked)
|
||||
|
||||
@ -208,8 +248,8 @@ Texts under 20 words or without sentence boundaries (`.!?` followed by whitespac
|
||||
|
||||
### Files
|
||||
|
||||
- `server.py` — `split_text()`, `_speak_single()`, `_speak_chunked()`, `_await_with_progress()`
|
||||
- `queue.py` — `suppress_exit_tone` field on `_WorkItem`
|
||||
- `server.py` — `split_text()`, unified `_speak()` (short + chunked share one path), status-aware `_await_with_progress()`
|
||||
- `queue.py` — `_Utterance` (streaming chunk channel + `_END` sentinel), `create_utterance()`, `_play_utterance()`, `_should_announce()`, call-waiting
|
||||
|
||||
## Cancellation
|
||||
|
||||
@ -235,8 +275,10 @@ Docker's `stop_grace_period` must exceed the total: `3s + shutdown_timeout + 5s
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
- Speech queue is serialized (one playback at a time) but synthesis is parallel
|
||||
- `speak()` blocks until playback finishes with live progress (5% → 30% → 35-99% → 100%)
|
||||
- Speech queue is serialized (one playback at a time) but synthesis is parallel; each `speak()` is one **streaming utterance** so concurrent projects never interleave
|
||||
- Secretary behavior: a project is announced by name on speaker-change / after a lull, in a **reserved voice** excluded from the project pool; a different project queuing mid-playback fires a **once-per-turn** call-waiting blip mixed over the current audio
|
||||
- Synthesized WAVs are reaped after playback (and on cancel/shutdown) so `/tmp/mcspeak` doesn't grow unbounded
|
||||
- `speak()` blocks until playback finishes with **status-aware** progress — a queued item reports "waiting in line", not a false "playing" percentage
|
||||
- Progress uses a background ticker task, NOT `asyncio.wait_for` polling (see below)
|
||||
- Entry tone is awaited in `speak()` before synthesis — covers latency gap
|
||||
- Explicit `cancel_speech()` kills pw-play + plays cancel tone; MCP disconnect lets playback finish
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user