/live: full-duplex voice (Kyutai streaming STT with semantic end of turn, or Silero and Whisper; Kokoro TTS) through a sidecar; /vibe: director mode that…

Voice and director modes for Claude Code, as a function-hook plugin (a mod) with a Python audio sidecar (bin/sidecar/, started by the mod with uv run --script bin/sidecar/main.py). The sidecar talks to the mod over stdout and a loopback HTTP port guarded by a per-run token.
| Command | What it does |
|---|---|
/live | Full-duplex voice with Claude. The mic stays open; you can talk over an answer to cut it off. What you say goes to Claude, and Claude's final answer is read aloud. Saying "switch to sonnet" runs /model. |
/vibe [request] | Director mode. Claude reads and directs worker subagents; every other tool is denied on the main loop. |
/livevibe | You talk with a small, fast voice model (the front). It chats with you directly, hands real work to Claude in vibe mode, and tells you the result in a sentence or two when it comes back. /livevibe model [name] lists or picks the front model; /livevibe url [url] points it at a server you run (/livevibe url managed goes back to the built-in one). |
/live and /livevibe turn each other off. /livevibe off puts vibe mode back the way it was.
Claude's answers print as usual. Around them, each line starts with live-vibe: and is dim:
| Line | What it is |
|---|---|
you (voice): ... | What you said, once per sentence: an utterance cut off mid-sentence ("Also, we need") waits up to 2.5 s for its rest, and the pieces go to the front as one. |
spoken: Let me check. | The front's acknowledgements, spoken and shown once; back-to-back ones share the line ((x2)). One still waiting when Claude's answer lands is dropped (the debug log keeps it). |
spoken summary: ... | The front's retelling of a result, at most two sentences plus a closing question. When it repeats the answer above, the line is cut short and marked (repeats the answer above); the debug log has it whole. It shows as soon as the front has written it, while the voice still reads it. |
voice: ... | The front's own reply to small talk. |
front prompt (3 lines, ctrl+o to expand): task: ... | The prompt the front handed Claude (your words, the front's reading, the instructions), drawn as one line; ctrl+o shows it whole. |
no update (nothing to add) | A Claude turn with nothing new, such as a worker's notification repeating a report: the front says nothing. |
Talking over the voice stops it on your first word. A word the echo guard places in what the speaker played in the last 2.5 s is the assistant's own echo and does not stop it (the sidecar log says reason=echo-ignored). Without a running echo canceller (AEC3 off or missing), the sidecar cannot tell echo from you, and three words stop it, as before. On the Whisper recognizer, the speech onset is transcribed until it shows a word.
The first word pauses the voice rather than ending it. If what you said turns out to be only a listening sound ("mm-hm", "uh-huh", "yeah", "okay", "right", "sure", "got it") and ends within 1.2 s, the voice picks up from the sample where it stopped, and your words reach Claude as a note, not a prompt. Anything else ends the answer there: the rest of the sentence and the sentences after it are dropped, and only the words actually played count as heard (a cut sentence up to its last whole word in the played share of its audio). Each barge-in writes one line to the sidecar log: barge-in: words=<n> at_ms=<ms of the clip played> resumed=<yes|no> reason=<backchannel|speech|echo-ignored>.
/vibe still works there. Headphones are best, because the mic stays open while the assistant speaks. On open speakers, the sidecar cancels the assistant's echo from the mic (WebRTC AEC3, through the livekit package), keeps barge-in strict until the speaker has gone quiet, and drops utterances that only repeat what it just said (see Barge-in below). Pick devices with the mic and speaker settings; --list-devices (below) shows the choices.brew install espeak-ng for Kokoro. Kyutai STT runs on MLX.sudo apt install espeak-ng libportaudio2. On x86_64, uv also installs PyTorch with CUDA and Kyutai's moshi package (about 3 GB the first time), and Kyutai STT runs on an Nvidia GPU (about 3.2 GB of VRAM; it needs 5 GB free when it starts, else Whisper runs). Without a usable GPU, Whisper runs instead: on an Nvidia GPU when CTranslate2 finds one, CUDA 12 cuBLAS and cuDNN 9 are installed and the GPU has room, otherwise on the CPU. Kokoro runs on the GPU through onnxruntime-gpu (a CUDA 12 build, sharing PyTorch's CUDA and cuDNN wheels), when about 1.4 GB of VRAM still leaves 3 GB free; otherwise on the CPU, which is fast enough. Every GPU backend checks free VRAM first, so a GPU already busy with another model (a local LLM server) does not get pushed into paging; the log says where each one runs (tts: Kokoro ... on CUDAExecutionProvider or CPUExecutionProvider).Under WSL2, WSLg carries audio over RDP, and that leg crackles: the speech leaves WSL clean (PulseAudio's RDPSink.monitor matches a fresh Kokoro render) yet sounds full of static on Windows. With speakerBackend on auto (the default) the sidecar plays speech on Windows instead: it runs win_player.exe, a small native player (Rust, in win_player/, shipped prebuilt as bin/sidecar/win_player.exe), through WSL interop and streams the PCM over its stdin, and the player plays it through WASAPI shared mode on the Windows default output (or the one whose name contains the speaker setting). No network, port or firewall rule is involved, and the mic stays on WSLg.
%LOCALAPPDATA%\live-vibe: a copy of the exe named by its content (win_player-<sha>.exe, so an update never overwrites one another session is running). Nothing is installed globally. /live setup stages it and reports speaker: Windows player via WASAPI on <device>, <rate> Hz, <latency> ms.hello in about 30 ms and opens the device in about 20 ms (measured through interop). The sidecar pings it from hello on, gives the open 30 s, and logs both timings and the player's opening progress, so a slow start shows which side stalled.bin/sidecar/win_player.py, the same player in Python, run by uv.exe (a pinned 0.10.8, sha256-checked, when none is on the PATH; uv's cache and any Python it downloads also live in %LOCALAPPDATA%\live-vibe, about 60 MB the first time).speakerBackend: local keeps the old path (WSLg), windows insists on Windows (still falling back with a warning).win_player/build.sh (needs rustup target add x86_64-pc-windows-gnu and sudo apt install gcc-mingw-w64-x86-64) runs the crate's tests, builds the host binary the --unit checks drive with --fake, cross-compiles the exe and copies it to bin/sidecar/win_player.exe; commit that file.Run /live setup once on each machine. uv installs the Python packages, then the sidecar downloads the models your settings name and tests each piece: PortAudio and the mic and speaker, espeak-ng, the synthesizer, the recognizer (it speaks a sentence into it and checks what comes back), one sentence through the speaker, and the front server (with frontUrl empty it downloads llama-server and the front model, starts the server once and stops it). Progress shows in the status line, and a report with a ✓, ! or ✗ per step lands in the transcript. Installing the plugin installs none of this, and setup never installs system packages: a ✗ line names the command to run (uv itself, espeak-ng, libportaudio2). Change a setting, run it again.
Without setup, the first /live or /livevibe downloads the models, and the status line shows loading until they are ready: Kyutai STT about 2.4 GB (Apple silicon, or Linux with an Nvidia GPU), the Whisper model (75 to 500 MB), Kokoro about 330 MB, and Silero VAD about 2 MB. Later starts take a few seconds. Kokoro and Silero are cached in ~/.cache/duplex_voice (LIVE_VIBE_CACHE overrides it); Kyutai and Whisper use the Hugging Face cache.
Everything the sidecar says (startup, which recognizer and synthesizer ran or fell back, whether the echo canceller loaded, front warm-up failures, warnings, and library output on stderr) is also appended to ~/.cache/duplex_voice/sidecar.log (LIVE_VIBE_CACHE or XDG_CACHE_HOME move it; the logFile setting sets the exact path). /live setup reports the path. Each session starts with a header line (plugin version, arguments, platform, Python), and each line is ISO-time LEVEL pid=<sidecar pid> text, with LEVEL one of INFO, WARNING or STDERR:
2026-10-02T23:14:42.826-06:00 INFO pid=786551 session start: live-vibe 0.3.1 argv=[...] platform=Linux-... python=3.12.3
2026-10-02T23:14:50.101-06:00 WARNING pid=786551 Kyutai STT on CUDA unavailable (...); using Whisper.
The file rotates at 5 MB and keeps one previous copy (sidecar.log.1). If the path cannot be written, the sidecar warns once and carries on without it. The managed front server's own output goes to a separate file, front-server.log beside it (the frontServerLog setting); the sidecar log gets its start, command line, port and stop.
/config)| Field | Default | Meaning |
|---|---|---|
stt | kyutai | kyutai (streaming, ends your turn on what you said; MLX on Apple silicon, CUDA on Linux with an Nvidia GPU, else Whisper) or whisper (Silero VAD + faster-whisper). |
asr | base.en | Whisper size: tiny.en, base.en, small.en. |
tts | kokoro | kokoro, or say (macOS only). |
voice | empty | A Kokoro voice (af_heart) or a macOS say voice. |
mic | empty | Input device: an index or part of its name (AirPods). Empty means the system default. |
speaker | empty | Output device, the same way. With the Windows speaker, part of a Windows device name. |
speakerBackend | auto | auto: the Windows player under WSL with interop, else local. local: always local (WSLg under WSL). windows: the Windows player, falling back to local with a warning. |
logFile | empty | Sidecar log path; empty means ~/.cache/duplex_voice/sidecar.log. |
endSilenceMs | 3000 | Whisper: the pause that ends your turn. Kyutai: the longest pause between words while the model is unsure you are done (1 s after a finished sentence). |
endSilenceLongMs | 4000 | Kyutai: the pause allowed when the model predicts you will keep talking or the sentence looks unfinished (a trailing comma, and, so, the). |
frontBackend | llamacpp | llamacpp (the managed llama-server, or any OpenAI-compatible server that takes response_format) or anthropic (needs ANTHROPIC_API_KEY or ant auth login). |
frontUrl | empty | Empty: a managed llama-server (below). Set: the front's server, which can be another host, such as a GPU box; nothing starts locally. |
frontServerBin | empty | Managed server: an existing llama-server to run instead of the download. |
frontServerModel | empty | Managed server: an existing GGUF to serve instead of the default model. |
frontServerLog | empty | Managed server: where its stdout and stderr go; empty means ~/.cache/duplex_voice/front-server.log. |
frontModel | empty | Sent as model on every request. Empty means the server's own model (or claude-haiku-4-5 for anthropic). |
experimental | true | The experimental voice path (below). Off restores the 0.6.4 behaviour. |
backchannels | false | Live vibe, with experimental on: the front says a short "mm-hm" during long dictation. Leave it off on open speakers. |
The experimental setting turns on changes from the GPT-Live design notes (docs/design/gpt-live-informed-voice.md), all at once. With it off, the sidecar and the relay prompt behave exactly as in 0.6.4, except that the 30 s maximum-utterance extension stays on regardless, since it is a bug fix.
Status: working, done, failed or cancelled and at most three spoken sentences; the front retells only those (500 tokens at most, counted by llama-server's /tokenize). The rest of the answer goes into the front's history as silent notes, so "how's it going?" right after a report is answered by the front itself, with no Claude turn; after a new delegation the front still says "Let me check."backchannels also on (Kyutai only): a cached Kokoro "mm-hm" when you have talked for over 6 s and pause with more to come, at most once per 8 s, never over the front's voice or while you answer its question. It plays through the same player as speech, so the echo canceller hears it as playback.The sidecar log gets one line per front turn, front turn: eot_to_first_token_ms=<n> first_audio_ms=<n> f_keep=<n/a> prefill_tokens=<n> cached_tokens=<n> kind=user|event, from llama-server's timings (its f_keep is only in front-server.log), with the setting on or off, and with it on, front speculate: used|discarded|cancelled saved_ms=<n>, front warm: ... and front backchannel: ... lines.
With frontUrl empty and the llamacpp backend, /livevibe runs its own llama.cpp llama-server. /live setup downloads a pinned prebuilt release (llama.cpp v0.5.0, build b11146, checked against GitHub's sha256 digests) and the default model, Qwen3-4B-Instruct-2507 Q4_K_M (2.5 GB, checked against the Hugging Face sha256), into ~/.cache/duplex_voice, starts the server once, and reports front: managed llama-server <version> with <model> on <GPU|CPU>. After that, nothing goes to the network. Which build it downloads:
| Machine | Build | Download |
|---|---|---|
| macOS, Apple silicon | Metal | 11 MB |
Linux x86_64 with an Nvidia GPU (nvidia-smi works, WSL2 too) | CUDA 12.8 plus its runtime libraries | 760 MB |
| Linux x86_64, no Nvidia GPU, a Vulkan loader, not WSL | Vulkan | 31 MB |
| Other Linux x86_64 or arm64, macOS Intel | CPU | 11 to 17 MB |
| Windows x86_64 (untested) | CUDA 12.4 with an Nvidia GPU, else CPU | 645 / 19 MB |
CUDA 12.8 rather than 13: 12.8 already has Blackwell kernels, while CUDA 13 needs a newer driver and drops pre-Turing GPUs, and the rest of the sidecar (PyTorch for Kyutai, onnxruntime for Kokoro) runs CUDA 12 anyway. Vulkan is not used on WSL2, which has no Nvidia Vulkan driver.
Each /livevibe start picks a free port on 127.0.0.1 and runs llama-server -m <model> --host 127.0.0.1 --port <port> --jinja -c 16384 -np 1 --no-webui -ngl 99, waits for /health, and points the front at it; the server stops when live vibe stops or the sidecar exits. If the sidecar is killed outright, Linux takes the server down with it (PR_SET_PDEATHSIG); elsewhere the next start stops a server whose sidecar is gone. The log (frontServerLog) gets a header with the command line, port, model and placement at every start. The server only answers requests that carry a key made fresh for each start (passed in LLAMA_API_KEY, never on the command line), since llama-server otherwise accepts calls from any web page.
GPU memory is shared with the voice models through one budget. The front asks first, before the recognizer and Kokoro, since a front model on the CPU costs seconds on every reply, while the recognizer falls back to Whisper and Kokoro is fine on the CPU. Its need is worked out from the GGUF (weights, f16 KV cache at the context, about 0.75 GB for the CUDA context): 5.3 GiB for the default model. If the GPU would keep less than 3 GiB free, the server runs on the CPU (-ngl 0 --device none) with one warning that replies will be slow. On Apple silicon it always uses Metal.
To skip the downloads, point frontServerBin at a llama-server you built and frontServerModel at a GGUF you have. If anything fails, the warning says why and speech goes straight to Claude, as when any front server is down. To use a server you run instead, set frontUrl (or /livevibe url <url>), as below.
Any OpenAI-compatible chat server that takes response_format with a JSON schema works (llama-server does). Each front turn is one JSON object, {"delegate": ..., "say": ...}, held to that schema by the server's grammar: small models offered a delegate tool tend to say they are on it and call nothing, but they fill a required field. The delegation goes out as soon as its field closes, and the say text is spoken as it streams, except on a turn that delegates or keeps a user's question (below): there only a short first sentence is spoken, then "On it, I've asked." or "I've asked for the details.", since the answer is Claude's to give. A small model is enough; Qwen3-4B-Instruct-2507 delegated every work request in testing, in about 5 GB of VRAM with a 16k context (4 GB at 8k):
llama-server -m Qwen3-4B-Instruct-2507-Q4_K_M.gguf --jinja -c 16384 -np 1 -ngl 99 --flash-attn on --port 8080
To switch between small models without keeping a big one in memory, put a router in front (llama-swap, or llama-server's multi-model mode where your build has it) and pick the model with /livevibe model <name>. A plain llama-server ignores the model name. Thinking is turned off per request, so a voice turn does not wait for it.
What reaches Claude from the front:
User said: "...", then The voice front read it as: ...), so a small model's rewrite cannot lose the question.[The user cut off the voice front's retelling of your last answer; it was spoken up to: "...<its last 12 words heard>"] (or that they heard none of it), so Claude does not assume the rest was heard. A reply of the front's own that was cut reaches Claude inside its note, as heard: up to its last whole word, then ....[The user, by voice, said "Mm-hm." while listening; the voice played on (no task asked)].experimental on, Claude's relay prompt asks it to end every answer with a Status: working|done|failed|cancelled line and at most three spoken sentences. The front retells those lines; the rest reaches the front only as background. A status question the latest report answers, with nothing handed off since, is answered by the front and reaches Claude as a note rather than a question.Claude's results are announced when nobody is talking. A result that arrives while the user is speaking goes into the front's reply to them instead; results that waited together are one announcement, and none is announced twice.
uv run --script bin/sidecar/main.py --list-devices # the indexes and names for mic and speaker
uv run --script bin/sidecar/main.py --setup # what /live setup runs (with the default settings)
uv run --script bin/sidecar/main.py --unit # pure unit checks
uv run --script bin/sidecar/main.py --selftest # both modes, no mic, speaker or networkhooks/register.tsx 622 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, PluginOptions, Register } from 'claude-code'
3
4import type { Live, LiveMode, LiveState } from '../types'
5
6const PLUGIN = 'live-vibe'
7const OFF: Live = { isOn: false, mode: 'live', state: 'off', port: 0, token: '', turnId: null, vibeBefore: false, gen: 0 }
8const live = atom({ plugin: 'live-vibe', key: 'live' } as const, OFF)
9const isVibe = atom({ plugin: 'live-vibe', key: 'isVibe' } as const, false)
10
11// The director reads and directs; every other tool is a worker's job.
12const DIRECTOR_TOOLS = new Set([
13 'Read', 'Agent', 'SendMessage', 'ListAgents', 'TaskStop',
14 'TaskCreate', 'TaskUpdate', 'TaskList', 'ToolSearch', 'AskUserQuestion',
15])
16
17// After Oh My Pi's vibe-mode-active.md, with its vibe_* tools mapped onto Claude Code's own.
18const VIBE_SECTION = `<vibe-mode>
19Vibe mode ON. You are DIRECTOR: drive worker subagents, full coding agents with every normal tool; NEVER edit, run, grep, or build yourself. Verify work by reading files.
20
21Toolset: Read, Agent (spawn a worker, run_in_background: true), SendMessage (continue a worker by name), ListAgents, TaskStop (kill a worker), TaskCreate/TaskUpdate/TaskList (your own bookkeeping), AskUserQuestion.
22
23# Workers
24
25- fast: model "sonnet"; mechanical, well-specified work: renames, small fixes, boilerplate, data collection, tests and output reports.
26- good: model "opus"; design, tricky debugging, multi-file refactors, judgment-heavy work.
27
28Workers are persistent conversations: give each a name, one worker per workstream, keep it on that workstream. Spawn once, then SendMessage the SAME worker for follow-ups; NEVER respawn it.
29
30# Direction
31
321. Split requests into independent workstreams.
332. Spawn each with a complete self-contained brief: files, constraints, acceptance criteria. Workers start blank; they never see this conversation.
343. Spawns and sends return immediately; a worker's result arrives as a task notification when its turn finishes. Direct other workers meanwhile; never poll.
354. On each result, Read the touched files to verify claims before building on them; then SendMessage corrections, the next step, or a review request. Track parent tasks with TaskCreate/TaskUpdate; workers do not own this bookkeeping.
365. Route by difficulty: draft with fast; escalate to good if fast stalls or judgment is needed. good designs; fast executes mechanical parts.
376. TaskStop stuck workers or workers whose workstream is done; ListAgents if the roster is lost.
38
39Run workers concurrently, normally one fast and one good on different workstreams. The final outcome is yours: verify with Read; do not take a worker's word for it.
40</vibe-mode>`
41
42const VOICE_SECTION = `<live-voice>
43Live voice is ON. The user is speaking, and your final answer of each turn is read aloud by a speech synthesizer; the transcript is still on screen. Do the work with tools as usual, but make the final visible answer one to three short spoken sentences: no markdown, no lists, no code blocks, no URLs. Say file names and numbers plainly. Anything long belongs in a file the user can open, named in one sentence. Bracketed notes in the conversation such as "The user spoke over you" or "they heard only" are system annotations, not your words: the first is a live interruption to act on at once, the second says how much of an answer was heard, so do not repeat it.
44</live-voice>`
45
46// In live vibe the user talks to a small voice model; Claude's answers reach them only through its summary.
47const RELAY_SECTION = `<live-vibe-relay>
48Live vibe is ON. The user is talking to a voice front, a small fast model that hands their spoken requests to you and relays your final answer of each turn back to them, summarized in one or two spoken sentences. End every turn with a concise plain-text report the front can summarize: what you did, what you found or changed, whether it is verified, and whether work is still running in the background (workers spawned, results pending). No spoken style needed, but lead with the outcome and leave out long listings unless asked. A bracketed note such as "[The user, by voice, adds ...]" is a live addition to act on at once. A request that starts "User said:" quotes the user's own words, then the front's reading of them; where they differ, go by the user's words. A note "[The user, by voice, to the voice front (no task asked): ...]" is information, such as a confirmation, a correction or a decision: take it as fact from the user in your next answer; it needs no reply of its own. When a turn brings the user nothing new (a worker's notification that repeats what you already reported, a result already relayed), reply with exactly: (nothing to add). The front then says nothing, and the transcript shows one dim "no update" line.
49</live-vibe-relay>`
50
51// The experimental voice path (userConfig `experimental`): the front retells only a Status line and the sentences after
52// it, and keeps the rest as silent context, as GPT-Live's commentary and thinking appends. One constant per setting,
53// so the section stays byte-stable in Claude's prompt cache for the whole session.
54const RELAY_SECTION_EXPERIMENTAL = RELAY_SECTION.replace('</live-vibe-relay>', `End every answer, after the report, with one line \`Status: working\`, \`Status: done\`, \`Status: failed\` or \`Status: cancelled\` (working while anything you started still runs), then at most three short plain sentences for the voice to say: the outcome first, a question for the user last. The front retells only those lines; the rest of your answer stays on screen and reaches the front as background it can use for status questions. The "(nothing to add)" reply has no Status line.
55</live-vibe-relay>`)
56
57// A whole utterance that only asks for another model ("switch to sonnet", "change the model to opus please", "use haiku").
58// Anchored at both ends, so a request that merely mentions a model ("use opus to review this") stays a prompt.
59// The live vibe sidecar receives this source as --switch-pattern, so both modes match the same words.
60const SPOKEN_SWITCH = /^(?:(?:ok|okay|hey|alright) )?(?:claude )?(?:please )?(?:(?:switch|change|swap|set)(?: over)?(?: the)?(?: model)? to|use)(?: the)? (opus|sonnet|haiku|fable)(?: model)?(?: please)?$/
61
62function spokenModel(text: string): string | undefined {
63 return SPOKEN_SWITCH.exec(text.toLowerCase().replace(/[^a-z ]+/g, ' ').replace(/\s+/g, ' ').trim())?.[1]
64}
65
66// After a delegation's two readings: the front's gloss is a hint, never a task of its own (seen with Qwen3-4B: "how
67// are we doing?" handed off as "Review the current status of all running tasks and logs").
68const READING = "Answer the user's words; treat the front's reading as a hint, and if it is a question about status or "
69 + 'progress, answer from what you know rather than starting new work.'
70
71// A delegation as Claude reads it: the user's words, the front's reading, then READING, one per line. The reading
72// gets a full stop only when it lacks one (a reading that ended in "." once came out "methods.. Answer").
73function closed(text: string) {
74 const t = text.trim()
75 return /[.!?\u2026]["')\]]*$/.test(t) ? t : `${t}.`
76}
77
78function delegationPrompt(said: string, reading: string) {
79 return `User said: "${said}"\nThe voice front read it as a task: ${closed(reading)}\n${READING}`
80}
81
82function delegationSteer(said: string, reading: string) {
83 return `[The user, by voice, adds while you work: "${said}" The voice front read it as a task: ${closed(reading)} ${READING}]`
84}
85
86// -- what the transcript shows in live vibe ---------------------------------------------------------------------------
87// The engine leads every $.ui.log line with "live-vibe: " and draws it dim. A log row has no collapsed state, so the
88// lines that repeat what is already on screen are one short line, the collapsed form; the prompts the plugin submits
89// are UserMessage rows, which the ui.render hook below draws as one dim line until ctrl+o expands the transcript.
90const FILLER_FLUSH_MS = 1500 // an acknowledgement waits this long for another to join its line
91const ANSWERS_KEPT = 3 // Claude's last answers, which a retelling may repeat
92
93let answers: string[] = []
94let filler: { parts: { text: string; n: number }[]; gen: number } | null = null
95let fillerGen = 0
96const voicePrompts: string[] = [] // prompts submitted for the voice in live vibe, drawn collapsed
97
98function clip(text: string, max: number) {
99 const t = text.replace(/\s+/g, ' ').trim()
100 return t.length <= max ? t : `${t.slice(0, max - 1).replace(/\s+\S*$/, '')}\u2026`
101}
102
103function wordsOf(text: string) {
104 return text.toLowerCase().match(/[a-z0-9']+/g) ?? []
105}
106
107// A retelling that is nearly all words of an answer on screen says nothing new: seen live, the front read Claude's
108// whole answer back and the transcript showed it twice.
109function repeats(spoken: string, shown: readonly string[]) {
110 const words = wordsOf(spoken)
111 if (words.length === 0) return false
112 return shown.some(answer => {
113 const vocabulary = new Set(wordsOf(answer))
114 return words.filter(w => vocabulary.has(w)).length / words.length >= 0.8
115 })
116}
117
118function fillerLine(parts: { text: string; n: number }[]) {
119 return `spoken: ${parts.map(p => (p.n > 1 ? `${p.text} (x${p.n})` : p.text)).join(' / ')}`
120}
121
122function flushFiller($: EngineInterface) {
123 if (!filler) return
124 const { parts } = filler
125 filler = null
126 $.ui.log(fillerLine(parts))
127}
128
129// "Let me check.", "On it, I've asked.": spoken, then one short line; back-to-back ones share it.
130function holdFiller($: EngineInterface, text: string) {
131 if (!filler) {
132 const gen = ++fillerGen
133 filler = { parts: [], gen }
134 void $.clock.sleep(FILLER_FLUSH_MS).then(() => { if (filler?.gen === gen) flushFiller($) }).catch(() => {})
135 }
136 const last = filler.parts.at(-1)
137 if (last && last.text === text) last.n += 1
138 else filler.parts.push({ text, n: 1 })
139}
140
141function showTranscript($: EngineInterface, msg: Record<string, unknown>) {
142 const text = String(msg.text ?? '').trim()
143 if (!text) return
144 if (msg.role === 'user') {
145 flushFiller($)
146 $.ui.log(`you (voice): ${text}`)
147 return
148 }
149 if (msg.kind === 'filler') return holdFiller($, text)
150 flushFiller($)
151 if (msg.kind !== 'relay') return $.ui.log(`voice: ${text}`)
152 if (!repeats(text, answers)) return $.ui.log(`spoken summary: ${text}`)
153 $.ui.log(`spoken summary: ${clip(text, 72)} (repeats the answer above)`)
154 $.ui.log(`live: spoken in full: ${text}`, { to: 'debug' })
155}
156
157// The collapsed line for a prompt this plugin submitted for the voice, or undefined to draw the row as the engine does.
158function promptHeader(text: string) {
159 const lines = text.split('\n').length
160 const head = `front prompt (${lines} line${lines === 1 ? '' : 's'}, ctrl+o to expand)`
161 const task = /^User said: "[\s\S]*"\nThe voice front read it as a task: (.*)\n/.exec(text)
162 if (task) return `${head}: task: ${clip(task[1] ?? '', 100)}`
163 if (text.startsWith('User asked by voice: ')) return `${head}: your question, for Claude to answer`
164 if (text.startsWith('User said, answering your last question: ')) return `${head}: your answer to Claude's question`
165 return voicePrompts.includes(text) ? head : undefined
166}
167
168// Claude's whole answer to a voice question when the front's own holding line already says enough: never posted.
169const NOTHING_TO_ADD = '(nothing to add)'
170
171function nothingToAdd(answer: string) {
172 return answer.trim().toLowerCase().replace(/\./g, '') === NOTHING_TO_ADD
173}
174
175// Where Claude's words go: read aloud as they are (/live), or through the front's announcer (/livevibe).
176function replyPath(l: Live) {
177 return l.mode === 'livevibe' ? '/event' : '/speak'
178}
179
180// /model as if typed. The engine queues it until the session is idle, so a switch spoken mid-turn lands when the turn ends.
181// Not awaited: the sidecar loop must keep reading while the run waits for the turn.
182function switchModel($: EngineInterface, model: string, isBusy: boolean) {
183 if (isBusy) $.ui.toast(`live: switching to ${model} when this turn ends`)
184 void $.command.run({ command: 'model', args: model })
185 .then(async ({ text }) => post($, replyPath(await read($, live)), text?.trim() || `Model set to ${model}.`))
186 .catch(err => $.ui.toast(`live: model switch failed: ${String(err).slice(0, 120)}`))
187}
188
189// A sidecar that is gone (it quit, or never reached ready) is not an error worth more than a debug line.
190// The sidecar refuses a POST without the token it reported in `ready`, so no other local process can drive it.
191async function post($: EngineInterface, path: string, body = '') {
192 const { port, token } = await read($, live)
193 if (!port) return
194 try {
195 await $.http.fetch(`http://127.0.0.1:${port}${path}`, { method: 'POST', body, headers: { 'X-Live-Token': token } })
196 } catch (err) {
197 $.ui.log(`live: POST ${path} failed: ${String(err).slice(0, 120)}`, { to: 'debug' })
198 }
199}
200
201// The engine leads the line with "live-vibe: ", so it names no mode of its own: "listening · director" is live vibe,
202// "voice listening" is /live, "director" alone is /vibe.
203async function status($: EngineInterface) {
204 const [l, v] = [await read($, live), await read($, isVibe)]
205 const state = l.state.replace('_', ' ')
206 const parts = [l.isOn && (l.mode === 'livevibe' ? state : `voice ${state}`), v && 'director']
207 $.ui.status(parts.some(Boolean) ? parts.filter(Boolean).join(' · ') : undefined)
208}
209
210// Ends either mode: the sidecar quits, and live vibe hands vibe mode back as it found it.
211async function stopLive($: EngineInterface) {
212 const l = await read($, live)
213 if (!l.isOn) return
214 await post($, '/quit')
215 if (l.mode === 'livevibe') await update($, isVibe, () => l.vibeBefore)
216 await update($, live, x => ({ ...OFF, gen: x.gen || 0 }))
217 await status($)
218}
219
220// Steer a running main-loop turn (the row reaches the model at its next step; nothing is aborted), or start one.
221async function toClaude($: EngineInterface, text: string, steer: string) {
222 const { turnId, mode } = await read($, live)
223 if (turnId) {
224 await $.session.append({ message: { type: 'user', content: [{ type: 'text', text: steer }] } })
225 return
226 }
227 if (mode === 'livevibe') voicePrompts.splice(0, Math.max(0, voicePrompts.push(text) - 50))
228 void $.prompt.submit({ text, asUser: true })
229}
230
231async function onSidecar($: EngineInterface, msg: Record<string, unknown>) {
232 switch (msg.type) {
233 case 'ready':
234 await update($, live, l => ({ ...l, port: Number(msg.port), token: String(msg.token ?? '') }))
235 $.ui.toast(`${(await read($, live)).mode === 'livevibe' ? 'live vibe' : 'live'}: listening (headphones recommended)`)
236 break
237 case 'state':
238 await update($, live, l => ({ ...l, state: msg.state as LiveState }))
239 await status($)
240 break
241 case 'utterance': {
242 const text = String(msg.text)
243 const model = spokenModel(text)
244 if (model) {
245 switchModel($, model, (await read($, live)).turnId !== null)
246 break
247 }
248 await toClaude($, text, `[The user spoke over you: "${text}". Take this into account now: change course if it asks you to, answer it in your next reply, and keep that reply short.]`)
249 break
250 }
251 // The front's reading of a request, with the user's own words when the sidecar sends them (`said`): a small
252 // model's rewrite can drop the question, so Claude always sees what was actually said.
253 case 'delegate': {
254 const text = String(msg.text)
255 const said = typeof msg.said === 'string' ? msg.said.trim() : ''
256 if (!said) {
257 await toClaude($, text, `[The user, by voice, adds a request while you work: "${text}". Take it into account now.]`)
258 break
259 }
260 await toClaude($, delegationPrompt(said, text), delegationSteer(said, text))
261 break
262 }
263 // A user turn the front answered itself: a confirmation, a correction or a decision is still Claude's to know.
264 // It joins the conversation without starting a turn. Two kinds are prompts: a reply to the question Claude's last
265 // answer asked (`answer`), and a question the front did not delegate (`ask`): it said only a holding line, and
266 // Claude gives the real answer, or NOTHING_TO_ADD, which turn.complete keeps from the front.
267 case 'note': {
268 const said = typeof msg.said === 'string' ? msg.said.trim() : ''
269 if (!said) break
270 // "mm-hm" over the voice: it paused, then played on. Claude learns the user was listening, nothing more.
271 if (msg.backchannel === true) {
272 await $.session.append({ message: { type: 'user', content: [{
273 type: 'text', text: `[The user, by voice, said "${said}" while listening; the voice played on (no task asked)]`,
274 }] } })
275 break
276 }
277 const reply = typeof msg.reply === 'string' && msg.reply.trim() ? ` (the voice front answered: "${msg.reply.trim()}")` : ''
278 if (msg.answer === true) {
279 await toClaude($, `User said, answering your last question: "${said}"`,
280 `[The user, by voice, answers your last question: "${said}"${reply}. Take it into account now.]`)
281 break
282 }
283 if (msg.ask === true) {
284 await toClaude($, `User asked by voice: "${said}"${reply}\nThe voice front did not answer it. Give the real answer. If you have nothing to add beyond the voice front's line, reply with exactly: ${NOTHING_TO_ADD}`,
285 `[The user, by voice, asks while you work: "${said}"${reply}. The voice front did not answer it: answer it in your next reply.]`)
286 break
287 }
288 await $.session.append({ message: { type: 'user', content: [{
289 type: 'text', text: `[The user, by voice, to the voice front (no task asked): "${said}"${reply}]`,
290 }] } })
291 break
292 }
293 // The user cut the front's retelling of Claude's answer: Claude hears where it stopped, so it does not assume the
294 // rest was heard.
295 case 'cut': {
296 const heard = typeof msg.heard === 'string' ? msg.heard.trim() : ''
297 const text = heard
298 ? `[The user cut off the voice front's retelling of your last answer; it was spoken up to: "...${heard}"]`
299 : "[The user cut off the voice front before it retold your last answer; they did not hear it]"
300 await $.session.append({ message: { type: 'user', content: [{ type: 'text', text }] } })
301 break
302 }
303 case 'switch_model':
304 switchModel($, String(msg.model), (await read($, live)).turnId !== null)
305 break
306 case 'transcript':
307 showTranscript($, msg)
308 break
309 case 'spoken':
310 if (msg.cut) {
311 await $.session.append({ message: { type: 'user', content: [{
312 type: 'text',
313 text: `[The user interrupted. Of your last answer they heard only: "${String(msg.text)}". Do not repeat it.]`,
314 }] } })
315 }
316 break
317 case 'warn':
318 $.ui.toast(`live: ${String(msg.text)}`, { timeoutMs: 8000 })
319 $.ui.log(`live: ${String(msg.text)}`, { to: 'debug' })
320 break
321 case 'log':
322 $.ui.log(String(msg.text), { to: 'debug' })
323 break
324 }
325}
326
327// With frontUrl empty the sidecar runs its own llama-server; these settings point it at files already on disk.
328function frontServerArgv(options: PluginOptions) {
329 const argv: string[] = []
330 if (options.frontServerBin) argv.push('--front-server-bin', String(options.frontServerBin))
331 if (options.frontServerModel) argv.push('--front-server-model', String(options.frontServerModel))
332 if (options.frontServerLog) argv.push('--front-server-log', String(options.frontServerLog))
333 return argv
334}
335
336// The sidecar lives for the session: the loop runs on after the hook returns and ends with the child or the module.
337function startSidecar($: EngineInterface, options: PluginOptions, mode: LiveMode, vibeBefore: boolean) {
338 void (async () => {
339 let gen = 0
340 await update($, live, (l): Live => {
341 gen = (l.gen || 0) + 1
342 return { isOn: true, mode, state: 'loading', port: 0, token: '', turnId: null, vibeBefore, gen }
343 })
344 await status($)
345 const argv = ['uv', 'run', '--script', `${$.plugin.root}/bin/sidecar/main.py`,
346 '--mode', mode === 'livevibe' ? 'front' : 'live', '--stt', String(options.stt), '--asr', String(options.asr),
347 '--tts', String(options.tts), '--end-silence-ms', String(options.endSilenceMs), '--end-silence-long-ms', String(options.endSilenceLongMs ?? 4000)]
348 if (options.voice) argv.push('--voice', String(options.voice))
349 if (options.mic) argv.push('--mic', String(options.mic))
350 if (options.speaker) argv.push('--speaker', String(options.speaker))
351 if (options.speakerBackend) argv.push('--speaker-backend', String(options.speakerBackend))
352 if (options.logFile) argv.push('--log-file', String(options.logFile))
353 if (options.experimental) argv.push('--experimental')
354 if (options.experimental && options.backchannels && mode === 'livevibe') argv.push('--backchannels')
355 if (mode === 'livevibe') {
356 argv.push('--front-backend', String(options.frontBackend), '--front-url', String(options.frontUrl),
357 '--switch-pattern', SPOKEN_SWITCH.source, ...frontServerArgv(options))
358 if (options.frontModel) argv.push('--front-model', String(options.frontModel))
359 }
360 const child = $.process.spawn({ argv })
361 const isCurrent = async () => { const l = await read($, live); return l.isOn && l.gen === gen }
362 let buf = ''
363 for await (const chunk of child) {
364 if (!('stream' in chunk)) break
365 // Turned off or replaced while the child was still writing: leaving the loop kills it.
366 if (!(await isCurrent())) return
367 if (chunk.stream === 'stderr') { $.ui.log(chunk.text, { to: 'debug' }); continue }
368 buf += chunk.text
369 const lines = buf.split('\n')
370 buf = lines.pop() ?? ''
371 for (const line of lines) {
372 if (!line.trim()) continue
373 let msg: Record<string, unknown>
374 try { msg = JSON.parse(line) } catch { $.ui.log(line, { to: 'debug' }); continue }
375 await onSidecar($, msg)
376 }
377 }
378 // The child ended on its own (goodbye, a crash): turn the mode off as /live or /livevibe would.
379 if (await isCurrent()) {
380 await update($, live, l => ({ ...l, port: 0 }))
381 await stopLive($)
382 }
383 })().catch(err => sidecarFailed($, err).catch(() => {})) // a module unloading mid-loop lands here too: stay quiet
384}
385
386// The child never started (no uv, most often) or the loop broke: say why, and turn off a mode that never got going.
387async function sidecarFailed($: EngineInterface, err: unknown) {
388 $.ui.toast(await hasUv($) ? `live: sidecar failed: ${String(err).slice(0, 120)}. /live setup checks everything.` : `live: ${UV_FIX}`, { timeoutMs: 15000 })
389 const l = await read($, live)
390 if (l.isOn && !l.port) await stopLive($)
391}
392
393const UV_FIX = 'uv is not installed, or not on the PATH Claude Code started with. Install it (curl -LsSf https://astral.sh/uv/install.sh | sh; macOS also brew install uv; Windows: powershell -c "irm https://astral.sh/uv/install.ps1 | iex"), restart Claude Code, then /live setup.'
394
395async function hasUv($: EngineInterface) {
396 try {
397 return (await $.process.run(['uv', '--version'])).exitCode === 0
398 } catch {
399 return false
400 }
401}
402
403const MARK: Record<string, string> = { ok: '✓', warn: '!', fail: '✗' }
404let isSettingUp = false // one at a time; a reload kills the child and starts this over
405
406// /live setup: the sidecar's --setup installs the Python packages (uv does, before Python starts), downloads the
407// models the settings name, and tests the mic, speaker, synthesizer, recognizer and front. Progress goes to the
408// status line; the report is one transcript row when it ends. System packages stay the person's to install: the
409// report names the command.
410function setupVoice($: EngineInterface, options: PluginOptions) {
411 isSettingUp = true
412 void (async () => {
413 $.ui.status('setup: installing the Python packages (a minute or two on first run)')
414 const argv = ['uv', 'run', '--script', `${$.plugin.root}/bin/sidecar/main.py`, '--setup',
415 '--stt', String(options.stt), '--asr', String(options.asr), '--tts', String(options.tts),
416 '--end-silence-ms', String(options.endSilenceMs), '--end-silence-long-ms', String(options.endSilenceLongMs ?? 4000), '--front-backend', String(options.frontBackend),
417 '--front-url', String(options.frontUrl), ...frontServerArgv(options)]
418 if (options.voice) argv.push('--voice', String(options.voice))
419 if (options.mic) argv.push('--mic', String(options.mic))
420 if (options.speaker) argv.push('--speaker', String(options.speaker))
421 if (options.speakerBackend) argv.push('--speaker-backend', String(options.speakerBackend))
422 if (options.logFile) argv.push('--log-file', String(options.logFile))
423 if (options.frontModel) argv.push('--front-model', String(options.frontModel))
424 const rows: string[] = []
425 const errTail: string[] = []
426 let ok: boolean | null = null
427 let buf = ''
428 for await (const chunk of $.process.spawn({ argv })) {
429 if (!('stream' in chunk)) break
430 if (chunk.stream === 'stderr') {
431 $.ui.log(chunk.text, { to: 'debug' })
432 errTail.push(...chunk.text.split('\n').filter(x => x.trim()))
433 errTail.splice(0, Math.max(0, errTail.length - 8))
434 continue
435 }
436 buf += chunk.text
437 const lines = buf.split('\n')
438 buf = lines.pop() ?? ''
439 for (const line of lines) {
440 let msg: Record<string, unknown>
441 try { msg = JSON.parse(line) } catch { continue }
442 const text = String(msg.text ?? '')
443 if (msg.type === 'progress') $.ui.status(`setup: ${text}`)
444 else if (msg.type === 'check') rows.push(`${MARK[String(msg.status)] ?? '?'} ${String(msg.name)}: ${text}`)
445 else if (msg.type === 'warn') rows.push(`! ${text}`)
446 else if (msg.type === 'log' && /^(download|kyutai):/.test(text)) $.ui.status(`setup: ${text}`)
447 else if (msg.type === 'done') ok = Boolean(msg.ok)
448 }
449 }
450 const head = ok ? 'Live voice setup: ready. /live or /livevibe to start.'
451 : ok === false ? 'Live voice setup: fix the ✗ lines, then /live setup again.'
452 : 'Live voice setup stopped before it finished. The last output:'
453 $.ui.log([head, ...rows, ...(ok === null ? errTail : [])].join('\n'))
454 $.ui.toast(ok ? 'live setup: ready' : 'live setup: needs attention (see the transcript)')
455 })()
456 .catch(err => $.ui.toast(`live setup failed: ${String(err).slice(0, 160)}`))
457 .finally(() => { isSettingUp = false; void status($).catch(() => {}) })
458}
459
460// Live vibe on: vibe mode joins the voice front, and stopLive later hands vibe back as `vibeBefore` had it.
461async function startLiveVibe($: EngineInterface, options: PluginOptions, vibeBefore: boolean) {
462 await update($, isVibe, () => true)
463 startSidecar($, options, 'livevibe', vibeBefore)
464}
465
466// What the front server offers, in a line; a server that is down or slow says so instead of holding the command.
467async function servedModels($: EngineInterface, url: string) {
468 try {
469 const r = await Promise.race([$.http.fetch(`${url}/v1/models`), $.clock.sleep(3000).then(() => null)])
470 if (!r) return `${url} did not answer within 3 s.`
471 if (!r.ok) return `${url}/v1/models answered ${r.status}.`
472 const ids = ((JSON.parse(r.text) as { data?: { id?: unknown }[] }).data ?? []).map(m => String(m.id))
473 return `It lists: ${ids.join(', ') || 'nothing'}.`
474 } catch (err) {
475 return `${url} did not answer (${String(err).slice(0, 80)}).`
476 }
477}
478
479// /livevibe model|url [value]: list what the front server offers, or point the front elsewhere. The model name is
480// sent as each request's `model`, so a router (llama-swap, llama-server's multi-model mode) loads it on demand and a
481// plain llama-server ignores it. The value is stored as this plugin's own userConfig field (`$.config.set` takes
482// `<plugin>.<field>`), so it survives the session. `/livevibe url managed` empties frontUrl: the sidecar's own server.
483async function frontSetting($: EngineInterface, options: PluginOptions, field: 'model' | 'url', value: string) {
484 const url = String(options.frontUrl ?? '').replace(/\/+$/, '')
485 if (!value) {
486 const at = url ? `at ${url}` : `on a managed llama-server (${String(options.frontServerModel || 'the default model')})`
487 const current = `Front: ${String(options.frontBackend)} ${at}, model ${String(options.frontModel) || '(server default)'}.`
488 if (field === 'url') return { text: `${current} Change it with /livevibe url <url>, or /livevibe url managed.` }
489 if (!url) return { text: `${current} The managed server serves one model; frontServerModel picks it.` }
490 return { text: `${current} ${await servedModels($, url)} Pick one with /livevibe model <name>; /livevibe model default clears it.` }
491 }
492 if (field === 'url' && value !== 'managed' && !/^https?:\/\/\S+$/.test(value)) return { text: `Not a URL: ${value}` }
493 const next = (field === 'model' && value === 'default') || (field === 'url' && value === 'managed') ? '' : value
494 const { deny } = await $.config.set({ key: field === 'url' ? 'live-vibe.frontUrl' : 'live-vibe.frontModel', value: next })
495 if (deny) return { text: `Front ${field} not changed: ${deny}` }
496 const l = await read($, live)
497 if (l.isOn && l.mode === 'livevibe') {
498 // ponytail: a restart reloads Whisper and Kokoro too; a sidecar endpoint that hot-swaps the front brain is the upgrade.
499 await stopLive($)
500 await startLiveVibe($, { ...options, [field === 'url' ? 'frontUrl' : 'frontModel']: next }, l.vibeBefore)
501 return { text: `Front ${field} set to ${next || (field === 'url' ? '(managed llama-server)' : '(server default)')}; the voice restarts on it.` }
502 }
503 return { text: `Front ${field} set to ${next || (field === 'url' ? '(managed llama-server)' : '(server default)')}.` }
504}
505
506export const register: Register = (on, options) => {
507 on('session.start', async ($, e, next) => {
508 await $.command.register({ name: 'live', description: 'Toggle live voice mode: speak to Claude, hear the answers. /live setup installs and tests what it needs', argumentHint: '[setup]' })
509 await $.command.register({ name: 'livevibe', description: 'Toggle live vibe: talk with a fast voice front that hands the work to Claude as vibe director', argumentHint: '[model [name] | url [url|managed]]' })
510 await $.command.register({ name: 'vibe', description: 'Toggle vibe mode: Claude directs worker subagents instead of editing itself', argumentHint: '[first request]' })
511 const l = await read($, live)
512 if (l.isOn) startSidecar($, options, l.mode === 'livevibe' ? 'livevibe' : 'live', Boolean(l.vibeBefore)) // a hot reload killed the child; bring it back
513 await status($)
514 return next(e)
515 })
516
517 on('command.run', { command: 'live' }, async ($, e) => {
518 const args = e.args.trim()
519 if (args && args !== 'setup') return { text: 'Usage: /live toggles live voice; /live setup checks and installs what it needs.' }
520 if (args === 'setup') {
521 if ((await read($, live)).isOn) return { text: 'Turn /live or /livevibe off first: setup loads the same models.' }
522 if (isSettingUp) return { text: 'Setup is already running; progress is in the status line.' }
523 if (!(await hasUv($))) return { text: UV_FIX }
524 setupVoice($, options)
525 return { text: 'Setting up live voice: Python packages, models, then a test of the mic, speaker and voice. Progress is in the status line; the report lands here.' }
526 }
527 const l = await read($, live)
528 await stopLive($) // off, or switching over from live vibe
529 if (l.isOn && l.mode === 'live') return { text: 'Live voice off.' }
530 startSidecar($, options, 'live', false)
531 return { text: 'Live voice on: loading the models, then listening. Say what you want; /live again turns it off. First run on this machine? /live setup shows progress.' }
532 })
533
534 on('command.run', { command: 'livevibe' }, async ($, e) => {
535 const sub = /^(model|url)(?:\s+(\S+))?$/.exec(e.args.trim())
536 if (sub) return frontSetting($, options, sub[1] === 'url' ? 'url' : 'model', sub[2] ?? '')
537 if (e.args.trim()) return { text: 'Usage: /livevibe toggles; /livevibe model [name]; /livevibe url [url|managed].' }
538 const l = await read($, live)
539 await stopLive($) // off, or switching over from /live
540 if (l.isOn && l.mode === 'livevibe') return { text: 'Live vibe off: voice front stopped, vibe mode back as it was.' }
541 await startLiveVibe($, options, await read($, isVibe))
542 return { text: `Live vibe on: Claude directs, a voice front (${String(options.frontBackend)}) talks with you. Loading, then listening; /livevibe again turns it off.` }
543 })
544
545 on('command.run', { command: 'vibe' }, async ($, e) => {
546 const turnOn = !(await read($, isVibe))
547 await update($, isVibe, () => turnOn)
548 await status($)
549 if (turnOn && e.args.trim()) void $.prompt.submit({ text: e.args.trim(), asUser: true })
550 return { text: turnOn ? 'Vibe mode on: Claude directs, workers do the work. /vibe again exits.' : 'Vibe mode off: the full toolset is back.' }
551 })
552
553 on('prompt.compose', async ($, e, next) => {
554 const { sections } = await next(e)
555 const out = [...sections]
556 const l = await read($, live)
557 // Only the main loop lists Agent; a worker's own prompt must not be told it is the director.
558 const isMain = e.tools.includes('Agent')
559 if ((await read($, isVibe)) && isMain) out.push({ id: `${PLUGIN}:vibe`, text: VIBE_SECTION, scope: 'session' })
560 if (l.isOn && l.mode === 'livevibe' && isMain) {
561 out.push({ id: `${PLUGIN}:relay`, text: options.experimental ? RELAY_SECTION_EXPERIMENTAL : RELAY_SECTION, scope: 'session' })
562 }
563 if (l.isOn && l.mode !== 'livevibe') out.push({ id: `${PLUGIN}:voice`, text: VOICE_SECTION, scope: 'session' })
564 return { sections: out }
565 })
566
567 on('tool.call', async ($, e, next) => {
568 if (e.agentId || !(await read($, isVibe)) || DIRECTOR_TOOLS.has(String(e.tool))) return next(e)
569 return { deny: `Vibe mode: the director does not run ${String(e.tool)}. Spawn or SendMessage a worker for it, then Read to verify.` }
570 })
571
572 // Main loop only: a subagent's run raises no turn.start.
573 on('turn.start', async ($, e, next) => {
574 if ((await read($, live)).isOn) await update($, live, l => ({ ...l, turnId: e.turnId }))
575 return next(e)
576 })
577
578 on('turn.complete', async ($, e, next) => {
579 const l = await read($, live)
580 if (!l.isOn || e.agentId) return next(e)
581 await update($, live, x => ({ ...x, turnId: null }))
582 if (e.reason === 'answer' && e.answer.trim()) {
583 if (!nothingToAdd(e.answer)) void post($, replyPath(l), e.answer)
584 answers = [e.answer, ...answers].slice(0, ANSWERS_KEPT)
585 // An acknowledgement still waiting for its line would land under the answer it stood in for.
586 if (filler) $.ui.log(`live: ${fillerLine(filler.parts)} (dropped: the answer came first)`, { to: 'debug' })
587 filler = null
588 }
589 // The front is waiting on a result; tell it the work stopped rather than leave it silent.
590 else if (l.mode === 'livevibe' && (e.reason === 'error' || e.reason === 'refusal')) {
591 void post($, '/event', e.reason === 'error' ? 'The work stopped on an error before it finished.' : 'That request was declined.')
592 }
593 return next(e)
594 })
595
596 // The prompts live vibe submits for the voice (the user's words, the front's reading, the instructions) are one dim
597 // line in the transcript, as a thinking block is; ctrl+o draws them whole, and Claude reads them whole either way.
598 on('ui.render', { component: 'UserMessage', props: { origin: { kind: 'plugin' } } }, async ($, e, next) => {
599 const { origin, text, isExpanded } = e.props
600 const head = !isExpanded && origin.kind === 'plugin' && origin.name === PLUGIN ? promptHeader(text) : undefined
601 if (!head) return next(e)
602 const { Text } = $.ui.resolve(e)
603 return <Text dimColor>{head}</Text>
604 })
605
606 // A turn with nothing new ("(nothing to add)", from any turn: a voice question, a worker's notification) is one dim
607 // line; the front never hears it (turn.complete). ctrl+o shows the stored message.
608 on('ui.render', { component: 'AssistantMessage' }, async ($, e, next) => {
609 if (!nothingToAdd(e.props.text)) return next(e)
610 const { Text } = $.ui.resolve(e)
611 return <Text dimColor>no update (nothing to add)</Text>
612 })
613
614 on('prompt.submit', async ($, e, next) => {
615 // A request the user typed cuts the voice short, like a hand raised. A spoken one already did, in the sidecar,
616 // and a plugin's own (a delegation, a task notification) must not silence the front mid-sentence.
617 const byUser = e.origin.kind === 'composer' || e.origin.kind === 'bridge'
618 if (byUser && (await read($, live)).state === 'speaking') void post($, '/stop')
619 return next(e)
620 })
621}
622types/index.d.ts 27 lines1export type LiveState = 'off' | 'loading' | 'listening' | 'user_speaking' | 'transcribing' | 'thinking' | 'speaking'
2
3/** live: speak to Claude directly. livevibe: speak to a voice front that delegates to Claude as the vibe director. */
4export type LiveMode = 'live' | 'livevibe'
5
6export type Live = {
7 isOn: boolean
8 mode: LiveMode
9 state: LiveState
10 /** The sidecar's HTTP port on 127.0.0.1; 0 until it reports ready. */
11 port: number
12 /** The sidecar's per-launch secret from `ready`, sent as X-Live-Token on every POST; '' until then. */
13 token: string
14 /** The main loop's running turn, so a spoken interruption can steer it; null when idle. */
15 turnId: string | null
16 /** Vibe mode as it was before /livevibe turned it on, restored when live vibe ends. */
17 vibeBefore: boolean
18 /** Bumped per sidecar start, so a stale sidecar's loop leaves the current one's state alone. */
19 gen: number
20}
21
22declare module 'claude-code' {
23 interface PluginState {
24 'live-vibe': { live: Live; isVibe: boolean }
25 }
26}
27