Voice conversation mode: talk to this session and hear it talk back

<h1>parlar</h1> <a href="https://github.com/agent-sh/parlar/actions/workflows/ci.yml"><img src="https://github.com/agent-sh/parlar/actions/workflows/ci.yml/badge.svg" alt="CI"></a> <a href="https://crates.io/crates/parlar"><img src="https://img.shields.io/crates/v/parlar.svg" alt="crates.io"></a> <a href="https://www.npmjs.com/package/@agent-sh/parlar"><img src="https://img.shields.io/npm/v/@agent-sh/parlar.svg" alt="npm"></a> <a href="#license"><img src="https://img.shields.io/badge/License-MIT%20OR%20Apache--2.0-yellow.svg" alt="License: MIT OR Apache-2.0"></a>
Running open models for your company? Tiyuvta helps with model choice, deployment and optimization, and fine-tuning on your hardware or cloud account.
parlar is a voice conversation mode for Claude Code and Codex. You speak, and the session hears you whether it is idle or in the middle of work. It answers out loud in short sentences, while tool calls, diffs and logs stay in the terminal as usual. The spoken exchange is printed in the session, so you can read back what was said.
Everything runs locally on the CPU. No audio and no transcript leaves your machine.
/parlar:talk or picked from the indicator. Other sessions never get your speech.parlard idles at about 20 MB with the mic closed. Models load when a conversation starts and unload two minutes after it stops.gnome-extensions enable parlar@avifenesh brings it back).| Status | |
|---|---|
| Linux x86_64, aarch64 | supported (prebuilt binaries and crates) |
| Windows x86_64 | supported: daemon, voice, recognizer, named pipe, login service, plugin launchers and the floating indicator (parlar-overlay), tested on Windows 11 |
| Audio | PipeWire |
| Claude Code | full: idle wake, mid-turn steering, spoken stop, transcript in the session |
| Codex | say tool, mid-turn steering, blocking Stop waiter. A brand-new session hears you after its first turn |
| Indicator | GNOME Shell 50 on Linux; parlar-overlay on Windows and macOS. Other Linux desktops work without the floating indicator |
| macOS (Apple silicon) | in progress: on CI it builds, fetches, speaks, recognizes, loads its launchd agents and runs the plugin paths, and the floating indicator starts; the mic, the speaker, the indicator on a real desktop and Claude Code itself are untested on a real Mac |
All options end the same way: parlar and parlard on your PATH, the models downloaded once (about 800 MB) into $XDG_DATA_HOME/parlar, and the parlard user service running.
/plugin marketplace add agent-sh/parlar
/plugin install parlar@parlar
/parlar:setup
/parlar:setup checks your machine and installs what is missing, asking before each step that changes your system.
scripts/install.sh from a cloneNeeds cargo (rustup.rs), git, curl, clang, and the PipeWire and ALSA headers (Debian/Ubuntu libpipewire-0.3-dev libasound2-dev, Fedora pipewire-devel alsa-lib-devel, Arch pipewire alsa-lib).
git clone https://github.com/agent-sh/parlar && cd parlar
scripts/install.sh
It builds, installs into ~/.local (set PREFIX to change it), and downloads libmoonshine and the models. It also writes the plugins into a local marketplace, installs the GNOME indicator, and starts the service. --no-service skips the service.
npm install -g @agent-sh/parlar
parlard fetch # libmoonshine and the models, once
parlard service # systemd user service for this parlard
The package downloads the release binaries for your machine and checks their sha256. If your npm blocks install scripts, add --allow-scripts=@agent-sh/parlar.
On Windows, the npm package ships everything, including the Visual C++ runtime and the floating indicator. Then run parlard fetch and parlard service, which starts parlard and the indicator at login through your user's Run key.
cargo installcargo install parlar parlard
parlard fetch # libmoonshine and the models, once
parlard service # systemd user service for this parlard
On Windows, cargo install needs the Visual C++ Redistributable installed; parlard fetch adds ONNX Runtime. On macOS, parlard fetch adds ONNX Runtime and links it next to the parlard binary (macOS binds it at launch), and parlard service installs a launchd agent (~/Library/LaunchAgents/dev.agent-sh.parlard.plist).
Each release has parlar-<version>-<target>.tar.gz with a .sha256 next to it. Put parlar and parlard on your PATH, then run parlard fetch and parlard service.
The GNOME indicator is in shell/gnome/parlar@avifenesh. Copy it to ~/.local/share/gnome-shell/extensions/ and run gnome-extensions enable parlar@avifenesh. GNOME on Wayland loads new extensions at your next login.
On macOS the binaries are not notarized, so a downloaded tarball carries Apple's quarantine mark: they run from a terminal, but Gatekeeper refuses a double-click and spctl reports them as rejected. To clear the mark after unpacking:
xattr -dr com.apple.quarantine parlar-*-aarch64-apple-darwin
parlard fetch then downloads the models and links ONNX Runtime beside the binaries, and parlard service installs the launchd agent.
claude plugin marketplace add agent-sh/parlar # or the local one install.sh printed
claude plugin install parlar@parlar
parlar setup claude # say speaks without a permission prompt
While a conversation is on, Claude Code shows a voice strip above the prompt: what parlar is doing, a level meter, the last line heard or spoken, and mute, voice, stop and "talk here" buttons (ctrl+x tab, then the hotkey). It is a function-hooks module in the plugin (plugin/claude/hooks/strip.tsx), an early access Claude Code API, run on 2.1.287. The same module draws the spoken exchange in the transcript as "you ▸" and "parlar ▸" lines where it happened, and /parlar opens a voice pane with the conversation, the sessions to talk to, and the mic and speaker.
Optional status line: set statusLine.command to parlar status, with refreshInterval: 1.
codex plugin marketplace add ~/.local/share/parlar/marketplace/codex
codex plugin add parlar@parlar
Trust the parlar hooks once from Codex's hooks review prompt. Codex has no background wake, so a brand-new Codex session hears you after its first turn; after that, speaking continues the conversation. Start with parlar ctl talk --session <id> --harness codex, or from the indicator menu.
/parlar:talk | start talking to this session (it gets voice focus) |
/parlar:stop | stop; the mic closes |
/parlar:mute, /parlar:mute off | mute or unmute the mic; the conversation stays on |
| click the swarm | stop or start |
| right-click the swarm | pick the session, input and output devices, mute, voice off |
While the agent works you can steer it ("also update the tests"), and "stop" or "wait" blocks its next tool call. When it is idle, speaking wakes it.
From the shell:
parlar ctl state # what parlar is doing and which session has focus
parlar ctl devices # inputs and outputs
parlar ctl input <id> # switch the mic (remembered across restarts)
parlar ctl output <id> # switch the speaker
parlar ctl voice-off # captions only, no audio out
parlar status # one line, for a status line
Bluetooth headsets have two modes. High-quality playback (A2DP) turns the headset mic off, so parlar then needs another mic, such as the laptop's. Hands-free mode has a working mic and lower-quality audio, and parlar works with it. parlar ctl input pipewire:input_default makes the mic follow the system default.
parlard fetch downloads everything once. Each file is pinned to a revision or release and checked by size and checksum (sha256, or CRC32C from Moonshine's manifest for Kokoro).
| Model | Role | On disk | License | Source |
|---|---|---|---|---|
| Phonon-2 ONNX, int8 encoder | speech recognition | 657 MB | CC-BY-4.0 | tiyuvta/Phonon-2-ONNX |
| Kokoro-82M | voice | 104 MB | Apache-2.0 | through libmoonshine |
| Smart Turn v3.2 | end of turn | 8 MB | BSD-2-Clause | pipecat-ai/smart-turn-v3 |
| Silero VAD v6 | voice activity | 2 MB | MIT | snakers4/silero-vad |
| libmoonshine v0.1.5 + ONNX Runtime 1.23 | runtime | 33 MB | MIT | moonshine-ai/moonshine |
Phonon-2 is Fermion Research's English recognizer derived from NVIDIA's Parakeet TDT 0.6B v3. We exported it to ONNX and publish it on Hugging Face. The int8 encoder scores 4.35% WER on LibriSpeech dev-clean, against 4.26% for the exact fp32 export. The model card has the details.
Settings live in ~/.config/parlar/config.toml ($XDG_CONFIG_HOME/parlar/config.toml). Every key is optional; without the file parlar is English with Phonon-2 and Kokoro's af_heart voice.
language = "es" # what you speak and what the voice speaks
[recognizer]
model = "parakeet-tdt-0.6b-v3" # phonon-2 (English), parakeet-tdt-0.6b-v3, or a model directory
encoder = "int8" # int8, exact4x2 or fp32, when the model ships them
[voice]
name = "kokoro_ef_dora" # any voice libmoonshine has for the language
command = ["espeak-ng", "-v", "es"] # or speak through any program that reads text on stdin
After a change, run parlard fetch (it downloads only what the new settings need) and restart the service. parlard config prints what is in use.
| Language | Recognizer | Voice |
|---|---|---|
English (en, en-gb) | Phonon-2 (default) or Parakeet | Kokoro |
Spanish, French, Italian, Portuguese (es, fr, it, pt) | Parakeet | Kokoro |
German, Russian (de, ru) | Parakeet | Piper (through libmoonshine) |
Japanese, Chinese, Hindi (ja, zh, hi) | a model directory you supply | Kokoro |
18 more European languages Parakeet knows (pl, nl, sv, uk, ...) | Parakeet | your [voice] command |
A model directory is any ONNX export in the onnx-asr nemo-conformer-tdt layout (encoder, decoder_joint, vocab.txt, and a preprocessor). Hebrew is not covered by either recognizer or libmoonshine yet. Outside English the turn rules rely on punctuation and the Smart Turn model, which is multilingual, rather than on English filler words.
Measured on a Core Ultra 9 275HX laptop:
| State | RAM | CPU |
|---|---|---|
| idle (no conversation) | about 20 MB | none, mic closed |
| conversation on, nobody talking | about 1.4 GB | about 2% of one core |
| a turn ends | peak about 1.4 GB | 1 to 2 s of recognition on 4 cores |
| speaking | first audio in 0.3 to 0.5 s, synthesis about 2x faster than real time |
Two minutes after you stop, the models and the voice unload and the memory goes back to the system.
mic -> AEC/NS/AGC -> Silero VAD -> (pause) Phonon-2 -> endpointer + Smart Turn -> utterance
|
harness <- hooks + MCP (parlar) <- unix socket <- parlard (focus, routing, voice) <--+
|
+-> say tool -> parlard -> Kokoro -> speaker (barge-in cuts it)
parlar is the light binary: the MCP server, the hooks and ctl. Hooks run on every tool call, so they use no async runtime, return in milliseconds, and do nothing when parlard is not running.parlard is the daemon. It owns the mic, the speaker, the models and which session has focus, and it listens on $XDG_RUNTIME_DIR/parlar/parlar.sock.Audio is processed in memory and never written to disk, unless you set PARLAR_DUMP_TURNS for debugging. Transcripts go only to the focused session, through the plugin. Nothing is sent over the network after parlard fetch.
journalctl --user -u parlard -f shows what parlar hears (hearing:), where each utterance went (heard u3), and when models load and unload.parlar ctl state (is it on, and is a session focused?), then the mic with parlar ctl devices. A Bluetooth headset in A2DP mode has no mic./parlar:talk in the session you want.parlard fetch downloads missing files again. parlard speak "hello" writes a test line to parlar-speak.wav, and parlard final file.wav runs a recording through the recognizer.systemctl --user disable --now parlard
claude plugin uninstall parlar@parlar
rm -rf ~/.local/bin/parlar ~/.local/bin/parlard ~/.local/share/parlar \
~/.config/parlar ~/.config/systemd/user/parlard.service \
~/.local/share/gnome-shell/extensions/parlar@avifenesh
See CONTRIBUTING.md. Conventions for agents working on this repo are in AGENTS.md.
parlar is MIT or Apache-2.0, at your option (LICENSE-MIT, LICENSE-APACHE). The models and libraries it downloads keep their own licenses, listed under Models. It links sonora (BSD-3-Clause) and loads ONNX Runtime (MIT).
hooks/strip.tsx 636 lines1// The voice strip above the prompt: what parlard is doing, the last line heard or spoken, and
2// buttons for mute, voice, stop and focus. /parlar (or the strip's pane button) opens the voice
3// pane: the conversation so far, the sessions to talk to, and the mic and speaker. It follows `parlar ctl watch` and draws nothing while
4// the conversation is off or parlard is not running. The command hooks in hooks.json carry the
5// conversation itself; this module only shows it.
6import { atom, read, update } from 'claude-code'
7import type { EngineInterface, Register } from 'claude-code'
8
9import type { Device, HistoryLine, Panel, Phase, Rows, Section, SessionRow, Strip } from '../types'
10import { AGENT, registerRows, USER, withNote } from './rows'
11
12const IDLE: Strip = {
13 connected: false,
14 phase: 'stopped',
15 muted: false,
16 voiceOff: false,
17 focused: false,
18 levels: [],
19 caption: null,
20 flare: false,
21}
22const strip = atom({ plugin: 'parlar', key: 'strip' } as const, IDLE)
23// the rows rows.tsx draws; the strip adds parlar's notes, which only the watch stream carries
24const rows = atom({ plugin: 'parlar', key: 'rows' } as const, { wakes: {}, calls: {} } as Rows)
25const panel = atom(
26 { plugin: 'parlar', key: 'panel' } as const,
27 { history: [], sessions: [], devices: null } as Panel,
28)
29
30const PANE = 'parlar'
31// lines the pane keeps, and devices it lists per kind
32const HISTORY = 200
33const DEVICES = 6
34
35const BARS = '▁▂▃▄▅▆▇█'
36const METER = 8
37// level events come at 30 Hz; the terminal redraws at this pace at most
38const FRAME_MS = 125
39const CAPTION_MS = 15_000
40// the longest phase word ('connecting'), so the line does not shift between phases
41const WORD_WIDTH = 10
42// cells of the strip's fixed parts: '● voice ' and the padded word, then each label and button
43// as the terminal draws it ('m: mute'), each with the gap before it
44const CELLS = {
45 phase: 8 + 10,
46 caption: 1,
47 meter: 1 + 8,
48 notFocused: 1 + 16,
49 captionsOnly: 1 + 13,
50 talk: 1 + 12,
51 mute: 1 + 9,
52 voice: 1 + 12,
53 stop: 1 + 7,
54 pane: 1 + 13,
55}
56
57/** What fits in `width`: mute and stop always (and talk here off focus), the rest by priority. */
58export function fits(width: number, s: { focused: boolean; voiceOff: boolean }, metering: boolean) {
59 // the caption box is always in the row, so its gap is too; the button group's leading gap is
60 // the one its first button counts
61 let left = width - CELLS.phase - CELLS.caption - CELLS.mute - CELLS.stop - (s.focused ? 0 : CELLS.talk)
62 const take = (want: boolean, cells: number) => {
63 if (!want || left < cells) return false
64 left -= cells
65 return true
66 }
67 const meter = take(metering, CELLS.meter)
68 const voice = take(true, CELLS.voice)
69 const notFocused = take(!s.focused, CELLS.notFocused)
70 const captionsOnly = take(s.voiceOff, CELLS.captionsOnly)
71 const pane = take(true, CELLS.pane)
72 return { meter, voice, notFocused, captionsOnly, pane }
73}
74const FLARE_MS = 1_200
75
76// the phases in which someone is talking: the strip takes its own line only then
77const TALKING: ReadonlySet<Phase> = new Set(['listening', 'interrupting', 'speaking'])
78
79const WORDS: Record<Phase, string> = {
80 stopped: 'off',
81 connecting: 'connecting',
82 ready: 'ready',
83 listening: 'listening',
84 interrupting: 'listening',
85 working: 'working',
86 speaking: 'speaking',
87}
88
89type Event =
90 | { ui: 'phase'; phase: Phase; mic_muted: boolean; voice_off: boolean }
91 | { ui: 'levels'; user: number; agent: number }
92 | { ui: 'tool'; ok: boolean; name: string | null }
93 | { ui: 'caption'; who: string; text: string; session?: string; partial?: boolean; call?: string }
94 | { ui: 'notice'; text: string }
95
96export const meter = (levels: readonly number[]) =>
97 levels
98 .map(l => BARS[Math.min(BARS.length - 1, Math.floor(Math.sqrt(Math.max(0, l)) * BARS.length))])
99 .join('')
100
101// the module's own values: a reload starts them over, the strip itself lives in $.state
102const live = {
103 // the plugin's launcher, which finds parlar the way the command hooks do (PARLAR_BIN, the user
104 // install directories, npm); plain `parlar` from PATH where the launcher is a .cmd (Windows)
105 bin: 'parlar',
106 watching: false,
107 session: '',
108 phase: 'stopped' as Phase,
109 levels: [] as number[],
110 dirty: false,
111 captionAt: 0,
112 flareAt: 0,
113 // set once the watch has seen parlard, so the first connect of a session raises no toast
114 seen: false,
115}
116
117function set($: EngineInterface, patch: Partial<Strip>) {
118 return update($, strip, s => ({ ...s, ...patch }))
119}
120
121async function ctl($: EngineInterface, ...args: string[]) {
122 await $.process.run([live.bin, 'ctl', ...args], { timeoutMs: 5_000 }).catch(() => undefined)
123}
124
125async function refreshFocus($: EngineInterface) {
126 const r = await $.process.run([live.bin, 'ctl', 'state'], { timeoutMs: 2_000 }).catch(() => undefined)
127 if (!r || r.exitCode !== 0) return
128 try {
129 const st = JSON.parse(r.stdout) as { sessions?: Partial<SessionRow>[] }
130 const sessions = (st.sessions ?? []).flatMap(s =>
131 s.session
132 ? [{ session: s.session, cwd: s.cwd ?? '', harness: s.harness ?? '', focused: s.focused === true }]
133 : [],
134 )
135 const focused = sessions.some(s => s.session === live.session && s.focused)
136 const was = await read($, strip)
137 if (focused !== was.focused) {
138 await set($, { focused })
139 // typing elsewhere moves the voice without a sound here; say where it went
140 const to = sessions.find(s => s.focused)
141 if (was.focused && to && was.phase !== 'stopped') $.ui.toast(`Voice moved to ${folder(to.cwd)}`)
142 }
143 if (JSON.stringify(sessions) !== JSON.stringify((await read($, panel)).sessions)) {
144 await update($, panel, p => ({ ...p, sessions }))
145 }
146 } catch {
147 // a daemon mid-restart can answer half a line; the next poll reads it again
148 }
149}
150
151async function apply($: EngineInterface, ev: Event, now: number) {
152 switch (ev.ui) {
153 case 'phase': {
154 const back = live.seen && !(await read($, strip)).connected && ev.phase !== 'stopped'
155 live.seen = true
156 if (back) $.ui.toast('parlard is back')
157 live.phase = ev.phase
158 live.levels = []
159 await set($, {
160 connected: true,
161 phase: ev.phase,
162 muted: ev.mic_muted,
163 voiceOff: ev.voice_off,
164 levels: [],
165 })
166 void refreshFocus($).catch(quiet)
167 return
168 }
169 case 'levels': {
170 const phase = live.phase
171 const l =
172 phase === 'speaking' ? ev.agent : phase === 'listening' || phase === 'interrupting' ? ev.user : 0
173 live.levels = [...live.levels, l].slice(-METER)
174 live.dirty = true
175 return
176 }
177 case 'tool':
178 if (!ev.ok) {
179 live.flareAt = now
180 await set($, { flare: true })
181 }
182 return
183 case 'notice':
184 $.ui.toast(ev.text, { timeoutMs: 12_000 })
185 return
186 case 'caption': {
187 // a caption names its session; a partial, parlar's own line, or a daemon from before that
188 // field belongs to whoever has focus
189 const mine = ev.session !== undefined ? ev.session === live.session : (await read($, strip)).focused
190 if (!mine) return
191 live.captionAt = now
192 await set($, { caption: { who: ev.who, text: ev.text } })
193 // parlar's answer during a long tool call is drawn on that call's row
194 if (ev.call) {
195 const call = ev.call
196 await update($, rows, r => withNote(r, call, ev.text))
197 }
198 // the pane keeps finished lines, not the words so far
199 if (ev.partial) return
200 const line: HistoryLine = { who: ev.who === 'user' ? 'you' : 'parlar', text: ev.text }
201 await update($, panel, p => ({ ...p, history: [...p.history, line].slice(-HISTORY) }))
202 void toLatest($).catch(quiet)
203 return
204 }
205 }
206}
207
208async function watch($: EngineInterface) {
209 live.watching = true
210 let buf = ''
211 try {
212 for await (const { stream, text } of $.process.spawn({ argv: [live.bin, 'ctl', 'watch'] })) {
213 if (stream !== 'stdout') continue
214 buf += text
215 let nl: number
216 while ((nl = buf.indexOf('\n')) >= 0) {
217 const line = buf.slice(0, nl)
218 buf = buf.slice(nl + 1)
219 try {
220 await apply($, JSON.parse(line) as Event, await $.clock.now())
221 } catch {
222 // an event kind this build does not know yet
223 }
224 }
225 }
226 } catch {
227 // parlar is not on PATH: stay quiet, as the command hooks do
228 } finally {
229 live.watching = false
230 const was = await read($, strip)
231 await set($, { connected: false, levels: [] })
232 // the watch also ends when the module unloads; only a parlard that no longer answers is news
233 if (was.connected && was.phase !== 'stopped' && !(await answers($))) {
234 $.ui.toast('parlard stopped: voice is off until it is back', { timeoutMs: 8_000 })
235 }
236 }
237}
238
239async function answers($: EngineInterface) {
240 const r = await $.process.run([live.bin, 'ctl', 'state'], { timeoutMs: 2_000 }).catch(() => undefined)
241 return r?.exitCode === 0
242}
243
244// a timer that fires as the module unloads finds $ gone; the next load starts its own
245function quiet() {}
246
247async function frame($: EngineInterface) {
248 const now = await $.clock.now()
249 const s = await read($, strip)
250 const patch: Partial<Strip> = {}
251 if (live.dirty) {
252 live.dirty = false
253 patch.levels = live.levels
254 }
255 if (s.caption && now - live.captionAt > CAPTION_MS) patch.caption = null
256 if (s.flare && now - live.flareAt > FLARE_MS) patch.flare = false
257 if (Object.keys(patch).length > 0) await set($, patch)
258}
259
260async function refreshDevices($: EngineInterface) {
261 const r = await $.process.run([live.bin, 'ctl', 'devices'], { timeoutMs: 5_000 }).catch(() => undefined)
262 if (!r || r.exitCode !== 0) return
263 try {
264 const d = JSON.parse(r.stdout) as { inputs?: Device[]; outputs?: Device[] }
265 await update($, panel, p => ({ ...p, devices: { inputs: d.inputs ?? [], outputs: d.outputs ?? [] } }))
266 } catch {
267 // an answer this build cannot read leaves the pickers as they were
268 }
269}
270
271async function openPane($: EngineInterface) {
272 await set($, { paneOpen: true })
273 const placed = await $.ui.open({ id: PANE, title: 'Voice' })
274 await Promise.all([refreshFocus($), refreshDevices($)])
275 await toLatest($)
276 return placed
277}
278
279/** The pane opens on, and follows, the newest line of the conversation. */
280async function toLatest($: EngineInterface) {
281 if (!(await read($, strip)).paneOpen) return
282 await $.ui.scroll({ in: PANE, to: 'end' }).catch(quiet)
283}
284
285/** The pane button: opens the pane, or closes it when it is open. */
286async function togglePane($: EngineInterface) {
287 // kept in $.state, so a reload of the module still knows the pane is open
288 if (!(await read($, strip)).paneOpen) return void (await openPane($))
289 await set($, { paneOpen: false })
290 await $.ui.close({ id: PANE })
291}
292
293async function pickDevice($: EngineInterface, kind: 'input' | 'output', id: string) {
294 await ctl($, kind, id)
295 await update($, panel, p => ({ ...p, open: null }))
296 await refreshDevices($)
297}
298
299/** Give a session voice focus from the pane; this one is attached first if parlard lacks it. */
300async function talkTo($: EngineInterface, session: string, here: boolean) {
301 await (here ? ctl($, 'talk', '--session', session) : ctl($, 'focus', session))
302 await update($, panel, p => ({ ...p, open: null }))
303 await refreshFocus($)
304}
305
306function folder(cwd: string) {
307 return cwd.split(/[\\/]/).filter(Boolean).pop() ?? cwd
308}
309
310async function pollFocus($: EngineInterface) {
311 const s = await read($, strip)
312 if (s.connected && s.phase !== 'stopped') await refreshFocus($)
313}
314
315export const register: Register = on => {
316 registerRows(on)
317 on('session.start', async ($, e, next) => {
318 const started = await next(e)
319 if (!e.isInteractive) return started
320 // the hooks started from now on leave the heard and spoken lines to rows.tsx
321 await $.env.set('PARLAR_ROWS', '1')
322 live.session = await $.session.id()
323 const root = $.plugin.root
324 live.bin = root.includes('\\') ? 'parlar' : `${root}/bin/parlar`
325 // a reload starts the strip over, but a pane it opened is still open
326 await update($, strip, old => ({ ...IDLE, paneOpen: old.paneOpen }))
327 void watch($).catch(quiet)
328 // reconnect after parlard starts or restarts
329 $.clock.every(5_000, () => {
330 if (!live.watching) void watch($).catch(quiet)
331 })
332 $.clock.every(FRAME_MS, () => void frame($).catch(quiet))
333 // focus moves with typing in another session, which sends no event here
334 $.clock.every(2_000, () => void pollFocus($).catch(quiet))
335 await $.command.register({
336 name: 'parlar',
337 description: 'Open the voice pane: the conversation, the sessions, the mic and speaker',
338 immediate: true,
339 })
340 return started
341 })
342
343 // closed by the person (its mark, Esc) or an unload: the button opens it next time
344 on('ui.close', async ($, e, next) => {
345 const closed = await next(e)
346 if (e.id === PANE) {
347 await set($, { paneOpen: false })
348 }
349 return closed
350 })
351
352 on('command.run', { command: 'parlar' }, async $ => {
353 const { isPlaced } = await openPane($)
354 return { text: isPlaced ? 'Voice pane opened.' : 'The voice pane will open when there is room.' }
355 })
356
357 on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
358 const { Box, Button, Text } = $.ui.resolve(e)
359 const s = await read($, strip)
360 const p = await read($, panel)
361 const others = p.sessions.filter(x => x.session !== live.session)
362 const here = p.sessions.find(x => x.session === live.session)
363 const word = !s.connected ? 'parlard is not running' : s.muted ? 'mic muted' : WORDS[s.phase]
364 const focusedName = p.sessions.find(x => x.focused)
365 const conversing = s.connected && s.phase !== 'stopped'
366 // plain ALSA lists one card many times: one entry per name, the current one first
367 const named = (list: Device[]) => {
368 const one = list.filter((d, i) => d.current || list.findIndex(x => x.name === d.name) === i)
369 return [...one.filter(d => d.current), ...one.filter(d => !d.current)].slice(0, DEVICES)
370 }
371 const inputs = named(p.devices?.inputs ?? [])
372 const outputs = named(p.devices?.outputs ?? [])
373 const current = (list: Device[]) => list.find(d => d.current)?.name ?? 'none'
374 // one line per section until clicked open; picking closes it again
375 const header = (key: Section, title: string, value: string) => (
376 <Button
377 key={`open-${key}`}
378 label={`${p.open === key ? '▾' : '▸'} ${title}: ${value}`}
379 plain
380 onPress={() => update($, panel, x => ({ ...x, open: x.open === key ? null : key }))}
381 />
382 )
383 const pick = (kind: 'input' | 'output', d: Device) =>
384 d.current ? (
385 <Text> ● {d.name}</Text>
386 ) : (
387 <Button
388 key={`${kind}-${d.id}`}
389 label={` ○ ${d.name}`}
390 plain
391 onPress={() => pickDevice($, kind, d.id)}
392 />
393 )
394 return (
395 <Box flexDirection="column">
396 {/* the one place with every control in every state (idle, muted, stopped), on the
397 status line itself so the pane stays the conversation */}
398 <Box flexDirection="row" gap={2}>
399 <Text bold color={s.focused ? USER : undefined}>
400 {s.focused ? '●' : '○'} voice {word}
401 </Text>
402 {conversing && (
403 <Button
404 key="pane-mute"
405 label={s.muted ? 'unmute' : 'mute'}
406 plain
407 onPress={() => ctl($, s.muted ? 'unmute' : 'mute')}
408 />
409 )}
410 {conversing && (
411 <Button
412 key="pane-voice"
413 label={s.voiceOff ? 'voice on' : 'voice off'}
414 plain
415 onPress={() => ctl($, s.voiceOff ? 'voice-on' : 'voice-off')}
416 />
417 )}
418 {s.connected && (
419 <Button
420 key="pane-power"
421 label={conversing ? 'stop' : 'start'}
422 plain
423 onPress={() => ctl($, conversing ? 'off' : 'on')}
424 />
425 )}
426 <Button key="pane-close" label="close" plain onPress={() => togglePane($)} />
427 </Box>
428 {header('sessions', 'Talking to', focusedName ? folder(focusedName.cwd) : 'nobody')}
429 {p.open === 'sessions' && (
430 <Box flexDirection="column">
431 {here && (
432 <Box flexDirection="row" gap={1}>
433 <Text>
434 {' '}
435 {here.focused ? '●' : '○'} {folder(here.cwd)} (this one)
436 </Text>
437 {!here.focused && (
438 <Button
439 key="attach"
440 label="talk here"
441 plain
442 onPress={() => talkTo($, live.session, true)}
443 />
444 )}
445 </Box>
446 )}
447 {!here && s.connected && (
448 // talk attaches this session if parlard does not know it yet, then focuses it
449 <Button key="attach" label=" talk here" plain onPress={() => talkTo($, live.session, true)} />
450 )}
451 {others.map(o => (
452 <Box flexDirection="row" gap={1}>
453 <Text dimColor={!o.focused}>
454 {' '}
455 {o.focused ? '●' : '○'} {folder(o.cwd)} {o.harness}
456 </Text>
457 {!o.focused && (
458 <Button
459 key={`focus-${o.session}`}
460 label="talk there"
461 plain
462 onPress={() => talkTo($, o.session, false)}
463 />
464 )}
465 </Box>
466 ))}
467 </Box>
468 )}
469 {p.devices && header('input', 'Mic', current(inputs))}
470 {p.open === 'input' && <Box flexDirection="column">{inputs.map(d => pick('input', d))}</Box>}
471 {p.devices && header('output', 'Speaker', current(outputs))}
472 {p.open === 'output' && <Box flexDirection="column">{outputs.map(d => pick('output', d))}</Box>}
473 <Box flexDirection="column" marginTop={1}>
474 {p.history.length === 0 && <Text dimColor>Nothing said yet.</Text>}
475 {p.history.map(l => (
476 <Text color={l.who === 'you' ? USER : AGENT} wrap="wrap">
477 {l.who} ▸ {l.text}
478 </Text>
479 ))}
480 </Box>
481 </Box>
482 )
483 })
484
485 on('prompt.submit', async ($, e, next) => {
486 // typing here takes focus; show it without waiting for the poll
487 const r = await next(e)
488 void refreshFocus($).catch(quiet)
489 return r
490 })
491
492 // while a turn runs and nobody talks: a voice mark and the buttons after Claude's own spinner.
493 // The spinner's row takes no keys, so its buttons are for the pointer and carry no hotkeys;
494 // /parlar opens the pane for the keyboard
495 on('ui.render', { component: 'Spinner' }, async ($, e, next) => {
496 const s = await read($, strip)
497 const spinner = await next(e)
498 if (!s.connected || s.phase === 'stopped' || TALKING.has(s.phase)) return spinner
499 const { Box, Button, Text } = $.ui.resolve(e)
500 return (
501 // the engine's spinner opens with a blank row and may end with a tip: the controls sit on its
502 // second row, the spinner's own text
503 <Box flexDirection="row" gap={2}>
504 {spinner}
505 <Box flexDirection="row" gap={1} flexShrink={0} marginTop={1}>
506 <Text color={s.muted ? undefined : AGENT} dimColor>
507 {s.focused ? '●' : '○'} {s.muted ? 'mic muted' : 'voice'}
508 </Text>
509 {!s.focused && (
510 <Button
511 key="spin-focus"
512 label="talk here"
513 plain
514 onPress={() => ctl($, 'talk', '--session', live.session)}
515 />
516 )}
517 <Button
518 key="spin-mute"
519 label={s.muted ? 'unmute' : 'mute'}
520 plain
521 onPress={() => ctl($, s.muted ? 'unmute' : 'mute')}
522 />
523 <Button key="spin-stop" label="stop" plain onPress={() => ctl($, 'off')} />
524 <Button
525 key="spin-pane"
526 label={s.paneOpen ? 'close pane' : 'pane'}
527 plain
528 onPress={() => togglePane($)}
529 />
530 </Box>
531 </Box>
532 )
533 })
534
535 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
536 const s = await read($, strip)
537 if (e.props.hasSurvey || !s.connected || s.phase === 'stopped') return next(e)
538 const { Box, Button, Text } = $.ui.resolve(e)
539 // whatever else the band holds stays, under the strip
540 const below = await next(e)
541
542 const talking = s.phase === 'speaking' ? AGENT : USER
543 const color = s.flare ? 'red' : s.muted ? undefined : s.phase === 'working' ? AGENT : talking
544 const word = s.muted ? 'mic muted' : WORDS[s.phase]
545 const metering = s.phase === 'listening' || s.phase === 'interrupting' || s.phase === 'speaking'
546 const who = s.caption?.who === 'user' ? 'you' : s.caption?.who === 'parlar' ? 'parlar' : 'agent'
547 // the engine draws its collapse mark in the last columns
548 const width = Math.max(20, e.props.bodyColumns - 4)
549 const fit = fits(width, s, metering)
550 // while a turn runs and nobody talks, the spinner line carries the voice and its buttons
551 if (!metering && e.props.isWorking) return below
552
553 return (
554 <Box flexDirection="column">
555 <Box flexDirection="row" gap={1} width={width}>
556 {/* fixed widths and no wrapping: the strip stays one line, and its parts do not shift
557 as the phase and the meter change */}
558 <Box flexShrink={0}>
559 <Text color={color} dimColor={s.muted || !s.focused} bold={s.focused} wrap="truncate-end">
560 {s.focused ? '●' : '○'} voice {word.padEnd(WORD_WIDTH)}
561 </Text>
562 </Box>
563 {fit.meter && (
564 <Box flexShrink={0}>
565 <Text color={talking} wrap="truncate-end">
566 {meter(s.levels).padEnd(METER)}
567 </Text>
568 </Box>
569 )}
570 {fit.notFocused && (
571 <Box flexShrink={0}>
572 <Text dimColor wrap="truncate-end">
573 not focused here
574 </Text>
575 </Box>
576 )}
577 {fit.captionsOnly && (
578 <Box flexShrink={0}>
579 <Text dimColor wrap="truncate-end">
580 captions only
581 </Text>
582 </Box>
583 )}
584 {/* the caption takes what is left; a long one keeps its last words */}
585 <Box flexGrow={1} flexShrink={1} minWidth={0}>
586 {s.focused && s.caption && (
587 <Text dimColor wrap="truncate-start">
588 {who}: {s.caption.text}
589 </Text>
590 )}
591 </Box>
592 <Box flexShrink={0} gap={1}>
593 {!s.focused && (
594 // talk attaches a session that started before parlard, which focus alone cannot
595 <Button
596 key="focus"
597 label="talk here"
598 hotkey="t"
599 plain
600 onPress={() => ctl($, 'talk', '--session', live.session)}
601 />
602 )}
603 <Button
604 key="mute"
605 label={s.muted ? 'unmute' : 'mute'}
606 hotkey="m"
607 plain
608 onPress={() => ctl($, s.muted ? 'unmute' : 'mute')}
609 />
610 {fit.voice && (
611 <Button
612 key="voice"
613 label={s.voiceOff ? 'voice on' : 'voice off'}
614 hotkey="v"
615 plain
616 onPress={() => ctl($, s.voiceOff ? 'voice-on' : 'voice-off')}
617 />
618 )}
619 <Button key="stop" label="stop" hotkey="s" plain onPress={() => ctl($, 'off')} />
620 {fit.pane && (
621 <Button
622 key="pane"
623 label={s.paneOpen ? 'close pane' : 'pane'}
624 hotkey="p"
625 plain
626 onPress={() => togglePane($)}
627 />
628 )}
629 </Box>
630 </Box>
631 {below}
632 </Box>
633 )
634 })
635}
636hooks/rows.tsx 214 lines1// The spoken exchange in the transcript: each say call draws as a "parlar" line where the call
2// sits, and each utterance as a "you" line on the row that delivered it (the voice wake row, or
3// the tool call whose hook passed it on). Without this module the command hooks print the same
4// lines as hook messages at the next tool call; PARLAR_ROWS, set in strip.tsx's session.start, tells
5// them to leave those lines here.
6// ctrl+o shows the wake rows and say groups as the engine draws them.
7import { atom, read, update } from 'claude-code'
8import type { EngineInterface, On } from 'claude-code'
9
10import type { Rows } from '../types'
11
12export const USER = '#5fafff'
13export const AGENT = '#ffaf00'
14
15const SAY = 'mcp__plugin_parlar_parlar__say'
16// what format::utterances and the wake message in hooks.json write
17const WAKE = 'The user spoke to you by voice:'
18const ACTIVATED = 'Voice mode is on.'
19// how mcp.rs starts a say result that was not spoken
20const NOT_SPOKEN = ['Not spoken', 'Voice mode is off']
21// what mcp.rs and hook.rs put ahead of utterances in a tool result: any other tool's output is
22// the tool's own, even when it quotes a delivery
23const RESULT_CUES = ['The user said meanwhile:', 'The user asked you to stop']
24// rows kept per session; older ones fall back to the engine's drawing
25const KEEP = 300
26
27const rows = atom({ plugin: 'parlar', key: 'rows' } as const, { wakes: {}, calls: {} } as Rows)
28
29// the tool call a later PostToolUse context row belongs to: the last result appended
30const last = { call: '' }
31
32/** The utterances a delivery carries, cleaned of the reminder and the notes around them. */
33export function heardIn(text: string): string[] {
34 const out: string[] = []
35 for (const m of text.matchAll(/\[voice u\d+(?: continues u\d+)?\] ([\s\S]*?)(?=\n\[voice u|\n\(|$)/g)) {
36 const lines = (m[1] ?? '').split('\n').filter(l => !l.startsWith(ACTIVATED))
37 const said = lines.join(' ').trim()
38 if (said) out.push(said)
39 }
40 return out
41}
42
43type Block = { type: string; text?: string; tool_use_id?: string; content?: string | Block[] }
44
45function textOf(content: string | readonly Block[] | undefined): string {
46 if (typeof content === 'string') return content
47 return (content ?? [])
48 .map(b => (b.type === 'text' ? (b.text ?? '') : ''))
49 .filter(Boolean)
50 .join('\n')
51}
52
53function keep(map: Record<string, string[]>, key: string, heard: string[]) {
54 const next = { ...map, [key]: [...(map[key] ?? []), ...heard] }
55 const keys = Object.keys(next)
56 for (const k of keys.slice(0, Math.max(0, keys.length - KEEP))) delete next[k]
57 return next
58}
59
60function remember($: EngineInterface, kind: 'wakes' | 'calls', key: string, heard: string[]) {
61 return update($, rows, r => ({ ...r, [kind]: keep(r[kind], key, heard) }))
62}
63
64function sayText(input: unknown): string | undefined {
65 const t = (input as { text?: unknown } | null)?.text
66 return typeof t === 'string' && t.trim() ? t.trim() : undefined
67}
68
69function outputText(output: unknown): string {
70 return Array.isArray(output) ? textOf(output as Block[]) : typeof output === 'string' ? output : ''
71}
72
73/** Why a say was not spoken, from its result: the first sentence of mcp.rs's note. */
74function notSpoken(output: unknown): string | undefined {
75 const t = outputText(output)
76 if (!NOT_SPOKEN.some(n => t.startsWith(n))) return undefined
77 return t.split(/(?<=\.) /)[0]?.replace(/\.$/, '')
78}
79
80/** The utterances a tool result passed on, after the cue parlar writes ahead of them. */
81function heardInResult(text: string): string[] {
82 const cue = RESULT_CUES.map(c => text.indexOf(c)).find(i => i >= 0)
83 return cue === undefined ? [] : heardIn(text.slice(cue))
84}
85
86type Call = { tool_use_id?: string; tool: string; input: unknown; output?: unknown }
87type Line = { who: 'you' | 'parlar' | 'note'; text: string; note?: string }
88
89/** The lines one tool call adds: what it spoke, then what the user said that it passed on. */
90function linesOf(c: Call, r: Rows): Line[] {
91 const out: Line[] = []
92 if (c.tool === SAY) {
93 const t = sayText(c.input)
94 if (t) out.push({ who: 'parlar', text: t, note: notSpoken(c.output) })
95 }
96 // a say result or a refused call's reason carries them in the row's own output; a PostToolUse
97 // context row only in what session.append noted
98 const heard = [...heardInResult(outputText(c.output)), ...(r.calls[c.tool_use_id ?? ''] ?? [])]
99 for (const h of heard) out.push({ who: 'you', text: h })
100 // parlar's own answer while the call ran: a system note, not the agent speaking
101 for (const n of r.notes?.[c.tool_use_id ?? ''] ?? []) out.push({ who: 'note', text: n })
102 return out
103}
104
105type Appended = { door: string; origin: { kind: string }; uuid: string; message: unknown; agentId?: string }
106type Note = { kind: 'wakes' | 'calls'; key: string; heard: string[] }
107
108/** What one appended row tells about which row delivered which utterances. */
109export function notesOf(e: Appended, lastCall: string): { notes: Note[]; lastCall: string } {
110 const m = e.message as { name?: string; content?: string | Block[] }
111 const notes: Note[] = []
112 // speech goes to the main thread; a subagent's rows would move lastCall off it
113 if (e.agentId) return { notes, lastCall }
114 if (e.door === 'prompt' && e.origin.kind === 'task-notification') {
115 const text = textOf(m.content)
116 if (text.includes(WAKE))
117 notes.push({ kind: 'wakes', key: e.uuid, heard: heardIn(text.slice(text.indexOf(WAKE))) })
118 }
119 // a tool result's own utterances are read from the row's output; here it only marks the call
120 // a following PostToolUse context row belongs to
121 if (e.door === 'tool-result' && Array.isArray(m.content)) {
122 for (const b of m.content) if (b.type === 'tool_result' && b.tool_use_id) lastCall = b.tool_use_id
123 }
124 if (e.door === 'hook-context' && m.name === 'hook_additional_context') {
125 notes.push({ kind: 'calls', key: lastCall, heard: heardIn(textOf(m.content)) })
126 }
127 return { notes: notes.filter(n => n.key && n.heard.length > 0), lastCall }
128}
129
130type TextOf = ReturnType<EngineInterface['ui']['resolve']>['Text']
131
132/** One drawn line: the person in blue, the agent in amber, parlar's own notes dim and plain. */
133function drawLine(l: Line, Text: TextOf) {
134 if (l.who === 'note') {
135 return (
136 <Text dimColor italic>
137 · {l.text}
138 </Text>
139 )
140 }
141 return (
142 <Text color={l.who === 'you' ? USER : AGENT} dimColor={l.note !== undefined}>
143 {l.who} ▸ {l.text}
144 {l.note ? ` (${l.note})` : ''}
145 </Text>
146 )
147}
148
149/** Keep parlar's note on the tool call it answered during, for that call's row. */
150export function withNote(r: Rows, call: string, text: string): Rows {
151 const notes = keep(r.notes ?? {}, call, [text])
152 return { ...r, notes }
153}
154
155export function registerRows(on: On) {
156 on('session.append', async ($, e, next) => {
157 const { notes, lastCall } = notesOf(e, last.call)
158 last.call = lastCall
159 for (const n of notes) await remember($, n.kind, n.key, n.heard)
160 return next(e)
161 })
162
163 on('ui.render', { component: 'UserMessage' }, async ($, e, next) => {
164 if (e.props.isExpanded || e.props.origin.kind !== 'task-notification') return next(e)
165 const heard = (await read($, rows)).wakes[e.requestId ?? '']
166 if (!heard) return next(e)
167 const { Box, Text } = $.ui.resolve(e)
168 return (
169 <Box flexDirection="column">
170 {heard.map(said => (
171 <Text color={USER}>you ▸ {said}</Text>
172 ))}
173 </Box>
174 )
175 })
176
177 on('ui.render', { component: 'ToolGroup' }, async ($, e, next) => {
178 if (e.props.isExpanded) return next(e)
179 const r = await read($, rows)
180 const lines = e.props.calls.flatMap(c => linesOf(c, r))
181 if (lines.length === 0) return next(e)
182 const { Box, Text } = $.ui.resolve(e)
183 // a group of say calls alone is the lines; any other call keeps the engine's summary above
184 const onlySay = e.props.calls.every(c => c.tool === SAY)
185 const above = onlySay ? null : await next(e)
186 return (
187 <Box flexDirection="column">
188 {above}
189 {lines.map(l => drawLine(l, Text))}
190 </Box>
191 )
192 })
193
194 on('ui.render', { component: 'ToolUse' }, async ($, e, next) => {
195 const lines = linesOf(e.props, await read($, rows))
196 if (lines.length === 0) return next(e)
197 const { Box, Text } = $.ui.resolve(e)
198 const above = e.props.tool === SAY ? null : await next(e)
199 return (
200 <Box flexDirection="column">
201 {above}
202 {lines.map(l => drawLine(l, Text))}
203 </Box>
204 )
205 })
206
207 // a say call's result is "Said." plus what it passed on, which its row already shows
208 on('ui.render', { component: 'ToolResult' }, async ($, e, next) => {
209 if (e.props.tool !== SAY || e.props.isErrored) return next(e)
210 const { Box } = $.ui.resolve(e)
211 return <Box />
212 })
213}
214types/index.d.ts 58 lines1export type Phase = 'stopped' | 'connecting' | 'ready' | 'listening' | 'interrupting' | 'working' | 'speaking'
2
3/** What the strip above the prompt draws, fed by `parlar ctl watch`. */
4export type Strip = {
5 /** False while parlard is not running. */
6 connected: boolean
7 phase: Phase
8 muted: boolean
9 voiceOff: boolean
10 /** Whether this session has voice focus. */
11 focused: boolean
12 /** Recent levels, 0..1, of whoever is talking. */
13 levels: number[]
14 caption: { who: string; text: string } | null
15 /** A failed tool call, shown red until it fades. */
16 flare: boolean
17 /** The voice pane is open: its buttons are there, so the strip leaves them out. */
18 paneOpen?: boolean
19}
20
21/** One line of the conversation as the pane lists it. */
22export type HistoryLine = { who: 'you' | 'parlar'; text: string }
23
24/** A session parlard knows, as `parlar ctl state` lists it. */
25export type SessionRow = { session: string; cwd: string; harness: string; focused: boolean }
26
27/** An audio device, as `parlar ctl devices` lists it. */
28export type Device = { id: string; name: string; current: boolean }
29
30/** A pane section that opens on a click: the sessions, the mic or the speaker. */
31export type Section = 'sessions' | 'input' | 'output'
32
33/** What the voice pane draws. */
34export type Panel = {
35 /** The section clicked open, one at a time. */
36 open?: Section | null
37 /** The conversation heard and spoken while this session had focus, oldest first. */
38 history: HistoryLine[]
39 sessions: SessionRow[]
40 devices: { inputs: Device[]; outputs: Device[] } | null
41}
42
43/** Utterances by the transcript row that delivered them, for drawing that row. */
44export type Rows = {
45 /** By the uuid of a voice wake's user row. */
46 wakes: Record<string, string[]>
47 /** By the tool call whose PostToolUse context passed them on. */
48 calls: Record<string, string[]>
49 /** parlar's own notes (its answer to words heard during a long call), by that call. */
50 notes?: Record<string, string[]>
51}
52
53declare module 'claude-code' {
54 interface PluginState {
55 parlar: { strip: Strip; rows: Rows; panel: Panel }
56 }
57}
58