SLOPSHOPPER

speak-aloud

Reads Claude's responses aloud with local Kokoro neural voices (macOS say as fallback), detecting English and Spanish per paragraph: a speaker button on each…

newrowscommandtoaststatusprocess
v0.3.0no licenseupdated 2026-10-06tiger3645/claude-tts
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · speak-aloud
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ⟨Claude Code's own drawing⟩ [ 🔊 ] ✻ Worked for 42s · done 4:20 PM › /speak ⎿ speak-aloud: Reading the latest response aloud. ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Claude's reply
⟨Claude Code's own drawing⟩ [ 🔊 ]
README

speak-aloud

A Claude Code mod that reads Claude's responses aloud.

Every reply in the transcript gets a 🔊 button. Press it to hear that reply; it turns into ■ while reading, and pressing it again stops. Markdown is cleaned up before speaking: formatting marks and link URLs are dropped, and code blocks are skipped (a reply that is only code gets no button).

Speech uses Kokoro, a small neural TTS model that runs locally on Apple silicon through mlx-audio. If Kokoro isn't installed or its server fails, the mod falls back to macOS say automatically (one toast says so).

Each paragraph is read in its own language — English or Spanish — with a voice per language.

Requires macOS. Kokoro needs Apple silicon; say works everywhere on macOS.

Install

  1. In a Claude Code session in a terminal:
   /plugin install speak-aloud --marketplace tiger3645/claude-tts

Answer y to add the marketplace, then pick a scope (user scope loads it in every session). It works right away with macOS say.

  1. For the natural Kokoro voices, run:
   /speak-install

It needs uv (brew install uv, or curl -LsSf https://astral.sh/uv/install.sh | sh). The install runs in the background with its progress in the status line: it installs mlx-audio and its extras, the spaCy English model, starts the server and downloads the model (~680 MB the first time), warming English and Spanish. When it's done, reads switch to Kokoro — no restart. /speak-install --repair reinstalls everything.

Until Kokoro is installed, the mod reads with say and reminds you about /speak-install at most once a day.

How Kokoro runs

The mod starts mlx_audio.server on 127.0.0.1:8765 when a session starts (or reuses one already answering there) and loads the model, so reads start in well under a second. The server runs detached and is shared by every session on the machine: a /reload-plugins keeps it, and it is stopped when the last session using it ends. Its logs, the session registry and temporary WAVs live in ~/Library/Caches/speak-aloud/.

Text is split into paragraphs and chunks of a few sentences; the next chunk is generated while the current one plays (curl to the server, afplay to play).

Manual install / troubleshooting

/speak-install runs these two commands; run them yourself if it fails:

uv tool install mlx-audio --with uvicorn --with fastapi --with python-multipart --with webrtcvad-wheels --with "misaki[en]" --with num2words --with spacy --with phonemizer-fork --with espeakng-loader
uv pip install --python ~/.local/share/uv/tools/mlx-audio/bin/python https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.8.0/en_core_web_sm-3.8.0-py3-none-any.whl

Re-run the second command after any uv tool install --force or upgrade of mlx-audio.

The mod looks for ~/.local/bin/mlx_audio.server. If reads fall back to say, check /speak-engine and the server log in ~/Library/Caches/speak-aloud/logs/server.log.

Commands

CommandWhat it does
/speakRead Claude's latest response aloud
/speak-stopStop reading
/speak-installInstall Kokoro (needs uv); --repair reinstalls
/speak-engineShow the engine and whether Kokoro is installed / running
`/speak-engine kokoro\say`Choose the engine (default kokoro); default resets it
/speak-voiceShow the English and Spanish voices for the current engine, and the voices to choose from
`/speak-voice [en\es] <name>`Set the voice for a language (no language = English), e.g. /speak-voice es em_alex or /speak-voice Samantha
`/speak-voice [en\es] default`Reset a language's voice
/speak-rateShow the current speed
/speak-rate <wpm>Set the speed in words per minute (80–500); default resets to 190

Voices are kept per engine. Defaults:

EnglishSpanish
Kokoroaf_heartef_dora
saysystem voicefirst Spanish (es_*) voice installed

Kokoro voices: English af_*, am_* (American) and bf_*, bm_* (British); Spanish ef_dora, em_alex, em_santa. For say, any voice from say -v '?' (download more in System Settings → Accessibility → Spoken Content).

The speed applies to both engines: for Kokoro, words per minute map to its speed as wpm ÷ 153 (Kokoro's natural pace), clamped to 0.5–2.0.

Engine, voices and speed are saved across sessions.

Language detection

Each paragraph is classified as English or Spanish by counting common words of each language (the, and, is… vs el, que, de, para…), with ñ, ¿, ¡ and accented vowels counting toward Spanish. Inline code, identifiers (snake_case, camelCase, file.ts), paths and URLs are ignored, so "Hice commit del plugin y el build pasó" reads as Spanish. A paragraph too short to tell ("OK.") takes the language of the whole message; a message that can't be told reads as English.

Notes

  • When a reply is split by tool calls, each text part gets its own button.
  • /speak reads the whole latest reply but doesn't highlight a button while it reads; /speak-stop stops it.

Development

claude plugin validate .
claude plugin test .

To load a working copy without installing: claude --plugin-dir /path/to/claude-tts.

Source 8 files
hooks/register.tsx 815 lines
1import { atom, read, update } from 'claude-code'
2import type {
3  EngineInterface,
4  HookStream,
5  ProcessSpawnChunk,
6  ProcessSpawnRequest,
7  ProcessSpawnResult,
8  Register,
9} from 'claude-code'
10
11import {
12  INSTALL_NOTICE,
13  UV_HELP,
14  shouldShowInstallNotice,
15  spacyInstallArgv,
16  tail,
17  uvCandidates,
18  uvInstallArgv,
19} from './install'
20import type { Lang } from './language'
21import { stripMarkdown } from './markdown'
22import {
23  DEFAULT_ENGINE,
24  DEFAULT_KOKORO_VOICES,
25  DEFAULT_RATE,
26  LANG_NAMES,
27  RATE_RANGE_MESSAGE,
28  SETTING_KEYS,
29  VOICE_KEYS,
30  type Settings,
31  describeVoices,
32  findVoice,
33  firstSpanishVoice,
34  isEngine,
35  parseRate,
36  parseVoiceArgs,
37  parseVoices,
38  sayArgv,
39  toSettings,
40} from './settings'
41import {
42  HEALTH_ARGV,
43  HEARTBEAT_MS,
44  LISTENER_ARGV,
45  STALE_WAV_MINUTES,
46  cacheDir,
47  daemonArgv,
48  entryName,
49  hasOtherLiveSessions,
50  isHealthyReply,
51  isKokoroServerCommand,
52  kokoroServerArgv,
53  logDir,
54  parsePids,
55  registryDir,
56  wavDir,
57} from './server'
58import {
59  BUILTIN_KOKORO_VOICES,
60  KOKORO_PORT,
61  type Engine,
62  type Segment,
63  kokoroCurlArgv,
64  kokoroRequestBody,
65  kokoroServerPath,
66  kokoroSnapshotsPath,
67  kokoroVoicesFor,
68  parseKokoroVoiceFiles,
69  planSpeech,
70} from './speech'
71
72const latest = atom({ plugin: 'speak-aloud', key: 'latest' } as const, '')
73/** Which reply block is being read: its `requestId`, LATEST for /speak, '' for none. */
74const speaking = atom({ plugin: 'speak-aloud', key: 'speaking' } as const, '')
75const LATEST = '/speak'
76
77type Child = HookStream<ProcessSpawnChunk, ProcessSpawnResult>
78/** One read: every child it has running (curl, afplay, say), and whether it was stopped. */
79type Utterance = { children: Set<Child>; isStopped: boolean }
80
81/** How often, and how many times, a starting Kokoro server is asked whether it is up: 30 s in all. */
82const SERVER_POLL_MS = 250
83const SERVER_POLLS = 120
84/** Each install step may take up to ten minutes, `$.process.run`'s most. */
85const INSTALL_STEP_MS = 600_000
86/** When the install notice was last shown, in `$.store`. */
87const NOTICE_KEY = 'installNoticeAt'
88
89// Module variables reset on a reload, and an unload kills every child, so a
90// stale `speaking` left in $.state reads as idle. The Kokoro server is no
91// child: it runs detached, shared by the sessions in the registry.
92let current: Utterance | null = null
93let hasRegistered = false
94let hasWarned = false
95let serverStarting: Promise<boolean> | null = null
96/** This session's registry entry, once it uses the server. */
97let registryEntry: string | null = null
98let heartbeat: { cancel: () => void } | null = null
99let wavDirReady: string | null = null
100/** Keeps this load's WAV names apart from other sessions' in the shared directory. */
101const loadToken = Math.random().toString(36).slice(2, 10)
102let fileCount = 0
103/** The first Spanish `say` voice; undefined until looked up, null when there is none. */
104let spanishSayVoice: string | null | undefined
105let isInstalling = false
106
107function log($: EngineInterface, text: string) {
108  $.ui.log(`speak-aloud: ${text}`, { to: 'debug' })
109}
110
111async function registerCommands($: EngineInterface) {
112  hasRegistered = true
113  try {
114    await $.command.register({ name: 'speak', description: "Read Claude's latest response aloud" })
115    await $.command.register({ name: 'speak-stop', description: 'Stop reading aloud' })
116    await $.command.register({
117      name: 'speak-voice',
118      description: 'Set the English or Spanish voice for reading aloud; alone, list the voices',
119      argumentHint: '[en|es] [name|default]',
120    })
121    await $.command.register({
122      name: 'speak-rate',
123      description: 'Set the reading speed in words per minute (80-500)',
124      argumentHint: '[wpm|default]',
125    })
126    await $.command.register({
127      name: 'speak-install',
128      description: 'Install Kokoro, the natural local voices (needs uv); --repair reinstalls',
129      argumentHint: '[--repair]',
130    })
131    await $.command.register({
132      name: 'speak-engine',
133      description: 'Read aloud with Kokoro (local neural voices) or macOS say; alone, show the engine',
134      argumentHint: '[kokoro|say|default]',
135    })
136  } catch (error) {
137    log($, `could not register commands: ${String(error)}`)
138  }
139}
140
141/**
142 * A module loaded or reloaded mid-session has missed the turns that already
143 * ended, so `latest` starts from the transcript's last assistant text.
144 */
145async function seedLatest($: EngineInterface) {
146  const known = await read($, latest)
147  if (known !== '') return
148  const messages = await $.session.messages()
149  const last = messages.findLast(m => m.role === 'assistant' && m.text.trim() !== '')
150  if (last !== undefined) await update($, latest, () => last.text)
151}
152
153async function readSettings($: EngineInterface): Promise<Settings> {
154  try {
155    const raw: Record<string, unknown> = {}
156    for (const key of SETTING_KEYS) raw[key] = await $.store.get(key)
157    return toSettings(raw)
158  } catch (error) {
159    // Speech goes on at the defaults rather than not at all.
160    log($, `could not read settings: ${String(error)}`)
161    return toSettings({})
162  }
163}
164
165async function sayListing($: EngineInterface): Promise<string | undefined> {
166  try {
167    const { exitCode, stdout } = await $.process.run(['say', '-v', '?'])
168    return exitCode === 0 && parseVoices(stdout).length > 0 ? stdout : undefined
169  } catch (error) {
170    log($, `could not list voices: ${String(error)}`)
171    return undefined
172  }
173}
174
175async function home($: EngineInterface): Promise<string | undefined> {
176  try {
177    return (await $.env.get('HOME')) || undefined
178  } catch {
179    return undefined
180  }
181}
182
183/** Kokoro's voices from the model's own voices directory, or the built-in list. */
184async function kokoroVoices($: EngineInterface): Promise<readonly string[]> {
185  const dir = await home($)
186  if (dir === undefined) return BUILTIN_KOKORO_VOICES
187  try {
188    const snapshots = kokoroSnapshotsPath(dir)
189    for (const snapshot of await $.fs.list(snapshots)) {
190      const files = await $.fs.list(`${snapshots}/${snapshot.name}/voices`)
191      const names = parseKokoroVoiceFiles(files.map(f => f.name))
192      if (names.length > 0) return names
193    }
194  } catch (error) {
195    log($, `could not read the Kokoro voices, using the built-in list: ${String(error)}`)
196  }
197  return BUILTIN_KOKORO_VOICES
198}
199
200async function isKokoroInstalled($: EngineInterface): Promise<boolean> {
201  const dir = await home($)
202  if (dir === undefined) return false
203  try {
204    return await $.fs.exists(kokoroServerPath(dir))
205  } catch {
206    return false
207  }
208}
209
210async function isServerUp($: EngineInterface): Promise<boolean> {
211  try {
212    const { exitCode, stdout } = await $.process.run(HEALTH_ARGV, { timeoutMs: 3000 })
213    return isHealthyReply(exitCode, stdout)
214  } catch {
215    return false
216  }
217}
218
219/** Makes the plugin's cache directories; resolves whether they are there. */
220async function ensureCacheDirs($: EngineInterface, dir: string): Promise<boolean> {
221  try {
222    const made = await $.process.run(['/bin/mkdir', '-p', registryDir(dir), wavDir(dir), logDir(dir)])
223    return made.exitCode === 0
224  } catch (error) {
225    log($, `could not make the cache directories: ${String(error)}`)
226    return false
227  }
228}
229
230/** Starts the Kokoro server unless one answers on the port; resolves whether it is up. */
231async function ensureServer($: EngineInterface): Promise<boolean> {
232  const isUp = await isServerUp($)
233  if (!isUp) {
234    serverStarting ??= startServer($).finally(() => {
235      serverStarting = null
236    })
237    if (!(await serverStarting)) return false
238  }
239  await joinRegistry($)
240  return true
241}
242
243async function startServer($: EngineInterface): Promise<boolean> {
244  const dir = await home($)
245  if (dir === undefined || !(await isKokoroInstalled($)) || !(await ensureCacheDirs($, dir))) return false
246  try {
247    // Detached, in the cache directory: a reload leaves it running, and no `logs/` lands in the project.
248    const started = await $.process.run(daemonArgv(kokoroServerArgv(dir)), {
249      cwd: cacheDir(dir),
250      env: { SPEAK_ALOUD_LOG: `${logDir(dir)}/server.log` },
251    })
252    if (started.exitCode !== 0) {
253      log($, `could not start the Kokoro server: ${tail(started.stderr, 5)}`)
254      return false
255    }
256  } catch (error) {
257    log($, `could not start the Kokoro server: ${String(error)}`)
258    return false
259  }
260  for (let i = 0; i < SERVER_POLLS; i++) {
261    await $.clock.sleep(SERVER_POLL_MS)
262    if (await isServerUp($)) return true
263  }
264  log($, `the Kokoro server did not answer in time; see ${logDir(dir)}/server.log`)
265  return false
266}
267
268/** Writes this session's heartbeat into the registry. */
269async function beat($: EngineInterface, dir: string, name: string) {
270  try {
271    await $.fs.write(`${registryDir(dir)}/${name}`, String(await $.clock.now()))
272  } catch (error) {
273    log($, `could not write the session registry: ${String(error)}`)
274  }
275}
276
277/** Marks this session as one using the server, once per load, and keeps the mark fresh. */
278async function joinRegistry($: EngineInterface) {
279  if (registryEntry !== null) return
280  const dir = await home($)
281  if (dir === undefined) return
282  let id = 'session'
283  try {
284    id = await $.session.id()
285  } catch {
286    // One entry under a fixed name still keeps the server for this session.
287  }
288  const name = entryName(id)
289  registryEntry = name
290  if (!(await ensureCacheDirs($, dir))) return
291  await beat($, dir, name)
292  heartbeat ??= $.clock.every(HEARTBEAT_MS, () => void beat($, dir, name))
293}
294
295/**
296 * At the session's end: drops its registry entry and, when no other live
297 * session uses the server, stops it. `force` (a repair) stops it regardless.
298 */
299async function leaveRegistry($: EngineInterface, force = false) {
300  const name = registryEntry
301  registryEntry = null
302  heartbeat?.cancel()
303  heartbeat = null
304  const dir = await home($)
305  if (dir === undefined || (name === null && !force)) return
306  try {
307    if (name !== null) await $.process.run(['/bin/rm', '-f', `${registryDir(dir)}/${name}`])
308    if (!force) {
309      const entries: { name: string; text: string }[] = []
310      for (const entry of await $.fs.list(registryDir(dir))) {
311        if (entry.kind !== 'file') continue
312        entries.push({ name: entry.name, text: await $.fs.read(`${registryDir(dir)}/${entry.name}`) })
313      }
314      if (hasOtherLiveSessions(entries, name ?? '', await $.clock.now())) return
315    }
316    await killServer($)
317  } catch (error) {
318    log($, `could not stop the Kokoro server: ${String(error)}`)
319  }
320}
321
322/** Kills whatever mlx-audio server listens on the port (only an mlx-audio server). */
323async function killServer($: EngineInterface) {
324  const { stdout } = await $.process.run(LISTENER_ARGV)
325  for (const pid of parsePids(stdout)) {
326    const ps = await $.process.run(['/bin/ps', '-p', String(pid), '-o', 'command='])
327    if (isKokoroServerCommand(ps.stdout)) await $.process.run(['/bin/kill', String(pid)])
328  }
329}
330
331/**
332 * Has the server load the model for each language, so the first read is
333 * quick; resolves whether every request went through. A first run downloads
334 * the model, so the install gives it longer.
335 */
336async function warmModel($: EngineInterface, settings: Settings, langs: readonly Lang[], timeoutMs = 60_000) {
337  let isWarm = true
338  for (const lang of langs) {
339    try {
340      const said = lang === 'es' ? 'Listo.' : 'Ready.'
341      const { exitCode, stderr } = await $.process.run(kokoroCurlArgv('/dev/null', Math.floor(timeoutMs / 1000) - 5), {
342        stdin: kokoroRequestBody(said, lang, settings.voices.kokoro[lang], settings.rate),
343        timeoutMs,
344      })
345      if (exitCode !== 0) {
346        isWarm = false
347        log($, `could not warm the Kokoro model (${lang}): ${tail(stderr, 5)}`)
348      }
349    } catch (error) {
350      isWarm = false
351      log($, `could not warm the Kokoro model (${lang}): ${String(error)}`)
352    }
353  }
354  return isWarm
355}
356
357/**
358 * At load: with Kokoro chosen and installed, start its server and load the
359 * model; chosen but missing, say how to install it (at most once a day).
360 * With say chosen it does nothing.
361 */
362async function startKokoro($: EngineInterface) {
363  const settings = await readSettings($)
364  if (settings.engine !== 'kokoro' || isInstalling) return
365  if (!(await isKokoroInstalled($))) {
366    await noticeInstall($)
367    return
368  }
369  if (await ensureServer($)) await warmModel($, settings, ['en'])
370}
371
372/** The install notice, as a toast, unless shown in the last day; it stands in for the fallback toast. */
373async function noticeInstall($: EngineInterface) {
374  try {
375    const now = await $.clock.now()
376    if (!shouldShowInstallNotice(await $.store.get(NOTICE_KEY), now)) return
377    await $.store.set(NOTICE_KEY, now)
378  } catch (error) {
379    log($, `could not check the install notice: ${String(error)}`)
380    return
381  }
382  hasWarned = true
383  log($, INSTALL_NOTICE)
384  $.ui.toast(`speak-aloud: ${INSTALL_NOTICE}`, { timeoutMs: 15_000 })
385}
386
387/** Says once per load, in the debug log and a toast, that Kokoro gave way to `say`. */
388function warnFallback($: EngineInterface, why: string) {
389  if (hasWarned) return
390  hasWarned = true
391  log($, `Kokoro is unavailable (${why}); reading with say instead`)
392  $.ui.toast(`speak-aloud: Kokoro is unavailable (${why}); reading with say.`)
393}
394
395/**
396 * The shared WAV directory, made once per load; WAVs a read long done left
397 * behind (a reload mid-read) are cleared then.
398 */
399async function ensureWavDir($: EngineInterface): Promise<string | undefined> {
400  if (wavDirReady !== null) return wavDirReady
401  const dir = await home($)
402  if (dir === undefined || !(await ensureCacheDirs($, dir))) return undefined
403  wavDirReady = wavDir(dir)
404  void $.process
405    .run(['/usr/bin/find', wavDirReady, '-name', '*.wav', '-mmin', `+${STALE_WAV_MINUTES}`, '-delete'])
406    .catch(() => {})
407  return wavDirReady
408}
409
410function removeFile($: EngineInterface, path: string) {
411  void $.process.run(['/bin/rm', '-f', path]).catch(() => {})
412}
413
414/** The `say` voice for a language: the one set, else for Spanish the first Spanish voice. */
415async function sayVoice($: EngineInterface, settings: Settings, lang: Lang): Promise<string | undefined> {
416  const set = settings.voices.say[lang]
417  if (set !== undefined || lang === 'en') return set
418  if (spanishSayVoice === undefined) {
419    const listing = await sayListing($)
420    spanishSayVoice = (listing === undefined ? undefined : firstSpanishVoice(listing)) ?? null
421  }
422  return spanishSayVoice ?? undefined
423}
424
425/** Runs one child of the read to its end; resolves its exit code, or undefined if it was stopped or failed. */
426async function runChild($: EngineInterface, mine: Utterance, request: ProcessSpawnRequest) {
427  if (mine.isStopped) return undefined
428  const child = $.process.spawn(request)
429  mine.children.add(child)
430  try {
431    for await (const _chunk of child) {
432      // Nothing these children write is worth showing: the loop is the child's life.
433    }
434    return (await child.result).code
435  } catch (error) {
436    if (!mine.isStopped) log($, `${request.argv[0]} failed: ${String(error)}`)
437    return undefined
438  } finally {
439    mine.children.delete(child)
440  }
441}
442
443/** Has the server make one segment's WAV; resolves its path, or undefined when it failed. */
444async function generate($: EngineInterface, mine: Utterance, segment: Segment, settings: Settings) {
445  const dir = await ensureWavDir($)
446  if (dir === undefined) return undefined
447  fileCount += 1
448  const out = `${dir}/${loadToken}-${fileCount}.wav`
449  const voice = settings.voices.kokoro[segment.lang]
450  const code = await runChild($, mine, {
451    argv: kokoroCurlArgv(out),
452    input: kokoroRequestBody(segment.text, segment.lang, voice, settings.rate),
453  })
454  if (code === 0) return out
455  removeFile($, out)
456  return undefined
457}
458
459/** Ends a read at once: marks it stopped and kills every child it has running. */
460function halt(utterance: Utterance | null) {
461  if (utterance === null) return
462  utterance.isStopped = true
463  // Not awaited: ending each loop kills its child, and nothing waits on its last words.
464  for (const child of utterance.children) void child.return({ code: null, signal: 'SIGTERM' }).catch(() => {})
465}
466
467/** Stops the current read, if any; resolves whether one was running. */
468async function stop($: EngineInterface): Promise<boolean> {
469  const was = current
470  current = null
471  halt(was)
472  await update($, speaking, () => '')
473  return was !== null
474}
475
476/** Which engine reads now: Kokoro when chosen and its server is up, else `say`. */
477async function engineFor($: EngineInterface, settings: Settings): Promise<Engine> {
478  if (settings.engine !== 'kokoro') return 'say'
479  if (await ensureServer($)) return 'kokoro'
480  if (await isKokoroInstalled($)) warnFallback($, 'its server did not start')
481  else await noticeInstall($)
482  return 'say'
483}
484
485/**
486 * Reads `markdown` aloud as the block `key`, stopping any current read;
487 * resolves once it has all been read or the read was stopped. Each paragraph
488 * is read in its language; with Kokoro, the next chunk is made while one plays.
489 */
490async function speak($: EngineInterface, markdown: string, key: string): Promise<void> {
491  const text = stripMarkdown(markdown)
492  // Before any await: a second read started meanwhile must find this one current, to stop it.
493  const mine: Utterance = { children: new Set(), isStopped: false }
494  halt(current)
495  current = mine
496  if (text === '') {
497    current = null
498    await update($, speaking, () => '')
499    return
500  }
501  await update($, speaking, () => key)
502  const made = new Map<number, Promise<string | undefined>>()
503  try {
504    const settings = await readSettings($)
505    let engine = await engineFor($, settings)
506    const segments = planSpeech(text, engine)
507    const prepare = (i: number) => {
508      const segment = segments[i]
509      if (segment !== undefined && !made.has(i)) made.set(i, generate($, mine, segment, settings))
510    }
511    for (const [i, segment] of segments.entries()) {
512      if (mine.isStopped) break
513      if (engine === 'kokoro') {
514        prepare(i)
515        prepare(i + 1)
516        const wav = await made.get(i)
517        if (mine.isStopped) break
518        if (wav !== undefined) {
519          await runChild($, mine, { argv: ['/usr/bin/afplay', wav] })
520          removeFile($, wav)
521          continue
522        }
523        // The rest of this read goes to `say`, from this segment on.
524        engine = 'say'
525        warnFallback($, 'its server failed')
526      }
527      const voice = await sayVoice($, settings, segment.lang)
528      await runChild($, mine, { argv: sayArgv({ voice, rate: settings.rate }), input: segment.text })
529    }
530  } catch (error) {
531    log($, `reading failed: ${String(error)}`)
532  } finally {
533    for (const wav of made.values()) {
534      void wav.then(path => path === undefined || removeFile($, path))
535    }
536    if (current === mine) {
537      current = null
538      await update($, speaking, () => '')
539    }
540  }
541}
542
543/** Which block is being read, '' when none; a value left from before a reload reads as none. */
544async function readSpeaking($: EngineInterface) {
545  const key = await read($, speaking)
546  return current === null ? '' : key
547}
548
549/** `/speak-voice [en|es] [name|default]`: answers the command's text. */
550async function voiceCommand($: EngineInterface, args: string): Promise<string> {
551  const { lang, name } = parseVoiceArgs(args)
552  const settings = await readSettings($)
553  const { engine } = settings
554  const storeKey = VOICE_KEYS[engine][lang]
555
556  if (name.toLowerCase() === 'default') {
557    await $.store.delete(storeKey)
558    const fallback =
559      engine === 'kokoro'
560        ? DEFAULT_KOKORO_VOICES[lang]
561        : lang === 'en'
562          ? 'the system default'
563          : 'the first Spanish voice'
564    return `${LANG_NAMES[lang]} voice (${engine}) reset to ${fallback}.`
565  }
566
567  let all: readonly string[]
568  let choices: Record<Lang, readonly string[]>
569  if (engine === 'kokoro') {
570    all = await kokoroVoices($)
571    choices = { en: kokoroVoicesFor('en', all), es: kokoroVoicesFor('es', all) }
572  } else {
573    const listing = await sayListing($)
574    if (listing === undefined) return "Could not list voices: `say -v '?'` gave nothing."
575    all = parseVoices(listing)
576    choices = { en: all, es: all }
577  }
578
579  if (name === '') {
580    const current: Record<Lang, string> =
581      engine === 'kokoro'
582        ? settings.voices.kokoro
583        : {
584            en: settings.voices.say.en ?? 'system default',
585            es: settings.voices.say.es ?? `${(await sayVoice($, settings, 'es')) ?? 'system default'} (default)`,
586          }
587    return describeVoices(engine, current, choices)
588  }
589
590  const found = findVoice(name, choices[lang])
591  if ('error' in found) return found.error
592  await $.store.set(storeKey, found.name)
593  return `${LANG_NAMES[lang]} voice (${engine}) set to ${found.name}.`
594}
595
596/** `/speak-rate [wpm|default]`: answers the command's text. */
597async function rateCommand($: EngineInterface, args: string): Promise<string> {
598  const wanted = args.trim()
599  if (wanted === '') {
600    const { rate } = await readSettings($)
601    return `Speaking rate: ${rate} wpm${rate === DEFAULT_RATE ? ' (default)' : ''}.`
602  }
603  if (wanted.toLowerCase() === 'default') {
604    await $.store.delete('rate')
605    return `Speaking rate reset to the default, ${DEFAULT_RATE} wpm.`
606  }
607  const rate = parseRate(wanted)
608  if (rate === undefined) return RATE_RANGE_MESSAGE
609  await $.store.set('rate', rate)
610  return `Speaking rate set to ${rate} wpm.`
611}
612
613async function kokoroStatus($: EngineInterface): Promise<string> {
614  if (await isServerUp($)) return `Kokoro: installed, server running on 127.0.0.1:${KOKORO_PORT}.`
615  if (await isKokoroInstalled($)) return 'Kokoro: installed; its server starts on the first read.'
616  return 'Kokoro: not installed; run /speak-install. Reading falls back to say.'
617}
618
619/** `/speak-engine [kokoro|say|default]`: answers the command's text. */
620async function engineCommand($: EngineInterface, args: string): Promise<string> {
621  const wanted = args.trim().toLowerCase()
622  if (wanted === '') {
623    const { engine } = await readSettings($)
624    return `Engine: ${engine}${engine === DEFAULT_ENGINE ? ' (default)' : ''}.\n${await kokoroStatus($)}`
625  }
626  const engine = wanted === 'default' ? DEFAULT_ENGINE : wanted
627  if (!isEngine(engine)) return 'Engine must be kokoro or say, or "default" (kokoro).'
628  if (wanted === 'default') await $.store.delete('engine')
629  else await $.store.set('engine', engine)
630  if (engine === 'say') return 'Engine set to say.'
631  $.clock.after(0, () => void startKokoro($))
632  return `Engine set to kokoro.\n${await kokoroStatus($)}`
633}
634
635/** Finds uv where its installer or Homebrew puts it. */
636async function findUv($: EngineInterface, dir: string): Promise<string | undefined> {
637  for (const path of uvCandidates(dir)) {
638    try {
639      if (await $.fs.exists(path)) return path
640    } catch {
641      // Not readable: try the next.
642    }
643  }
644  return undefined
645}
646
647function installStatus($: EngineInterface, step: number, what: string) {
648  $.ui.status(`speak-aloud: installing Kokoro, step ${step}/4: ${what}…`)
649}
650
651/** Runs the install's steps in order; resolves the failure's message, or undefined when all went through. */
652async function installSteps($: EngineInterface, uv: string, dir: string, isRepair: boolean) {
653  const commands: [string, string[]][] = [
654    ['mlx-audio (uv tool install)', uvInstallArgv(uv, isRepair)],
655    ['spaCy English model', spacyInstallArgv(uv, dir)],
656  ]
657  for (const [i, [what, argv]] of commands.entries()) {
658    installStatus($, i + 1, what)
659    try {
660      const { exitCode, stdout, stderr } = await $.process.run(argv, { timeoutMs: INSTALL_STEP_MS })
661      if (exitCode !== 0) {
662        log($, `install step ${i + 1} failed (${exitCode}):\n${stderr || stdout}`)
663        return `step ${i + 1}/4 (${what}) failed: ${tail(stderr || stdout, 6) || `exit code ${exitCode}`}`
664      }
665    } catch (error) {
666      return `step ${i + 1}/4 (${what}) failed: ${String(error)}`
667    }
668  }
669  installStatus($, 3, 'starting the Kokoro server')
670  if (!(await ensureServer($))) return 'step 3/4 (starting the Kokoro server) failed: it did not answer within 30 s.'
671  installStatus($, 4, 'downloading the model (~680 MB the first time) and warming English and Spanish')
672  const settings = await readSettings($)
673  if (!(await warmModel($, settings, ['en', 'es'], INSTALL_STEP_MS))) {
674    return 'step 4/4 (downloading and warming the model) failed; see the debug log.'
675  }
676  return undefined
677}
678
679async function runInstall($: EngineInterface, uv: string, dir: string, isRepair: boolean) {
680  isInstalling = true
681  try {
682    // A repair replaces the Python the running server was started from.
683    if (isRepair) await leaveRegistry($, true)
684    const failure = await installSteps($, uv, dir, isRepair)
685    if (failure === undefined) {
686      await $.store.set('engine', 'kokoro')
687      await $.store.delete(NOTICE_KEY)
688      hasWarned = false
689      $.ui.toast('speak-aloud: Kokoro is installed — reading with natural voices.', { timeoutMs: 10_000 })
690    } else {
691      log($, `Kokoro install failed at ${failure}`)
692      $.ui.toast(`speak-aloud: Kokoro install failed at ${failure}\nReading with say until it is fixed; /speak-install --repair retries.`, {
693        timeoutMs: 30_000,
694      })
695    }
696  } catch (error) {
697    log($, `Kokoro install failed: ${String(error)}`)
698    $.ui.toast(`speak-aloud: Kokoro install failed: ${String(error)}`, { timeoutMs: 30_000 })
699  } finally {
700    isInstalling = false
701    $.ui.status(undefined)
702  }
703}
704
705/** `/speak-install [--repair]`: answers at once; the install runs on, its progress in the status line. */
706async function installCommand($: EngineInterface, args: string): Promise<string> {
707  const wanted = args.trim().toLowerCase()
708  if (wanted !== '' && wanted !== '--repair' && wanted !== 'repair') return 'Usage: /speak-install [--repair]'
709  const isRepair = wanted !== ''
710  if (isInstalling) return 'Kokoro is already being installed; progress is in the status line.'
711  const dir = await home($)
712  if (dir === undefined) return 'Could not install Kokoro: HOME is not set.'
713
714  if (!isRepair && (await isKokoroInstalled($))) {
715    await $.store.set('engine', 'kokoro')
716    if (await isServerUp($)) return 'Kokoro is already installed and its server is running. /speak-install --repair reinstalls it.'
717    $.clock.after(0, () => void startKokoro($))
718    return 'Kokoro is already installed; starting its server. /speak-install --repair reinstalls it.'
719  }
720
721  const uv = await findUv($, dir)
722  if (uv === undefined) return UV_HELP
723  // On a timer, so the install outlives this command's dispatch.
724  $.clock.after(0, () => void runInstall($, uv, dir, isRepair))
725  return `${isRepair ? 'Reinstalling' : 'Installing'} Kokoro… progress in the status line.`
726}
727
728export const register: Register = on => {
729  on('session.start', async ($, e, next) => {
730    const started = await next(e)
731    await registerCommands($)
732    try {
733      await seedLatest($)
734    } catch (error) {
735      log($, `could not read the transcript: ${String(error)}`)
736    }
737    // On a timer, so the server outlives this dispatch.
738    $.clock.after(0, () => void startKokoro($))
739    return started
740  })
741
742  on('session.end', async ($, e, next) => {
743    await stop($)
744    // A /clear or a resume goes on in this process, still reading aloud.
745    if (e.reason !== 'clear' && e.reason !== 'resume') await leaveRegistry($)
746    if (e.reason === 'clear') await update($, latest, () => '')
747    return next(e)
748  })
749
750  on('turn.complete', async ($, e, next) => {
751    const done = await next(e)
752    if (e.agentId === undefined && e.answer.trim() !== '') {
753      await update($, latest, () => e.answer)
754    }
755    // A mod enabled or reloaded mid-session never sees session.start.
756    if (!hasRegistered) {
757      await registerCommands($)
758      $.clock.after(0, () => void startKokoro($))
759    }
760    return done
761  })
762
763  on('command.run', { command: 'speak' }, async $ => {
764    const text = await read($, latest)
765    if (text.trim() === '') return { text: 'No response to read yet.' }
766    // On a timer, so the children outlive this command's dispatch.
767    $.clock.after(0, () => void speak($, text, LATEST))
768    return { text: 'Reading the latest response aloud.' }
769  })
770
771  on('command.run', { command: 'speak-stop' }, async $ => {
772    const wasSpeaking = await stop($)
773    return { text: wasSpeaking ? 'Stopped reading.' : 'Nothing is being read.' }
774  })
775
776  on('command.run', { command: 'speak-voice' }, async ($, e) => ({ text: await voiceCommand($, e.args) }))
777
778  on('command.run', { command: 'speak-rate' }, async ($, e) => ({ text: await rateCommand($, e.args) }))
779
780  on('command.run', { command: 'speak-engine' }, async ($, e) => ({ text: await engineCommand($, e.args) }))
781
782  on('command.run', { command: 'speak-install' }, async ($, e) => ({ text: await installCommand($, e.args) }))
783
784  // A speaker button beside each reply block, the block itself the engine's own drawing.
785  on('ui.render', { component: 'AssistantMessage' }, async ($, e, next) => {
786    const drawn = await next(e)
787    const text = e.props.text
788    if (e.props.isSummary === true || stripMarkdown(text) === '') return drawn
789
790    const { Box, Button } = $.ui.resolve(e)
791    if (Button === undefined) return drawn
792    const key = e.requestId
793    const isOn = (await readSpeaking($)) === key
794    return (
795      <Box flexDirection="row">
796        <Box flexDirection="column" flexGrow={1} flexShrink={1}>
797          {drawn}
798        </Box>
799        {isOn ? (
800          <Button key="stop" label="■" onPress={() => stop($)} />
801        ) : (
802          <Button
803            key="speak"
804            label="🔊"
805            // On a timer, so the read and its children outlive the press's dispatch.
806            onPress={() => {
807              $.clock.after(0, () => void speak($, text, key))
808            }}
809          />
810        )}
811      </Box>
812    )
813  })
814}
815
hooks/install.ts 58 lines
1/** Installing Kokoro (mlx-audio and what it needs) with uv, as /speak-install does. */
2
3export const SPACY_MODEL_URL =
4  'https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.8.0/en_core_web_sm-3.8.0-py3-none-any.whl'
5
6const EXTRAS = [
7  'uvicorn',
8  'fastapi',
9  'python-multipart',
10  'webrtcvad-wheels',
11  'misaki[en]',
12  'num2words',
13  'spacy',
14  'phonemizer-fork',
15  'espeakng-loader',
16]
17
18export const UV_HELP =
19  'uv is needed to install Kokoro. Install it with `brew install uv` or `curl -LsSf https://astral.sh/uv/install.sh | sh`, then run /speak-install again.'
20
21export const INSTALL_NOTICE =
22  'Kokoro not installed — run /speak-install for natural voices (using `say` until then).'
23
24/** The install notice comes back at most once a day while Kokoro is missing. */
25export const NOTICE_INTERVAL_MS = 24 * 60 * 60 * 1000
26
27/** Where uv may be: its own installer's bin, Homebrew's (Apple silicon, Intel). */
28export function uvCandidates(home: string): string[] {
29  return [`${home}/.local/bin/uv`, '/opt/homebrew/bin/uv', '/usr/local/bin/uv']
30}
31
32/** `uv tool install mlx-audio` with its extras; `--force` to repair. */
33export function uvInstallArgv(uv: string, isRepair: boolean): string[] {
34  return [uv, 'tool', 'install', ...(isRepair ? ['--force'] : []), 'mlx-audio', ...EXTRAS.flatMap(x => ['--with', x])]
35}
36
37/** The spaCy English model, into the mlx-audio tool's own Python. */
38export function spacyInstallArgv(uv: string, home: string): string[] {
39  return [uv, 'pip', 'install', '--python', `${home}/.local/share/uv/tools/mlx-audio/bin/python`, SPACY_MODEL_URL]
40}
41
42/** The last `lines` non-empty lines of a command's output, at most 400 characters. */
43export function tail(text: string, lines: number): string {
44  const kept = text
45    .split('\n')
46    .map(l => l.trimEnd())
47    .filter(l => l !== '')
48    .slice(-lines)
49    .join('\n')
50  return kept.length > 400 ? kept.slice(-400) : kept
51}
52
53/** Whether to show the install notice, given when it was last shown (a stored value) and now. */
54export function shouldShowInstallNotice(lastShown: unknown, now: number): boolean {
55  if (typeof lastShown !== 'number' || !Number.isFinite(lastShown)) return true
56  return now - lastShown >= NOTICE_INTERVAL_MS
57}
58
hooks/language.ts 74 lines
1/** The languages speech tells apart. */
2export type Lang = 'en' | 'es'
3
4const SPANISH = new Set(
5  'el la los las que de y en un una para por con es está no se lo como más pero sus al del esto puedes quieres'.split(' '),
6)
7const ENGLISH = new Set('the and is are to of in that it you for with this on be not can what i'.split(' '))
8
9/** A paragraph needs this much evidence before it decides for itself. */
10const MIN_EVIDENCE = 2
11
12/**
13 * Text with what is not prose taken out: inline code, URLs, paths, and
14 * identifiers (snake_case, camelCase, dotted names, anything with a digit).
15 */
16function proseWords(text: string): string[] {
17  const prose = text
18    .replace(/`[^`]*`/g, ' ')
19    .replace(/\b(?:https?:\/\/|www\.)\S+/gi, ' ')
20  const words: string[] = []
21  for (const token of prose.split(/\s+/)) {
22    const bare = token.replace(/^[^\p{L}\d_]+|[^\p{L}\d_]+$/gu, '')
23    if (bare === '') continue
24    if (/[_/\\\d]/.test(bare) || /\p{L}\.\p{L}/u.test(bare) || /\p{Ll}\p{Lu}/u.test(bare)) continue
25    words.push(...(bare.toLowerCase().match(/[\p{L}]+/gu) ?? []))
26  }
27  return words
28}
29
30/** The evidence for each language: stopwords, plus Spanish's own letters and marks. */
31export function scoreLanguage(text: string): { en: number; es: number } {
32  let en = 0
33  let es = 0
34  const words = proseWords(text)
35  for (const word of words) {
36    if (ENGLISH.has(word)) en += 1
37    if (SPANISH.has(word)) es += 1
38    if (/ñ/.test(word)) es += 2
39    else if (/[áéíóú]/.test(word)) es += 1
40  }
41  // ¿ and ¡ count only where prose survived the filter.
42  if (words.length > 0) es += 2 * ((text.match(/[¿¡]/g) ?? []).length)
43  return { en, es }
44}
45
46function decide({ en, es }: { en: number; es: number }): Lang | undefined {
47  if (en + es < MIN_EVIDENCE || en === es) return undefined
48  return es > en ? 'es' : 'en'
49}
50
51/** The paragraph's language, or undefined when it is too short or too even to say. */
52export function classify(text: string): Lang | undefined {
53  return decide(scoreLanguage(text))
54}
55
56/** The message's language over all its text; English when unclear. */
57export function detectLanguage(text: string): Lang {
58  return classify(text) ?? 'en'
59}
60
61/** The paragraphs of plain text: blocks between blank lines, trimmed, empty ones dropped. */
62export function splitParagraphs(text: string): string[] {
63  return text
64    .split(/\n[ \t]*\n/)
65    .map(p => p.trim())
66    .filter(p => p !== '')
67}
68
69/** Each paragraph with its language; one that cannot say takes the message's. */
70export function labelParagraphs(text: string): { text: string; lang: Lang }[] {
71  const overall = detectLanguage(text)
72  return splitParagraphs(text).map(p => ({ text: p, lang: classify(p) ?? overall }))
73}
74
hooks/markdown.ts 65 lines
1/**
2 * Turns a markdown answer into plain text worth hearing: code blocks are
3 * skipped silently, URLs are dropped (link text kept), and the
4 * markers of headings, lists, quotes, emphasis and tables go.
5 */
6export function stripMarkdown(text: string): string {
7  const lines = text.replace(/\r\n?/g, '\n').split('\n')
8  const kept: string[] = []
9  let fence: string | null = null
10
11  for (const line of lines) {
12    const opener = /^\s*(`{3,}|~{3,})/.exec(line)?.[1]
13    if (fence !== null) {
14      if (opener !== undefined && opener[0] === fence[0] && opener.length >= fence.length) fence = null
15      continue
16    }
17    if (opener !== undefined) {
18      fence = opener
19      // Skipped silently; the blank line keeps the text around it apart.
20      kept.push('')
21      continue
22    }
23    const plain = stripLine(line)
24    if (plain !== null) kept.push(plain)
25  }
26
27  return kept.join('\n').replace(/\n{3,}/g, '\n\n').trim()
28}
29
30/** One line outside a code fence; null for a line that says nothing. */
31function stripLine(raw: string): string | null {
32  let line = raw
33
34  // A horizontal rule is a paragraph break; a table's separator row nothing.
35  if (/^\s*([-*_])(\s*\1){2,}\s*$/.test(line)) return ''
36  if (/^\s*\|?\s*:?-{2,}:?\s*(\|\s*:?-{2,}:?\s*)*\|?\s*$/.test(line)) return null
37
38  // Table rows: cells read as a list.
39  if (/^\s*\|.*\|\s*$/.test(line)) {
40    line = line
41      .trim()
42      .replace(/^\||\|$/g, '')
43      .split('|')
44      .map(cell => cell.trim())
45      .join(', ')
46  }
47
48  return line
49    .replace(/^\s{0,3}#{1,6}\s+/, '') // heading
50    .replace(/^(\s*>\s?)+/, '') // quote
51    .replace(/^\s*(?:[-*+]|\d+[.)])\s+(\[[ xX]\]\s+)?/, '') // list item, task box
52    .replace(/!\[([^\]]*)\]\([^)]*\)/g, '$1') // image: alt text
53    .replace(/\[([^\]]+)\]\([^)]*\)/g, '$1') // link: its text
54    .replace(/\[([^\]]+)\]\[[^\]]*\]/g, '$1') // reference link
55    .replace(/<(https?:\/\/[^>]+)>/g, 'link') // autolink
56    .replace(/<\/?[A-Za-z][^>]*>/g, '') // html tags
57    .replace(/https?:\/\/[^\s)]+?(?=[.,;:!?]?(\s|$))/g, 'link') // bare URL
58    .replace(/`+([^`]*)`+/g, '$1') // inline code
59    .replace(/(\*\*|__)(?=\S)(.+?)(?<=\S)\1/g, '$2') // bold
60    .replace(/(^|[^\w*])\*(?=\S)(.+?)(?<=\S)\*(?!\w)/g, '$1$2') // *em*
61    .replace(/(^|[^\w])_(?=\S)(.+?)(?<=\S)_(?!\w)/g, '$1$2') // _em_
62    .replace(/~~(.+?)~~/g, '$1') // strikethrough
63    .trimEnd()
64}
65
hooks/settings.ts 130 lines
1import type { Lang } from './language'
2import type { Engine } from './speech'
3
4/** The rate when the person has set none, in words per minute. */
5export const DEFAULT_RATE = 190
6export const MIN_RATE = 80
7export const MAX_RATE = 500
8
9export const DEFAULT_ENGINE: Engine = 'kokoro'
10export const DEFAULT_KOKORO_VOICES: Record<Lang, string> = { en: 'af_heart', es: 'ef_dora' }
11export const LANG_NAMES: Record<Lang, string> = { en: 'English', es: 'Spanish' }
12
13/** The `$.store` key of each engine's voice per language; `voice` predates languages. */
14export const VOICE_KEYS: Record<Engine, Record<Lang, string>> = {
15  say: { en: 'voice', es: 'sayVoiceEs' },
16  kokoro: { en: 'kokoroVoiceEn', es: 'kokoroVoiceEs' },
17}
18export const SETTING_KEYS = ['engine', 'rate', ...Object.values(VOICE_KEYS).flatMap(Object.values)] as const
19
20/**
21 * What is kept in `$.store`, across sessions. A `say` voice left unset is the
22 * system default for English, and the first Spanish voice for Spanish.
23 */
24export type Settings = {
25  engine: Engine
26  rate: number
27  voices: { say: Record<Lang, string | undefined>; kokoro: Record<Lang, string> }
28}
29
30export function isRate(value: unknown): value is number {
31  return typeof value === 'number' && Number.isInteger(value) && value >= MIN_RATE && value <= MAX_RATE
32}
33
34export function isEngine(value: unknown): value is Engine {
35  return value === 'kokoro' || value === 'say'
36}
37
38function nonEmpty(value: unknown): string | undefined {
39  return typeof value === 'string' && value !== '' ? value : undefined
40}
41
42/** The settings from the store's raw values by key, one of the wrong shape counting as unset. */
43export function toSettings(raw: Readonly<Record<string, unknown>>): Settings {
44  const { say, kokoro } = VOICE_KEYS
45  return {
46    engine: isEngine(raw.engine) ? raw.engine : DEFAULT_ENGINE,
47    rate: isRate(raw.rate) ? raw.rate : DEFAULT_RATE,
48    voices: {
49      say: { en: nonEmpty(raw[say.en]), es: nonEmpty(raw[say.es]) },
50      kokoro: {
51        en: nonEmpty(raw[kokoro.en]) ?? DEFAULT_KOKORO_VOICES.en,
52        es: nonEmpty(raw[kokoro.es]) ?? DEFAULT_KOKORO_VOICES.es,
53      },
54    },
55  }
56}
57
58/** The `say` command line, reading its text from stdin. */
59export function sayArgv({ voice, rate }: { voice: string | undefined; rate: number }): string[] {
60  return ['say', ...(voice === undefined ? [] : ['-v', voice]), '-r', String(rate), '-f', '-']
61}
62
63/** The voices in `say -v '?'` output: each line is `<name> <locale> # <sample>`. */
64export function parseVoiceEntries(listing: string): { name: string; locale: string }[] {
65  const voices: { name: string; locale: string }[] = []
66  for (const line of listing.split('\n')) {
67    const match = /^(.+?)\s+([a-z]{2,3}_[A-Za-z0-9]+)\s+#/.exec(line)
68    const [, name, locale] = match ?? []
69    if (name !== undefined && locale !== undefined && !voices.some(v => v.name === name)) voices.push({ name, locale })
70  }
71  return voices
72}
73
74/** The voice names in `say -v '?'` output. */
75export function parseVoices(listing: string): string[] {
76  return parseVoiceEntries(listing).map(v => v.name)
77}
78
79/** The first Spanish (`es_*`) voice in `say -v '?'` output, if any. */
80export function firstSpanishVoice(listing: string): string | undefined {
81  return parseVoiceEntries(listing).find(v => v.locale.startsWith('es_'))?.name
82}
83
84/** `/speak-voice`'s arguments: an optional `en`/`es` first, English when absent, then the name. */
85export function parseVoiceArgs(args: string): { lang: Lang; name: string } {
86  const match = /^(en|es)(?:\s+([\s\S]*))?$/i.exec(args.trim())
87  if (match === null) return { lang: 'en', name: args.trim() }
88  return { lang: match[1]!.toLowerCase() as Lang, name: (match[2] ?? '').trim() }
89}
90
91/** A typed rate, or undefined when it is not a whole number in range. */
92export function parseRate(text: string): number | undefined {
93  const rate = /^\d+$/.test(text) ? Number(text) : NaN
94  return isRate(rate) ? rate : undefined
95}
96
97export const RATE_RANGE_MESSAGE = `Rate must be a whole number of words per minute from ${MIN_RATE} to ${MAX_RATE}, or "default" (${DEFAULT_RATE}).`
98
99/** `/speak-voice` alone: the engine, each language's voice, and the voices to pick from. */
100export function describeVoices(
101  engine: Engine,
102  current: Record<Lang, string>,
103  choices: Record<Lang, readonly string[]> | undefined,
104): string {
105  const lines = [
106    `Engine: ${engine}`,
107    `English voice: ${current.en}`,
108    `Spanish voice: ${current.es}`,
109  ]
110  if (choices !== undefined) {
111    if (engine === 'kokoro') {
112      lines.push(`English voices: ${choices.en.join(', ')}`, `Spanish voices: ${choices.es.join(', ')}`)
113    } else {
114      lines.push(`Voices (${choices.en.length}): ${choices.en.join(', ')}`)
115    }
116  }
117  lines.push('Set one with /speak-voice [en|es] <name>, or /speak-voice [en|es] default.')
118  return lines.join('\n')
119}
120
121/** The voice `wanted` names, matched without case, or a message saying it is unknown. */
122export function findVoice(wanted: string, voices: readonly string[]): { name: string } | { error: string } {
123  const lower = wanted.toLowerCase()
124  const found = voices.find(name => name.toLowerCase() === lower)
125  if (found !== undefined) return { name: found }
126  const close = voices.filter(name => name.toLowerCase().includes(lower)).slice(0, 5)
127  const hint = close.length > 0 ? ` Did you mean: ${close.join(', ')}?` : ''
128  return { error: `Unknown voice "${wanted}".${hint} Run /speak-voice to list the voices.` }
129}
130
hooks/server.ts 105 lines
1/**
2 * The Kokoro server's life across sessions, as pure helpers.
3 *
4 * The server runs detached (nohup), so a plugin reload never kills it and
5 * every session on the machine shares the one on the port. Each session that
6 * uses it keeps an entry in a registry directory, its heartbeat the time
7 * written in it; the session that ends last stops the server.
8 */
9
10import { KOKORO_PORT, KOKORO_URL } from './speech'
11
12/** How often a session refreshes its registry entry, and when an entry counts as dead. */
13export const HEARTBEAT_MS = 60_000
14export const STALE_MS = 3 * HEARTBEAT_MS
15/** WAVs older than this, in the shared WAV directory, are left over from a read long done. */
16export const STALE_WAV_MINUTES = 10
17
18/** The plugin's own directory under the user's caches. */
19export function cacheDir(home: string): string {
20  return `${home}/Library/Caches/speak-aloud`
21}
22
23export function registryDir(home: string): string {
24  return `${cacheDir(home)}/sessions`
25}
26
27export function wavDir(home: string): string {
28  return `${cacheDir(home)}/wav`
29}
30
31export function logDir(home: string): string {
32  return `${cacheDir(home)}/logs`
33}
34
35/** A file name safe for a registry entry, from a session id. */
36export function entryName(sessionId: string): string {
37  return sessionId.replace(/[^A-Za-z0-9._-]/g, '_') || 'session'
38}
39
40/** curl asking the server for its model list: an mlx-audio server answers `{"object":"list", ...}`. */
41export const HEALTH_ARGV: readonly string[] = ['/usr/bin/curl', '-sf', '-m', '1', `${KOKORO_URL}/v1/models`]
42
43/** Whether the health check's output is an mlx-audio server's, not any other process on the port. */
44export function isHealthyReply(exitCode: number, stdout: string): boolean {
45  if (exitCode !== 0) return false
46  try {
47    const body = JSON.parse(stdout) as { object?: unknown; data?: unknown }
48    return body.object === 'list' && Array.isArray(body.data)
49  } catch {
50    return false
51  }
52}
53
54/** The server's command line: logs into the plugin's cache, not the session's directory. */
55export function kokoroServerArgv(home: string): string[] {
56  return [
57    `${home}/.local/bin/mlx_audio.server`,
58    '--host',
59    '127.0.0.1',
60    '--port',
61    String(KOKORO_PORT),
62    '--log-dir',
63    logDir(home),
64  ]
65}
66
67/**
68 * Starts `argv` detached from the session (nohup, in the background, output
69 * to `$SPEAK_ALOUD_LOG`); the shell exits at once, so a `$.process.run` of it
70 * returns, and an unload has no child left to kill.
71 */
72export function daemonArgv(argv: readonly string[]): string[] {
73  return ['/bin/sh', '-c', 'nohup "$@" >>"$SPEAK_ALOUD_LOG" 2>&1 </dev/null &', 'speak-aloud', ...argv]
74}
75
76/** lsof listing the pids listening on the server's port. */
77export const LISTENER_ARGV: readonly string[] = ['/usr/sbin/lsof', '-t', '-n', `-iTCP:${KOKORO_PORT}`, '-sTCP:LISTEN']
78
79/** The pids in lsof's `-t` output. */
80export function parsePids(stdout: string): number[] {
81  return stdout
82    .split('\n')
83    .map(l => l.trim())
84    .filter(l => /^\d+$/.test(l))
85    .map(Number)
86}
87
88/** Whether a `ps -o command=` line is the mlx-audio server's, so killing it kills nothing else. */
89export function isKokoroServerCommand(command: string): boolean {
90  return /mlx_audio[./]server/.test(command)
91}
92
93/** Whether any session other than `self` has a live registry entry. */
94export function hasOtherLiveSessions(
95  entries: readonly { name: string; text: string }[],
96  self: string,
97  now: number,
98): boolean {
99  return entries.some(({ name, text }) => {
100    if (name === self) return false
101    const beat = Number(text.trim())
102    return Number.isFinite(beat) && now - beat < STALE_MS
103  })
104}
105
hooks/speech.ts 158 lines
1import { labelParagraphs, type Lang } from './language'
2
3/** What speaks a language: the local Kokoro model, or macOS `say`. */
4export type Engine = 'kokoro' | 'say'
5
6/** One piece of an utterance, in one language. */
7export type Segment = { text: string; lang: Lang }
8
9export const KOKORO_MODEL = 'mlx-community/Kokoro-82M-bf16'
10export const KOKORO_PORT = 8765
11export const KOKORO_URL = `http://127.0.0.1:${KOKORO_PORT}`
12/** Kokoro's own pace at speed 1.0, measured: 122 words in 47.7 s. */
13export const KOKORO_NATURAL_WPM = 153
14
15/** The first Kokoro chunk is kept short, so the first audio comes soon. */
16export const FIRST_CHUNK_CHARS = 200
17export const CHUNK_CHARS = 400
18
19/** Kokoro's `speed` for a rate in words per minute, 0.5 to 2.0, two decimals. */
20export function kokoroSpeed(wpm: number): number {
21  const speed = Math.min(2, Math.max(0.5, wpm / KOKORO_NATURAL_WPM))
22  return Math.round(speed * 100) / 100
23}
24
25/** Kokoro's `lang_code` for a voice: British voices (bf_, bm_) speak `b`. */
26export function kokoroLangCode(lang: Lang, voice: string): string {
27  if (lang === 'es') return 'e'
28  return /^b[fm]_/.test(voice) ? 'b' : 'a'
29}
30
31/** The JSON body of a `/v1/audio/speech` request. */
32export function kokoroRequestBody(text: string, lang: Lang, voice: string, wpm: number): string {
33  return JSON.stringify({
34    model: KOKORO_MODEL,
35    input: text,
36    voice,
37    lang_code: kokoroLangCode(lang, voice),
38    response_format: 'wav',
39    speed: kokoroSpeed(wpm),
40  })
41}
42
43/** curl posting the body on stdin to the server, the WAV written to `out`, giving up after `maxSeconds`. */
44export function kokoroCurlArgv(out: string, maxSeconds = 120): string[] {
45  return [
46    '/usr/bin/curl',
47    '-sS',
48    '-f',
49    '-m',
50    String(maxSeconds),
51    '-X',
52    'POST',
53    '-H',
54    'Content-Type: application/json',
55    '--data-binary',
56    '@-',
57    '-o',
58    out,
59    `${KOKORO_URL}/v1/audio/speech`,
60  ]
61}
62
63/** Where uv puts the mlx-audio tools, under the home directory. */
64export function kokoroServerPath(home: string): string {
65  return `${home}/.local/bin/mlx_audio.server`
66}
67
68/** The Kokoro voices directory's parent, in the Hugging Face cache. */
69export function kokoroSnapshotsPath(home: string): string {
70  return `${home}/.cache/huggingface/hub/models--mlx-community--Kokoro-82M-bf16/snapshots`
71}
72
73/** Kokoro's voices as the model ships them, for when the cache cannot be read. */
74export const BUILTIN_KOKORO_VOICES: readonly string[] = [
75  'af_alloy', 'af_aoede', 'af_bella', 'af_heart', 'af_jessica', 'af_kore', 'af_nicole', 'af_nova', 'af_river',
76  'af_sarah', 'af_sky', 'am_adam', 'am_echo', 'am_eric', 'am_fenrir', 'am_liam', 'am_michael', 'am_onyx', 'am_puck',
77  'am_santa', 'bf_alice', 'bf_emma', 'bf_isabella', 'bf_lily', 'bm_daniel', 'bm_fable', 'bm_george', 'bm_lewis',
78  'ef_dora', 'em_alex', 'em_santa',
79]
80
81const KOKORO_PREFIXES: Record<Lang, RegExp> = { en: /^(af|am|bf|bm)_/, es: /^(ef|em)_/ }
82
83/** The voice names among a voices directory's file names, sorted. */
84export function parseKokoroVoiceFiles(files: readonly string[]): string[] {
85  return files
86    .map(f => /^([a-z]{2}_[a-z0-9]+)\.(safetensors|pt)$/.exec(f)?.[1])
87    .filter((n): n is string => n !== undefined)
88    .filter((n, i, all) => all.indexOf(n) === i)
89    .sort()
90}
91
92/** The Kokoro voices that speak `lang`. */
93export function kokoroVoicesFor(lang: Lang, voices: readonly string[]): string[] {
94  return voices.filter(v => KOKORO_PREFIXES[lang].test(v))
95}
96
97/** Splits a paragraph into sentence groups of at most `max` characters (a longer sentence on word breaks). */
98export function splitLong(text: string, max: number): string[] {
99  if (text.length <= max) return [text]
100  const sentences = text.split(/(?<=[.!?…:;])\s+/)
101  const pieces: string[] = []
102  for (const sentence of sentences) {
103    if (sentence.length <= max) {
104      pieces.push(sentence)
105      continue
106    }
107    let line = ''
108    for (const word of sentence.split(/\s+/)) {
109      if (line !== '' && line.length + 1 + word.length > max) {
110        pieces.push(line)
111        line = ''
112      }
113      line = line === '' ? word : `${line} ${word}`
114    }
115    if (line !== '') pieces.push(line)
116  }
117  const groups: string[] = []
118  for (const piece of pieces) {
119    const last = groups.at(-1)
120    if (last !== undefined && last.length + 1 + piece.length <= max) groups[groups.length - 1] = `${last} ${piece}`
121    else groups.push(piece)
122  }
123  return groups
124}
125
126/** The opening paragraph's pieces: a short first one, then the usual size. */
127function splitFirst(text: string): string[] {
128  const [head, ...rest] = splitLong(text, FIRST_CHUNK_CHARS)
129  if (head === undefined) return []
130  return rest.length === 0 ? [head] : [head, ...splitLong(rest.join(' '), CHUNK_CHARS)]
131}
132
133/**
134 * The utterance as segments: each paragraph in its language. `say` takes a
135 * run of same-language paragraphs whole; Kokoro takes
136 * chunks of a few sentences, the first one short, so audio starts soon and the
137 * next chunk is made while one plays.
138 */
139export function planSpeech(text: string, engine: Engine): Segment[] {
140  const segments: Segment[] = []
141  for (const { text: paragraph, lang } of labelParagraphs(text)) {
142    const last = segments.at(-1)
143    if (engine === 'say') {
144      if (last !== undefined && last.lang === lang) last.text = `${last.text}\n\n${paragraph}`
145      else segments.push({ text: paragraph, lang })
146      continue
147    }
148    for (const piece of segments.length === 0 ? splitFirst(paragraph) : splitLong(paragraph, CHUNK_CHARS)) {
149      const now = segments.at(-1)
150      const max = segments.length <= 1 ? FIRST_CHUNK_CHARS : CHUNK_CHARS
151      if (now !== undefined && now.lang === lang && now.text.length + 2 + piece.length <= max) {
152        now.text = `${now.text}\n\n${piece}`
153      } else segments.push({ text: piece, lang })
154    }
155  }
156  return segments
157}
158
types/index.d.ts 9 lines
1/** What /speak reads: the main loop's latest non-empty answer, as markdown. */
2export type SpeakAloudText = string
3
4declare module 'claude-code' {
5  interface PluginState {
6    'speak-aloud': { latest: SpeakAloudText; speaking: string }
7  }
8}
9