SLOPSHOPPER

claudio-tts

Spoken replies for Claude Code. Per-session mute (new sessions start muted), shared volume, speed and output devices.

newcommandstatusprocess
v0.5.0MITupdated 2026-10-07restante/claudio-tts/src/claudio_tts/mod
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · claudio-tts
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /tts ⎿ claudio-tts: TTS on (this session), volume 10/10, speed 1, voice default, language auto, output default ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⚠ claudio-tts: TTS on
README

<sub>🌍 English (US) · Italiano · Polski · Français · Deutsch · 日本語 · 简体中文 · हिन्दी · Русский · Español · Português</sub>

🔊 claudio-tts

Give Claude Code a voice. Locally. For free. With zero tokens.

Natural spoken replies for Claude Code, powered by the open-source Kokoro voice model and built on Claude Code's new mod system. No API key, no account, no per-word cost, and your text never leaves your machine.

CI Release License: MIT macOS Windows Voices Languages No API key Tokens PRs welcome

<sub>Claude Code text-to-speech · Claude talks back · voice output for Claude Code · offline TTS · Kokoro · Claude Code plugin / mod · works with 54 voices in 9 languages</sub>

<img src="docs/demo.gif" alt="Terminal demo: one-command install, 54 voices, spoken test phrases, and the bug-report helper" width="760">

<sub>A real recording, nothing faked: install, configure, then a live Claude Code session where Claude talks back. A GIF can't carry sound, so <a href="https://restante.github.io/claudio-tts/">listen to the voices in the player</a>. Re-record it any time with <code>scripts/make-hero.sh</code>.</sub>

Install · Voices · Commands · Why Kokoro · Mods · Contribute · Report a bug


✨ Highlights

  • 🎧 Hear Claude while you do something else. Read a diff, make coffee, rest your eyes. Claude's replies, and the narration before each tool call, are spoken as they arrive.
  • 🆓 Free forever, no tokens, no APIs. Speech is generated on your own computer. There is nothing to sign up for and nothing to pay for.
  • 🔒 Private and offline. After the one-time model download, it works with no internet. Your code and your conversations are never sent anywhere to be voiced.
  • 🎯 Always the latest reply. It listens to Claude Code's own turn events instead of scraping the transcript file, so it can never read the message before the one you just got.
  • 🧑‍🤝‍🧑 Made for many sessions. Mute is per session, new sessions start muted, sessions take turns instead of talking over each other, and each one only ever stops its own voice.
  • 🗣️ 54 voices, 9 languages. Change the voice with one command: /tts voice af_heart.
  • 🔈 Your speakers, your rules. Volume, speed and output device (one, several, or all at once).
  • 🩺 Easy to support. claudio-tts doctor --report writes a ready-to-paste bug report.

🧩 Built on Claude Code mods

claudio-tts is built on the new mod system in Claude Code, not on the older "run a shell script on every event" hooks. A mod is a small plugin of typed functions that runs inside Claude Code, sees what the model is doing as it happens, can add commands and status-line entries, and hot-reloads while you work. That is exactly what a good voice needs:

Mod featureWhat claudio-tts does with it
Turn events (turn.complete, turn.step, turn.start, session.end)Receives the model's final text and the narration before tool calls directly in the event, so it is never one reply behind and never has to re-read a transcript file
Per-session stateEach session remembers its own mute switch, so ten open sessions don't become ten voices
Persistent mod storeVolume, speed, voice and output device survive restarts
Slash-command registrationAdds /tts with everything below, right inside Claude Code
Status lineShows TTS on or TTS muted for the session you are looking at
Process APIHands the text to the local speech engine without ever blocking Claude
Typed contract and toolingShips a type contract, and is checked with claude plugin validate and claude plugin test
Hot reloadEdit the mod and it reloads, no restart needed while developing

The mod itself is a thin TypeScript layer (src/claudio_tts/mod). The audio work lives in a small Python package so macOS and Windows share one code path.


🚀 Install

You need Claude Code. The installer handles everything else, including Python, the dependencies and the voice model.

macOS

curl -fsSL https://raw.githubusercontent.com/restante/claudio-tts/main/install.sh | bash

Windows (PowerShell, beta)

irm https://raw.githubusercontent.com/restante/claudio-tts/main/install.ps1 | iex

Then restart Claude Code and, in a session:

/tts unmute

That's it. Send a message and listen. 🎉 (New sessions start muted on purpose; see many sessions.)

  1. Installs uv if you don't have it. uv also downloads a private Python 3.12, so you don't need Python installed.
  2. Creates an isolated environment (macOS: ~/Library/Application Support/claudio-tts, Windows: %LOCALAPPDATA%\claudio-tts) and installs this package and its dependencies into it.
  3. Downloads the Kokoro model (326 MB, or 92 MB with --lite) and verifies its SHA-256 checksum.
  4. Copies the Claude Code mod into ~/.claude/mods/claudio-tts and registers it in ~/.claude/settings.json. Your settings are backed up to settings.json.claudio-tts.bak and merged, never overwritten.
  5. Runs doctor to check the model, audio devices and Claude Code.

It is safe to run again: re-running updates, and nothing is duplicated.

macOS / Linux (bash -s -- …)Windows (-…)Effect
--lite-LiteSmaller 92 MB model (a little less natural, faster to download and run)
--no-model-NoModelSkip the model download
--ref <ref>-Ref <ref>Install a branch, tag or commit
--uninstall-UninstallRemove everything (add --keep-models / -KeepModels to keep the voice files)
--local-LocalInstall from the checkout you are standing in

With PowerShell's irm | iex you cannot pass switches; set CLAUDIO_TTS_LITE=1, CLAUDIO_TTS_NO_MODEL=1, CLAUDIO_TTS_REF=<ref> or CLAUDIO_TTS_UNINSTALL=1 first, or use & ([scriptblock]::Create((irm <url>))) -Lite.


🎙️ Talk to Claude, and hear Claude

Voice works in both directions in Claude Code, with two independent pieces:

DirectionFeaturePowered by
🗣️ You → ClaudeClaude Code's built-in voice input: run /voice to toggle it (hold the key to talk)Claude Code itself
🔊 Claude → youSpoken replies, narration before tool calls, per-session controlclaudio-tts

Turn on both and you can have a hands-free conversation with Claude Code. To switch voice input on by default, put this in ~/.claude/settings.json (the install already added the claudio-tts part):

{
  "voice": { "enabled": true, "mode": "hold" },
  "env": { "KOKORO_VOICE": "af_heart" }
}

Then, in a Claude Code session:

/tts unmute          # new sessions start muted on purpose
/tts voice af_heart  # pick the voice (it says hello)
> explain this stack trace      <- type it, or hold the voice key and say it

Claude's reply is spoken as it arrives. /tts mute silences just this session.


🎮 Commands

Everything is one slash command inside Claude Code:

CommandWhat it does
/ttsToggle speech for this session
/tts mute · /tts unmuteTurn it off / on for this session only
/tts statusShow mute state, volume, speed, voice and output
/tts default on · /tts default offWhether new sessions start speaking (off = start muted, the default)
/tts disable · /tts enableTurn claudio-tts fully off (no speech, no status line, no update notice) or back on, for this session only
/tts voiceList every voice and show the current one
/tts voice af_heartChange the voice (it says hello in the new voice). /tts voice default resets it
/tts lang de · /tts lang autoMake the voice read another language (any of 140), or go back to its own
/tts updateCheck for a newer release and install it (nothing installs until you type this). /tts update check only looks, /tts update off stops the daily check
/tts volume 1-10Loudness, shared by all sessions. /tts volume shows it
/tts speed 0.5-1.5Speaking pace (1 is normal). /tts pace is an alias
/tts deviceList output devices and show the current choice
/tts device airpodsSpeak on one device (partial names work)
/tts device airpods,macbook…on several devices at the same time
/tts device all…on every real output (virtual devices such as Zoom and Teams are skipped)
/tts device defaultBack to the system default
/tts micList microphones; /tts mic <name> saves a preference

This is what it looks like in a session:

> /tts unmute
claudio-tts: TTS on (this session), volume 10/10, speed 1, voice default, output default

> /tts voice bf_emma
claudio-tts: Voice: bf_emma

> /tts speed 0.9
claudio-tts: TTS speed 0.9

Volume, speed, voice and device describe your setup, so they are shared. Mute describes a conversation, so it is per session.


🗣️ Voices

Kokoro ships 54 voices in 9 languages. Pick one with /tts voice <name>, or list them all with /tts voice (or claudio-tts voices in a terminal). The first letter of a voice name is its language and the second is its gender, and claudio-tts picks the right language automatically from the name.

These are examples, not limits. Any voice Kokoro supports can be used, you can add your own voice files, and a voice can read text in 140 languages (see Use any other voice or language).

LanguageFemaleMale
🇺🇸American Englishaf_alloy af_aoede af_bella af_heart af_jessica af_kore af_nicole af_nova af_river af_sarah af_skyam_adam am_echo am_eric am_fenrir am_liam am_michael am_onyx am_puck am_santa
🇬🇧British Englishbf_alice bf_emma bf_isabella bf_lilybm_daniel bm_fable bm_george bm_lewis
🇪🇸Spanishef_doraem_alex em_santa
🇫🇷Frenchff_siwis
🇮🇳Hindihf_alpha hf_betahm_omega hm_psi
🇮🇹Italianif_saraim_nicola
🇯🇵Japanesejf_alpha jf_gongitsune jf_nezumi jf_tebukurojm_kumo
🇧🇷Brazilian Portuguesepf_dorapm_alex pm_santa
🇨🇳Mandarin Chinesezf_xiaobei zf_xiaoni zf_xiaoxiao zf_xiaoyizm_yunjian zm_yunxi zm_yunxia zm_yunyang

German, Polish and Russian: Kokoro has no native German, Polish or Russian voices yet. The installer, the commands and this documentation are fully available in all three (see the links at the top), but the spoken voice will be an English or other-language one reading your text with a noticeable accent. Try claudio-tts say "…" --voice af_heart --lang de (or pl, ru) to hear it. If Kokoro adds those voices, claudio-tts will pick them up through the same /tts voice command.

🎧 Hear the voices

▶ Open the voice player to listen to all 54 voices right in your browser, with one-click play buttons. (GitHub can't play audio inside a README, so the player lives on a small web page.) Or click a name below to jump straight to it. Samples are generated by Kokoro itself.

VoiceListenVoiceListen
af_heart ⭐▶ listenbf_emma▶ listen
af_bella ⭐▶ listenbf_isabella▶ listen
af_nicole▶ listenbm_george▶ listen
af_sarah▶ listenbm_fable▶ listen
af_sky▶ listenef_dora 🇪🇸▶ listen
am_michael▶ listenff_siwis 🇫🇷▶ listen
am_fenrir▶ listenif_sara 🇮🇹▶ listen
am_puck▶ listenjf_alpha 🇯🇵▶ listen
hf_alpha 🇮🇳▶ listenzf_xiaoxiao 🇨🇳▶ listen
pf_dora 🇧🇷▶ listen

⭐ af_heart and af_bella are generally regarded as the most natural English voices; start there.

Tips

  • The voice and the text should match. A Spanish voice reading English will sound odd, because the voice decides how the text is pronounced. If you chat with Claude in Spanish, pick ef_dora or em_alex.
  • Set a default for every session without a command: put "KOKORO_VOICE": "bf_emma" in the env block of ~/.claude/settings.json. /tts voice overrides it, and /tts voice default goes back to it.
  • Too fast, too slow? /tts speed 0.85 slows it down, /tts speed 1.2 speeds it up.
  • Want it quieter? /tts volume 4. Volume is applied per sample, so it doesn't touch your system volume.
  • The first sentence of a session can take a moment while the model loads; later ones are quick. The --lite model starts faster and uses less memory.

🔧 Use any other voice or language

The 54 built-in voices and 9 native languages are just what ships in the box:

/tts voice af_mix                # a voice you added (a .npy file in the voices folder)
/tts lang de                     # read German (or any of 140 languages) through the current voice
/tts lang auto                   # back to the voice's own language
{ "env": { "CLAUDIO_TTS_MODEL": "/path/to/model.onnx", "CLAUDIO_TTS_VOICES": "/path/to/voices.bin" } }
  • Add your own voice: save a Kokoro style vector as voices/<name>.npy in the install folder, and <name> appears in /tts voice. You can even blend two voices into a new one.
  • Use another Kokoro model or voices pack (a newer release, a community pack): set the two environment variables above in ~/.claude/settings.json.
  • Read any language: /tts lang <code> makes the current voice read that language (140 codes, see claudio-tts languages). Languages without a native voice are read with an accent.

Step-by-step, with a blending script: docs/voices.md.


🆓 Why Kokoro? No tokens, no APIs, no bill

Most "make it talk" setups send every reply to a cloud text-to-speech service. claudio-tts doesn't. It runs Kokoro, a small open-weight voice model, on your own CPU.

☁️ Cloud text-to-speech🔊 claudio-tts + Kokoro
API key / accountRequiredNone
CostPer character or per minute, foreverFree
Claude tokens usedOften extra, if a model writes the scriptZero extra (see note)
PrivacyYour replies are sent to a third partyNothing leaves your computer
OfflineNoYes (after the one-time download)
LatencyNetwork round trip plus queueingStarts speaking as soon as the first sentence is ready
Rate limits / outagesYesNone
LicenceTerms of serviceApache-2.0 model, MIT code
Sizen/a326 MB (92 MB --lite)

Honest notes. The one-time model download is 326 MB. Generating speech uses some CPU while it talks. The very best paid cloud voices can sound richer than Kokoro, but Kokoro is remarkably natural for a model this small. And the optional spoken summary asks Claude to write one extra sentence or two per reply, which costs a handful of output tokens, only if you turn it on. Without it, claudio-tts reads text Claude already wrote and uses no extra tokens at all.


🧑‍🤝‍🧑 Many sessions, one pair of ears

flowchart LR
  A["Session A<br/>unmuted"] -->|reply| Q{{"one speaker<br/>at a time"}}
  B["Session B<br/>muted"] -. silent .-> Q
  C["Session C<br/>unmuted"] -->|reply| Q
  Q --> D1["Headphones"]
  Q --> D2["Speakers"]
  • New sessions start muted. Unmute only the one you are watching (/tts unmute), or run /tts default on if you would rather they all start speaking.
  • If two unmuted sessions answer at once, the second waits for its turn instead of cutting the first off.
  • Sending a prompt, or muting, stops only that session's voice.

✂️ Make it speak less: the optional spoken summary

Long answers are tiring to listen to. If a reply contains a summary block, claudio-tts reads only that block and skips the rest. Add something like this to your CLAUDE.md and Claude will write one every time:

At the END of EVERY response add a short, conversational summary for text-to-speech, in this exact form.
Avoid URLs, file paths, code and variable names.

<!-- TTS_SUMMARY Two or three plain sentences about the result. TTS_SUMMARY -->

No block? It reads the whole reply, with markdown, code blocks, links and emoji cleaned up so it sounds natural.


🛠️ How it works

flowchart LR
  E["Claude Code<br/>turn events"] --> M["claudio-tts mod<br/>decides what and when"]
  M -->|"claudio-tts speak"| P["Python worker<br/>detached"]
  P --> K["Kokoro<br/>local neural voice"]
  K --> O["Your output<br/>devices"]

The mod listens for the model's text and tool calls, tracks per-session mute, and hands text to the claudio-tts command. All the OS-specific work (audio, process control, locking, ducking your music while it talks) lives in the Python package, so macOS and Windows share one code path. Details in docs/how-it-works.md.

⚙️ Configuration

Environment variables (put them in the env block of ~/.claude/settings.json):

VariableDefaultMeaning
KOKORO_VOICEaf_skyDefault voice for every session (see Voices)
CLAUDIO_TTS_MODEL, CLAUDIO_TTS_VOICESbuilt-inUse another Kokoro model / voices pack (both required)
AUDIO_DUCK_ENABLEDtrueLower Apple Music / Spotify while speaking (macOS only)
DUCK_LEVEL5Percent of the original music volume to duck to
CLAUDIO_TTS_HOMEper OSWhere the install lives

💻 Command line

The package also installs a claudio-tts command (inside its private environment):

claudio-tts say "Hello there" --voice bf_emma   speak now and wait
claudio-tts voices                              list all voices (the 54 built-in plus yours)
claudio-tts languages                           list the 140 languages a voice can read
claudio-tts devices [--inputs]                  list audio devices
claudio-tts doctor [--speak | --report]         check the install, or write a bug report
claudio-tts download-model [--lite]             fetch and verify the voice files
claudio-tts install-mod / uninstall-mod

🖥️ Platform support

Status
macOS (Apple silicon and Intel)✅ Supported and tested, including music ducking
Windows 10/11🧪 Beta. Tested in CI; real-audio feedback welcome. No music ducking yet
Linux🤷 Best effort, untested on real hardware

🐞 Report a bug

Found something odd? That is genuinely useful, thank you.

  1. Run this and copy the output:
   claudio-tts doctor --report

(If claudio-tts isn't on your PATH, use the full path the installer printed, ending in python -m claudio_tts doctor --report.) It contains your OS, versions and device names, but never anything you have spoken or any secrets.

  1. Open a bug report and paste it in. Say what you expected and what happened.

Quick fixes live in docs/troubleshooting.md and docs/windows.md. The most common one: silence usually means the session is still muted, so type /tts unmute.

🤝 Contribute

Contributors are very welcome. This started as a one-person weekend project, and it gets better with more people: more voices tested, more platforms covered, more ideas.

Great places to jump in:

  • 🪟 Windows: try it on real hardware, report what you hear, or build music ducking (per-app volume).
  • 🐧 Linux: polish and test the installer on your distro.
  • 🎛️ Voice blending and previews: mix two voices, or audition a voice before choosing.
  • 📦 Packaging: pipx, Homebrew, winget.
  • 🌍 Languages: better reading of code, numbers and mixed-language text.
  • 📚 Docs and demos: better GIFs, translations, tutorials.

Look for good first issue and help wanted, and read CONTRIBUTING.md to get set up in five minutes.

Want to talk first? Start a discussion or reach me through my GitHub profile, @restante. Ideas, questions, "I tried it on X and…" stories, and offers to help are all welcome. And if claudio-tts made your day, a ⭐ helps others find it.

📱 Control Claude from your phone

claudio-vibecode is a companion project built on claudio-tts: a local web page for your phone (iPhone or Android) where you read the conversation, send a message, stop Claude, approve tool calls, hold a button to talk, and hear Claude's voice on the phone while the Mac stays quiet. It installs claudio-tts for you and appears as an extra output in /tts device.

curl -fsSL https://raw.githubusercontent.com/restante/claudio-vibecode/main/install.sh | bash

🗺️ Roadmap

  • ☐ Music ducking on Windows
  • ☐ Vo
Source 3 files
hooks/register.ts 297 lines
1import type { Hook, Register } from 'claude-code'
2
3import {
4  DEFAULT_SPEED,
5  DEFAULT_VOLUME,
6  MAX_SPEED,
7  MIN_SPEED,
8  hasSummary,
9  parseTtsArgs,
10  speechText,
11} from './speech'
12
13const MIN_NARRATION = 10
14const USAGE =
15  'Usage: /tts [mute|unmute|status|default on|off|volume 1-10|speed 0.5-1.5|voice <name>|lang <code|auto>|device <names|all|default>|mic <name|default>|update [check|on|off]|disable|enable]'
16
17type Dollar = Parameters<Hook<'turn.start'>>[0]
18
19// Every OS-specific job (audio, process control, music ducking) lives in the `claudio-tts` Python
20// package; this mod only decides what to say and when. CLAUDIO_TTS_PYTHON is set by `install-mod`.
21async function python($: Dollar) {
22  try {
23    return (await $.env.get('CLAUDIO_TTS_PYTHON')) ?? 'python'
24  } catch {
25    return 'python'
26  }
27}
28
29async function cli($: Dollar, args: string[], stdin?: string, timeoutMs = 15000) {
30  const exe = await python($)
31  try {
32    return await $.process.run([exe, '-m', 'claudio_tts', ...args], { stdin, timeoutMs })
33  } catch (error) {
34    $.ui.log(`claudio-tts: ${String(error)}`, { to: 'debug' })
35    return undefined
36  }
37}
38
39async function sessionId($: Dollar) {
40  try {
41    return await $.session.id()
42  } catch {
43    return 'default' // only when the engine cannot name the session (tests); speech still works
44  }
45}
46
47async function stop($: Dollar) {
48  await cli($, ['stop', '--session', await sessionId($)])
49}
50
51const DISABLED = { plugin: 'claudio-tts', key: 'disabled' } as const
52
53// `/tts disable` switches the mod off for this session only: no speech, no status line, no update notice.
54async function isDisabled($: Dollar) {
55  const { value } = await $.state.get(DISABLED)
56  return value === true
57}
58
59async function speak($: Dollar, text: string) {
60  if (await isDisabled($)) return
61  if (await isMuted($)) return
62  const volume = Number((await $.store.get('volume')) ?? DEFAULT_VOLUME)
63  const speed = Number((await $.store.get('speed')) ?? DEFAULT_SPEED)
64  const devices = String((await $.store.get('devices')) ?? 'default')
65  const voice = String((await $.store.get('voice')) ?? '')
66  const lang = String((await $.store.get('lang')) ?? '')
67  await cli(
68    $,
69    [
70      'speak',
71      '--session', await sessionId($),
72      '--volume', String(volume),
73      '--speed', String(speed),
74      '--devices', devices,
75      ...(voice ? ['--voice', voice] : []),
76      ...(lang ? ['--lang', lang] : []),
77    ],
78    text,
79  )
80}
81
82const MUTED = { plugin: 'claudio-tts', key: 'muted' } as const
83
84// This session's switch if it was ever set here, else the startup default (muted unless `/tts default on`).
85async function isMuted($: Dollar) {
86  const { value } = await $.state.get(MUTED)
87  if (value !== undefined) return value
88  return ((await $.store.get('defaultMuted')) ?? true) === true
89}
90
91// Once a day (the Python side caches), say if a newer release exists. Never installs anything.
92async function noticeUpdate($: Dollar) {
93  if (((await $.store.get('updateCheck')) ?? true) !== true) return
94  if (await isDisabled($)) return
95  const found = await cli($, ['update', '--quiet'], undefined, 8000)
96  const line = found?.stdout.trim()
97  if (line) $.ui.status(`${(await isMuted($)) ? 'TTS muted' : 'TTS on'} | update available: /tts update`)
98}
99
100async function showStatus($: Dollar) {
101  if (await isDisabled($)) {
102    $.ui.status(undefined)
103    return
104  }
105  const muted = await isMuted($)
106  $.ui.status(muted ? 'TTS muted' : 'TTS on')
107  // A small file other add-ons (claudio-vibecode) can read to show this session's sound state.
108  await cli($, ['session-state', '--session', await sessionId($), '--muted', muted ? 'on' : 'off'])
109}
110
111export const register: Register = on => {
112  on('session.start', async ($, e, next) => {
113    await $.command.register({
114      name: 'tts',
115      description: 'Spoken replies (per session): /tts [mute|unmute|status|default|volume|speed|voice|lang|device|mic|update]',
116    })
117    await showStatus($)
118    await noticeUpdate($)
119    return next(e)
120  })
121
122  on('command.run', { command: 'tts' }, async ($, e) => {
123    const cmd = parseTtsArgs(e.args)
124    if (!cmd) return { text: USAGE }
125    const volume = Number((await $.store.get('volume')) ?? DEFAULT_VOLUME)
126
127    if (cmd.kind === 'disable' || cmd.kind === 'enable') {
128      await $.state.set(DISABLED, cmd.kind === 'disable')
129      if (cmd.kind === 'disable') await stop($)
130      await showStatus($)
131      return { text: cmd.kind === 'disable' ? 'TTS disabled for this session (/tts enable to turn it back on)' : 'TTS enabled for this session' }
132    }
133
134    if (await isDisabled($)) return { text: 'TTS is disabled. Use /tts enable to turn it back on' }
135
136    if (cmd.kind === 'device') {
137      const current = String((await $.store.get('devices')) ?? 'default')
138      if (cmd.value === undefined) {
139        const listed = await cli($, ['devices'])
140        return {
141          text: `TTS output: ${current}\nAvailable:\n${listed?.stdout.trim() ?? '(could not list devices)'}\nUse /tts device <name,name|all|default>`,
142        }
143      }
144      await $.store.set('devices', cmd.value)
145      await speak($, 'Speech output changed.')
146      return { text: `TTS output: ${cmd.value}` }
147    }
148
149    if (cmd.kind === 'mic') {
150      const current = String((await $.store.get('mic')) ?? 'default')
151      if (cmd.value === undefined) {
152        const inputs = await cli($, ['devices', '--inputs'])
153        return {
154          text: `Microphone: ${current}\nAvailable:\n${inputs?.stdout.trim() ?? '(could not list devices)'}\nUse /tts mic <name|default>`,
155        }
156      }
157      await $.store.set('mic', cmd.value)
158      return { text: `Microphone: ${cmd.value}` }
159    }
160
161    if (cmd.kind === 'update') {
162      if (cmd.value === 'on' || cmd.value === 'off') {
163        await $.store.set('updateCheck', cmd.value === 'on')
164        return { text: `Update check at session start: ${cmd.value}` }
165      }
166      // Typing /tts update is the approval; `check` only looks.
167      const args = cmd.value === 'check' ? ['update', '--check'] : ['update', '--yes']
168      const done = await cli($, args, undefined, 300000)
169      return { text: done?.stdout.trim() || done?.stderr.trim() || 'Could not run the update.' }
170    }
171
172    if (cmd.kind === 'default') {
173      const startMuted = ((await $.store.get('defaultMuted')) ?? true) === true
174      if (cmd.value === undefined) {
175        return { text: `New sessions start ${startMuted ? 'muted' : 'speaking'} (/tts default on|off to change)` }
176      }
177      await $.store.set('defaultMuted', cmd.value === 'off')
178      return { text: `New sessions will start ${cmd.value === 'off' ? 'muted' : 'speaking'}` }
179    }
180
181    if (cmd.kind === 'lang') {
182      const current = String((await $.store.get('lang')) ?? '')
183      if (cmd.value === undefined) {
184        const listed = await cli($, ['languages'])
185        return {
186          text: `Language: ${current || 'auto (the voice\'s own)'}\n${listed?.stdout.trim() ?? '(could not list languages)'}`,
187        }
188      }
189      if (cmd.value === 'auto' || cmd.value === 'default') {
190        await $.store.delete('lang')
191        return { text: 'Language: auto (the voice\'s own)' }
192      }
193      const check = await cli($, ['languages', '--check', cmd.value])
194      if (check && check.exitCode !== 0) {
195        return { text: check.stdout.trim() || `Unknown language '${cmd.value}'. Try /tts lang to list them` }
196      }
197      await $.store.set('lang', cmd.value.toLowerCase())
198      return { text: `Language: ${cmd.value.toLowerCase()} (the current voice will read text as ${cmd.value.toLowerCase()}, with an accent if it has no native voice)` }
199    }
200
201    if (cmd.kind === 'voice') {
202      const current = String((await $.store.get('voice')) ?? '')
203      if (cmd.value === undefined) {
204        const listed = await cli($, ['voices'])
205        return {
206          text: `Voice: ${current || 'default (KOKORO_VOICE or af_sky)'}\n${listed?.stdout.trim() ?? '(could not list voices)'}`,
207        }
208      }
209      if (cmd.value === 'default') {
210        await $.store.delete('voice')
211        return { text: 'Voice reset to the default' }
212      }
213      const check = await cli($, ['voices', '--check', cmd.value])
214      if (check && check.exitCode !== 0) {
215        return { text: check.stdout.trim() || `Unknown voice '${cmd.value}'. Try /tts voice to list them` }
216      }
217      await $.store.set('voice', cmd.value)
218      await speak($, `Hi, this is ${cmd.value.slice(3)}.`)
219      return { text: `Voice: ${cmd.value}` }
220    }
221
222    if (cmd.kind === 'speed') {
223      const speed = Number((await $.store.get('speed')) ?? DEFAULT_SPEED)
224      if (cmd.value === undefined) {
225        return { text: `TTS speed ${speed} (${MIN_SPEED} slow to ${MAX_SPEED} fast, 1 normal)` }
226      }
227      await $.store.set('speed', cmd.value)
228      await speak($, `Speed ${cmd.value}`)
229      return { text: `TTS speed ${cmd.value}` }
230    }
231
232    if (cmd.kind === 'volume') {
233      if (cmd.value === undefined) return { text: `TTS volume ${volume}/10` }
234      await $.store.set('volume', cmd.value)
235      await speak($, `Volume ${cmd.value}`)
236      return { text: `TTS volume ${cmd.value}/10` }
237    }
238
239    const muted = await isMuted($)
240    const next =
241      cmd.kind === 'toggle' ? !muted : cmd.kind === 'mute' ? true : cmd.kind === 'unmute' ? false : muted
242    if (cmd.kind !== 'status') {
243      await $.state.set(MUTED, next)
244      await showStatus($)
245      if (next) await stop($)
246    }
247    const where = String((await $.store.get('devices')) ?? 'default')
248    const pace = Number((await $.store.get('speed')) ?? DEFAULT_SPEED)
249    const chosen = String((await $.store.get('voice')) ?? '') || 'default'
250    const reads = String((await $.store.get('lang')) ?? '') || 'auto'
251    return {
252      text: `${next ? 'TTS muted' : 'TTS on'} (this session), volume ${volume}/10, speed ${pace}, voice ${chosen}, language ${reads}, output ${where}`,
253    }
254  })
255
256  // New prompt: cut off whatever this session is still saying.
257  on('turn.start', async ($, e, next) => {
258    await stop($)
259    return next(e)
260  })
261
262  let narratedTurn: string | undefined
263
264  // Before the first tool call of a turn, read the narration that came ahead of it.
265  on('turn.step', async function* ($, e, next) {
266    let narration = ''
267    for await (const chunk of next(e)) {
268      if (chunk.kind === 'text') narration += chunk.text
269      if (
270        chunk.kind === 'tool' &&
271        e.agentId === undefined &&
272        narratedTurn !== e.turnId &&
273        narration.trim().length > MIN_NARRATION &&
274        !hasSummary(narration)
275      ) {
276        narratedTurn = e.turnId
277        await speak($, speechText(narration))
278      }
279      yield chunk
280    }
281  })
282
283  // The final answer comes from the event itself, so it is always the latest one.
284  on('turn.complete', async ($, e, next) => {
285    if (e.agentId === undefined && e.reason === 'answer' && e.answer.trim()) {
286      await speak($, speechText(e.answer))
287    }
288    return next(e)
289  })
290
291  on('session.end', async ($, e, next) => {
292    await stop($)
293    await cli($, ['session-state', '--session', await sessionId($), '--clear'])
294    return next(e)
295  })
296}
297
hooks/speech.ts 79 lines
1const OPEN = '<!-- TTS_SUMMARY'
2const CLOSE = 'TTS_SUMMARY -->'
3
4/** The text to read aloud for a reply: the TTS_SUMMARY block if present, else the whole text. */
5export const speechText = (text: string): string => {
6  const start = text.indexOf(OPEN)
7  if (start !== -1) {
8    const rest = text.slice(start + OPEN.length)
9    const end = rest.indexOf(CLOSE)
10    const summary = end === -1 ? '' : rest.slice(0, end).trim()
11    if (summary) return summary
12  }
13  return text.trim().slice(0, 5000)
14}
15
16/** Narration before a tool call is skipped when it carries a summary block; the final answer reads that. */
17export const hasSummary = (text: string): boolean => text.includes(OPEN)
18
19export type TtsCommand =
20  | { kind: 'toggle' | 'mute' | 'unmute' | 'status' | 'disable' | 'enable' }
21  | { kind: 'volume'; value?: number }
22  | { kind: 'device'; value?: string }
23  | { kind: 'mic'; value?: string }
24  | { kind: 'speed'; value?: number }
25  | { kind: 'voice'; value?: string }
26  | { kind: 'lang'; value?: string }
27  | { kind: 'default'; value?: 'on' | 'off' }
28  | { kind: 'update'; value?: 'check' | 'on' | 'off' }
29
30export const DEFAULT_VOLUME = 10
31export const DEFAULT_SPEED = 1
32export const MIN_SPEED = 0.5
33export const MAX_SPEED = 1.5
34
35/** Parses `/tts` arguments; undefined when they make no sense. Volume is a whole number 1-10. */
36export const parseTtsArgs = (args: string): TtsCommand | undefined => {
37  const trimmed = args.trim()
38  const [word = '', ...restWords] = trimmed.toLowerCase().split(/\s+/)
39  const rest = restWords.join(' ')
40  if (word === '') return { kind: 'toggle' }
41  if (word === 'mute' || word === 'off') return { kind: 'mute' }
42  if (word === 'unmute' || word === 'on') return { kind: 'unmute' }
43  if (word === 'disable') return { kind: 'disable' }
44  if (word === 'enable') return { kind: 'enable' }
45  if (word === 'status') return { kind: 'status' }
46  if (word === 'volume' || word === 'vol') {
47    if (rest === '') return { kind: 'volume' }
48    const value = Number(rest)
49    return Number.isInteger(value) && value >= 1 && value <= 10 ? { kind: 'volume', value } : undefined
50  }
51  if (word === 'device' || word === 'devices' || word === 'output') {
52    return rest === '' ? { kind: 'device' } : { kind: 'device', value: rest.replace(/\s*,\s*/g, ',') }
53  }
54  if (word === 'mic' || word === 'microphone') {
55    const name = rest.replace(/\s*\(default\)\s*$/, '').trim()
56    return name === '' ? { kind: 'mic' } : { kind: 'mic', value: name }
57  }
58  if (word === 'speed' || word === 'pace') {
59    if (rest === '') return { kind: 'speed' }
60    const value = Number(rest)
61    return Number.isFinite(value) && value >= MIN_SPEED && value <= MAX_SPEED ? { kind: 'speed', value } : undefined
62  }
63  if (word === 'default' || word === 'startup') {
64    if (rest === '') return { kind: 'default' }
65    if (rest === 'on' || rest === 'off') return { kind: 'default', value: rest }
66  }
67  if (word === 'voice' || word === 'voices') {
68    return rest === '' || word === 'voices' ? { kind: 'voice' } : { kind: 'voice', value: rest }
69  }
70  if (word === 'lang' || word === 'language') {
71    return rest === '' ? { kind: 'lang' } : { kind: 'lang', value: rest }
72  }
73  if (word === 'update' || word === 'upgrade') {
74    if (rest === '') return { kind: 'update' }
75    if (rest === 'check' || rest === 'on' || rest === 'off') return { kind: 'update', value: rest }
76  }
77  return undefined
78}
79
types/index.d.ts 8 lines
1declare module 'claude-code' {
2  interface PluginState {
3    // `disabled`: this session is switched off completely (/tts disable).
4    // Whether this session is muted; unset until the person mutes or unmutes it here.
5    'claudio-tts': { muted: boolean; disabled: boolean }
6  }
7}
8