Reads Claude's replies aloud as they stream, sentence by sentence, through VOICEVOX, Irodori-TTS or any command.

Claude Code の mod(関数フックのプラグイン)。Claude の返答をストリーミング中に一文ずつ読み上げます。
turn.step の text チャンクを文に分割し、合成と再生をパイプラインで流す(コードブロック・表・URL は読み上げから除外してます)読み上げ用のテキストでは、箇条書きの印を除き、括弧を空白に置き換えて中身を読みます。明確なファイルパスはファイル名と行番号に短縮します。負数・小数・C++・--help など、意味を持つ記号は残します。
Claude Code 2.1.287 以降(Mods が既定で有効なバージョン)。どちらか一方で使えるようにできます。
ターミナルの Claude Code のプロンプトで:
/plugin install tts --marketplace kajidog/cc-mods-tts
「Add marketplace?」に y、スコープは user(既定)のまま Enter。そのセッションからすぐ有効になり、以降のセッションでも読み込まれます。更新は claude plugin update tts のあと /reload-plugins。
デスクトップアプリの Code タブではこのコマンドは使えません。ターミナルから user スコープで入れれば、デスクトップアプリのローカルセッションでも読み込まれます。
git clone https://github.com/kajidog/cc-mods-tts.git ~/cc-mods-tts
claude --plugin-dir ~/cc-mods-tts
常に読み込むなら ~/.claude/settings.json の env に "CLAUDE_CODE_PLUGIN_DIRS": "~/cc-mods-tts"。この方法ではクローンしたフォルダのコードがそのまま動き、編集するとセッション中に読み込み直されます。
/tts engine で voicevox が ● になっているか確認/tts voice で選ぶ(既定は「ずんだもん・ノーマル」)/tts test で鳴るか確認Windows ネイティブの Claude Code で使えます。VOICEVOX も同じ Windows 上で起動すれば、既定の http://localhost:50021 のまま接続できます。
player: auto は ffplay があればそれを使い、なければ Windows 標準の PowerShell で再生します。この Mod のために WSL や sh・mktemp を追加する必要はありません。PowerShell 再生時の音量変更と、音声の前後の無音除去には ffmpeg が必要です。ない場合は元の音声を再生します。
WSL 内のターミナルで動かす場合は Linux と同じ構成で、powershell.exe による Windows 側の再生も使えます。
sample.wav)をサーバー側の voices/ に置く。拡張子を除いたファイル名(この例では sample)が声の ID になります/tts engine irodori で切り替え、/tts voice で追加した声を選ぶ(/tts voice sample で直接指定もできます)/tts test で鳴るか確認voices/ は Irodori-TTS-Server 側のフォルダです。IRODORI_VOICES_DIR を設定している場合は、そのフォルダに置いてください。Irodori は声を選ぶまで読み上げができません。
対応する音声形式やアップロード API など、詳しくは公式の音声追加・管理手順(Voice Management)を参照してください。
A で入れた場合:
claude plugin uninstall tts@cc-mods-tts
claude plugin marketplace remove cc-mods-tts # マーケットプレイスも消すなら
起動中のセッションには /reload-plugins で反映されます。
B の場合は --plugin-dir を付けずに起動するか、CLAUDE_CODE_PLUGIN_DIRS を settings.json から消して、クローンしたフォルダを削除します。
どちらの場合も、/config で変えた設定は ~/.claude/settings.json の pluginConfigs に残ります(A は tts@cc-mods-tts、B は tts@inline のキー)。不要なら手で消してください。
| コマンド | 内容 | ||
|---|---|---|---|
/tts | 状態表示(on/off・モード・エンジン・速度・音量) | ||
/tts help | コマンド一覧 | ||
/tts config | 設定の一覧と今の値 | ||
/tts on / /tts off | 有効・無効(/config の Enabled と同じ) | ||
/tts stop | 合成・再生中の処理とキューを止め、そのターンの残りの返答も読まない。次のターンから再開 | ||
/tts test [text] | 通常の返答と同じ整形を通してテスト読み上げ | ||
| `/tts mode [stream\ | final\ | notify]` | 読み上げモード切替。引数なしでモードを選ぶペインを開く |
/tts engine [name] | エンジン切替。引数なしで各エンジンが応答するか(ok / down / unset)を表示 | ||
/tts voice [id] | 引数なしで声を選ぶペインを開く(↑↓で移動、Enter で決定、Esc で閉じる)。id 指定で直接設定 | ||
/tts speed [n] / /tts volume [n] | 速度・音量(1 が等倍)。引数なしで候補を選ぶペインを開く |
声の候補が多い場合は64件ずつ表示します。p / n(Previous / Next)でページを切り替え、↑↓と Enter で選択できます。最初は現在の声があるページを開きます。
| モード | 読むもの |
|---|---|
stream(既定) | 返答をストリーミング中に一文ずつ |
final | ターン終了時に最終的な返答だけ(途中のツール前の文は読まない) |
notify | 短い通知だけ:「完了しました」「エラーで止まりました」など |
どのモードでも、以下の操作待ちのダイアログが出たら読み上げを止めて通知します:
| ダイアログ | 読む文(既定。/config の TTS notice で変えられる) |
|---|---|
| 質問(AskUserQuestion) | 1件: 質問があります。+質問文 / 複数: 質問が3件あります。 |
| プランの承認(ExitPlanMode) | プランの確認をお願いします。 |
| Bash の許可 | コマンドの確認をお願いします。+コマンドの説明(description) |
| その他のツールの許可 | ツールの許可をお願いします。 |
許可は classic.PermissionRequest(ダイアログを出すときに発火)で拾うので、ルールや auto モードで自動判定された呼び出しでは鳴らない。通知の音声は初回だけ合成し、以降は使い回す。
停止すると Mod 側の合成リクエストも中断します。ただし、TTS サーバーが切断後も計算を続ける場合、その計算の終了までは次の合成が待たされることがあります。
ffplay / powershell では、再生終了時に語尾が途切れるのを抑えるため、無音除去後の WAV に0.2秒の無音を足します。この処理には ffmpeg が必要です。ffmpeg がない場合は、無音除去・追加を行わず元の音声を再生します。
WSL で ffplay を使うと、音声が WSLg(PulseAudio → RDP)を経由するため、負荷によって「ジリジリ」とノイズが混じることがあります。気になる場合は /config の TTS: Player を powershell にすると、Windows 側で直接再生するので改善します(文ごとに PowerShell を起動するため、文と文の間は少し空きます)。
設定はすべて /config の tts にまとめてます。項目名はどれも TTS: / TTS notice: で始まります。
| 項目(キー) | 内容 | 既定 |
|---|---|---|
TTS: Enabled(enabled) | 読み上げる / 読み上げない | true |
TTS: Read mode(mode) | stream / final / notify(上記) | stream |
TTS: Engine(engine) | voicevox / irodori / command | voicevox |
TTS: Fallback engine(fallback) | メインのエンジンが失敗した文(サーバー停止・未設定など)を読むエンジン。メインが使えない間は、ステータス行に ⚠ tts: voicevox down のように出る | none |
TTS: Speed(speed) | 速度。エンジン側で合成時にかける(Irodori は 0.25〜4、VOICEVOX は speedScale)。command では無視 | 1.1 |
TTS: Volume(volume) | 音量。再生時にかけるのでどのエンジンでも効く。ffplay / paplay / afplay は再生コマンドで、aplay / powershell は ffmpeg があれば再生前に変える | 1 |
TTS: Player(player) | 再生コマンド。auto は Windows ネイティブでは ffplay → powershell.exe、それ以外では ffplay → paplay → aplay → afplay → powershell.exe の順に最初に見つかったもの | auto |
TTS: Max characters per response(maxCharsPerTurn) | 1回の応答で読む最大文字数。超えたら「途中を省略します。」と言って残りを飛ばし、応答の最後の1文だけ読む。0 で無制限 | 600 |
TTS: Stop on new prompt(interruptOnPrompt) | プロンプト送信で読み上げを止める | true |
TTS: Prompt for the model(modelPrompt) | 読み上げ中(stream / final モード)に Claude のシステムプロンプトへ足す文。説明・進捗報告は原則日本語にし、コード・引用・ユーザーの言語指定は保持する。空にすると何も足さない | 応答は音声で読み上げられています。説明やツール呼び出しの合間の進捗報告は原則として日本語で書いてください。ただし、コード・識別子・引用はそのまま保持し、ユーザーが指定した言語を優先してください。 |
TTS: Irodori URL(irodoriUrl) | Irodori-TTS の URL | http://localhost:8088 |
TTS: Irodori voice(irodoriVoice) | Irodori の声(/tts voice で選べる) | (未設定) |
TTS: VOICEVOX URL(voicevoxUrl) | VOICEVOX 互換エンジンの URL(AivisSpeech は http://localhost:10101) | http://localhost:50021 |
TTS: VOICEVOX speaker(voicevoxSpeaker) | VOICEVOX のスタイル id(/tts voice で選べる) | 3 |
TTS: Command(command) | engine=command のときのコマンド。空白区切り、{out} が書き出す wav のパスに置き換わり、本文は標準入力で渡る | (未設定) |
TTS notice: done / error / refusal(noticeDone / noticeError / noticeRefusal) | notify モードでターン終了時に言う文 | 完了しました。/ エラーで止まりました。/ 応答できませんでした。 |
TTS notice: question(noticeQuestion) | 質問が1件出たとき、質問文の前に言う文 | 質問があります。 |
TTS notice: questions(noticeQuestions) | 質問が複数出たときに言う文。{n} が件数になる | 質問が{n}件あります。 |
TTS notice: plan(noticePlan) | プランの承認待ちで言う文 | プランの確認をお願いします。 |
TTS notice: command(noticeCommand) | Bash の許可ダイアログで、コマンドの説明の前に言う文 | コマンドの確認をお願いします。 |
TTS notice: tool(noticeTool) | その他のツールの許可ダイアログで言う文 | ツールの許可をお願いします。 |
TTS notice: cut(noticeCut) | 文字数の上限で残りを飛ばすときに言う文。この後に応答の最後の1文を読む | 途中を省略します。 |
TTS notice: code(noticeCode) | コードブロックを飛ばすところで言う文(stream / final モード) | コードは省略します。 |
通知の文はどれも空にするとその通知の読み上げをスキップできます(注:スキップしても既存の再生は止まります)。 質問やコマンドの通知を空にした場合は、質問文・コマンドの説明も読みません。
Irodori の速度の対応範囲は 0.25〜4 です。Mod ではこの下限による入力エラーは出しませんが、範囲外の値は Irodori サーバーが拒否する場合があります。文字数の上限に、省略やコードを知らせる通知文と、省略後に読む最後の1文は含みません。上限をまたぐ文は途中まで読まずに飛ばします。最後の1文が300文字を超える場合は読みません。
音声ファイルはセッション用の一時ディレクトリに保存し、再生後に削除します。macOS / Linux / WSL では所有者のみアクセスできる権限で作成し、Windows ネイティブではユーザーの一時フォルダ(%TEMP%)のアクセス権を引き継ぎます。通知キャッシュはセッション終了時・Mod の再読み込み時に削除します。プロセスの強制終了などで清掃が実行できなかった場合は、一時ディレクトリに残ることがあります。
curl(Windows では curl.exe)mktemp(macOS / Linux / WSL の一時ディレクトリ作成に使用)powershell.exe、一時ファイルの管理と再生に使用)ffplay / paplay / aplay / afplay / powershell.exe(Windows・WSL)のいずれか(player: auto で自動選択)エンジンは hooks/engines/ に1ファイルずつある(irodori.ts / voicevox.ts / command.ts)。追加するには:
hooks/engines/<name>.ts に Engine(engine.ts の型)を書く。load(options, isWindows) で自分の設定を読み、synthesis(text, out) で wav を out に書くコマンド列を返す。声の一覧があれば voices もhooks/engines/index.ts の ENGINES に足す.claude-plugin/plugin.json の userConfig.engine.options に名前を足し、そのエンジンの設定項目を追加するnpm install
CLAUDE_CODE_PLUGIN_DIR_WATCH=1 claude --init-only --plugin-dir . # .claude-plugin/types/ を生成(型チェックに必要)
npm run check # lint・validate・型チェック・test をまとめて実行
個別には npm run lint / npm run validate / npm run typecheck / npm test で実行できます。Biome と TypeScript は devDependencies で固定し、claude は PATH にあるものを使います。
CI(.github/workflows/ci.yml)は main への push と PR で lint を実行し、Linux・Windows で validate・型チェック・test を実行します。型チェックの前に、上記の初期化コマンドで .claude-plugin/types/ を生成します。対話セッションや API キーは不要です。
API は early access のため、Claude Code の更新で変わることがある(CI で確認しているバージョン: 2.1.296)。
変更履歴は CHANGELOG.md にまとめています。Changesets で変更メモを集め、GitHub Actions がリリース用 PR を作成・更新します。
開発時は Node.js 22.11 以降の 22 系、24 系、または 26 以降と、npm 10.9 以降を使います。
npm ci
npm run changeset
patch(修正)・minor(機能追加)・major(互換性のない変更)を選び、変更内容を日本語で入力します。生成された .changeset/*.md をコードと一緒にコミットしてください。ドキュメントや CI だけの変更ではメモは不要です。
chore: release tts という PR を作成・更新する。CHANGELOG への追記と、package.json・package-lock.json・.claude-plugin/plugin.json のバージョン更新がまとまります。claude plugin update tts と /reload-plugins で更新できます。最初に GitHub リポジトリの Settings → Actions → General → Workflow permissions で Allow GitHub Actions to create and approve pull requests を有効にしてください(Changesets の設定手順)。標準の GITHUB_TOKEN を使うため、追加のトークンは不要です。自動作成された PR の CI が承認待ちになった場合は、PR 画面の Approve workflows to run で実行を承認してください(GitHub の仕様)。
手元でリリース差分を生成する場合は npm run release:version、バージョンの整合性確認は npm run check:version を使います。履歴の生成時はプラグイン側も同期するため、changeset version 単体では実行しません。バージョン管理用の package.json は private: true で、npm 公開や GitHub Release の作成は行いません。
hooks/register.tsx 621 lines1import { atom, read, update } from "claude-code"
2import type { EngineInterface, Register } from "claude-code"
3
4import type { Picker, Setting } from "../types"
5
6import { Channel } from "./channel"
7import { type Notices, type ReadMode, type TtsConfig, toConfig } from "./config"
8import { DIALOG_TOOLS, dialogSpeech } from "./dialog-text"
9import { type EngineName, ENGINES, type LoadedEngine } from "./engines"
10import { type Pin, PINS_KEY, type Pins, pinnedOptions, toPins } from "./pins"
11import { removeDirectoryCommand, removeFilesCommand, tempDirectoryCommand } from "./platform"
12import { findPlayerCommand, playCommand, prepareAudioCommand } from "./player"
13import { SpeechProcess } from "./process"
14import { type Change, MODE_LABELS, MODES, optionLine, planChange, SPEEDS, VOLUMES } from "./settings"
15import { capLength, SKIPPED_CODE, type Spoken, SpeechSplitter, toSentences } from "./speech-text"
16
17const HELP = [
18 "/tts state: on/off, mode, engine, speed, volume",
19 "/tts on | off read replies aloud or not",
20 "/tts stop stop speech and mute the rest of this turn",
21 "/tts test [text] read a test sentence",
22 "/tts mode [name] stream | final | notify; no name opens a picker",
23 "/tts engine [name] switch engines; no name checks whether each one answers",
24 "/tts voice [id] pick a voice; no id opens a picker",
25 "/tts speed [n] speech speed (1: as is, Irodori: 0.25–4); no value opens a picker",
26 "/tts volume [n] playback volume (1: as is); no value opens a picker",
27 "/tts pin | unpin keep this directory's engine and voice apart from the rest, or stop",
28 "/tts config every setting and its value; change them in /config → tts",
29 "/tts help this list",
30].join("\n")
31/** Which notice notify mode says as a turn ends; an interrupted turn stays silent. */
32const TURN_NOTICES: Partial<Record<string, keyof Notices>> = { answer: "done", error: "error", refusal: "refusal" }
33const PICKER_PANE = "tts-picker"
34// Claude Code refuses a Select with more than 64 options.
35const PICKER_PAGE_SIZE = 64
36const picker = atom({ plugin: "tts", key: "picker" } as const, null as Picker | null)
37// Survives a reload so the next module can remove the old module's audio files.
38const audioDirectory = atom({ plugin: "tts", key: "audioDirectory" } as const, null as string | null)
39
40/** Irodori runs diffusion with a concurrency of 1, so one request can take minutes. */
41const SYNTH_TIMEOUT_MS = 300_000
42/** A last sentence longer than this is not said after the cut: a long line without a stop is no question. */
43const MAX_TAIL = 300
44
45/**
46 * How much of one model response has been queued; past the limit the rest is dropped,
47 * save its last sentence (`tail`), which `finishResponse` says: it often asks what to do next.
48 */
49type Budget = { spoken: number; isCut: boolean; tail?: string }
50
51/** `cacheAs` keeps the synthesized clip at that path for the next time. */
52type Sentence = { text: string; generation: number; cacheAs?: string }
53type Clip = { file: string; generation: number; isCached?: boolean }
54
55export const register: Register = (on, options) => {
56 let isWindows = false
57 let config = toConfig(options)
58 let engine = ENGINES[config.engine].load(options)
59 let fallback = config.fallback === undefined ? undefined : ENGINES[config.fallback].load(options)
60 // The working directory's pin, if any, lays its engine and voice over the options.
61 let cwd = ""
62 let pins: Pins = {}
63
64 // `generation` moves on at every stop: queued work of an older one is dropped.
65 let generation = 0
66 let isTurnMuted = false
67 const { isEnabled } = config
68 let player: string | undefined
69 const synthesisProcess = new SpeechProcess()
70 const playbackProcess = new SpeechProcess()
71 const sentences = new Channel<Sentence>()
72 const clips = new Channel<Clip>()
73 // Notices are few and fixed: each is synthesized once per load, then replayed at once.
74 const cachedNotices = new Map<string, string>()
75 let directory = ""
76 let noticeCount = 0
77 // Set while the engine in use fails ("voicevox down"); shown beside playback until it reads again.
78 let engineIssue: string | undefined
79 // Set by session.start: probes the engine in use and shows the result.
80 let recheckEngine = () => {}
81
82 /** Rebuilds the engines from the options with this directory's pin laid over them. */
83 const applyPins = (stored: Pins) => {
84 pins = stored
85 const effective = pinnedOptions(options, pins[cwd])
86 config = toConfig(effective)
87 engine = ENGINES[config.engine].load(effective, isWindows)
88 fallback = config.fallback === undefined ? undefined : ENGINES[config.fallback].load(effective, isWindows)
89 // Cached notices were spoken in the voice before.
90 cachedNotices.clear()
91 recheckEngine()
92 }
93
94 /** What saving `value` for `setting` changes, against the settings in effect here. */
95 const changeOf = (setting: Setting, value: string) => planChange(setting, value, { options, config, engine, cwd, pin: pins[cwd] })
96
97 /** Invalidates queued speech and closes the active synthesis and playback streams. */
98 const stop = async () => {
99 generation += 1
100 await Promise.all([synthesisProcess.stop(), playbackProcess.stop()])
101 }
102
103 /** Queues `text` to be said next; called right after a stop, it plays first. */
104 const queueNotice = (text: string) => {
105 if (text === "") return
106 const cached = cachedNotices.get(text)
107 if (cached !== undefined) return clips.push({ file: cached, generation, isCached: true })
108 sentences.push({ text, generation, cacheAs: `${directory}/notice-${noticeCount++}.wav` })
109 }
110
111 /** Queues `text` for synthesis, in pieces an engine takes. */
112 const say = (text: string) => {
113 if (text === "") return
114 for (const piece of capLength(text)) sentences.push({ text: piece, generation })
115 }
116
117 /**
118 * Queues sentences; with a budget, stops before the sentence that would pass the limit
119 * (never reading half of one) and says once that the rest is skipped.
120 */
121 const enqueue = (spoken: readonly Spoken[], budget?: Budget) => {
122 for (const text of spoken) {
123 // The code notice is the mod's, not the response's: it uses no budget and is never the last sentence.
124 if (text === SKIPPED_CODE) {
125 if (budget?.isCut !== true) say(config.notices.code)
126 continue
127 }
128 if (budget !== undefined && config.maxCharsPerTurn > 0) {
129 const length = Array.from(text).length
130 if (!budget.isCut && budget.spoken + length > config.maxCharsPerTurn) {
131 budget.isCut = true
132 say(config.notices.cut)
133 }
134 // Each skipped sentence may be the last, so the one that passes the limit counts too.
135 if (budget.isCut) {
136 budget.tail = text
137 continue
138 }
139 budget.spoken += length
140 }
141 say(text)
142 }
143 }
144
145 /** Ends a response: past the limit, says its last sentence after the cut notice. */
146 const finishResponse = (budget: Budget) => {
147 if (budget.tail !== undefined && Array.from(budget.tail).length <= MAX_TAIL) say(budget.tail)
148 budget.tail = undefined
149 }
150
151 on("session.start", async ($, e, next) => {
152 const started = await next(e)
153 isWindows = (await $.env.get("OS")) === "Windows_NT"
154 cwd = await $.session.cwd()
155 applyPins(toPins(await $.store.get(PINS_KEY)))
156 player = config.player === "auto" ? (await $.process.run(findPlayerCommand(isWindows))).stdout.trim() || undefined : config.player
157 if (player === undefined) $.ui.toast("tts: no audio player found (ffplay, paplay, aplay, afplay, powershell.exe)")
158
159 directory = await prepareDirectory($, isWindows)
160 let clipCount = 0
161 let lastErrorAt = 0
162
163 const noteEngine = (error?: unknown) => {
164 const issue = error === undefined ? undefined : `${config.engine} ${/is not set/.test(String(error)) ? "unset" : "down"}`
165 showEngineIssue(issue)
166 }
167 // The status line shows only a failing engine; Claude Code draws it as "⚠ tts: voicevox down".
168 const showEngineIssue = (issue: string | undefined) => {
169 if (issue === engineIssue) return
170 engineIssue = issue
171 $.ui.status(issue)
172 }
173 // The last module's line stays after a reload.
174 $.ui.status(undefined)
175
176 // The fallback reads on quietly: the status line says the engine is failing.
177 const reportFallback = (error: unknown) => {
178 $.ui.log(`tts: ${config.engine} failed, read with ${config.fallback}: ${String(error)}`, { to: "debug" })
179 }
180
181 const reportError = (error: unknown) => {
182 const message = error instanceof Error ? error.message : String(error)
183 $.ui.log(`tts: ${message}`, { to: "debug" })
184 // One toast per burst: a server that is down fails every sentence of a reply.
185 if (Date.now() - lastErrorAt > 30_000) $.ui.toast(`tts: ${message}`)
186 lastErrorAt = Date.now()
187 }
188
189 // Synthesis runs ahead of playback, one request at a time.
190 void (async () => {
191 for (;;) {
192 const sentence = await sentences.take()
193 if (sentence.generation !== generation) continue
194 const file = sentence.cacheAs ?? `${directory}/${String(clipCount++).padStart(6, "0")}.wav`
195 const isCurrent = () => sentence.generation === generation
196 let isFallback = false
197 try {
198 try {
199 await synthesize($, synthesisProcess, engine, sentence.text, file, isCurrent)
200 if (isCurrent()) noteEngine()
201 } catch (error) {
202 if (isCurrent()) noteEngine(error)
203 if (!isCurrent() || fallback === undefined || config.fallback === undefined) throw error
204 isFallback = true
205 await synthesize($, synthesisProcess, fallback, sentence.text, file, isCurrent)
206 reportFallback(error)
207 }
208 if (!isCurrent()) {
209 await $.process.run(removeFilesCommand([file], isWindows))
210 continue
211 }
212 // Done while the clip before plays; an untrimmed clip still plays.
213 await runProcess($, synthesisProcess, prepareAudioCommand(file, player, isWindows), undefined, 30_000).catch(() => undefined)
214 if (!isCurrent()) {
215 await $.process.run(removeFilesCommand([file, `${file}.t.wav`], isWindows))
216 continue
217 }
218 // A notice the fallback read is kept for this time only: the engine may be back next time.
219 const isCached = sentence.cacheAs !== undefined && !isFallback
220 if (isCached) cachedNotices.set(sentence.text, file)
221 clips.push({ file, generation: sentence.generation, isCached })
222 } catch (error) {
223 await $.process.run(removeFilesCommand([file], isWindows))
224 if (isCurrent()) reportError(error)
225 }
226 }
227 })()
228
229 void (async () => {
230 for (;;) {
231 const clip = await clips.take()
232 if (clip.generation === generation && player !== undefined) {
233 try {
234 await playbackProcess.read($.process.spawn({ argv: playCommand(player, clip.file, config.volume, isWindows) }))
235 } catch (error) {
236 if (clip.generation === generation) reportError(error)
237 }
238 }
239 const files = [`${clip.file}.v.wav`]
240 if (clip.isCached !== true) files.push(clip.file)
241 await $.process.run(removeFilesCommand(files, isWindows))
242 }
243 })()
244
245 // Checked at once, so a start or an engine switch shows a dead engine before anything is read.
246 recheckEngine = () => {
247 showEngineIssue(undefined)
248 if (!isEnabled) return
249 const checked = engine
250 void checkEngine($, checked).then(state => {
251 if (checked === engine && state !== "ok") showEngineIssue(`${config.engine} ${state}`)
252 })
253 }
254 recheckEngine()
255
256 await $.command.register({
257 name: "tts",
258 description: "Read replies aloud: on, off, stop, test, mode, engine, voice, speed, volume, pin, unpin, config, help",
259 argumentHint: "[on|off|stop|test [text]|mode|engine|voice|speed|volume [value]|pin|unpin|config|help]",
260 // `/tts stop` has to land while a reply is still being read.
261 immediate: true,
262 })
263 return started
264 })
265
266 on("session.end", async ($, e, next) => {
267 await stop()
268 cachedNotices.clear()
269 if (directory !== "") await $.process.run(removeDirectoryCommand(directory, isWindows))
270 directory = ""
271 await update($, audioDirectory, () => null)
272 return next(e)
273 }).catch(($, e, next) => next(e))
274
275 // /clear, resume and /branch keep this module alive without another session.start, and reset
276 // $.state, which holds the directory for the next reload to remove.
277 on("classic.SessionStart", { source: ["clear", "resume", "fork"] }, async ($, e, next) => {
278 if (directory === "") directory = await prepareDirectory($, isWindows)
279 else await update($, audioDirectory, () => directory)
280 return next(e)
281 }).catch(($, e, next) => next(e))
282
283 on("command.run", { command: "tts" }, async ($, e) => {
284 const [action = "", ...rest] = e.args.trim().split(/\s+/)
285 switch (action) {
286 case "on":
287 case "off": {
288 if (action === "off") await stop()
289 const set = await $.config.set({ key: "tts.enabled", value: action === "on" })
290 return { text: "deny" in set && set.deny ? set.deny : action }
291 }
292 case "stop": {
293 isTurnMuted = true
294 await stop()
295 return { text: "stopped" }
296 }
297 case "test": {
298 const text = e.args.trim().replace(/^test(?:\s+|$)/, "") || "読み上げのテストです。"
299 const texts = toSentences(text)
300 if (texts.length === 0) return { text: "no text to read after formatting" }
301 try {
302 engine.synthesis(text, "probe.wav")
303 } catch (error) {
304 if (fallback === undefined) return { text: error instanceof Error ? error.message : String(error) }
305 }
306 enqueue(texts)
307 return { text: `testing ${config.engine} (the first Irodori request loads the model and can take a while)` }
308 }
309 case "voice": {
310 const { voices } = engine
311 if (voices === undefined) return { text: `engine=${config.engine} has no voices; set it up in /config` }
312 const voice = rest.join(" ")
313 if (voice === "") {
314 const listed = await $.http.fetch(voices.url)
315 if (!listed.ok) return { text: `GET ${voices.url} failed with ${listed.status}` }
316 const options = voices.parse(listed.text)
317 const title = `${config.engine} voice`
318 if (await openPicker($, { setting: "voice", title, options, value: voices.current })) {
319 return { text: "pick a voice in the pane (Esc closes it)" }
320 }
321 return { text: [`${config.engine} voices (/tts voice <id>):`, ...options.map(optionLine)].join("\n") }
322 }
323 return { text: await commit($, changeOf("voice", voice), applyPins) }
324 }
325 case "mode": {
326 const name = rest[0] as ReadMode | undefined
327 if (name === undefined) {
328 const options = MODES.map(mode => ({ value: mode, label: MODE_LABELS[mode] }))
329 if (await openPicker($, { setting: "mode", title: "read mode", options, value: config.mode })) {
330 return { text: "pick a mode in the pane (Esc closes it)" }
331 }
332 return { text: `mode: ${config.mode} (${MODES.join(" | ")})` }
333 }
334 return { text: await commit($, changeOf("mode", name), applyPins) }
335 }
336 case "speed":
337 case "volume": {
338 const value = rest[0]
339 if (value === undefined) {
340 const current = String(action === "speed" ? options.speed : config.volume)
341 const choices = (action === "speed" ? SPEEDS : VOLUMES).map(each => ({ value: each, label: `× ${each}` }))
342 if (await openPicker($, { setting: action, title: action, options: choices, value: current })) {
343 return { text: `pick a ${action} in the pane (Esc closes it)` }
344 }
345 return { text: `${action}: ${current}` }
346 }
347 return { text: await commit($, changeOf(action, value), applyPins) }
348 }
349 case "pin": {
350 const pin: Pin = { engine: config.engine, voices: engine.voices ? { [config.engine]: engine.voices.current } : {} }
351 const stored = { ...toPins(await $.store.get(PINS_KEY)), [cwd]: pin }
352 await $.store.set(PINS_KEY, stored)
353 applyPins(stored)
354 return {
355 text: `pinned ${cwd}: ${config.engine} (${engine.voice})\nengine and voice changes here stay in this directory; /tts unpin undoes it`,
356 }
357 }
358 case "unpin": {
359 if (pins[cwd] === undefined) return { text: "this directory is not pinned" }
360 const before = `${config.engine} (${engine.voice})`
361 const { [cwd]: _unpinned, ...stored } = toPins(await $.store.get(PINS_KEY))
362 await $.store.set(PINS_KEY, stored)
363 applyPins(stored)
364 return { text: `unpinned ${cwd}: ${before} → ${config.engine} (${engine.voice})` }
365 }
366 case "engine": {
367 const name = rest[0]
368 if (name === undefined) {
369 const names = Object.keys(ENGINES) as EngineName[]
370 const rows = await Promise.all(names.map(each => engineRow($, each, ENGINES[each].load(options, isWindows), config)))
371 return { text: ["engines (● answers, ○ down or unset; /tts engine <name> switches):", ...rows].join("\n") }
372 }
373 return { text: await commit($, changeOf("engine", name), applyPins) }
374 }
375 case "help":
376 return { text: HELP }
377 case "config": {
378 const width = Math.max(...Object.keys(options).map(key => key.length))
379 const rows = Object.entries(options).map(([key, value]) => ` ${key.padEnd(width)} ${JSON.stringify(value)}`)
380 const pinned = pins[cwd] === undefined ? [] : [`engine and voice are pinned here: ${config.engine} (${engine.voice})`]
381 return { text: ["settings (change them in /config → tts):", ...rows, ...pinned].join("\n") }
382 }
383 case "": {
384 const marks = await Promise.all([answerMark($, engine), fallback && answerMark($, fallback)])
385 return {
386 text: [
387 `${isEnabled ? "on" : "off"} (mode: ${config.mode})`,
388 `engine: ${marks[0]} ${config.engine} (${engine.voice})${config.fallback ? `, fallback: ${marks[1]} ${config.fallback}` : ""}`,
389 `speed: ${options.speed}, volume: ${config.volume}`,
390 ...(pins[cwd] === undefined ? [] : [`engine and voice pinned to ${cwd} (/tts unpin)`]),
391 `player: ${player ?? "none"}`,
392 "/tts help lists the commands",
393 ].join("\n"),
394 }
395 }
396 default:
397 return { text: `unknown: ${action}\n\n${HELP}` }
398 }
399 })
400
401 // The picker a `/tts` command without a value opens: a pick saves the setting and closes the pane.
402 on("ui.render", { component: "Pane", requestId: PICKER_PANE }, async ($, e) => {
403 const shown = await read($, picker)
404 if (e.surface === "mobile" || shown === null || shown.options.length === 0) {
405 const { Text } = $.ui.resolve(e)
406 return <Text dimColor>{shown === null || shown.options.length === 0 ? "Nothing to pick." : shown.options.map(optionLine).join("\n")}</Text>
407 }
408 const { Box, Text, Select, Button } = $.ui.resolve(e)
409 const page = shown.page ?? Math.floor(Math.max(0, shown.options.findIndex(option => option.value === shown.value)) / PICKER_PAGE_SIZE)
410 const pages = Math.ceil(shown.options.length / PICKER_PAGE_SIZE)
411 const choices = shown.options.slice(page * PICKER_PAGE_SIZE, (page + 1) * PICKER_PAGE_SIZE)
412 const turnPage = (page: number) => update($, picker, () => ({ ...shown, page }))
413 return (
414 <Box flexDirection="column">
415 <Text dimColor>{`${shown.title}: ↑↓ move, Enter picks, Esc closes`}</Text>
416 {pages > 1 && (
417 <Box>
418 {page > 0 && <Button key="previous-page" hotkey="p" onPress={() => turnPage(page - 1)}>Previous</Button>}
419 <Text dimColor>{` Page ${page + 1}/${pages} `}</Text>
420 {page + 1 < pages && <Button key="next-page" hotkey="n" onPress={() => turnPage(page + 1)}>Next</Button>}
421 </Box>
422 )}
423 <Select
424 key={`${shown.setting}-${page}`}
425 options={choices.map(option => ({ value: option.value, label: option.label ?? option.value }))}
426 value={choices.some(option => option.value === shown.value) ? shown.value : undefined}
427 autoFocus
428 onSelect={async value => {
429 await $.ui.close({ id: PICKER_PANE })
430 $.ui.toast(`tts: ${await commit($, changeOf(shown.setting, value), applyPins)}`)
431 }}
432 />
433 </Box>
434 )
435 })
436
437 on("prompt.submit", async ($, e, next) => {
438 if (config.interruptOnPrompt) await stop()
439 return next(e)
440 }).catch(($, e, next) => next(e))
441
442 // Asks the model to write what reads well aloud; notify mode reads none of its text.
443 on("prompt.compose", async ($, e, next) => {
444 const composed = await next(e)
445 if (!isEnabled || config.mode === "notify" || config.modelPrompt === "") return composed
446 return { sections: [...composed.sections, { id: "tts:spoken", text: config.modelPrompt, scope: "session" as const }] }
447 }).catch(($, e, next) => next(e))
448
449 on("turn.start", async ($, e, next) => {
450 isTurnMuted = false
451 return next(e)
452 })
453
454 on("turn.complete", async ($, e, next) => {
455 if (e.reason === "aborted" && e.agentId === undefined) await stop()
456 if (isEnabled && !isTurnMuted && e.agentId === undefined) {
457 if (config.mode === "final" && e.reason === "answer") {
458 const budget: Budget = { spoken: 0, isCut: false }
459 enqueue(toSentences(e.answer), budget)
460 finishResponse(budget)
461 }
462 const notice = TURN_NOTICES[e.reason]
463 if (config.mode === "notify" && notice !== undefined && config.notices[notice] !== "") enqueue([config.notices[notice]])
464 }
465 return next(e)
466 })
467
468 // Whatever the mode, a dialog waiting for the person cuts off the reading and is announced.
469 on("tool.call", async ($, e, next) => {
470 if (isEnabled && DIALOG_TOOLS.has(e.tool)) {
471 await stop()
472 const { notice, detail } = dialogSpeech(e.tool, e, config.notices)
473 queueNotice(notice)
474 enqueue(detail)
475 }
476 return next(e)
477 }).catch(($, e, next) => next(e))
478
479 // Raised as a permission dialog is about to be shown, not for a call a rule or the mode settles.
480 on("classic.PermissionRequest", async ($, e, next) => {
481 if (isEnabled && !DIALOG_TOOLS.has(e.tool_name)) {
482 await stop()
483 const { notice, detail } = dialogSpeech(e.tool_name, e.tool_input, config.notices)
484 queueNotice(notice)
485 enqueue(detail)
486 }
487 return next(e)
488 }).catch(($, e, next) => next(e))
489
490 // Speaks the main loop's text as it streams; thinking, tool calls and subagents stay silent.
491 on("turn.step", async function* ($, e, next) {
492 const stream = next(e)
493 if (!isEnabled || config.mode !== "stream" || e.agentId !== undefined) return yield* stream
494
495 // Each response has its own budget: text before a tool call does not use up the answer's.
496 const budget: Budget = { spoken: 0, isCut: false }
497 const splitter = new SpeechSplitter()
498 const responseGeneration = generation
499 const queueResponse = (texts: readonly Spoken[]) => {
500 if (!isTurnMuted && responseGeneration === generation) enqueue(texts, budget)
501 }
502 let block: number | undefined
503 try {
504 for await (const chunk of stream) {
505 if (chunk.kind === "text") {
506 if (block !== undefined && block !== chunk.index) queueResponse(splitter.flush())
507 block = chunk.index
508 queueResponse(splitter.push(chunk.text))
509 }
510 yield chunk
511 }
512 } finally {
513 // Runs even when the engine stops pulling after the last chunk: the line without a newline is said too.
514 queueResponse(splitter.flush())
515 if (!isTurnMuted && responseGeneration === generation) finishResponse(budget)
516 }
517 return stream.result
518 })
519}
520
521/** Creates a temporary directory, removing the previous module's files after a reload. */
522async function prepareDirectory($: EngineInterface, isWindows: boolean): Promise<string> {
523 const previous = await read($, audioDirectory)
524 if (previous !== null) await $.process.run(removeDirectoryCommand(previous, isWindows))
525 const created = await $.process.run(tempDirectoryCommand(isWindows))
526 if (created.exitCode !== 0 || created.stdout.trim() === "") throw new Error("tts: could not create a private audio directory")
527 const directory = created.stdout.trim()
528 await update($, audioDirectory, () => directory)
529 return directory
530}
531
532
533// Everything handed $ stays in this file: the validator follows $ into this file's functions only.
534
535/** Runs an engine's commands for one sentence; throws on the first that fails. */
536async function synthesize(
537 $: EngineInterface,
538 process: SpeechProcess,
539 engine: LoadedEngine,
540 text: string,
541 file: string,
542 isCurrent: () => boolean,
543) {
544 let stdout = ""
545 for (const step of engine.synthesis(text, file)) {
546 if (!isCurrent()) return
547 const stdin = typeof step.stdin === "function" ? step.stdin(stdout) : step.stdin
548 const ran = await runProcess($, process, step.argv, stdin, SYNTH_TIMEOUT_MS)
549 if (ran.code !== 0) {
550 throw new Error(`${step.argv[0]} exited ${ran.code}: ${(ran.stderr || ran.stdout).trim().slice(0, 200)}`)
551 }
552 stdout = ran.stdout
553 }
554}
555
556/** A cancellable command with the same time limit synthesis had under process.run. */
557async function runProcess($: EngineInterface, process: SpeechProcess, argv: string[], input: string | undefined, timeoutMs: number) {
558 const running = process.read($.process.spawn({ argv, input }))
559 let timedOut = false
560 const timer = $.clock.after(timeoutMs, () => {
561 timedOut = true
562 void process.stop()
563 })
564 try {
565 const result = await running
566 if (timedOut) throw new Error(`${argv[0]} timed out`)
567 return result
568 } finally {
569 timer.cancel()
570 }
571}
572
573/** Whether the engine can read now: set up (its synthesis builds) and answering its probe. */
574async function checkEngine($: EngineInterface, engine: LoadedEngine): Promise<"ok" | "down" | "unset"> {
575 if (engine.probe === undefined) return "unset"
576 try {
577 engine.synthesis("", "probe.wav")
578 } catch {
579 return "unset"
580 }
581 const probed = await $.process.run(engine.probe, { timeoutMs: 5_000 }).catch(() => undefined)
582 return probed?.exitCode === 0 ? "ok" : "down"
583}
584
585/** ● when the engine answers now, ○ when it is down or not set up. */
586async function answerMark($: EngineInterface, engine: LoadedEngine): Promise<string> {
587 return (await checkEngine($, engine)) === "ok" ? "●" : "○"
588}
589
590/** One line of `/tts engine`: whether the engine answers now, where, and its role. */
591async function engineRow($: EngineInterface, name: EngineName, engine: LoadedEngine, config: TtsConfig): Promise<string> {
592 const role = name === config.engine ? " ← in use" : name === config.fallback ? " ← fallback" : ""
593 return ` ${await answerMark($, engine)} ${name.padEnd(9)} ${engine.location}${role}`
594}
595
596/** Opens the picker pane on `shown`; false when the surface cannot place it. */
597async function openPicker($: EngineInterface, shown: Picker): Promise<boolean> {
598 await update($, picker, () => shown)
599 const rows = Math.min(shown.options.length + 3, 14)
600 const opened = await $.ui.open({ id: PICKER_PANE, title: `tts ${shown.setting}`, focus: true, closeOnEscape: true, rows })
601 return opened.isPlaced
602}
603
604
605/**
606 * Saves a change, to `/config` (which reloads the module) or to this
607 * directory's pin; answers how it reads, before and after, or why it did not.
608 */
609async function commit($: EngineInterface, change: Change, applyPins: (stored: Pins) => void): Promise<string> {
610 if ("error" in change) return change.error
611 if ("config" in change) {
612 const set = await $.config.set(change.config)
613 if ("deny" in set && set.deny) return set.deny
614 } else {
615 const stored = change.pins(toPins(await $.store.get(PINS_KEY)))
616 await $.store.set(PINS_KEY, stored)
617 applyPins(stored)
618 }
619 return `${change.label}: ${change.previous} → ${change.next}`
620}
621hooks/channel.ts 21 lines1/** An unbounded queue whose `take` waits for the next item. */
2export class Channel<T> {
3 private items: T[] = []
4 private waiter: (() => void) | undefined
5
6 get size(): number {
7 return this.items.length
8 }
9
10 push(item: T) {
11 this.items.push(item)
12 this.waiter?.()
13 this.waiter = undefined
14 }
15
16 async take(): Promise<T> {
17 while (this.items.length === 0) await new Promise<void>(resolve => (this.waiter = resolve))
18 return this.items.shift() as T
19 }
20}
21hooks/config.ts 70 lines1import type { PluginOptions } from "claude-code"
2
3import { type EngineName, ENGINES } from "./engines"
4
5/** stream: every sentence as it streams. final: the turn's answer once it ends. notify: short notices alone. */
6export type ReadMode = "stream" | "final" | "notify"
7
8/** The lines said as notices, each a `/config` row; an empty one stays silent. */
9export type Notices = {
10 done: string
11 error: string
12 refusal: string
13 question: string
14 /** `{n}` is the number of questions. */
15 questions: string
16 plan: string
17 command: string
18 tool: string
19 cut: string
20 code: string
21}
22
23const NOTICE_OPTIONS: Record<keyof Notices, string> = {
24 done: "noticeDone",
25 error: "noticeError",
26 refusal: "noticeRefusal",
27 question: "noticeQuestion",
28 questions: "noticeQuestions",
29 plan: "noticePlan",
30 command: "noticeCommand",
31 tool: "noticeTool",
32 cut: "noticeCut",
33 code: "noticeCode",
34}
35
36/** The settings every engine shares; each engine reads its own in `load`. */
37export type TtsConfig = {
38 isEnabled: boolean
39 mode: ReadMode
40 engine: EngineName
41 /** The engine tried when `engine` fails a sentence; undefined for none. */
42 fallback: EngineName | undefined
43 player: string
44 volume: number
45 maxCharsPerTurn: number
46 interruptOnPrompt: boolean
47 /** Added to the system prompt while replies are read; empty: nothing. */
48 modelPrompt: string
49 notices: Notices
50}
51
52export function toConfig(options: PluginOptions): TtsConfig {
53 const engine = String(options.engine)
54 const fallback = String(options.fallback)
55 return {
56 isEnabled: options.enabled !== false,
57 mode: options.mode as ReadMode,
58 engine: engine in ENGINES ? (engine as EngineName) : "voicevox",
59 fallback: fallback in ENGINES && fallback !== engine ? (fallback as EngineName) : undefined,
60 player: String(options.player),
61 volume: Number(options.volume),
62 maxCharsPerTurn: Number(options.maxCharsPerTurn),
63 interruptOnPrompt: options.interruptOnPrompt === true,
64 modelPrompt: String(options.modelPrompt ?? "").trim(),
65 notices: Object.fromEntries(
66 Object.entries(NOTICE_OPTIONS).map(([notice, option]) => [notice, String(options[option] ?? "").trim()]),
67 ) as Notices,
68 }
69}
70hooks/dialog-text.ts 34 lines1import type { Notices } from "./config"
2import { type Spoken, toSentences } from "./speech-text"
3
4/** Tools whose call itself opens a dialog for the person; the rest open one only to ask permission. */
5export const DIALOG_TOOLS: ReadonlySet<string> = new Set(["AskUserQuestion", "ExitPlanMode"])
6
7/**
8 * What is said as a dialog opens: `notice`, a short line from the settings
9 * that is cached and plays at once (empty: none), then `detail`, synthesized
10 * each time.
11 */
12export type DialogSpeech = { notice: string; detail: Spoken[] }
13
14export function dialogSpeech(tool: string, input: unknown, notices: Notices): DialogSpeech {
15 const fields = (typeof input === "object" && input !== null ? input : {}) as Record<string, unknown>
16 switch (tool) {
17 case "AskUserQuestion": {
18 const questions = Array.isArray(fields.questions) ? (fields.questions as { question?: unknown }[]) : []
19 // Several questions are counted, not read: the dialog shows them one at a time.
20 if (questions.length > 1) return { notice: notices.questions.replaceAll("{n}", String(questions.length)), detail: [] }
21 const question = questions[0]?.question
22 return { notice: notices.question, detail: notices.question !== "" && typeof question === "string" ? toSentences(question) : [] }
23 }
24 case "ExitPlanMode":
25 return { notice: notices.plan, detail: [] }
26 case "Bash": {
27 const { description } = fields
28 return { notice: notices.command, detail: notices.command !== "" && typeof description === "string" ? toSentences(description) : [] }
29 }
30 default:
31 return { notice: notices.tool, detail: [] }
32 }
33}
34hooks/engines/index.ts 12 lines1import { command } from "./command"
2import type { Engine } from "./engine"
3import { irodori } from "./irodori"
4import { voicevox } from "./voicevox"
5
6export type { Engine, LoadedEngine, Step } from "./engine"
7
8/** Every engine by the name `/tts engine` and the `engine` setting take. */
9export const ENGINES = { voicevox, irodori, command } satisfies Record<string, Engine>
10
11export type EngineName = keyof typeof ENGINES
12hooks/pins.ts 31 lines1import type { PluginOptions } from "claude-code"
2
3import { type EngineName, ENGINES } from "./engines"
4
5/** What a working directory pins: its engine, and the voice it chose for each engine. */
6export type Pin = { engine: EngineName; voices: Partial<Record<EngineName, string>> }
7
8/** Pins by working directory, as `$.store` keeps them. */
9export type Pins = Record<string, Pin>
10
11export const PINS_KEY = "pins"
12
13/** The options with a directory's pin laid over them; the options alone without one. */
14export function pinnedOptions(options: PluginOptions, pin: Pin | undefined): PluginOptions {
15 if (pin === undefined) return options
16 const pinned: Record<string, PluginOptions[string]> = { ...options, engine: pin.engine }
17 for (const [name, voice] of Object.entries(pin.voices) as [EngineName, string][]) {
18 const selected = ENGINES[name].load(options).voices?.select(voice)
19 if (selected !== undefined) pinned[selected.key.replace(/^tts\./, "")] = selected.value
20 }
21 return pinned
22}
23
24/** Reads `$.store`'s value as pins, dropping what is not one. */
25export function toPins(stored: unknown): Pins {
26 if (typeof stored !== "object" || stored === null) return {}
27 return Object.fromEntries(
28 Object.entries(stored).filter(([, pin]) => typeof pin === "object" && pin !== null && (pin as Pin).engine in ENGINES),
29 ) as Pins
30}
31hooks/platform.ts 36 lines1/** Quotes a literal for PowerShell, including the smart quotes it treats as delimiters. */
2export function quotePowerShell(value: string): string {
3 return `'${value.replace(/['‘’‚‛]/g, quote => quote + quote)}'`
4}
5
6/** Keep paths containing Japanese readable when Windows PowerShell writes to a pipe. */
7export function powershellCommand(script: string): string[] {
8 return [
9 "powershell.exe", "-NoProfile", "-NonInteractive", "-Command",
10 `$ErrorActionPreference = 'Stop'; [Console]::OutputEncoding = [Text.UTF8Encoding]::new(); ${script}`,
11 ]
12}
13
14/** Unix uses mode 0700; Windows inherits the user's temporary folder permissions. */
15export function tempDirectoryCommand(isWindows: boolean): string[] {
16 return isWindows
17 ? powershellCommand("$path = Join-Path ([IO.Path]::GetTempPath()) ('claude-tts-' + [Guid]::NewGuid().ToString('N')); [IO.Directory]::CreateDirectory($path).FullName")
18 : ["mktemp", "-d", "/tmp/claude-tts-XXXXXXXXXX"]
19}
20
21export function removeFilesCommand(files: string[], isWindows: boolean): string[] {
22 return isWindows
23 ? powershellCommand(`Remove-Item -LiteralPath ${files.map(quotePowerShell).join(", ")} -Force -ErrorAction SilentlyContinue; exit 0`)
24 : ["rm", "-f", ...files]
25}
26
27export function removeDirectoryCommand(directory: string, isWindows: boolean): string[] {
28 return isWindows
29 ? powershellCommand(`Remove-Item -LiteralPath ${quotePowerShell(directory)} -Recurse -Force -ErrorAction SilentlyContinue; exit 0`)
30 : ["rm", "-rf", "--", directory]
31}
32
33export function findProgramCommand(program: string, isWindows: boolean): string[] {
34 return isWindows ? ["where.exe", program] : ["sh", "-c", 'command -v "$1" >/dev/null', "probe", program]
35}
36hooks/player.ts 85 lines1import { powershellCommand, quotePowerShell } from "./platform"
2
3const TAIL_PADDING_SECONDS = 0.2
4
5/** Prints the first audio player found on this machine. */
6export function findPlayerCommand(isWindows: boolean): string[] {
7 if (isWindows) return powershellCommand("if (Get-Command ffplay -CommandType Application -ErrorAction SilentlyContinue) { 'ffplay' } else { 'powershell' }")
8 return [
9 "sh",
10 "-c",
11 // biome-ignore lint/suspicious/noTemplateCurlyInString: shell parameter expansion
12 "for p in ffplay paplay aplay afplay powershell.exe; do command -v $p >/dev/null && { echo ${p%.exe}; exit; }; done",
13 ]
14}
15
16/**
17 * Cuts long engine pauses, then gives ffplay and SoundPlayer time to finish speech.
18 * Padding must be in the WAV: ffplay does not flush an apad filter at EOF.
19 * Leaves the file as it was where ffmpeg is missing.
20 */
21export function prepareAudioCommand(file: string, player: string | undefined, isWindows = false): string[] {
22 const edge = (keep: number) => `silenceremove=start_periods=1:start_threshold=-45dB:start_silence=${keep}`
23 const padding = player === "ffplay" || player === "powershell" ? `,apad=pad_dur=${TAIL_PADDING_SECONDS}` : ""
24 const filter = `${edge(0.05)},areverse,${edge(0.2)},areverse${padding}`
25 if (isWindows) return powershellCommand([
26 "if (-not (Get-Command ffmpeg -CommandType Application -ErrorAction SilentlyContinue)) { exit 0 }",
27 `$file = ${quotePowerShell(file)}`,
28 `ffmpeg -loglevel error -y -i $file -af '${filter}' ($file + '.t.wav')`,
29 "if ($LASTEXITCODE -eq 0) { Move-Item -LiteralPath ($file + '.t.wav') -Destination $file -Force }",
30 ].join("; "))
31 return [
32 "sh",
33 "-c",
34 `command -v ffmpeg >/dev/null || exit 0; ffmpeg -loglevel error -y -i "$1" -af "${filter}" "$1.t.wav" && mv -f "$1.t.wav" "$1"`,
35 "tts-trim",
36 file,
37 ]
38}
39
40/**
41 * PowerShell that plays the wav at `path` (a PowerShell expression) to its end. The file goes in
42 * as a stream: SoundPlayer given a \\wsl.localhost path with capitals plays nothing, without error.
43 */
44function soundPlayerScript(path: string, variable = "$"): string {
45 const stream = `${variable}stream`
46 return `${stream} = [IO.File]::OpenRead(${path}); try { (New-Object Media.SoundPlayer ${stream}).PlaySync() } finally { ${stream}.Dispose() }`
47}
48
49/**
50 * Plays `file` at `volume` (1 as synthesized), preserving the cached original.
51 */
52export function playCommand(player: string, file: string, volume: number, isWindows = false): string[] {
53 const shell = (script: string) => ["sh", "-c", script, "tts-play", file]
54 const isScaled = volume !== 1
55 // aplay and SoundPlayer have no volume: play a scaled copy when ffmpeg is present.
56 const scaleFile = isScaled
57 ? `command -v ffmpeg >/dev/null && ffmpeg -loglevel error -y -i "$1" -af volume=${volume} "$1.v.wav" && set -- "$1.v.wav"; `
58 : ""
59 switch (player) {
60 case "ffplay":
61 return ["ffplay", "-nodisp", "-autoexit", "-loglevel", "error", ...(isScaled ? ["-af", `volume=${volume}`] : []), file]
62 case "paplay":
63 return ["paplay", `--volume=${Math.round(65536 * volume)}`, file]
64 case "afplay":
65 return ["afplay", "-v", String(volume), file]
66 case "aplay":
67 return shell(`${scaleFile}exec aplay -q "$1"`)
68 case "powershell":
69 if (isWindows) return powershellCommand([
70 `$file = ${quotePowerShell(file)}`,
71 ...(isScaled ? [
72 `if (Get-Command ffmpeg -CommandType Application -ErrorAction SilentlyContinue) { ffmpeg -loglevel error -y -i $file -af volume=${volume} ($file + '.v.wav'); if ($LASTEXITCODE -eq 0) { $file += '.v.wav' } }`,
73 ] : []),
74 soundPlayerScript("$file"),
75 ].join("; "))
76 // WSL: Windows reads the file through its \\wsl.localhost path.
77 return shell(
78 // `\$` keeps sh from expanding PowerShell's variable.
79 `${scaleFile}exec powershell.exe -NoProfile -NonInteractive -Command "${soundPlayerScript(`'$(wslpath -w "$1")'`, "\\$")}"`,
80 )
81 default:
82 return [player, file]
83 }
84}
85hooks/process.ts 33 lines1import type { HookStream, ProcessSpawnChunk, ProcessSpawnResult } from "claude-code"
2
3/** One synthesis or playback process; closing its stream also kills the child. */
4export class SpeechProcess {
5 private stream: HookStream<ProcessSpawnChunk, ProcessSpawnResult> | undefined
6
7 get isRunning(): boolean {
8 return this.stream !== undefined
9 }
10
11 async stop(): Promise<void> {
12 await this.stream?.return({ code: null, signal: "SIGTERM" }).catch(() => undefined)
13 }
14
15 async read(stream: HookStream<ProcessSpawnChunk, ProcessSpawnResult>) {
16 this.stream = stream
17 // Closing a stream rejects its result; observe that rejection even during cancellation.
18 const result = stream.result.catch(() => ({ code: null, signal: "SIGTERM" }))
19 let stdout = ""
20 let stderr = ""
21 try {
22 for await (const piece of stream) {
23 // Bound output so a noisy command cannot fill memory.
24 if (piece.stream === "stdout") stdout = (stdout + piece.text).slice(0, 4_194_304)
25 else stderr = (stderr + piece.text).slice(0, 4_194_304)
26 }
27 return { ...(await result), stdout, stderr }
28 } finally {
29 if (this.stream === stream) this.stream = undefined
30 }
31 }
32}
33hooks/settings.ts 69 lines1import type { PluginOptions } from "claude-code"
2
3import type { PickerOption, Setting } from "../types"
4
5import type { ReadMode, TtsConfig } from "./config"
6import { type EngineName, ENGINES, type LoadedEngine } from "./engines"
7import type { Pin, Pins } from "./pins"
8
9export const MODES: readonly ReadMode[] = ["stream", "final", "notify"]
10export const MODE_LABELS: Record<ReadMode, string> = {
11 stream: "stream 返答を一文ずつ、流れてくるそばから",
12 final: "final ターンが終わってから最終的な返答だけ",
13 notify: "notify 完了・エラーなどの短い通知だけ",
14}
15export const SPEEDS = ["0.8", "1", "1.1", "1.2", "1.5", "2"]
16export const VOLUMES = ["0.5", "0.75", "1", "1.25", "1.5"]
17
18/** A setting change ready to save: where it goes, and how it reads before and after. */
19export type Change =
20 | { label: string; previous: string; next: string; config: { key: string; value: string | number } }
21 | { label: string; previous: string; next: string; pins: (stored: Pins) => Pins }
22 | { error: string }
23
24/** The settings in effect, which a change is measured against. */
25export type Current = { options: PluginOptions; config: TtsConfig; engine: LoadedEngine; cwd: string; pin: Pin | undefined }
26
27/** What saving `value` for `setting` changes; while pinned, engine and voice go to the pin. */
28export function planChange(setting: Setting, value: string, { options, config, engine, cwd, pin }: Current): Change {
29 const pinned = (update: (base: Pin) => Pin) => (stored: Pins) => {
30 const base = stored[cwd] ?? pin
31 return base === undefined ? stored : { ...stored, [cwd]: update(base) }
32 }
33 switch (setting) {
34 case "mode":
35 if (!MODES.includes(value as ReadMode)) return { error: `mode is one of ${MODES.join(", ")}` }
36 return { label: "mode", previous: config.mode, next: value, config: { key: "tts.mode", value } }
37 case "speed":
38 case "volume": {
39 const scale = Number(value)
40 if (!(scale > 0 && scale <= 4)) return { error: `${setting} is a number above 0, up to 4` }
41 const previous = String(setting === "speed" ? options.speed : config.volume)
42 return { label: setting, previous, next: String(scale), config: { key: `tts.${setting}`, value: scale } }
43 }
44 case "engine": {
45 if (!(value in ENGINES)) return { error: `engine is one of ${Object.keys(ENGINES).join(", ")}` }
46 const name = value as EngineName
47 if (pin === undefined) return { label: "engine", previous: config.engine, next: name, config: { key: "tts.engine", value: name } }
48 return { label: "engine (this directory)", previous: config.engine, next: name, pins: pinned(base => ({ ...base, engine: name })) }
49 }
50 case "voice": {
51 const { voices } = engine
52 if (voices === undefined) return { error: `engine=${config.engine} has no voices; set it up in /config` }
53 if (pin === undefined) return { label: "voice", previous: voices.current, next: value, config: voices.select(value) }
54 const name = config.engine
55 return {
56 label: "voice (this directory)",
57 previous: voices.current,
58 next: value,
59 pins: pinned(base => ({ ...base, voices: { ...base.voices, [name]: value } })),
60 }
61 }
62 }
63}
64
65/** A choice as a text listing shows it: its id, then its name where it has one. */
66export function optionLine(option: PickerOption): string {
67 return option.label === undefined ? ` ${option.value}` : ` ${option.value}: ${option.label}`
68}
69hooks/speech-text.ts 214 lines1/** Sentence ends: Japanese/full-width punctuation with any closing brackets, or a period before whitespace. */
2const TERMINATOR = /[。!?!?]+[」』))"]*|\.(?=\s)/g
3const SPEAKABLE = /[\p{L}\p{N}]/u
4const MAX_SENTENCE = 160
5const LIST_MARKER = /^\s*(?:[-*+–—•]|\d+[.)])\s+(?:\[[ xX]\]\s+)?/
6const PATH_NAME = String.raw`[\p{L}\p{N}\p{M}_@+-][\p{L}\p{N}\p{M}_.@+-]*\.[A-Za-z][A-Za-z0-9]{0,9}`
7const PATH_LINE = String.raw`(?:(?::|#L)(\d+)(?:-L?(\d+))?(?::\d+)?)?`
8/**
9 * A file path, optionally in backticks and with a line (`:12`, `:12:5`,
10 * `:12-20`, `#L12-L20`). Folders need a separator and the name an extension
11 * that starts with a letter, so `and/or`, `1/2` and `v1.2/v1.3` never match.
12 * Folders have no spaces, so a match never runs across prose.
13 */
14const PATH = new RegExp(
15 String.raw`(?<![\p{L}\p{N}\p{M}_.:/\\~-])(\x60?)(~|[A-Za-z]:)?((?:[\p{L}\p{N}\p{M}_.@+-]*[/\\])+)(${PATH_NAME})${PATH_LINE}(\x60?)(?![\p{L}\p{N}\p{M}_/\\])`,
16 "gu",
17)
18/** Ends what came before, so a path may follow at once: Japanese punctuation and brackets. */
19const PATH_SEPARATORS = /[\s、。,:;([{〈《「『【]/
20/** May open a path when they start a word: ASCII brackets and quotes. */
21const PATH_OPENERS = /[([{<"'\x60]/
22/** A whole inline code span that is a path; only there, its folders may have spaces. */
23const CODE_PATH = new RegExp(String.raw`^(~|[A-Za-z]:)?((?:[\p{L}\p{N}\p{M}_.@+ -]*[/\\])+)(${PATH_NAME})${PATH_LINE}$`, "u")
24
25/** Where fenced code was skipped; the reader decides what, if anything, to say there. */
26export const SKIPPED_CODE: unique symbol = Symbol("skipped code")
27
28/** What a reply turns into for a listener: its sentences, and marks where code was skipped. */
29export type Spoken = string | typeof SKIPPED_CODE
30
31/**
32 * Turns streamed Markdown into speakable sentences as early as possible:
33 * complete sentences of the current line are released before its newline
34 * arrives, and fenced code and tables are never spoken.
35 */
36export class SpeechSplitter {
37 private partial = ""
38 private fence: string | undefined
39
40 push(text: string): Spoken[] {
41 this.partial += text
42 const sentences: Spoken[] = []
43
44 let newline = this.partial.indexOf("\n")
45 while (newline >= 0) {
46 sentences.push(...this.line(this.partial.slice(0, newline)))
47 this.partial = this.partial.slice(newline + 1)
48 newline = this.partial.indexOf("\n")
49 }
50
51 // Keep unfinished links and URLs whole: punctuation inside them is not a sentence end.
52 // Lines that may open a fence or a table also wait for their newline.
53 if (this.fence === undefined && !/^\s*[`~|]/.test(this.partial)) {
54 const markup = this.partial.search(/https?:\/\/|!?\[/)
55 // The period in "1. " belongs to the list marker, not a spoken sentence.
56 const start = LIST_MARKER.exec(this.partial)?.[0].length ?? 0
57 const end = markup < 0 ? this.partial.length : Math.max(start, markup)
58 const cut = lastTerminatorEnd(this.partial.slice(start, end))
59 if (cut > 0) {
60 sentences.push(...this.line(this.partial.slice(0, start + cut)))
61 this.partial = this.partial.slice(start + cut)
62 }
63 }
64 return sentences
65 }
66
67 flush(): Spoken[] {
68 const sentences = this.line(this.partial)
69 this.partial = ""
70 this.fence = undefined
71 return sentences
72 }
73
74 private line(text: string): Spoken[] {
75 const marker = /^\s*(`{3,}|~{3,})(.*)$/.exec(text)
76 if (this.fence !== undefined) {
77 if (marker && marker[1]!.startsWith(this.fence) && marker[2]!.trim() === "") {
78 this.fence = undefined
79 }
80 return []
81 }
82 if (marker) {
83 this.fence = marker[1]
84 return [SKIPPED_CODE]
85 }
86 return splitSentences(toSpeakable(text))
87 }
88}
89
90/** A whole Markdown text as speakable sentences. */
91export function toSentences(text: string): Spoken[] {
92 const splitter = new SpeechSplitter()
93 return [...splitter.push(text), ...splitter.flush()]
94}
95
96/** Strips the Markdown a listener should not hear from one line. */
97export function toSpeakable(line: string): string {
98 const text = line.trim()
99 if (text.startsWith("|") || /^[-*_]{3,}$/.test(text)) return ""
100 return text
101 .replace(/^#{1,6}\s+/, "")
102 .replace(/^>\s?/, "")
103 .replace(LIST_MARKER, "")
104 .replace(/!\[[^\]]*\]\([^)]*\)/g, "")
105 .replace(/\[([^\]]+)\]\([^)]*\)/g, "$1")
106 .replace(/https?:\/\/\S+/g, "URL")
107 .replace(/`([^`]+)`/g, speakCodePath)
108 .replace(PATH, speakPath)
109 .replace(/`([^`]*)`/g, "$1")
110 .replace(/\*\*|~~|(?<!\w)\*(?!\s)|(?<!\s)\*(?!\w)/g, "")
111 .replace(/[()[\]{}()[]{}「」『』【】〈〉《》]/g, " ")
112 .replace(/\s+/g, " ")
113 .trim()
114}
115
116/** Reads a path as its file name and line: "speech-text.tsの12行目". */
117function spokenPath(name: string, line: string | undefined, end: string | undefined): string {
118 return line === undefined ? name : end === undefined ? `${name}の${line}行目` : `${name}の${line}から${end}行目`
119}
120
121/** Whether a path may start right after `before`, not in the middle of a longer one. */
122function startsPath(before: string): boolean {
123 const last = before.at(-1)
124 if (last === undefined || PATH_SEPARATORS.test(last)) return true
125 if (!PATH_OPENERS.test(last)) return false
126 const prior = before.at(-2)
127 return prior === undefined || !/[\x21-\x7e]/.test(prior)
128}
129
130function isRootedPath(root: string | undefined, dirs: string): boolean {
131 return root !== undefined || /^\.{0,2}[/\\]/.test(dirs)
132}
133
134/**
135 * Shortens a code span that is a path with spaces in its folders, when it is
136 * rooted: `cat src/a.ts` or `My Project/src/a.ts` could be a command or a
137 * folder, so they stay whole.
138 */
139function speakCodePath(span: string, inner: string): string {
140 const match = CODE_PATH.exec(inner)
141 if (match === null) return span
142 const [, root, dirs = "", name = "", line, end] = match
143 return isRootedPath(root, dirs) ? spokenPath(name, line, end) : span
144}
145
146/**
147 * Shortens a path without spaces, only when it is clearly a path: rooted
148 * (`/`, `~/`, `./`, `C:\\`), with a line, or a whole inline code span. A
149 * bare `React/Next.js` stays. A path joined to the text before it, or an
150 * unrooted one right after a word and a space, may be the tail of a longer
151 * path (`/tmp/foo&bar/a.ts`, `My Project/src/a.ts`), so it stays whole.
152 */
153function speakPath(
154 match: string,
155 open: string,
156 root: string | undefined,
157 dirs: string,
158 name: string,
159 line: string | undefined,
160 end: string | undefined,
161 close: string,
162 offset: number,
163 text: string,
164): string {
165 const code = open !== "" && close !== ""
166 const isRooted = isRootedPath(root, dirs)
167 if (!(code || line !== undefined || isRooted)) return match
168 if (open === "") {
169 const before = text.slice(0, offset)
170 // Text joined to the match (`foo&` in `/tmp/foo&bar/a.ts`) means it is the tail of a longer path.
171 if (!startsPath(before)) return match
172 if (!isRooted && /[\p{L}\p{N}\p{M}] $/u.test(before)) return match
173 }
174 const spoken = spokenPath(name, line, end)
175 return code ? spoken : `${open}${spoken}${close}`
176}
177
178function splitSentences(text: string): string[] {
179 const sentences: string[] = []
180 let start = 0
181 for (const match of text.matchAll(TERMINATOR)) {
182 const end = match.index + match[0].length
183 sentences.push(text.slice(start, end))
184 start = end
185 }
186 sentences.push(text.slice(start))
187 return sentences
188 .map(sentence => sentence.trim())
189 .filter(sentence => SPEAKABLE.test(sentence))
190}
191
192/**
193 * Engines choke on very long input: split at a comma near the limit, else hard.
194 * Applied as a sentence is queued, so the response budget sees whole sentences.
195 */
196export function capLength(sentence: string): string[] {
197 const pieces: string[] = []
198 let rest = Array.from(sentence)
199 while (rest.length > MAX_SENTENCE) {
200 const comma = Math.max(rest.lastIndexOf("、", MAX_SENTENCE - 1), rest.lastIndexOf(",", MAX_SENTENCE - 1))
201 const cut = comma > MAX_SENTENCE / 2 ? comma + 1 : MAX_SENTENCE
202 pieces.push(rest.slice(0, cut).join(""))
203 rest = rest.slice(cut)
204 }
205 pieces.push(rest.join(""))
206 return pieces
207}
208
209function lastTerminatorEnd(text: string): number {
210 let end = 0
211 for (const match of text.matchAll(TERMINATOR)) end = match.index + match[0].length
212 return end
213}
214hooks/engines/command.ts 21 lines1import { findProgramCommand } from "../platform"
2import type { Engine } from "./engine"
3
4/** Any command: argv split on spaces, `{out}` the wav to write, the text on stdin. */
5export const command: Engine = {
6 load: (options, isWindows = false) => {
7 const template = String(options.command)
8 const program = template.split(/\s+/).find(Boolean)
9 return {
10 voice: template || "(unset)",
11 location: template || "(unset)",
12 probe: program === undefined ? undefined : findProgramCommand(program, isWindows),
13 synthesis: (text, out) => {
14 const argv = template.split(/\s+/).filter(Boolean).map(arg => arg.replaceAll("{out}", out))
15 if (argv.length === 0) throw new Error("command is not set (/config → tts)")
16 return [{ argv, stdin: text }]
17 },
18 }
19 },
20}
21