SLOPSHOPPER

tts

Reads Claude's replies aloud as they stream, sentence by sentence, through VOICEVOX, Irodori-TTS or any command.

newpaneguardcommandtoaststatus
v0.4.0MITupdated 2026-10-10kajidog/cc-mods-tts
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · tts
│ ┃ tts-picker ✕ › fix the failing auth╭────────────────────────────────────────────╮ │ ┃ Nothing to pick. │ tts │ │ ⏺ Read(src/auth.ts) │ tts: no audio player found (ffplay, │ │ ⎿ Read 6 lines │ paplay, aplay, afplay, powershell.exe) │ │ ⏺ Update(src/auth.ts) ╰────────────────────────────────────────────╯ │ ⎿ Added 2 lines, removed 1 line │ ⏺ Bash(bun test) │ ⎿ 3 pass, 1 fail │ │ ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ │ │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · tts-picker
Nothing to pick.
README

cc-mods-tts

Claude Code の mod(関数フックのプラグイン)。Claude の返答をストリーミング中に一文ずつ読み上げます。

  • 対応プラットフォーム: macOS / Linux / Windows(ネイティブ・WSL)
  • turn.step の text チャンクを文に分割し、合成と再生をパイプラインで流す(コードブロック・表・URL は読み上げから除外してます)
  • エンジン: VOICEVOX 互換 / Irodori-TTS / 任意コマンド

読み上げ用のテキストでは、箇条書きの印を除き、括弧を空白に置き換えて中身を読みます。明確なファイルパスはファイル名と行番号に短縮します。負数・小数・C++・--help など、意味を持つ記号は残します。

セットアップ

Claude Code 2.1.287 以降(Mods が既定で有効なバージョン)。どちらか一方で使えるようにできます。

A. GitHub からインストール(使うだけなら)

ターミナルの Claude Code のプロンプトで:

/plugin install tts --marketplace kajidog/cc-mods-tts

「Add marketplace?」に y、スコープは user(既定)のまま Enter。そのセッションからすぐ有効になり、以降のセッションでも読み込まれます。更新は claude plugin update tts のあと /reload-plugins。

デスクトップアプリの Code タブではこのコマンドは使えません。ターミナルから user スコープで入れれば、デスクトップアプリのローカルセッションでも読み込まれます。

B. クローンして読み込む(手元で改造するなら)

git clone https://github.com/kajidog/cc-mods-tts.git ~/cc-mods-tts
claude --plugin-dir ~/cc-mods-tts

常に読み込むなら ~/.claude/settings.json の env に "CLAUDE_CODE_PLUGIN_DIRS": "~/cc-mods-tts"。この方法ではクローンしたフォルダのコードがそのまま動き、編集するとセッション中に読み込み直されます。

読み込んだら

  1. VOICEVOX を起動しておく(インストール時の既定エンジン)
  2. /tts engine で voicevox が ● になっているか確認
  3. 声を変えたい場合は /tts voice で選ぶ(既定は「ずんだもん・ノーマル」)
  4. /tts test で鳴るか確認

Windows で使う場合

Windows ネイティブの Claude Code で使えます。VOICEVOX も同じ Windows 上で起動すれば、既定の http://localhost:50021 のまま接続できます。

player: auto は ffplay があればそれを使い、なければ Windows 標準の PowerShell で再生します。この Mod のために WSL や sh・mktemp を追加する必要はありません。PowerShell 再生時の音量変更と、音声の前後の無音除去には ffmpeg が必要です。ない場合は元の音声を再生します。

WSL 内のターミナルで動かす場合は Linux と同じ構成で、powershell.exe による Windows 側の再生も使えます。

Irodori-TTS を使う場合

  1. Irodori-TTS-Server を起動する
  2. 参照音声(例: sample.wav)をサーバー側の voices/ に置く。拡張子を除いたファイル名(この例では sample)が声の ID になります
  3. /tts engine irodori で切り替え、/tts voice で追加した声を選ぶ(/tts voice sample で直接指定もできます)
  4. /tts test で鳴るか確認

voices/ は Irodori-TTS-Server 側のフォルダです。IRODORI_VOICES_DIR を設定している場合は、そのフォルダに置いてください。Irodori は声を選ぶまで読み上げができません。

対応する音声形式やアップロード API など、詳しくは公式の音声追加・管理手順(Voice Management)を参照してください。

アンインストール

A で入れた場合:

claude plugin uninstall tts@cc-mods-tts
claude plugin marketplace remove cc-mods-tts   # マーケットプレイスも消すなら

起動中のセッションには /reload-plugins で反映されます。

B の場合は --plugin-dir を付けずに起動するか、CLAUDE_CODE_PLUGIN_DIRS を settings.json から消して、クローンしたフォルダを削除します。

どちらの場合も、/config で変えた設定は ~/.claude/settings.json の pluginConfigs に残ります(A は tts@cc-mods-tts、B は tts@inline のキー)。不要なら手で消してください。

使い方

コマンド内容
/tts状態表示(on/off・モード・エンジン・速度・音量)
/tts helpコマンド一覧
/tts config設定の一覧と今の値
/tts on / /tts off有効・無効(/config の Enabled と同じ)
/tts stop合成・再生中の処理とキューを止め、そのターンの残りの返答も読まない。次のターンから再開
/tts test [text]通常の返答と同じ整形を通してテスト読み上げ
`/tts mode [stream\final\notify]`読み上げモード切替。引数なしでモードを選ぶペインを開く
/tts engine [name]エンジン切替。引数なしで各エンジンが応答するか(ok / down / unset)を表示
/tts voice [id]引数なしで声を選ぶペインを開く(↑↓で移動、Enter で決定、Esc で閉じる)。id 指定で直接設定
/tts speed [n] / /tts volume [n]速度・音量(1 が等倍)。引数なしで候補を選ぶペインを開く

声の候補が多い場合は64件ずつ表示します。p / n(Previous / Next)でページを切り替え、↑↓と Enter で選択できます。最初は現在の声があるページを開きます。

読み上げモード

モード読むもの
stream(既定)返答をストリーミング中に一文ずつ
finalターン終了時に最終的な返答だけ(途中のツール前の文は読まない)
notify短い通知だけ:「完了しました」「エラーで止まりました」など

どのモードでも、以下の操作待ちのダイアログが出たら読み上げを止めて通知します:

ダイアログ読む文(既定。/config の TTS notice で変えられる)
質問(AskUserQuestion)1件: 質問があります。+質問文 / 複数: 質問が3件あります。
プランの承認(ExitPlanMode)プランの確認をお願いします。
Bash の許可コマンドの確認をお願いします。+コマンドの説明(description)
その他のツールの許可ツールの許可をお願いします。

許可は classic.PermissionRequest(ダイアログを出すときに発火)で拾うので、ルールや auto モードで自動判定された呼び出しでは鳴らない。通知の音声は初回だけ合成し、以降は使い回す。

停止すると Mod 側の合成リクエストも中断します。ただし、TTS サーバーが切断後も計算を続ける場合、その計算の終了までは次の合成が待たされることがあります。

ffplay / powershell では、再生終了時に語尾が途切れるのを抑えるため、無音除去後の WAV に0.2秒の無音を足します。この処理には ffmpeg が必要です。ffmpeg がない場合は、無音除去・追加を行わず元の音声を再生します。

WSL で ffplay を使うと、音声が WSLg(PulseAudio → RDP)を経由するため、負荷によって「ジリジリ」とノイズが混じることがあります。気になる場合は /config の TTS: Player を powershell にすると、Windows 側で直接再生するので改善します(文ごとに PowerShell を起動するため、文と文の間は少し空きます)。

設定

設定はすべて /config の tts にまとめてます。項目名はどれも TTS: / TTS notice: で始まります。

項目(キー)内容既定
TTS: Enabled(enabled)読み上げる / 読み上げないtrue
TTS: Read mode(mode)stream / final / notify(上記)stream
TTS: Engine(engine)voicevox / irodori / commandvoicevox
TTS: Fallback engine(fallback)メインのエンジンが失敗した文(サーバー停止・未設定など)を読むエンジン。メインが使えない間は、ステータス行に ⚠ tts: voicevox down のように出るnone
TTS: Speed(speed)速度。エンジン側で合成時にかける(Irodori は 0.25〜4、VOICEVOX は speedScale)。command では無視1.1
TTS: Volume(volume)音量。再生時にかけるのでどのエンジンでも効く。ffplay / paplay / afplay は再生コマンドで、aplay / powershell は ffmpeg があれば再生前に変える1
TTS: Player(player)再生コマンド。auto は Windows ネイティブでは ffplay → powershell.exe、それ以外では ffplay → paplay → aplay → afplay → powershell.exe の順に最初に見つかったものauto
TTS: Max characters per response(maxCharsPerTurn)1回の応答で読む最大文字数。超えたら「途中を省略します。」と言って残りを飛ばし、応答の最後の1文だけ読む。0 で無制限600
TTS: Stop on new prompt(interruptOnPrompt)プロンプト送信で読み上げを止めるtrue
TTS: Prompt for the model(modelPrompt)読み上げ中(stream / final モード)に Claude のシステムプロンプトへ足す文。説明・進捗報告は原則日本語にし、コード・引用・ユーザーの言語指定は保持する。空にすると何も足さない応答は音声で読み上げられています。説明やツール呼び出しの合間の進捗報告は原則として日本語で書いてください。ただし、コード・識別子・引用はそのまま保持し、ユーザーが指定した言語を優先してください。
TTS: Irodori URL(irodoriUrl)Irodori-TTS の URLhttp://localhost:8088
TTS: Irodori voice(irodoriVoice)Irodori の声(/tts voice で選べる)(未設定)
TTS: VOICEVOX URL(voicevoxUrl)VOICEVOX 互換エンジンの URL(AivisSpeech は http://localhost:10101)http://localhost:50021
TTS: VOICEVOX speaker(voicevoxSpeaker)VOICEVOX のスタイル id(/tts voice で選べる)3
TTS: Command(command)engine=command のときのコマンド。空白区切り、{out} が書き出す wav のパスに置き換わり、本文は標準入力で渡る(未設定)
TTS notice: done / error / refusal(noticeDone / noticeError / noticeRefusal)notify モードでターン終了時に言う文完了しました。/ エラーで止まりました。/ 応答できませんでした。
TTS notice: question(noticeQuestion)質問が1件出たとき、質問文の前に言う文質問があります。
TTS notice: questions(noticeQuestions)質問が複数出たときに言う文。{n} が件数になる質問が{n}件あります。
TTS notice: plan(noticePlan)プランの承認待ちで言う文プランの確認をお願いします。
TTS notice: command(noticeCommand)Bash の許可ダイアログで、コマンドの説明の前に言う文コマンドの確認をお願いします。
TTS notice: tool(noticeTool)その他のツールの許可ダイアログで言う文ツールの許可をお願いします。
TTS notice: cut(noticeCut)文字数の上限で残りを飛ばすときに言う文。この後に応答の最後の1文を読む途中を省略します。
TTS notice: code(noticeCode)コードブロックを飛ばすところで言う文(stream / final モード)コードは省略します。

通知の文はどれも空にするとその通知の読み上げをスキップできます(注:スキップしても既存の再生は止まります)。 質問やコマンドの通知を空にした場合は、質問文・コマンドの説明も読みません。

Irodori の速度の対応範囲は 0.25〜4 です。Mod ではこの下限による入力エラーは出しませんが、範囲外の値は Irodori サーバーが拒否する場合があります。文字数の上限に、省略やコードを知らせる通知文と、省略後に読む最後の1文は含みません。上限をまたぐ文は途中まで読まずに飛ばします。最後の1文が300文字を超える場合は読みません。

音声ファイルはセッション用の一時ディレクトリに保存し、再生後に削除します。macOS / Linux / WSL では所有者のみアクセスできる権限で作成し、Windows ネイティブではユーザーの一時フォルダ(%TEMP%)のアクセス権を引き継ぎます。通知キャッシュはセッション終了時・Mod の再読み込み時に削除します。プロセスの強制終了などで清掃が実行できなかった場合は、一時ディレクトリに残ることがあります。

前提

  • curl(Windows では curl.exe)
  • mktemp(macOS / Linux / WSL の一時ディレクトリ作成に使用)
  • Windows ネイティブ: Windows PowerShell(powershell.exe、一時ファイルの管理と再生に使用)
  • 再生: ffplay / paplay / aplay / afplay / powershell.exe(Windows・WSL)のいずれか(player: auto で自動選択)

エンジンを追加する

エンジンは hooks/engines/ に1ファイルずつある(irodori.ts / voicevox.ts / command.ts)。追加するには:

  1. hooks/engines/<name>.ts に Engine(engine.ts の型)を書く。load(options, isWindows) で自分の設定を読み、synthesis(text, out) で wav を out に書くコマンド列を返す。声の一覧があれば voices も
  2. hooks/engines/index.ts の ENGINES に足す
  3. .claude-plugin/plugin.json の userConfig.engine.options に名前を足し、そのエンジンの設定項目を追加する

開発

npm install
CLAUDE_CODE_PLUGIN_DIR_WATCH=1 claude --init-only --plugin-dir .  # .claude-plugin/types/ を生成(型チェックに必要)
npm run check  # lint・validate・型チェック・test をまとめて実行

個別には npm run lint / npm run validate / npm run typecheck / npm test で実行できます。Biome と TypeScript は devDependencies で固定し、claude は PATH にあるものを使います。

CI(.github/workflows/ci.yml)は main への push と PR で lint を実行し、Linux・Windows で validate・型チェック・test を実行します。型チェックの前に、上記の初期化コマンドで .claude-plugin/types/ を生成します。対話セッションや API キーは不要です。

API は early access のため、Claude Code の更新で変わることがある(CI で確認しているバージョン: 2.1.296)。

変更履歴とリリース

変更履歴は CHANGELOG.md にまとめています。Changesets で変更メモを集め、GitHub Actions がリリース用 PR を作成・更新します。

開発時は Node.js 22.11 以降の 22 系、24 系、または 26 以降と、npm 10.9 以降を使います。

npm ci
npm run changeset

patch(修正)・minor(機能追加)・major(互換性のない変更)を選び、変更内容を日本語で入力します。生成された .changeset/*.md をコードと一緒にコミットしてください。ドキュメントや CI だけの変更ではメモは不要です。

  1. 変更メモを含む PR を main にマージする。
  2. Actions が chore: release tts という PR を作成・更新する。CHANGELOG への追記と、package.json・package-lock.json・.claude-plugin/plugin.json のバージョン更新がまとまります。
  3. 内容を確認してリリース用 PR をマージする。利用者は claude plugin update tts と /reload-plugins で更新できます。

最初に GitHub リポジトリの Settings → Actions → General → Workflow permissions で Allow GitHub Actions to create and approve pull requests を有効にしてください(Changesets の設定手順)。標準の GITHUB_TOKEN を使うため、追加のトークンは不要です。自動作成された PR の CI が承認待ちになった場合は、PR 画面の Approve workflows to run で実行を承認してください(GitHub の仕様)。

手元でリリース差分を生成する場合は npm run release:version、バージョンの整合性確認は npm run check:version を使います。履歴の生成時はプラグイン側も同期するため、changeset version 単体では実行しません。バージョン管理用の package.json は private: true で、npm 公開や GitHub Release の作成は行いません。

ライセンス

MIT License

Source 16 files
hooks/register.tsx 621 lines
1import { atom, read, update } from "claude-code"
2import type { EngineInterface, Register } from "claude-code"
3
4import type { Picker, Setting } from "../types"
5
6import { Channel } from "./channel"
7import { type Notices, type ReadMode, type TtsConfig, toConfig } from "./config"
8import { DIALOG_TOOLS, dialogSpeech } from "./dialog-text"
9import { type EngineName, ENGINES, type LoadedEngine } from "./engines"
10import { type Pin, PINS_KEY, type Pins, pinnedOptions, toPins } from "./pins"
11import { removeDirectoryCommand, removeFilesCommand, tempDirectoryCommand } from "./platform"
12import { findPlayerCommand, playCommand, prepareAudioCommand } from "./player"
13import { SpeechProcess } from "./process"
14import { type Change, MODE_LABELS, MODES, optionLine, planChange, SPEEDS, VOLUMES } from "./settings"
15import { capLength, SKIPPED_CODE, type Spoken, SpeechSplitter, toSentences } from "./speech-text"
16
17const HELP = [
18  "/tts                   state: on/off, mode, engine, speed, volume",
19  "/tts on | off          read replies aloud or not",
20  "/tts stop              stop speech and mute the rest of this turn",
21  "/tts test [text]       read a test sentence",
22  "/tts mode [name]       stream | final | notify; no name opens a picker",
23  "/tts engine [name]     switch engines; no name checks whether each one answers",
24  "/tts voice [id]        pick a voice; no id opens a picker",
25  "/tts speed [n]         speech speed (1: as is, Irodori: 0.25–4); no value opens a picker",
26  "/tts volume [n]        playback volume (1: as is); no value opens a picker",
27  "/tts pin | unpin       keep this directory's engine and voice apart from the rest, or stop",
28  "/tts config            every setting and its value; change them in /config → tts",
29  "/tts help              this list",
30].join("\n")
31/** Which notice notify mode says as a turn ends; an interrupted turn stays silent. */
32const TURN_NOTICES: Partial<Record<string, keyof Notices>> = { answer: "done", error: "error", refusal: "refusal" }
33const PICKER_PANE = "tts-picker"
34// Claude Code refuses a Select with more than 64 options.
35const PICKER_PAGE_SIZE = 64
36const picker = atom({ plugin: "tts", key: "picker" } as const, null as Picker | null)
37// Survives a reload so the next module can remove the old module's audio files.
38const audioDirectory = atom({ plugin: "tts", key: "audioDirectory" } as const, null as string | null)
39
40/** Irodori runs diffusion with a concurrency of 1, so one request can take minutes. */
41const SYNTH_TIMEOUT_MS = 300_000
42/** A last sentence longer than this is not said after the cut: a long line without a stop is no question. */
43const MAX_TAIL = 300
44
45/**
46 * How much of one model response has been queued; past the limit the rest is dropped,
47 * save its last sentence (`tail`), which `finishResponse` says: it often asks what to do next.
48 */
49type Budget = { spoken: number; isCut: boolean; tail?: string }
50
51/** `cacheAs` keeps the synthesized clip at that path for the next time. */
52type Sentence = { text: string; generation: number; cacheAs?: string }
53type Clip = { file: string; generation: number; isCached?: boolean }
54
55export const register: Register = (on, options) => {
56  let isWindows = false
57  let config = toConfig(options)
58  let engine = ENGINES[config.engine].load(options)
59  let fallback = config.fallback === undefined ? undefined : ENGINES[config.fallback].load(options)
60  // The working directory's pin, if any, lays its engine and voice over the options.
61  let cwd = ""
62  let pins: Pins = {}
63
64  // `generation` moves on at every stop: queued work of an older one is dropped.
65  let generation = 0
66  let isTurnMuted = false
67  const { isEnabled } = config
68  let player: string | undefined
69  const synthesisProcess = new SpeechProcess()
70  const playbackProcess = new SpeechProcess()
71  const sentences = new Channel<Sentence>()
72  const clips = new Channel<Clip>()
73  // Notices are few and fixed: each is synthesized once per load, then replayed at once.
74  const cachedNotices = new Map<string, string>()
75  let directory = ""
76  let noticeCount = 0
77  // Set while the engine in use fails ("voicevox down"); shown beside playback until it reads again.
78  let engineIssue: string | undefined
79  // Set by session.start: probes the engine in use and shows the result.
80  let recheckEngine = () => {}
81
82  /** Rebuilds the engines from the options with this directory's pin laid over them. */
83  const applyPins = (stored: Pins) => {
84    pins = stored
85    const effective = pinnedOptions(options, pins[cwd])
86    config = toConfig(effective)
87    engine = ENGINES[config.engine].load(effective, isWindows)
88    fallback = config.fallback === undefined ? undefined : ENGINES[config.fallback].load(effective, isWindows)
89    // Cached notices were spoken in the voice before.
90    cachedNotices.clear()
91    recheckEngine()
92  }
93
94  /** What saving `value` for `setting` changes, against the settings in effect here. */
95  const changeOf = (setting: Setting, value: string) => planChange(setting, value, { options, config, engine, cwd, pin: pins[cwd] })
96
97  /** Invalidates queued speech and closes the active synthesis and playback streams. */
98  const stop = async () => {
99    generation += 1
100    await Promise.all([synthesisProcess.stop(), playbackProcess.stop()])
101  }
102
103  /** Queues `text` to be said next; called right after a stop, it plays first. */
104  const queueNotice = (text: string) => {
105    if (text === "") return
106    const cached = cachedNotices.get(text)
107    if (cached !== undefined) return clips.push({ file: cached, generation, isCached: true })
108    sentences.push({ text, generation, cacheAs: `${directory}/notice-${noticeCount++}.wav` })
109  }
110
111  /** Queues `text` for synthesis, in pieces an engine takes. */
112  const say = (text: string) => {
113    if (text === "") return
114    for (const piece of capLength(text)) sentences.push({ text: piece, generation })
115  }
116
117  /**
118   * Queues sentences; with a budget, stops before the sentence that would pass the limit
119   * (never reading half of one) and says once that the rest is skipped.
120   */
121  const enqueue = (spoken: readonly Spoken[], budget?: Budget) => {
122    for (const text of spoken) {
123      // The code notice is the mod's, not the response's: it uses no budget and is never the last sentence.
124      if (text === SKIPPED_CODE) {
125        if (budget?.isCut !== true) say(config.notices.code)
126        continue
127      }
128      if (budget !== undefined && config.maxCharsPerTurn > 0) {
129        const length = Array.from(text).length
130        if (!budget.isCut && budget.spoken + length > config.maxCharsPerTurn) {
131          budget.isCut = true
132          say(config.notices.cut)
133        }
134        // Each skipped sentence may be the last, so the one that passes the limit counts too.
135        if (budget.isCut) {
136          budget.tail = text
137          continue
138        }
139        budget.spoken += length
140      }
141      say(text)
142    }
143  }
144
145  /** Ends a response: past the limit, says its last sentence after the cut notice. */
146  const finishResponse = (budget: Budget) => {
147    if (budget.tail !== undefined && Array.from(budget.tail).length <= MAX_TAIL) say(budget.tail)
148    budget.tail = undefined
149  }
150
151  on("session.start", async ($, e, next) => {
152    const started = await next(e)
153    isWindows = (await $.env.get("OS")) === "Windows_NT"
154    cwd = await $.session.cwd()
155    applyPins(toPins(await $.store.get(PINS_KEY)))
156    player = config.player === "auto" ? (await $.process.run(findPlayerCommand(isWindows))).stdout.trim() || undefined : config.player
157    if (player === undefined) $.ui.toast("tts: no audio player found (ffplay, paplay, aplay, afplay, powershell.exe)")
158
159    directory = await prepareDirectory($, isWindows)
160    let clipCount = 0
161    let lastErrorAt = 0
162
163    const noteEngine = (error?: unknown) => {
164      const issue = error === undefined ? undefined : `${config.engine} ${/is not set/.test(String(error)) ? "unset" : "down"}`
165      showEngineIssue(issue)
166    }
167    // The status line shows only a failing engine; Claude Code draws it as "⚠ tts: voicevox down".
168    const showEngineIssue = (issue: string | undefined) => {
169      if (issue === engineIssue) return
170      engineIssue = issue
171      $.ui.status(issue)
172    }
173    // The last module's line stays after a reload.
174    $.ui.status(undefined)
175
176    // The fallback reads on quietly: the status line says the engine is failing.
177    const reportFallback = (error: unknown) => {
178      $.ui.log(`tts: ${config.engine} failed, read with ${config.fallback}: ${String(error)}`, { to: "debug" })
179    }
180
181    const reportError = (error: unknown) => {
182      const message = error instanceof Error ? error.message : String(error)
183      $.ui.log(`tts: ${message}`, { to: "debug" })
184      // One toast per burst: a server that is down fails every sentence of a reply.
185      if (Date.now() - lastErrorAt > 30_000) $.ui.toast(`tts: ${message}`)
186      lastErrorAt = Date.now()
187    }
188
189    // Synthesis runs ahead of playback, one request at a time.
190    void (async () => {
191      for (;;) {
192        const sentence = await sentences.take()
193        if (sentence.generation !== generation) continue
194        const file = sentence.cacheAs ?? `${directory}/${String(clipCount++).padStart(6, "0")}.wav`
195        const isCurrent = () => sentence.generation === generation
196        let isFallback = false
197        try {
198          try {
199            await synthesize($, synthesisProcess, engine, sentence.text, file, isCurrent)
200            if (isCurrent()) noteEngine()
201          } catch (error) {
202            if (isCurrent()) noteEngine(error)
203            if (!isCurrent() || fallback === undefined || config.fallback === undefined) throw error
204            isFallback = true
205            await synthesize($, synthesisProcess, fallback, sentence.text, file, isCurrent)
206            reportFallback(error)
207          }
208          if (!isCurrent()) {
209            await $.process.run(removeFilesCommand([file], isWindows))
210            continue
211          }
212          // Done while the clip before plays; an untrimmed clip still plays.
213          await runProcess($, synthesisProcess, prepareAudioCommand(file, player, isWindows), undefined, 30_000).catch(() => undefined)
214          if (!isCurrent()) {
215            await $.process.run(removeFilesCommand([file, `${file}.t.wav`], isWindows))
216            continue
217          }
218          // A notice the fallback read is kept for this time only: the engine may be back next time.
219          const isCached = sentence.cacheAs !== undefined && !isFallback
220          if (isCached) cachedNotices.set(sentence.text, file)
221          clips.push({ file, generation: sentence.generation, isCached })
222        } catch (error) {
223          await $.process.run(removeFilesCommand([file], isWindows))
224          if (isCurrent()) reportError(error)
225        }
226      }
227    })()
228
229    void (async () => {
230      for (;;) {
231        const clip = await clips.take()
232        if (clip.generation === generation && player !== undefined) {
233          try {
234            await playbackProcess.read($.process.spawn({ argv: playCommand(player, clip.file, config.volume, isWindows) }))
235          } catch (error) {
236            if (clip.generation === generation) reportError(error)
237          }
238        }
239        const files = [`${clip.file}.v.wav`]
240        if (clip.isCached !== true) files.push(clip.file)
241        await $.process.run(removeFilesCommand(files, isWindows))
242      }
243    })()
244
245    // Checked at once, so a start or an engine switch shows a dead engine before anything is read.
246    recheckEngine = () => {
247      showEngineIssue(undefined)
248      if (!isEnabled) return
249      const checked = engine
250      void checkEngine($, checked).then(state => {
251        if (checked === engine && state !== "ok") showEngineIssue(`${config.engine} ${state}`)
252      })
253    }
254    recheckEngine()
255
256    await $.command.register({
257      name: "tts",
258      description: "Read replies aloud: on, off, stop, test, mode, engine, voice, speed, volume, pin, unpin, config, help",
259      argumentHint: "[on|off|stop|test [text]|mode|engine|voice|speed|volume [value]|pin|unpin|config|help]",
260      // `/tts stop` has to land while a reply is still being read.
261      immediate: true,
262    })
263    return started
264  })
265
266  on("session.end", async ($, e, next) => {
267    await stop()
268    cachedNotices.clear()
269    if (directory !== "") await $.process.run(removeDirectoryCommand(directory, isWindows))
270    directory = ""
271    await update($, audioDirectory, () => null)
272    return next(e)
273  }).catch(($, e, next) => next(e))
274
275  // /clear, resume and /branch keep this module alive without another session.start, and reset
276  // $.state, which holds the directory for the next reload to remove.
277  on("classic.SessionStart", { source: ["clear", "resume", "fork"] }, async ($, e, next) => {
278    if (directory === "") directory = await prepareDirectory($, isWindows)
279    else await update($, audioDirectory, () => directory)
280    return next(e)
281  }).catch(($, e, next) => next(e))
282
283  on("command.run", { command: "tts" }, async ($, e) => {
284    const [action = "", ...rest] = e.args.trim().split(/\s+/)
285    switch (action) {
286      case "on":
287      case "off": {
288        if (action === "off") await stop()
289        const set = await $.config.set({ key: "tts.enabled", value: action === "on" })
290        return { text: "deny" in set && set.deny ? set.deny : action }
291      }
292      case "stop": {
293        isTurnMuted = true
294        await stop()
295        return { text: "stopped" }
296      }
297      case "test": {
298        const text = e.args.trim().replace(/^test(?:\s+|$)/, "") || "読み上げのテストです。"
299        const texts = toSentences(text)
300        if (texts.length === 0) return { text: "no text to read after formatting" }
301        try {
302          engine.synthesis(text, "probe.wav")
303        } catch (error) {
304          if (fallback === undefined) return { text: error instanceof Error ? error.message : String(error) }
305        }
306        enqueue(texts)
307        return { text: `testing ${config.engine} (the first Irodori request loads the model and can take a while)` }
308      }
309      case "voice": {
310        const { voices } = engine
311        if (voices === undefined) return { text: `engine=${config.engine} has no voices; set it up in /config` }
312        const voice = rest.join(" ")
313        if (voice === "") {
314          const listed = await $.http.fetch(voices.url)
315          if (!listed.ok) return { text: `GET ${voices.url} failed with ${listed.status}` }
316          const options = voices.parse(listed.text)
317          const title = `${config.engine} voice`
318          if (await openPicker($, { setting: "voice", title, options, value: voices.current })) {
319            return { text: "pick a voice in the pane (Esc closes it)" }
320          }
321          return { text: [`${config.engine} voices (/tts voice <id>):`, ...options.map(optionLine)].join("\n") }
322        }
323        return { text: await commit($, changeOf("voice", voice), applyPins) }
324      }
325      case "mode": {
326        const name = rest[0] as ReadMode | undefined
327        if (name === undefined) {
328          const options = MODES.map(mode => ({ value: mode, label: MODE_LABELS[mode] }))
329          if (await openPicker($, { setting: "mode", title: "read mode", options, value: config.mode })) {
330            return { text: "pick a mode in the pane (Esc closes it)" }
331          }
332          return { text: `mode: ${config.mode} (${MODES.join(" | ")})` }
333        }
334        return { text: await commit($, changeOf("mode", name), applyPins) }
335      }
336      case "speed":
337      case "volume": {
338        const value = rest[0]
339        if (value === undefined) {
340          const current = String(action === "speed" ? options.speed : config.volume)
341          const choices = (action === "speed" ? SPEEDS : VOLUMES).map(each => ({ value: each, label: `× ${each}` }))
342          if (await openPicker($, { setting: action, title: action, options: choices, value: current })) {
343            return { text: `pick a ${action} in the pane (Esc closes it)` }
344          }
345          return { text: `${action}: ${current}` }
346        }
347        return { text: await commit($, changeOf(action, value), applyPins) }
348      }
349      case "pin": {
350        const pin: Pin = { engine: config.engine, voices: engine.voices ? { [config.engine]: engine.voices.current } : {} }
351        const stored = { ...toPins(await $.store.get(PINS_KEY)), [cwd]: pin }
352        await $.store.set(PINS_KEY, stored)
353        applyPins(stored)
354        return {
355          text: `pinned ${cwd}: ${config.engine} (${engine.voice})\nengine and voice changes here stay in this directory; /tts unpin undoes it`,
356        }
357      }
358      case "unpin": {
359        if (pins[cwd] === undefined) return { text: "this directory is not pinned" }
360        const before = `${config.engine} (${engine.voice})`
361        const { [cwd]: _unpinned, ...stored } = toPins(await $.store.get(PINS_KEY))
362        await $.store.set(PINS_KEY, stored)
363        applyPins(stored)
364        return { text: `unpinned ${cwd}: ${before} → ${config.engine} (${engine.voice})` }
365      }
366      case "engine": {
367        const name = rest[0]
368        if (name === undefined) {
369          const names = Object.keys(ENGINES) as EngineName[]
370          const rows = await Promise.all(names.map(each => engineRow($, each, ENGINES[each].load(options, isWindows), config)))
371          return { text: ["engines (● answers, ○ down or unset; /tts engine <name> switches):", ...rows].join("\n") }
372        }
373        return { text: await commit($, changeOf("engine", name), applyPins) }
374      }
375      case "help":
376        return { text: HELP }
377      case "config": {
378        const width = Math.max(...Object.keys(options).map(key => key.length))
379        const rows = Object.entries(options).map(([key, value]) => `  ${key.padEnd(width)}  ${JSON.stringify(value)}`)
380        const pinned = pins[cwd] === undefined ? [] : [`engine and voice are pinned here: ${config.engine} (${engine.voice})`]
381        return { text: ["settings (change them in /config → tts):", ...rows, ...pinned].join("\n") }
382      }
383      case "": {
384        const marks = await Promise.all([answerMark($, engine), fallback && answerMark($, fallback)])
385        return {
386          text: [
387            `${isEnabled ? "on" : "off"} (mode: ${config.mode})`,
388            `engine: ${marks[0]} ${config.engine} (${engine.voice})${config.fallback ? `, fallback: ${marks[1]} ${config.fallback}` : ""}`,
389            `speed: ${options.speed}, volume: ${config.volume}`,
390            ...(pins[cwd] === undefined ? [] : [`engine and voice pinned to ${cwd} (/tts unpin)`]),
391            `player: ${player ?? "none"}`,
392            "/tts help lists the commands",
393          ].join("\n"),
394        }
395      }
396      default:
397        return { text: `unknown: ${action}\n\n${HELP}` }
398    }
399  })
400
401  // The picker a `/tts` command without a value opens: a pick saves the setting and closes the pane.
402  on("ui.render", { component: "Pane", requestId: PICKER_PANE }, async ($, e) => {
403    const shown = await read($, picker)
404    if (e.surface === "mobile" || shown === null || shown.options.length === 0) {
405      const { Text } = $.ui.resolve(e)
406      return <Text dimColor>{shown === null || shown.options.length === 0 ? "Nothing to pick." : shown.options.map(optionLine).join("\n")}</Text>
407    }
408    const { Box, Text, Select, Button } = $.ui.resolve(e)
409    const page = shown.page ?? Math.floor(Math.max(0, shown.options.findIndex(option => option.value === shown.value)) / PICKER_PAGE_SIZE)
410    const pages = Math.ceil(shown.options.length / PICKER_PAGE_SIZE)
411    const choices = shown.options.slice(page * PICKER_PAGE_SIZE, (page + 1) * PICKER_PAGE_SIZE)
412    const turnPage = (page: number) => update($, picker, () => ({ ...shown, page }))
413    return (
414      <Box flexDirection="column">
415        <Text dimColor>{`${shown.title}: ↑↓ move, Enter picks, Esc closes`}</Text>
416        {pages > 1 && (
417          <Box>
418            {page > 0 && <Button key="previous-page" hotkey="p" onPress={() => turnPage(page - 1)}>Previous</Button>}
419            <Text dimColor>{` Page ${page + 1}/${pages} `}</Text>
420            {page + 1 < pages && <Button key="next-page" hotkey="n" onPress={() => turnPage(page + 1)}>Next</Button>}
421          </Box>
422        )}
423        <Select
424          key={`${shown.setting}-${page}`}
425          options={choices.map(option => ({ value: option.value, label: option.label ?? option.value }))}
426          value={choices.some(option => option.value === shown.value) ? shown.value : undefined}
427          autoFocus
428          onSelect={async value => {
429            await $.ui.close({ id: PICKER_PANE })
430            $.ui.toast(`tts: ${await commit($, changeOf(shown.setting, value), applyPins)}`)
431          }}
432        />
433      </Box>
434    )
435  })
436
437  on("prompt.submit", async ($, e, next) => {
438    if (config.interruptOnPrompt) await stop()
439    return next(e)
440  }).catch(($, e, next) => next(e))
441
442  // Asks the model to write what reads well aloud; notify mode reads none of its text.
443  on("prompt.compose", async ($, e, next) => {
444    const composed = await next(e)
445    if (!isEnabled || config.mode === "notify" || config.modelPrompt === "") return composed
446    return { sections: [...composed.sections, { id: "tts:spoken", text: config.modelPrompt, scope: "session" as const }] }
447  }).catch(($, e, next) => next(e))
448
449  on("turn.start", async ($, e, next) => {
450    isTurnMuted = false
451    return next(e)
452  })
453
454  on("turn.complete", async ($, e, next) => {
455    if (e.reason === "aborted" && e.agentId === undefined) await stop()
456    if (isEnabled && !isTurnMuted && e.agentId === undefined) {
457      if (config.mode === "final" && e.reason === "answer") {
458        const budget: Budget = { spoken: 0, isCut: false }
459        enqueue(toSentences(e.answer), budget)
460        finishResponse(budget)
461      }
462      const notice = TURN_NOTICES[e.reason]
463      if (config.mode === "notify" && notice !== undefined && config.notices[notice] !== "") enqueue([config.notices[notice]])
464    }
465    return next(e)
466  })
467
468  // Whatever the mode, a dialog waiting for the person cuts off the reading and is announced.
469  on("tool.call", async ($, e, next) => {
470    if (isEnabled && DIALOG_TOOLS.has(e.tool)) {
471      await stop()
472      const { notice, detail } = dialogSpeech(e.tool, e, config.notices)
473      queueNotice(notice)
474      enqueue(detail)
475    }
476    return next(e)
477  }).catch(($, e, next) => next(e))
478
479  // Raised as a permission dialog is about to be shown, not for a call a rule or the mode settles.
480  on("classic.PermissionRequest", async ($, e, next) => {
481    if (isEnabled && !DIALOG_TOOLS.has(e.tool_name)) {
482      await stop()
483      const { notice, detail } = dialogSpeech(e.tool_name, e.tool_input, config.notices)
484      queueNotice(notice)
485      enqueue(detail)
486    }
487    return next(e)
488  }).catch(($, e, next) => next(e))
489
490  // Speaks the main loop's text as it streams; thinking, tool calls and subagents stay silent.
491  on("turn.step", async function* ($, e, next) {
492    const stream = next(e)
493    if (!isEnabled || config.mode !== "stream" || e.agentId !== undefined) return yield* stream
494
495    // Each response has its own budget: text before a tool call does not use up the answer's.
496    const budget: Budget = { spoken: 0, isCut: false }
497    const splitter = new SpeechSplitter()
498    const responseGeneration = generation
499    const queueResponse = (texts: readonly Spoken[]) => {
500      if (!isTurnMuted && responseGeneration === generation) enqueue(texts, budget)
501    }
502    let block: number | undefined
503    try {
504      for await (const chunk of stream) {
505        if (chunk.kind === "text") {
506          if (block !== undefined && block !== chunk.index) queueResponse(splitter.flush())
507          block = chunk.index
508          queueResponse(splitter.push(chunk.text))
509        }
510        yield chunk
511      }
512    } finally {
513      // Runs even when the engine stops pulling after the last chunk: the line without a newline is said too.
514      queueResponse(splitter.flush())
515      if (!isTurnMuted && responseGeneration === generation) finishResponse(budget)
516    }
517    return stream.result
518  })
519}
520
521/** Creates a temporary directory, removing the previous module's files after a reload. */
522async function prepareDirectory($: EngineInterface, isWindows: boolean): Promise<string> {
523  const previous = await read($, audioDirectory)
524  if (previous !== null) await $.process.run(removeDirectoryCommand(previous, isWindows))
525  const created = await $.process.run(tempDirectoryCommand(isWindows))
526  if (created.exitCode !== 0 || created.stdout.trim() === "") throw new Error("tts: could not create a private audio directory")
527  const directory = created.stdout.trim()
528  await update($, audioDirectory, () => directory)
529  return directory
530}
531
532
533// Everything handed $ stays in this file: the validator follows $ into this file's functions only.
534
535/** Runs an engine's commands for one sentence; throws on the first that fails. */
536async function synthesize(
537  $: EngineInterface,
538  process: SpeechProcess,
539  engine: LoadedEngine,
540  text: string,
541  file: string,
542  isCurrent: () => boolean,
543) {
544  let stdout = ""
545  for (const step of engine.synthesis(text, file)) {
546    if (!isCurrent()) return
547    const stdin = typeof step.stdin === "function" ? step.stdin(stdout) : step.stdin
548    const ran = await runProcess($, process, step.argv, stdin, SYNTH_TIMEOUT_MS)
549    if (ran.code !== 0) {
550      throw new Error(`${step.argv[0]} exited ${ran.code}: ${(ran.stderr || ran.stdout).trim().slice(0, 200)}`)
551    }
552    stdout = ran.stdout
553  }
554}
555
556/** A cancellable command with the same time limit synthesis had under process.run. */
557async function runProcess($: EngineInterface, process: SpeechProcess, argv: string[], input: string | undefined, timeoutMs: number) {
558  const running = process.read($.process.spawn({ argv, input }))
559  let timedOut = false
560  const timer = $.clock.after(timeoutMs, () => {
561    timedOut = true
562    void process.stop()
563  })
564  try {
565    const result = await running
566    if (timedOut) throw new Error(`${argv[0]} timed out`)
567    return result
568  } finally {
569    timer.cancel()
570  }
571}
572
573/** Whether the engine can read now: set up (its synthesis builds) and answering its probe. */
574async function checkEngine($: EngineInterface, engine: LoadedEngine): Promise<"ok" | "down" | "unset"> {
575  if (engine.probe === undefined) return "unset"
576  try {
577    engine.synthesis("", "probe.wav")
578  } catch {
579    return "unset"
580  }
581  const probed = await $.process.run(engine.probe, { timeoutMs: 5_000 }).catch(() => undefined)
582  return probed?.exitCode === 0 ? "ok" : "down"
583}
584
585/** ● when the engine answers now, ○ when it is down or not set up. */
586async function answerMark($: EngineInterface, engine: LoadedEngine): Promise<string> {
587  return (await checkEngine($, engine)) === "ok" ? "●" : "○"
588}
589
590/** One line of `/tts engine`: whether the engine answers now, where, and its role. */
591async function engineRow($: EngineInterface, name: EngineName, engine: LoadedEngine, config: TtsConfig): Promise<string> {
592  const role = name === config.engine ? "  ← in use" : name === config.fallback ? "  ← fallback" : ""
593  return `  ${await answerMark($, engine)} ${name.padEnd(9)} ${engine.location}${role}`
594}
595
596/** Opens the picker pane on `shown`; false when the surface cannot place it. */
597async function openPicker($: EngineInterface, shown: Picker): Promise<boolean> {
598  await update($, picker, () => shown)
599  const rows = Math.min(shown.options.length + 3, 14)
600  const opened = await $.ui.open({ id: PICKER_PANE, title: `tts ${shown.setting}`, focus: true, closeOnEscape: true, rows })
601  return opened.isPlaced
602}
603
604
605/**
606 * Saves a change, to `/config` (which reloads the module) or to this
607 * directory's pin; answers how it reads, before and after, or why it did not.
608 */
609async function commit($: EngineInterface, change: Change, applyPins: (stored: Pins) => void): Promise<string> {
610  if ("error" in change) return change.error
611  if ("config" in change) {
612    const set = await $.config.set(change.config)
613    if ("deny" in set && set.deny) return set.deny
614  } else {
615    const stored = change.pins(toPins(await $.store.get(PINS_KEY)))
616    await $.store.set(PINS_KEY, stored)
617    applyPins(stored)
618  }
619  return `${change.label}: ${change.previous} → ${change.next}`
620}
621
hooks/channel.ts 21 lines
1/** An unbounded queue whose `take` waits for the next item. */
2export class Channel<T> {
3  private items: T[] = []
4  private waiter: (() => void) | undefined
5
6  get size(): number {
7    return this.items.length
8  }
9
10  push(item: T) {
11    this.items.push(item)
12    this.waiter?.()
13    this.waiter = undefined
14  }
15
16  async take(): Promise<T> {
17    while (this.items.length === 0) await new Promise<void>(resolve => (this.waiter = resolve))
18    return this.items.shift() as T
19  }
20}
21
hooks/config.ts 70 lines
1import type { PluginOptions } from "claude-code"
2
3import { type EngineName, ENGINES } from "./engines"
4
5/** stream: every sentence as it streams. final: the turn's answer once it ends. notify: short notices alone. */
6export type ReadMode = "stream" | "final" | "notify"
7
8/** The lines said as notices, each a `/config` row; an empty one stays silent. */
9export type Notices = {
10  done: string
11  error: string
12  refusal: string
13  question: string
14  /** `{n}` is the number of questions. */
15  questions: string
16  plan: string
17  command: string
18  tool: string
19  cut: string
20  code: string
21}
22
23const NOTICE_OPTIONS: Record<keyof Notices, string> = {
24  done: "noticeDone",
25  error: "noticeError",
26  refusal: "noticeRefusal",
27  question: "noticeQuestion",
28  questions: "noticeQuestions",
29  plan: "noticePlan",
30  command: "noticeCommand",
31  tool: "noticeTool",
32  cut: "noticeCut",
33  code: "noticeCode",
34}
35
36/** The settings every engine shares; each engine reads its own in `load`. */
37export type TtsConfig = {
38  isEnabled: boolean
39  mode: ReadMode
40  engine: EngineName
41  /** The engine tried when `engine` fails a sentence; undefined for none. */
42  fallback: EngineName | undefined
43  player: string
44  volume: number
45  maxCharsPerTurn: number
46  interruptOnPrompt: boolean
47  /** Added to the system prompt while replies are read; empty: nothing. */
48  modelPrompt: string
49  notices: Notices
50}
51
52export function toConfig(options: PluginOptions): TtsConfig {
53  const engine = String(options.engine)
54  const fallback = String(options.fallback)
55  return {
56    isEnabled: options.enabled !== false,
57    mode: options.mode as ReadMode,
58    engine: engine in ENGINES ? (engine as EngineName) : "voicevox",
59    fallback: fallback in ENGINES && fallback !== engine ? (fallback as EngineName) : undefined,
60    player: String(options.player),
61    volume: Number(options.volume),
62    maxCharsPerTurn: Number(options.maxCharsPerTurn),
63    interruptOnPrompt: options.interruptOnPrompt === true,
64    modelPrompt: String(options.modelPrompt ?? "").trim(),
65    notices: Object.fromEntries(
66      Object.entries(NOTICE_OPTIONS).map(([notice, option]) => [notice, String(options[option] ?? "").trim()]),
67    ) as Notices,
68  }
69}
70
hooks/dialog-text.ts 34 lines
1import type { Notices } from "./config"
2import { type Spoken, toSentences } from "./speech-text"
3
4/** Tools whose call itself opens a dialog for the person; the rest open one only to ask permission. */
5export const DIALOG_TOOLS: ReadonlySet<string> = new Set(["AskUserQuestion", "ExitPlanMode"])
6
7/**
8 * What is said as a dialog opens: `notice`, a short line from the settings
9 * that is cached and plays at once (empty: none), then `detail`, synthesized
10 * each time.
11 */
12export type DialogSpeech = { notice: string; detail: Spoken[] }
13
14export function dialogSpeech(tool: string, input: unknown, notices: Notices): DialogSpeech {
15  const fields = (typeof input === "object" && input !== null ? input : {}) as Record<string, unknown>
16  switch (tool) {
17    case "AskUserQuestion": {
18      const questions = Array.isArray(fields.questions) ? (fields.questions as { question?: unknown }[]) : []
19      // Several questions are counted, not read: the dialog shows them one at a time.
20      if (questions.length > 1) return { notice: notices.questions.replaceAll("{n}", String(questions.length)), detail: [] }
21      const question = questions[0]?.question
22      return { notice: notices.question, detail: notices.question !== "" && typeof question === "string" ? toSentences(question) : [] }
23    }
24    case "ExitPlanMode":
25      return { notice: notices.plan, detail: [] }
26    case "Bash": {
27      const { description } = fields
28      return { notice: notices.command, detail: notices.command !== "" && typeof description === "string" ? toSentences(description) : [] }
29    }
30    default:
31      return { notice: notices.tool, detail: [] }
32  }
33}
34
hooks/engines/index.ts 12 lines
1import { command } from "./command"
2import type { Engine } from "./engine"
3import { irodori } from "./irodori"
4import { voicevox } from "./voicevox"
5
6export type { Engine, LoadedEngine, Step } from "./engine"
7
8/** Every engine by the name `/tts engine` and the `engine` setting take. */
9export const ENGINES = { voicevox, irodori, command } satisfies Record<string, Engine>
10
11export type EngineName = keyof typeof ENGINES
12
hooks/pins.ts 31 lines
1import type { PluginOptions } from "claude-code"
2
3import { type EngineName, ENGINES } from "./engines"
4
5/** What a working directory pins: its engine, and the voice it chose for each engine. */
6export type Pin = { engine: EngineName; voices: Partial<Record<EngineName, string>> }
7
8/** Pins by working directory, as `$.store` keeps them. */
9export type Pins = Record<string, Pin>
10
11export const PINS_KEY = "pins"
12
13/** The options with a directory's pin laid over them; the options alone without one. */
14export function pinnedOptions(options: PluginOptions, pin: Pin | undefined): PluginOptions {
15  if (pin === undefined) return options
16  const pinned: Record<string, PluginOptions[string]> = { ...options, engine: pin.engine }
17  for (const [name, voice] of Object.entries(pin.voices) as [EngineName, string][]) {
18    const selected = ENGINES[name].load(options).voices?.select(voice)
19    if (selected !== undefined) pinned[selected.key.replace(/^tts\./, "")] = selected.value
20  }
21  return pinned
22}
23
24/** Reads `$.store`'s value as pins, dropping what is not one. */
25export function toPins(stored: unknown): Pins {
26  if (typeof stored !== "object" || stored === null) return {}
27  return Object.fromEntries(
28    Object.entries(stored).filter(([, pin]) => typeof pin === "object" && pin !== null && (pin as Pin).engine in ENGINES),
29  ) as Pins
30}
31
hooks/platform.ts 36 lines
1/** Quotes a literal for PowerShell, including the smart quotes it treats as delimiters. */
2export function quotePowerShell(value: string): string {
3  return `'${value.replace(/['‘’‚‛]/g, quote => quote + quote)}'`
4}
5
6/** Keep paths containing Japanese readable when Windows PowerShell writes to a pipe. */
7export function powershellCommand(script: string): string[] {
8  return [
9    "powershell.exe", "-NoProfile", "-NonInteractive", "-Command",
10    `$ErrorActionPreference = 'Stop'; [Console]::OutputEncoding = [Text.UTF8Encoding]::new(); ${script}`,
11  ]
12}
13
14/** Unix uses mode 0700; Windows inherits the user's temporary folder permissions. */
15export function tempDirectoryCommand(isWindows: boolean): string[] {
16  return isWindows
17    ? powershellCommand("$path = Join-Path ([IO.Path]::GetTempPath()) ('claude-tts-' + [Guid]::NewGuid().ToString('N')); [IO.Directory]::CreateDirectory($path).FullName")
18    : ["mktemp", "-d", "/tmp/claude-tts-XXXXXXXXXX"]
19}
20
21export function removeFilesCommand(files: string[], isWindows: boolean): string[] {
22  return isWindows
23    ? powershellCommand(`Remove-Item -LiteralPath ${files.map(quotePowerShell).join(", ")} -Force -ErrorAction SilentlyContinue; exit 0`)
24    : ["rm", "-f", ...files]
25}
26
27export function removeDirectoryCommand(directory: string, isWindows: boolean): string[] {
28  return isWindows
29    ? powershellCommand(`Remove-Item -LiteralPath ${quotePowerShell(directory)} -Recurse -Force -ErrorAction SilentlyContinue; exit 0`)
30    : ["rm", "-rf", "--", directory]
31}
32
33export function findProgramCommand(program: string, isWindows: boolean): string[] {
34  return isWindows ? ["where.exe", program] : ["sh", "-c", 'command -v "$1" >/dev/null', "probe", program]
35}
36
hooks/player.ts 85 lines
1import { powershellCommand, quotePowerShell } from "./platform"
2
3const TAIL_PADDING_SECONDS = 0.2
4
5/** Prints the first audio player found on this machine. */
6export function findPlayerCommand(isWindows: boolean): string[] {
7  if (isWindows) return powershellCommand("if (Get-Command ffplay -CommandType Application -ErrorAction SilentlyContinue) { 'ffplay' } else { 'powershell' }")
8  return [
9    "sh",
10    "-c",
11    // biome-ignore lint/suspicious/noTemplateCurlyInString: shell parameter expansion
12    "for p in ffplay paplay aplay afplay powershell.exe; do command -v $p >/dev/null && { echo ${p%.exe}; exit; }; done",
13  ]
14}
15
16/**
17 * Cuts long engine pauses, then gives ffplay and SoundPlayer time to finish speech.
18 * Padding must be in the WAV: ffplay does not flush an apad filter at EOF.
19 * Leaves the file as it was where ffmpeg is missing.
20 */
21export function prepareAudioCommand(file: string, player: string | undefined, isWindows = false): string[] {
22  const edge = (keep: number) => `silenceremove=start_periods=1:start_threshold=-45dB:start_silence=${keep}`
23  const padding = player === "ffplay" || player === "powershell" ? `,apad=pad_dur=${TAIL_PADDING_SECONDS}` : ""
24  const filter = `${edge(0.05)},areverse,${edge(0.2)},areverse${padding}`
25  if (isWindows) return powershellCommand([
26    "if (-not (Get-Command ffmpeg -CommandType Application -ErrorAction SilentlyContinue)) { exit 0 }",
27    `$file = ${quotePowerShell(file)}`,
28    `ffmpeg -loglevel error -y -i $file -af '${filter}' ($file + '.t.wav')`,
29    "if ($LASTEXITCODE -eq 0) { Move-Item -LiteralPath ($file + '.t.wav') -Destination $file -Force }",
30  ].join("; "))
31  return [
32    "sh",
33    "-c",
34    `command -v ffmpeg >/dev/null || exit 0; ffmpeg -loglevel error -y -i "$1" -af "${filter}" "$1.t.wav" && mv -f "$1.t.wav" "$1"`,
35    "tts-trim",
36    file,
37  ]
38}
39
40/**
41 * PowerShell that plays the wav at `path` (a PowerShell expression) to its end. The file goes in
42 * as a stream: SoundPlayer given a \\wsl.localhost path with capitals plays nothing, without error.
43 */
44function soundPlayerScript(path: string, variable = "$"): string {
45  const stream = `${variable}stream`
46  return `${stream} = [IO.File]::OpenRead(${path}); try { (New-Object Media.SoundPlayer ${stream}).PlaySync() } finally { ${stream}.Dispose() }`
47}
48
49/**
50 * Plays `file` at `volume` (1 as synthesized), preserving the cached original.
51 */
52export function playCommand(player: string, file: string, volume: number, isWindows = false): string[] {
53  const shell = (script: string) => ["sh", "-c", script, "tts-play", file]
54  const isScaled = volume !== 1
55  // aplay and SoundPlayer have no volume: play a scaled copy when ffmpeg is present.
56  const scaleFile = isScaled
57    ? `command -v ffmpeg >/dev/null && ffmpeg -loglevel error -y -i "$1" -af volume=${volume} "$1.v.wav" && set -- "$1.v.wav"; `
58    : ""
59  switch (player) {
60    case "ffplay":
61      return ["ffplay", "-nodisp", "-autoexit", "-loglevel", "error", ...(isScaled ? ["-af", `volume=${volume}`] : []), file]
62    case "paplay":
63      return ["paplay", `--volume=${Math.round(65536 * volume)}`, file]
64    case "afplay":
65      return ["afplay", "-v", String(volume), file]
66    case "aplay":
67      return shell(`${scaleFile}exec aplay -q "$1"`)
68    case "powershell":
69      if (isWindows) return powershellCommand([
70        `$file = ${quotePowerShell(file)}`,
71        ...(isScaled ? [
72          `if (Get-Command ffmpeg -CommandType Application -ErrorAction SilentlyContinue) { ffmpeg -loglevel error -y -i $file -af volume=${volume} ($file + '.v.wav'); if ($LASTEXITCODE -eq 0) { $file += '.v.wav' } }`,
73        ] : []),
74        soundPlayerScript("$file"),
75      ].join("; "))
76      // WSL: Windows reads the file through its \\wsl.localhost path.
77      return shell(
78        // `\$` keeps sh from expanding PowerShell's variable.
79        `${scaleFile}exec powershell.exe -NoProfile -NonInteractive -Command "${soundPlayerScript(`'$(wslpath -w "$1")'`, "\\$")}"`,
80      )
81    default:
82      return [player, file]
83  }
84}
85
hooks/process.ts 33 lines
1import type { HookStream, ProcessSpawnChunk, ProcessSpawnResult } from "claude-code"
2
3/** One synthesis or playback process; closing its stream also kills the child. */
4export class SpeechProcess {
5  private stream: HookStream<ProcessSpawnChunk, ProcessSpawnResult> | undefined
6
7  get isRunning(): boolean {
8    return this.stream !== undefined
9  }
10
11  async stop(): Promise<void> {
12    await this.stream?.return({ code: null, signal: "SIGTERM" }).catch(() => undefined)
13  }
14
15  async read(stream: HookStream<ProcessSpawnChunk, ProcessSpawnResult>) {
16    this.stream = stream
17    // Closing a stream rejects its result; observe that rejection even during cancellation.
18    const result = stream.result.catch(() => ({ code: null, signal: "SIGTERM" }))
19    let stdout = ""
20    let stderr = ""
21    try {
22      for await (const piece of stream) {
23        // Bound output so a noisy command cannot fill memory.
24        if (piece.stream === "stdout") stdout = (stdout + piece.text).slice(0, 4_194_304)
25        else stderr = (stderr + piece.text).slice(0, 4_194_304)
26      }
27      return { ...(await result), stdout, stderr }
28    } finally {
29      if (this.stream === stream) this.stream = undefined
30    }
31  }
32}
33
hooks/settings.ts 69 lines
1import type { PluginOptions } from "claude-code"
2
3import type { PickerOption, Setting } from "../types"
4
5import type { ReadMode, TtsConfig } from "./config"
6import { type EngineName, ENGINES, type LoadedEngine } from "./engines"
7import type { Pin, Pins } from "./pins"
8
9export const MODES: readonly ReadMode[] = ["stream", "final", "notify"]
10export const MODE_LABELS: Record<ReadMode, string> = {
11  stream: "stream  返答を一文ずつ、流れてくるそばから",
12  final: "final   ターンが終わってから最終的な返答だけ",
13  notify: "notify  完了・エラーなどの短い通知だけ",
14}
15export const SPEEDS = ["0.8", "1", "1.1", "1.2", "1.5", "2"]
16export const VOLUMES = ["0.5", "0.75", "1", "1.25", "1.5"]
17
18/** A setting change ready to save: where it goes, and how it reads before and after. */
19export type Change =
20  | { label: string; previous: string; next: string; config: { key: string; value: string | number } }
21  | { label: string; previous: string; next: string; pins: (stored: Pins) => Pins }
22  | { error: string }
23
24/** The settings in effect, which a change is measured against. */
25export type Current = { options: PluginOptions; config: TtsConfig; engine: LoadedEngine; cwd: string; pin: Pin | undefined }
26
27/** What saving `value` for `setting` changes; while pinned, engine and voice go to the pin. */
28export function planChange(setting: Setting, value: string, { options, config, engine, cwd, pin }: Current): Change {
29  const pinned = (update: (base: Pin) => Pin) => (stored: Pins) => {
30    const base = stored[cwd] ?? pin
31    return base === undefined ? stored : { ...stored, [cwd]: update(base) }
32  }
33  switch (setting) {
34    case "mode":
35      if (!MODES.includes(value as ReadMode)) return { error: `mode is one of ${MODES.join(", ")}` }
36      return { label: "mode", previous: config.mode, next: value, config: { key: "tts.mode", value } }
37    case "speed":
38    case "volume": {
39      const scale = Number(value)
40      if (!(scale > 0 && scale <= 4)) return { error: `${setting} is a number above 0, up to 4` }
41      const previous = String(setting === "speed" ? options.speed : config.volume)
42      return { label: setting, previous, next: String(scale), config: { key: `tts.${setting}`, value: scale } }
43    }
44    case "engine": {
45      if (!(value in ENGINES)) return { error: `engine is one of ${Object.keys(ENGINES).join(", ")}` }
46      const name = value as EngineName
47      if (pin === undefined) return { label: "engine", previous: config.engine, next: name, config: { key: "tts.engine", value: name } }
48      return { label: "engine (this directory)", previous: config.engine, next: name, pins: pinned(base => ({ ...base, engine: name })) }
49    }
50    case "voice": {
51      const { voices } = engine
52      if (voices === undefined) return { error: `engine=${config.engine} has no voices; set it up in /config` }
53      if (pin === undefined) return { label: "voice", previous: voices.current, next: value, config: voices.select(value) }
54      const name = config.engine
55      return {
56        label: "voice (this directory)",
57        previous: voices.current,
58        next: value,
59        pins: pinned(base => ({ ...base, voices: { ...base.voices, [name]: value } })),
60      }
61    }
62  }
63}
64
65/** A choice as a text listing shows it: its id, then its name where it has one. */
66export function optionLine(option: PickerOption): string {
67  return option.label === undefined ? `  ${option.value}` : `  ${option.value}: ${option.label}`
68}
69
hooks/speech-text.ts 214 lines
1/** Sentence ends: Japanese/full-width punctuation with any closing brackets, or a period before whitespace. */
2const TERMINATOR = /[。!?!?]+[」』))"]*|\.(?=\s)/g
3const SPEAKABLE = /[\p{L}\p{N}]/u
4const MAX_SENTENCE = 160
5const LIST_MARKER = /^\s*(?:[-*+–—•]|\d+[.)])\s+(?:\[[ xX]\]\s+)?/
6const PATH_NAME = String.raw`[\p{L}\p{N}\p{M}_@+-][\p{L}\p{N}\p{M}_.@+-]*\.[A-Za-z][A-Za-z0-9]{0,9}`
7const PATH_LINE = String.raw`(?:(?::|#L)(\d+)(?:-L?(\d+))?(?::\d+)?)?`
8/**
9 * A file path, optionally in backticks and with a line (`:12`, `:12:5`,
10 * `:12-20`, `#L12-L20`). Folders need a separator and the name an extension
11 * that starts with a letter, so `and/or`, `1/2` and `v1.2/v1.3` never match.
12 * Folders have no spaces, so a match never runs across prose.
13 */
14const PATH = new RegExp(
15  String.raw`(?<![\p{L}\p{N}\p{M}_.:/\\~-])(\x60?)(~|[A-Za-z]:)?((?:[\p{L}\p{N}\p{M}_.@+-]*[/\\])+)(${PATH_NAME})${PATH_LINE}(\x60?)(?![\p{L}\p{N}\p{M}_/\\])`,
16  "gu",
17)
18/** Ends what came before, so a path may follow at once: Japanese punctuation and brackets. */
19const PATH_SEPARATORS = /[\s、。,:;([{〈《「『【]/
20/** May open a path when they start a word: ASCII brackets and quotes. */
21const PATH_OPENERS = /[([{<"'\x60]/
22/** A whole inline code span that is a path; only there, its folders may have spaces. */
23const CODE_PATH = new RegExp(String.raw`^(~|[A-Za-z]:)?((?:[\p{L}\p{N}\p{M}_.@+ -]*[/\\])+)(${PATH_NAME})${PATH_LINE}$`, "u")
24
25/** Where fenced code was skipped; the reader decides what, if anything, to say there. */
26export const SKIPPED_CODE: unique symbol = Symbol("skipped code")
27
28/** What a reply turns into for a listener: its sentences, and marks where code was skipped. */
29export type Spoken = string | typeof SKIPPED_CODE
30
31/**
32 * Turns streamed Markdown into speakable sentences as early as possible:
33 * complete sentences of the current line are released before its newline
34 * arrives, and fenced code and tables are never spoken.
35 */
36export class SpeechSplitter {
37  private partial = ""
38  private fence: string | undefined
39
40  push(text: string): Spoken[] {
41    this.partial += text
42    const sentences: Spoken[] = []
43
44    let newline = this.partial.indexOf("\n")
45    while (newline >= 0) {
46      sentences.push(...this.line(this.partial.slice(0, newline)))
47      this.partial = this.partial.slice(newline + 1)
48      newline = this.partial.indexOf("\n")
49    }
50
51    // Keep unfinished links and URLs whole: punctuation inside them is not a sentence end.
52    // Lines that may open a fence or a table also wait for their newline.
53    if (this.fence === undefined && !/^\s*[`~|]/.test(this.partial)) {
54      const markup = this.partial.search(/https?:\/\/|!?\[/)
55      // The period in "1. " belongs to the list marker, not a spoken sentence.
56      const start = LIST_MARKER.exec(this.partial)?.[0].length ?? 0
57      const end = markup < 0 ? this.partial.length : Math.max(start, markup)
58      const cut = lastTerminatorEnd(this.partial.slice(start, end))
59      if (cut > 0) {
60        sentences.push(...this.line(this.partial.slice(0, start + cut)))
61        this.partial = this.partial.slice(start + cut)
62      }
63    }
64    return sentences
65  }
66
67  flush(): Spoken[] {
68    const sentences = this.line(this.partial)
69    this.partial = ""
70    this.fence = undefined
71    return sentences
72  }
73
74  private line(text: string): Spoken[] {
75    const marker = /^\s*(`{3,}|~{3,})(.*)$/.exec(text)
76    if (this.fence !== undefined) {
77      if (marker && marker[1]!.startsWith(this.fence) && marker[2]!.trim() === "") {
78        this.fence = undefined
79      }
80      return []
81    }
82    if (marker) {
83      this.fence = marker[1]
84      return [SKIPPED_CODE]
85    }
86    return splitSentences(toSpeakable(text))
87  }
88}
89
90/** A whole Markdown text as speakable sentences. */
91export function toSentences(text: string): Spoken[] {
92  const splitter = new SpeechSplitter()
93  return [...splitter.push(text), ...splitter.flush()]
94}
95
96/** Strips the Markdown a listener should not hear from one line. */
97export function toSpeakable(line: string): string {
98  const text = line.trim()
99  if (text.startsWith("|") || /^[-*_]{3,}$/.test(text)) return ""
100  return text
101    .replace(/^#{1,6}\s+/, "")
102    .replace(/^>\s?/, "")
103    .replace(LIST_MARKER, "")
104    .replace(/!\[[^\]]*\]\([^)]*\)/g, "")
105    .replace(/\[([^\]]+)\]\([^)]*\)/g, "$1")
106    .replace(/https?:\/\/\S+/g, "URL")
107    .replace(/`([^`]+)`/g, speakCodePath)
108    .replace(PATH, speakPath)
109    .replace(/`([^`]*)`/g, "$1")
110    .replace(/\*\*|~~|(?<!\w)\*(?!\s)|(?<!\s)\*(?!\w)/g, "")
111    .replace(/[()[\]{}()[]{}「」『』【】〈〉《》]/g, " ")
112    .replace(/\s+/g, " ")
113    .trim()
114}
115
116/** Reads a path as its file name and line: "speech-text.tsの12行目". */
117function spokenPath(name: string, line: string | undefined, end: string | undefined): string {
118  return line === undefined ? name : end === undefined ? `${name}の${line}行目` : `${name}の${line}から${end}行目`
119}
120
121/** Whether a path may start right after `before`, not in the middle of a longer one. */
122function startsPath(before: string): boolean {
123  const last = before.at(-1)
124  if (last === undefined || PATH_SEPARATORS.test(last)) return true
125  if (!PATH_OPENERS.test(last)) return false
126  const prior = before.at(-2)
127  return prior === undefined || !/[\x21-\x7e]/.test(prior)
128}
129
130function isRootedPath(root: string | undefined, dirs: string): boolean {
131  return root !== undefined || /^\.{0,2}[/\\]/.test(dirs)
132}
133
134/**
135 * Shortens a code span that is a path with spaces in its folders, when it is
136 * rooted: `cat src/a.ts` or `My Project/src/a.ts` could be a command or a
137 * folder, so they stay whole.
138 */
139function speakCodePath(span: string, inner: string): string {
140  const match = CODE_PATH.exec(inner)
141  if (match === null) return span
142  const [, root, dirs = "", name = "", line, end] = match
143  return isRootedPath(root, dirs) ? spokenPath(name, line, end) : span
144}
145
146/**
147 * Shortens a path without spaces, only when it is clearly a path: rooted
148 * (`/`, `~/`, `./`, `C:\\`), with a line, or a whole inline code span. A
149 * bare `React/Next.js` stays. A path joined to the text before it, or an
150 * unrooted one right after a word and a space, may be the tail of a longer
151 * path (`/tmp/foo&bar/a.ts`, `My Project/src/a.ts`), so it stays whole.
152 */
153function speakPath(
154  match: string,
155  open: string,
156  root: string | undefined,
157  dirs: string,
158  name: string,
159  line: string | undefined,
160  end: string | undefined,
161  close: string,
162  offset: number,
163  text: string,
164): string {
165  const code = open !== "" && close !== ""
166  const isRooted = isRootedPath(root, dirs)
167  if (!(code || line !== undefined || isRooted)) return match
168  if (open === "") {
169    const before = text.slice(0, offset)
170    // Text joined to the match (`foo&` in `/tmp/foo&bar/a.ts`) means it is the tail of a longer path.
171    if (!startsPath(before)) return match
172    if (!isRooted && /[\p{L}\p{N}\p{M}] $/u.test(before)) return match
173  }
174  const spoken = spokenPath(name, line, end)
175  return code ? spoken : `${open}${spoken}${close}`
176}
177
178function splitSentences(text: string): string[] {
179  const sentences: string[] = []
180  let start = 0
181  for (const match of text.matchAll(TERMINATOR)) {
182    const end = match.index + match[0].length
183    sentences.push(text.slice(start, end))
184    start = end
185  }
186  sentences.push(text.slice(start))
187  return sentences
188    .map(sentence => sentence.trim())
189    .filter(sentence => SPEAKABLE.test(sentence))
190}
191
192/**
193 * Engines choke on very long input: split at a comma near the limit, else hard.
194 * Applied as a sentence is queued, so the response budget sees whole sentences.
195 */
196export function capLength(sentence: string): string[] {
197  const pieces: string[] = []
198  let rest = Array.from(sentence)
199  while (rest.length > MAX_SENTENCE) {
200    const comma = Math.max(rest.lastIndexOf("、", MAX_SENTENCE - 1), rest.lastIndexOf(",", MAX_SENTENCE - 1))
201    const cut = comma > MAX_SENTENCE / 2 ? comma + 1 : MAX_SENTENCE
202    pieces.push(rest.slice(0, cut).join(""))
203    rest = rest.slice(cut)
204  }
205  pieces.push(rest.join(""))
206  return pieces
207}
208
209function lastTerminatorEnd(text: string): number {
210  let end = 0
211  for (const match of text.matchAll(TERMINATOR)) end = match.index + match[0].length
212  return end
213}
214
hooks/engines/command.ts 21 lines
1import { findProgramCommand } from "../platform"
2import type { Engine } from "./engine"
3
4/** Any command: argv split on spaces, `{out}` the wav to write, the text on stdin. */
5export const command: Engine = {
6  load: (options, isWindows = false) => {
7    const template = String(options.command)
8    const program = template.split(/\s+/).find(Boolean)
9    return {
10      voice: template || "(unset)",
11      location: template || "(unset)",
12      probe: program === undefined ? undefined : findProgramCommand(program, isWindows),
13      synthesis: (text, out) => {
14        const argv = template.split(/\s+/).filter(Boolean).map(arg => arg.replaceAll("{out}", out))
15        if (argv.length === 0) throw new Error("command is not set (/config → tts)")
16        return [{ argv, stdin: text }]
17      },
18    }
19  },
20}
21