SLOPSHOPPER

tts

在每段回复下加一个🔊按钮,用火山引擎豆包语音(知性女声 2.0)流式朗读,可暂停、继续、按节跳转、从头念、倍速,音色可自行添加切换

newpanerowscommandtoaststatus
v0.4.3MITupdated 2026-10-07gnehiur/claude-code-volc-tts
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · tts
│ ┃ 🎙 朗读音色 ✕ › fix the failing auth test and add an audit log call │ ┃ 添加音色 │ ┃ 音色 ID : 例如 ICL_uranus_zh_female_xingganm ⏺ Read(src/auth.ts) │ ┃ 名字 : 可选,不填就用 ID ⏎ 验证 ⎿ Read 6 lines │ ┃ [ 验证并保存 ] ⏺ Update(src/auth.ts) │ ┃ ⎿ Added 2 lines, removed 1 line │ ┃ 音色 ID ⏺ Bash(bun test) │ ┃ 在火山控制台「音色库」里找;验证会真的合成一 ⎿ 3 pass, 1 fail │ ┃ 句试听语(约 10 个字计费)。 │ ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ │ › /voice │ ⎿ tts: 已打开朗读音色面板。 │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · 🎙 朗读音色
添加音色 音色 ID : 例如 ICL_uranus_zh_female_xingganmeihuo_tob ⏎ 验证 名字 : 可选,不填就用 ID ⏎ 验证 [ 验证并保存 ] 音色 ID 在火山控制台「音色库」里找;验证会真的合成一句试听语(约 10 个字计费)。
README

claude-code-volc-tts

一个 Claude Code mod:在每轮回复的最终回复下面加一个 🔊 按钮,用火山引擎豆包语音流式朗读。音色可以在面板里自己添加、切换(默认「知性女声 2.0」)。桌面 App 的 Code 标签页和终端版都能用。

没在念:   🔊
正在念:   ⏸   ⏮   ⏭   ⟲   ⏹   1×   🎙
已暂停:   ▶️   ⏮   ⏭   ⟲   ⏹   1×   🎙

功能

  • 流式朗读:边合成边播放,首声约 0.8 秒,和回复长短无关
  • 三态主按钮:🔊 开始 → ⏸ 暂停 → ▶️ 从停下的地方继续
  • ⏮ ⏭ 按节跳转:节的边界是 ## 标题、--- 分隔线和单独一行的粗体(例如结尾的 这次攒下了什么);没有标题的回复按段落切。⏮ 在本节念了 3 秒以上时回到本节开头,否则回到上一节;在最后一节按 ⏭ 结束朗读。状态栏显示进度,例如 🔊 火山朗读中 2/5
  • ⟲ 从头念 / ⏹ 停止:只在念或暂停时出现
  • 代码块略过,表格逐行念:代码块念成"这里有一段代码,略过";表格每行念成一句,第一格当这一行的名字,其余格念成"表头:内容",例如 | 装 mpv | 0.41.0 |(表头"结果")念成"装 mpv,结果:0.41.0。"
  • 倍速:1× → 1.25× → 1.5× → 2× 循环,音调不变,记住上次的选择
  • 只在最终回复显示:Claude 干活过程中的说明文字不加按钮
  • 全局只有一个声音:在任何会话里开始新的朗读,旧的自动结束
  • 切走会话自动暂停(仅桌面 App):切回来点 ▶️ 继续
  • 本地缓存:按节缓存,同一节再念不调用火山、不花钱,0.2 秒出声;保留 7 天、上限 200MB
  • 音色面板:点 🎙 或输入 /voice 打开;下拉切换、试听、删除,填音色 ID 验证通过才保存
  • 后备模式:mpv 或 Python 不可用时,退回整段合成后播放(只能停止,不能暂停)

依赖

  • macOS(用到了 /usr/bin/python3、afplay 和 Claude 桌面 App 的日志路径)
  • Claude Code 2.1.259 及以上(需要 mod / function hooks 功能)
  • mpv:brew install mpv
  • 火山引擎账号,开通「豆包语音合成大模型 2.0」,在控制台 API Key 管理 里创建一个 API Key

安装

  1. 克隆到任意位置,例如:
   git clone https://github.com/gnehiur/claude-code-volc-tts.git ~/Projects/claude-mods/volc-tts
  1. 放好 API Key(只有你自己可读):
   mkdir -p ~/.config/volc-tts && chmod 700 ~/.config/volc-tts
   (umask 077; printf '%s' '你的 API Key' > ~/.config/volc-tts/api_key)
  1. 在 ~/.claude/settings.json 里让 Claude Code 加载它(多个 mod 用 : 分隔):
   {
     "env": {
       "CLAUDE_CODE_PLUGIN_DIRS": "/Users/<你>/Projects/claude-mods/volc-tts",
       "CLAUDE_CODE_PLUGIN_DIR_WATCH": "1"
     }
   }

CLAUDE_CODE_PLUGIN_DIR_WATCH 可选:开了以后改代码会自动重载。桌面 App 的会话默认不监视这个目录。

  1. 新开一个会话,等 Claude 回复完,最终回复下面就会出现 🔊。

换音色

点朗读时那排按钮最后的 🎙,或者在输入框输入 /voice,打开音色面板:

┌ 🎙 朗读音色 ─────────────────────────────┐
│ 当前音色  [ 知性女声 2.0            ▾ ]  │  选中即切换,下一次朗读生效
│           🔊 试听   🗑 删除               │
│ 添加音色                                  │
│ 音色 ID   [ ICL_uranus_zh_female_…     ]  │
│ 名字      [ 性感魅惑 ]         (可选)   │
│           [ 验证并保存 ]                  │
└───────────────────────────────────────────┘
  • 音色 ID 在火山控制台的 音色库 里找,例如 zh_female_zhixingnv_uranus_bigtts、ICL_uranus_zh_female_xingganmeihuo_tob。
  • 验证就是真的合成一句「你好,我是 xxx。」(约 10 个字,会计费):先用 seed-tts-2.0,火山回复"音色和资源对不上"(错误码 55000000)时再试 seed-icl-2.0(声音复刻音色);都不行就判定不存在或你的账号没开通,不会保存。
  • 验证通过的音色存进 ~/.config/volc-tts/voices.json,所有会话共用;验证时合成的那句直接写进缓存,所以紧接着的试听不再计费。
  • 正在念的时候换音色,当前这段不变,下一次朗读用新音色;想马上听,点 ⟲。

命令行也能管:python3 bin/voices.py list | use <ID> | remove <ID> | add <ID> [名字]。

工作原理

🔊 按钮(hooks/register.tsx,运行在 Claude Code 里)
  │ $.process.spawn,把整理好的文字从 stdin 交给
  ▼
bin/stream.py ──HTTP Chunked──▶ 火山 /api/v3/tts/unidirectional
  │ 收到一块 PCM 就写一块
  ▼
mpv(开着 IPC 遥控口 ~/.config/volc-tts/mpv.sock)
  ▲
  │ pause / resume / stop / speed
bin/ctl.py ◀── ⏸ ▶️ ⏹ 倍速 按钮
  • mod 不能直接收发流式数据($.http.fetch 要等全部内容收完才返回),所以"边收边播"交给一个本地 Python 脚本和 mpv。
  • stream.py 订阅 mpv 的暂停、倍速和播放位置,在 stdout 报告 STATE playing|paused、SPEED x、SECTION 2/5,mod 据此重画按钮和状态栏。
  • 按节合成,语气连贯:每节单独请求、单独缓存;同一次朗读的所有请求带同一个 section_id,由一个下载线程按顺序一次发一个,服务端据此记住前文,下一节接着前面的语气念(仅 2.0 音色和声音复刻 2.0 音色支持)。正在念第 k 节时预取到第 k+2 节(火山流式接口大约按实时速度返回,不预取的话跳过去要等)。跳节时停掉当前 mpv,从目标节的音频重新开一个,因为 mpv 读的是管道,不能往前跳。
  • 只给最终回复加按钮:用 turn.complete 事件里的 answer 认出每轮的最终文字;会话启动时用 $.session.messages() 补上历史回复。
  • 朗读前会去掉代码块、图片和链接网址;剩下的 Markdown 符号交给火山服务端过滤(disable_markdown_filter)。
  • bin/ctl.py 在终端里也能用:python3 bin/ctl.py pause、python3 bin/ctl.py next、python3 bin/ctl.py speed 1.25。

排查问题

每次朗读和添加音色的过程都记在 ~/.cache/volc-tts/stream.log(音色、请求了几段、首声时间、错误原因),超过 1MB 自动轮换成 stream.log.1:

tail -20 ~/.cache/volc-tts/stream.log

已知局限

  • 仅 macOS。
  • "切走会话自动暂停"依赖 Claude 桌面 App 未公开的日志格式(~/Library/Logs/Claude/main.log 里的 LocalSessions.setFocusedSession:)。App 更新后可能失效,表现为切走后继续念,不会报错。
  • mod 接口仍处于早期阶段,Anthropic 说明它可能在版本之间变化,本插件随时可能需要跟着调整。
  • 火山按合成的字数计费,Markdown 符号也算字数;用过的内容走本地缓存,不会重复计费。

文件

文件作用
.claude-plugin/plugin.json插件清单
hooks/hooks.json声明 hooks 模块
hooks/register.tsx按钮、状态、事件钩子
types/index.d.ts$.state 里存的值的类型
bin/stream.py流式合成 + mpv 播放 + 切走暂停 + 缓存
bin/volc.py火山接口、缓存路径、音色列表(两个脚本共用)
bin/voices.py音色列表的查看、切换、删除、验证并添加
hooks/speech.ts把回复整理成朗读文字(代码块略过、表格转句子)
hooks/speech_test.ts上面那个的单元测试:deno test hooks/speech_test.ts
hooks/voice-pane.test.tsx界面冒烟测试:claude plugin test .
bin/ctl.py遥控正在进行的朗读
tsconfig.json编辑器类型检查用;它引用的 .claude-plugin/types/ 由引擎生成、不在仓库里,在 Claude Code 里运行 /plugin-types 即可生成

许可

MIT

Source 3 files
hooks/register.tsx 497 lines
1import { atom, read, update } from 'claude-code'
2import type { Register } from 'claude-code'
3
4import type { TtsNow, TtsVoiceList } from '../types'
5
6import { toSpeech } from './speech'
7
8// 火山引擎豆包语音:单向流式合成(HTTP Chunked)
9const ENDPOINT = 'https://openspeech.bytedance.com/api/v3/tts/unidirectional'
10const DEFAULT_VOICE = { id: 'zh_female_zhixingnv_uranus_bigtts', resource: 'seed-tts-2.0' } // 知性女声 2.0
11const VOICES_FILE = '.config/volc-tts/voices.json' // 相对 $HOME,音色列表,由 bin/voices.py 维护
12const KEY_FILE = '.config/volc-tts/api_key' // 相对 $HOME,权限 600
13const SPEED_FILE = '.config/volc-tts/speed' // 相对 $HOME,记住上次的倍速
14const CHUNK_CHARS = 800 // 后备方案每次请求的最大字数,长回复切段依次合成
15const FIRST_CHARS = 60 // 后备方案第一段的字数上限
16const PYTHON = '/usr/bin/python3'
17const SPEEDS = [1, 1.25, 1.5, 2] // 倍速按钮依次循环
18const VOICE_PANE = 'tts-voice'
19const SAMPLE_KEY = 'voice-sample' // 试听用的朗读,不属于任何一段回复
20
21// 会话级状态:哪一段在念、什么状态;当前倍速。写入会让读它的按钮重画
22const now = atom({ plugin: 'tts', key: 'now' } as const, null)
23const speed = atom({ plugin: 'tts', key: 'speed' } as const, 1)
24const finals = atom({ plugin: 'tts', key: 'finals' } as const, [])
25const FINALS_MAX = 200 // 只记最近这么多轮
26const voices = atom({ plugin: 'tts', key: 'voices' } as const, null)
27const draftId = atom({ plugin: 'tts', key: 'draftId' } as const, '')
28const draftName = atom({ plugin: 'tts', key: 'draftName' } as const, '')
29const voiceMsg = atom({ plugin: 'tts', key: 'voiceMsg' } as const, '')
30const voiceBusy = atom({ plugin: 'tts', key: 'voiceBusy' } as const, false)
31
32// 每次开始朗读加一;旧的朗读循环发现自己不是最新一轮,就不再改状态
33let run = 0
34let buffered: AbortController | null = null // 后备模式的停止开关
35
36// 一段回复的指纹:同一段文字永远得到同一个 key
37function fingerprint(text: string): string {
38  let h = 5381
39  for (let i = 0; i < text.length; i++) h = ((h << 5) + h + text.charCodeAt(i)) | 0
40  return `${(h >>> 0).toString(36)}-${text.length}`
41}
42
43// 比对前去掉多余空白:answer 和文字块的换行、缩进可能不完全一样
44function normalize(text: string): string {
45  return text.replace(/\s+/g, ' ').trim()
46}
47
48// 这个文字块是不是某一轮的最终回复。最终回复若被拆成几块,answer 是拼起来的全文,所以认“结尾那块”
49function isFinal(list: readonly string[], text: string): boolean {
50  const t = normalize(text)
51  return t.length > 0 && list.some(answer => answer.endsWith(t))
52}
53
54function speedLabel(x: number): string {
55  return `${x}×`
56}
57
58// 先切成句子,再装箱:第一段不超过 FIRST_CHARS 字(尽快出声),之后每段不超过 CHUNK_CHARS 字
59function split(text: string): string[] {
60  const sentences = text.split(/(?<=[。!?;.!?;\n])/).filter(s => s.trim())
61  const parts: string[] = []
62  let buf = ''
63  for (const s of sentences) {
64    const limit = parts.length ? CHUNK_CHARS : FIRST_CHARS
65    if (buf && buf.length + s.length > limit) {
66      parts.push(buf)
67      buf = ''
68    }
69    buf += s
70  }
71  if (buf.trim()) parts.push(buf)
72  return parts
73}
74
75// 服务端返回若干个首尾相接的 JSON 对象,每个的 data 是一小段 base64 mp3
76function joinAudio(body: string): string {
77  let bin = ''
78  let depth = 0
79  let begin = -1
80  let inStr = false
81  for (let i = 0; i < body.length; i++) {
82    const c = body[i]
83    if (inStr) {
84      if (c === '\\') i++
85      else if (c === '"') inStr = false
86      continue
87    }
88    if (c === '"') inStr = true
89    else if (c === '{') {
90      if (depth++ === 0) begin = i
91    } else if (c === '}' && --depth === 0) {
92      const o = JSON.parse(body.slice(begin, i + 1))
93      if (o.code !== 0 && o.code !== 20000000) throw new Error(`火山返回错误 ${o.code}: ${o.message}`)
94      if (o.data) bin += atob(o.data)
95    }
96  }
97  if (!bin) throw new Error('火山没有返回音频')
98  return btoa(bin)
99}
100
101const cache = new Map<string, string>() // 后备模式:文本 → mp3 base64,避免重复点击重复计费
102let apiKey: string | undefined
103
104// 后备模式用的当前音色:直接读 voices.json,读不到就用默认音色
105async function currentVoice($: any): Promise<{ id: string; resource: string }> {
106  try {
107    const home = await $.env.get('HOME')
108    const list = JSON.parse(await $.fs.read(`${home}/${VOICES_FILE}`))
109    return list.voices.find((v: { id: string }) => v.id === list.current) ?? DEFAULT_VOICE
110  } catch {
111    return DEFAULT_VOICE
112  }
113}
114
115async function synth($: any, text: string): Promise<string> {
116  const voice = await currentVoice($)
117  const hit = cache.get(`${voice.id}\n${text}`)
118  if (hit) return hit
119  if (!apiKey) {
120    const home = await $.env.get('HOME')
121    apiKey = (await $.fs.read(`${home}/${KEY_FILE}`)).trim()
122  }
123  const res = await $.http.fetch(ENDPOINT, {
124    method: 'POST',
125    headers: {
126      'X-Api-Key': apiKey!,
127      'X-Api-Resource-Id': voice.resource,
128      'X-Api-Request-Id': crypto.randomUUID(),
129      'Content-Type': 'application/json',
130    },
131    body: JSON.stringify({
132      req_params: {
133        text,
134        speaker: voice.id,
135        audio_params: { format: 'mp3', sample_rate: 24000 },
136        additions: JSON.stringify({ disable_markdown_filter: true, disable_emoji_filter: true }),
137      },
138    }),
139  })
140  if (!res.ok) throw new Error(`HTTP ${res.status}: ${res.text.slice(0, 200)}`)
141  const audio = joinAudio(res.text)
142  cache.set(`${voice.id}\n${text}`, audio)
143  return audio
144}
145
146// 改状态,同时更新状态栏:只留“🔊 火山朗读中”或“⏸ 已暂停”
147async function setNow($: any, value: TtsNow | null) {
148  await update($, now, () => value)
149  const progress = value?.section && value.section.total > 1 ? ` ${value.section.index}/${value.section.total}` : ''
150  $.ui.status(value === null ? undefined : (value.status === 'paused' ? '⏸ 已暂停' : '🔊 火山朗读中') + progress)
151}
152
153// 遥控正在进行的流式朗读:pause / resume / stop / next / prev / speed <x>
154async function ctl($: any, ...args: string[]) {
155  await $.process.run([PYTHON, `${$.plugin.root}/bin/ctl.py`, ...args])
156}
157
158async function addFinals($: any, answers: string[]) {
159  const add = answers.map(normalize).filter(Boolean)
160  if (!add.length) return
161  await update($, finals, list => [...list.filter(a => !add.includes(a)), ...add].slice(-FINALS_MAX))
162}
163
164// 历史回复:你每次提问之后、下一次提问之前,我的最后一段文字就是那一轮的最终回复
165async function loadHistoryFinals($: any) {
166  const rows = await $.session.messages()
167  const answers: string[] = []
168  let last = ''
169  for (const row of rows) {
170    const isPrompt = row.role === 'user' && row.text.trim() && !row.toolResults?.length
171    if (isPrompt) {
172      if (last) answers.push(last)
173      last = ''
174    } else if (row.role === 'assistant' && row.text.trim()) {
175      last = row.text
176    }
177  }
178  if (last) answers.push(last)
179  await addFinals($, answers)
180}
181
182async function loadSpeed($: any) {
183  try {
184    const home = await $.env.get('HOME')
185    const saved = Number((await $.fs.read(`${home}/${SPEED_FILE}`)).trim())
186    if (SPEEDS.includes(saved)) await update($, speed, () => saved)
187  } catch {
188    // 还没调过倍速
189  }
190}
191
192// 结束本会话里正在进行的朗读,按钮回到 🔊
193async function stopCurrent($: any) {
194  const cur = await read($, now)
195  run++
196  if (cur?.mode === 'buffered') buffered?.abort()
197  else if (cur) await ctl($, 'stop')
198  if (cur) await setNow($, null)
199}
200
201// 后备方案:整段合成完再播(mpv 或 Python 不可用时),只能停止,不能暂停
202async function speakBuffered($: any, text: string, key: string, id: number) {
203  const ctrl = new AbortController()
204  buffered = ctrl
205  await setNow($, { key, status: 'playing', mode: 'buffered' })
206  const parts = split(text)
207  let next = synth($, parts[0]) // 边播当前段边合成下一段
208  for (let i = 0; i < parts.length && !ctrl.signal.aborted && id === run; i++) {
209    const audio = await next
210    if (i + 1 < parts.length) next = synth($, parts[i + 1])
211    await $.audio.play({ base64: audio, mime: 'audio/mpeg' }, { signal: ctrl.signal })
212  }
213}
214
215// 主方案:bin/stream.py 流式合成,mpv 边收边播;它在 stdout 报告状态(见 stream.py 顶部)
216async function speakStreaming($: any, text: string, key: string, id: number) {
217  const child = $.process.spawn({
218    argv: [PYTHON, `${$.plugin.root}/bin/stream.py`],
219    input: text,
220  })
221  let out = ''
222  let stderr = ''
223  let started = false
224  for await (const piece of child) {
225    if (id !== run) break // 已被新的一轮取代;离开循环引擎会结束子进程
226    if (piece.stream === 'stderr') {
227      stderr += piece.text
228      continue
229    }
230    out += piece.text
231    let nl: number
232    while ((nl = out.indexOf('\n')) >= 0) {
233      const line = out.slice(0, nl)
234      out = out.slice(nl + 1)
235      if (line === 'STATE playing' || line === 'STATE paused') {
236        started = true
237        const section = (await read($, now))?.section
238        await setNow($, { key, status: line === 'STATE paused' ? 'paused' : 'playing', mode: 'stream', section })
239      } else if (line.startsWith('SECTION ')) {
240        const [index, total] = line.slice(8).split('/').map(Number)
241        const cur = await read($, now)
242        if (cur?.key === key && index && total) await setNow($, { ...cur, section: { index, total } })
243      } else if (line.startsWith('SPEED ')) {
244        const x = Number(line.slice(6))
245        if (x) await update($, speed, () => x)
246      }
247    }
248  }
249  if (!started && /Traceback|No such file/.test(stderr)) throw new Error(stderr.slice(-300))
250  const failed = stderr.match(/火山朗读失败:.*/)
251  if (failed && id === run) $.ui.toast(failed[0])
252}
253
254async function start($: any, markdown: string, key: string) {
255  const text = toSpeech(markdown)
256  if (!text) return
257  await stopCurrent($)
258  const id = ++run
259  await setNow($, { key, status: 'loading', mode: 'stream' })
260  try {
261    try {
262      await speakStreaming($, text, key, id)
263    } catch (err) {
264      // Python 或 mpv 起不来,退回整段合成
265      if (id !== run) return
266      $.ui.log(`tts: 流式朗读不可用,改用后备模式:${(err as Error).message}`, { to: 'debug' })
267      await speakBuffered($, text, key, id)
268    }
269  } catch (err) {
270    if (id === run && !buffered?.signal.aborted) $.ui.toast(`火山朗读失败:${(err as Error).message}`)
271  } finally {
272    if (id === run) await setNow($, null)
273  }
274}
275
276// 主按钮:🔊 开始 → ⏸ 暂停 → ▶️ 继续;后备模式下是 ⏹ 停止
277async function pressMain($: any, markdown: string, key: string) {
278  const cur = await read($, now)
279  if (!cur || cur.key !== key) return start($, markdown, key)
280  if (cur.mode === 'buffered') return stopCurrent($)
281  // 先改按钮(立刻有反馈),stream.py 随后会报告真实状态
282  if (cur.status === 'paused') {
283    await setNow($, { ...cur, status: 'playing' })
284    await ctl($, 'resume')
285  } else {
286    await setNow($, { ...cur, status: 'paused' })
287    await ctl($, 'pause')
288  }
289}
290
291async function pressSpeed($: any) {
292  const cur = await read($, speed)
293  const x = SPEEDS[(SPEEDS.indexOf(cur) + 1) % SPEEDS.length]
294  await update($, speed, () => x)
295  await ctl($, 'speed', String(x))
296}
297
298// 跑 bin/voices.py,它输出一行 JSON(含最新的音色列表),顺手刷新面板
299async function voicesCmd($: any, ...args: string[]) {
300  const { stdout, stderr } = await $.process.run([PYTHON, `${$.plugin.root}/bin/voices.py`, ...args], {
301    timeoutMs: 60000,
302  })
303  try {
304    const res = JSON.parse(stdout.trim().split('\n').pop() || '{}')
305    if (res.voices) await update($, voices, () => ({ current: res.current, voices: res.voices }))
306    return res
307  } catch {
308    return { ok: false, error: (stderr || stdout).slice(-200) || 'voices.py 没有输出' }
309  }
310}
311
312async function openVoicePane($: any) {
313  await update($, voiceMsg, () => '')
314  await voicesCmd($, 'list')
315  await $.ui.open({ id: VOICE_PANE, title: '🎙 朗读音色', focus: true })
316}
317
318function voiceName(list: TtsVoiceList | null, id: string): string {
319  return list?.voices.find(v => v.id === id)?.name ?? id
320}
321
322async function pickVoice($: any, id: string) {
323  const res = await voicesCmd($, 'use', id)
324  await update($, voiceMsg, () =>
325    res.ok ? `✅ 已切换为「${voiceName(res, id)}」,下一次朗读生效` : `❌ ${res.error}`,
326  )
327}
328
329async function previewVoice($: any) {
330  const list = await read($, voices)
331  if (!list) return
332  await start($, `你好,我是${voiceName(list, list.current)}。`, SAMPLE_KEY)
333}
334
335async function removeVoice($: any) {
336  const list = await read($, voices)
337  if (!list) return
338  const name = voiceName(list, list.current)
339  const res = await voicesCmd($, 'remove', list.current)
340  await update($, voiceMsg, () => (res.ok ? `🗑 已删除「${name}」` : `❌ ${res.error}`))
341}
342
343// 验证并保存:voices.py 真的合成一句试听语,成功才保存;随后播放这句(已在缓存里,不再计费)
344async function addVoice($: any) {
345  if (await read($, voiceBusy)) return
346  const id = (await read($, draftId)).trim()
347  const name = (await read($, draftName)).trim()
348  if (!id) {
349    await update($, voiceMsg, () => '请先填音色 ID')
350    return
351  }
352  await update($, voiceBusy, () => true)
353  await update($, voiceMsg, () => '⏳ 正在向火山验证…')
354  try {
355    const res = await voicesCmd($, 'add', id, name)
356    if (!res.ok) {
357      await update($, voiceMsg, () => `❌ ${res.error}`)
358      return
359    }
360    await update($, draftId, () => '')
361    await update($, draftName, () => '')
362    await update($, voiceMsg, () => `✅ 已添加「${res.added.name}」,并切换为当前音色`)
363    void start($, res.sample, SAMPLE_KEY)
364  } finally {
365    await update($, voiceBusy, () => false)
366  }
367}
368
369export const register: Register = on => {
370  on('session.start', async ($, e, next) => {
371    const started = await next(e)
372    // 重新加载后没有朗读循环了,旧状态作废
373    await update($, now, () => null)
374    $.ui.status(undefined)
375    await loadSpeed($)
376    await loadHistoryFinals($)
377    await $.command.register({ name: 'voice', description: '打开朗读音色面板:切换、试听、添加火山音色' })
378    return started
379  })
380
381  on('command.run', { command: 'voice' }, async $ => {
382    await openVoicePane($)
383    return { text: '已打开朗读音色面板。' }
384  })
385
386  on('ui.render', { component: 'Pane', requestId: VOICE_PANE }, async ($, e) => {
387    const { Box, Text, Button, Select, Input } = $.ui.resolve(e)
388    const list = await read($, voices)
389    const busy = await read($, voiceBusy)
390    const msg = await read($, voiceMsg)
391    return (
392      <Box flexDirection="column" gap={1}>
393        {list && (
394          <Box flexDirection="column">
395            <Select
396              key="voice-pick"
397              label="当前音色 "
398              options={list.voices.map(v => ({ value: v.id, label: v.name }))}
399              value={list.current}
400              onSelect={id => void pickVoice($, id)}
401            />
402            <Box flexDirection="row" gap={2}>
403              <Button key="voice-preview" label="🔊 试听" plain onPress={() => void previewVoice($)} />
404              <Button key="voice-remove" label="🗑 删除" plain dimColor onPress={() => void removeVoice($)} />
405            </Box>
406          </Box>
407        )}
408        <Box flexDirection="column">
409          <Text bold>添加音色</Text>
410          <Input
411            key="voice-id"
412            label="音色 ID "
413            placeholder="例如 ICL_uranus_zh_female_xingganmeihuo_tob"
414            value={await read($, draftId)}
415            submitLabel="验证"
416            onInput={v => void update($, draftId, () => v)}
417            onSubmit={v => void update($, draftId, () => v).then(() => addVoice($))}
418          />
419          <Input
420            key="voice-name"
421            label="名字   "
422            placeholder="可选,不填就用 ID"
423            value={await read($, draftName)}
424            submitLabel="验证"
425            onInput={v => void update($, draftName, () => v)}
426            onSubmit={v => void update($, draftName, () => v).then(() => addVoice($))}
427          />
428          <Button key="voice-add" label={busy ? '⏳ 验证中…' : '验证并保存'} onPress={() => void addVoice($)} />
429        </Box>
430        {msg ? <Text>{msg}</Text> : null}
431        <Text dimColor>音色 ID 在火山控制台「音色库」里找;验证会真的合成一句试听语(约 10 个字计费)。</Text>
432      </Box>
433    )
434  })
435
436  // 一轮结束:把这轮的最终回复记进名单,它下面才出现 🔊(子代理的回合不算)
437  on('turn.complete', async ($, e, next) => {
438    const done = await next(e)
439    if (!e.agentId && e.answer.trim()) await addFinals($, [e.answer])
440    return done
441  })
442
443  on('ui.render', { component: 'AssistantMessage' }, async ($, e, next) => {
444    const drawn = await next(e)
445    const text = e.props.text
446    if (!text.trim()) return drawn
447    // 过程中的说明文字不加按钮,只有一轮的最终回复才有
448    if (!isFinal(await read($, finals), text)) return drawn
449    const key = fingerprint(text)
450    const cur = await read($, now)
451    const { Box, Button } = $.ui.resolve(e)
452
453    if (!cur || cur.key !== key) {
454      return (
455        <Box flexDirection="column">
456          {drawn}
457          <Button key="tts-main" label="🔊" plain dimColor onPress={() => void pressMain($, text, key)} />
458        </Box>
459      )
460    }
461
462    if (cur.mode === 'buffered') {
463      return (
464        <Box flexDirection="column">
465          {drawn}
466          <Button key="tts-main" label="⏹" plain onPress={() => void pressMain($, text, key)} />
467        </Box>
468      )
469    }
470
471    const x = await read($, speed)
472    return (
473      <Box flexDirection="column">
474        {drawn}
475        <Box flexDirection="row" gap={2}>
476          <Button
477            key="tts-main"
478            label={cur.status === 'paused' ? '▶️' : '⏸'}
479            plain
480            onPress={() => void pressMain($, text, key)}
481          />
482          {(cur.section?.total ?? 1) > 1 && (
483            <Button key="tts-prev" label="⏮" plain dimColor onPress={() => void ctl($, 'prev')} />
484          )}
485          {(cur.section?.total ?? 1) > 1 && (
486            <Button key="tts-next" label="⏭" plain dimColor onPress={() => void ctl($, 'next')} />
487          )}
488          <Button key="tts-restart" label="⟲" plain dimColor onPress={() => void start($, text, key)} />
489          <Button key="tts-stop" label="⏹" plain dimColor onPress={() => void stopCurrent($)} />
490          <Button key="tts-speed" label={speedLabel(x)} plain dimColor onPress={() => void pressSpeed($)} />
491          <Button key="tts-voice" label="🎙" plain dimColor onPress={() => void openVoicePane($)} />
492        </Box>
493      </Box>
494    )
495  })
496}
497
hooks/speech.ts 76 lines
1// 把 Claude 的 Markdown 回复整理成适合朗读的文字。纯函数,不碰 $,可单独测试:deno test hooks/speech_test.ts
2//
3// - 代码块略过(念成“这里有一段代码,略过”)
4// - 表格逐行念:每行一句,第一格当这一行的名字,其余格念成“表头:内容”
5// - 标题、分隔线、粗体行原样留着,stream.py 靠它们切节
6// - 剩下的 Markdown 符号(** ` # 等)交给火山服务端过滤
7
8const TABLE_ROW = /^[ \t]*\|.*\|[ \t]*$/
9const TABLE_RULE = /^[ \t]*\|?[ \t]*:?-{2,}:?[ \t]*(?:\|[ \t]*:?-{2,}:?[ \t]*)*\|?[ \t]*$/
10
11function cells(row: string): string[] {
12  return row
13    .trim()
14    .replace(/^\|/, '')
15    .replace(/\|$/, '')
16    .split('|')
17    .map(c => c.trim())
18}
19
20// 去掉格子里只给眼睛看的符号:粗体星号、反引号、行内链接网址
21function clean(cell: string): string {
22  return cell
23    .replace(/\*\*([^*]+)\*\*/g, '$1')
24    .replace(/`([^`]+)`/g, '$1')
25    .replace(/\[([^\]]+)\]\([^)]*\)/g, '$1')
26    .trim()
27}
28
29function endSentence(s: string): string {
30  return /[。!?.!?]$/.test(s) ? s : `${s}。`
31}
32
33export function tableToSpeech(lines: string[]): string {
34  const hasHeader = lines.length > 1 && TABLE_RULE.test(lines[1])
35  const header = hasHeader ? cells(lines[0]).map(clean) : []
36  const body = (hasHeader ? lines.slice(2) : lines).filter(l => !TABLE_RULE.test(l))
37  const sentences: string[] = []
38  for (const row of body) {
39    const [first, ...rest] = cells(row).map(clean)
40    const parts = first ? [first] : []
41    rest.forEach((value, i) => {
42      if (!value) return
43      const name = header[i + 1]
44      parts.push(name ? `${name}:${value}` : value)
45    })
46    if (parts.length) sentences.push(endSentence(parts.join(',')))
47  }
48  return sentences.join('\n')
49}
50
51export function toSpeech(md: string): string {
52  const out: string[] = []
53  let table: string[] = []
54  const flushTable = () => {
55    if (table.length) out.push(tableToSpeech(table))
56    table = []
57  }
58  const noCode = md.replace(/```[\s\S]*?```/g, '(这里有一段代码,略过)')
59  for (const line of noCode.split('\n')) {
60    if (TABLE_ROW.test(line) || (table.length && TABLE_RULE.test(line))) {
61      table.push(line)
62      continue
63    }
64    flushTable()
65    out.push(line)
66  }
67  flushTable()
68  return out
69    .join('\n')
70    .replace(/!\[[^\]]*\]\([^)]*\)/g, '')
71    .replace(/\[([^\]]+)\]\([^)]*\)/g, '$1')
72    .replace(/https?:\/\/\S+/g, '链接')
73    .replace(/\n{3,}/g, '\n\n')
74    .trim()
75}
76
types/index.d.ts 41 lines
1// tts 这个 mod 在 $.state 里存的值(会话级)
2export type TtsNow = {
3  /** 正在念的那段回复的指纹(文本哈希),按钮靠它认出“是不是我这一段” */
4  key: string
5  status: 'loading' | 'playing' | 'paused'
6  /** stream:Python + mpv 流式;buffered:后备的整段合成 */
7  mode: 'stream' | 'buffered'
8  /** 第几节 / 共几节(流式模式下由 stream.py 报告),例如 2/5 */
9  section?: { index: number; total: number }
10}
11
12/** 一个保存过的音色,和 ~/.config/volc-tts/voices.json 里的一项相同 */
13export type TtsVoice = {
14  id: string
15  name: string
16  /** 验证时试出来的资源:seed-tts-2.0(官方音色)或 seed-icl-2.0(声音复刻) */
17  resource: string
18}
19
20export type TtsVoiceList = {
21  current: string
22  voices: TtsVoice[]
23}
24
25declare module 'claude-code' {
26  interface PluginState {
27    tts: {
28      now: TtsNow | null
29      speed: number
30      /** 本会话每轮的最终回复(去掉多余空白后的全文),只有它们下面才画 🔊 */
31      finals: string[]
32      /** 音色面板:音色列表(读自 voices.json)、两个输入框的草稿、提示语、是否正在验证 */
33      voices: TtsVoiceList | null
34      draftId: string
35      draftName: string
36      voiceMsg: string
37      voiceBusy: boolean
38    }
39  }
40}
41