SLOPSHOPPER

tts

在每段回复下加一个🔊按钮,用火山引擎豆包语音(知性女声 2.0)流式朗读,可暂停、继续、从头念、倍速

newrowstoaststatusprocessnetwork
A shopper browsing a rack in a slop shop
README

claude-code-volc-tts

一个 Claude Code mod:在每轮回复的最终回复下面加一个 🔊 按钮,用火山引擎豆包语音(默认「知性女声 2.0」)流式朗读。桌面 App 的 Code 标签页和终端版都能用。

没在念:   🔊
正在念:   ⏸   ⟲   ⏹   1×
已暂停:   ▶️   ⟲   ⏹   1×

功能

  • 流式朗读:边合成边播放,首声约 0.8 秒,和回复长短无关
  • 三态主按钮:🔊 开始 → ⏸ 暂停 → ▶️ 从停下的地方继续
  • ⟲ 从头念 / ⏹ 停止:只在念或暂停时出现
  • 倍速:1× → 1.25× → 1.5× → 2× 循环,音调不变,记住上次的选择
  • 只在最终回复显示:Claude 干活过程中的说明文字不加按钮
  • 全局只有一个声音:在任何会话里开始新的朗读,旧的自动结束
  • 切走会话自动暂停(仅桌面 App):切回来点 ▶️ 继续
  • 本地缓存:同一段再念不调用火山、不花钱,0.2 秒出声;保留 7 天、上限 200MB
  • 后备模式:mpv 或 Python 不可用时,退回整段合成后播放(只能停止,不能暂停)

依赖

  • macOS(用到了 /usr/bin/python3、afplay 和 Claude 桌面 App 的日志路径)
  • Claude Code 2.1.259 及以上(需要 mod / function hooks 功能)
  • mpv:brew install mpv
  • 火山引擎账号,开通「豆包语音合成大模型 2.0」,在控制台 API Key 管理 里创建一个 API Key

安装

  1. 克隆到任意位置,例如:
   git clone https://github.com/gnehiur/claude-code-volc-tts.git ~/Projects/claude-mods/volc-tts
  1. 放好 API Key(只有你自己可读):
   mkdir -p ~/.config/volc-tts && chmod 700 ~/.config/volc-tts
   (umask 077; printf '%s' '你的 API Key' > ~/.config/volc-tts/api_key)
  1. 在 ~/.claude/settings.json 里让 Claude Code 加载它(多个 mod 用 : 分隔):
   {
     "env": {
       "CLAUDE_CODE_PLUGIN_DIRS": "/Users/<你>/Projects/claude-mods/volc-tts",
       "CLAUDE_CODE_PLUGIN_DIR_WATCH": "1"
     }
   }

CLAUDE_CODE_PLUGIN_DIR_WATCH 可选:开了以后改代码会自动重载。桌面 App 的会话默认不监视这个目录。

  1. 新开一个会话,等 Claude 回复完,最终回复下面就会出现 🔊。

换音色

在火山引擎控制台的 音色库 找到音色 ID(2.0 音色形如 zh_female_xxx_uranus_bigtts),改 hooks/register.tsx 顶部的 SPEAKER。

工作原理

🔊 按钮(hooks/register.tsx,运行在 Claude Code 里)
  │ $.process.spawn,把整理好的文字从 stdin 交给
  ▼
bin/stream.py ──HTTP Chunked──▶ 火山 /api/v3/tts/unidirectional
  │ 收到一块 PCM 就写一块
  ▼
mpv(开着 IPC 遥控口 ~/.config/volc-tts/mpv.sock)
  ▲
  │ pause / resume / stop / speed
bin/ctl.py ◀── ⏸ ▶️ ⏹ 倍速 按钮
  • mod 不能直接收发流式数据($.http.fetch 要等全部内容收完才返回),所以"边收边播"交给一个本地 Python 脚本和 mpv。
  • stream.py 订阅 mpv 的暂停和倍速变化,在 stdout 报告 STATE playing|paused、SPEED x,mod 据此重画按钮。
  • 只给最终回复加按钮:用 turn.complete 事件里的 answer 认出每轮的最终文字;会话启动时用 $.session.messages() 补上历史回复。
  • 朗读前会去掉代码块、图片和链接网址;剩下的 Markdown 符号交给火山服务端过滤(disable_markdown_filter)。
  • bin/ctl.py 在终端里也能用:python3 bin/ctl.py pause、python3 bin/ctl.py speed 1.25。

已知局限

  • 仅 macOS。
  • "切走会话自动暂停"依赖 Claude 桌面 App 未公开的日志格式(~/Library/Logs/Claude/main.log 里的 LocalSessions.setFocusedSession:)。App 更新后可能失效,表现为切走后继续念,不会报错。
  • mod 接口仍处于早期阶段,Anthropic 说明它可能在版本之间变化,本插件随时可能需要跟着调整。
  • 火山按合成的字数计费,Markdown 符号也算字数;用过的内容走本地缓存,不会重复计费。

文件

文件作用
.claude-plugin/plugin.json插件清单
hooks/hooks.json声明 hooks 模块
hooks/register.tsx按钮、状态、事件钩子
types/index.d.ts$.state 里存的值的类型
bin/stream.py流式合成 + mpv 播放 + 切走暂停 + 缓存
bin/ctl.py遥控正在进行的朗读
tsconfig.json编辑器类型检查用;它引用的 .claude-plugin/types/ 由引擎生成、不在仓库里,在 Claude Code 里运行 /plugin-types 即可生成

许可

MIT

Source 2 files
hooks/register.tsx 348 lines
1import { atom, read, update } from 'claude-code'
2import type { Register } from 'claude-code'
3
4import type { TtsNow } from '../types'
5
6// 火山引擎豆包语音:单向流式合成(HTTP Chunked)
7const ENDPOINT = 'https://openspeech.bytedance.com/api/v3/tts/unidirectional'
8const RESOURCE_ID = 'seed-tts-2.0'
9const SPEAKER = 'zh_female_zhixingnv_uranus_bigtts' // 知性女声 2.0
10const KEY_FILE = '.config/volc-tts/api_key' // 相对 $HOME,权限 600
11const SPEED_FILE = '.config/volc-tts/speed' // 相对 $HOME,记住上次的倍速
12const CHUNK_CHARS = 800 // 后备方案每次请求的最大字数,长回复切段依次合成
13const FIRST_CHARS = 60 // 后备方案第一段的字数上限
14const PYTHON = '/usr/bin/python3'
15const SPEEDS = [1, 1.25, 1.5, 2] // 倍速按钮依次循环
16
17// 会话级状态:哪一段在念、什么状态;当前倍速。写入会让读它的按钮重画
18const now = atom({ plugin: 'tts', key: 'now' } as const, null)
19const speed = atom({ plugin: 'tts', key: 'speed' } as const, 1)
20const finals = atom({ plugin: 'tts', key: 'finals' } as const, [])
21const FINALS_MAX = 200 // 只记最近这么多轮
22
23// 每次开始朗读加一;旧的朗读循环发现自己不是最新一轮,就不再改状态
24let run = 0
25let buffered: AbortController | null = null // 后备模式的停止开关
26
27// 把 Markdown 回复整理成适合朗读的纯文本;剩余的 Markdown 符号交给服务端过滤
28function toSpeech(md: string): string {
29  return md
30    .replace(/```[\s\S]*?```/g, '(这里有一段代码,略过)')
31    .replace(/!\[[^\]]*\]\([^)]*\)/g, '')
32    .replace(/\[([^\]]+)\]\([^)]*\)/g, '$1')
33    .replace(/https?:\/\/\S+/g, '链接')
34    .replace(/^\s*\|?[\s:|-]+\|[\s:|-]*$/gm, '')
35    .replace(/\n{3,}/g, '\n\n')
36    .trim()
37}
38
39// 一段回复的指纹:同一段文字永远得到同一个 key
40function fingerprint(text: string): string {
41  let h = 5381
42  for (let i = 0; i < text.length; i++) h = ((h << 5) + h + text.charCodeAt(i)) | 0
43  return `${(h >>> 0).toString(36)}-${text.length}`
44}
45
46// 比对前去掉多余空白:answer 和文字块的换行、缩进可能不完全一样
47function normalize(text: string): string {
48  return text.replace(/\s+/g, ' ').trim()
49}
50
51// 这个文字块是不是某一轮的最终回复。最终回复若被拆成几块,answer 是拼起来的全文,所以认“结尾那块”
52function isFinal(list: readonly string[], text: string): boolean {
53  const t = normalize(text)
54  return t.length > 0 && list.some(answer => answer.endsWith(t))
55}
56
57function speedLabel(x: number): string {
58  return `${x}×`
59}
60
61// 先切成句子,再装箱:第一段不超过 FIRST_CHARS 字(尽快出声),之后每段不超过 CHUNK_CHARS 字
62function split(text: string): string[] {
63  const sentences = text.split(/(?<=[。!?;.!?;\n])/).filter(s => s.trim())
64  const parts: string[] = []
65  let buf = ''
66  for (const s of sentences) {
67    const limit = parts.length ? CHUNK_CHARS : FIRST_CHARS
68    if (buf && buf.length + s.length > limit) {
69      parts.push(buf)
70      buf = ''
71    }
72    buf += s
73  }
74  if (buf.trim()) parts.push(buf)
75  return parts
76}
77
78// 服务端返回若干个首尾相接的 JSON 对象,每个的 data 是一小段 base64 mp3
79function joinAudio(body: string): string {
80  let bin = ''
81  let depth = 0
82  let begin = -1
83  let inStr = false
84  for (let i = 0; i < body.length; i++) {
85    const c = body[i]
86    if (inStr) {
87      if (c === '\\') i++
88      else if (c === '"') inStr = false
89      continue
90    }
91    if (c === '"') inStr = true
92    else if (c === '{') {
93      if (depth++ === 0) begin = i
94    } else if (c === '}' && --depth === 0) {
95      const o = JSON.parse(body.slice(begin, i + 1))
96      if (o.code !== 0 && o.code !== 20000000) throw new Error(`火山返回错误 ${o.code}: ${o.message}`)
97      if (o.data) bin += atob(o.data)
98    }
99  }
100  if (!bin) throw new Error('火山没有返回音频')
101  return btoa(bin)
102}
103
104const cache = new Map<string, string>() // 后备模式:文本 → mp3 base64,避免重复点击重复计费
105let apiKey: string | undefined
106
107async function synth($: any, text: string): Promise<string> {
108  const hit = cache.get(text)
109  if (hit) return hit
110  if (!apiKey) {
111    const home = await $.env.get('HOME')
112    apiKey = (await $.fs.read(`${home}/${KEY_FILE}`)).trim()
113  }
114  const res = await $.http.fetch(ENDPOINT, {
115    method: 'POST',
116    headers: {
117      'X-Api-Key': apiKey!,
118      'X-Api-Resource-Id': RESOURCE_ID,
119      'X-Api-Request-Id': crypto.randomUUID(),
120      'Content-Type': 'application/json',
121    },
122    body: JSON.stringify({
123      req_params: {
124        text,
125        speaker: SPEAKER,
126        audio_params: { format: 'mp3', sample_rate: 24000 },
127        additions: JSON.stringify({ disable_markdown_filter: true, disable_emoji_filter: true }),
128      },
129    }),
130  })
131  if (!res.ok) throw new Error(`HTTP ${res.status}: ${res.text.slice(0, 200)}`)
132  const audio = joinAudio(res.text)
133  cache.set(text, audio)
134  return audio
135}
136
137// 改状态,同时更新状态栏:只留“🔊 火山朗读中”或“⏸ 已暂停”
138async function setNow($: any, value: TtsNow | null) {
139  await update($, now, () => value)
140  $.ui.status(value === null ? undefined : value.status === 'paused' ? '⏸ 已暂停' : '🔊 火山朗读中')
141}
142
143// 遥控正在进行的流式朗读:pause / resume / stop / speed <x>
144async function ctl($: any, ...args: string[]) {
145  await $.process.run([PYTHON, `${$.plugin.root}/bin/ctl.py`, ...args])
146}
147
148async function addFinals($: any, answers: string[]) {
149  const add = answers.map(normalize).filter(Boolean)
150  if (!add.length) return
151  await update($, finals, list => [...list.filter(a => !add.includes(a)), ...add].slice(-FINALS_MAX))
152}
153
154// 历史回复:你每次提问之后、下一次提问之前,我的最后一段文字就是那一轮的最终回复
155async function loadHistoryFinals($: any) {
156  const rows = await $.session.messages()
157  const answers: string[] = []
158  let last = ''
159  for (const row of rows) {
160    const isPrompt = row.role === 'user' && row.text.trim() && !row.toolResults?.length
161    if (isPrompt) {
162      if (last) answers.push(last)
163      last = ''
164    } else if (row.role === 'assistant' && row.text.trim()) {
165      last = row.text
166    }
167  }
168  if (last) answers.push(last)
169  await addFinals($, answers)
170}
171
172async function loadSpeed($: any) {
173  try {
174    const home = await $.env.get('HOME')
175    const saved = Number((await $.fs.read(`${home}/${SPEED_FILE}`)).trim())
176    if (SPEEDS.includes(saved)) await update($, speed, () => saved)
177  } catch {
178    // 还没调过倍速
179  }
180}
181
182// 结束本会话里正在进行的朗读,按钮回到 🔊
183async function stopCurrent($: any) {
184  const cur = await read($, now)
185  run++
186  if (cur?.mode === 'buffered') buffered?.abort()
187  else if (cur) await ctl($, 'stop')
188  if (cur) await setNow($, null)
189}
190
191// 后备方案:整段合成完再播(mpv 或 Python 不可用时),只能停止,不能暂停
192async function speakBuffered($: any, text: string, key: string, id: number) {
193  const ctrl = new AbortController()
194  buffered = ctrl
195  await setNow($, { key, status: 'playing', mode: 'buffered' })
196  const parts = split(text)
197  let next = synth($, parts[0]) // 边播当前段边合成下一段
198  for (let i = 0; i < parts.length && !ctrl.signal.aborted && id === run; i++) {
199    const audio = await next
200    if (i + 1 < parts.length) next = synth($, parts[i + 1])
201    await $.audio.play({ base64: audio, mime: 'audio/mpeg' }, { signal: ctrl.signal })
202  }
203}
204
205// 主方案:bin/stream.py 流式合成,mpv 边收边播;它在 stdout 报告状态(见 stream.py 顶部)
206async function speakStreaming($: any, text: string, key: string, id: number) {
207  const child = $.process.spawn({
208    argv: [PYTHON, `${$.plugin.root}/bin/stream.py`, SPEAKER],
209    input: text,
210  })
211  let out = ''
212  let stderr = ''
213  let started = false
214  for await (const piece of child) {
215    if (id !== run) break // 已被新的一轮取代;离开循环引擎会结束子进程
216    if (piece.stream === 'stderr') {
217      stderr += piece.text
218      continue
219    }
220    out += piece.text
221    let nl: number
222    while ((nl = out.indexOf('\n')) >= 0) {
223      const line = out.slice(0, nl)
224      out = out.slice(nl + 1)
225      if (line === 'STATE playing' || line === 'STATE paused') {
226        started = true
227        await setNow($, { key, status: line === 'STATE paused' ? 'paused' : 'playing', mode: 'stream' })
228      } else if (line.startsWith('SPEED ')) {
229        const x = Number(line.slice(6))
230        if (x) await update($, speed, () => x)
231      }
232    }
233  }
234  if (!started && /Traceback|No such file/.test(stderr)) throw new Error(stderr.slice(-300))
235  const failed = stderr.match(/火山朗读失败:.*/)
236  if (failed && id === run) $.ui.toast(failed[0])
237}
238
239async function start($: any, markdown: string, key: string) {
240  const text = toSpeech(markdown)
241  if (!text) return
242  await stopCurrent($)
243  const id = ++run
244  await setNow($, { key, status: 'loading', mode: 'stream' })
245  try {
246    try {
247      await speakStreaming($, text, key, id)
248    } catch (err) {
249      // Python 或 mpv 起不来,退回整段合成
250      if (id !== run) return
251      $.ui.log(`tts: 流式朗读不可用,改用后备模式:${(err as Error).message}`, { to: 'debug' })
252      await speakBuffered($, text, key, id)
253    }
254  } catch (err) {
255    if (id === run && !buffered?.signal.aborted) $.ui.toast(`火山朗读失败:${(err as Error).message}`)
256  } finally {
257    if (id === run) await setNow($, null)
258  }
259}
260
261// 主按钮:🔊 开始 → ⏸ 暂停 → ▶️ 继续;后备模式下是 ⏹ 停止
262async function pressMain($: any, markdown: string, key: string) {
263  const cur = await read($, now)
264  if (!cur || cur.key !== key) return start($, markdown, key)
265  if (cur.mode === 'buffered') return stopCurrent($)
266  // 先改按钮(立刻有反馈),stream.py 随后会报告真实状态
267  if (cur.status === 'paused') {
268    await setNow($, { ...cur, status: 'playing' })
269    await ctl($, 'resume')
270  } else {
271    await setNow($, { ...cur, status: 'paused' })
272    await ctl($, 'pause')
273  }
274}
275
276async function pressSpeed($: any) {
277  const cur = await read($, speed)
278  const x = SPEEDS[(SPEEDS.indexOf(cur) + 1) % SPEEDS.length]
279  await update($, speed, () => x)
280  await ctl($, 'speed', String(x))
281}
282
283export const register: Register = on => {
284  on('session.start', async ($, e, next) => {
285    const started = await next(e)
286    // 重新加载后没有朗读循环了,旧状态作废
287    await update($, now, () => null)
288    $.ui.status(undefined)
289    await loadSpeed($)
290    await loadHistoryFinals($)
291    return started
292  })
293
294  // 一轮结束:把这轮的最终回复记进名单,它下面才出现 🔊(子代理的回合不算)
295  on('turn.complete', async ($, e, next) => {
296    const done = await next(e)
297    if (!e.agentId && e.answer.trim()) await addFinals($, [e.answer])
298    return done
299  })
300
301  on('ui.render', { component: 'AssistantMessage' }, async ($, e, next) => {
302    const drawn = await next(e)
303    const text = e.props.text
304    if (!text.trim()) return drawn
305    // 过程中的说明文字不加按钮,只有一轮的最终回复才有
306    if (!isFinal(await read($, finals), text)) return drawn
307    const key = fingerprint(text)
308    const cur = await read($, now)
309    const { Box, Button } = $.ui.resolve(e)
310
311    if (!cur || cur.key !== key) {
312      return (
313        <Box flexDirection="column">
314          {drawn}
315          <Button key="tts-main" label="🔊" plain dimColor onPress={() => void pressMain($, text, key)} />
316        </Box>
317      )
318    }
319
320    if (cur.mode === 'buffered') {
321      return (
322        <Box flexDirection="column">
323          {drawn}
324          <Button key="tts-main" label="⏹" plain onPress={() => void pressMain($, text, key)} />
325        </Box>
326      )
327    }
328
329    const x = await read($, speed)
330    return (
331      <Box flexDirection="column">
332        {drawn}
333        <Box flexDirection="row" gap={2}>
334          <Button
335            key="tts-main"
336            label={cur.status === 'paused' ? '▶️' : '⏸'}
337            plain
338            onPress={() => void pressMain($, text, key)}
339          />
340          <Button key="tts-restart" label="⟲" plain dimColor onPress={() => void start($, text, key)} />
341          <Button key="tts-stop" label="⏹" plain dimColor onPress={() => void stopCurrent($)} />
342          <Button key="tts-speed" label={speedLabel(x)} plain dimColor onPress={() => void pressSpeed($)} />
343        </Box>
344      </Box>
345    )
346  })
347}
348
types/index.d.ts 16 lines
1// tts 这个 mod 在 $.state 里存的值(会话级)
2export type TtsNow = {
3  /** 正在念的那段回复的指纹(文本哈希),按钮靠它认出“是不是我这一段” */
4  key: string
5  status: 'loading' | 'playing' | 'paused'
6  /** stream:Python + mpv 流式;buffered:后备的整段合成 */
7  mode: 'stream' | 'buffered'
8}
9
10declare module 'claude-code' {
11  interface PluginState {
12    /** finals:本会话每轮的最终回复(去掉多余空白后的全文),只有它们下面才画 🔊 */
13    tts: { now: TtsNow | null; speed: number; finals: string[] }
14  }
15}
16