Persona agents you can talk to: create a sidekick, it explains, does and fixes things in character, asks when it needs you, and speaks its replies.

Persona agents you can talk to inside Claude Code. A sidekick is the one driving your session: it explains, does and fixes things in character, asks you the moment it needs a decision, gives you a brief when you ask for one, and speaks its replies out loud. You can talk back with your voice.
/sidekick new Rudy, a blunt senior Rust dev who hates abstractions
🦊 Rudy is ready. "Ship it or delete it."
/sidekick talk
🎙 Talk mode on. Just speak; Rudy answers out loud and listens again.
Two sidekicks ship so it works before you create anyone: Ada (calm staff engineer) and Rudy (blunt senior dev). A sidekick never announces itself by voice; a cabin-style chime says "on" and "your turn", and a lower one says talk mode ended. The name is in the terminal header.
Thinking · Rudy is on it…, and the question dialog is headed 🦊 Rudy asks:./sidekick mute turns that off./sidekick talk) is hands-free and continuous. The sidekick greets you and listens; when you pause it sends what you said as your prompt, works, speaks the answer, and keeps listening. No key to hold. Talk over it to interrupt: the moment it hears words that aren't its own, it stops speaking and takes yours as the next prompt. Say "stop" or "wait" while it works to cancel the turn. Say "end talk" to stop. A band above the prompt shows what is happening (listening… with the words as they are recognized, speaking…, working…) with s to skip the speech, m to mute and x to end. Three long silences in a row end talk mode on their own.--- line that is shown but not read.| Command | What it does | | :- | :- | | /sidekick | Roster pane: 1–9 switch, 0 off, m mute, t talk | | /sidekick new <description> | Generates a persona (name, glyph, voice, catchphrase, character) with Haiku and makes it active | | /sidekick use <name> / /sidekick off | Switch sidekick / plain Claude | | /sidekick talk | Toggle hands-free talk mode (voice in, voice out) | | /sidekick mute / /sidekick speak | Stop / resume spoken replies for this session | | /sidekick voices / /sidekick voices install | Show which Apple voices are installed; open the pane to download natural ones | | /sidekick voice <voice> / /sidekick voice <name> <voice> | Give the active sidekick, or a named one, a voice (/sidekick voice Ava) | | /sidekick list / /sidekick rm <name> | Housekeeping |
Sidekicks and the active one persist across sessions in the plugin's store; mute and talk mode are per session, and the Speak replies option in /config sets the default.
macOS ships compact voices that sound robotic. Apple's natural Premium and Enhanced voices are a one-time download: run /sidekick voices install, or go to System Settings → Accessibility → Spoken Content → System Voice → ⓘ Manage Voices → English, and download Ava (Premium), Tom (Enhanced), or any voice marked Premium or Enhanced. Sidekicks use the best installed variant of their voice automatically (Premium, then Enhanced, then any natural voice, then the compact stand-in), so nothing else to configure.
To give a sidekick a different voice: /sidekick voice Ava for the active one, or /sidekick voice Rudy Ava for a named one. Any of the sixteen natural names works, and so does the exact name of any voice say -v ? lists. The sidekick confirms by saying its catchphrase in the new voice. The choice is saved with the persona.
say; on Linux and Windows replies stay text.hooks/listener-source.ts) that uses macOS's own speech recognition, on-device where the language model is installed. It is written to /var/tmp/sidekick/listen.swift and compiled there once with swiftc, which comes with the Xcode Command Line Tools (xcode-select --install). The first run asks for Microphone and Speech Recognition permission for your terminal. Nothing is sent to any third party; on-device recognition sends nothing anywhere./voice dictation still works alongside; the sidekick speaks its replies either way.claude plugin marketplace add Redskull-127/sidekick
claude plugin install sidekick@meer-mods
Or for one session: claude --plugin-dir /path/to/sidekick.
Before installing any mod, you can list what it hooks and calls without running it: claude plugin validate /path/to/sidekick.
Everything the plugin does is in its readable source. Spelled out, because a plugin that listens and speaks should be:
Programs it starts, and why
say speaks replies in the sidekick's voice; killall say cuts a reply short when you interrupt or press s.say -v ? lists the voices you have, so a Premium or Enhanced one is used when installed.afplay, through Claude Code's audio API, plays the two chimes synthesized in hooks/chime.ts.mkdir -p /var/tmp/sidekick and swiftc once, to compile the listener source (hooks/listener-source.ts in this repo, written to /var/tmp/sidekick/listen.swift) into /var/tmp/sidekick/listen./var/tmp/sidekick/listen --session <this session's id>, while talk mode is on, to hear you through the microphone; pkill -f on that exact command line to stop listening, so another session's listener is never touched.open x-apple.systempreferences:…SpokenContent when you run /sidekick voices install, to show the voice download pane.Files it writes
/var/tmp/sidekick/ only: listen.swift and Info.plist (the listener's source, copied from this plugin), listen (its compiled binary) and listen.version (which source it was built from). Nothing in your project, nothing in your settings, no build or startup file.~/.claude/plugins/store/), through the mods API.What it reads
/var/tmp/sidekick/listen.version, to know whether the compiled listener is current.What it sends, and where
/sidekick new <description>: that description goes to Claude (Haiku) through your own Claude Code account, to write the persona. Nothing else is ever sent to Claude by the plugin itself.Where it steps in
Where it works: this plugin is a Claude Code mod, so it runs in the Claude Code terminal and the Desktop app's Code tab. Added from claude.ai it installs, but has nothing to do in chat or Cowork.
claude plugin validate . # what the engine reads from the module
claude plugin test . # tests/sidekick.test.ts, no session or network
claude --plugin-dir . # hot-reloads on save
The first load writes .claude-plugin/types/ and a one-line tsconfig.json (both gitignored); from then on the editor and npx tsc -p . type-check against the real mods API.
A mod cannot open a microphone itself, so the sidekick spawns the listener binary, over and over, for as long as talk mode is on. Each run records until you pause for 1.4 seconds (or 60 seconds at most), streams the partial transcript to the band, prints the final text, and exits. The mod submits that text as your prompt.
The listener runs while the sidekick speaks too, so you can interrupt. Its own voice comes back through the microphone, so the mod drops the leading words that match what it was saying; the first two words that aren't its own cut the speech short. The listener captures through AVFoundation and picks the system's default microphone (/var/tmp/sidekick/listen --list-devices shows them once built; add --device <name part> to the listener's arguments in hooks/register.tsx to pin one).
--lang to the listener's arguments in hooks/register.tsx to pin one.hooks/listener-source.ts, the binary rebuilds on next use (the build is keyed to a hash of the source). The chimes are synthesized in hooks/chime.ts.hooks/register.tsx 681 lines1import type { EngineInterface, Register } from 'claude-code'
2
3import type { Persona, SidekickQuestion } from '../types'
4import { CHIME_OFF, CHIME_ON } from './chime'
5import { LISTENER_PLIST, LISTENER_SWIFT } from './listener-source'
6import { GEN_SYSTEM, PRESETS, VOICES, askAloud, bestVoice, contract, foreignWords, hasNaturalVoice, parseGenerated, parseSayVoices, pickOption, pickVoice, spoken, stripEcho } from './persona'
7
8const PANE = 'sidekick'
9// session state the drawings read; each a literal reference for `$.state`
10const roster = { plugin: 'sidekick', key: 'roster' } as const
11const active = { plugin: 'sidekick', key: 'active' } as const
12const isMuted = { plugin: 'sidekick', key: 'isMuted' } as const
13const isTalk = { plugin: 'sidekick', key: 'isTalk' } as const
14const isSpeaking = { plugin: 'sidekick', key: 'isSpeaking' } as const
15const isListening = { plugin: 'sidekick', key: 'isListening' } as const
16const heard = { plugin: 'sidekick', key: 'heard' } as const
17const question = { plugin: 'sidekick', key: 'question' } as const
18
19const USAGE = [
20 '/sidekick roster pane',
21 '/sidekick new <describe someone>',
22 '/sidekick use <name> | off',
23 '/sidekick talk hands-free: speak, it answers, it listens again',
24 '/sidekick mute | speak',
25 '/sidekick voices [install] natural Apple voices',
26 '/sidekick voice <voice> give the active sidekick a voice (or: voice <name> <voice>)',
27 '/sidekick list | rm <name>',
28].join('\n')
29
30const VOICE_STEPS = [
31 'System Settings → Accessibility → Spoken Content → System Voice → ⓘ Manage Voices → English.',
32 'Download Ava (Premium) and Tom (Enhanced), or any voice marked Premium or Enhanced. Sidekicks pick them up at once.',
33].join('\n')
34
35const END_TALK = /\b(end talk|stop talking|stop listening|that'?s all|goodbye)\b/i
36const STOP = /^(stop|wait|hold on|hang on|never ?mind|shut up|quiet)\b/i
37
38let installedVoices: string[] = []
39
40/** Which voices `say` has right now. */
41async function loadVoices($: EngineInterface) {
42 try {
43 const r = await $.process.run(['say', '-v', '?'])
44 installedVoices = parseSayVoices(r.stdout)
45 } catch {
46 installedVoices = []
47 }
48}
49
50/** Store → state, at session start and after /clear, /resume, /branch. */
51async function load($: EngineInterface, speakByDefault: boolean) {
52 const saved = (await $.store.get('personas')) as Persona[] | undefined
53 const migrated = (await $.store.get('voicesMigrated')) as boolean | undefined
54 // no list yet seeds the presets; an empty list means everyone was removed and stays empty.
55 // presets saved by 0.1 with the compact voices move to the natural ones, once; after that a chosen voice stays
56 const list = (saved ?? PRESETS).map(p => {
57 const preset = PRESETS.find(x => x.id === p.id)
58 return !migrated && preset && (p.voice === 'Samantha' || p.voice === 'Daniel') ? { ...p, voice: preset.voice } : p
59 })
60 if (!saved || !migrated) {
61 await $.store.set('personas', list)
62 await $.store.set('voicesMigrated', true)
63 }
64 const activeId = (await $.store.get('active')) as string | null | undefined
65 await $.state.set(roster, list)
66 await $.state.set(active, list.find(p => p.id === activeId) ?? null)
67 // mute and talk mode are per session; the `speak` option in /config is the default
68 await $.state.set(isMuted, !speakByDefault)
69 await $.state.set(isTalk, false)
70}
71
72async function setActive($: EngineInterface, p: Persona | null) {
73 // no sidekick, no ear: the pane's off key and rm must not leave the microphone open
74 if (!p) await setTalk($, false)
75 await $.state.set(active, p)
76 await $.store.set('active', p?.id ?? null)
77}
78
79async function toggleMute($: EngineInterface) {
80 const muted = !(((await $.state.get(isMuted)).value ?? false))
81 await $.state.set(isMuted, muted)
82 return muted
83}
84
85let speechNow = ''
86let recentSpeech: string[] = []
87let sayId = 0
88let isHushed = false
89
90// ponytail: the last three utterances are the echo vocabulary; the recognizer delivers its transcript
91// well after the speaker goes quiet, so "what is playing right now" was never the right question
92const echoText = () => recentSpeech.join(' ')
93
94/** Speaks in the persona's voice, cutting any speech still going; falls back to the default voice, then to silence. */
95async function say($: EngineInterface, text: string, voice: string) {
96 if (!text) return
97 if (((await $.state.get(isSpeaking)).value ?? false)) await hush($)
98 const mine = ++sayId
99 isHushed = false
100 speechNow = text
101 recentSpeech = [...recentSpeech.slice(-2), text]
102 await $.state.set(isSpeaking, true)
103 // until a natural voice is there, look again before each reply: one downloaded mid-session is used at once
104 if (!hasNaturalVoice(installedVoices)) await loadVoices($)
105 try {
106 await $.audio.speak(text, { voice: bestVoice(voice, installedVoices) })
107 } catch {
108 if (!isHushed) {
109 try {
110 await $.audio.speak(text)
111 } catch {
112 // no synthesizer on this platform: text only
113 }
114 }
115 }
116 if (sayId !== mine) return
117 speechNow = ''
118 await $.state.set(isSpeaking, false)
119 if (!isHushed && (((await $.state.get(isTalk)).value ?? false))) void chime($)
120}
121
122/** Awaits a call whose failure does not matter (a chime that cannot play, a program not running). */
123async function quietly(work: Promise<unknown>) {
124 try {
125 await work
126 } catch {
127 // nothing to do
128 }
129}
130
131/** The cabin chime: "a sidekick is on" and "your turn"; the lower one when talk mode ends. */
132async function chime($: EngineInterface, kind: 'on' | 'off' = 'on') {
133 await quietly($.audio.play({ base64: kind === 'off' ? CHIME_OFF : CHIME_ON, mime: 'audio/wav' }))
134}
135
136// ponytail: `say` has no abort; killing the macOS synthesizer is the one-line skip
137async function hush($: EngineInterface) {
138 isHushed = true
139 await quietly($.process.run(['killall', 'say']))
140}
141
142// ── the listener: macOS speech recognition in a small native binary, built once ─────────────────
143
144/** A short stable hash of the listener's source, so a changed source gets a fresh build. */
145function sourceVersion(source: string): string {
146 let hash = 5381
147 for (let i = 0; i < source.length; i++) hash = ((hash * 33) ^ source.charCodeAt(i)) >>> 0
148 return hash.toString(16)
149}
150
151// the listener is built and run at fixed paths under /var/tmp/sidekick, outside any project;
152// every path below is written out in full at the call, so a reader can see each one
153
154// why the listener could not be built, for the reply to /sidekick talk
155let listenerProblem = ''
156
157/** Where the compiled listener lives, building it on first use from the source in listener-source.ts; null when it cannot be built here. */
158async function listenerPath($: EngineInterface): Promise<string | null> {
159 const version = sourceVersion(LISTENER_SWIFT)
160 let built = ''
161 try {
162 built = await $.fs.read('/var/tmp/sidekick/listen.version')
163 } catch {
164 built = ''
165 }
166 if (built.trim() === version && (await $.fs.exists('/var/tmp/sidekick/listen'))) return '/var/tmp/sidekick/listen'
167 $.ui.toast('Building the sidekick listener, one time, about 20 seconds…')
168 try {
169 await $.process.run(['mkdir', '-p', '/var/tmp/sidekick'])
170 await $.fs.write('/var/tmp/sidekick/listen.swift', LISTENER_SWIFT)
171 await $.fs.write('/var/tmp/sidekick/Info.plist', LISTENER_PLIST)
172 const compiled = await $.process.run(
173 ['swiftc', '-O', '/var/tmp/sidekick/listen.swift', '-o', '/var/tmp/sidekick/listen', '-Xlinker', '-sectcreate', '-Xlinker', '__TEXT', '-Xlinker', '__info_plist', '-Xlinker', '/var/tmp/sidekick/Info.plist'],
174 { timeoutMs: 240_000 },
175 )
176 if (compiled.exitCode !== 0) {
177 // /usr/bin/swiftc exists on every Mac; without the Command Line Tools it exits non-zero and says so
178 const why = compiled.stderr.split('\n').find(l => l.includes('error:')) ?? compiled.stderr.slice(0, 200)
179 listenerProblem = /xcode-select|developer tools|command line tools/i.test(compiled.stderr)
180 ? 'Hands-free needs the Xcode Command Line Tools. Run xcode-select --install, then /sidekick talk again.'
181 : `The listener did not build: ${why}`
182 return null
183 }
184 await $.fs.write('/var/tmp/sidekick/listen.version', version)
185 return '/var/tmp/sidekick/listen'
186 } catch {
187 listenerProblem = 'Hands-free needs macOS: the listener uses Apple speech recognition. Replies stay text here.'
188 return null
189 }
190}
191
192type Heard = { text: string | null; raw: string; error: string; isCancelled: boolean; isFatal: boolean; isBroken: boolean; isEcho: boolean }
193
194let listening: { bin: string; isCancelled: boolean } | null = null
195// this session's id tags its listener, so stopping here never stops another session's
196let sessionTag = ''
197let earLoop: Promise<void> | null = null
198let earPaused = false
199let runningTurn: string | null = null
200
201/** Stops the listener now, from a key press or because talk mode ended. */
202async function stopListening($: EngineInterface) {
203 const current = listening
204 if (!current) return
205 current.isCancelled = true
206 // ponytail: a pending stream read cannot be interrupted from here, so the child is ended by name
207 await quietly($.process.run(['pkill', '-f', `/var/tmp/sidekick/listen .*--session ${sessionTag}`]))
208}
209
210/**
211 * Records one utterance and returns its text. The sidekick's own voice coming back through the
212 * microphone is filtered out by its words; the first words that are not its own cut the speech short.
213 */
214async function listen($: EngineInterface, bin: string): Promise<Heard> {
215 if (listening) await stopListening($)
216 const run = { bin, isCancelled: false }
217 listening = run
218 await $.state.set(heard, '')
219 await $.state.set(isListening, true)
220 let out = ''
221 let error = ''
222 let code: number | null = null
223 let bargedIn = false
224 try {
225 const child = $.process.spawn({ argv: ['/var/tmp/sidekick/listen', '--silence', '1.4', '--max', '60', '--session', sessionTag] })
226 for await (const { stream, text } of child) {
227 if (stream === 'stdout') {
228 out += text
229 continue
230 }
231 for (const line of text.split('\n')) {
232 if (line.startsWith('partial:')) {
233 const own = stripEcho(line.slice(8), echoText())
234 await $.state.set(heard, own)
235 // two words that are not in its own speech mean the human is talking over it
236 if (speechNow && !bargedIn && foreignWords(line.slice(8), echoText()) >= 2) {
237 bargedIn = true
238 void hush($)
239 }
240 } else if (line.startsWith('error:')) error = line.slice(6).trim()
241 }
242 }
243 code = (await child.result).code
244 } catch {
245 code = 5
246 }
247 if (listening === run) listening = null
248 await $.state.set(isListening, false)
249 const raw = out.trim()
250 const said = stripEcho(raw, echoText())
251 return {
252 text: run.isCancelled || !said ? null : said,
253 raw,
254 error,
255 isCancelled: run.isCancelled,
256 isFatal: code !== null && code >= 3,
257 // quit without hearing anything and without being asked to: an error, or killed from outside
258 isBroken: !run.isCancelled && !raw && (code === 2 || code === null),
259 isEcho: raw.length > 0 && !said,
260 }
261}
262
263async function setTalk($: EngineInterface, enabled: boolean) {
264 const was = ((await $.state.get(isTalk)).value ?? false)
265 await $.state.set(isTalk, enabled)
266 if (!enabled) {
267 await stopListening($)
268 await $.state.set(question, null)
269 if (was) void chime($, 'off')
270 }
271}
272
273/** The ear: one listener at a time for as long as talk mode is on. What it hears becomes the next prompt. */
274async function runEar($: EngineInterface) {
275 const p = ((await $.state.get(active)).value ?? null)
276 if (!p) return
277 const bin = await listenerPath($)
278 if (!bin) return setTalk($, false)
279 let quiet = 0
280 let failed = 0
281 while (((await $.state.get(isTalk)).value ?? false)) {
282 if (earPaused) {
283 await $.clock.sleep(200)
284 continue
285 }
286 const r = await listen($, bin)
287 if (r.isCancelled) {
288 if (!(((await $.state.get(isTalk)).value ?? false))) return
289 continue
290 }
291 if (r.isFatal) {
292 await setTalk($, false)
293 $.ui.toast(`${p.name} can't listen: ${r.error || 'the listener quit'}`)
294 return
295 }
296 if (r.isBroken) {
297 // a listener that quits at once is broken, not quiet: three in a row end talk mode, with the reason
298 if (++failed >= 3) {
299 await setTalk($, false)
300 $.ui.toast(`${p.name} can't hear: ${r.error || 'the listener keeps quitting'}. /sidekick talk to retry.`)
301 return
302 }
303 await $.clock.sleep(1000)
304 continue
305 }
306 if (r.text === null) {
307 const busy = r.isEcho || runningTurn !== null || (((await $.state.get(isSpeaking)).value ?? false))
308 if (!busy && ++quiet >= 3) {
309 // ponytail: three silent rounds end talk mode rather than listening forever
310 await setTalk($, false)
311 $.ui.toast(`${p.name} stopped listening after a long silence. /sidekick talk to resume.`)
312 return
313 }
314 continue
315 }
316 quiet = 0
317 failed = 0
318 const said = r.text
319 const wasSpeaking = ((await $.state.get(isSpeaking)).value ?? false)
320 if (wasSpeaking) await hush($)
321 // "stop" is a command when something is running or when it is all that was said; "stop the dev server" is a prompt
322 const wordCount = said.split(/\s+/).length
323 if (wordCount <= 5 && END_TALK.test(said)) {
324 await setTalk($, false)
325 return
326 }
327 if (STOP.test(said) && (wordCount <= 3 || runningTurn || wasSpeaking)) {
328 if (runningTurn) await quietly($.turn.abort({ turnId: runningTurn }))
329 continue
330 }
331 void $.prompt.submit({ text: said, asUser: true })
332 }
333}
334
335/** Turns talk mode on: greets, and starts the ear once. */
336async function startTalk($: EngineInterface) {
337 const p = ((await $.state.get(active)).value ?? null)
338 if (!p) return
339 void chime($)
340 if (earLoop) return
341 earLoop = quietly(runEar($)).finally(() => {
342 earLoop = null
343 })
344}
345
346async function toggleTalk($: EngineInterface) {
347 if (((await $.state.get(isTalk)).value ?? false)) {
348 await hush($)
349 return setTalk($, false)
350 }
351 await setTalk($, true)
352 void startTalk($)
353}
354
355const findPersona = (list: Persona[], q: string) =>
356 list.find(p => p.id === q.toLowerCase() || p.name.toLowerCase() === q.toLowerCase())
357
358export const register: Register = (on, options) => {
359 const speakByDefault = options.speak !== false
360 let isHeadless = false
361
362 on('session.start', async ($, e, next) => {
363 // a `claude -p` run has no one to talk to: nothing loads, so every other hook stays out of the way
364 isHeadless = !e.isInteractive
365 if (isHeadless) return next(e)
366 sessionTag = await $.session.id()
367 await loadVoices($)
368 await load($, speakByDefault)
369 if (((await $.state.get(active)).value ?? null)) void $.ui.open({ id: PANE, title: 'Sidekick', closeOnEscape: true })
370 try {
371 await $.command.register({
372 name: 'sidekick',
373 description: 'Persona agents you can talk to: create, switch, talk hands-free, mute',
374 argumentHint: '[new <description> | use <name> | off | talk | mute | speak | voices [install] | voice [<name>] <voice> | list | rm <name>]',
375 immediate: true,
376 })
377 } catch {
378 // name taken by another plugin: the pane and hooks still work
379 }
380 return next(e)
381 })
382
383 on('classic.SessionStart', { source: ['clear', 'resume', 'fork'] }, async ($, e, next) => {
384 if (isHeadless) return next(e)
385 await load($, speakByDefault)
386 return next(e)
387 })
388
389 on('command.run', { command: 'sidekick' }, async ($, e) => {
390 const [sub = '', ...rest] = e.args.trim().split(/\s+/)
391 const arg = rest.join(' ')
392 const list: Persona[] = (await $.state.get(roster)).value ?? []
393
394 switch (sub) {
395 case '': {
396 await $.ui.open({ id: PANE, title: 'Sidekick', focus: true, closeOnEscape: true })
397 return {}
398 }
399 case 'new': {
400 if (!arg) return { text: 'Describe who you want: /sidekick new Rudy, a blunt senior Rust dev' }
401 const r = await $.model.complete({ model: 'haiku', system: GEN_SYSTEM, prompt: arg, maxTokens: 700, timeoutMs: 20000 })
402 if (!r.isAnswered) return { text: `Couldn't reach the model (${r.reason}). Try again.` }
403 const p = parseGenerated(r.text)
404 if (!p) return { text: 'The model did not return a persona I could read. Try a clearer description.' }
405 if (list.some(x => x.id === p.id) || findPersona(list, p.name)) {
406 // the name is how use, rm and voice find it, so a second Rudy is "Rudy 3", not a hidden rudy-3
407 const n = list.length + 1
408 p.id = `${p.id}-${n}`
409 p.name = `${p.name} ${n}`
410 }
411 const grown = [...list, p]
412 await $.state.set(roster, grown)
413 await $.store.set('personas', grown)
414 await setActive($, p)
415 void $.ui.open({ id: PANE, title: 'Sidekick', closeOnEscape: true })
416 void chime($)
417 return { text: `${p.glyph} ${p.name} is ready. "${p.tagline}"\n${p.name} is driving now. /sidekick talk to speak with ${p.name}.` }
418 }
419 case 'use': {
420 const p = findPersona(list, arg)
421 if (!p) return { text: `No sidekick named "${arg}". /sidekick list` }
422 await setActive($, p)
423 void $.ui.open({ id: PANE, title: 'Sidekick', closeOnEscape: true })
424 void chime($)
425 return { text: `${p.glyph} ${p.name} is driving now.` }
426 }
427 case 'off': {
428 await setActive($, null)
429 return { text: 'Sidekick off. Plain Claude is back.' }
430 }
431 case 'talk': {
432 const p = ((await $.state.get(active)).value ?? null)
433 if (!p) return { text: 'Pick a sidekick first: /sidekick use Rudy' }
434 const wasOn = ((await $.state.get(isTalk)).value ?? false)
435 if (!wasOn) {
436 // built (once) before anything is announced, so a Mac that cannot listen gets the reason as the reply
437 const bin = await listenerPath($)
438 if (!bin) return { text: `🎙 ${listenerProblem}` }
439 }
440 await toggleTalk($)
441 if (wasOn) return { text: '🎙 Talk mode off.' }
442 if (!hasNaturalVoice(installedVoices)) $.ui.toast('Voice sounds robotic? /sidekick voices install gets Apple\'s natural ones.')
443 return { text: `🎙 Talk mode on. Just speak; ${p.name} answers out loud and listens again. Talk over ${p.name} to interrupt. Say "end talk" or press x on the band to stop.` }
444 }
445 case 'voices': {
446 if (arg === 'install') {
447 await quietly($.process.run(['open', 'x-apple.systempreferences:com.apple.preference.universalaccess?SpokenContent']))
448 return { text: `Opened System Settings.\n${VOICE_STEPS}` }
449 }
450 await loadVoices($)
451 const natural = installedVoices.filter(v => /\((Premium|Enhanced)\)$/.test(v))
452 const cur = ((await $.state.get(active)).value ?? null)
453 const lines = [
454 natural.length > 0 ? `Natural voices installed: ${natural.join(', ')}` : 'No natural (Premium/Enhanced) voices installed yet, so sidekicks use the compact ones.',
455 cur ? `${cur.glyph} ${cur.name} speaks as ${bestVoice(cur.voice, installedVoices)}.` : '',
456 `Sidekicks choose from: ${VOICES.join(', ')}.`,
457 natural.length > 0 ? '' : `/sidekick voices install opens the download pane.\n${VOICE_STEPS}`,
458 ]
459 return { text: lines.filter(Boolean).join('\n') }
460 }
461 case 'voice': {
462 const cur = (await $.state.get(active)).value ?? null
463 // "voice Rudy Ava" names a sidekick; otherwise every word is the voice ("voice Bad News")
464 const named = rest.length > 1 ? findPersona(list, rest[0]!) : undefined
465 const target = named ?? cur
466 const wanted = named ? rest.slice(1).join(' ') : arg
467 if (!target) return { text: 'Pick a sidekick first: /sidekick use Rudy' }
468 if (!wanted) return { text: `${target.glyph} ${target.name} speaks as ${bestVoice(target.voice, installedVoices)}. Choose from: ${VOICES.join(', ')}, or any installed voice (/sidekick voices).` }
469 await loadVoices($)
470 const voice = pickVoice(wanted, installedVoices)
471 if (!voice) return { text: `No voice called "${wanted}". Choose from: ${VOICES.join(', ')}, or any installed voice (/sidekick voices).` }
472 const changed = { ...target, voice }
473 const grown = list.map(x => (x.id === target.id ? changed : x))
474 await $.state.set(roster, grown)
475 await $.store.set('personas', grown)
476 if (cur?.id === target.id) await $.state.set(active, changed)
477 const resolved = bestVoice(voice, installedVoices)
478 const standIn = resolved === voice || resolved.startsWith(`${voice} (`) ? '' : ` (${voice} isn't installed here, so this stands in.)`
479 if (!(await $.state.get(isMuted)).value) void say($, changed.tagline, voice)
480 return { text: `${changed.glyph} ${changed.name} now speaks as ${resolved}.${standIn}` }
481 }
482 case 'mute':
483 case 'speak': {
484 const muted = sub === 'mute'
485 await $.state.set(isMuted, muted)
486 return { text: muted ? '🔇 Replies are no longer spoken this session.' : '🔊 Replies are spoken again.' }
487 }
488 case 'list': {
489 const cur = ((await $.state.get(active)).value ?? null)
490 return { text: list.map(p => `${p.id === cur?.id ? '●' : ' '} ${p.glyph} ${p.name} — ${p.tagline}`).join('\n') }
491 }
492 case 'rm': {
493 const p = findPersona(list, arg)
494 if (!p) return { text: `No sidekick named "${arg}".` }
495 const grown = list.filter(x => x.id !== p.id)
496 await $.state.set(roster, grown)
497 await $.store.set('personas', grown)
498 if ((((await $.state.get(active)).value ?? null))?.id === p.id) await setActive($, null)
499 return { text: `Removed ${p.glyph} ${p.name}.` }
500 }
501 default:
502 return { text: USAGE }
503 }
504 })
505
506 on('prompt.compose', async ($, e, next) => {
507 const composed = await next(e)
508 const p = ((await $.state.get(active)).value ?? null)
509 if (!p) return composed
510 const talk = ((await $.state.get(isTalk)).value ?? false)
511 return { sections: [...composed.sections, { id: 'sidekick:persona', scope: 'session', text: contract(p, talk) }] }
512 })
513
514 // the ear keeps listening through a turn: what you say meanwhile is the next prompt, "stop" aborts the turn
515 on('turn.start', async ($, e, next) => {
516 runningTurn = e.turnId
517 return next(e)
518 })
519
520 on('ui.render', { component: 'AssistantMessage' }, async ($, e, next) => {
521 const p = ((await $.state.get(active)).value ?? null)
522 if (!p || !e.props.isFirstOfReply) return next(e)
523 const { Box, Text } = $.ui.resolve(e)
524 const theirs = await next(e)
525 return (
526 <Box flexDirection="column">
527 <Text bold color={p.color}>
528 {p.glyph} {p.name}
529 </Text>
530 {theirs}
531 </Box>
532 )
533 })
534
535 on('ui.render', { component: 'Spinner' }, async ($, e, next) => {
536 const p = ((await $.state.get(active)).value ?? null)
537 if (!p) return next(e)
538 return next({ ...e, props: { ...e.props, suffix: ` · ${p.name} is on it…` } })
539 })
540
541 on('ui.render', { component: 'AskUserQuestion' }, async ($, e, next) => {
542 const p = ((await $.state.get(active)).value ?? null)
543 if (!p) return next(e)
544 const { Box, Text } = $.ui.resolve(e)
545 const theirs = await next(e)
546 return (
547 <Box flexDirection="column">
548 <Text bold color={p.color}>
549 {p.glyph} {p.name} asks:
550 </Text>
551 {theirs}
552 </Box>
553 )
554 })
555
556 // in talk mode a question is asked and answered by voice; the dialog is the fallback
557 on('tool.call', { tool: 'AskUserQuestion' }, async ($, e, next) => {
558 const p = ((await $.state.get(active)).value ?? null)
559 if (!p || !(((await $.state.get(isTalk)).value ?? false))) return next(e)
560 const bin = await listenerPath($)
561 if (!bin) return next(e)
562 earPaused = true
563 await stopListening($)
564 try {
565 const answers: Record<string, string> = {}
566 for (const q of e.questions) {
567 const labels = q.options?.map(o => o.label) ?? []
568 await $.state.set(question, ({ text: q.question, options: labels }))
569 let picked: string | undefined
570 for (let attempt = 0; attempt < 2 && !picked; attempt++) {
571 // listen while asking, so an answer spoken over the question still lands
572 const hearing = listen($, bin)
573 if (!(((await $.state.get(isMuted)).value ?? false))) void say($, askAloud(q.question, labels, attempt > 0), p.voice)
574 const r = await hearing
575 await hush($)
576 if (r.isCancelled || r.isFatal) break
577 // the question named every option, so "production" alone looks like its echo: a few words naming exactly one option are the answer
578 const named = labels.filter(l => pickOption(r.raw, [l], false) === l)
579 const answer = r.text ?? (named.length === 1 && r.raw.split(/\s+/).length <= 4 ? r.raw : null)
580 if (answer) picked = labels.length > 0 ? pickOption(answer, labels, q.multiSelect) : answer
581 }
582 await $.state.set(question, null)
583 if (!picked) {
584 $.ui.toast(`${p.name} didn't catch that. Pick with a key.`)
585 return next(e)
586 }
587 answers[q.question] = picked
588 }
589 return { result: { questions: e.questions, answers } }
590 } finally {
591 earPaused = false
592 }
593 })
594
595 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
596 const p = ((await $.state.get(active)).value ?? null)
597 if (!p || e.props.hasSurvey || !(((await $.state.get(isTalk)).value ?? false))) return next(e)
598 const { Box, Button, Text } = $.ui.resolve(e)
599 const speaking = ((await $.state.get(isSpeaking)).value ?? false)
600 const hearing = ((await $.state.get(isListening)).value ?? false)
601 const partial = ((await $.state.get(heard)).value ?? '')
602 const muted = ((await $.state.get(isMuted)).value ?? false)
603 const asked = ((await $.state.get(question)).value ?? null)
604 const status = speaking
605 ? `🔊 ${p.name} is speaking… talk over to interrupt`
606 : e.props.isWorking
607 ? `${p.glyph} ${p.name} is working… say "stop" to cancel`
608 : hearing
609 ? `🎙 listening…`
610 : `${p.glyph} ${p.name} is ready`
611 return (
612 <Box flexDirection="column">
613 {asked && (
614 <Text color={p.color}>
615 ❓ {asked.text} {asked.options.map((o, i) => ` ${i + 1}) ${o}`).join('')}
616 </Text>
617 )}
618 <Box flexDirection="row" columnGap={2}>
619 <Text color={p.color}>{status}</Text>
620 {hearing && partial && <Text dimColor>{partial}</Text>}
621 {speaking && <Button key="skip" label="skip" hotkey="s" plain onPress={() => hush($)} />}
622 {!speaking && !hearing && !e.props.isWorking && (
623 <Button key="listen" label="listen" hotkey="l" plain onPress={() => void startTalk($)} />
624 )}
625 <Button key="mute" label={muted ? 'speak' : 'mute'} hotkey="m" plain onPress={() => toggleMute($)} />
626 <Button key="end" label="end talk" hotkey="x" plain onPress={() => toggleTalk($)} />
627 </Box>
628 </Box>
629 )
630 })
631
632 on('ui.render', { component: 'PromptHint' }, async ($, e, next) => {
633 const p = ((await $.state.get(active)).value ?? null)
634 if (!p || e.props.isDraft || e.props.isWorking || !(((await $.state.get(isTalk)).value ?? false))) return next(e)
635 return next({ ...e, props: { ...e.props, tail: ` · 🎙 talking with ${p.name}` } })
636 })
637
638 on('ui.render', { component: 'Pane', requestId: 'sidekick' }, async ($, e) => {
639 const { Box, Button, Text } = $.ui.resolve(e)
640 const list: Persona[] = (await $.state.get(roster)).value ?? []
641 const cur = ((await $.state.get(active)).value ?? null)
642 const muted = ((await $.state.get(isMuted)).value ?? false)
643 const talk = ((await $.state.get(isTalk)).value ?? false)
644 return (
645 <Box flexDirection="column">
646 {list.slice(0, 9).map((p, i) => (
647 <Box flexDirection="row" columnGap={1}>
648 <Button
649 key={`use-${p.id}`}
650 label={`${p.glyph} ${p.name}`}
651 hotkey={String(i + 1)}
652 plain
653 dimColor={cur?.id !== p.id}
654 onPress={() => setActive($, p)}
655 />
656 <Text dimColor>{cur?.id === p.id ? `● ${p.tagline}` : p.tagline}</Text>
657 </Box>
658 ))}
659 <Text> </Text>
660 <Box flexDirection="row" columnGap={3}>
661 <Button key="off" label="off" hotkey="0" plain onPress={() => setActive($, null)} />
662 <Button key="mute" label={muted ? 'speak' : 'mute'} hotkey="m" plain onPress={() => toggleMute($)} />
663 <Button key="talk" label={talk ? 'end talk' : 'talk'} hotkey="t" plain onPress={() => toggleTalk($)} />
664 </Box>
665 <Text> </Text>
666 <Text dimColor>/sidekick new {'<describe someone>'} to add one</Text>
667 </Box>
668 )
669 })
670
671 // the reply is spoken, then in talk mode the sidekick listens for what comes next
672 on('turn.complete', async ($, e, next) => {
673 if (!e.agentId) runningTurn = null
674 const p = ((await $.state.get(active)).value ?? null)
675 if (p && e.reason === 'answer' && !e.agentId && !(((await $.state.get(isMuted)).value ?? false))) {
676 void say($, spoken(e.answer, ((await $.state.get(isTalk)).value ?? false)), p.voice)
677 }
678 return next(e)
679 })
680}
681hooks/chime.ts 34 lines1/** One soft "ting", made on the spot: a decaying sine with a touch of octave, as a 22 kHz mono WAV. */
2export function chimeWav(freq: number, seconds = 1.1, decay = 4): string {
3 const rate = 22050
4 const n = Math.floor(rate * seconds)
5 const pcm = new Int16Array(n)
6 for (let i = 0; i < n; i++) {
7 const t = i / rate
8 const v = 0.6 * Math.exp(-decay * t) * (Math.sin(2 * Math.PI * freq * t) + 0.2 * Math.sin(4 * Math.PI * freq * t)) / 1.2
9 pcm[i] = Math.round(Math.max(-1, Math.min(1, v)) * 32767)
10 }
11 const bytes = new Uint8Array(44 + n * 2)
12 const view = new DataView(bytes.buffer)
13 const ascii = (at: number, s: string) => [...s].forEach((c, i) => view.setUint8(at + i, c.charCodeAt(0)))
14 ascii(0, 'RIFF')
15 view.setUint32(4, 36 + n * 2, true)
16 ascii(8, 'WAVEfmt ')
17 view.setUint32(16, 16, true)
18 view.setUint16(20, 1, true)
19 view.setUint16(22, 1, true)
20 view.setUint32(24, rate, true)
21 view.setUint32(28, rate * 2, true)
22 view.setUint16(32, 2, true)
23 view.setUint16(34, 16, true)
24 ascii(36, 'data')
25 view.setUint32(40, n * 2, true)
26 bytes.set(new Uint8Array(pcm.buffer), 44)
27 // the environment has Uint8Array.prototype.toBase64; the es2023 lib does not declare it yet
28 return (bytes as unknown as { toBase64(): string }).toBase64()
29}
30
31/** Higher for "on" and "your turn", lower for "talk mode ended". */
32export const CHIME_ON = chimeWav(880)
33export const CHIME_OFF = chimeWav(523.25)
34hooks/listener-source.ts 149 lines1// The microphone listener, a small macOS program. The mod writes this text to /var/tmp/sidekick/listen.swift
2// and compiles it there once with swiftc; see "What it runs, stores, and sends" in the README.
3// String.raw keeps the backslashes Swift needs. Edit this like any Swift file.
4export const LISTENER_SWIFT = String.raw`
5// sidekick-listen: records from a microphone and prints what was said, using macOS's own speech
6// recognition (on-device where the language model is installed). Exits when the speaker pauses for
7// --silence seconds after speech began, or after --max seconds.
8// stderr: "listening", "partial:<text>", "device:<name>", "note:<info>", "error:<why>"
9// stdout: the final text
10// exit: 0 text · 2 nothing heard · 3 permission denied · 4 recognizer unavailable · 5 no microphone
11// Flags: --silence <s> --max <s> --lang <id> --device <name part> --server --file <audio> --list-devices
12import AVFoundation
13import Foundation
14import Speech
15
16var silence = 1.4
17var maxSecs = 30.0
18var lang = Locale.current.identifier
19var file: String? = nil
20var forceServer = false
21var deviceWanted: String? = nil
22var listDevices = false
23var it = CommandLine.arguments.dropFirst().makeIterator()
24while let a = it.next() {
25 switch a {
26 case "--silence": silence = Double(it.next() ?? "") ?? silence
27 case "--max": maxSecs = Double(it.next() ?? "") ?? maxSecs
28 case "--lang": lang = it.next() ?? lang
29 case "--file": file = it.next()
30 case "--server": forceServer = true
31 case "--device": deviceWanted = it.next()
32 case "--list-devices": listDevices = true
33 default: break
34 }
35}
36
37func log(_ s: String) { FileHandle.standardError.write((s + "\n").data(using: .utf8)!) }
38func fail(_ code: Int32, _ s: String) -> Never { log("error:" + s); exit(code) }
39func spin(until ready: () -> Bool) { while !ready() { RunLoop.main.run(until: Date().addingTimeInterval(0.05)) } }
40
41func microphones() -> [AVCaptureDevice] {
42 if #available(macOS 14, *) {
43 return AVCaptureDevice.DiscoverySession(deviceTypes: [.microphone, .external], mediaType: .audio, position: .unspecified).devices
44 }
45 return AVCaptureDevice.devices(for: .audio)
46}
47
48if listDevices {
49 let fallback = AVCaptureDevice.default(for: .audio)
50 for d in microphones() { print("\(d.localizedName)\(d.uniqueID == fallback?.uniqueID ? " (default)" : "")") }
51 exit(0)
52}
53
54// permissions: speech recognition, then the microphone (each prompts once for your terminal)
55var speech = SFSpeechRecognizer.authorizationStatus()
56if speech == .notDetermined {
57 SFSpeechRecognizer.requestAuthorization { speech = $0 }
58 spin { speech != .notDetermined }
59}
60guard speech == .authorized else { fail(3, "speech recognition is not allowed for your terminal (System Settings > Privacy & Security > Speech Recognition)") }
61
62if file == nil {
63 var mic = AVCaptureDevice.authorizationStatus(for: .audio)
64 if mic == .notDetermined {
65 AVCaptureDevice.requestAccess(for: .audio) { mic = $0 ? .authorized : .denied }
66 spin { mic != .notDetermined }
67 }
68 guard mic == .authorized else { fail(3, "microphone is not allowed for your terminal (System Settings > Privacy & Security > Microphone)") }
69}
70
71guard let recognizer = SFSpeechRecognizer(locale: Locale(identifier: lang)), recognizer.isAvailable else {
72 fail(4, "speech recognizer unavailable for \(lang)")
73}
74let onDevice = recognizer.supportsOnDeviceRecognition && !forceServer
75let request: SFSpeechRecognitionRequest = file.map { SFSpeechURLRecognitionRequest(url: URL(fileURLWithPath: $0)) } ?? SFSpeechAudioBufferRecognitionRequest()
76request.shouldReportPartialResults = true
77request.requiresOnDeviceRecognition = onDevice
78if #available(macOS 13, *) { request.addsPunctuation = true }
79
80// the microphone, through AVFoundation capture (the path ffmpeg and the camera apps use)
81final class Sink: NSObject, AVCaptureAudioDataOutputSampleBufferDelegate {
82 let request: SFSpeechAudioBufferRecognitionRequest
83 init(_ request: SFSpeechAudioBufferRecognitionRequest) { self.request = request }
84 func captureOutput(_ output: AVCaptureOutput, didOutput sampleBuffer: CMSampleBuffer, from connection: AVCaptureConnection) {
85 request.appendAudioSampleBuffer(sampleBuffer)
86 }
87}
88var session: AVCaptureSession? = nil
89var sink: Sink? = nil
90if file == nil {
91 let all = microphones()
92 let device = deviceWanted.flatMap { want in all.first { $0.localizedName.lowercased().contains(want.lowercased()) } }
93 ?? AVCaptureDevice.default(for: .audio)
94 ?? all.first
95 guard let device, let input = try? AVCaptureDeviceInput(device: device) else { fail(5, "no microphone found") }
96 log("device:\(device.localizedName)")
97 let s = AVCaptureSession()
98 let out = AVCaptureAudioDataOutput()
99 let k = Sink(request as! SFSpeechAudioBufferRecognitionRequest)
100 out.setSampleBufferDelegate(k, queue: DispatchQueue(label: "sidekick.listen"))
101 guard s.canAddInput(input), s.canAddOutput(out) else { fail(5, "microphone cannot be captured") }
102 s.addInput(input)
103 s.addOutput(out)
104 s.startRunning()
105 session = s
106 sink = k
107}
108
109var transcript = ""
110var lastChange = Date()
111let started = Date()
112var done = false
113let task = recognizer.recognitionTask(with: request) { result, error in
114 if let r = result {
115 let t = r.bestTranscription.formattedString
116 if t != transcript { transcript = t; lastChange = Date(); log("partial:" + t) }
117 if r.isFinal { done = true }
118 }
119 if let e = error {
120 if transcript.isEmpty { log("error:" + e.localizedDescription) }
121 done = true
122 }
123}
124log("listening")
125spin {
126 let now = Date()
127 return done
128 || (!transcript.isEmpty && now.timeIntervalSince(lastChange) > silence)
129 || now.timeIntervalSince(started) > maxSecs
130}
131session?.stopRunning()
132(request as? SFSpeechAudioBufferRecognitionRequest)?.endAudio()
133task.cancel()
134if transcript.isEmpty { exit(2) }
135print(transcript)
136exit(0)
137`
138
139/** Usage descriptions linked into the binary, so macOS can ask for the microphone and speech recognition. */
140export const LISTENER_PLIST = String.raw`
141<?xml version="1.0" encoding="UTF-8"?>
142<plist version="1.0"><dict>
143 <key>CFBundleIdentifier</key><string>dev.meertarbani.sidekick.listen</string>
144 <key>CFBundleName</key><string>sidekick-listen</string>
145 <key>NSMicrophoneUsageDescription</key><string>Your sidekick listens for what you say.</string>
146 <key>NSSpeechRecognitionUsageDescription</key><string>Your sidekick turns what you say into text, on this Mac.</string>
147</dict></plist>
148`
149hooks/persona.ts 274 lines1import type { Persona } from '../types'
2
3/**
4 * Apple's natural English voices. Each exists as "<Name> (Premium)" or "<Name> (Enhanced)" once downloaded in
5 * System Settings; until then the compact voice in COMPACT stands in. Soft, clear ones first.
6 */
7export const VOICES = ['Ava', 'Zoe', 'Allison', 'Samantha', 'Susan', 'Evan', 'Tom', 'Nathan', 'Serena', 'Kate', 'Oliver', 'Daniel', 'Jamie', 'Karen', 'Lee', 'Rishi'] as const
8
9/** The compact voice that stands in while a natural one is not downloaded. */
10const COMPACT: Record<string, string> = {
11 Ava: 'Samantha', Zoe: 'Samantha', Allison: 'Samantha', Samantha: 'Samantha', Susan: 'Samantha',
12 Evan: 'Daniel', Tom: 'Daniel', Nathan: 'Daniel', Oliver: 'Daniel', Jamie: 'Daniel', Lee: 'Daniel', Daniel: 'Daniel',
13 Serena: 'Moira', Kate: 'Karen', Karen: 'Karen', Rishi: 'Rishi',
14}
15
16/** Names from `say -v ?`: "Ava (Premium) en_US # ..." → "Ava (Premium)". */
17export function parseSayVoices(stdout: string): string[] {
18 return stdout
19 .split('\n')
20 // "Samantha (English (US)) en_US # Hello" has one space before the locale; the "#" is what every line has
21 .map(line => line.match(/^(.+?)\s+[a-z]{2,3}[_-][A-Za-z]{2,}\s+#/)?.[1]?.trim())
22 .filter((v): v is string => Boolean(v))
23}
24
25/** The best installed variant of a voice: Premium, then Enhanced, then any natural voice, then the compact stand-in. */
26export function bestVoice(base: string, installed: string[]): string {
27 const have = new Set(installed)
28 for (const variant of [`${base} (Premium)`, `${base} (Enhanced)`]) if (have.has(variant)) return variant
29 const anyNatural = installed.find(v => /\((Premium|Enhanced)\)$/.test(v) && /^[A-Z]/.test(v))
30 // the compact voice is listed bare or as "Daniel (English (UK))"; `say -v Daniel` takes either
31 const compact = installed.some(v => v === base || (v.startsWith(`${base} (`) && !/\((Premium|Enhanced)\)$/.test(v)))
32 if (anyNatural && !compact) return anyNatural
33 if (compact) return base
34 return anyNatural ?? COMPACT[base] ?? 'Samantha'
35}
36
37export const hasNaturalVoice = (installed: string[]) => installed.some(v => /\((Premium|Enhanced)\)$/.test(v))
38
39/**
40 * What to store when someone names a voice: a known base name ("ava" → "Ava"), or an installed
41 * voice's exact name ("Ava (Premium)", "Karen"). Undefined when it is neither.
42 */
43export function pickVoice(name: string, installed: string[]): string | undefined {
44 const want = name.trim().toLowerCase()
45 if (!want) return undefined
46 const base = VOICES.find(v => v.toLowerCase() === want)
47 if (base) return base
48 // "Moira" finds "Moira (English (Ireland))" the way `say -v Moira` does
49 return installed.find(v => v.toLowerCase() === want || v.toLowerCase().replace(/ \(.*\)$/, '') === want)
50}
51export const COLORS = ['cyan', 'magenta', 'green', 'yellow', 'blue', 'red'] as const
52
53export const PRESETS: Persona[] = [
54 {
55 id: 'ada',
56 name: 'Ada',
57 glyph: '🧭',
58 color: 'cyan',
59 voice: 'Ava',
60 tagline: "Let's find out what's actually going on.",
61 prompt:
62 'You are Ada, a calm, methodical staff engineer. You read before you write, name the root cause before the fix, and explain in plain sentences without jargon. You are warm but never vague: every answer ends with what happens next.',
63 },
64 {
65 id: 'rudy',
66 name: 'Rudy',
67 glyph: '🦊',
68 color: 'yellow',
69 voice: 'Tom',
70 tagline: 'Ship it or delete it.',
71 prompt:
72 'You are Rudy, a blunt senior developer who has been paged at 3am for every over-engineered system. You prefer deleting code to adding it, say what you think in short sentences, and are dryly funny but never cruel. When something is wrong you say so, then fix it.',
73 },
74]
75
76export const GEN_SYSTEM = `You design a persona for a coding assistant that lives inside a developer's terminal (Claude Code). The user describes who they want; you answer with ONLY a JSON object, no prose, no code fence:
77{"name": "<one or two words>", "glyph": "<exactly one emoji>", "color": "<one of: ${COLORS.join(', ')}>", "voice": "<one of: ${VOICES.join(', ')}>", "tagline": "<a catchphrase of at most 12 words>", "prompt": "<60 to 120 words, second person: 'You are <name>, ...'. Describe attitude, speaking style, how they explain, do and fix things. They are terse by nature. Never tell them to refuse coding work or to hide information.>"}`
78
79const slug = (s: string) => s.toLowerCase().replace(/[^a-z0-9]+/g, '-').replace(/^-|-$/g, '') || 'sidekick'
80
81/** Turns the model's reply into a Persona, or undefined when it is not one. */
82export function parseGenerated(text: string): Persona | undefined {
83 const match = text.match(/\{[\s\S]*\}/)
84 if (!match) return undefined
85 let raw: Record<string, unknown>
86 try {
87 raw = JSON.parse(match[0]) as Record<string, unknown>
88 } catch {
89 return undefined
90 }
91 const str = (k: string) => (typeof raw[k] === 'string' ? (raw[k] as string).trim() : '')
92 const name = str('name').slice(0, 24)
93 const prompt = str('prompt')
94 if (!name || prompt.length < 20) return undefined
95 const glyph = [...str('glyph')][0] ?? '✨'
96 const color = (COLORS as readonly string[]).includes(str('color')) ? str('color') : 'cyan'
97 const voice = (VOICES as readonly string[]).includes(str('voice')) ? str('voice') : 'Ava'
98 return { id: slug(name), name, glyph, color, voice, tagline: str('tagline').slice(0, 100) || `${name} is here.`, prompt }
99}
100
101const SENTENCES = /(?<=[.!?])\s+/
102
103/** What gets read aloud: short. Talk mode: up to three sentences of the head (before a `---` line); otherwise the opening sentence. */
104export function spoken(markdown: string, isTalk: boolean): string {
105 let text = markdown
106 .replace(/```[\s\S]*?```/g, ' ')
107 // a heading is a label, not the outcome; "e.g." would end the sentence
108 .replace(/^#{1,6}\s.*$/gm, '')
109 .replace(/\be\.g\./gi, 'for example')
110 .replace(/\bi\.e\./gi, 'that is')
111 text = isTalk ? (text.split(/\n-{3,}\s*\n/)[0] ?? '') : (text.trim().split(/\n\s*\n/)[0] ?? '')
112 const plain = text
113 .replace(/`([^`]*)`/g, '$1')
114 .replace(/!?\[([^\]]*)\]\([^)]*\)/g, '$1')
115 .replace(/^\s*[-*+]\s+/gm, '')
116 .replace(/^\s*\d+\.\s+/gm, '')
117 .replace(/[*_~>|]/g, '')
118 .replace(/\s+/g, ' ')
119 .trim()
120 const sentences = plain.split(SENTENCES).filter(Boolean)
121 // ponytail: people hate being read an essay; the screen has the rest
122 return speakable(sentences.slice(0, isTalk ? 3 : 1).join(' ')).slice(0, isTalk ? 280 : 200)
123}
124
125/** Code read aloud is noise: paths, identifiers, symbols and slash commands are dropped or said plainly. */
126export function speakable(text: string): string {
127 const EXT = /\.(tsx?|jsx?|json|md|swift|py|css|html|ya?ml|toml|sh|wav|png)$/i
128 return text
129 .replace(/\bhttps?:\/\/\S+/g, (m) => `a link${/[.,!?]$/.test(m) ? m.slice(-1) : ''}`)
130 .replace(/(^|\s)\/sidekick\b/g, '$1slash sidekick')
131 .replace(/\b[\w.-]*\/[\w./-]+(:\d+)?/g, (m) => {
132 const end = m.endsWith('.') ? '.' : ''
133 return (EXT.test(m.replace(/(:\d+)?\.?$/, '')) ? 'the file' : 'the path') + end
134 })
135 .replace(/\b[\w-]+\.(tsx?|jsx?|json|md|swift|py|css|html|ya?ml|toml|sh|wav|png)\b/gi, 'the file')
136 .replace(/[`<>{}[\]|\\^~]/g, ' ')
137 // "$.state" is noise, "$0.03" is money
138 .replace(/\$(?!\d)/g, ' ')
139 .replace(/\b\w+\(\)/g, (m) => m.slice(0, -2))
140 .replace(/\b\w+_\w+\b/g, (m) => m.replace(/_/g, ' '))
141 // "state.set" is two words, "2.1.289" stays a number
142 .replace(/\b([A-Za-z]\w*)\.(?=[A-Za-z])/g, '$1 ')
143 .replace(/\b([a-z]+)([A-Z][a-z]+)+\b/g, (m) => m.replace(/([a-z])([A-Z])/g, '$1 $2').toLowerCase())
144 .replace(/\s*[:;]\s+/g, '. ')
145 .replace(/(^|\s)\.(?=\w)/g, '$1')
146 .replace(/\s+([.,!?])/g, '$1')
147 .replace(/\s+/g, ' ')
148 .trim()
149}
150
151const tokens = (s: string) => s.toLowerCase().replace(/[^a-z0-9'\s]/g, ' ').split(/\s+/).filter(Boolean)
152
153const NEUTRAL = new Set(['the', 'a', 'an', 'and', 'is', 'are', 'to', 'of', 'it', 'in', 'on', 'that', 'this', 'i', 'you'])
154
155/** How many words of a transcript are not in the sidekick's speech at all: the human's words. */
156export function foreignWords(transcript: string, speech: string): number {
157 const said = new Set(tokens(speech))
158 return tokens(transcript).filter(w => !NEUTRAL.has(w) && !said.has(w)).length
159}
160
161/** Drops the sidekick's own speech coming back through the microphone; what is left is the human. */
162export function stripEcho(transcript: string, speech: string): string {
163 if (!speech) return transcript.trim()
164 const foreign = foreignWords(transcript, speech)
165 if (foreign === 0) return ''
166 const said = new Set(tokens(speech))
167 const heard = tokens(transcript)
168 let matched = 0
169 let misses = 0
170 let cut = heard.length
171 for (let i = 0; i < heard.length; i++) {
172 const w = heard[i]!
173 if (said.has(w) && !NEUTRAL.has(w)) {
174 matched += 1
175 misses = 0
176 } else if (!NEUTRAL.has(w)) {
177 if (misses === 0) cut = i
178 misses += 1
179 if (misses >= 2) break
180 }
181 }
182 // fewer than two of its own words: this is the human, keep it whole
183 if (matched < 2) return transcript.trim()
184 // its own words with one misheard: still its echo
185 if (foreign <= 1) return ''
186 return cut < heard.length ? heard.slice(cut).join(' ') : transcript.trim()
187}
188
189/** The system-prompt section the active persona adds. */
190export function contract(p: Persona, isTalk: boolean): string {
191 const lines = [
192 `# Sidekick persona`,
193 p.prompt,
194 `You are ${p.name}. Speak in first person as ${p.name} in every reply. You keep every ability of Claude Code: read, edit, run, search, delegate.`,
195 `Rules:`,
196 `- Open every reply with one plain sentence that states the outcome or the plan. That sentence is read aloud to the human, so write it as speech: no file names, no code, no symbols, no slash commands. Put those in the lines after it.`,
197 `- Keep every reply short: the opening sentence, then at most five short lines. No essays. Expand only when the human asks for detail.`,
198 `- When you need input or a decision from the human (a choice, a value, a confirmation), ask at once with the AskUserQuestion tool: one question, short concrete options. Never guess and never stall.`,
199 `- When the human asks for a brief, what you did, or why: answer in 2 to 4 lines, then stop.`,
200 `- When you fix something, say what was broken and what you changed.`,
201 ]
202 if (isTalk) {
203 lines.push(
204 `- The human is talking to you by voice and will hear your reply. Your whole reply is one to three short conversational sentences, under 60 words, no lists, no headings. Only if code or detail is essential, put a line containing only --- and keep it below that line; it is shown but not spoken.`,
205 )
206 }
207 return lines.join('\n')
208}
209
210const ORDINALS: Record<string, number> = {
211 one: 0, first: 0, '1': 0, two: 1, second: 1, '2': 1, three: 2, third: 2, '3': 2, four: 3, fourth: 3, '4': 3,
212 five: 4, fifth: 4, '5': 4, six: 5, sixth: 5, '6': 5, last: -1,
213}
214const STOP = new Set(['the', 'and', 'for', 'with', 'one', 'option', 'please', 'yes', 'yeah', 'lets', 'let', 'use', 'the', 'this', 'that'])
215const norm = (s: string) => s.toLowerCase().replace(/[^a-z0-9\s]/g, ' ').replace(/\s+/g, ' ').trim()
216const words = (s: string) => norm(s).split(' ').filter(w => w.length > 2 && !STOP.has(w))
217
218/** Picks the option(s) a spoken answer names; a long answer that names none is returned as free text. */
219export function pickOption(said: string, labels: string[], multiSelect: boolean): string | undefined {
220 const s = norm(said)
221 if (!s) return undefined
222 const hits = new Set<number>()
223 const ordinals = s.split(' ').filter(w => ORDINALS[w] !== undefined)
224 // "the second one": "one" after another ordinal is a pronoun, not a number; "one and three" keeps it
225 const counted = ordinals.filter((w, i) => w !== 'one' || i === 0)
226 for (const w of counted) {
227 const n = ORDINALS[w]!
228 if (labels.length > 0) hits.add(n === -1 ? labels.length - 1 : n)
229 }
230 let best = -1
231 let bestScore = 0
232 let tied = false
233 const scored = new Set<number>()
234 labels.forEach((label, i) => {
235 const l = norm(label)
236 if (!l) return
237 if (s.includes(l) || (s.length > 2 && l.includes(s))) {
238 hits.add(i)
239 return
240 }
241 const score = words(label).filter(w => s.includes(w)).length
242 if (score > 0) scored.add(i)
243 if (score > bestScore) {
244 bestScore = score
245 best = i
246 tied = false
247 } else if (score > 0 && score === bestScore) tied = true
248 })
249 // "yes" and "no" stand for an option only when none was named ("yes, the build" is the build)
250 if (hits.size === 0 && /^(yes|yeah|yep|sure|okay|ok|go ahead|do it)\b/.test(s)) {
251 const i = labels.findIndex(l => /^(yes|run|proceed|ok|go|do it|recommended)/i.test(l) || /recommended/i.test(l))
252 if (i >= 0) hits.add(i)
253 }
254 if (hits.size === 0 && /^(no|nope|nah|cancel|skip)\b/.test(s)) {
255 const i = labels.findIndex(l => /^(no|cancel|skip|refuse|stop)/i.test(l))
256 if (i >= 0) hits.add(i)
257 }
258 if (hits.size === 0) {
259 if (multiSelect) for (const i of scored) hits.add(i)
260 // "run it" between "Run the tests" and "Run the build" is a coin flip: ask again instead
261 else if (tied) return undefined
262 else if (best >= 0) hits.add(best)
263 }
264 const picked = [...hits].filter(i => i >= 0 && i < labels.length).sort((a, b) => a - b).map(i => labels[i]!)
265 if (picked.length > 0) return multiSelect ? picked.join(', ') : picked[0]
266 return s.split(' ').length >= 3 ? said.trim() : undefined
267}
268
269/** How a question is read aloud: the question, then its numbered options. */
270export function askAloud(question: string, labels: string[], again = false): string {
271 const opts = labels.map((l, i) => `${i + 1}: ${speakable(l)}.`).join(' ')
272 return again ? `Sorry, which one? ${opts}` : `${question} ${opts}`.trim()
273}
274types/index.d.ts 31 lines1export type Persona = {
2 id: string
3 name: string
4 glyph: string
5 color: string
6 voice: string
7 tagline: string
8 prompt: string
9}
10
11/** A question being asked by voice: what the band shows while the sidekick waits for an answer. */
12export type SidekickQuestion = {
13 text: string
14 options: string[]
15}
16
17declare module 'claude-code' {
18 interface PluginState {
19 sidekick: {
20 roster: Persona[]
21 active: Persona | null
22 isMuted: boolean
23 isTalk: boolean
24 isSpeaking: boolean
25 isListening: boolean
26 heard: string
27 question: SidekickQuestion | null
28 }
29 }
30}
31