SLOPSHOPPER

sidekick

Persona agents you can talk to: create a sidekick, it explains, does and fixes things in character, asks when it needs you, and speaks its replies.

newpanebandspinnerrowsguard
v0.3.0MITupdated 2026-10-04Redskull-127/sidekick
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · sidekick
│ ┃ Sidekick ✕ › fix the failing auth test and add an audit log call │ ┃ 1: 🧭 Ada Let's find out what's actually goi │ ┃ 2: 🦊 Rudy Ship it or delete it. ⏺ Read(src/auth.ts) │ ┃ ⎿ Read 6 lines │ ┃ 0: off m: mute t: talk ⏺ Update(src/auth.ts) │ ┃ ⎿ Added 2 lines, removed 1 line │ ┃ /sidekick new <describe someone> to add one ⏺ Bash(bun test) │ ⎿ 3 pass, 1 fail │ │ ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ │ › /sidekick │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · Sidekick
1: 🧭 Ada Let's find out what's actually going on. 2: 🦊 Rudy Ship it or delete it. 0: off m: mute t: talk /sidekick new <describe someone> to add one
README

sidekick

Persona agents you can talk to inside Claude Code. A sidekick is the one driving your session: it explains, does and fixes things in character, asks you the moment it needs a decision, gives you a brief when you ask for one, and speaks its replies out loud. You can talk back with your voice.

/sidekick new Rudy, a blunt senior Rust dev who hates abstractions
  🦊 Rudy is ready. "Ship it or delete it."

/sidekick talk
  🎙 Talk mode on. Just speak; Rudy answers out loud and listens again.

Two sidekicks ship so it works before you create anyone: Ada (calm staff engineer) and Rudy (blunt senior dev). A sidekick never announces itself by voice; a cabin-style chime says "on" and "your turn", and a lower one says talk mode ended. The name is in the terminal header.

What changes in your session

  • Every reply is headed with the sidekick's glyph and name, the spinner reads Thinking · Rudy is on it…, and the question dialog is headed 🦊 Rudy asks:.
  • The sidekick is told to ask you with a short, concrete question whenever it needs input instead of guessing, to answer "brief?" in a few lines, and to say what broke and what it changed when it fixes something.
  • The opening sentence of each reply is read aloud in the sidekick's macOS voice. /sidekick mute turns that off.
  • Talk mode (/sidekick talk) is hands-free and continuous. The sidekick greets you and listens; when you pause it sends what you said as your prompt, works, speaks the answer, and keeps listening. No key to hold. Talk over it to interrupt: the moment it hears words that aren't its own, it stops speaking and takes yours as the next prompt. Say "stop" or "wait" while it works to cancel the turn. Say "end talk" to stop. A band above the prompt shows what is happening (listening… with the words as they are recognized, speaking…, working…) with s to skip the speech, m to mute and x to end. Three long silences in a row end talk mode on their own.
  • Short answers. Spoken replies are one sentence (up to three in talk mode); the persona is told to keep the whole reply to a few lines and to expand only when asked.
  • Questions by voice. When the sidekick needs a decision in talk mode it reads the question and the numbered options aloud, listens, and takes your answer: "the second one", "production", "yes", or a free-text answer of a few words. If it can't match what you said after two tries, the normal dialog opens and you pick with a key.
  • In talk mode the sidekick keeps the spoken part short and puts code after a --- line that is shown but not read.

Commands

| Command | What it does | | :- | :- | | /sidekick | Roster pane: 1–9 switch, 0 off, m mute, t talk | | /sidekick new <description> | Generates a persona (name, glyph, voice, catchphrase, character) with Haiku and makes it active | | /sidekick use <name> / /sidekick off | Switch sidekick / plain Claude | | /sidekick talk | Toggle hands-free talk mode (voice in, voice out) | | /sidekick mute / /sidekick speak | Stop / resume spoken replies for this session | | /sidekick voices / /sidekick voices install | Show which Apple voices are installed; open the pane to download natural ones | | /sidekick voice <voice> / /sidekick voice <name> <voice> | Give the active sidekick, or a named one, a voice (/sidekick voice Ava) | | /sidekick list / /sidekick rm <name> | Housekeeping |

Sidekicks and the active one persist across sessions in the plugin's store; mute and talk mode are per session, and the Speak replies option in /config sets the default.

Voices

macOS ships compact voices that sound robotic. Apple's natural Premium and Enhanced voices are a one-time download: run /sidekick voices install, or go to System Settings → Accessibility → Spoken Content → System Voice → ⓘ Manage Voices → English, and download Ava (Premium), Tom (Enhanced), or any voice marked Premium or Enhanced. Sidekicks use the best installed variant of their voice automatically (Premium, then Enhanced, then any natural voice, then the compact stand-in), so nothing else to configure.

To give a sidekick a different voice: /sidekick voice Ava for the active one, or /sidekick voice Rudy Ava for a named one. Any of the sixteen natural names works, and so does the exact name of any voice say -v ? lists. The sidekick confirms by saying its catchphrase in the new voice. The choice is saved with the persona.

Requirements

  • Claude Code 2.1.287 or later (tested on 2.1.289). Mods run in the terminal and the Desktop app's Code tab.
  • Speech out uses macOS say; on Linux and Windows replies stay text.
  • Voice in is a small native listener (the Swift program in hooks/listener-source.ts) that uses macOS's own speech recognition, on-device where the language model is installed. It is written to /var/tmp/sidekick/listen.swift and compiled there once with swiftc, which comes with the Xcode Command Line Tools (xcode-select --install). The first run asks for Microphone and Speech Recognition permission for your terminal. Nothing is sent to any third party; on-device recognition sends nothing anywhere.
  • Prefer push-to-talk? Claude Code's own /voice dictation still works alongside; the sidekick speaks its replies either way.

Install

claude plugin marketplace add Redskull-127/sidekick
claude plugin install sidekick@meer-mods

Or for one session: claude --plugin-dir /path/to/sidekick.

Before installing any mod, you can list what it hooks and calls without running it: claude plugin validate /path/to/sidekick.

What it runs, stores, and sends

Everything the plugin does is in its readable source. Spelled out, because a plugin that listens and speaks should be:

Programs it starts, and why

  • say speaks replies in the sidekick's voice; killall say cuts a reply short when you interrupt or press s.
  • say -v ? lists the voices you have, so a Premium or Enhanced one is used when installed.
  • afplay, through Claude Code's audio API, plays the two chimes synthesized in hooks/chime.ts.
  • mkdir -p /var/tmp/sidekick and swiftc once, to compile the listener source (hooks/listener-source.ts in this repo, written to /var/tmp/sidekick/listen.swift) into /var/tmp/sidekick/listen.
  • /var/tmp/sidekick/listen --session <this session's id>, while talk mode is on, to hear you through the microphone; pkill -f on that exact command line to stop listening, so another session's listener is never touched.
  • open x-apple.systempreferences:…SpokenContent when you run /sidekick voices install, to show the voice download pane.

Files it writes

  • In /var/tmp/sidekick/ only: listen.swift and Info.plist (the listener's source, copied from this plugin), listen (its compiled binary) and listen.version (which source it was built from). Nothing in your project, nothing in your settings, no build or startup file.
  • Your sidekicks and preferences in Claude Code's plugin store (~/.claude/plugins/store/), through the mods API.

What it reads

  • The transcript's final answer of each turn, to speak its first sentence or two.
  • Each AskUserQuestion call's questions, to read them aloud in talk mode.
  • /var/tmp/sidekick/listen.version, to know whether the compiled listener is current.

What it sends, and where

  • /sidekick new <description>: that description goes to Claude (Haiku) through your own Claude Code account, to write the persona. Nothing else is ever sent to Claude by the plugin itself.
  • Your voice goes to Apple's speech recognition, on-device where the language model is installed, otherwise to Apple's servers, exactly as macOS dictation does. The plugin only sees the text that comes back.
  • The prompts it submits are exactly the words you spoke, as transcribed, sent as your own message. It never composes a prompt of its own.
  • Nothing goes anywhere else. No telemetry, no analytics, no network calls of its own.

Where it steps in

  • It adds one section to the system prompt: the active persona and the rules above (ask promptly, keep it short). That is all it changes about what Claude reads.
  • In talk mode it answers the AskUserQuestion tool in the dialog's place when it understood your spoken answer; otherwise the normal dialog opens.
  • "Stop" or "wait" spoken in talk mode cancels the running turn. Nothing else touches tool calls or permissions.

Where it works: this plugin is a Claude Code mod, so it runs in the Claude Code terminal and the Desktop app's Code tab. Added from claude.ai it installs, but has nothing to do in chat or Cowork.

Develop

claude plugin validate .   # what the engine reads from the module
claude plugin test .       # tests/sidekick.test.ts, no session or network
claude --plugin-dir .      # hot-reloads on save

The first load writes .claude-plugin/types/ and a one-line tsconfig.json (both gitignored); from then on the editor and npx tsc -p . type-check against the real mods API.

How hands-free works

A mod cannot open a microphone itself, so the sidekick spawns the listener binary, over and over, for as long as talk mode is on. Each run records until you pause for 1.4 seconds (or 60 seconds at most), streams the partial transcript to the band, prints the final text, and exits. The mod submits that text as your prompt.

The listener runs while the sidekick speaks too, so you can interrupt. Its own voice comes back through the microphone, so the mod drops the leading words that match what it was saying; the first two words that aren't its own cut the speech short. The listener captures through AVFoundation and picks the system's default microphone (/var/tmp/sidekick/listen --list-devices shows them once built; add --device <name part> to the listener's arguments in hooks/register.tsx to pin one).

Known limits

  • Echo is filtered by words, not by acoustics. Repeating the sidekick's own words back to it right after it says them won't register as a new prompt.
  • Languages: the listener uses your macOS locale. Add --lang to the listener's arguments in hooks/register.tsx to pin one.
  • After editing the Swift in hooks/listener-source.ts, the binary rebuilds on next use (the build is keyed to a hash of the source). The chimes are synthesized in hooks/chime.ts.
Source 5 files
hooks/register.tsx 681 lines
1import type { EngineInterface, Register } from 'claude-code'
2
3import type { Persona, SidekickQuestion } from '../types'
4import { CHIME_OFF, CHIME_ON } from './chime'
5import { LISTENER_PLIST, LISTENER_SWIFT } from './listener-source'
6import { GEN_SYSTEM, PRESETS, VOICES, askAloud, bestVoice, contract, foreignWords, hasNaturalVoice, parseGenerated, parseSayVoices, pickOption, pickVoice, spoken, stripEcho } from './persona'
7
8const PANE = 'sidekick'
9// session state the drawings read; each a literal reference for `$.state`
10const roster = { plugin: 'sidekick', key: 'roster' } as const
11const active = { plugin: 'sidekick', key: 'active' } as const
12const isMuted = { plugin: 'sidekick', key: 'isMuted' } as const
13const isTalk = { plugin: 'sidekick', key: 'isTalk' } as const
14const isSpeaking = { plugin: 'sidekick', key: 'isSpeaking' } as const
15const isListening = { plugin: 'sidekick', key: 'isListening' } as const
16const heard = { plugin: 'sidekick', key: 'heard' } as const
17const question = { plugin: 'sidekick', key: 'question' } as const
18
19const USAGE = [
20  '/sidekick                 roster pane',
21  '/sidekick new <describe someone>',
22  '/sidekick use <name> | off',
23  '/sidekick talk            hands-free: speak, it answers, it listens again',
24  '/sidekick mute | speak',
25  '/sidekick voices [install]  natural Apple voices',
26  '/sidekick voice <voice>     give the active sidekick a voice (or: voice <name> <voice>)',
27  '/sidekick list | rm <name>',
28].join('\n')
29
30const VOICE_STEPS = [
31  'System Settings → Accessibility → Spoken Content → System Voice → ⓘ Manage Voices → English.',
32  'Download Ava (Premium) and Tom (Enhanced), or any voice marked Premium or Enhanced. Sidekicks pick them up at once.',
33].join('\n')
34
35const END_TALK = /\b(end talk|stop talking|stop listening|that'?s all|goodbye)\b/i
36const STOP = /^(stop|wait|hold on|hang on|never ?mind|shut up|quiet)\b/i
37
38let installedVoices: string[] = []
39
40/** Which voices `say` has right now. */
41async function loadVoices($: EngineInterface) {
42  try {
43    const r = await $.process.run(['say', '-v', '?'])
44    installedVoices = parseSayVoices(r.stdout)
45  } catch {
46    installedVoices = []
47  }
48}
49
50/** Store → state, at session start and after /clear, /resume, /branch. */
51async function load($: EngineInterface, speakByDefault: boolean) {
52  const saved = (await $.store.get('personas')) as Persona[] | undefined
53  const migrated = (await $.store.get('voicesMigrated')) as boolean | undefined
54  // no list yet seeds the presets; an empty list means everyone was removed and stays empty.
55  // presets saved by 0.1 with the compact voices move to the natural ones, once; after that a chosen voice stays
56  const list = (saved ?? PRESETS).map(p => {
57    const preset = PRESETS.find(x => x.id === p.id)
58    return !migrated && preset && (p.voice === 'Samantha' || p.voice === 'Daniel') ? { ...p, voice: preset.voice } : p
59  })
60  if (!saved || !migrated) {
61    await $.store.set('personas', list)
62    await $.store.set('voicesMigrated', true)
63  }
64  const activeId = (await $.store.get('active')) as string | null | undefined
65  await $.state.set(roster, list)
66  await $.state.set(active, list.find(p => p.id === activeId) ?? null)
67  // mute and talk mode are per session; the `speak` option in /config is the default
68  await $.state.set(isMuted, !speakByDefault)
69  await $.state.set(isTalk, false)
70}
71
72async function setActive($: EngineInterface, p: Persona | null) {
73  // no sidekick, no ear: the pane's off key and rm must not leave the microphone open
74  if (!p) await setTalk($, false)
75  await $.state.set(active, p)
76  await $.store.set('active', p?.id ?? null)
77}
78
79async function toggleMute($: EngineInterface) {
80  const muted = !(((await $.state.get(isMuted)).value ?? false))
81  await $.state.set(isMuted, muted)
82  return muted
83}
84
85let speechNow = ''
86let recentSpeech: string[] = []
87let sayId = 0
88let isHushed = false
89
90// ponytail: the last three utterances are the echo vocabulary; the recognizer delivers its transcript
91// well after the speaker goes quiet, so "what is playing right now" was never the right question
92const echoText = () => recentSpeech.join(' ')
93
94/** Speaks in the persona's voice, cutting any speech still going; falls back to the default voice, then to silence. */
95async function say($: EngineInterface, text: string, voice: string) {
96  if (!text) return
97  if (((await $.state.get(isSpeaking)).value ?? false)) await hush($)
98  const mine = ++sayId
99  isHushed = false
100  speechNow = text
101  recentSpeech = [...recentSpeech.slice(-2), text]
102  await $.state.set(isSpeaking, true)
103  // until a natural voice is there, look again before each reply: one downloaded mid-session is used at once
104  if (!hasNaturalVoice(installedVoices)) await loadVoices($)
105  try {
106    await $.audio.speak(text, { voice: bestVoice(voice, installedVoices) })
107  } catch {
108    if (!isHushed) {
109      try {
110        await $.audio.speak(text)
111      } catch {
112        // no synthesizer on this platform: text only
113      }
114    }
115  }
116  if (sayId !== mine) return
117  speechNow = ''
118  await $.state.set(isSpeaking, false)
119  if (!isHushed && (((await $.state.get(isTalk)).value ?? false))) void chime($)
120}
121
122/** Awaits a call whose failure does not matter (a chime that cannot play, a program not running). */
123async function quietly(work: Promise<unknown>) {
124  try {
125    await work
126  } catch {
127    // nothing to do
128  }
129}
130
131/** The cabin chime: "a sidekick is on" and "your turn"; the lower one when talk mode ends. */
132async function chime($: EngineInterface, kind: 'on' | 'off' = 'on') {
133  await quietly($.audio.play({ base64: kind === 'off' ? CHIME_OFF : CHIME_ON, mime: 'audio/wav' }))
134}
135
136// ponytail: `say` has no abort; killing the macOS synthesizer is the one-line skip
137async function hush($: EngineInterface) {
138  isHushed = true
139  await quietly($.process.run(['killall', 'say']))
140}
141
142// ── the listener: macOS speech recognition in a small native binary, built once ─────────────────
143
144/** A short stable hash of the listener's source, so a changed source gets a fresh build. */
145function sourceVersion(source: string): string {
146  let hash = 5381
147  for (let i = 0; i < source.length; i++) hash = ((hash * 33) ^ source.charCodeAt(i)) >>> 0
148  return hash.toString(16)
149}
150
151// the listener is built and run at fixed paths under /var/tmp/sidekick, outside any project;
152// every path below is written out in full at the call, so a reader can see each one
153
154// why the listener could not be built, for the reply to /sidekick talk
155let listenerProblem = ''
156
157/** Where the compiled listener lives, building it on first use from the source in listener-source.ts; null when it cannot be built here. */
158async function listenerPath($: EngineInterface): Promise<string | null> {
159  const version = sourceVersion(LISTENER_SWIFT)
160  let built = ''
161  try {
162    built = await $.fs.read('/var/tmp/sidekick/listen.version')
163  } catch {
164    built = ''
165  }
166  if (built.trim() === version && (await $.fs.exists('/var/tmp/sidekick/listen'))) return '/var/tmp/sidekick/listen'
167  $.ui.toast('Building the sidekick listener, one time, about 20 seconds…')
168  try {
169    await $.process.run(['mkdir', '-p', '/var/tmp/sidekick'])
170    await $.fs.write('/var/tmp/sidekick/listen.swift', LISTENER_SWIFT)
171    await $.fs.write('/var/tmp/sidekick/Info.plist', LISTENER_PLIST)
172    const compiled = await $.process.run(
173      ['swiftc', '-O', '/var/tmp/sidekick/listen.swift', '-o', '/var/tmp/sidekick/listen', '-Xlinker', '-sectcreate', '-Xlinker', '__TEXT', '-Xlinker', '__info_plist', '-Xlinker', '/var/tmp/sidekick/Info.plist'],
174      { timeoutMs: 240_000 },
175    )
176    if (compiled.exitCode !== 0) {
177      // /usr/bin/swiftc exists on every Mac; without the Command Line Tools it exits non-zero and says so
178      const why = compiled.stderr.split('\n').find(l => l.includes('error:')) ?? compiled.stderr.slice(0, 200)
179      listenerProblem = /xcode-select|developer tools|command line tools/i.test(compiled.stderr)
180        ? 'Hands-free needs the Xcode Command Line Tools. Run xcode-select --install, then /sidekick talk again.'
181        : `The listener did not build: ${why}`
182      return null
183    }
184    await $.fs.write('/var/tmp/sidekick/listen.version', version)
185    return '/var/tmp/sidekick/listen'
186  } catch {
187    listenerProblem = 'Hands-free needs macOS: the listener uses Apple speech recognition. Replies stay text here.'
188    return null
189  }
190}
191
192type Heard = { text: string | null; raw: string; error: string; isCancelled: boolean; isFatal: boolean; isBroken: boolean; isEcho: boolean }
193
194let listening: { bin: string; isCancelled: boolean } | null = null
195// this session's id tags its listener, so stopping here never stops another session's
196let sessionTag = ''
197let earLoop: Promise<void> | null = null
198let earPaused = false
199let runningTurn: string | null = null
200
201/** Stops the listener now, from a key press or because talk mode ended. */
202async function stopListening($: EngineInterface) {
203  const current = listening
204  if (!current) return
205  current.isCancelled = true
206  // ponytail: a pending stream read cannot be interrupted from here, so the child is ended by name
207  await quietly($.process.run(['pkill', '-f', `/var/tmp/sidekick/listen .*--session ${sessionTag}`]))
208}
209
210/**
211 * Records one utterance and returns its text. The sidekick's own voice coming back through the
212 * microphone is filtered out by its words; the first words that are not its own cut the speech short.
213 */
214async function listen($: EngineInterface, bin: string): Promise<Heard> {
215  if (listening) await stopListening($)
216  const run = { bin, isCancelled: false }
217  listening = run
218  await $.state.set(heard, '')
219  await $.state.set(isListening, true)
220  let out = ''
221  let error = ''
222  let code: number | null = null
223  let bargedIn = false
224  try {
225    const child = $.process.spawn({ argv: ['/var/tmp/sidekick/listen', '--silence', '1.4', '--max', '60', '--session', sessionTag] })
226    for await (const { stream, text } of child) {
227      if (stream === 'stdout') {
228        out += text
229        continue
230      }
231      for (const line of text.split('\n')) {
232        if (line.startsWith('partial:')) {
233          const own = stripEcho(line.slice(8), echoText())
234          await $.state.set(heard, own)
235          // two words that are not in its own speech mean the human is talking over it
236          if (speechNow && !bargedIn && foreignWords(line.slice(8), echoText()) >= 2) {
237            bargedIn = true
238            void hush($)
239          }
240        } else if (line.startsWith('error:')) error = line.slice(6).trim()
241      }
242    }
243    code = (await child.result).code
244  } catch {
245    code = 5
246  }
247  if (listening === run) listening = null
248  await $.state.set(isListening, false)
249  const raw = out.trim()
250  const said = stripEcho(raw, echoText())
251  return {
252    text: run.isCancelled || !said ? null : said,
253    raw,
254    error,
255    isCancelled: run.isCancelled,
256    isFatal: code !== null && code >= 3,
257    // quit without hearing anything and without being asked to: an error, or killed from outside
258    isBroken: !run.isCancelled && !raw && (code === 2 || code === null),
259    isEcho: raw.length > 0 && !said,
260  }
261}
262
263async function setTalk($: EngineInterface, enabled: boolean) {
264  const was = ((await $.state.get(isTalk)).value ?? false)
265  await $.state.set(isTalk, enabled)
266  if (!enabled) {
267    await stopListening($)
268    await $.state.set(question, null)
269    if (was) void chime($, 'off')
270  }
271}
272
273/** The ear: one listener at a time for as long as talk mode is on. What it hears becomes the next prompt. */
274async function runEar($: EngineInterface) {
275  const p = ((await $.state.get(active)).value ?? null)
276  if (!p) return
277  const bin = await listenerPath($)
278  if (!bin) return setTalk($, false)
279  let quiet = 0
280  let failed = 0
281  while (((await $.state.get(isTalk)).value ?? false)) {
282    if (earPaused) {
283      await $.clock.sleep(200)
284      continue
285    }
286    const r = await listen($, bin)
287    if (r.isCancelled) {
288      if (!(((await $.state.get(isTalk)).value ?? false))) return
289      continue
290    }
291    if (r.isFatal) {
292      await setTalk($, false)
293      $.ui.toast(`${p.name} can't listen: ${r.error || 'the listener quit'}`)
294      return
295    }
296    if (r.isBroken) {
297      // a listener that quits at once is broken, not quiet: three in a row end talk mode, with the reason
298      if (++failed >= 3) {
299        await setTalk($, false)
300        $.ui.toast(`${p.name} can't hear: ${r.error || 'the listener keeps quitting'}. /sidekick talk to retry.`)
301        return
302      }
303      await $.clock.sleep(1000)
304      continue
305    }
306    if (r.text === null) {
307      const busy = r.isEcho || runningTurn !== null || (((await $.state.get(isSpeaking)).value ?? false))
308      if (!busy && ++quiet >= 3) {
309        // ponytail: three silent rounds end talk mode rather than listening forever
310        await setTalk($, false)
311        $.ui.toast(`${p.name} stopped listening after a long silence. /sidekick talk to resume.`)
312        return
313      }
314      continue
315    }
316    quiet = 0
317    failed = 0
318    const said = r.text
319    const wasSpeaking = ((await $.state.get(isSpeaking)).value ?? false)
320    if (wasSpeaking) await hush($)
321    // "stop" is a command when something is running or when it is all that was said; "stop the dev server" is a prompt
322    const wordCount = said.split(/\s+/).length
323    if (wordCount <= 5 && END_TALK.test(said)) {
324      await setTalk($, false)
325      return
326    }
327    if (STOP.test(said) && (wordCount <= 3 || runningTurn || wasSpeaking)) {
328      if (runningTurn) await quietly($.turn.abort({ turnId: runningTurn }))
329      continue
330    }
331    void $.prompt.submit({ text: said, asUser: true })
332  }
333}
334
335/** Turns talk mode on: greets, and starts the ear once. */
336async function startTalk($: EngineInterface) {
337  const p = ((await $.state.get(active)).value ?? null)
338  if (!p) return
339  void chime($)
340  if (earLoop) return
341  earLoop = quietly(runEar($)).finally(() => {
342    earLoop = null
343  })
344}
345
346async function toggleTalk($: EngineInterface) {
347  if (((await $.state.get(isTalk)).value ?? false)) {
348    await hush($)
349    return setTalk($, false)
350  }
351  await setTalk($, true)
352  void startTalk($)
353}
354
355const findPersona = (list: Persona[], q: string) =>
356  list.find(p => p.id === q.toLowerCase() || p.name.toLowerCase() === q.toLowerCase())
357
358export const register: Register = (on, options) => {
359  const speakByDefault = options.speak !== false
360  let isHeadless = false
361
362  on('session.start', async ($, e, next) => {
363    // a `claude -p` run has no one to talk to: nothing loads, so every other hook stays out of the way
364    isHeadless = !e.isInteractive
365    if (isHeadless) return next(e)
366    sessionTag = await $.session.id()
367    await loadVoices($)
368    await load($, speakByDefault)
369    if (((await $.state.get(active)).value ?? null)) void $.ui.open({ id: PANE, title: 'Sidekick', closeOnEscape: true })
370    try {
371      await $.command.register({
372        name: 'sidekick',
373        description: 'Persona agents you can talk to: create, switch, talk hands-free, mute',
374        argumentHint: '[new <description> | use <name> | off | talk | mute | speak | voices [install] | voice [<name>] <voice> | list | rm <name>]',
375        immediate: true,
376      })
377    } catch {
378      // name taken by another plugin: the pane and hooks still work
379    }
380    return next(e)
381  })
382
383  on('classic.SessionStart', { source: ['clear', 'resume', 'fork'] }, async ($, e, next) => {
384    if (isHeadless) return next(e)
385    await load($, speakByDefault)
386    return next(e)
387  })
388
389  on('command.run', { command: 'sidekick' }, async ($, e) => {
390    const [sub = '', ...rest] = e.args.trim().split(/\s+/)
391    const arg = rest.join(' ')
392    const list: Persona[] = (await $.state.get(roster)).value ?? []
393
394    switch (sub) {
395      case '': {
396        await $.ui.open({ id: PANE, title: 'Sidekick', focus: true, closeOnEscape: true })
397        return {}
398      }
399      case 'new': {
400        if (!arg) return { text: 'Describe who you want: /sidekick new Rudy, a blunt senior Rust dev' }
401        const r = await $.model.complete({ model: 'haiku', system: GEN_SYSTEM, prompt: arg, maxTokens: 700, timeoutMs: 20000 })
402        if (!r.isAnswered) return { text: `Couldn't reach the model (${r.reason}). Try again.` }
403        const p = parseGenerated(r.text)
404        if (!p) return { text: 'The model did not return a persona I could read. Try a clearer description.' }
405        if (list.some(x => x.id === p.id) || findPersona(list, p.name)) {
406          // the name is how use, rm and voice find it, so a second Rudy is "Rudy 3", not a hidden rudy-3
407          const n = list.length + 1
408          p.id = `${p.id}-${n}`
409          p.name = `${p.name} ${n}`
410        }
411        const grown = [...list, p]
412        await $.state.set(roster, grown)
413        await $.store.set('personas', grown)
414        await setActive($, p)
415        void $.ui.open({ id: PANE, title: 'Sidekick', closeOnEscape: true })
416        void chime($)
417        return { text: `${p.glyph} ${p.name} is ready. "${p.tagline}"\n${p.name} is driving now. /sidekick talk to speak with ${p.name}.` }
418      }
419      case 'use': {
420        const p = findPersona(list, arg)
421        if (!p) return { text: `No sidekick named "${arg}". /sidekick list` }
422        await setActive($, p)
423        void $.ui.open({ id: PANE, title: 'Sidekick', closeOnEscape: true })
424        void chime($)
425        return { text: `${p.glyph} ${p.name} is driving now.` }
426      }
427      case 'off': {
428        await setActive($, null)
429        return { text: 'Sidekick off. Plain Claude is back.' }
430      }
431      case 'talk': {
432        const p = ((await $.state.get(active)).value ?? null)
433        if (!p) return { text: 'Pick a sidekick first: /sidekick use Rudy' }
434        const wasOn = ((await $.state.get(isTalk)).value ?? false)
435        if (!wasOn) {
436          // built (once) before anything is announced, so a Mac that cannot listen gets the reason as the reply
437          const bin = await listenerPath($)
438          if (!bin) return { text: `🎙 ${listenerProblem}` }
439        }
440        await toggleTalk($)
441        if (wasOn) return { text: '🎙 Talk mode off.' }
442        if (!hasNaturalVoice(installedVoices)) $.ui.toast('Voice sounds robotic? /sidekick voices install gets Apple\'s natural ones.')
443        return { text: `🎙 Talk mode on. Just speak; ${p.name} answers out loud and listens again. Talk over ${p.name} to interrupt. Say "end talk" or press x on the band to stop.` }
444      }
445      case 'voices': {
446        if (arg === 'install') {
447          await quietly($.process.run(['open', 'x-apple.systempreferences:com.apple.preference.universalaccess?SpokenContent']))
448          return { text: `Opened System Settings.\n${VOICE_STEPS}` }
449        }
450        await loadVoices($)
451        const natural = installedVoices.filter(v => /\((Premium|Enhanced)\)$/.test(v))
452        const cur = ((await $.state.get(active)).value ?? null)
453        const lines = [
454          natural.length > 0 ? `Natural voices installed: ${natural.join(', ')}` : 'No natural (Premium/Enhanced) voices installed yet, so sidekicks use the compact ones.',
455          cur ? `${cur.glyph} ${cur.name} speaks as ${bestVoice(cur.voice, installedVoices)}.` : '',
456          `Sidekicks choose from: ${VOICES.join(', ')}.`,
457          natural.length > 0 ? '' : `/sidekick voices install opens the download pane.\n${VOICE_STEPS}`,
458        ]
459        return { text: lines.filter(Boolean).join('\n') }
460      }
461      case 'voice': {
462        const cur = (await $.state.get(active)).value ?? null
463        // "voice Rudy Ava" names a sidekick; otherwise every word is the voice ("voice Bad News")
464        const named = rest.length > 1 ? findPersona(list, rest[0]!) : undefined
465        const target = named ?? cur
466        const wanted = named ? rest.slice(1).join(' ') : arg
467        if (!target) return { text: 'Pick a sidekick first: /sidekick use Rudy' }
468        if (!wanted) return { text: `${target.glyph} ${target.name} speaks as ${bestVoice(target.voice, installedVoices)}. Choose from: ${VOICES.join(', ')}, or any installed voice (/sidekick voices).` }
469        await loadVoices($)
470        const voice = pickVoice(wanted, installedVoices)
471        if (!voice) return { text: `No voice called "${wanted}". Choose from: ${VOICES.join(', ')}, or any installed voice (/sidekick voices).` }
472        const changed = { ...target, voice }
473        const grown = list.map(x => (x.id === target.id ? changed : x))
474        await $.state.set(roster, grown)
475        await $.store.set('personas', grown)
476        if (cur?.id === target.id) await $.state.set(active, changed)
477        const resolved = bestVoice(voice, installedVoices)
478        const standIn = resolved === voice || resolved.startsWith(`${voice} (`) ? '' : ` (${voice} isn't installed here, so this stands in.)`
479        if (!(await $.state.get(isMuted)).value) void say($, changed.tagline, voice)
480        return { text: `${changed.glyph} ${changed.name} now speaks as ${resolved}.${standIn}` }
481      }
482      case 'mute':
483      case 'speak': {
484        const muted = sub === 'mute'
485        await $.state.set(isMuted, muted)
486        return { text: muted ? '🔇 Replies are no longer spoken this session.' : '🔊 Replies are spoken again.' }
487      }
488      case 'list': {
489        const cur = ((await $.state.get(active)).value ?? null)
490        return { text: list.map(p => `${p.id === cur?.id ? '●' : ' '} ${p.glyph} ${p.name} — ${p.tagline}`).join('\n') }
491      }
492      case 'rm': {
493        const p = findPersona(list, arg)
494        if (!p) return { text: `No sidekick named "${arg}".` }
495        const grown = list.filter(x => x.id !== p.id)
496        await $.state.set(roster, grown)
497        await $.store.set('personas', grown)
498        if ((((await $.state.get(active)).value ?? null))?.id === p.id) await setActive($, null)
499        return { text: `Removed ${p.glyph} ${p.name}.` }
500      }
501      default:
502        return { text: USAGE }
503    }
504  })
505
506  on('prompt.compose', async ($, e, next) => {
507    const composed = await next(e)
508    const p = ((await $.state.get(active)).value ?? null)
509    if (!p) return composed
510    const talk = ((await $.state.get(isTalk)).value ?? false)
511    return { sections: [...composed.sections, { id: 'sidekick:persona', scope: 'session', text: contract(p, talk) }] }
512  })
513
514  // the ear keeps listening through a turn: what you say meanwhile is the next prompt, "stop" aborts the turn
515  on('turn.start', async ($, e, next) => {
516    runningTurn = e.turnId
517    return next(e)
518  })
519
520  on('ui.render', { component: 'AssistantMessage' }, async ($, e, next) => {
521    const p = ((await $.state.get(active)).value ?? null)
522    if (!p || !e.props.isFirstOfReply) return next(e)
523    const { Box, Text } = $.ui.resolve(e)
524    const theirs = await next(e)
525    return (
526      <Box flexDirection="column">
527        <Text bold color={p.color}>
528          {p.glyph} {p.name}
529        </Text>
530        {theirs}
531      </Box>
532    )
533  })
534
535  on('ui.render', { component: 'Spinner' }, async ($, e, next) => {
536    const p = ((await $.state.get(active)).value ?? null)
537    if (!p) return next(e)
538    return next({ ...e, props: { ...e.props, suffix: ` · ${p.name} is on it…` } })
539  })
540
541  on('ui.render', { component: 'AskUserQuestion' }, async ($, e, next) => {
542    const p = ((await $.state.get(active)).value ?? null)
543    if (!p) return next(e)
544    const { Box, Text } = $.ui.resolve(e)
545    const theirs = await next(e)
546    return (
547      <Box flexDirection="column">
548        <Text bold color={p.color}>
549          {p.glyph} {p.name} asks:
550        </Text>
551        {theirs}
552      </Box>
553    )
554  })
555
556  // in talk mode a question is asked and answered by voice; the dialog is the fallback
557  on('tool.call', { tool: 'AskUserQuestion' }, async ($, e, next) => {
558    const p = ((await $.state.get(active)).value ?? null)
559    if (!p || !(((await $.state.get(isTalk)).value ?? false))) return next(e)
560    const bin = await listenerPath($)
561    if (!bin) return next(e)
562    earPaused = true
563    await stopListening($)
564    try {
565      const answers: Record<string, string> = {}
566      for (const q of e.questions) {
567        const labels = q.options?.map(o => o.label) ?? []
568        await $.state.set(question, ({ text: q.question, options: labels }))
569        let picked: string | undefined
570        for (let attempt = 0; attempt < 2 && !picked; attempt++) {
571          // listen while asking, so an answer spoken over the question still lands
572          const hearing = listen($, bin)
573          if (!(((await $.state.get(isMuted)).value ?? false))) void say($, askAloud(q.question, labels, attempt > 0), p.voice)
574          const r = await hearing
575          await hush($)
576          if (r.isCancelled || r.isFatal) break
577          // the question named every option, so "production" alone looks like its echo: a few words naming exactly one option are the answer
578          const named = labels.filter(l => pickOption(r.raw, [l], false) === l)
579          const answer = r.text ?? (named.length === 1 && r.raw.split(/\s+/).length <= 4 ? r.raw : null)
580          if (answer) picked = labels.length > 0 ? pickOption(answer, labels, q.multiSelect) : answer
581        }
582        await $.state.set(question, null)
583        if (!picked) {
584          $.ui.toast(`${p.name} didn't catch that. Pick with a key.`)
585          return next(e)
586        }
587        answers[q.question] = picked
588      }
589      return { result: { questions: e.questions, answers } }
590    } finally {
591      earPaused = false
592    }
593  })
594
595  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
596    const p = ((await $.state.get(active)).value ?? null)
597    if (!p || e.props.hasSurvey || !(((await $.state.get(isTalk)).value ?? false))) return next(e)
598    const { Box, Button, Text } = $.ui.resolve(e)
599    const speaking = ((await $.state.get(isSpeaking)).value ?? false)
600    const hearing = ((await $.state.get(isListening)).value ?? false)
601    const partial = ((await $.state.get(heard)).value ?? '')
602    const muted = ((await $.state.get(isMuted)).value ?? false)
603    const asked = ((await $.state.get(question)).value ?? null)
604    const status = speaking
605      ? `🔊 ${p.name} is speaking… talk over to interrupt`
606      : e.props.isWorking
607        ? `${p.glyph} ${p.name} is working… say "stop" to cancel`
608        : hearing
609          ? `🎙 listening…`
610          : `${p.glyph} ${p.name} is ready`
611    return (
612      <Box flexDirection="column">
613        {asked && (
614          <Text color={p.color}>
615            ❓ {asked.text} {asked.options.map((o, i) => `  ${i + 1}) ${o}`).join('')}
616          </Text>
617        )}
618        <Box flexDirection="row" columnGap={2}>
619          <Text color={p.color}>{status}</Text>
620          {hearing && partial && <Text dimColor>{partial}</Text>}
621          {speaking && <Button key="skip" label="skip" hotkey="s" plain onPress={() => hush($)} />}
622          {!speaking && !hearing && !e.props.isWorking && (
623            <Button key="listen" label="listen" hotkey="l" plain onPress={() => void startTalk($)} />
624          )}
625          <Button key="mute" label={muted ? 'speak' : 'mute'} hotkey="m" plain onPress={() => toggleMute($)} />
626          <Button key="end" label="end talk" hotkey="x" plain onPress={() => toggleTalk($)} />
627        </Box>
628      </Box>
629    )
630  })
631
632  on('ui.render', { component: 'PromptHint' }, async ($, e, next) => {
633    const p = ((await $.state.get(active)).value ?? null)
634    if (!p || e.props.isDraft || e.props.isWorking || !(((await $.state.get(isTalk)).value ?? false))) return next(e)
635    return next({ ...e, props: { ...e.props, tail: ` · 🎙 talking with ${p.name}` } })
636  })
637
638  on('ui.render', { component: 'Pane', requestId: 'sidekick' }, async ($, e) => {
639    const { Box, Button, Text } = $.ui.resolve(e)
640    const list: Persona[] = (await $.state.get(roster)).value ?? []
641    const cur = ((await $.state.get(active)).value ?? null)
642    const muted = ((await $.state.get(isMuted)).value ?? false)
643    const talk = ((await $.state.get(isTalk)).value ?? false)
644    return (
645      <Box flexDirection="column">
646        {list.slice(0, 9).map((p, i) => (
647          <Box flexDirection="row" columnGap={1}>
648            <Button
649              key={`use-${p.id}`}
650              label={`${p.glyph} ${p.name}`}
651              hotkey={String(i + 1)}
652              plain
653              dimColor={cur?.id !== p.id}
654              onPress={() => setActive($, p)}
655            />
656            <Text dimColor>{cur?.id === p.id ? `● ${p.tagline}` : p.tagline}</Text>
657          </Box>
658        ))}
659        <Text> </Text>
660        <Box flexDirection="row" columnGap={3}>
661          <Button key="off" label="off" hotkey="0" plain onPress={() => setActive($, null)} />
662          <Button key="mute" label={muted ? 'speak' : 'mute'} hotkey="m" plain onPress={() => toggleMute($)} />
663          <Button key="talk" label={talk ? 'end talk' : 'talk'} hotkey="t" plain onPress={() => toggleTalk($)} />
664        </Box>
665        <Text> </Text>
666        <Text dimColor>/sidekick new {'<describe someone>'} to add one</Text>
667      </Box>
668    )
669  })
670
671  // the reply is spoken, then in talk mode the sidekick listens for what comes next
672  on('turn.complete', async ($, e, next) => {
673    if (!e.agentId) runningTurn = null
674    const p = ((await $.state.get(active)).value ?? null)
675    if (p && e.reason === 'answer' && !e.agentId && !(((await $.state.get(isMuted)).value ?? false))) {
676      void say($, spoken(e.answer, ((await $.state.get(isTalk)).value ?? false)), p.voice)
677    }
678    return next(e)
679  })
680}
681
hooks/chime.ts 34 lines
1/** One soft "ting", made on the spot: a decaying sine with a touch of octave, as a 22 kHz mono WAV. */
2export function chimeWav(freq: number, seconds = 1.1, decay = 4): string {
3  const rate = 22050
4  const n = Math.floor(rate * seconds)
5  const pcm = new Int16Array(n)
6  for (let i = 0; i < n; i++) {
7    const t = i / rate
8    const v = 0.6 * Math.exp(-decay * t) * (Math.sin(2 * Math.PI * freq * t) + 0.2 * Math.sin(4 * Math.PI * freq * t)) / 1.2
9    pcm[i] = Math.round(Math.max(-1, Math.min(1, v)) * 32767)
10  }
11  const bytes = new Uint8Array(44 + n * 2)
12  const view = new DataView(bytes.buffer)
13  const ascii = (at: number, s: string) => [...s].forEach((c, i) => view.setUint8(at + i, c.charCodeAt(0)))
14  ascii(0, 'RIFF')
15  view.setUint32(4, 36 + n * 2, true)
16  ascii(8, 'WAVEfmt ')
17  view.setUint32(16, 16, true)
18  view.setUint16(20, 1, true)
19  view.setUint16(22, 1, true)
20  view.setUint32(24, rate, true)
21  view.setUint32(28, rate * 2, true)
22  view.setUint16(32, 2, true)
23  view.setUint16(34, 16, true)
24  ascii(36, 'data')
25  view.setUint32(40, n * 2, true)
26  bytes.set(new Uint8Array(pcm.buffer), 44)
27  // the environment has Uint8Array.prototype.toBase64; the es2023 lib does not declare it yet
28  return (bytes as unknown as { toBase64(): string }).toBase64()
29}
30
31/** Higher for "on" and "your turn", lower for "talk mode ended". */
32export const CHIME_ON = chimeWav(880)
33export const CHIME_OFF = chimeWav(523.25)
34
hooks/listener-source.ts 149 lines
1// The microphone listener, a small macOS program. The mod writes this text to /var/tmp/sidekick/listen.swift
2// and compiles it there once with swiftc; see "What it runs, stores, and sends" in the README.
3// String.raw keeps the backslashes Swift needs. Edit this like any Swift file.
4export const LISTENER_SWIFT = String.raw`
5// sidekick-listen: records from a microphone and prints what was said, using macOS's own speech
6// recognition (on-device where the language model is installed). Exits when the speaker pauses for
7// --silence seconds after speech began, or after --max seconds.
8//   stderr: "listening", "partial:<text>", "device:<name>", "note:<info>", "error:<why>"
9//   stdout: the final text
10//   exit:   0 text · 2 nothing heard · 3 permission denied · 4 recognizer unavailable · 5 no microphone
11// Flags: --silence <s> --max <s> --lang <id> --device <name part> --server --file <audio> --list-devices
12import AVFoundation
13import Foundation
14import Speech
15
16var silence = 1.4
17var maxSecs = 30.0
18var lang = Locale.current.identifier
19var file: String? = nil
20var forceServer = false
21var deviceWanted: String? = nil
22var listDevices = false
23var it = CommandLine.arguments.dropFirst().makeIterator()
24while let a = it.next() {
25  switch a {
26  case "--silence": silence = Double(it.next() ?? "") ?? silence
27  case "--max": maxSecs = Double(it.next() ?? "") ?? maxSecs
28  case "--lang": lang = it.next() ?? lang
29  case "--file": file = it.next()
30  case "--server": forceServer = true
31  case "--device": deviceWanted = it.next()
32  case "--list-devices": listDevices = true
33  default: break
34  }
35}
36
37func log(_ s: String) { FileHandle.standardError.write((s + "\n").data(using: .utf8)!) }
38func fail(_ code: Int32, _ s: String) -> Never { log("error:" + s); exit(code) }
39func spin(until ready: () -> Bool) { while !ready() { RunLoop.main.run(until: Date().addingTimeInterval(0.05)) } }
40
41func microphones() -> [AVCaptureDevice] {
42  if #available(macOS 14, *) {
43    return AVCaptureDevice.DiscoverySession(deviceTypes: [.microphone, .external], mediaType: .audio, position: .unspecified).devices
44  }
45  return AVCaptureDevice.devices(for: .audio)
46}
47
48if listDevices {
49  let fallback = AVCaptureDevice.default(for: .audio)
50  for d in microphones() { print("\(d.localizedName)\(d.uniqueID == fallback?.uniqueID ? "  (default)" : "")") }
51  exit(0)
52}
53
54// permissions: speech recognition, then the microphone (each prompts once for your terminal)
55var speech = SFSpeechRecognizer.authorizationStatus()
56if speech == .notDetermined {
57  SFSpeechRecognizer.requestAuthorization { speech = $0 }
58  spin { speech != .notDetermined }
59}
60guard speech == .authorized else { fail(3, "speech recognition is not allowed for your terminal (System Settings > Privacy & Security > Speech Recognition)") }
61
62if file == nil {
63  var mic = AVCaptureDevice.authorizationStatus(for: .audio)
64  if mic == .notDetermined {
65    AVCaptureDevice.requestAccess(for: .audio) { mic = $0 ? .authorized : .denied }
66    spin { mic != .notDetermined }
67  }
68  guard mic == .authorized else { fail(3, "microphone is not allowed for your terminal (System Settings > Privacy & Security > Microphone)") }
69}
70
71guard let recognizer = SFSpeechRecognizer(locale: Locale(identifier: lang)), recognizer.isAvailable else {
72  fail(4, "speech recognizer unavailable for \(lang)")
73}
74let onDevice = recognizer.supportsOnDeviceRecognition && !forceServer
75let request: SFSpeechRecognitionRequest = file.map { SFSpeechURLRecognitionRequest(url: URL(fileURLWithPath: $0)) } ?? SFSpeechAudioBufferRecognitionRequest()
76request.shouldReportPartialResults = true
77request.requiresOnDeviceRecognition = onDevice
78if #available(macOS 13, *) { request.addsPunctuation = true }
79
80// the microphone, through AVFoundation capture (the path ffmpeg and the camera apps use)
81final class Sink: NSObject, AVCaptureAudioDataOutputSampleBufferDelegate {
82  let request: SFSpeechAudioBufferRecognitionRequest
83  init(_ request: SFSpeechAudioBufferRecognitionRequest) { self.request = request }
84  func captureOutput(_ output: AVCaptureOutput, didOutput sampleBuffer: CMSampleBuffer, from connection: AVCaptureConnection) {
85    request.appendAudioSampleBuffer(sampleBuffer)
86  }
87}
88var session: AVCaptureSession? = nil
89var sink: Sink? = nil
90if file == nil {
91  let all = microphones()
92  let device = deviceWanted.flatMap { want in all.first { $0.localizedName.lowercased().contains(want.lowercased()) } }
93    ?? AVCaptureDevice.default(for: .audio)
94    ?? all.first
95  guard let device, let input = try? AVCaptureDeviceInput(device: device) else { fail(5, "no microphone found") }
96  log("device:\(device.localizedName)")
97  let s = AVCaptureSession()
98  let out = AVCaptureAudioDataOutput()
99  let k = Sink(request as! SFSpeechAudioBufferRecognitionRequest)
100  out.setSampleBufferDelegate(k, queue: DispatchQueue(label: "sidekick.listen"))
101  guard s.canAddInput(input), s.canAddOutput(out) else { fail(5, "microphone cannot be captured") }
102  s.addInput(input)
103  s.addOutput(out)
104  s.startRunning()
105  session = s
106  sink = k
107}
108
109var transcript = ""
110var lastChange = Date()
111let started = Date()
112var done = false
113let task = recognizer.recognitionTask(with: request) { result, error in
114  if let r = result {
115    let t = r.bestTranscription.formattedString
116    if t != transcript { transcript = t; lastChange = Date(); log("partial:" + t) }
117    if r.isFinal { done = true }
118  }
119  if let e = error {
120    if transcript.isEmpty { log("error:" + e.localizedDescription) }
121    done = true
122  }
123}
124log("listening")
125spin {
126  let now = Date()
127  return done
128    || (!transcript.isEmpty && now.timeIntervalSince(lastChange) > silence)
129    || now.timeIntervalSince(started) > maxSecs
130}
131session?.stopRunning()
132(request as? SFSpeechAudioBufferRecognitionRequest)?.endAudio()
133task.cancel()
134if transcript.isEmpty { exit(2) }
135print(transcript)
136exit(0)
137`
138
139/** Usage descriptions linked into the binary, so macOS can ask for the microphone and speech recognition. */
140export const LISTENER_PLIST = String.raw`
141<?xml version="1.0" encoding="UTF-8"?>
142<plist version="1.0"><dict>
143  <key>CFBundleIdentifier</key><string>dev.meertarbani.sidekick.listen</string>
144  <key>CFBundleName</key><string>sidekick-listen</string>
145  <key>NSMicrophoneUsageDescription</key><string>Your sidekick listens for what you say.</string>
146  <key>NSSpeechRecognitionUsageDescription</key><string>Your sidekick turns what you say into text, on this Mac.</string>
147</dict></plist>
148`
149
hooks/persona.ts 274 lines
1import type { Persona } from '../types'
2
3/**
4 * Apple's natural English voices. Each exists as "<Name> (Premium)" or "<Name> (Enhanced)" once downloaded in
5 * System Settings; until then the compact voice in COMPACT stands in. Soft, clear ones first.
6 */
7export const VOICES = ['Ava', 'Zoe', 'Allison', 'Samantha', 'Susan', 'Evan', 'Tom', 'Nathan', 'Serena', 'Kate', 'Oliver', 'Daniel', 'Jamie', 'Karen', 'Lee', 'Rishi'] as const
8
9/** The compact voice that stands in while a natural one is not downloaded. */
10const COMPACT: Record<string, string> = {
11  Ava: 'Samantha', Zoe: 'Samantha', Allison: 'Samantha', Samantha: 'Samantha', Susan: 'Samantha',
12  Evan: 'Daniel', Tom: 'Daniel', Nathan: 'Daniel', Oliver: 'Daniel', Jamie: 'Daniel', Lee: 'Daniel', Daniel: 'Daniel',
13  Serena: 'Moira', Kate: 'Karen', Karen: 'Karen', Rishi: 'Rishi',
14}
15
16/** Names from `say -v ?`: "Ava (Premium)   en_US   # ..." → "Ava (Premium)". */
17export function parseSayVoices(stdout: string): string[] {
18  return stdout
19    .split('\n')
20    // "Samantha (English (US)) en_US    # Hello" has one space before the locale; the "#" is what every line has
21    .map(line => line.match(/^(.+?)\s+[a-z]{2,3}[_-][A-Za-z]{2,}\s+#/)?.[1]?.trim())
22    .filter((v): v is string => Boolean(v))
23}
24
25/** The best installed variant of a voice: Premium, then Enhanced, then any natural voice, then the compact stand-in. */
26export function bestVoice(base: string, installed: string[]): string {
27  const have = new Set(installed)
28  for (const variant of [`${base} (Premium)`, `${base} (Enhanced)`]) if (have.has(variant)) return variant
29  const anyNatural = installed.find(v => /\((Premium|Enhanced)\)$/.test(v) && /^[A-Z]/.test(v))
30  // the compact voice is listed bare or as "Daniel (English (UK))"; `say -v Daniel` takes either
31  const compact = installed.some(v => v === base || (v.startsWith(`${base} (`) && !/\((Premium|Enhanced)\)$/.test(v)))
32  if (anyNatural && !compact) return anyNatural
33  if (compact) return base
34  return anyNatural ?? COMPACT[base] ?? 'Samantha'
35}
36
37export const hasNaturalVoice = (installed: string[]) => installed.some(v => /\((Premium|Enhanced)\)$/.test(v))
38
39/**
40 * What to store when someone names a voice: a known base name ("ava" → "Ava"), or an installed
41 * voice's exact name ("Ava (Premium)", "Karen"). Undefined when it is neither.
42 */
43export function pickVoice(name: string, installed: string[]): string | undefined {
44  const want = name.trim().toLowerCase()
45  if (!want) return undefined
46  const base = VOICES.find(v => v.toLowerCase() === want)
47  if (base) return base
48  // "Moira" finds "Moira (English (Ireland))" the way `say -v Moira` does
49  return installed.find(v => v.toLowerCase() === want || v.toLowerCase().replace(/ \(.*\)$/, '') === want)
50}
51export const COLORS = ['cyan', 'magenta', 'green', 'yellow', 'blue', 'red'] as const
52
53export const PRESETS: Persona[] = [
54  {
55    id: 'ada',
56    name: 'Ada',
57    glyph: '🧭',
58    color: 'cyan',
59    voice: 'Ava',
60    tagline: "Let's find out what's actually going on.",
61    prompt:
62      'You are Ada, a calm, methodical staff engineer. You read before you write, name the root cause before the fix, and explain in plain sentences without jargon. You are warm but never vague: every answer ends with what happens next.',
63  },
64  {
65    id: 'rudy',
66    name: 'Rudy',
67    glyph: '🦊',
68    color: 'yellow',
69    voice: 'Tom',
70    tagline: 'Ship it or delete it.',
71    prompt:
72      'You are Rudy, a blunt senior developer who has been paged at 3am for every over-engineered system. You prefer deleting code to adding it, say what you think in short sentences, and are dryly funny but never cruel. When something is wrong you say so, then fix it.',
73  },
74]
75
76export const GEN_SYSTEM = `You design a persona for a coding assistant that lives inside a developer's terminal (Claude Code). The user describes who they want; you answer with ONLY a JSON object, no prose, no code fence:
77{"name": "<one or two words>", "glyph": "<exactly one emoji>", "color": "<one of: ${COLORS.join(', ')}>", "voice": "<one of: ${VOICES.join(', ')}>", "tagline": "<a catchphrase of at most 12 words>", "prompt": "<60 to 120 words, second person: 'You are <name>, ...'. Describe attitude, speaking style, how they explain, do and fix things. They are terse by nature. Never tell them to refuse coding work or to hide information.>"}`
78
79const slug = (s: string) => s.toLowerCase().replace(/[^a-z0-9]+/g, '-').replace(/^-|-$/g, '') || 'sidekick'
80
81/** Turns the model's reply into a Persona, or undefined when it is not one. */
82export function parseGenerated(text: string): Persona | undefined {
83  const match = text.match(/\{[\s\S]*\}/)
84  if (!match) return undefined
85  let raw: Record<string, unknown>
86  try {
87    raw = JSON.parse(match[0]) as Record<string, unknown>
88  } catch {
89    return undefined
90  }
91  const str = (k: string) => (typeof raw[k] === 'string' ? (raw[k] as string).trim() : '')
92  const name = str('name').slice(0, 24)
93  const prompt = str('prompt')
94  if (!name || prompt.length < 20) return undefined
95  const glyph = [...str('glyph')][0] ?? '✨'
96  const color = (COLORS as readonly string[]).includes(str('color')) ? str('color') : 'cyan'
97  const voice = (VOICES as readonly string[]).includes(str('voice')) ? str('voice') : 'Ava'
98  return { id: slug(name), name, glyph, color, voice, tagline: str('tagline').slice(0, 100) || `${name} is here.`, prompt }
99}
100
101const SENTENCES = /(?<=[.!?])\s+/
102
103/** What gets read aloud: short. Talk mode: up to three sentences of the head (before a `---` line); otherwise the opening sentence. */
104export function spoken(markdown: string, isTalk: boolean): string {
105  let text = markdown
106    .replace(/```[\s\S]*?```/g, ' ')
107    // a heading is a label, not the outcome; "e.g." would end the sentence
108    .replace(/^#{1,6}\s.*$/gm, '')
109    .replace(/\be\.g\./gi, 'for example')
110    .replace(/\bi\.e\./gi, 'that is')
111  text = isTalk ? (text.split(/\n-{3,}\s*\n/)[0] ?? '') : (text.trim().split(/\n\s*\n/)[0] ?? '')
112  const plain = text
113    .replace(/`([^`]*)`/g, '$1')
114    .replace(/!?\[([^\]]*)\]\([^)]*\)/g, '$1')
115    .replace(/^\s*[-*+]\s+/gm, '')
116    .replace(/^\s*\d+\.\s+/gm, '')
117    .replace(/[*_~>|]/g, '')
118    .replace(/\s+/g, ' ')
119    .trim()
120  const sentences = plain.split(SENTENCES).filter(Boolean)
121  // ponytail: people hate being read an essay; the screen has the rest
122  return speakable(sentences.slice(0, isTalk ? 3 : 1).join(' ')).slice(0, isTalk ? 280 : 200)
123}
124
125/** Code read aloud is noise: paths, identifiers, symbols and slash commands are dropped or said plainly. */
126export function speakable(text: string): string {
127  const EXT = /\.(tsx?|jsx?|json|md|swift|py|css|html|ya?ml|toml|sh|wav|png)$/i
128  return text
129    .replace(/\bhttps?:\/\/\S+/g, (m) => `a link${/[.,!?]$/.test(m) ? m.slice(-1) : ''}`)
130    .replace(/(^|\s)\/sidekick\b/g, '$1slash sidekick')
131    .replace(/\b[\w.-]*\/[\w./-]+(:\d+)?/g, (m) => {
132      const end = m.endsWith('.') ? '.' : ''
133      return (EXT.test(m.replace(/(:\d+)?\.?$/, '')) ? 'the file' : 'the path') + end
134    })
135    .replace(/\b[\w-]+\.(tsx?|jsx?|json|md|swift|py|css|html|ya?ml|toml|sh|wav|png)\b/gi, 'the file')
136    .replace(/[`<>{}[\]|\\^~]/g, ' ')
137    // "$.state" is noise, "$0.03" is money
138    .replace(/\$(?!\d)/g, ' ')
139    .replace(/\b\w+\(\)/g, (m) => m.slice(0, -2))
140    .replace(/\b\w+_\w+\b/g, (m) => m.replace(/_/g, ' '))
141    // "state.set" is two words, "2.1.289" stays a number
142    .replace(/\b([A-Za-z]\w*)\.(?=[A-Za-z])/g, '$1 ')
143    .replace(/\b([a-z]+)([A-Z][a-z]+)+\b/g, (m) => m.replace(/([a-z])([A-Z])/g, '$1 $2').toLowerCase())
144    .replace(/\s*[:;]\s+/g, '. ')
145    .replace(/(^|\s)\.(?=\w)/g, '$1')
146    .replace(/\s+([.,!?])/g, '$1')
147    .replace(/\s+/g, ' ')
148    .trim()
149}
150
151const tokens = (s: string) => s.toLowerCase().replace(/[^a-z0-9'\s]/g, ' ').split(/\s+/).filter(Boolean)
152
153const NEUTRAL = new Set(['the', 'a', 'an', 'and', 'is', 'are', 'to', 'of', 'it', 'in', 'on', 'that', 'this', 'i', 'you'])
154
155/** How many words of a transcript are not in the sidekick's speech at all: the human's words. */
156export function foreignWords(transcript: string, speech: string): number {
157  const said = new Set(tokens(speech))
158  return tokens(transcript).filter(w => !NEUTRAL.has(w) && !said.has(w)).length
159}
160
161/** Drops the sidekick's own speech coming back through the microphone; what is left is the human. */
162export function stripEcho(transcript: string, speech: string): string {
163  if (!speech) return transcript.trim()
164  const foreign = foreignWords(transcript, speech)
165  if (foreign === 0) return ''
166  const said = new Set(tokens(speech))
167  const heard = tokens(transcript)
168  let matched = 0
169  let misses = 0
170  let cut = heard.length
171  for (let i = 0; i < heard.length; i++) {
172    const w = heard[i]!
173    if (said.has(w) && !NEUTRAL.has(w)) {
174      matched += 1
175      misses = 0
176    } else if (!NEUTRAL.has(w)) {
177      if (misses === 0) cut = i
178      misses += 1
179      if (misses >= 2) break
180    }
181  }
182  // fewer than two of its own words: this is the human, keep it whole
183  if (matched < 2) return transcript.trim()
184  // its own words with one misheard: still its echo
185  if (foreign <= 1) return ''
186  return cut < heard.length ? heard.slice(cut).join(' ') : transcript.trim()
187}
188
189/** The system-prompt section the active persona adds. */
190export function contract(p: Persona, isTalk: boolean): string {
191  const lines = [
192    `# Sidekick persona`,
193    p.prompt,
194    `You are ${p.name}. Speak in first person as ${p.name} in every reply. You keep every ability of Claude Code: read, edit, run, search, delegate.`,
195    `Rules:`,
196    `- Open every reply with one plain sentence that states the outcome or the plan. That sentence is read aloud to the human, so write it as speech: no file names, no code, no symbols, no slash commands. Put those in the lines after it.`,
197    `- Keep every reply short: the opening sentence, then at most five short lines. No essays. Expand only when the human asks for detail.`,
198    `- When you need input or a decision from the human (a choice, a value, a confirmation), ask at once with the AskUserQuestion tool: one question, short concrete options. Never guess and never stall.`,
199    `- When the human asks for a brief, what you did, or why: answer in 2 to 4 lines, then stop.`,
200    `- When you fix something, say what was broken and what you changed.`,
201  ]
202  if (isTalk) {
203    lines.push(
204      `- The human is talking to you by voice and will hear your reply. Your whole reply is one to three short conversational sentences, under 60 words, no lists, no headings. Only if code or detail is essential, put a line containing only --- and keep it below that line; it is shown but not spoken.`,
205    )
206  }
207  return lines.join('\n')
208}
209
210const ORDINALS: Record<string, number> = {
211  one: 0, first: 0, '1': 0, two: 1, second: 1, '2': 1, three: 2, third: 2, '3': 2, four: 3, fourth: 3, '4': 3,
212  five: 4, fifth: 4, '5': 4, six: 5, sixth: 5, '6': 5, last: -1,
213}
214const STOP = new Set(['the', 'and', 'for', 'with', 'one', 'option', 'please', 'yes', 'yeah', 'lets', 'let', 'use', 'the', 'this', 'that'])
215const norm = (s: string) => s.toLowerCase().replace(/[^a-z0-9\s]/g, ' ').replace(/\s+/g, ' ').trim()
216const words = (s: string) => norm(s).split(' ').filter(w => w.length > 2 && !STOP.has(w))
217
218/** Picks the option(s) a spoken answer names; a long answer that names none is returned as free text. */
219export function pickOption(said: string, labels: string[], multiSelect: boolean): string | undefined {
220  const s = norm(said)
221  if (!s) return undefined
222  const hits = new Set<number>()
223  const ordinals = s.split(' ').filter(w => ORDINALS[w] !== undefined)
224  // "the second one": "one" after another ordinal is a pronoun, not a number; "one and three" keeps it
225  const counted = ordinals.filter((w, i) => w !== 'one' || i === 0)
226  for (const w of counted) {
227    const n = ORDINALS[w]!
228    if (labels.length > 0) hits.add(n === -1 ? labels.length - 1 : n)
229  }
230  let best = -1
231  let bestScore = 0
232  let tied = false
233  const scored = new Set<number>()
234  labels.forEach((label, i) => {
235    const l = norm(label)
236    if (!l) return
237    if (s.includes(l) || (s.length > 2 && l.includes(s))) {
238      hits.add(i)
239      return
240    }
241    const score = words(label).filter(w => s.includes(w)).length
242    if (score > 0) scored.add(i)
243    if (score > bestScore) {
244      bestScore = score
245      best = i
246      tied = false
247    } else if (score > 0 && score === bestScore) tied = true
248  })
249  // "yes" and "no" stand for an option only when none was named ("yes, the build" is the build)
250  if (hits.size === 0 && /^(yes|yeah|yep|sure|okay|ok|go ahead|do it)\b/.test(s)) {
251    const i = labels.findIndex(l => /^(yes|run|proceed|ok|go|do it|recommended)/i.test(l) || /recommended/i.test(l))
252    if (i >= 0) hits.add(i)
253  }
254  if (hits.size === 0 && /^(no|nope|nah|cancel|skip)\b/.test(s)) {
255    const i = labels.findIndex(l => /^(no|cancel|skip|refuse|stop)/i.test(l))
256    if (i >= 0) hits.add(i)
257  }
258  if (hits.size === 0) {
259    if (multiSelect) for (const i of scored) hits.add(i)
260    // "run it" between "Run the tests" and "Run the build" is a coin flip: ask again instead
261    else if (tied) return undefined
262    else if (best >= 0) hits.add(best)
263  }
264  const picked = [...hits].filter(i => i >= 0 && i < labels.length).sort((a, b) => a - b).map(i => labels[i]!)
265  if (picked.length > 0) return multiSelect ? picked.join(', ') : picked[0]
266  return s.split(' ').length >= 3 ? said.trim() : undefined
267}
268
269/** How a question is read aloud: the question, then its numbered options. */
270export function askAloud(question: string, labels: string[], again = false): string {
271  const opts = labels.map((l, i) => `${i + 1}: ${speakable(l)}.`).join(' ')
272  return again ? `Sorry, which one? ${opts}` : `${question} ${opts}`.trim()
273}
274
types/index.d.ts 31 lines
1export type Persona = {
2  id: string
3  name: string
4  glyph: string
5  color: string
6  voice: string
7  tagline: string
8  prompt: string
9}
10
11/** A question being asked by voice: what the band shows while the sidekick waits for an answer. */
12export type SidekickQuestion = {
13  text: string
14  options: string[]
15}
16
17declare module 'claude-code' {
18  interface PluginState {
19    sidekick: {
20      roster: Persona[]
21      active: Persona | null
22      isMuted: boolean
23      isTalk: boolean
24      isSpeaking: boolean
25      isListening: boolean
26      heard: string
27      question: SidekickQuestion | null
28    }
29  }
30}
31