SLOPSHOPPER

Aside

Ask Claude for advice in a side pane without touching the session's context: any model, optional read-only tools, an ongoing side thread, cache-aware cost…

newpaneguardcommandtoastprompt
★ 1v0.1.0MITupdated 2026-10-10astrosteveo/claude-plugins/plugins/aside
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · aside
│ ┃ Aside ✕ › fix the failing auth test and add an audit log call │ ┃ Ask: a side question ⏎ ask │ ┃ Model: Session (Opus 5.5) ≈$0.050 ▾ Effort: ⏺ Read(src/auth.ts) │ ┃ 97k tokens in context, cache warm (5m cache) ⎿ Read 6 lines │ ┃ Nothing asked yet. Nothing here reaches the ⏺ Update(src/auth.ts) │ ┃ unless you send it. ⎿ Added 2 lines, removed 1 line │ ⏺ Bash(bun test) │ ⎿ 3 pass, 1 fail │ │ ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ │ › /aside │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · Aside
Ask: a side question ⏎ ask Model: Session (Opus 5.5) ≈$0.050 ▾ Effort: default ▾ [ To 97k tokens in context, cache warm (5m cache). Nothing asked yet. Nothing here reaches the session unless you send it.
README

Aside

Aside lets you ask Claude for advice in the middle of a session without touching the session's context. The question and the answer stay in a pane of their own. The session's model never sees either unless you hand the answer over.

It does what /btw does, plus four things:

  • Any model. The answer can come from the session's own model, or from Fable, Opus, Sonnet or Haiku at any effort.
  • Read-only tools. An aside can read files, search, and run read-only commands before it answers.
  • A side thread. Each question carries the earlier answered ones, so you can go back and forth.
  • A hand-off. One key sends an answer into the session, or puts it in the prompt for you to edit first.

It also keeps an eye on the prompt cache, and tells you when another model would cost more than the session's own.

Asking

  • /aside <question>: asks with the pane's current settings and opens the pane. It works while Claude is busy.
  • /aside: opens the pane. Type a question in its field and press Enter.
  • Flags in front of a question override the pane's settings for that question:
  • -m <model>: session, fable, opus, sonnet or haiku.
  • -e <effort>: low, medium, high, xhigh or max.
  • -t turns tools on, and -T turns them off.

For example: /aside -m fable -e high -t is this migration safe to run twice?

  • /aside model <m>, /aside effort <e> and /aside tools on|off change the pane's settings.
  • /aside clear empties the side thread.

Handing an answer over

Only the newest answer has buttons, with keys: s sends it to the session as a prompt, e puts it in the prompt box, and c copies it. Each older answer shows its number. To hand over answer n, run /aside send n or /aside edit n. Without a number, they act on the newest answer.

How each aside is asked, and what it costs

SettingsHow it runsCache
Session model, default effort, no toolsIt forks the session's own last request and puts the question after it.It reads the session's cache, so it is usually the cheapest.
Session model, default effort, toolsIt forks a subagent from the session.It reads the session's cache, plus whatever its tool calls cost.
Another model or effort, no toolsIt makes a request of its own, with the conversation as a transcript.It writes its own cache, and later asides on that model read it.
Another model or effort, toolsIt starts a fresh subagent and hands it the transcript.It caches nothing across asides.

A transcript holds up to transcriptTokens of the conversation, newest first, and cuts tool results short. A model that reads one sees less than the session's own model does.

The model picker shows an estimated cost next to each choice. Switching to a cheaper model is not always cheaper. On a long conversation with a warm cache, the session's model reads the whole context at the cache-read price. Another model has to write that context to its own cache first. When the model you pick would cost more than the session's own, Aside holds the question and shows why, for example:

Fable 5.1 costs about $0.70: a request of its own reads 49k tokens of this conversation cold. Opus 5.5, the session's model, reads its whole 320k-token context from cache for about $0.11.

You can then press y to ask the session's model, a to ask anyway, or x to drop the question. Set confirmSwitch to false to skip this.

The estimates use Claude API list prices. Each one starts as a rough guess and gets more accurate as Aside learns from real answers:

  • Tokens: a transcript's token count starts as an estimate from its length. Each answer from another model reports the real count, and later estimates scale to match it.
  • Answer length: each model's answer length is learned from its answers, starting from 1,500 tokens.
  • Subagents: a subagent run with tools first counts only its first request. After that, each run is scaled by what earlier runs on the same model cost against their estimates.

What Aside learns is kept across sessions. Once an aside is answered, its line in the pane shows what it actually cost and how much it read from the cache.

The read-only guard

An aside's subagent can call Read, Grep, Glob, WebSearch, WebFetch, and Bash for read-only commands: git status, git log, git diff, rg, ls, cat, sed -n, and the like. It refuses any other tool and any command it does not recognize. It also refuses redirection, chaining, substitution, and flags that write, such as sort -o, find -delete or git log --output. An aside never stops to ask you for permission: anything that would ask is refused.

Settings

Change them in /config, or under pluginConfigs.aside in your settings.

SettingDefaultWhat it does
modelsessionThe model the pane starts on.
toolsfalseWhether the pane starts with read-only tools on.
confirmSwitchtrueHold a question when the model you picked would cost more than the session's own.
cacheTtlautoHow long the session's cache lives. auto reads it from the session transcript's last cache write at the end of a turn. If the transcript is over 4 MiB, it starts at 5 minutes and switches to 1 hour once a fork finds the cache still held after more than 5 minutes.
transcriptTokens100000The most of the conversation a transcript holds.
threadTurns10How many earlier answered asides each question carries.

Requirements

Claude Code v2.1.296 or later. Mods are an early access part of Claude Code, so a Claude Code release can break the plugin until it is updated.

Install

/plugin install aside --marketplace astrosteveo/claude-plugins
Source 4 files
hooks/register.tsx 660 lines
1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, ModelTextBlock, ModelUsage, PluginOptions, Register } from 'claude-code'
3
4import type { Choice, Effort, Exchange, Route, Picks, Spent, Ttl } from '../types'
5import * as cost from './cost'
6import * as p from './prompt'
7
8const PANE = 'aside'
9const thread = atom({ plugin: 'aside', key: 'thread' } as const, [])
10const chosen = atom({ plugin: 'aside', key: 'defaults' } as const, null)
11const pending = atom({ plugin: 'aside', key: 'pending' } as const, null)
12const quote = atom({ plugin: 'aside', key: 'quote' } as const, null)
13const caches = atom({ plugin: 'aside', key: 'caches' } as const, {})
14const agents = atom({ plugin: 'aside', key: 'agents' } as const, [])
15const main = atom({ plugin: 'aside', key: 'main' } as const, { at: null, isRunning: false })
16
17const MAX_THREAD = 50
18const MAX_TOKENS = 8_000
19
20type Options = {
21  model: Choice
22  tools: boolean
23  confirmSwitch: boolean
24  cacheTtl: 'auto' | Ttl
25  transcriptTokens: number
26  threadTurns: number
27}
28
29function optionsOf(o: PluginOptions): Options {
30  const num = (v: unknown, d: number) => (typeof v === 'number' && Number.isFinite(v) && v >= 0 ? v : d)
31
32  return {
33    model: p.CHOICES.includes(o.model as Choice) ? (o.model as Choice) : 'session',
34    tools: o.tools === true,
35    confirmSwitch: o.confirmSwitch !== false,
36    cacheTtl: o.cacheTtl === '5m' || o.cacheTtl === '1h' ? o.cacheTtl : 'auto',
37    transcriptTokens: num(o.transcriptTokens, 100_000),
38    threadTurns: num(o.threadTurns, 10),
39  }
40}
41
42// The subagents asides started: their tool calls are held to reading, and
43// their reports are kept out of the session. Restored from state on each load.
44const ours = new Set<string>()
45// Spawns in flight, whose subagent may call a tool before its id is known.
46const spawning = new Set<Promise<unknown>>()
47let counter = 0
48
49const ADVISOR = 'advisor'
50const advisorType = (effort: Effort) => `aside:${effort === 'default' ? ADVISOR : `${ADVISOR}-${effort}`}`
51
52function spentOf(u: ModelUsage | undefined): Spent | undefined {
53  if (u === undefined) return undefined
54
55  return { input: u.input_tokens, output: u.output_tokens, read: u.cache_read_input_tokens, write: u.cache_creation_input_tokens }
56}
57
58async function settingsOf($: EngineInterface, o: Options): Promise<Picks> {
59  return (await read($, chosen)) ?? { model: o.model, effort: 'default', tools: o.tools }
60}
61
62// Figures by model id kept across sessions: `scale`, real tokens per
63// estimated one; `answers`, how long answers run; `runs`, what subagent runs
64// cost against their estimate.
65async function figures($: EngineInterface, key: 'scale' | 'answers' | 'runs'): Promise<Record<string, number>> {
66  const v = await $.store.get(key)
67
68  return typeof v === 'object' && v !== null ? (v as Record<string, number>) : {}
69}
70
71const scaleOf = ($: EngineInterface) => figures($, 'scale')
72
73async function ttlOf($: EngineInterface, o: Options): Promise<Ttl> {
74  if (o.cacheTtl !== 'auto') return o.cacheTtl
75
76  return (await $.store.get('ttl')) === '1h' ? '1h' : '5m'
77}
78
79async function sessionModel($: EngineInterface): Promise<cost.Model> {
80  const name = await $.session.model()
81  // A model of no family the table knows is priced as the newest Opus.
82  const opus = cost.resolve('opus') as cost.Model
83
84  return cost.resolve(name) ?? { ...opus, id: name, name }
85}
86
87// Who answers and how, for these settings: the session's own model at its
88// own effort forks the session's request and reads its cache; anything else
89// is a request or a subagent of the aside's own.
90function planOf(s: Picks, session: cost.Model): { plan: cost.Plan; modelArg: string } {
91  const model = s.model === 'session' ? session : (cost.resolve(s.model) ?? session)
92  const isFork = model.id === session.id && s.effort === 'default'
93  const route: Route = isFork ? (s.tools ? 'fork-agent' : 'fork') : s.tools ? 'agent' : 'model'
94
95  return { plan: { route, model }, modelArg: model.id === session.id ? session.id : s.model }
96}
97
98type Situation = { s: cost.Situation; rendered: string[]; now: number; mainAt: number | null; isMainRunning: boolean }
99
100async function situation($: EngineInterface, o: Options, askTokens: number): Promise<Situation> {
101  const [session, usage, messages, m, tracks, ttl, now, scale, answers, runs] = await Promise.all([
102    sessionModel($),
103    $.session.usage(),
104    $.session.messages(),
105    read($, main),
106    read($, caches),
107    ttlOf($, o),
108    $.clock.now(),
109    scaleOf($),
110    figures($, 'answers'),
111    figures($, 'runs'),
112  ])
113  const rendered = messages.map(p.renderMessage)
114  const cached: Record<string, number> = {}
115  for (const [id, track] of Object.entries(tracks)) cached[id] = p.cachedTokens(rendered, track, now)
116
117  return {
118    s: {
119      session,
120      contextTokens: usage.context.tokens ?? 0,
121      isMainWarm: m.isRunning || (m.at !== null && now - m.at < cost.TTL_MS[ttl]),
122      ttl,
123      transcriptTokens: p.transcript(rendered, o.transcriptTokens, undefined, now).tokens,
124      cached,
125      askTokens,
126      scale,
127      answers,
128      runs,
129    },
130    rendered,
131    now,
132    mainAt: m.at,
133    isMainRunning: m.isRunning,
134  }
135}
136
137const stayPlan = (s: Picks, session: cost.Model): cost.Plan => ({ route: s.tools ? 'fork-agent' : 'fork', model: session })
138
139// What the next question would cost on each choice, for the pane.
140async function refreshQuote($: EngineInterface, o: Options) {
141  const [s, list] = await Promise.all([settingsOf($, o), read($, thread)])
142  const at = await situation($, o, p.estimateTokens(p.threadText(list, o.threadTurns)) + 100)
143  const costs: Partial<Record<Choice, number>> = {}
144  for (const c of p.CHOICES) costs[c] = cost.estimate(planOf({ ...s, model: c }, at.s.session).plan, at.s)
145  await update($, quote, () => ({
146    at: at.now,
147    session: at.s.session.name,
148    contextTokens: at.s.contextTokens,
149    transcriptTokens: at.s.transcriptTokens,
150    isMainWarm: at.s.isMainWarm,
151    ttl: at.s.ttl,
152    costs,
153  }))
154}
155
156async function patch($: EngineInterface, id: string, change: Partial<Exchange>) {
157  await update($, thread, list => list.map(x => (x.id === id ? { ...x, ...change } : x)))
158}
159
160async function settle($: EngineInterface, o: Options, id: string, change: Partial<Exchange>) {
161  const now = await $.clock.now()
162  const was = (await read($, thread)).find(x => x.id === id)
163  await patch($, id, { ...change, ms: was === undefined ? undefined : now - was.at })
164  // A single request's answer length, for the next estimate; a subagent's
165  // output runs through its tool calls too, so it says nothing of one answer.
166  const model = was === undefined ? undefined : cost.resolve(was.model)
167  const route = change.route ?? was?.route
168  if (change.status === 'done' && change.spent !== undefined && model !== undefined && (route === 'fork' || route === 'model')) {
169    await $.store.set('answers', cost.lengthen(await figures($, 'answers'), model.id, change.spent.output))
170  }
171  if (!isShown) $.ui.toast(change.status === 'done' ? `Aside answered. /aside shows it.` : `Aside failed: ${change.error ?? 'no answer'}`)
172  void refreshQuote($, o).catch(() => {})
173}
174
175// Asks one question with these settings. A pricier pick than the session's
176// own cache waits for the person's word unless `force` gives it.
177async function ask($: EngineInterface, o: Options, question: string, settings: Picks, force = false) {
178  const list = await read($, thread)
179  const carried = p.threadText(list, o.threadTurns)
180  const at = await situation($, o, p.estimateTokens(carried + question) + 100)
181  const { plan, modelArg } = planOf(settings, at.s.session)
182  const mine = cost.estimate(plan, at.s)
183  const stay = cost.estimate(stayPlan(settings, at.s.session), at.s)
184  if (!force && o.confirmSwitch && plan.route !== 'fork' && plan.route !== 'fork-agent') {
185    const why = cost.warning(plan, mine, stay, at.s)
186    if (why !== undefined) {
187      await update($, pending, () => ({ question, settings, usd: mine, stayUsd: stay, model: plan.model.name, session: at.s.session.name, why }))
188      await openPane($)
189      if (!isShown) $.ui.toast(`Aside held: ${why} /aside to choose.`)
190
191      return
192    }
193  }
194  await update($, pending, () => null)
195  const ex: Exchange = {
196    id: `${at.now.toString(36)}-${++counter}`,
197    at: at.now,
198    question,
199    model: plan.model.name,
200    route: plan.route,
201    tools: settings.tools,
202    effort: settings.effort,
203    status: 'running',
204    estimate: mine,
205    stayEstimate: stay,
206    ...(force ? { forced: true } : {}),
207  }
208  await update($, thread, l => [...l, ex].slice(-MAX_THREAD))
209  void run($, o, ex, plan, modelArg, carried, at).catch(err => settle($, o, ex.id, { status: 'failed', error: String(err) }))
210}
211
212async function run($: EngineInterface, o: Options, ex: Exchange, plan: cost.Plan, modelArg: string, carried: string, at: Situation) {
213  if (plan.route === 'fork') {
214    const gap = at.isMainRunning ? 0 : at.mainAt === null ? undefined : at.now - at.mainAt
215    const r = await $.model.fork({ prompt: p.forkPrompt(carried, ex.question, false) })
216    if (!r.isAnswered && r.reason === 'nothing-to-fork') {
217      // No conversation yet: ask the session's model on its own.
218      await patch($, ex.id, { route: 'model' })
219
220      return run($, o, { ...ex, route: 'model' }, { route: 'model', model: plan.model }, modelArg, carried, at)
221    }
222    const spent = spentOf(r.usage)
223    // A fork reads the session's cache entry, which refreshes it.
224    if (spent !== undefined && spent.read > 0) await update($, main, m => ({ ...m, at: Math.max(m.at ?? 0, ex.at) }))
225    await learnTtl($, o, gap, spent, at.s.contextTokens)
226    const usd = spent === undefined ? undefined : cost.usd(plan.model, spent, at.s.ttl)
227    if (r.isAnswered) return settle($, o, ex.id, { status: 'done', answer: r.text, spent, usd })
228
229    return settle($, o, ex.id, { status: 'failed', error: failure(r), spent, usd })
230  }
231
232  if (plan.route === 'model') {
233    const tracks = await read($, caches)
234    const t = p.transcript(at.rendered, o.transcriptTokens, tracks[plan.model.id], at.now)
235    const marked = (i: number) => i >= t.stretches.length - 2
236    const prompt: ModelTextBlock[] = [
237      { text: p.TRANSCRIPT_HEAD },
238      ...t.stretches.map((text, i) => (marked(i) ? { text: text || '…', cache: true as const } : { text: text || '…' })),
239      { text: p.askText(carried, ex.question) },
240    ]
241    const r = await $.model.complete({
242      model: modelArg,
243      system: p.SYSTEM,
244      prompt,
245      maxTokens: MAX_TOKENS,
246      ...(ex.effort === 'default' ? {} : { effort: ex.effort }),
247    })
248    const spent = spentOf(r.usage)
249    const usd = spent === undefined ? undefined : cost.usd(plan.model, spent, '5m')
250    if (!r.isAnswered) return settle($, o, ex.id, { status: 'failed', error: failure(r), spent, usd })
251    await update($, caches, all => ({ ...all, [plan.model.id]: t.track }))
252    // What the model counted against what was estimated, for the next estimate.
253    if (spent !== undefined) {
254      const estimated = cost.SYSTEM_TOKENS + t.tokens + p.estimateTokens(p.TRANSCRIPT_HEAD + p.askText(carried, ex.question))
255      await $.store.set('scale', cost.rescale(await scaleOf($), plan.model.id, spent.input + spent.read + spent.write, estimated))
256    }
257
258    return settle($, o, ex.id, { status: 'done', answer: r.text, spent, usd })
259  }
260
261  // A subagent with read-only tools: a fork of the session, which reads its
262  // cache, or a fresh advisor on another model or effort, handed a transcript.
263  const brief = () => {
264    const t = p.transcript(at.rendered, o.transcriptTokens, undefined, at.now)
265
266    return `${p.TRANSCRIPT_HEAD}${t.stretches.join('\n')}${p.askText(carried, ex.question)}`
267  }
268  const description = `Aside: ${ex.question.slice(0, 40)}`
269  const spawn = async (args: Parameters<EngineInterface['agent']['spawn']>[0]) => {
270    const call = $.agent.spawn(args)
271    spawning.add(call)
272    try {
273      const r = await call
274      if (r.agentId !== undefined) {
275        ours.add(r.agentId)
276        await update($, agents, l => [...l, r.agentId as string].slice(-100))
277      }
278
279      return r
280    } finally {
281      spawning.delete(call)
282    }
283  }
284  let r = plan.route === 'fork-agent'
285    ? await spawn({ subagentType: 'fork', prompt: p.forkPrompt(carried, ex.question, true), description }).catch(() => undefined)
286    : await spawn({ subagentType: advisorType(ex.effort), model: modelArg, prompt: brief(), description })
287  if (plan.route === 'fork-agent' && (r === undefined || r.agentId === undefined)) {
288    // No fork to be had: an advisor on the session's model, handed a transcript.
289    await patch($, ex.id, { route: 'agent' })
290    r = await spawn({ subagentType: advisorType(ex.effort), model: modelArg, prompt: brief(), description })
291  }
292  if (r === undefined || r.agentId === undefined) {
293    return settle($, o, ex.id, { status: 'failed', error: r?.deny ?? 'The subagent did not start.' })
294  }
295  await patch($, ex.id, { agentId: r.agentId })
296}
297
298function failure(r: { reason: string; status?: number | null; error?: unknown }): string {
299  if (r.reason === 'api-error') return `API error${r.status ? ` ${r.status}` : ''}${r.error ? `: ${String(r.error)}` : ''}`
300  if (r.reason === 'empty-reply') return 'The model gave no answer.'
301  if (r.reason === 'aborted') return 'Cut short.'
302
303  return r.reason
304}
305
306// Whether this load has read the session cache's lifetime from the transcript.
307let isTtlRead = false
308
309// The session's cache lifetime, as the transcript records its last cache
310// write: no call on `$` reports it. A transcript past what one read takes
311// (4 MiB) is left to what the forks show.
312async function readTtl($: EngineInterface, path: string) {
313  const lines = (await $.fs.read(path)).split('\n')
314  for (let i = lines.length - 1; i >= 0; i--) {
315    const line = lines[i] ?? ''
316    if (!line.includes('"ephemeral_')) continue
317    const hour = Number(/"ephemeral_1h_input_tokens":\s*(\d+)/.exec(line)?.[1] ?? 0)
318    const five = Number(/"ephemeral_5m_input_tokens":\s*(\d+)/.exec(line)?.[1] ?? 0)
319    if (hour === 0 && five === 0) continue
320    await $.store.set('ttl', hour > 0 ? '1h' : '5m')
321    isTtlRead = true
322
323    return
324  }
325}
326
327// What a fork's cache read says about the session cache's lifetime, where
328// the transcript did not say: a read past five minutes means an hour; a miss
329// well inside an hour means five.
330async function learnTtl($: EngineInterface, o: Options, gap: number | undefined, spent: Spent | undefined, prefix: number) {
331  if (o.cacheTtl !== 'auto' || isTtlRead || gap === undefined || spent === undefined || prefix === 0) return
332  if (gap <= cost.TTL_MS['5m'] || gap >= cost.TTL_MS['1h']) return
333  if (spent.read > prefix / 2) await $.store.set('ttl', '1h')
334  else if (spent.read < prefix / 10) await $.store.set('ttl', '5m')
335}
336
337// Hands an answer to the session: sent as a prompt of its own, or put in the
338// prompt box to edit first.
339async function handOff($: EngineInterface, how: 'send' | 'edit', index: number | undefined) {
340  const list = await read($, thread)
341  const ex = index === undefined ? list.filter(x => x.status === 'done').at(-1) : list[index - 1]
342  if (ex === undefined || ex.status !== 'done' || ex.answer === undefined) {
343    $.ui.toast(index === undefined ? 'No answered aside to hand over yet.' : `Aside ${index} has no answer.`)
344
345    return
346  }
347  const text = `I asked ${ex.model} on the side: ${ex.question}\n\nIts answer:\n\n${ex.answer}`
348  if (how === 'edit') {
349    const r = await $.prompt.fill({ text, mode: 'insert' })
350    if (!r.isFilled) {
351      $.ui.toast('The prompt box could not take the answer.')
352
353      return
354    }
355  } else void $.prompt.submit({ text })
356  await patch($, ex.id, { handedOff: how === 'send' ? 'sent' : 'edited' })
357}
358
359let isShown = false
360
361async function openPane($: EngineInterface, focus = false) {
362  const r = await $.ui.open({ id: PANE, title: 'Aside', ...(focus ? { focus: true as const } : {}) })
363  isShown = r.isPlaced
364}
365
366async function setSettings($: EngineInterface, o: Options, change: Partial<Picks>) {
367  const s = await settingsOf($, o)
368  await update($, chosen, () => ({ ...s, ...change }))
369  void refreshQuote($, o).catch(() => {})
370}
371
372function choiceName(c: Choice, session: string | undefined): string {
373  if (c === 'session') return session === undefined ? 'Session' : `Session (${session})`
374
375  return cost.resolve(c)?.name ?? c
376}
377
378function meta(x: Exchange): string {
379  const parts = [x.model]
380  if (x.tools) parts.push('tools')
381  if (x.effort !== 'default') parts.push(x.effort)
382  if (x.route === 'fork' || x.route === 'fork-agent') parts.push('session cache')
383  if (x.ms !== undefined) parts.push(`${Math.round(x.ms / 1000)}s`)
384  if (x.usd !== undefined) parts.push(x.route === 'model' || x.route === 'fork' ? cost.money(x.usd) : `${cost.money(x.usd)} incl. tools`)
385  // Off the session's cache, what it was estimated at against staying, and
386  // whether a hold was overridden, so a pricier answer says how it got here.
387  if (x.route === 'model' || x.route === 'agent') {
388    const vs = x.stayEstimate === undefined ? '' : ` vs ≈${cost.money(x.stayEstimate)} staying`
389    if (x.estimate !== undefined) parts.push(`est. ≈${cost.money(x.estimate)}${vs}`)
390    if (x.forced) parts.push('asked anyway after a hold')
391  } else if (x.usd === undefined && x.estimate !== undefined) parts.push(`≈${cost.money(x.estimate)}`)
392  if (x.spent !== undefined) {
393    const { input, read, write, output } = x.spent
394    parts.push(`${cost.tokens(input + read + write)} in${read > 0 ? ` (${cost.tokens(read)} cached)` : ''} · ${cost.tokens(output)} out`)
395  }
396  if (x.handedOff !== undefined) parts.push(x.handedOff === 'sent' ? 'sent to the session' : 'put in the prompt')
397
398  return parts.join(' · ')
399}
400
401export const register: Register = (on, options) => {
402  const o = optionsOf(options)
403
404  on('session.start', async ($, e, next) => {
405    const result = await next(e)
406    for (const id of await read($, agents)) ours.add(id)
407    try {
408      await $.command.register({
409        name: 'aside',
410        description: 'Ask a side question without touching the session: any model, read-only tools, a pane of its own',
411        argumentHint: '[-m model] [-e effort] [-t] question | send | edit | clear',
412        immediate: true,
413      })
414    } catch {
415      // A clash with another command leaves the pane's own input working.
416    }
417    for (const effort of p.EFFORTS) {
418      await $.agent
419        .register({
420          name: effort === 'default' ? ADVISOR : `${ADVISOR}-${effort}`,
421          description: 'Answers a side question for /aside. Started by the aside mod only.',
422          prompt: p.AGENT_SYSTEM,
423          permissionMode: 'dontAsk',
424          disallowedTools: ['Edit', 'Write', 'NotebookEdit', 'Agent'],
425          maxTurns: 40,
426          ...(effort === 'default' ? {} : { effort }),
427        })
428        .catch(() => {})
429    }
430    void refreshQuote($, o).catch(() => {})
431
432    return result
433  })
434
435  // The advisor types are the mod's own: the session's model never sees them.
436  on('agent.offer', async ($, e, next) => (e.agent.startsWith('aside:') ? { isOffered: false } : next(e)))
437
438  on('turn.start', async ($, e, next) => {
439    await update($, main, m => ({ ...m, isRunning: true }))
440
441    return next(e)
442  })
443
444  on('turn.complete', async ($, e, next) => {
445    const result = await next(e)
446    if (e.agentId === undefined) {
447      const now = await $.clock.now()
448      await update($, main, () => ({ at: now, isRunning: false }))
449      void refreshQuote($, o).catch(() => {})
450
451      return result
452    }
453    if (!ours.has(e.agentId)) return result
454    const ex = (await read($, thread)).find(x => x.agentId === e.agentId)
455    if (ex === undefined || ex.status !== 'running') return result
456    const spent = spentOf(e.usage)
457    const model = (e.usage?.model !== undefined ? cost.resolve(e.usage.model) : undefined) ?? cost.resolve(ex.model)
458    const ttl = ex.route === 'fork-agent' ? await ttlOf($, o) : '5m'
459    const usd = spent === undefined || model === undefined ? undefined : cost.usd(model, spent, ttl)
460    if (e.reason === 'answer' && usd !== undefined && model !== undefined && ex.estimate !== undefined) {
461      await $.store.set('runs', cost.rerun(await figures($, 'runs'), ex.route, model.id, usd, ex.estimate))
462    }
463    if (e.reason === 'answer' && e.answer.trim() !== '') await settle($, o, ex.id, { status: 'done', answer: e.answer, spent, usd })
464    else await settle($, o, ex.id, { status: 'failed', error: e.reason === 'aborted' ? 'Stopped.' : 'The subagent gave no answer.', spent, usd })
465
466    return result
467  })
468
469  // A subagent's report would reach the session as a task notification:
470  // an aside's stays in its pane.
471  on('prompt.submit', async ($, e, next) => {
472    if (e.origin.kind === 'task-notification') {
473      for (const id of ours) if (e.text.includes(id)) return { drop: 'An aside finished: its answer is in the aside pane.' }
474    }
475
476    return next(e)
477  })
478
479  // An aside's subagent only reads. While one of its spawns is in flight, a
480  // subagent the mod cannot place yet waits for it before it is judged.
481  const judged = (agentId: string | undefined, tool: string, input: unknown): string | undefined =>
482    agentId !== undefined && (ours.has(agentId) || spawning.size > 0) ? p.refusal(tool, input) : undefined
483  on('tool.call', async ($, e, next) => {
484    if (e.agentId !== undefined && !ours.has(e.agentId) && spawning.size > 0) await Promise.allSettled([...spawning])
485    const why = e.agentId !== undefined && ours.has(e.agentId) ? p.refusal(String(e.tool), e) : undefined
486
487    return why === undefined ? next(e) : { deny: why }
488  }).catch(($, e, next) => {
489    if (next.called) return next(e)
490    const why = judged(e.agentId, String(e.tool), e)
491
492    return why === undefined ? next(e) : { deny: why }
493  })
494
495  // Nor does it ever stop to ask the person for permission.
496  on('tool.check', async ($, e, next) => {
497    const verdict = await next(e)
498    if (e.agentId === undefined || !ours.has(e.agentId) || verdict.decision !== 'ask') return verdict
499
500    return { decision: 'deny', reason: 'An aside never asks for permission.' }
501  })
502
503  on('command.run', { command: 'aside' }, async ($, e) => {
504    const c = p.parse(e.args)
505    if (c.kind === 'open') await openPane($, true)
506    else if (c.kind === 'ask') {
507      const s = await settingsOf($, o)
508      const settings: Picks = { model: c.model ?? s.model, effort: c.effort ?? s.effort, tools: c.tools ?? s.tools }
509      await openPane($)
510      void ask($, o, c.question, settings).catch(err => $.ui.toast(`Aside failed: ${String(err)}`))
511    } else if (c.kind === 'send' || c.kind === 'edit') await handOff($, c.kind, c.index)
512    else if (c.kind === 'clear') {
513      await update($, thread, () => [])
514      await update($, pending, () => null)
515      $.ui.toast('Aside thread cleared.')
516    } else if (c.kind === 'model') await setSettings($, o, { model: c.model })
517    else if (c.kind === 'effort') await setSettings($, o, { effort: c.effort })
518    else if (c.kind === 'tools') await setSettings($, o, { tools: c.tools })
519    else $.ui.toast(c.text)
520
521    // Nothing of the command reaches the transcript.
522    return {}
523  })
524
525  // At the end of a main turn, until this load has it, the cache lifetime.
526  on('classic.Stop', async ($, e, next) => {
527    const result = await next(e)
528    if (o.cacheTtl === 'auto' && !isTtlRead) {
529      await readTtl($, e.transcript_path).catch(() => {})
530      void refreshQuote($, o).catch(() => {})
531    }
532
533    return result
534  })
535
536  on('ui.close', async ($, e, next) => {
537    if (e.id === PANE) isShown = false
538
539    return next(e)
540  })
541
542  on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
543    const { Box, Text, Button, Markdown } = $.ui.resolve(e)
544    const [list, s, held, q] = await Promise.all([read($, thread), settingsOf($, o), read($, pending), read($, quote)])
545    const asked = (question: string, settings: Picks, force = false) =>
546      void ask($, o, question, settings, force).catch(err => $.ui.toast(`Aside failed: ${String(err)}`))
547    const priced = (c: Choice) => {
548      const usd = q?.costs[c]
549
550      return `${choiceName(c, q?.session)}${usd === undefined ? '' : ` ≈${cost.money(usd)}`}`
551    }
552
553    const controls =
554      e.surface === 'mobile' ? (
555        <Text dimColor>Ask with /aside from a terminal or the desktop.</Text>
556      ) : (
557        (() => {
558          const { Input, Select } = $.ui.resolve(e)
559
560          return (
561            <Box flexDirection="column">
562              <Input key="ask" label="Ask" placeholder="a side question" submitLabel="ask" autoFocus onSubmit={v => (v.trim() === '' ? undefined : asked(v.trim(), s))} />
563              <Box flexDirection="row" flexWrap="wrap" columnGap={2}>
564                <Select
565                  key="model"
566                  label="Model"
567                  value={s.model}
568                  options={p.CHOICES.map(c => ({ value: c, label: priced(c) }))}
569                  onSelect={v => void setSettings($, o, { model: v as Choice })}
570                />
571                <Select
572                  key="effort"
573                  label="Effort"
574                  value={s.effort}
575                  options={p.EFFORTS.map(x => ({ value: x }))}
576                  onSelect={v => void setSettings($, o, { effort: v as Effort })}
577                />
578                <Button key="tools" hotkey="t" onPress={() => void setSettings($, o, { tools: !s.tools })}>
579                  {s.tools ? 'Tools: read-only' : 'Tools: off'}
580                </Button>
581              </Box>
582            </Box>
583          )
584        })()
585      )
586
587    const cacheLine =
588      q === null
589        ? undefined
590        : q.contextTokens === 0
591          ? 'No conversation yet.'
592          : `${cost.tokens(q.contextTokens)} tokens in context, ${q.isMainWarm ? 'cache warm' : 'cache likely cold'} (${q.ttl} cache).`
593
594    const latest = list.filter(x => x.status === 'done').at(-1)?.id
595    const rows = list
596      .map((x, i) => ({ x, n: i + 1 }))
597      .reverse()
598      .map(({ x, n }) => {
599        const isLatest = x.id === latest
600
601        return (
602          <Box key={x.id} flexDirection="column" marginTop={1}>
603            <Text bold>
604              {n}. {x.question}
605            </Text>
606            {x.status === 'running' && <Text dimColor>{x.model} is thinking…</Text>}
607            {x.status === 'done' && <Markdown text={x.answer ?? ''} />}
608            {x.status === 'failed' && <Text color="error">{x.error ?? 'No answer.'}</Text>}
609            <Text dimColor>{meta(x)}</Text>
610            {/* Only the newest answer has buttons, so a press can't land on an older one by mistake. */}
611            {x.status === 'done' && isLatest && (
612              <Box flexDirection="row" columnGap={1}>
613                <Button key="send" hotkey="s" onPress={() => void handOff($, 'send', n)}>
614                  {x.handedOff === 'sent' ? 'Send again' : 'Send to session'}
615                </Button>
616                <Button key="edit" hotkey="e" onPress={() => void handOff($, 'edit', n)}>
617                  Edit in prompt
618                </Button>
619                <Button key="copy" hotkey="c" onPress={pe => void $.ui.copy({ text: x.answer ?? '', surface: pe.surface })}>
620                  Copy
621                </Button>
622              </Box>
623            )}
624            {x.status === 'done' && !isLatest && (
625              <Text dimColor>
626                /aside send {n} · /aside edit {n}
627              </Text>
628            )}
629          </Box>
630        )
631      })
632
633    return (
634      <Box flexDirection="column" width={e.props.bodyColumns}>
635        {controls}
636        {cacheLine !== undefined && <Text dimColor>{cacheLine}</Text>}
637        {held !== null && (
638          <Box flexDirection="column" marginTop={1} borderStyle="round" borderColor="warning" paddingX={1}>
639            <Text color="warning">{held.why}</Text>
640            <Text>Held: {held.question}</Text>
641            <Box flexDirection="row" columnGap={1}>
642              <Button key="stay" hotkey="y" variant="primary" onPress={() => asked(held.question, { ...held.settings, model: 'session', effort: 'default' })}>
643                Ask {held.session} instead
644              </Button>
645              <Button key="anyway" hotkey="a" onPress={() => asked(held.question, held.settings, true)}>
646                Ask {held.model} anyway
647              </Button>
648              <Button key="drop" hotkey="x" onPress={() => void update($, pending, () => null)}>
649                Drop it
650              </Button>
651            </Box>
652          </Box>
653        )}
654        {list.length === 0 && held === null && <Text dimColor>Nothing asked yet. Nothing here reaches the session unless you send it.</Text>}
655        {rows}
656      </Box>
657    )
658  })
659}
660
hooks/cost.ts 212 lines
1import type { Spent, Ttl } from '../types'
2
3// List prices on the Claude API in US dollars per million tokens: uncached
4// input, output and a cache read. A cache write is input times WRITE.
5export type Price = { input: number; output: number; read: number }
6
7export type Model = {
8  id: string
9  name: string
10  family: string
11  price: Price
12  // A price that takes over once a request's prompt passes `at` tokens.
13  over?: { at: number; price: Price }
14  // The shortest prefix the API caches.
15  minCache: number
16}
17
18const FABLE_51 = { input: 10, output: 50, read: 0.25 }
19const FABLE = { input: 10, output: 50, read: 1 }
20const OPUS_55 = { input: 4, output: 20, read: 0.2 }
21const OPUS = { input: 5, output: 25, read: 0.5 }
22const SONNET_5 = { input: 2, output: 10, read: 0.2 }
23const SONNET_4 = { input: 3, output: 15, read: 0.3 }
24
25// Newest first within each family, so an alias resolves to its first entry.
26const MODELS: Model[] = [
27  { id: 'claude-fable-5-1', name: 'Fable 5.1', family: 'fable', price: FABLE_51, minCache: 512 },
28  { id: 'claude-fable-5', name: 'Fable 5', family: 'fable', price: FABLE, minCache: 512 },
29  { id: 'claude-mythos-5-1', name: 'Mythos 5.1', family: 'mythos', price: FABLE_51, minCache: 512 },
30  { id: 'claude-mythos-5', name: 'Mythos 5', family: 'mythos', price: FABLE, minCache: 512 },
31  { id: 'claude-opus-5-5', name: 'Opus 5.5', family: 'opus', price: OPUS_55, minCache: 512 },
32  { id: 'claude-opus-5', name: 'Opus 5', family: 'opus', price: OPUS, minCache: 512 },
33  { id: 'claude-opus-4-8', name: 'Opus 4.8', family: 'opus', price: OPUS, minCache: 1024 },
34  { id: 'claude-opus-4-7', name: 'Opus 4.7', family: 'opus', price: OPUS, minCache: 2048 },
35  { id: 'claude-opus-4-6', name: 'Opus 4.6', family: 'opus', price: OPUS, minCache: 4096 },
36  { id: 'claude-opus-4-5', name: 'Opus 4.5', family: 'opus', price: OPUS, minCache: 4096 },
37  { id: 'claude-sonnet-5-5', name: 'Sonnet 5.5', family: 'sonnet', price: SONNET_5, minCache: 512 },
38  { id: 'claude-sonnet-5', name: 'Sonnet 5', family: 'sonnet', price: SONNET_5, minCache: 1024 },
39  { id: 'claude-sonnet-4-6', name: 'Sonnet 4.6', family: 'sonnet', price: SONNET_4, minCache: 1024 },
40  { id: 'claude-sonnet-4-5', name: 'Sonnet 4.5', family: 'sonnet', price: SONNET_4, minCache: 1024 },
41  {
42    id: 'claude-haiku-5-5',
43    name: 'Haiku 5.5',
44    family: 'haiku',
45    price: { input: 0.1, output: 0.5, read: 0.01 },
46    over: { at: 100_000, price: { input: 0.5, output: 2.5, read: 0.05 } },
47    minCache: 512,
48  },
49  { id: 'claude-haiku-4-5', name: 'Haiku 4.5', family: 'haiku', price: { input: 1, output: 5, read: 0.1 }, minCache: 4096 },
50]
51
52const WRITE: Record<Ttl, number> = { '5m': 1.25, '1h': 2 }
53export const TTL_MS: Record<Ttl, number> = { '5m': 5 * 60_000, '1h': 60 * 60_000 }
54
55// The model a name means: a full id (any provider prefix, date or context
56// suffix), a display name, or a bare alias, which means the family's newest.
57// Undefined for a name of no family it knows.
58export function resolve(name: string): Model | undefined {
59  const m = /(fable|mythos|opus|sonnet|haiku)(?:[-_ ]?(\d)(?:[-._](\d)(?!\d))?)?/i.exec(name)
60  if (m === null) return undefined
61  const family = (m[1] ?? '').toLowerCase()
62  const inFamily = MODELS.filter(x => x.family === family)
63  if (m[2] === undefined) return inFamily[0]
64  const id = `claude-${family}-${m[2]}${m[3] === undefined ? '' : `-${m[3]}`}`
65
66  return inFamily.find(x => x.id === id) ?? inFamily[0]
67}
68
69function priceAt(model: Model, promptTokens: number): Price {
70  return model.over !== undefined && promptTokens > model.over.at ? model.over.price : model.price
71}
72
73// What a request cost: its four token counts at the model's prices, cache
74// writes at the given lifetime's rate.
75export function usd(model: Model, t: Spent, ttl: Ttl = '5m'): number {
76  const p = priceAt(model, t.input + t.read + t.write)
77
78  return (t.input * p.input + t.read * p.read + t.write * p.input * WRITE[ttl] + t.output * p.output) / 1e6
79}
80
81// What an answer is assumed to run to, for the estimate; thinking included.
82export const ANSWER_TOKENS = 1_500
83// What a fresh subagent's own system prompt and tool list come to.
84export const AGENT_OVERHEAD = 10_000
85// The aside's own instructions on a request of its own.
86export const SYSTEM_TOKENS = 300
87
88// The figures an estimate reads.
89export type Situation = {
90  session: Model
91  // What the main thread's last request carried, and whether its cache is
92  // likely still held.
93  contextTokens: number
94  isMainWarm: boolean
95  ttl: Ttl
96  // The conversation as a transcript for another model, and how much of it
97  // each model's cache likely holds now, by model id.
98  transcriptTokens: number
99  cached: Record<string, number>
100  // The question and the side thread it carries.
101  askTokens: number
102  // Real tokens per estimated one, as earlier answers on each model counted
103  // them, by model id; `*` across models. A model with no count uses `*`.
104  scale: Record<string, number>
105  // How long each model's answers have run, in output tokens, by model id.
106  answers: Record<string, number>
107  // What subagent runs cost against their first-request estimate, by
108  // `<route>:<model id>`, and by route across models.
109  runs: Record<string, number>
110}
111
112export function runFactor(s: Pick<Situation, 'runs'>, route: string, id: string): number {
113  return s.runs[`${route}:${id}`] ?? s.runs[route] ?? 1
114}
115
116// Folds one subagent run's cost, against what was estimated with the factor
117// then in force, into the factor for its route and model.
118export function rerun(runs: Record<string, number>, route: string, id: string, actual: number, estimated: number): Record<string, number> {
119  if (actual <= 0 || estimated <= 0) return runs
120  const used = runFactor({ runs }, route, id)
121  const target = Math.min(5, Math.max(0.5, (used * actual) / estimated))
122  const fold = (was: number | undefined) => (was === undefined ? target : (was + target) / 2)
123
124  return { ...runs, [`${route}:${id}`]: fold(runs[`${route}:${id}`]), [route]: fold(runs[route]) }
125}
126
127// Folds one answer's length into a model's running figure.
128export function lengthen(answers: Record<string, number>, id: string, output: number): Record<string, number> {
129  if (output <= 0) return answers
130  const was = answers[id]
131
132  return { ...answers, [id]: was === undefined ? output : Math.round((was + output) / 2) }
133}
134
135export function scaleFor(s: Pick<Situation, 'scale'>, id: string): number {
136  return s.scale[id] ?? s.scale['*'] ?? 1
137}
138
139// Folds one answer's count of real against estimated tokens into the scale.
140export function rescale(scale: Record<string, number>, id: string, actual: number, estimated: number): Record<string, number> {
141  if (actual <= 0 || estimated <= 0) return scale
142  const ratio = Math.min(3, Math.max(0.5, actual / estimated))
143  const fold = (was: number | undefined) => (was === undefined ? ratio : (was + ratio) / 2)
144
145  return { ...scale, [id]: fold(scale[id]), '*': fold(scale['*']) }
146}
147
148export type Plan = { route: 'fork' | 'model' | 'fork-agent' | 'agent'; model: Model }
149
150// The estimated cost of one aside before it is asked. A subagent's is its
151// first request scaled by what earlier runs on its route and model cost
152// against theirs, which takes in its tool calls once it has run before.
153export function estimate(plan: Plan, s: Situation): number {
154  const first = request(plan, s)
155
156  return plan.route === 'agent' || plan.route === 'fork-agent' ? first * runFactor(s, plan.route, plan.model.id) : first
157}
158
159function request(plan: Plan, s: Situation): number {
160  const output = s.answers[plan.model.id] ?? ANSWER_TOKENS
161  if (plan.route === 'fork' || plan.route === 'fork-agent') {
162    const prefix = s.contextTokens
163    const t = s.isMainWarm
164      ? { input: s.askTokens, read: prefix, write: 0, output }
165      : { input: s.askTokens, read: 0, write: prefix, output }
166
167    return usd(plan.model, t, s.ttl)
168  }
169  const k = scaleFor(s, plan.model.id)
170  const overhead = plan.route === 'agent' ? AGENT_OVERHEAD : SYSTEM_TOKENS
171  const prefix = overhead + s.transcriptTokens * k
172  const input = s.askTokens * k
173  if (prefix < plan.model.minCache) return usd(plan.model, { input: prefix + input, read: 0, write: 0, output })
174  const read = plan.route === 'model' ? Math.min((s.cached[plan.model.id] ?? 0) * k, prefix) : 0
175
176  return usd(plan.model, { input, read, write: prefix - read, output })
177}
178
179// Under a tenth of a cent reads as such; under ten cents to the tenth of a
180// cent, since that is where asides differ; dollars to the cent past that.
181export function money(x: number): string {
182  if (x < 0.001) return '<$0.001'
183  if (x < 0.1) return `$${x.toFixed(3)}`
184
185  return `$${x.toFixed(2)}`
186}
187
188export function tokens(n: number): string {
189  if (n < 1_000) return String(Math.round(n))
190  if (n < 10_000) return `${(n / 1_000).toFixed(1).replace(/\.0$/, '')}k`
191  if (n < 1_000_000) return `${Math.round(n / 1_000)}k`
192
193  return `${(n / 1_000_000).toFixed(1).replace(/\.0$/, '')}M`
194}
195
196// Why an aside on `model` costs more than the session's own model would,
197// in a sentence for the person, or undefined when it does not.
198export function warning(plan: Plan, mine: number, stay: number, s: Situation): string | undefined {
199  if (mine <= stay * 1.05 || mine - stay < 0.001) return undefined
200  const how = plan.route === 'agent' ? 'a fresh subagent' : 'a request of its own'
201  const cold = plan.route === 'model' && (s.cached[plan.model.id] ?? 0) > 0 ? '' : ' cold'
202  const own = plan.model.id === s.session.id ? `${plan.model.name} at another effort` : plan.model.name
203  const read = tokens(s.transcriptTokens * scaleFor(s, plan.model.id))
204  const context = `its whole ${tokens(s.contextTokens)}-token context`
205  const stays = s.isMainWarm ? `reads ${context} from cache` : `would re-cache ${context}`
206
207  return (
208    `${own} costs about ${money(mine)}: ${how} reads ${read} tokens of this conversation${cold}. ` +
209    `${s.session.name}, the session's model, ${stays} for about ${money(stay)}.`
210  )
211}
212
hooks/prompt.ts 284 lines
1import type { SessionMessage } from 'claude-code'
2
3import type { CacheTrack, Choice, Effort, Exchange } from '../types'
4
5// A token is taken as 3.5 characters: code and prose sit either side of it.
6export function estimateTokens(text: string): number {
7  return Math.ceil(text.length / 3.5)
8}
9
10const RESULT_CHARS = 2_000
11const INPUT_CHARS = 300
12
13function clip(text: string, max: number): string {
14  return text.length <= max ? text : `${text.slice(0, max)}… (${text.length - max} more characters)`
15}
16
17// One message of the conversation as transcript text. A tool's result is
18// drawn under its call, so a user message that only carries results draws
19// nothing. The same message always draws the same text, which is what keeps
20// a cached transcript's prefix stable.
21export function renderMessage(m: SessionMessage): string {
22  if (m.role === 'user') {
23    const text = m.text.trim()
24
25    return text === '' ? '' : `## Person\n${text}\n`
26  }
27  const parts = m.text.trim() === '' ? [] : [m.text.trim()]
28  for (const use of m.toolUses) {
29    parts.push(`[${use.tool}] ${clip(JSON.stringify(use.input), INPUT_CHARS)}`)
30    if (use.text !== undefined) parts.push(`${use.isError ? '(error) ' : ''}${clip(use.text.trim(), RESULT_CHARS)}`)
31  }
32
33  return parts.length === 0 ? '' : `## Claude\n${parts.join('\n')}\n`
34}
35
36// 32-bit FNV-1a: enough to tell a transcript prefix that changed under a
37// compaction from one that did not.
38export function hash(text: string): number {
39  let h = 0x811c9dc5
40  for (let i = 0; i < text.length; i++) {
41    h ^= text.charCodeAt(i)
42    h = Math.imul(h, 0x01000193)
43  }
44
45  return h >>> 0
46}
47
48const sum = (xs: number[]) => xs.reduce((a, b) => a + b, 0)
49const span = (rendered: string[], a: number, b: number) => rendered.slice(a, b).join('\n')
50
51// Past this many stretches a transcript starts over as one.
52const MAX_STRETCHES = 24
53
54// The transcript cut into stretches for one model's request. Each earlier
55// request on the model ended its transcript at an anchor; the stretches run
56// between them, then on to the newest message. The request marks the last two
57// for the cache: the first ends where the previous request's cache entry
58// ended, so it is read, and the second is written for the next. `track` is
59// what the next request should remember.
60export type Transcript = { stretches: string[]; tokens: number; track: CacheTrack; isFresh: boolean }
61
62export function transcript(rendered: string[], maxTokens: number, was: CacheTrack | undefined, now: number): Transcript {
63  const count = rendered.length
64  const sizes = rendered.map(estimateTokens)
65  let from = 0
66  for (let i = count - 1, total = 0; i >= 0; i--) {
67    total += sizes[i] ?? 0
68    if (total > maxTokens) {
69      from = i + 1
70      break
71    }
72  }
73
74  // An earlier track holds while its text is unchanged (no compaction, no
75  // clear) and the conversation since has not outgrown the room by a quarter.
76  const last = was?.anchors.at(-1)
77  const holds =
78    was !== undefined &&
79    last !== undefined &&
80    last <= count &&
81    was.anchors.length < MAX_STRETCHES &&
82    sum(sizes.slice(was.from)) <= maxTokens * 1.25 &&
83    hash(span(rendered, was.from, last)) === was.hash
84  const start = holds ? was.from : from
85  const anchors = holds ? was.anchors : []
86  const edges = [start, ...anchors.filter(a => a < count), count]
87  const stretches: string[] = []
88  for (let i = 0; i < edges.length - 1; i++) stretches.push(span(rendered, edges[i] ?? 0, edges[i + 1] ?? count))
89  const next = anchors.at(-1) === count ? anchors : [...anchors, count]
90
91  return {
92    stretches,
93    tokens: sum(sizes.slice(start)),
94    track: { from: start, anchors: next, hash: hash(span(rendered, start, count)), at: now },
95    isFresh: !holds,
96  }
97}
98
99// How much of the transcript a model's cache likely holds now: what its last
100// request covered, for the five minutes such an entry lives.
101export function cachedTokens(rendered: string[], track: CacheTrack | undefined, now: number): number {
102  const last = track?.anchors.at(-1)
103  if (track === undefined || last === undefined || now - track.at > 5 * 60_000) return 0
104  if (last > rendered.length || hash(span(rendered, track.from, last)) !== track.hash) return 0
105
106  return sum(rendered.slice(track.from, last).map(estimateTokens))
107}
108
109const ROLE = `You are answering a side question the person asked about their Claude Code session. \
110This is a separate thread: your answer goes to a side pane, the session's main conversation never sees it, \
111and nothing you say changes what the session does next. Answer as an advisor: give your own judgment and \
112concrete recommendations, say plainly where you disagree with the direction the session is taking, and keep it \
113as short as the question allows.`
114
115const NO_TOOLS = `You cannot use tools here: answer from what the conversation already shows, and name what you \
116would need to check where it matters.`
117
118const TOOLS = `You may read files, search the code and the web, and run read-only commands to check things \
119before you answer. Change nothing: no edits, no writes, no commands that alter files, git state or anything else. \
120Your final message is your answer, written to the person.`
121
122export const SYSTEM = `${ROLE}\n\n${NO_TOOLS}`
123export const AGENT_SYSTEM = `${ROLE}\n\n${TOOLS}`
124
125// The earlier exchanges a question carries, oldest first.
126export function threadText(thread: Exchange[], turns: number): string {
127  const done = turns <= 0 ? [] : thread.filter(x => x.status === 'done' && x.answer !== undefined).slice(-turns)
128  if (done.length === 0) return ''
129  const rows = done.map(x => `Q: ${x.question}\nA: ${clip(x.answer ?? '', 4_000)}`)
130
131  return `Earlier questions in this side thread, oldest first:\n\n${rows.join('\n\n')}\n\n`
132}
133
134// The question as a fork reads it, after the session's own conversation.
135export function forkPrompt(thread: string, question: string, tools: boolean): string {
136  return `<aside>\n${ROLE}\n\n${tools ? TOOLS : NO_TOOLS}\n\n${thread}The question:\n${question}\n</aside>`
137}
138
139// The opening of a request of the aside's own, ahead of the transcript.
140export const TRANSCRIPT_HEAD = 'The session so far, as a transcript (tool results cut short):\n\n'
141
142export function askText(thread: string, question: string): string {
143  return `\n\n${thread}The question:\n${question}`
144}
145
146// What `/aside` was asked to do.
147export type Command =
148  | { kind: 'open' }
149  | { kind: 'ask'; question: string; model?: Choice; effort?: Effort; tools?: boolean }
150  | { kind: 'send'; index?: number }
151  | { kind: 'edit'; index?: number }
152  | { kind: 'clear' }
153  | { kind: 'model'; model: Choice }
154  | { kind: 'effort'; effort: Effort }
155  | { kind: 'tools'; tools: boolean }
156  | { kind: 'error'; text: string }
157
158export const CHOICES: readonly Choice[] = ['session', 'fable', 'opus', 'sonnet', 'haiku']
159export const EFFORTS: readonly Effort[] = ['default', 'low', 'medium', 'high', 'xhigh', 'max']
160
161const isChoice = (v: string): v is Choice => (CHOICES as readonly string[]).includes(v)
162const isEffort = (v: string): v is Effort => (EFFORTS as readonly string[]).includes(v)
163const noModel = (v: string): Command => ({ kind: 'error', text: `No model "${v}": ${CHOICES.join(', ')}.` })
164const noEffort = (v: string): Command => ({ kind: 'error', text: `No effort "${v}": ${EFFORTS.join(', ')}.` })
165
166export function parse(args: string): Command {
167  const text = args.trim()
168  if (text === '') return { kind: 'open' }
169  const sub = /^(send|edit)(?:\s+(\d+))?$/.exec(text)
170  if (sub !== null) return { kind: sub[1] === 'send' ? 'send' : 'edit', index: sub[2] === undefined ? undefined : Number(sub[2]) }
171  if (text === 'clear') return { kind: 'clear' }
172  const set = /^(model|effort|tools)\s+(\S+)$/.exec(text)
173  if (set !== null) {
174    const key = set[1]
175    const value = set[2] ?? ''
176    if (key === 'model') return isChoice(value) ? { kind: 'model', model: value } : noModel(value)
177    if (key === 'effort') return isEffort(value) ? { kind: 'effort', effort: value } : noEffort(value)
178    if (value === 'on' || value === 'off') return { kind: 'tools', tools: value === 'on' }
179
180    return { kind: 'error', text: 'Tools are on or off.' }
181  }
182
183  // Flags lead the question: -m <model>, -e <effort>, -t (tools), -T (no
184  // tools). The question after them keeps its own spacing and lines.
185  const ask: Extract<Command, { kind: 'ask' }> = { kind: 'ask', question: '' }
186  let rest = text
187  for (;;) {
188    const valued = /^(-m|--model|-e|--effort)\s+(\S+)(?:\s+|$)/.exec(rest)
189    const bare = /^(-t|--tools|-T|--no-tools)(?:\s+|$)/.exec(rest)
190    if (valued !== null) {
191      const flag = valued[1]
192      const value = valued[2] ?? ''
193      if (flag === '-m' || flag === '--model') {
194        if (!isChoice(value)) return noModel(value)
195        ask.model = value
196      } else {
197        if (!isEffort(value)) return noEffort(value)
198        ask.effort = value
199      }
200      rest = rest.slice(valued[0].length)
201    } else if (bare !== null) {
202      ask.tools = bare[1] === '-t' || bare[1] === '--tools'
203      rest = rest.slice(bare[0].length)
204    } else break
205  }
206  ask.question = rest.trim()
207
208  return ask.question === '' ? { kind: 'error', text: 'Ask a question after the flags.' } : ask
209}
210
211// Shell commands that only read, each with the arguments that would make it
212// write or run something else. Anything not listed is refused.
213const READERS: Record<string, RegExp | null> = {
214  cat: null, head: null, tail: null, wc: null, ls: null, pwd: null, file: null, stat: null, du: null, df: null,
215  echo: null, printf: null, grep: null, egrep: null, ag: null, which: null, type: null, basename: null,
216  dirname: null, realpath: null, readlink: null, cut: null, tr: null, column: null, diff: null, cmp: null,
217  jq: null, nl: null, od: null, hexdump: null, strings: null, sha256sum: null, md5sum: null, ps: null,
218  tree: /^-o$/,
219  sort: /^(-o|--output)/,
220  rg: /^--pre/,
221  fd: /^(-x|-X|--exec|--exec-batch)$/,
222  find: /^-(exec|execdir|ok|okdir|delete|fprint|fprint0|fprintf|fls)$/,
223  awk: /system|getline/,
224  uniq: null,
225  sed: null,
226  git: null,
227}
228const GIT_READERS = new Set([
229  'status', 'log', 'diff', 'show', 'blame', 'branch', 'tag', 'rev-parse', 'ls-files', 'ls-tree', 'shortlog',
230  'describe', 'grep', 'remote', 'reflog', 'cat-file', 'merge-base', 'name-rev', 'config',
231])
232
233const unquote = (s: string) => s.replace(/^(['"])(.*)\1$/, '$2')
234
235function readsOnly(name: string, args: string[]): boolean {
236  if (!(name in READERS)) return false
237  const deny = READERS[name]
238  if (deny !== null && deny !== undefined && args.some(a => deny.test(a))) return false
239  const positional = args.filter(a => !a.startsWith('-'))
240  // uniq writes its second operand; sed reads only as -n with line ranges.
241  if (name === 'uniq') return positional.length <= 1
242  if (name === 'sed') return args.includes('-n') && positional.length >= 1 && /^[0-9,$]+p$/.test(unquote(positional[0] ?? ''))
243  if (name === 'git') {
244    const verb = positional[0]
245    if (verb === undefined || !GIT_READERS.has(verb)) return false
246    if (args.some(a => /^--output|^-O$|^--open-files-in-pager|^--ext-diff/.test(a))) return false
247    if (verb === 'config') return args.some(a => a === '--list' || a === '-l' || a.startsWith('--get'))
248    if (verb === 'branch' || verb === 'tag' || verb === 'remote') {
249      return args.every(a => a === verb || /^-(v|vv|a|r|l|-list|-all|-verbose|-show-current|-contains|-merged|-no-merged)$/.test(a))
250    }
251  }
252
253  return true
254}
255
256export function isReadOnlyCommand(command: string): boolean {
257  // No redirection, substitution, background jobs, sequencing or subshells,
258  // quoted or not; only plain pipes between readers.
259  if (/[<>`;&(){}\n\\]/.test(command)) return false
260  for (const segment of command.split('|')) {
261    const [name, ...args] = segment.trim().split(/\s+/)
262    if (name === undefined || !readsOnly(name, args)) return false
263  }
264
265  return true
266}
267
268// The tools an aside's subagent may call; Bash only for read-only commands.
269const READ_TOOLS = new Set(['Read', 'Grep', 'Glob', 'LS', 'WebSearch', 'WebFetch', 'ToolSearch', 'LSP', 'NotebookRead', 'TodoWrite'])
270
271// Why an aside's subagent may not make this call, or undefined when it may.
272export function refusal(tool: string, input: unknown): string | undefined {
273  if (READ_TOOLS.has(tool)) return undefined
274  if (tool === 'Bash') {
275    const command = (input as { command?: unknown } | null)?.command
276
277    return typeof command === 'string' && isReadOnlyCommand(command)
278      ? undefined
279      : 'An aside only runs read-only commands: no writes, redirection, chaining or substitution.'
280  }
281
282  return `An aside only reads: ${tool} is not one of its tools.`
283}
284
types/index.d.ts 94 lines
1// Who answers: the session's own model, or another by its alias.
2export type Choice = 'session' | 'fable' | 'opus' | 'sonnet' | 'haiku'
3
4// How hard the answering model thinks; default leaves it as the model has it.
5export type Effort = 'default' | 'low' | 'medium' | 'high' | 'xhigh' | 'max'
6
7// How an aside is asked. fork: the session's own last request with the
8// question after it, read from its prompt cache, no tools. model: a request of
9// the aside's own with the conversation as a transcript. fork-agent: a forked
10// subagent with read-only tools. agent: a fresh subagent with read-only tools.
11export type Route = 'fork' | 'model' | 'fork-agent' | 'agent'
12
13// What the session's prompt cache lives for.
14export type Ttl = '5m' | '1h'
15
16// The four token counts an answer cost, as the API reports them.
17export type Spent = { input: number; output: number; read: number; write: number }
18
19// One question of the side thread and its answer.
20export type Exchange = {
21  id: string
22  at: number
23  question: string
24  // The answering model's name (`Opus 5.5`), and how it was asked.
25  model: string
26  route: Route
27  tools: boolean
28  effort: Effort
29  status: 'running' | 'done' | 'failed'
30  answer?: string
31  error?: string
32  spent?: Spent
33  usd?: number
34  // What the cost estimate said before it was asked, against the session's
35  // own model reading its cache, and whether it was asked anyway after a hold.
36  estimate?: number
37  stayEstimate?: number
38  forced?: boolean
39  agentId?: string
40  // When the answer was handed to the session, and how.
41  handedOff?: 'sent' | 'edited'
42  ms?: number
43}
44
45// The defaults each new question is asked with.
46export type Picks = { model: Choice; effort: Effort; tools: boolean }
47
48// A question held while the person decides on a pricier model.
49export type Pending = {
50  question: string
51  settings: Picks
52  usd: number
53  stayUsd: number
54  // The model that would answer, and the session's own, by name.
55  model: string
56  session: string
57  why: string
58}
59
60// What each choice is estimated to cost for the next question.
61export type Quote = {
62  at: number
63  session: string
64  contextTokens: number
65  transcriptTokens: number
66  isMainWarm: boolean
67  ttl: Ttl
68  costs: Partial<Record<Choice, number>>
69}
70
71// The conversation as one model's cache holds it: the transcript's first
72// message, the message counts each earlier request's blocks ended at, a hash
73// of the text up to the last of them, and when it was last read or written.
74export type CacheTrack = { from: number; anchors: number[]; hash: number; at: number }
75
76declare module 'claude-code' {
77  interface PluginState {
78    aside: {
79      // The side thread, oldest first.
80      thread: Exchange[]
81      // The settings the person picked; null until they pick.
82      defaults: Picks | null
83      pending: Pending | null
84      quote: Quote | null
85      // Each model's cached transcript, by model id.
86      caches: Record<string, CacheTrack>
87      // The subagents asides started, whose reports stay out of the session.
88      agents: string[]
89      // When the main thread last sent a request, and whether a turn runs now.
90      main: { at: number | null; isRunning: boolean }
91    }
92  }
93}
94