Ask Claude for advice in a side pane without touching the session's context: any model, optional read-only tools, an ongoing side thread, cache-aware cost…

Aside lets you ask Claude for advice in the middle of a session without touching the session's context. The question and the answer stay in a pane of their own. The session's model never sees either unless you hand the answer over.
It does what /btw does, plus four things:
It also keeps an eye on the prompt cache, and tells you when another model would cost more than the session's own.
/aside <question>: asks with the pane's current settings and opens the pane. It works while Claude is busy./aside: opens the pane. Type a question in its field and press Enter.-m <model>: session, fable, opus, sonnet or haiku.-e <effort>: low, medium, high, xhigh or max.-t turns tools on, and -T turns them off.For example: /aside -m fable -e high -t is this migration safe to run twice?
/aside model <m>, /aside effort <e> and /aside tools on|off change the pane's settings./aside clear empties the side thread.Only the newest answer has buttons, with keys: s sends it to the session as a prompt, e puts it in the prompt box, and c copies it. Each older answer shows its number. To hand over answer n, run /aside send n or /aside edit n. Without a number, they act on the newest answer.
| Settings | How it runs | Cache |
|---|---|---|
| Session model, default effort, no tools | It forks the session's own last request and puts the question after it. | It reads the session's cache, so it is usually the cheapest. |
| Session model, default effort, tools | It forks a subagent from the session. | It reads the session's cache, plus whatever its tool calls cost. |
| Another model or effort, no tools | It makes a request of its own, with the conversation as a transcript. | It writes its own cache, and later asides on that model read it. |
| Another model or effort, tools | It starts a fresh subagent and hands it the transcript. | It caches nothing across asides. |
A transcript holds up to transcriptTokens of the conversation, newest first, and cuts tool results short. A model that reads one sees less than the session's own model does.
The model picker shows an estimated cost next to each choice. Switching to a cheaper model is not always cheaper. On a long conversation with a warm cache, the session's model reads the whole context at the cache-read price. Another model has to write that context to its own cache first. When the model you pick would cost more than the session's own, Aside holds the question and shows why, for example:
Fable 5.1 costs about $0.70: a request of its own reads 49k tokens of this conversation cold. Opus 5.5, the session's model, reads its whole 320k-token context from cache for about $0.11.
You can then press y to ask the session's model, a to ask anyway, or x to drop the question. Set confirmSwitch to false to skip this.
The estimates use Claude API list prices. Each one starts as a rough guess and gets more accurate as Aside learns from real answers:
What Aside learns is kept across sessions. Once an aside is answered, its line in the pane shows what it actually cost and how much it read from the cache.
An aside's subagent can call Read, Grep, Glob, WebSearch, WebFetch, and Bash for read-only commands: git status, git log, git diff, rg, ls, cat, sed -n, and the like. It refuses any other tool and any command it does not recognize. It also refuses redirection, chaining, substitution, and flags that write, such as sort -o, find -delete or git log --output. An aside never stops to ask you for permission: anything that would ask is refused.
Change them in /config, or under pluginConfigs.aside in your settings.
| Setting | Default | What it does |
|---|---|---|
model | session | The model the pane starts on. |
tools | false | Whether the pane starts with read-only tools on. |
confirmSwitch | true | Hold a question when the model you picked would cost more than the session's own. |
cacheTtl | auto | How long the session's cache lives. auto reads it from the session transcript's last cache write at the end of a turn. If the transcript is over 4 MiB, it starts at 5 minutes and switches to 1 hour once a fork finds the cache still held after more than 5 minutes. |
transcriptTokens | 100000 | The most of the conversation a transcript holds. |
threadTurns | 10 | How many earlier answered asides each question carries. |
Claude Code v2.1.296 or later. Mods are an early access part of Claude Code, so a Claude Code release can break the plugin until it is updated.
/plugin install aside --marketplace astrosteveo/claude-pluginshooks/register.tsx 660 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, ModelTextBlock, ModelUsage, PluginOptions, Register } from 'claude-code'
3
4import type { Choice, Effort, Exchange, Route, Picks, Spent, Ttl } from '../types'
5import * as cost from './cost'
6import * as p from './prompt'
7
8const PANE = 'aside'
9const thread = atom({ plugin: 'aside', key: 'thread' } as const, [])
10const chosen = atom({ plugin: 'aside', key: 'defaults' } as const, null)
11const pending = atom({ plugin: 'aside', key: 'pending' } as const, null)
12const quote = atom({ plugin: 'aside', key: 'quote' } as const, null)
13const caches = atom({ plugin: 'aside', key: 'caches' } as const, {})
14const agents = atom({ plugin: 'aside', key: 'agents' } as const, [])
15const main = atom({ plugin: 'aside', key: 'main' } as const, { at: null, isRunning: false })
16
17const MAX_THREAD = 50
18const MAX_TOKENS = 8_000
19
20type Options = {
21 model: Choice
22 tools: boolean
23 confirmSwitch: boolean
24 cacheTtl: 'auto' | Ttl
25 transcriptTokens: number
26 threadTurns: number
27}
28
29function optionsOf(o: PluginOptions): Options {
30 const num = (v: unknown, d: number) => (typeof v === 'number' && Number.isFinite(v) && v >= 0 ? v : d)
31
32 return {
33 model: p.CHOICES.includes(o.model as Choice) ? (o.model as Choice) : 'session',
34 tools: o.tools === true,
35 confirmSwitch: o.confirmSwitch !== false,
36 cacheTtl: o.cacheTtl === '5m' || o.cacheTtl === '1h' ? o.cacheTtl : 'auto',
37 transcriptTokens: num(o.transcriptTokens, 100_000),
38 threadTurns: num(o.threadTurns, 10),
39 }
40}
41
42// The subagents asides started: their tool calls are held to reading, and
43// their reports are kept out of the session. Restored from state on each load.
44const ours = new Set<string>()
45// Spawns in flight, whose subagent may call a tool before its id is known.
46const spawning = new Set<Promise<unknown>>()
47let counter = 0
48
49const ADVISOR = 'advisor'
50const advisorType = (effort: Effort) => `aside:${effort === 'default' ? ADVISOR : `${ADVISOR}-${effort}`}`
51
52function spentOf(u: ModelUsage | undefined): Spent | undefined {
53 if (u === undefined) return undefined
54
55 return { input: u.input_tokens, output: u.output_tokens, read: u.cache_read_input_tokens, write: u.cache_creation_input_tokens }
56}
57
58async function settingsOf($: EngineInterface, o: Options): Promise<Picks> {
59 return (await read($, chosen)) ?? { model: o.model, effort: 'default', tools: o.tools }
60}
61
62// Figures by model id kept across sessions: `scale`, real tokens per
63// estimated one; `answers`, how long answers run; `runs`, what subagent runs
64// cost against their estimate.
65async function figures($: EngineInterface, key: 'scale' | 'answers' | 'runs'): Promise<Record<string, number>> {
66 const v = await $.store.get(key)
67
68 return typeof v === 'object' && v !== null ? (v as Record<string, number>) : {}
69}
70
71const scaleOf = ($: EngineInterface) => figures($, 'scale')
72
73async function ttlOf($: EngineInterface, o: Options): Promise<Ttl> {
74 if (o.cacheTtl !== 'auto') return o.cacheTtl
75
76 return (await $.store.get('ttl')) === '1h' ? '1h' : '5m'
77}
78
79async function sessionModel($: EngineInterface): Promise<cost.Model> {
80 const name = await $.session.model()
81 // A model of no family the table knows is priced as the newest Opus.
82 const opus = cost.resolve('opus') as cost.Model
83
84 return cost.resolve(name) ?? { ...opus, id: name, name }
85}
86
87// Who answers and how, for these settings: the session's own model at its
88// own effort forks the session's request and reads its cache; anything else
89// is a request or a subagent of the aside's own.
90function planOf(s: Picks, session: cost.Model): { plan: cost.Plan; modelArg: string } {
91 const model = s.model === 'session' ? session : (cost.resolve(s.model) ?? session)
92 const isFork = model.id === session.id && s.effort === 'default'
93 const route: Route = isFork ? (s.tools ? 'fork-agent' : 'fork') : s.tools ? 'agent' : 'model'
94
95 return { plan: { route, model }, modelArg: model.id === session.id ? session.id : s.model }
96}
97
98type Situation = { s: cost.Situation; rendered: string[]; now: number; mainAt: number | null; isMainRunning: boolean }
99
100async function situation($: EngineInterface, o: Options, askTokens: number): Promise<Situation> {
101 const [session, usage, messages, m, tracks, ttl, now, scale, answers, runs] = await Promise.all([
102 sessionModel($),
103 $.session.usage(),
104 $.session.messages(),
105 read($, main),
106 read($, caches),
107 ttlOf($, o),
108 $.clock.now(),
109 scaleOf($),
110 figures($, 'answers'),
111 figures($, 'runs'),
112 ])
113 const rendered = messages.map(p.renderMessage)
114 const cached: Record<string, number> = {}
115 for (const [id, track] of Object.entries(tracks)) cached[id] = p.cachedTokens(rendered, track, now)
116
117 return {
118 s: {
119 session,
120 contextTokens: usage.context.tokens ?? 0,
121 isMainWarm: m.isRunning || (m.at !== null && now - m.at < cost.TTL_MS[ttl]),
122 ttl,
123 transcriptTokens: p.transcript(rendered, o.transcriptTokens, undefined, now).tokens,
124 cached,
125 askTokens,
126 scale,
127 answers,
128 runs,
129 },
130 rendered,
131 now,
132 mainAt: m.at,
133 isMainRunning: m.isRunning,
134 }
135}
136
137const stayPlan = (s: Picks, session: cost.Model): cost.Plan => ({ route: s.tools ? 'fork-agent' : 'fork', model: session })
138
139// What the next question would cost on each choice, for the pane.
140async function refreshQuote($: EngineInterface, o: Options) {
141 const [s, list] = await Promise.all([settingsOf($, o), read($, thread)])
142 const at = await situation($, o, p.estimateTokens(p.threadText(list, o.threadTurns)) + 100)
143 const costs: Partial<Record<Choice, number>> = {}
144 for (const c of p.CHOICES) costs[c] = cost.estimate(planOf({ ...s, model: c }, at.s.session).plan, at.s)
145 await update($, quote, () => ({
146 at: at.now,
147 session: at.s.session.name,
148 contextTokens: at.s.contextTokens,
149 transcriptTokens: at.s.transcriptTokens,
150 isMainWarm: at.s.isMainWarm,
151 ttl: at.s.ttl,
152 costs,
153 }))
154}
155
156async function patch($: EngineInterface, id: string, change: Partial<Exchange>) {
157 await update($, thread, list => list.map(x => (x.id === id ? { ...x, ...change } : x)))
158}
159
160async function settle($: EngineInterface, o: Options, id: string, change: Partial<Exchange>) {
161 const now = await $.clock.now()
162 const was = (await read($, thread)).find(x => x.id === id)
163 await patch($, id, { ...change, ms: was === undefined ? undefined : now - was.at })
164 // A single request's answer length, for the next estimate; a subagent's
165 // output runs through its tool calls too, so it says nothing of one answer.
166 const model = was === undefined ? undefined : cost.resolve(was.model)
167 const route = change.route ?? was?.route
168 if (change.status === 'done' && change.spent !== undefined && model !== undefined && (route === 'fork' || route === 'model')) {
169 await $.store.set('answers', cost.lengthen(await figures($, 'answers'), model.id, change.spent.output))
170 }
171 if (!isShown) $.ui.toast(change.status === 'done' ? `Aside answered. /aside shows it.` : `Aside failed: ${change.error ?? 'no answer'}`)
172 void refreshQuote($, o).catch(() => {})
173}
174
175// Asks one question with these settings. A pricier pick than the session's
176// own cache waits for the person's word unless `force` gives it.
177async function ask($: EngineInterface, o: Options, question: string, settings: Picks, force = false) {
178 const list = await read($, thread)
179 const carried = p.threadText(list, o.threadTurns)
180 const at = await situation($, o, p.estimateTokens(carried + question) + 100)
181 const { plan, modelArg } = planOf(settings, at.s.session)
182 const mine = cost.estimate(plan, at.s)
183 const stay = cost.estimate(stayPlan(settings, at.s.session), at.s)
184 if (!force && o.confirmSwitch && plan.route !== 'fork' && plan.route !== 'fork-agent') {
185 const why = cost.warning(plan, mine, stay, at.s)
186 if (why !== undefined) {
187 await update($, pending, () => ({ question, settings, usd: mine, stayUsd: stay, model: plan.model.name, session: at.s.session.name, why }))
188 await openPane($)
189 if (!isShown) $.ui.toast(`Aside held: ${why} /aside to choose.`)
190
191 return
192 }
193 }
194 await update($, pending, () => null)
195 const ex: Exchange = {
196 id: `${at.now.toString(36)}-${++counter}`,
197 at: at.now,
198 question,
199 model: plan.model.name,
200 route: plan.route,
201 tools: settings.tools,
202 effort: settings.effort,
203 status: 'running',
204 estimate: mine,
205 stayEstimate: stay,
206 ...(force ? { forced: true } : {}),
207 }
208 await update($, thread, l => [...l, ex].slice(-MAX_THREAD))
209 void run($, o, ex, plan, modelArg, carried, at).catch(err => settle($, o, ex.id, { status: 'failed', error: String(err) }))
210}
211
212async function run($: EngineInterface, o: Options, ex: Exchange, plan: cost.Plan, modelArg: string, carried: string, at: Situation) {
213 if (plan.route === 'fork') {
214 const gap = at.isMainRunning ? 0 : at.mainAt === null ? undefined : at.now - at.mainAt
215 const r = await $.model.fork({ prompt: p.forkPrompt(carried, ex.question, false) })
216 if (!r.isAnswered && r.reason === 'nothing-to-fork') {
217 // No conversation yet: ask the session's model on its own.
218 await patch($, ex.id, { route: 'model' })
219
220 return run($, o, { ...ex, route: 'model' }, { route: 'model', model: plan.model }, modelArg, carried, at)
221 }
222 const spent = spentOf(r.usage)
223 // A fork reads the session's cache entry, which refreshes it.
224 if (spent !== undefined && spent.read > 0) await update($, main, m => ({ ...m, at: Math.max(m.at ?? 0, ex.at) }))
225 await learnTtl($, o, gap, spent, at.s.contextTokens)
226 const usd = spent === undefined ? undefined : cost.usd(plan.model, spent, at.s.ttl)
227 if (r.isAnswered) return settle($, o, ex.id, { status: 'done', answer: r.text, spent, usd })
228
229 return settle($, o, ex.id, { status: 'failed', error: failure(r), spent, usd })
230 }
231
232 if (plan.route === 'model') {
233 const tracks = await read($, caches)
234 const t = p.transcript(at.rendered, o.transcriptTokens, tracks[plan.model.id], at.now)
235 const marked = (i: number) => i >= t.stretches.length - 2
236 const prompt: ModelTextBlock[] = [
237 { text: p.TRANSCRIPT_HEAD },
238 ...t.stretches.map((text, i) => (marked(i) ? { text: text || '…', cache: true as const } : { text: text || '…' })),
239 { text: p.askText(carried, ex.question) },
240 ]
241 const r = await $.model.complete({
242 model: modelArg,
243 system: p.SYSTEM,
244 prompt,
245 maxTokens: MAX_TOKENS,
246 ...(ex.effort === 'default' ? {} : { effort: ex.effort }),
247 })
248 const spent = spentOf(r.usage)
249 const usd = spent === undefined ? undefined : cost.usd(plan.model, spent, '5m')
250 if (!r.isAnswered) return settle($, o, ex.id, { status: 'failed', error: failure(r), spent, usd })
251 await update($, caches, all => ({ ...all, [plan.model.id]: t.track }))
252 // What the model counted against what was estimated, for the next estimate.
253 if (spent !== undefined) {
254 const estimated = cost.SYSTEM_TOKENS + t.tokens + p.estimateTokens(p.TRANSCRIPT_HEAD + p.askText(carried, ex.question))
255 await $.store.set('scale', cost.rescale(await scaleOf($), plan.model.id, spent.input + spent.read + spent.write, estimated))
256 }
257
258 return settle($, o, ex.id, { status: 'done', answer: r.text, spent, usd })
259 }
260
261 // A subagent with read-only tools: a fork of the session, which reads its
262 // cache, or a fresh advisor on another model or effort, handed a transcript.
263 const brief = () => {
264 const t = p.transcript(at.rendered, o.transcriptTokens, undefined, at.now)
265
266 return `${p.TRANSCRIPT_HEAD}${t.stretches.join('\n')}${p.askText(carried, ex.question)}`
267 }
268 const description = `Aside: ${ex.question.slice(0, 40)}`
269 const spawn = async (args: Parameters<EngineInterface['agent']['spawn']>[0]) => {
270 const call = $.agent.spawn(args)
271 spawning.add(call)
272 try {
273 const r = await call
274 if (r.agentId !== undefined) {
275 ours.add(r.agentId)
276 await update($, agents, l => [...l, r.agentId as string].slice(-100))
277 }
278
279 return r
280 } finally {
281 spawning.delete(call)
282 }
283 }
284 let r = plan.route === 'fork-agent'
285 ? await spawn({ subagentType: 'fork', prompt: p.forkPrompt(carried, ex.question, true), description }).catch(() => undefined)
286 : await spawn({ subagentType: advisorType(ex.effort), model: modelArg, prompt: brief(), description })
287 if (plan.route === 'fork-agent' && (r === undefined || r.agentId === undefined)) {
288 // No fork to be had: an advisor on the session's model, handed a transcript.
289 await patch($, ex.id, { route: 'agent' })
290 r = await spawn({ subagentType: advisorType(ex.effort), model: modelArg, prompt: brief(), description })
291 }
292 if (r === undefined || r.agentId === undefined) {
293 return settle($, o, ex.id, { status: 'failed', error: r?.deny ?? 'The subagent did not start.' })
294 }
295 await patch($, ex.id, { agentId: r.agentId })
296}
297
298function failure(r: { reason: string; status?: number | null; error?: unknown }): string {
299 if (r.reason === 'api-error') return `API error${r.status ? ` ${r.status}` : ''}${r.error ? `: ${String(r.error)}` : ''}`
300 if (r.reason === 'empty-reply') return 'The model gave no answer.'
301 if (r.reason === 'aborted') return 'Cut short.'
302
303 return r.reason
304}
305
306// Whether this load has read the session cache's lifetime from the transcript.
307let isTtlRead = false
308
309// The session's cache lifetime, as the transcript records its last cache
310// write: no call on `$` reports it. A transcript past what one read takes
311// (4 MiB) is left to what the forks show.
312async function readTtl($: EngineInterface, path: string) {
313 const lines = (await $.fs.read(path)).split('\n')
314 for (let i = lines.length - 1; i >= 0; i--) {
315 const line = lines[i] ?? ''
316 if (!line.includes('"ephemeral_')) continue
317 const hour = Number(/"ephemeral_1h_input_tokens":\s*(\d+)/.exec(line)?.[1] ?? 0)
318 const five = Number(/"ephemeral_5m_input_tokens":\s*(\d+)/.exec(line)?.[1] ?? 0)
319 if (hour === 0 && five === 0) continue
320 await $.store.set('ttl', hour > 0 ? '1h' : '5m')
321 isTtlRead = true
322
323 return
324 }
325}
326
327// What a fork's cache read says about the session cache's lifetime, where
328// the transcript did not say: a read past five minutes means an hour; a miss
329// well inside an hour means five.
330async function learnTtl($: EngineInterface, o: Options, gap: number | undefined, spent: Spent | undefined, prefix: number) {
331 if (o.cacheTtl !== 'auto' || isTtlRead || gap === undefined || spent === undefined || prefix === 0) return
332 if (gap <= cost.TTL_MS['5m'] || gap >= cost.TTL_MS['1h']) return
333 if (spent.read > prefix / 2) await $.store.set('ttl', '1h')
334 else if (spent.read < prefix / 10) await $.store.set('ttl', '5m')
335}
336
337// Hands an answer to the session: sent as a prompt of its own, or put in the
338// prompt box to edit first.
339async function handOff($: EngineInterface, how: 'send' | 'edit', index: number | undefined) {
340 const list = await read($, thread)
341 const ex = index === undefined ? list.filter(x => x.status === 'done').at(-1) : list[index - 1]
342 if (ex === undefined || ex.status !== 'done' || ex.answer === undefined) {
343 $.ui.toast(index === undefined ? 'No answered aside to hand over yet.' : `Aside ${index} has no answer.`)
344
345 return
346 }
347 const text = `I asked ${ex.model} on the side: ${ex.question}\n\nIts answer:\n\n${ex.answer}`
348 if (how === 'edit') {
349 const r = await $.prompt.fill({ text, mode: 'insert' })
350 if (!r.isFilled) {
351 $.ui.toast('The prompt box could not take the answer.')
352
353 return
354 }
355 } else void $.prompt.submit({ text })
356 await patch($, ex.id, { handedOff: how === 'send' ? 'sent' : 'edited' })
357}
358
359let isShown = false
360
361async function openPane($: EngineInterface, focus = false) {
362 const r = await $.ui.open({ id: PANE, title: 'Aside', ...(focus ? { focus: true as const } : {}) })
363 isShown = r.isPlaced
364}
365
366async function setSettings($: EngineInterface, o: Options, change: Partial<Picks>) {
367 const s = await settingsOf($, o)
368 await update($, chosen, () => ({ ...s, ...change }))
369 void refreshQuote($, o).catch(() => {})
370}
371
372function choiceName(c: Choice, session: string | undefined): string {
373 if (c === 'session') return session === undefined ? 'Session' : `Session (${session})`
374
375 return cost.resolve(c)?.name ?? c
376}
377
378function meta(x: Exchange): string {
379 const parts = [x.model]
380 if (x.tools) parts.push('tools')
381 if (x.effort !== 'default') parts.push(x.effort)
382 if (x.route === 'fork' || x.route === 'fork-agent') parts.push('session cache')
383 if (x.ms !== undefined) parts.push(`${Math.round(x.ms / 1000)}s`)
384 if (x.usd !== undefined) parts.push(x.route === 'model' || x.route === 'fork' ? cost.money(x.usd) : `${cost.money(x.usd)} incl. tools`)
385 // Off the session's cache, what it was estimated at against staying, and
386 // whether a hold was overridden, so a pricier answer says how it got here.
387 if (x.route === 'model' || x.route === 'agent') {
388 const vs = x.stayEstimate === undefined ? '' : ` vs ≈${cost.money(x.stayEstimate)} staying`
389 if (x.estimate !== undefined) parts.push(`est. ≈${cost.money(x.estimate)}${vs}`)
390 if (x.forced) parts.push('asked anyway after a hold')
391 } else if (x.usd === undefined && x.estimate !== undefined) parts.push(`≈${cost.money(x.estimate)}`)
392 if (x.spent !== undefined) {
393 const { input, read, write, output } = x.spent
394 parts.push(`${cost.tokens(input + read + write)} in${read > 0 ? ` (${cost.tokens(read)} cached)` : ''} · ${cost.tokens(output)} out`)
395 }
396 if (x.handedOff !== undefined) parts.push(x.handedOff === 'sent' ? 'sent to the session' : 'put in the prompt')
397
398 return parts.join(' · ')
399}
400
401export const register: Register = (on, options) => {
402 const o = optionsOf(options)
403
404 on('session.start', async ($, e, next) => {
405 const result = await next(e)
406 for (const id of await read($, agents)) ours.add(id)
407 try {
408 await $.command.register({
409 name: 'aside',
410 description: 'Ask a side question without touching the session: any model, read-only tools, a pane of its own',
411 argumentHint: '[-m model] [-e effort] [-t] question | send | edit | clear',
412 immediate: true,
413 })
414 } catch {
415 // A clash with another command leaves the pane's own input working.
416 }
417 for (const effort of p.EFFORTS) {
418 await $.agent
419 .register({
420 name: effort === 'default' ? ADVISOR : `${ADVISOR}-${effort}`,
421 description: 'Answers a side question for /aside. Started by the aside mod only.',
422 prompt: p.AGENT_SYSTEM,
423 permissionMode: 'dontAsk',
424 disallowedTools: ['Edit', 'Write', 'NotebookEdit', 'Agent'],
425 maxTurns: 40,
426 ...(effort === 'default' ? {} : { effort }),
427 })
428 .catch(() => {})
429 }
430 void refreshQuote($, o).catch(() => {})
431
432 return result
433 })
434
435 // The advisor types are the mod's own: the session's model never sees them.
436 on('agent.offer', async ($, e, next) => (e.agent.startsWith('aside:') ? { isOffered: false } : next(e)))
437
438 on('turn.start', async ($, e, next) => {
439 await update($, main, m => ({ ...m, isRunning: true }))
440
441 return next(e)
442 })
443
444 on('turn.complete', async ($, e, next) => {
445 const result = await next(e)
446 if (e.agentId === undefined) {
447 const now = await $.clock.now()
448 await update($, main, () => ({ at: now, isRunning: false }))
449 void refreshQuote($, o).catch(() => {})
450
451 return result
452 }
453 if (!ours.has(e.agentId)) return result
454 const ex = (await read($, thread)).find(x => x.agentId === e.agentId)
455 if (ex === undefined || ex.status !== 'running') return result
456 const spent = spentOf(e.usage)
457 const model = (e.usage?.model !== undefined ? cost.resolve(e.usage.model) : undefined) ?? cost.resolve(ex.model)
458 const ttl = ex.route === 'fork-agent' ? await ttlOf($, o) : '5m'
459 const usd = spent === undefined || model === undefined ? undefined : cost.usd(model, spent, ttl)
460 if (e.reason === 'answer' && usd !== undefined && model !== undefined && ex.estimate !== undefined) {
461 await $.store.set('runs', cost.rerun(await figures($, 'runs'), ex.route, model.id, usd, ex.estimate))
462 }
463 if (e.reason === 'answer' && e.answer.trim() !== '') await settle($, o, ex.id, { status: 'done', answer: e.answer, spent, usd })
464 else await settle($, o, ex.id, { status: 'failed', error: e.reason === 'aborted' ? 'Stopped.' : 'The subagent gave no answer.', spent, usd })
465
466 return result
467 })
468
469 // A subagent's report would reach the session as a task notification:
470 // an aside's stays in its pane.
471 on('prompt.submit', async ($, e, next) => {
472 if (e.origin.kind === 'task-notification') {
473 for (const id of ours) if (e.text.includes(id)) return { drop: 'An aside finished: its answer is in the aside pane.' }
474 }
475
476 return next(e)
477 })
478
479 // An aside's subagent only reads. While one of its spawns is in flight, a
480 // subagent the mod cannot place yet waits for it before it is judged.
481 const judged = (agentId: string | undefined, tool: string, input: unknown): string | undefined =>
482 agentId !== undefined && (ours.has(agentId) || spawning.size > 0) ? p.refusal(tool, input) : undefined
483 on('tool.call', async ($, e, next) => {
484 if (e.agentId !== undefined && !ours.has(e.agentId) && spawning.size > 0) await Promise.allSettled([...spawning])
485 const why = e.agentId !== undefined && ours.has(e.agentId) ? p.refusal(String(e.tool), e) : undefined
486
487 return why === undefined ? next(e) : { deny: why }
488 }).catch(($, e, next) => {
489 if (next.called) return next(e)
490 const why = judged(e.agentId, String(e.tool), e)
491
492 return why === undefined ? next(e) : { deny: why }
493 })
494
495 // Nor does it ever stop to ask the person for permission.
496 on('tool.check', async ($, e, next) => {
497 const verdict = await next(e)
498 if (e.agentId === undefined || !ours.has(e.agentId) || verdict.decision !== 'ask') return verdict
499
500 return { decision: 'deny', reason: 'An aside never asks for permission.' }
501 })
502
503 on('command.run', { command: 'aside' }, async ($, e) => {
504 const c = p.parse(e.args)
505 if (c.kind === 'open') await openPane($, true)
506 else if (c.kind === 'ask') {
507 const s = await settingsOf($, o)
508 const settings: Picks = { model: c.model ?? s.model, effort: c.effort ?? s.effort, tools: c.tools ?? s.tools }
509 await openPane($)
510 void ask($, o, c.question, settings).catch(err => $.ui.toast(`Aside failed: ${String(err)}`))
511 } else if (c.kind === 'send' || c.kind === 'edit') await handOff($, c.kind, c.index)
512 else if (c.kind === 'clear') {
513 await update($, thread, () => [])
514 await update($, pending, () => null)
515 $.ui.toast('Aside thread cleared.')
516 } else if (c.kind === 'model') await setSettings($, o, { model: c.model })
517 else if (c.kind === 'effort') await setSettings($, o, { effort: c.effort })
518 else if (c.kind === 'tools') await setSettings($, o, { tools: c.tools })
519 else $.ui.toast(c.text)
520
521 // Nothing of the command reaches the transcript.
522 return {}
523 })
524
525 // At the end of a main turn, until this load has it, the cache lifetime.
526 on('classic.Stop', async ($, e, next) => {
527 const result = await next(e)
528 if (o.cacheTtl === 'auto' && !isTtlRead) {
529 await readTtl($, e.transcript_path).catch(() => {})
530 void refreshQuote($, o).catch(() => {})
531 }
532
533 return result
534 })
535
536 on('ui.close', async ($, e, next) => {
537 if (e.id === PANE) isShown = false
538
539 return next(e)
540 })
541
542 on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
543 const { Box, Text, Button, Markdown } = $.ui.resolve(e)
544 const [list, s, held, q] = await Promise.all([read($, thread), settingsOf($, o), read($, pending), read($, quote)])
545 const asked = (question: string, settings: Picks, force = false) =>
546 void ask($, o, question, settings, force).catch(err => $.ui.toast(`Aside failed: ${String(err)}`))
547 const priced = (c: Choice) => {
548 const usd = q?.costs[c]
549
550 return `${choiceName(c, q?.session)}${usd === undefined ? '' : ` ≈${cost.money(usd)}`}`
551 }
552
553 const controls =
554 e.surface === 'mobile' ? (
555 <Text dimColor>Ask with /aside from a terminal or the desktop.</Text>
556 ) : (
557 (() => {
558 const { Input, Select } = $.ui.resolve(e)
559
560 return (
561 <Box flexDirection="column">
562 <Input key="ask" label="Ask" placeholder="a side question" submitLabel="ask" autoFocus onSubmit={v => (v.trim() === '' ? undefined : asked(v.trim(), s))} />
563 <Box flexDirection="row" flexWrap="wrap" columnGap={2}>
564 <Select
565 key="model"
566 label="Model"
567 value={s.model}
568 options={p.CHOICES.map(c => ({ value: c, label: priced(c) }))}
569 onSelect={v => void setSettings($, o, { model: v as Choice })}
570 />
571 <Select
572 key="effort"
573 label="Effort"
574 value={s.effort}
575 options={p.EFFORTS.map(x => ({ value: x }))}
576 onSelect={v => void setSettings($, o, { effort: v as Effort })}
577 />
578 <Button key="tools" hotkey="t" onPress={() => void setSettings($, o, { tools: !s.tools })}>
579 {s.tools ? 'Tools: read-only' : 'Tools: off'}
580 </Button>
581 </Box>
582 </Box>
583 )
584 })()
585 )
586
587 const cacheLine =
588 q === null
589 ? undefined
590 : q.contextTokens === 0
591 ? 'No conversation yet.'
592 : `${cost.tokens(q.contextTokens)} tokens in context, ${q.isMainWarm ? 'cache warm' : 'cache likely cold'} (${q.ttl} cache).`
593
594 const latest = list.filter(x => x.status === 'done').at(-1)?.id
595 const rows = list
596 .map((x, i) => ({ x, n: i + 1 }))
597 .reverse()
598 .map(({ x, n }) => {
599 const isLatest = x.id === latest
600
601 return (
602 <Box key={x.id} flexDirection="column" marginTop={1}>
603 <Text bold>
604 {n}. {x.question}
605 </Text>
606 {x.status === 'running' && <Text dimColor>{x.model} is thinking…</Text>}
607 {x.status === 'done' && <Markdown text={x.answer ?? ''} />}
608 {x.status === 'failed' && <Text color="error">{x.error ?? 'No answer.'}</Text>}
609 <Text dimColor>{meta(x)}</Text>
610 {/* Only the newest answer has buttons, so a press can't land on an older one by mistake. */}
611 {x.status === 'done' && isLatest && (
612 <Box flexDirection="row" columnGap={1}>
613 <Button key="send" hotkey="s" onPress={() => void handOff($, 'send', n)}>
614 {x.handedOff === 'sent' ? 'Send again' : 'Send to session'}
615 </Button>
616 <Button key="edit" hotkey="e" onPress={() => void handOff($, 'edit', n)}>
617 Edit in prompt
618 </Button>
619 <Button key="copy" hotkey="c" onPress={pe => void $.ui.copy({ text: x.answer ?? '', surface: pe.surface })}>
620 Copy
621 </Button>
622 </Box>
623 )}
624 {x.status === 'done' && !isLatest && (
625 <Text dimColor>
626 /aside send {n} · /aside edit {n}
627 </Text>
628 )}
629 </Box>
630 )
631 })
632
633 return (
634 <Box flexDirection="column" width={e.props.bodyColumns}>
635 {controls}
636 {cacheLine !== undefined && <Text dimColor>{cacheLine}</Text>}
637 {held !== null && (
638 <Box flexDirection="column" marginTop={1} borderStyle="round" borderColor="warning" paddingX={1}>
639 <Text color="warning">{held.why}</Text>
640 <Text>Held: {held.question}</Text>
641 <Box flexDirection="row" columnGap={1}>
642 <Button key="stay" hotkey="y" variant="primary" onPress={() => asked(held.question, { ...held.settings, model: 'session', effort: 'default' })}>
643 Ask {held.session} instead
644 </Button>
645 <Button key="anyway" hotkey="a" onPress={() => asked(held.question, held.settings, true)}>
646 Ask {held.model} anyway
647 </Button>
648 <Button key="drop" hotkey="x" onPress={() => void update($, pending, () => null)}>
649 Drop it
650 </Button>
651 </Box>
652 </Box>
653 )}
654 {list.length === 0 && held === null && <Text dimColor>Nothing asked yet. Nothing here reaches the session unless you send it.</Text>}
655 {rows}
656 </Box>
657 )
658 })
659}
660hooks/cost.ts 212 lines1import type { Spent, Ttl } from '../types'
2
3// List prices on the Claude API in US dollars per million tokens: uncached
4// input, output and a cache read. A cache write is input times WRITE.
5export type Price = { input: number; output: number; read: number }
6
7export type Model = {
8 id: string
9 name: string
10 family: string
11 price: Price
12 // A price that takes over once a request's prompt passes `at` tokens.
13 over?: { at: number; price: Price }
14 // The shortest prefix the API caches.
15 minCache: number
16}
17
18const FABLE_51 = { input: 10, output: 50, read: 0.25 }
19const FABLE = { input: 10, output: 50, read: 1 }
20const OPUS_55 = { input: 4, output: 20, read: 0.2 }
21const OPUS = { input: 5, output: 25, read: 0.5 }
22const SONNET_5 = { input: 2, output: 10, read: 0.2 }
23const SONNET_4 = { input: 3, output: 15, read: 0.3 }
24
25// Newest first within each family, so an alias resolves to its first entry.
26const MODELS: Model[] = [
27 { id: 'claude-fable-5-1', name: 'Fable 5.1', family: 'fable', price: FABLE_51, minCache: 512 },
28 { id: 'claude-fable-5', name: 'Fable 5', family: 'fable', price: FABLE, minCache: 512 },
29 { id: 'claude-mythos-5-1', name: 'Mythos 5.1', family: 'mythos', price: FABLE_51, minCache: 512 },
30 { id: 'claude-mythos-5', name: 'Mythos 5', family: 'mythos', price: FABLE, minCache: 512 },
31 { id: 'claude-opus-5-5', name: 'Opus 5.5', family: 'opus', price: OPUS_55, minCache: 512 },
32 { id: 'claude-opus-5', name: 'Opus 5', family: 'opus', price: OPUS, minCache: 512 },
33 { id: 'claude-opus-4-8', name: 'Opus 4.8', family: 'opus', price: OPUS, minCache: 1024 },
34 { id: 'claude-opus-4-7', name: 'Opus 4.7', family: 'opus', price: OPUS, minCache: 2048 },
35 { id: 'claude-opus-4-6', name: 'Opus 4.6', family: 'opus', price: OPUS, minCache: 4096 },
36 { id: 'claude-opus-4-5', name: 'Opus 4.5', family: 'opus', price: OPUS, minCache: 4096 },
37 { id: 'claude-sonnet-5-5', name: 'Sonnet 5.5', family: 'sonnet', price: SONNET_5, minCache: 512 },
38 { id: 'claude-sonnet-5', name: 'Sonnet 5', family: 'sonnet', price: SONNET_5, minCache: 1024 },
39 { id: 'claude-sonnet-4-6', name: 'Sonnet 4.6', family: 'sonnet', price: SONNET_4, minCache: 1024 },
40 { id: 'claude-sonnet-4-5', name: 'Sonnet 4.5', family: 'sonnet', price: SONNET_4, minCache: 1024 },
41 {
42 id: 'claude-haiku-5-5',
43 name: 'Haiku 5.5',
44 family: 'haiku',
45 price: { input: 0.1, output: 0.5, read: 0.01 },
46 over: { at: 100_000, price: { input: 0.5, output: 2.5, read: 0.05 } },
47 minCache: 512,
48 },
49 { id: 'claude-haiku-4-5', name: 'Haiku 4.5', family: 'haiku', price: { input: 1, output: 5, read: 0.1 }, minCache: 4096 },
50]
51
52const WRITE: Record<Ttl, number> = { '5m': 1.25, '1h': 2 }
53export const TTL_MS: Record<Ttl, number> = { '5m': 5 * 60_000, '1h': 60 * 60_000 }
54
55// The model a name means: a full id (any provider prefix, date or context
56// suffix), a display name, or a bare alias, which means the family's newest.
57// Undefined for a name of no family it knows.
58export function resolve(name: string): Model | undefined {
59 const m = /(fable|mythos|opus|sonnet|haiku)(?:[-_ ]?(\d)(?:[-._](\d)(?!\d))?)?/i.exec(name)
60 if (m === null) return undefined
61 const family = (m[1] ?? '').toLowerCase()
62 const inFamily = MODELS.filter(x => x.family === family)
63 if (m[2] === undefined) return inFamily[0]
64 const id = `claude-${family}-${m[2]}${m[3] === undefined ? '' : `-${m[3]}`}`
65
66 return inFamily.find(x => x.id === id) ?? inFamily[0]
67}
68
69function priceAt(model: Model, promptTokens: number): Price {
70 return model.over !== undefined && promptTokens > model.over.at ? model.over.price : model.price
71}
72
73// What a request cost: its four token counts at the model's prices, cache
74// writes at the given lifetime's rate.
75export function usd(model: Model, t: Spent, ttl: Ttl = '5m'): number {
76 const p = priceAt(model, t.input + t.read + t.write)
77
78 return (t.input * p.input + t.read * p.read + t.write * p.input * WRITE[ttl] + t.output * p.output) / 1e6
79}
80
81// What an answer is assumed to run to, for the estimate; thinking included.
82export const ANSWER_TOKENS = 1_500
83// What a fresh subagent's own system prompt and tool list come to.
84export const AGENT_OVERHEAD = 10_000
85// The aside's own instructions on a request of its own.
86export const SYSTEM_TOKENS = 300
87
88// The figures an estimate reads.
89export type Situation = {
90 session: Model
91 // What the main thread's last request carried, and whether its cache is
92 // likely still held.
93 contextTokens: number
94 isMainWarm: boolean
95 ttl: Ttl
96 // The conversation as a transcript for another model, and how much of it
97 // each model's cache likely holds now, by model id.
98 transcriptTokens: number
99 cached: Record<string, number>
100 // The question and the side thread it carries.
101 askTokens: number
102 // Real tokens per estimated one, as earlier answers on each model counted
103 // them, by model id; `*` across models. A model with no count uses `*`.
104 scale: Record<string, number>
105 // How long each model's answers have run, in output tokens, by model id.
106 answers: Record<string, number>
107 // What subagent runs cost against their first-request estimate, by
108 // `<route>:<model id>`, and by route across models.
109 runs: Record<string, number>
110}
111
112export function runFactor(s: Pick<Situation, 'runs'>, route: string, id: string): number {
113 return s.runs[`${route}:${id}`] ?? s.runs[route] ?? 1
114}
115
116// Folds one subagent run's cost, against what was estimated with the factor
117// then in force, into the factor for its route and model.
118export function rerun(runs: Record<string, number>, route: string, id: string, actual: number, estimated: number): Record<string, number> {
119 if (actual <= 0 || estimated <= 0) return runs
120 const used = runFactor({ runs }, route, id)
121 const target = Math.min(5, Math.max(0.5, (used * actual) / estimated))
122 const fold = (was: number | undefined) => (was === undefined ? target : (was + target) / 2)
123
124 return { ...runs, [`${route}:${id}`]: fold(runs[`${route}:${id}`]), [route]: fold(runs[route]) }
125}
126
127// Folds one answer's length into a model's running figure.
128export function lengthen(answers: Record<string, number>, id: string, output: number): Record<string, number> {
129 if (output <= 0) return answers
130 const was = answers[id]
131
132 return { ...answers, [id]: was === undefined ? output : Math.round((was + output) / 2) }
133}
134
135export function scaleFor(s: Pick<Situation, 'scale'>, id: string): number {
136 return s.scale[id] ?? s.scale['*'] ?? 1
137}
138
139// Folds one answer's count of real against estimated tokens into the scale.
140export function rescale(scale: Record<string, number>, id: string, actual: number, estimated: number): Record<string, number> {
141 if (actual <= 0 || estimated <= 0) return scale
142 const ratio = Math.min(3, Math.max(0.5, actual / estimated))
143 const fold = (was: number | undefined) => (was === undefined ? ratio : (was + ratio) / 2)
144
145 return { ...scale, [id]: fold(scale[id]), '*': fold(scale['*']) }
146}
147
148export type Plan = { route: 'fork' | 'model' | 'fork-agent' | 'agent'; model: Model }
149
150// The estimated cost of one aside before it is asked. A subagent's is its
151// first request scaled by what earlier runs on its route and model cost
152// against theirs, which takes in its tool calls once it has run before.
153export function estimate(plan: Plan, s: Situation): number {
154 const first = request(plan, s)
155
156 return plan.route === 'agent' || plan.route === 'fork-agent' ? first * runFactor(s, plan.route, plan.model.id) : first
157}
158
159function request(plan: Plan, s: Situation): number {
160 const output = s.answers[plan.model.id] ?? ANSWER_TOKENS
161 if (plan.route === 'fork' || plan.route === 'fork-agent') {
162 const prefix = s.contextTokens
163 const t = s.isMainWarm
164 ? { input: s.askTokens, read: prefix, write: 0, output }
165 : { input: s.askTokens, read: 0, write: prefix, output }
166
167 return usd(plan.model, t, s.ttl)
168 }
169 const k = scaleFor(s, plan.model.id)
170 const overhead = plan.route === 'agent' ? AGENT_OVERHEAD : SYSTEM_TOKENS
171 const prefix = overhead + s.transcriptTokens * k
172 const input = s.askTokens * k
173 if (prefix < plan.model.minCache) return usd(plan.model, { input: prefix + input, read: 0, write: 0, output })
174 const read = plan.route === 'model' ? Math.min((s.cached[plan.model.id] ?? 0) * k, prefix) : 0
175
176 return usd(plan.model, { input, read, write: prefix - read, output })
177}
178
179// Under a tenth of a cent reads as such; under ten cents to the tenth of a
180// cent, since that is where asides differ; dollars to the cent past that.
181export function money(x: number): string {
182 if (x < 0.001) return '<$0.001'
183 if (x < 0.1) return `$${x.toFixed(3)}`
184
185 return `$${x.toFixed(2)}`
186}
187
188export function tokens(n: number): string {
189 if (n < 1_000) return String(Math.round(n))
190 if (n < 10_000) return `${(n / 1_000).toFixed(1).replace(/\.0$/, '')}k`
191 if (n < 1_000_000) return `${Math.round(n / 1_000)}k`
192
193 return `${(n / 1_000_000).toFixed(1).replace(/\.0$/, '')}M`
194}
195
196// Why an aside on `model` costs more than the session's own model would,
197// in a sentence for the person, or undefined when it does not.
198export function warning(plan: Plan, mine: number, stay: number, s: Situation): string | undefined {
199 if (mine <= stay * 1.05 || mine - stay < 0.001) return undefined
200 const how = plan.route === 'agent' ? 'a fresh subagent' : 'a request of its own'
201 const cold = plan.route === 'model' && (s.cached[plan.model.id] ?? 0) > 0 ? '' : ' cold'
202 const own = plan.model.id === s.session.id ? `${plan.model.name} at another effort` : plan.model.name
203 const read = tokens(s.transcriptTokens * scaleFor(s, plan.model.id))
204 const context = `its whole ${tokens(s.contextTokens)}-token context`
205 const stays = s.isMainWarm ? `reads ${context} from cache` : `would re-cache ${context}`
206
207 return (
208 `${own} costs about ${money(mine)}: ${how} reads ${read} tokens of this conversation${cold}. ` +
209 `${s.session.name}, the session's model, ${stays} for about ${money(stay)}.`
210 )
211}
212hooks/prompt.ts 284 lines1import type { SessionMessage } from 'claude-code'
2
3import type { CacheTrack, Choice, Effort, Exchange } from '../types'
4
5// A token is taken as 3.5 characters: code and prose sit either side of it.
6export function estimateTokens(text: string): number {
7 return Math.ceil(text.length / 3.5)
8}
9
10const RESULT_CHARS = 2_000
11const INPUT_CHARS = 300
12
13function clip(text: string, max: number): string {
14 return text.length <= max ? text : `${text.slice(0, max)}… (${text.length - max} more characters)`
15}
16
17// One message of the conversation as transcript text. A tool's result is
18// drawn under its call, so a user message that only carries results draws
19// nothing. The same message always draws the same text, which is what keeps
20// a cached transcript's prefix stable.
21export function renderMessage(m: SessionMessage): string {
22 if (m.role === 'user') {
23 const text = m.text.trim()
24
25 return text === '' ? '' : `## Person\n${text}\n`
26 }
27 const parts = m.text.trim() === '' ? [] : [m.text.trim()]
28 for (const use of m.toolUses) {
29 parts.push(`[${use.tool}] ${clip(JSON.stringify(use.input), INPUT_CHARS)}`)
30 if (use.text !== undefined) parts.push(`${use.isError ? '(error) ' : ''}${clip(use.text.trim(), RESULT_CHARS)}`)
31 }
32
33 return parts.length === 0 ? '' : `## Claude\n${parts.join('\n')}\n`
34}
35
36// 32-bit FNV-1a: enough to tell a transcript prefix that changed under a
37// compaction from one that did not.
38export function hash(text: string): number {
39 let h = 0x811c9dc5
40 for (let i = 0; i < text.length; i++) {
41 h ^= text.charCodeAt(i)
42 h = Math.imul(h, 0x01000193)
43 }
44
45 return h >>> 0
46}
47
48const sum = (xs: number[]) => xs.reduce((a, b) => a + b, 0)
49const span = (rendered: string[], a: number, b: number) => rendered.slice(a, b).join('\n')
50
51// Past this many stretches a transcript starts over as one.
52const MAX_STRETCHES = 24
53
54// The transcript cut into stretches for one model's request. Each earlier
55// request on the model ended its transcript at an anchor; the stretches run
56// between them, then on to the newest message. The request marks the last two
57// for the cache: the first ends where the previous request's cache entry
58// ended, so it is read, and the second is written for the next. `track` is
59// what the next request should remember.
60export type Transcript = { stretches: string[]; tokens: number; track: CacheTrack; isFresh: boolean }
61
62export function transcript(rendered: string[], maxTokens: number, was: CacheTrack | undefined, now: number): Transcript {
63 const count = rendered.length
64 const sizes = rendered.map(estimateTokens)
65 let from = 0
66 for (let i = count - 1, total = 0; i >= 0; i--) {
67 total += sizes[i] ?? 0
68 if (total > maxTokens) {
69 from = i + 1
70 break
71 }
72 }
73
74 // An earlier track holds while its text is unchanged (no compaction, no
75 // clear) and the conversation since has not outgrown the room by a quarter.
76 const last = was?.anchors.at(-1)
77 const holds =
78 was !== undefined &&
79 last !== undefined &&
80 last <= count &&
81 was.anchors.length < MAX_STRETCHES &&
82 sum(sizes.slice(was.from)) <= maxTokens * 1.25 &&
83 hash(span(rendered, was.from, last)) === was.hash
84 const start = holds ? was.from : from
85 const anchors = holds ? was.anchors : []
86 const edges = [start, ...anchors.filter(a => a < count), count]
87 const stretches: string[] = []
88 for (let i = 0; i < edges.length - 1; i++) stretches.push(span(rendered, edges[i] ?? 0, edges[i + 1] ?? count))
89 const next = anchors.at(-1) === count ? anchors : [...anchors, count]
90
91 return {
92 stretches,
93 tokens: sum(sizes.slice(start)),
94 track: { from: start, anchors: next, hash: hash(span(rendered, start, count)), at: now },
95 isFresh: !holds,
96 }
97}
98
99// How much of the transcript a model's cache likely holds now: what its last
100// request covered, for the five minutes such an entry lives.
101export function cachedTokens(rendered: string[], track: CacheTrack | undefined, now: number): number {
102 const last = track?.anchors.at(-1)
103 if (track === undefined || last === undefined || now - track.at > 5 * 60_000) return 0
104 if (last > rendered.length || hash(span(rendered, track.from, last)) !== track.hash) return 0
105
106 return sum(rendered.slice(track.from, last).map(estimateTokens))
107}
108
109const ROLE = `You are answering a side question the person asked about their Claude Code session. \
110This is a separate thread: your answer goes to a side pane, the session's main conversation never sees it, \
111and nothing you say changes what the session does next. Answer as an advisor: give your own judgment and \
112concrete recommendations, say plainly where you disagree with the direction the session is taking, and keep it \
113as short as the question allows.`
114
115const NO_TOOLS = `You cannot use tools here: answer from what the conversation already shows, and name what you \
116would need to check where it matters.`
117
118const TOOLS = `You may read files, search the code and the web, and run read-only commands to check things \
119before you answer. Change nothing: no edits, no writes, no commands that alter files, git state or anything else. \
120Your final message is your answer, written to the person.`
121
122export const SYSTEM = `${ROLE}\n\n${NO_TOOLS}`
123export const AGENT_SYSTEM = `${ROLE}\n\n${TOOLS}`
124
125// The earlier exchanges a question carries, oldest first.
126export function threadText(thread: Exchange[], turns: number): string {
127 const done = turns <= 0 ? [] : thread.filter(x => x.status === 'done' && x.answer !== undefined).slice(-turns)
128 if (done.length === 0) return ''
129 const rows = done.map(x => `Q: ${x.question}\nA: ${clip(x.answer ?? '', 4_000)}`)
130
131 return `Earlier questions in this side thread, oldest first:\n\n${rows.join('\n\n')}\n\n`
132}
133
134// The question as a fork reads it, after the session's own conversation.
135export function forkPrompt(thread: string, question: string, tools: boolean): string {
136 return `<aside>\n${ROLE}\n\n${tools ? TOOLS : NO_TOOLS}\n\n${thread}The question:\n${question}\n</aside>`
137}
138
139// The opening of a request of the aside's own, ahead of the transcript.
140export const TRANSCRIPT_HEAD = 'The session so far, as a transcript (tool results cut short):\n\n'
141
142export function askText(thread: string, question: string): string {
143 return `\n\n${thread}The question:\n${question}`
144}
145
146// What `/aside` was asked to do.
147export type Command =
148 | { kind: 'open' }
149 | { kind: 'ask'; question: string; model?: Choice; effort?: Effort; tools?: boolean }
150 | { kind: 'send'; index?: number }
151 | { kind: 'edit'; index?: number }
152 | { kind: 'clear' }
153 | { kind: 'model'; model: Choice }
154 | { kind: 'effort'; effort: Effort }
155 | { kind: 'tools'; tools: boolean }
156 | { kind: 'error'; text: string }
157
158export const CHOICES: readonly Choice[] = ['session', 'fable', 'opus', 'sonnet', 'haiku']
159export const EFFORTS: readonly Effort[] = ['default', 'low', 'medium', 'high', 'xhigh', 'max']
160
161const isChoice = (v: string): v is Choice => (CHOICES as readonly string[]).includes(v)
162const isEffort = (v: string): v is Effort => (EFFORTS as readonly string[]).includes(v)
163const noModel = (v: string): Command => ({ kind: 'error', text: `No model "${v}": ${CHOICES.join(', ')}.` })
164const noEffort = (v: string): Command => ({ kind: 'error', text: `No effort "${v}": ${EFFORTS.join(', ')}.` })
165
166export function parse(args: string): Command {
167 const text = args.trim()
168 if (text === '') return { kind: 'open' }
169 const sub = /^(send|edit)(?:\s+(\d+))?$/.exec(text)
170 if (sub !== null) return { kind: sub[1] === 'send' ? 'send' : 'edit', index: sub[2] === undefined ? undefined : Number(sub[2]) }
171 if (text === 'clear') return { kind: 'clear' }
172 const set = /^(model|effort|tools)\s+(\S+)$/.exec(text)
173 if (set !== null) {
174 const key = set[1]
175 const value = set[2] ?? ''
176 if (key === 'model') return isChoice(value) ? { kind: 'model', model: value } : noModel(value)
177 if (key === 'effort') return isEffort(value) ? { kind: 'effort', effort: value } : noEffort(value)
178 if (value === 'on' || value === 'off') return { kind: 'tools', tools: value === 'on' }
179
180 return { kind: 'error', text: 'Tools are on or off.' }
181 }
182
183 // Flags lead the question: -m <model>, -e <effort>, -t (tools), -T (no
184 // tools). The question after them keeps its own spacing and lines.
185 const ask: Extract<Command, { kind: 'ask' }> = { kind: 'ask', question: '' }
186 let rest = text
187 for (;;) {
188 const valued = /^(-m|--model|-e|--effort)\s+(\S+)(?:\s+|$)/.exec(rest)
189 const bare = /^(-t|--tools|-T|--no-tools)(?:\s+|$)/.exec(rest)
190 if (valued !== null) {
191 const flag = valued[1]
192 const value = valued[2] ?? ''
193 if (flag === '-m' || flag === '--model') {
194 if (!isChoice(value)) return noModel(value)
195 ask.model = value
196 } else {
197 if (!isEffort(value)) return noEffort(value)
198 ask.effort = value
199 }
200 rest = rest.slice(valued[0].length)
201 } else if (bare !== null) {
202 ask.tools = bare[1] === '-t' || bare[1] === '--tools'
203 rest = rest.slice(bare[0].length)
204 } else break
205 }
206 ask.question = rest.trim()
207
208 return ask.question === '' ? { kind: 'error', text: 'Ask a question after the flags.' } : ask
209}
210
211// Shell commands that only read, each with the arguments that would make it
212// write or run something else. Anything not listed is refused.
213const READERS: Record<string, RegExp | null> = {
214 cat: null, head: null, tail: null, wc: null, ls: null, pwd: null, file: null, stat: null, du: null, df: null,
215 echo: null, printf: null, grep: null, egrep: null, ag: null, which: null, type: null, basename: null,
216 dirname: null, realpath: null, readlink: null, cut: null, tr: null, column: null, diff: null, cmp: null,
217 jq: null, nl: null, od: null, hexdump: null, strings: null, sha256sum: null, md5sum: null, ps: null,
218 tree: /^-o$/,
219 sort: /^(-o|--output)/,
220 rg: /^--pre/,
221 fd: /^(-x|-X|--exec|--exec-batch)$/,
222 find: /^-(exec|execdir|ok|okdir|delete|fprint|fprint0|fprintf|fls)$/,
223 awk: /system|getline/,
224 uniq: null,
225 sed: null,
226 git: null,
227}
228const GIT_READERS = new Set([
229 'status', 'log', 'diff', 'show', 'blame', 'branch', 'tag', 'rev-parse', 'ls-files', 'ls-tree', 'shortlog',
230 'describe', 'grep', 'remote', 'reflog', 'cat-file', 'merge-base', 'name-rev', 'config',
231])
232
233const unquote = (s: string) => s.replace(/^(['"])(.*)\1$/, '$2')
234
235function readsOnly(name: string, args: string[]): boolean {
236 if (!(name in READERS)) return false
237 const deny = READERS[name]
238 if (deny !== null && deny !== undefined && args.some(a => deny.test(a))) return false
239 const positional = args.filter(a => !a.startsWith('-'))
240 // uniq writes its second operand; sed reads only as -n with line ranges.
241 if (name === 'uniq') return positional.length <= 1
242 if (name === 'sed') return args.includes('-n') && positional.length >= 1 && /^[0-9,$]+p$/.test(unquote(positional[0] ?? ''))
243 if (name === 'git') {
244 const verb = positional[0]
245 if (verb === undefined || !GIT_READERS.has(verb)) return false
246 if (args.some(a => /^--output|^-O$|^--open-files-in-pager|^--ext-diff/.test(a))) return false
247 if (verb === 'config') return args.some(a => a === '--list' || a === '-l' || a.startsWith('--get'))
248 if (verb === 'branch' || verb === 'tag' || verb === 'remote') {
249 return args.every(a => a === verb || /^-(v|vv|a|r|l|-list|-all|-verbose|-show-current|-contains|-merged|-no-merged)$/.test(a))
250 }
251 }
252
253 return true
254}
255
256export function isReadOnlyCommand(command: string): boolean {
257 // No redirection, substitution, background jobs, sequencing or subshells,
258 // quoted or not; only plain pipes between readers.
259 if (/[<>`;&(){}\n\\]/.test(command)) return false
260 for (const segment of command.split('|')) {
261 const [name, ...args] = segment.trim().split(/\s+/)
262 if (name === undefined || !readsOnly(name, args)) return false
263 }
264
265 return true
266}
267
268// The tools an aside's subagent may call; Bash only for read-only commands.
269const READ_TOOLS = new Set(['Read', 'Grep', 'Glob', 'LS', 'WebSearch', 'WebFetch', 'ToolSearch', 'LSP', 'NotebookRead', 'TodoWrite'])
270
271// Why an aside's subagent may not make this call, or undefined when it may.
272export function refusal(tool: string, input: unknown): string | undefined {
273 if (READ_TOOLS.has(tool)) return undefined
274 if (tool === 'Bash') {
275 const command = (input as { command?: unknown } | null)?.command
276
277 return typeof command === 'string' && isReadOnlyCommand(command)
278 ? undefined
279 : 'An aside only runs read-only commands: no writes, redirection, chaining or substitution.'
280 }
281
282 return `An aside only reads: ${tool} is not one of its tools.`
283}
284types/index.d.ts 94 lines1// Who answers: the session's own model, or another by its alias.
2export type Choice = 'session' | 'fable' | 'opus' | 'sonnet' | 'haiku'
3
4// How hard the answering model thinks; default leaves it as the model has it.
5export type Effort = 'default' | 'low' | 'medium' | 'high' | 'xhigh' | 'max'
6
7// How an aside is asked. fork: the session's own last request with the
8// question after it, read from its prompt cache, no tools. model: a request of
9// the aside's own with the conversation as a transcript. fork-agent: a forked
10// subagent with read-only tools. agent: a fresh subagent with read-only tools.
11export type Route = 'fork' | 'model' | 'fork-agent' | 'agent'
12
13// What the session's prompt cache lives for.
14export type Ttl = '5m' | '1h'
15
16// The four token counts an answer cost, as the API reports them.
17export type Spent = { input: number; output: number; read: number; write: number }
18
19// One question of the side thread and its answer.
20export type Exchange = {
21 id: string
22 at: number
23 question: string
24 // The answering model's name (`Opus 5.5`), and how it was asked.
25 model: string
26 route: Route
27 tools: boolean
28 effort: Effort
29 status: 'running' | 'done' | 'failed'
30 answer?: string
31 error?: string
32 spent?: Spent
33 usd?: number
34 // What the cost estimate said before it was asked, against the session's
35 // own model reading its cache, and whether it was asked anyway after a hold.
36 estimate?: number
37 stayEstimate?: number
38 forced?: boolean
39 agentId?: string
40 // When the answer was handed to the session, and how.
41 handedOff?: 'sent' | 'edited'
42 ms?: number
43}
44
45// The defaults each new question is asked with.
46export type Picks = { model: Choice; effort: Effort; tools: boolean }
47
48// A question held while the person decides on a pricier model.
49export type Pending = {
50 question: string
51 settings: Picks
52 usd: number
53 stayUsd: number
54 // The model that would answer, and the session's own, by name.
55 model: string
56 session: string
57 why: string
58}
59
60// What each choice is estimated to cost for the next question.
61export type Quote = {
62 at: number
63 session: string
64 contextTokens: number
65 transcriptTokens: number
66 isMainWarm: boolean
67 ttl: Ttl
68 costs: Partial<Record<Choice, number>>
69}
70
71// The conversation as one model's cache holds it: the transcript's first
72// message, the message counts each earlier request's blocks ended at, a hash
73// of the text up to the last of them, and when it was last read or written.
74export type CacheTrack = { from: number; anchors: number[]; hash: number; at: number }
75
76declare module 'claude-code' {
77 interface PluginState {
78 aside: {
79 // The side thread, oldest first.
80 thread: Exchange[]
81 // The settings the person picked; null until they pick.
82 defaults: Picks | null
83 pending: Pending | null
84 quote: Quote | null
85 // Each model's cached transcript, by model id.
86 caches: Record<string, CacheTrack>
87 // The subagents asides started, whose reports stay out of the session.
88 agents: string[]
89 // When the main thread last sent a request, and whether a turn runs now.
90 main: { at: number | null; isRunning: boolean }
91 }
92 }
93}
94