Prompt-cache meter above the Claude Code prompt: how many tokens each request read from, wrote to and sent past the cache, a live countdown to the cache's…

A prompt-cache meter above the Claude Code prompt. Every request Claude makes reports how much of its prompt the cache served, how much it wrote and how much went uncached; this mod keeps those numbers per request and per turn, counts down to the moment the cache lapses and tells you what to do about it: keep going, /compact or /clear.
cache ██████████ 98% read 80k · wrote 1k · new 300 ⏱ 3:41 5m · warm: keep going
cache ░░░░░░░░░░ 0% read 0 · wrote 52k · new 300 ⏱ 4:58 5m · cache missed: model changed (…)
cache ██████████ 98% read 150k · wrote 1k · new 300 ⏱ 0:00 5m · expired: the next message rewrites 151k tokens. /compact first, or /clear if the task is done
/cache opens a pane: the time left with a solid bar that shrinks as the cache runs out (green, yellow below 40% of the lifetime, red from the warning threshold), a stacked read / wrote / new bar for the last request, and a colour-coded table with one row per turn. Bars are filled cells and columns have fixed widths with no-break spaces, so the terminal and the Desktop (HTML) pane render the same.
The band and the /cache pane carry a Clear button (hotkey c once the band or pane has focus: ctrl+x tab for the band). Pressing it:
~/.claude/projects/<folder>/<session>.jsonl), keeps the user and assistant text and one line per tool call, and asks haiku for a brief with fixed headings: Goal, Current state, Decisions made, Files and commands touched, Open items, Next step;/clear, so the next request starts from an empty, cheap prefix;If the model gives no brief, a fallback with your last prompts and a claude --resume <id> line is sent instead. If /clear or the submit is refused, the brief is put in the prompt box for you to send. The button is hidden while a turn is running. The old session id is in the brief, so claude --resume gets the full history back.
From Anthropic's prompt caching documentation:
input_tokens (uncached remainder) + cache_read_input_tokens + cache_creation_input_tokens./compact./clear starts a new conversation in the same process, so the meter and the /cache table start over with it. A change in the prefix (model, effort or thinking settings, tool set, system prompt, CLAUDE.md) makes the next request write instead of read. The mod names the cause when it sees a miss: model changed, the cache had lapsed, or the prefix changed.The mod follows Claude Code's own rules (prompt caching: cache lifetime, Claude Code 2.1.242 or later). For the main conversation the TTL is the first match of:
| # | Source | Result |
|---|---|---|
| 1 | the mod's ttl option (5m / 1h) | what you set |
| 2 | FORCE_PROMPT_CACHING_5M=1 | 5 minutes |
| 3 | CLAUDE_CODE_PROMPT_CACHE_TTL | 5m or 1h |
| 4 | the promptCacheTtl setting (local, project or user settings file) | 5m or 1h |
| 5 | ENABLE_PROMPT_CACHING_1H=1 | 1 hour |
| 6 | the account | 1 hour on a Claude subscription within its plan usage; 5 minutes on usage credits, an API key or a cloud provider |
The account comes from the rate-limit windows the last response reported: a five_hour or seven_day window means a subscription, and one at 100% means requests now draw on usage credits. An API key or a cloud provider reports no such window, and before the first response nothing is known, so the mod starts from 5 minutes there. Managed settings are not readable from a mod.
On top of that the mod watches the traffic, which beats rows 2 to 6: a request that hits the cache more than 5 minutes after the previous one proves the 1-hour lifetime (a later miss does not undo it, since a changed prefix looks the same), and a miss 5 to 60 minutes after the previous request, with the same model and a prompt that did not shrink, says the entry lapsed, so 5 minutes (a later hit overrules it). That covers what the mod cannot see: managed settings, a gateway that rewrites the TTL, or a subscription that ran out of plan usage mid-session. The pane header names the source in use.
Why the mod infers instead of reading it: the API names the TTL of each write (cache_creation.ephemeral_5m_input_tokens / ephemeral_1h_input_tokens) and Claude Code's status line exposes it as prompt_cache.ttl, but the mod API passes on only the four token counts. To check by hand, claude -p "hello" --output-format json and read usage.cache_creation.
Other switches read from the environment at session start:
| Variable | Effect on the meter |
|---|---|
DISABLE_PROMPT_CACHING=1 (and _HAIKU, _SONNET, _OPUS) | the band says caching is off for that model |
turn.step: reads each main-loop request's usage (subagents have their own prefixes and are left out)$.clock.every(1000): redraws the countdown, and only while its text changes, so an idle expired session costs nothingui.render on AbovePrompt (the band) and on Pane (/cache)$.ui.toast: once per cache entry at the warning threshold (60 s by default) and again at 10, 3, 2 and 1 seconds left, for prompts of 20k tokens or more ttl: string "auto" | "5m" | "1h" (default auto)
warnSeconds: number countdown threshold for the yellow state and the toast (default 60)
compactAtTokens: number prompt size that makes an expired cache suggest /compact (default 100000)
band: boolean row above the prompt (default true)
status: boolean entry under the prompt, "cache 98% · 3:41" (default false)
toast: boolean toasts at the threshold, 10, 3, 2 and 1 s (default true)
The 100k compactAtTokens is a judgement, not a figure from the documentation: lower it if your model's cache writes are expensive for you.
In a Claude Code terminal session (2.1.287 or later, where mods load by default):
/plugin install prompt-cache-control --marketplace Nasrallah-Adel/claude-prompt-cache-control
Answer y to add the marketplace, then pick the user scope (Enter) so it loads in every project. It is active at once in that session and in each session started after.
To try it from a clone instead, for one session with hot reload:
git clone https://github.com/Nasrallah-Adel/claude-prompt-cache-control.git
claude --plugin-dir ./claude-prompt-cache-control
It is written to .claude/skills/prompt-cache-control/, which Claude Code auto-loads as prompt-cache-control@skills-dir once the workspace trust prompt is accepted. For one session with hot reload: claude --plugin-dir .claude/skills/prompt-cache-control. claude plugin validate .claude/skills/prompt-cache-control prints every event it hooks and every $ call it makes; claude plugin test .claude/skills/prompt-cache-control runs its tests.
Options are read from user settings (~/.claude/settings.json, never project settings), --settings <file> or managed settings, under the plugin's full id:
{ "pluginConfigs": { "prompt-cache-control@skills-dir": { "options": { } } } }
Requirements. Mods are on by default in Claude Code 2.1.287+. Typed against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods
MIT. The cache meter started from Daniel Ávila's prompt-cache-control mod (davila7/claude-code-templates); the Clear button, the handoff flow and the band layout are this repo's. LICENSE carries both copyright lines, as MIT requires.
hooks/prompt-cache-control.tsx 514 lines1/**
2 * prompt-cache-control — Claude Mod
3 *
4 * A prompt-cache meter for Claude Code. Every main-loop request reports how
5 * many prompt tokens the cache served (`cache_read_input_tokens`), wrote
6 * (`cache_creation_input_tokens`) and sent uncached (`input_tokens`); this mod
7 * keeps those per request and per turn, counts down to the moment the cache
8 * lapses, and says what to do about it: keep going, /compact or /clear.
9 *
10 * - `turn.step` reads each main-loop request's usage (subagents have their
11 * own prefixes and are left out)
12 * - `$.clock.every(1000)` redraws the countdown, and only while its text
13 * changes: an idle, expired session costs nothing
14 * - a row above the prompt (the AbovePrompt component), an optional status
15 * line entry, and `/cache`, a pane with one row per turn
16 *
17 * The lifetime is counted from the start of the request that last wrote or read
18 * the cache, as Anthropic documents it. Which lifetime Claude Code asked for
19 * follows its documented rules (see decideTtl in ./cache.ts): FORCE_PROMPT_CACHING_5M,
20 * CLAUDE_CODE_PROMPT_CACHE_TTL, the promptCacheTtl setting, ENABLE_PROMPT_CACHING_1H,
21 * then the account (1 hour on a Claude subscription, 5 minutes otherwise). The
22 * API names the TTL of a write but the mod API passes on only the token counts,
23 * so the mod also watches the gaps between requests (a hit after more than 5
24 * minutes proves 1 hour; see observeTtl). `ttl: "5m" | "1h"` pins it.
25 *
26 * Needs Claude Code >= 2.1.287.
27 *
28 * Options (pluginConfigs["prompt-cache-control@skills-dir"].options):
29 * ttl: "auto" | "5m" | "1h" cache lifetime (default auto)
30 * warnSeconds: number countdown threshold for the warning (default 60)
31 * compactAtTokens: number prompt size that makes an expired cache suggest /compact (default 100000)
32 * band: boolean row above the prompt (default true)
33 * status: boolean entry under the prompt (default false)
34 * toast: boolean toasts near expiry: at warnSeconds, then 10, 3, 2 and 1 s (default true)
35 */
36import type { EngineInterface, Register } from 'claude-code'
37import {
38 advise,
39 COUNTDOWN_MARKS,
40 bar,
41 byTurn,
42 fit,
43 fmtClock,
44 fmtTokens,
45 hitRatio,
46 isCachingDisabled,
47 accountOf,
48 decideTtl,
49 observeTtl,
50 lifeColor,
51 lifeRatio,
52 nextToastMark,
53 positive,
54 promptTokens,
55 remainingMs,
56 rowRatio,
57 segments,
58} from './cache.ts'
59import type { Account, Advice, CacheEnv, Sample, Ttl } from './cache.ts'
60import { digestTranscript, fallbackBrief, HANDOFF_SYSTEM, handoffPrompt, parseRows, wrapForNewSession } from './handoff.ts'
61
62const HANDOFF_MODEL = 'haiku'
63const HANDOFF_MAX_TOKENS = 1200
64const HANDOFF_TIMEOUT_MS = 90_000
65
66const PANE = 'cache'
67const COMMAND = 'cache'
68const KEEP = 200
69// below this a lapsed cache costs too little to interrupt anyone about
70const TOAST_MIN_TOKENS = 20_000
71
72let samples: Sample[] = []
73let ttl: Ttl = '5m'
74let baseTtl: Ttl = '5m'
75let pinned = false
76let observed: Ttl | undefined
77let setting: unknown
78let account: Account = 'other'
79let ttlSource = 'default'
80let envSource = 'default'
81let env: CacheEnv = {}
82let timer: { cancel: () => void } | undefined
83let lastKey = ''
84let toastedFor = 0
85let toastLevel = Infinity
86let isPaneOpen = false
87
88type Policy = { warnMs: number; compactAtTokens: number }
89
90function current(policy: Policy, now: number) {
91 const last = samples[samples.length - 1]
92 const prev = samples[samples.length - 2]
93 const disabled = last ? isCachingDisabled(last.model, env) : isCachingDisabled('', env)
94 const advice: Advice = advise(last, prev, { ttl, ...policy }, now, disabled)
95 const left = last ? remainingMs(last, ttl, now) : 0
96 return { last, advice, left }
97}
98
99const COLOR: Record<Advice['kind'], string | undefined> = {
100 warm: 'green',
101 soon: 'yellow',
102 expired: 'red',
103 miss: 'red',
104 off: undefined,
105 cold: undefined,
106 uncached: undefined,
107}
108
109function shortLine(policy: Policy, now: number): string {
110 const { last, advice, left } = current(policy, now)
111 if (!last || advice.kind === 'off') return `cache: ${advice.text}`
112 const clock = left > 0 ? ` · ${fmtClock(left)}` : ''
113 return `cache ${Math.round(hitRatio(last) * 100)}%${clock}`
114}
115
116// the promptCacheTtl setting, from the settings files that can carry it (local over project over user)
117async function readSetting($: EngineInterface): Promise<unknown> {
118 const home = await $.env.get('HOME').catch(() => undefined)
119 const cwd = await $.session.cwd().catch(() => undefined)
120 const files = [cwd && `${cwd}/.claude/settings.local.json`, cwd && `${cwd}/.claude/settings.json`, home && `${home}/.claude/settings.json`]
121 for (const file of files) {
122 if (!file) continue
123 try {
124 const value = JSON.parse(await $.fs.read(file)).promptCacheTtl
125 if (value === '5m' || value === '1h') return value
126 } catch {
127 // missing or unreadable: the next file
128 }
129 }
130 return undefined
131}
132
133// The Clear button's progress text while it runs; undefined when idle.
134let clearing: string | undefined
135
136/** Read this session's transcript file, if the engine wrote one where it usually does. */
137async function readTranscript($: EngineInterface, sessionId: string): Promise<string | undefined> {
138 const home = await $.env.get('HOME').catch(() => undefined)
139 if (!home) return undefined
140 const projects = `${home}/.claude/projects`
141 const dirs = await $.fs.list(projects).catch(() => [])
142 for (const dir of dirs) {
143 if (dir.kind !== 'dir') continue
144 const path = `${projects}/${dir.name}/${sessionId}.jsonl`
145 if (await $.fs.exists(path).catch(() => false)) {
146 return $.fs.read(path).catch(() => undefined) as Promise<string | undefined>
147 }
148 }
149 return undefined
150}
151
152/** Clear: write a handoff brief from the transcript, /clear, then hand the brief to the fresh session. */
153async function clearAndHandoff($: EngineInterface): Promise<void> {
154 if (clearing) return
155 const step = (text: string | undefined) => {
156 clearing = text
157 $.ui.invalidate('ui.render')
158 }
159 step('writing handoff…')
160 $.ui.toast('cache: writing a handoff brief, then clearing')
161 const sessionId = await $.session.id().catch(() => 'unknown')
162 let brief: string
163 try {
164 const digest = digestTranscript(parseRows((await readTranscript($, sessionId)) ?? ''))
165 const r = await $.model.complete({
166 model: HANDOFF_MODEL,
167 system: HANDOFF_SYSTEM,
168 prompt: handoffPrompt(digest),
169 maxTokens: HANDOFF_MAX_TOKENS,
170 effort: 'low',
171 timeoutMs: HANDOFF_TIMEOUT_MS,
172 })
173 brief = r.isAnswered && r.text.trim() ? r.text : fallbackBrief(digest, sessionId)
174 if (!r.isAnswered) $.ui.log(`prompt-cache-control: handoff brief not generated (${r.reason}); sent the fallback`, { to: 'debug' })
175 } catch (err) {
176 brief = fallbackBrief(digestTranscript([]), sessionId)
177 $.ui.log(`prompt-cache-control: handoff failed (${String(err)}); sent the fallback`, { to: 'debug' })
178 }
179 const text = wrapForNewSession(brief, sessionId)
180 try {
181 step('clearing…')
182 await $.command.run({ command: 'clear' })
183 step('sending handoff…')
184 await $.prompt.submit({ text, asUser: true })
185 $.ui.toast('cache: cleared; the handoff brief is in the new conversation')
186 } catch (err) {
187 $.ui.log(`prompt-cache-control: clear or submit failed (${String(err)})`, { to: 'debug' })
188 const filled = await $.prompt.fill({ text }).catch(() => ({ isFilled: false }))
189 $.ui.toast(filled.isFilled ? 'cache: handoff is in the prompt box, press Enter to send it' : 'cache: could not clear; run /clear yourself')
190 } finally {
191 step(undefined)
192 }
193}
194
195export const register: Register = (on, options) => {
196 const policy: Policy = {
197 warnMs: positive(options.warnSeconds, 60) * 1000,
198 compactAtTokens: positive(options.compactAtTokens, 100_000),
199 }
200 const showBand = options.band !== false
201 const showStatus = options.status === true
202 const wantToast = options.toast !== false
203
204 on('session.start', async ($, e, next) => {
205 const r = await next(e)
206 samples = []
207 lastKey = ''
208 toastedFor = 0
209 const none = () => undefined
210 env = {
211 enable1h: await $.env.get('ENABLE_PROMPT_CACHING_1H').catch(none),
212 force5m: await $.env.get('FORCE_PROMPT_CACHING_5M').catch(none),
213 ttlVar: await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL').catch(none),
214 disableAll: await $.env.get('DISABLE_PROMPT_CACHING').catch(none),
215 disableHaiku: await $.env.get('DISABLE_PROMPT_CACHING_HAIKU').catch(none),
216 disableSonnet: await $.env.get('DISABLE_PROMPT_CACHING_SONNET').catch(none),
217 disableOpus: await $.env.get('DISABLE_PROMPT_CACHING_OPUS').catch(none),
218 }
219 pinned = options.ttl === '5m' || options.ttl === '1h'
220 observed = undefined
221 setting = await readSetting($)
222 account = accountOf((await $.session.usage().catch(() => undefined))?.rateLimits ?? [])
223 const choice = decideTtl(options.ttl, env, setting, account)
224 baseTtl = choice.ttl
225 ttl = baseTtl
226 envSource = choice.source
227 ttlSource = envSource
228
229 await $.command
230 .register({
231 name: COMMAND,
232 description: 'Prompt-cache usage per turn and the time left before it lapses (stop closes)',
233 argumentHint: '[stop]',
234 immediate: true,
235 })
236 .catch(err => $.ui.log(`prompt-cache-control: /${COMMAND} not registered: ${err}`))
237 $.ui.log(`prompt-cache-control loaded: ${ttl} cache (${ttlSource}), /${COMMAND} opens the table`, { to: 'debug' })
238
239 timer?.cancel()
240 timer = $.clock.every(1000, () => {
241 const now = Date.now()
242 const { last, advice, left } = current(policy, now)
243 const key = `${advice.kind}|${advice.text}|${left > 0 ? fmtClock(left) : ''}`
244 if (key !== lastKey) {
245 lastKey = key
246 if (showStatus) $.ui.status(shortLine(policy, now))
247 $.ui.invalidate('ui.render')
248 }
249 if (wantToast && last && left > 0 && promptTokens(last) >= TOAST_MIN_TOKENS) {
250 if (toastedFor !== last.startedAt) {
251 toastedFor = last.startedAt
252 toastLevel = Infinity
253 }
254 // the first toast comes at warnSeconds, then 10, 3, 2 and 1 seconds; a late tick skips to the newest one
255 const secs = Math.ceil(left / 1000)
256 const mark = nextToastMark(secs, policy.warnMs / 1000, toastLevel)
257 if (mark !== undefined) {
258 toastLevel = mark
259 const tail = secs <= COUNTDOWN_MARKS[0] ? 'send a message now' : `send a message to keep ${fmtTokens(promptTokens(last))} tokens warm`
260 $.ui.toast(`cache expires in ${secs >= 60 ? fmtClock(left) : `${secs}s`}: ${tail}`)
261 }
262 }
263 })
264 return r
265 })
266
267 on('session.end', async ($, e, next) => {
268 // /clear starts a new conversation in the same process: its cache is a new one
269 if (e.reason === 'clear') {
270 samples = []
271 lastKey = ''
272 toastedFor = 0
273 observed = undefined
274 ttl = baseTtl
275 ttlSource = envSource
276 $.ui.invalidate('ui.render')
277 return next(e)
278 }
279 timer?.cancel()
280 timer = undefined
281 return next(e)
282 })
283
284 // each main-loop request: what the cache did with it
285 on('turn.step', async function* ($, e, next) {
286 if (e.agentId) return yield* next(e)
287 const startedAt = Date.now()
288 const r = yield* next(e)
289 if (r.usage) {
290 samples.push({
291 turnId: e.turnId,
292 index: e.index,
293 model: r.usage.model || e.model,
294 startedAt,
295 read: r.usage.cache_read_input_tokens,
296 write: r.usage.cache_creation_input_tokens,
297 fresh: r.usage.input_tokens,
298 output: r.usage.output_tokens,
299 })
300 if (samples.length > KEEP) samples = samples.slice(-KEEP)
301 if (!pinned) {
302 // the account can change under a session: a subscription running out of plan usage moves to usage credits
303 account = accountOf((await $.session.usage().catch(() => undefined))?.rateLimits ?? [])
304 const choice = decideTtl(options.ttl, env, setting, account)
305 baseTtl = choice.ttl
306 envSource = choice.source
307 if (observed === undefined) {
308 ttl = baseTtl
309 ttlSource = envSource
310 }
311 const seen = observeTtl(samples[samples.length - 2], samples[samples.length - 1], observed)
312 if (seen !== observed) {
313 observed = seen
314 ttl = seen ?? baseTtl
315 ttlSource = `observed from request timing; ${envSource} said ${baseTtl}`
316 $.ui.log(`prompt-cache-control: cache lifetime is ${ttl} (${ttlSource})`, { to: 'debug' })
317 }
318 }
319 lastKey = ''
320 if (showStatus) $.ui.status(shortLine(policy, Date.now()))
321 $.ui.invalidate('ui.render')
322 }
323 return r
324 })
325
326 on('command.run', { command: COMMAND }, async ($, e) => {
327 if (e.args.trim().toLowerCase() === 'stop') {
328 await $.ui.close({ id: PANE }).catch(() => undefined)
329 isPaneOpen = false
330 return { text: 'cache table closed' }
331 }
332 isPaneOpen = true
333 await $.ui.open({ id: PANE, title: 'cache', focus: true })
334 $.ui.invalidate('ui.render')
335 const { advice } = current(policy, Date.now())
336 return { text: `${ttl} cache (${ttlSource}) · ${advice.text} · /${COMMAND} stop closes` }
337 })
338
339 on('ui.close', async ($, e, next) => {
340 if (e.id !== PANE) return next(e)
341 isPaneOpen = false
342 return next(e)
343 })
344
345 on('ui.press', async ($, e, next) => {
346 if (e.plugin !== $.plugin.name || e.requestId !== PANE) return next(e)
347 if (e.element === 'close') await $.ui.close({ id: PANE }).catch(() => undefined)
348 return next(e)
349 })
350
351 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
352 if (!showBand || e.props.hasSurvey || isPaneOpen) return next(e)
353 const { last, advice, left } = current(policy, Date.now())
354 if (!last && advice.kind !== 'off') return next(e)
355 const { Box, Text, Button } = $.ui.resolve(e)
356 const columns = e.viewport?.columns ?? 100
357 const color = COLOR[advice.kind]
358 // The band is one instance: whatever the plugins beneath draw (another mod's band) stays, under this line.
359 const below = await next(e)
360 const stack = (own: ReturnType<typeof Text>) => (
361 <Box flexDirection="column">
362 {own}
363 {below}
364 </Box>
365 )
366 // Clear: brief + /clear + handoff. Always drawn; mid-turn it only says to wait (the brief would go stale).
367 const busy = e.props.isWorking
368 const clearControl = clearing ? (
369 <Text color="yellow">{clearing}</Text>
370 ) : (
371 <Button
372 key="clear"
373 label="Clear"
374 hotkey="c"
375 onPress={() => (busy ? $.ui.toast('cache: wait for the turn to finish, then press Clear') : void clearAndHandoff($))}
376 />
377 )
378
379 if (!last) return stack(<Text dimColor>{fit(`cache: ${advice.text}`, columns)}</Text>)
380
381 const ratio = hitRatio(last)
382 const wide = columns >= 90
383 return stack(
384 <Box flexDirection="row" columnGap={1}>
385 <Text bold color={color}>{advice.kind === 'warm' ? '●' : advice.kind === 'soon' ? '▲' : advice.kind === 'off' || advice.kind === 'cold' || advice.kind === 'uncached' ? '○' : '✖'}</Text>
386 <Text bold color="cyan">cache</Text>
387 <Text color={color}>{bar(ratio, wide ? 10 : 6)}</Text>
388 <Text bold>{`${Math.round(ratio * 100)}%`}</Text>
389 {wide && <Text color="green">{`read ${fmtTokens(last.read)}`}</Text>}
390 {wide && <Text color="yellow">{`wrote ${fmtTokens(last.write)}`}</Text>}
391 {wide && <Text color="cyan">{`new ${fmtTokens(last.fresh)}`}</Text>}
392 {!wide && <Text dimColor>{`${fmtTokens(promptTokens(last))} tok`}</Text>}
393 {advice.kind !== 'uncached' && advice.kind !== 'off' && (
394 <Text bold color={left > 0 ? lifeColor(left, ttl, policy.warnMs) : 'red'}>{left > 0 ? `⏱ ${fmtClock(left)}` : '⏱ 0:00'}</Text>
395 )}
396 {clearControl}
397 <Text dimColor wrap="truncate-end">{`${ttl} · ${advice.text}`}</Text>
398 </Box>,
399 )
400 })
401
402 on('ui.render', { component: 'Pane' }, async ($, e, next) => {
403 if (e.requestId !== PANE) return next(e)
404 const { Box, Text, Button } = $.ui.resolve(e)
405 const width = Math.max(30, e.props.bodyColumns - 1)
406 // HTML collapses runs of spaces and trims a text's ends; a no-break space keeps them
407 const sp = (t: string) => (e.surface === 'terminal' ? t : t.replace(/ /g, ' '))
408 const now = Date.now()
409 const { last, advice, left } = current(policy, now)
410 const all = byTurn(samples)
411 const counting = !!last && advice.kind !== 'uncached' && advice.kind !== 'off'
412 // the countdown goes green, then yellow, then red as the cache runs out
413 const clockColor = counting ? lifeColor(left, ttl, policy.warnMs) : undefined
414 const stateColor = advice.kind === 'expired' || advice.kind === 'miss' ? 'red' : (clockColor ?? COLOR[advice.kind])
415 const hitColor = (pct: number) => (pct >= 80 ? 'green' : pct >= 40 ? 'yellow' : 'red')
416 // solid bars are filled Boxes, not block characters, so HTML draws no seams between cells
417 const solid = (key: string, parts: [number, string | undefined][]) => (
418 <Box key={key} flexDirection="row" height={1} flexShrink={0}>
419 {parts.map(([w, c], i) => (w > 0 ? <Box key={`${key}:${i}`} width={w} height={1} flexShrink={0} backgroundColor={c} /> : null))}
420 </Box>
421 )
422 const cell = (key: string, w: number, text: string, c?: string, bold = false) => (
423 <Box key={key} width={w} flexShrink={0} justifyContent="flex-end">
424 <Text color={c} bold={bold} dimColor={!c}>{sp(text)}</Text>
425 </Box>
426 )
427
428 const barW = Math.min(width, 40)
429 const life = lifeRatio(left, ttl)
430 const lifeFilled = Math.round(life * barW)
431 const [sr, sw, sn] = last ? segments(last.read, last.write, last.fresh, barW) : [0, 0, 0]
432 const rows = all.slice(-Math.max(3, (e.viewport?.rows ?? 24) - 16))
433 const icon = advice.kind === 'warm' ? '●' : advice.kind === 'soon' ? '▲' : advice.kind === 'expired' || advice.kind === 'miss' ? '✖' : '○'
434
435 return (
436 <Box flexDirection="column">
437 <Box key="title" flexDirection="row" columnGap={1}>
438 <Text bold color="cyan">{sp('⚡ PROMPT CACHE')}</Text>
439 <Text dimColor>{sp(`· ${ttl} lifetime (${ttlSource})`)}</Text>
440 </Box>
441
442 <Box key="clock" flexDirection="column" marginTop={1}>
443 <Text bold color={clockColor}>{sp(counting ? `⏱ ${left > 0 ? fmtClock(left) : '0:00'}` : '⏱ --:--')}</Text>
444 {counting ? (
445 <Box flexDirection="row" columnGap={1}>
446 {solid('life', [[lifeFilled, clockColor], [barW - lifeFilled, 'gray']])}
447 <Text dimColor>{sp(`${Math.round(life * 100)}%`)}</Text>
448 </Box>
449 ) : null}
450 </Box>
451
452 <Box key="advice" marginTop={1} flexDirection="column">
453 <Text bold color={stateColor}>{sp(`${icon} ${advice.text}`)}</Text>
454 {last ? <Text dimColor>{sp(fit(`${last.model} · prompt ${fmtTokens(promptTokens(last))} tokens`, width))}</Text> : null}
455 </Box>
456
457 {last ? (
458 <Box key="stack" flexDirection="column" marginTop={1}>
459 <Box flexDirection="row" columnGap={1}>
460 {solid('stack', [[sr, 'green'], [sw, 'yellow'], [sn, 'cyan']])}
461 <Text bold color={hitColor(Math.round(hitRatio(last) * 100))}>{sp(`${Math.round(hitRatio(last) * 100)}% hit`)}</Text>
462 </Box>
463 <Box flexDirection="row" columnGap={2}>
464 <Text color="green">{sp(`■ read ${fmtTokens(last.read)}`)}</Text>
465 <Text color="yellow">{sp(`■ wrote ${fmtTokens(last.write)}`)}</Text>
466 <Text color="cyan">{sp(`■ new ${fmtTokens(last.fresh)}`)}</Text>
467 </Box>
468 </Box>
469 ) : null}
470
471 <Box key="table" flexDirection="column" marginTop={1}>
472 <Box key="head" flexDirection="row" columnGap={1}>
473 {cell('h:turn', 4, 'turn', 'cyan', true)}
474 {cell('h:steps', 5, 'steps', 'cyan', true)}
475 {cell('h:read', 6, 'read', 'green', true)}
476 {cell('h:wrote', 6, 'wrote', 'yellow', true)}
477 {cell('h:new', 5, 'new', 'cyan', true)}
478 {cell('h:hit', 4, 'hit', 'magenta', true)}
479 </Box>
480 {rows.length === 0 ? <Text dimColor>{sp('no requests yet')}</Text> : null}
481 {rows.map((row, i) => {
482 const n = all.length - rows.length + i + 1
483 const pct = Math.round(rowRatio(row) * 100)
484 return (
485 <Box key={`t:${row.turnId}`} flexDirection="row" columnGap={1}>
486 {cell(`c:turn:${row.turnId}`, 4, String(n))}
487 {cell(`c:steps:${row.turnId}`, 5, String(row.steps))}
488 {cell(`c:read:${row.turnId}`, 6, fmtTokens(row.read), 'green')}
489 {cell(`c:wrote:${row.turnId}`, 6, fmtTokens(row.write), 'yellow')}
490 {cell(`c:new:${row.turnId}`, 5, fmtTokens(row.fresh), 'cyan')}
491 {cell(`c:hit:${row.turnId}`, 4, `${pct}%`, hitColor(pct), true)}
492 </Box>
493 )
494 })}
495 </Box>
496
497 <Box key="foot" marginTop={1} flexDirection="column">
498 <Button key="close" label="close" onPress={() => {}} />
499 {clearing ? (
500 <Text color="yellow">{sp(clearing)}</Text>
501 ) : (
502 <Button key="clear" label="Clear: write a handoff brief, /clear, and send it to the new conversation" hotkey="c" onPress={() => void clearAndHandoff($)} />
503 )}
504 <Box key="legend" marginTop={1} flexDirection="column">
505 <Text color="green">{sp('■ read: served by the cache')}</Text>
506 <Text color="yellow">{sp('■ wrote: new cache entry')}</Text>
507 <Text color="cyan">{sp('■ new: sent uncached')}</Text>
508 </Box>
509 </Box>
510 </Box>
511 )
512 })
513}
514hooks/cache.ts 318 lines1/**
2 * cache.ts — the pure half of prompt-cache-control: no `$`, no engine.
3 *
4 * What it models, from Anthropic's prompt-caching documentation:
5 * - the cache lives 5 minutes by default, 1 hour when asked for; a read
6 * refreshes the entry at no extra cost, and the lifetime is measured from
7 * the START of the request that wrote or read it
8 * - a request's prompt is `input_tokens` (uncached remainder) +
9 * `cache_read_input_tokens` + `cache_creation_input_tokens`
10 * - writes cost 1.25x base input for 5m and 2x for 1h; reads about 0.1x
11 * (less on some models), so an expired cache on a large context is the
12 * expensive moment
13 * - a prefix change (model, effort/thinking settings, tool set, system
14 * prompt) makes the next request write instead of read
15 *
16 * Claude Code's own switches (read from the environment):
17 * ENABLE_PROMPT_CACHING_1H=1 ask for the 1-hour TTL
18 * FORCE_PROMPT_CACHING_5M=1 force the 5-minute TTL, beating the above
19 * DISABLE_PROMPT_CACHING=1 no caching; DISABLE_PROMPT_CACHING_{HAIKU,SONNET,OPUS}
20 * turn it off for that model family only
21 */
22
23export type Ttl = '5m' | '1h'
24
25export type CacheEnv = {
26 enable1h?: string
27 force5m?: string
28 /** CLAUDE_CODE_PROMPT_CACHE_TTL: "5m" or "1h" for the main conversation */
29 ttlVar?: string
30 disableAll?: string
31 disableHaiku?: string
32 disableSonnet?: string
33 disableOpus?: string
34}
35
36/** One main-loop request, as the API reported it. */
37export type Sample = {
38 turnId: string
39 index: number
40 model: string
41 /** ms since the epoch when the request started: the cache's lifetime is counted from here */
42 startedAt: number
43 read: number
44 write: number
45 fresh: number
46 output: number
47}
48
49export type AdviceKind = 'off' | 'cold' | 'uncached' | 'warm' | 'soon' | 'expired' | 'miss'
50
51export type Advice = {
52 kind: AdviceKind
53 /** one sentence for the band */
54 text: string
55}
56
57export type Policy = {
58 ttl: Ttl
59 warnMs: number
60 compactAtTokens: number
61}
62
63export const isOn = (v: string | undefined) => v === '1' || v?.toLowerCase() === 'true'
64
65/** What the account is billed as, as far as the mod can tell. */
66export type Account = 'subscription' | 'credits' | 'other'
67
68export type TtlChoice = { ttl: Ttl; source: string }
69
70const asTtl = (v: unknown): Ttl | undefined => (v === '5m' || v === '1h' ? v : undefined)
71
72/**
73 * Which lifetime Claude Code asks for on the main conversation, in the order
74 * its documentation gives (https://code.claude.com/docs/en/prompt-caching,
75 * "Choose the TTL yourself"), after the mod's own `ttl` option:
76 *
77 * FORCE_PROMPT_CACHING_5M, CLAUDE_CODE_PROMPT_CACHE_TTL, the promptCacheTtl
78 * setting, ENABLE_PROMPT_CACHING_1H, then the default of the account: one
79 * hour on a Claude subscription within its plan usage, five minutes on usage
80 * credits, an API key or a cloud provider.
81 */
82export function decideTtl(option: unknown, env: CacheEnv, setting?: unknown, account?: Account): TtlChoice {
83 const pinned = asTtl(option)
84 if (pinned) return { ttl: pinned, source: 'the ttl option' }
85 if (isOn(env.force5m)) return { ttl: '5m', source: 'FORCE_PROMPT_CACHING_5M' }
86 const fromVar = asTtl(env.ttlVar)
87 if (fromVar) return { ttl: fromVar, source: 'CLAUDE_CODE_PROMPT_CACHE_TTL' }
88 const fromSetting = asTtl(setting)
89 if (fromSetting) return { ttl: fromSetting, source: 'the promptCacheTtl setting' }
90 if (isOn(env.enable1h)) return { ttl: '1h', source: 'ENABLE_PROMPT_CACHING_1H' }
91 if (account === 'subscription') return { ttl: '1h', source: 'Claude subscription default' }
92 if (account === 'credits') return { ttl: '5m', source: 'usage credits default' }
93 return { ttl: '5m', source: 'default' }
94}
95
96export const resolveTtl = (option: unknown, env: CacheEnv, setting?: unknown, account?: Account): Ttl =>
97 decideTtl(option, env, setting, account).ttl
98
99/**
100 * The account, from the rate-limit windows the last response reported: a
101 * five-hour or seven-day window means a Claude subscription, and one that is
102 * full means the next requests draw on usage credits. No window (an API key,
103 * a cloud provider, or no response yet) says nothing.
104 */
105export function accountOf(windows: readonly { kind: string; percentUsed: number }[]): Account {
106 const plan = windows.filter(w => w.kind === 'five_hour' || w.kind === 'seven_day')
107 if (plan.length === 0) return 'other'
108 return plan.some(w => w.percentUsed >= 100) ? 'credits' : 'subscription'
109}
110
111export function ttlMs(ttl: Ttl): number {
112 return ttl === '1h' ? 3_600_000 : 300_000
113}
114
115/** Caching switched off for this model by the environment. */
116export function isCachingDisabled(model: string, env: CacheEnv): boolean {
117 if (isOn(env.disableAll)) return true
118 const name = model.toLowerCase()
119 if (name.includes('haiku')) return isOn(env.disableHaiku)
120 if (name.includes('sonnet')) return isOn(env.disableSonnet)
121 if (name.includes('opus')) return isOn(env.disableOpus)
122 return false
123}
124
125export const promptTokens = (s: Sample) => s.read + s.write + s.fresh
126
127/** Share of the prompt the cache served, 0 to 1; 0 for an empty prompt. */
128export function hitRatio(s: Sample): number {
129 const total = promptTokens(s)
130 return total === 0 ? 0 : s.read / total
131}
132
133/** When the cache entry the sample touched lapses, ms since the epoch. */
134export const expiresAt = (s: Sample, ttl: Ttl) => s.startedAt + ttlMs(ttl)
135
136/** Zero for a request that read and wrote nothing: it created or refreshed no entry, so there is nothing to count down. */
137export function remainingMs(s: Sample, ttl: Ttl, now: number): number {
138 if (s.read + s.write === 0) return 0
139 return Math.max(0, expiresAt(s, ttl) - now)
140}
141
142/**
143 * Why a request that should have read the cache wrote it instead; undefined
144 * when it did not miss. A prompt that shrank is a /compact or /clear, not a
145 * miss, and the first request of a session has nothing to read.
146 */
147export function missReason(prev: Sample | undefined, cur: Sample, ttl: Ttl): string | undefined {
148 if (!prev) return undefined
149 const before = promptTokens(prev)
150 if (before === 0 || promptTokens(cur) < before * 0.7) return undefined
151 if (cur.read >= before * 0.5 || cur.write === 0) return undefined
152 if (cur.model !== prev.model) return `model changed (${prev.model} to ${cur.model})`
153 if (cur.startedAt - prev.startedAt > ttlMs(ttl)) return `the ${ttl} cache had lapsed`
154 return 'the prompt prefix changed (effort, tools, system prompt or CLAUDE.md)'
155}
156
157export function advise(last: Sample | undefined, prev: Sample | undefined, policy: Policy, now: number, disabled: boolean): Advice {
158 if (disabled) return { kind: 'off', text: 'prompt caching is off for this model (DISABLE_PROMPT_CACHING*)' }
159 if (!last) return { kind: 'cold', text: 'no request yet: the first one writes the cache' }
160 if (last.read + last.write === 0) {
161 return { kind: 'uncached', text: 'this request was not cached (prompt under the model minimum, or caching off)' }
162 }
163 const miss = missReason(prev, last, policy.ttl)
164 const left = remainingMs(last, policy.ttl, now)
165 const size = promptTokens(last)
166 if (left <= 0) {
167 const big = size >= policy.compactAtTokens
168 return {
169 kind: 'expired',
170 text: big
171 ? `expired: the next message rewrites ${fmtTokens(size)} tokens. /compact first, or /clear if the task is done`
172 : `expired: only ${fmtTokens(size)} tokens to rebuild, just keep going`,
173 }
174 }
175 if (left <= policy.warnMs) {
176 return { kind: 'soon', text: 'expires soon: any message refreshes it for free' }
177 }
178 if (miss) return { kind: 'miss', text: `cache missed: ${miss}` }
179 return { kind: 'warm', text: 'warm: keep going' }
180}
181
182export function fmtTokens(n: number): string {
183 if (n < 1000) return String(n)
184 if (n < 100_000) return `${(n / 1000).toFixed(1).replace(/\.0$/, '')}k`
185 if (n < 1_000_000) return `${Math.round(n / 1000)}k`
186 return `${(n / 1_000_000).toFixed(1).replace(/\.0$/, '')}M`
187}
188
189/** m:ss, or h:mm:ss from an hour up. */
190export function fmtClock(ms: number): string {
191 const total = Math.max(0, Math.ceil(ms / 1000))
192 const h = Math.floor(total / 3600)
193 const m = Math.floor((total % 3600) / 60)
194 const s = total % 60
195 const pad = (n: number) => String(n).padStart(2, '0')
196 return h > 0 ? `${h}:${pad(m)}:${pad(s)}` : `${m}:${pad(s)}`
197}
198
199export function bar(ratio: number, width: number): string {
200 const filled = Math.round(Math.min(1, Math.max(0, ratio)) * width)
201 return '█'.repeat(filled) + '░'.repeat(width - filled)
202}
203
204export type TurnRow = {
205 turnId: string
206 steps: number
207 read: number
208 write: number
209 fresh: number
210 output: number
211}
212
213/** Samples grouped by turn, oldest first, each turn's requests summed. */
214export function byTurn(samples: readonly Sample[]): TurnRow[] {
215 const rows: TurnRow[] = []
216 for (const s of samples) {
217 let row = rows[rows.length - 1]
218 if (!row || row.turnId !== s.turnId) {
219 row = { turnId: s.turnId, steps: 0, read: 0, write: 0, fresh: 0, output: 0 }
220 rows.push(row)
221 }
222 row.steps += 1
223 row.read += s.read
224 row.write += s.write
225 row.fresh += s.fresh
226 row.output += s.output
227 }
228 return rows
229}
230
231export const rowRatio = (r: TurnRow) => {
232 const total = r.read + r.write + r.fresh
233 return total === 0 ? 0 : r.read / total
234}
235
236export function fit(text: string, width: number): string {
237 return text.length <= width ? text : `${text.slice(0, Math.max(0, width - 1))}…`
238}
239
240export function positive(v: unknown, fallback: number): number {
241 return typeof v === 'number' && Number.isFinite(v) && v > 0 ? v : fallback
242}
243
244/** Share of the cache lifetime left, 0 to 1. */
245export function lifeRatio(leftMs: number, ttl: Ttl): number {
246 return Math.min(1, Math.max(0, leftMs / ttlMs(ttl)))
247}
248
249/**
250 * Widths of the three stacked-bar segments (read, wrote, new) over `width`
251 * cells: proportional, each non-empty part at least one cell, summing to width.
252 */
253export function segments(read: number, write: number, fresh: number, width: number): [number, number, number] {
254 const total = read + write + fresh
255 if (total === 0 || width <= 0) return [0, 0, 0]
256 const parts = [read, write, fresh]
257 const cells = parts.map(p => (p > 0 ? Math.max(1, Math.round((p / total) * width)) : 0))
258 let over = cells.reduce((a, b) => a + b, 0) - width
259 while (over !== 0) {
260 const i = over > 0 ? cells.indexOf(Math.max(...cells)) : parts.indexOf(Math.max(...parts))
261 cells[i] += over > 0 ? -1 : 1
262 over += over > 0 ? -1 : 1
263 }
264 return [cells[0], cells[1], cells[2]]
265}
266
267/** Seconds left at which a toast counts down after the one at the warning threshold. */
268export const COUNTDOWN_MARKS = [10, 3, 2, 1]
269
270/**
271 * The toast mark to fire now, or undefined. `level` is the mark last fired for
272 * this cache entry (Infinity before any); a late tick skips straight to the
273 * newest mark crossed, so a stalled clock never replays old ones.
274 */
275export function nextToastMark(secsLeft: number, warnSecs: number, level: number): number | undefined {
276 const marks = [warnSecs, ...COUNTDOWN_MARKS].filter(m => m <= warnSecs)
277 const due = marks.filter(m => secsLeft <= m && m < level)
278 return due.length ? Math.min(...due) : undefined
279}
280
281export type LifeColor = 'green' | 'yellow' | 'red'
282
283/** Countdown colour: green while there is plenty, yellow below 40% of the lifetime, red from the warning threshold down. */
284export function lifeColor(leftMs: number, ttl: Ttl, warnMs: number): LifeColor {
285 if (leftMs <= warnMs) return 'red'
286 return leftMs / ttlMs(ttl) <= 0.4 ? 'yellow' : 'green'
287}
288
289// requests are timed from their start, so a little slack keeps a hit that
290// landed just inside the lifetime from reading as proof of the longer one
291const SLACK_MS = 10_000
292
293/**
294 * What the traffic says about the cache lifetime, given the request before and
295 * `known`, what earlier requests already showed.
296 *
297 * - a hit (the cache served at least half of the previous prompt) more than
298 * 5 minutes after the previous request began proves the 1-hour lifetime,
299 * and nothing later undoes it: a miss afterwards is more likely a changed
300 * prefix than a lapse
301 * - a miss with the same model and a prompt that did not shrink, 5 minutes to
302 * an hour after the previous request, says the entry lapsed: 5 minutes
303 * (weaker: a changed prefix looks the same, so a later hit overrules it)
304 *
305 * Needed because the API names the TTL of a write (`cache_creation.ephemeral_*`)
306 * but Claude Code's mod API passes on only the four token counts.
307 */
308export function observeTtl(prev: Sample | undefined, cur: Sample, known: Ttl | undefined): Ttl | undefined {
309 if (!prev || prev.read + prev.write === 0 || cur.model !== prev.model) return known
310 const gap = cur.startedAt - prev.startedAt
311 const before = promptTokens(prev)
312 if (gap <= ttlMs('5m') + SLACK_MS) return known
313 if (cur.read >= before * 0.5) return '1h'
314 if (known === '1h') return known
315 const lapsed = cur.write > 0 && promptTokens(cur) >= before * 0.7 && gap < ttlMs('1h') + SLACK_MS
316 return lapsed ? '5m' : known
317}
318hooks/handoff.ts 140 lines1// Pure helpers for the Clear button: turn a transcript into a digest a cheap
2// model can brief from, and build the text the fresh session receives.
3// No `$` calls here, so every function tests on its own.
4
5export type TranscriptRow = {
6 type?: string
7 isSidechain?: boolean
8 isMeta?: boolean
9 message?: { role?: string; content?: unknown }
10}
11
12export type Digest = {
13 /** Turns in order, "user: ..." / "assistant: ...", trimmed to the cap from the start. */
14 text: string
15 /** Files and commands the tools touched, most recent last, de-duplicated. */
16 touched: string[]
17 /** The person's prompts, verbatim, oldest first. */
18 prompts: string[]
19}
20
21export const DIGEST_CAP = 60_000
22const PROMPT_KEEP = 8
23const TOUCHED_KEEP = 20
24const COMMAND_PEEK = 80
25
26/** How much of a transcript's tail is parsed: a long session's file runs to many MB. */
27export const TRANSCRIPT_TAIL_CHARS = 3_000_000
28
29/** The last `cap` characters of a JSONL text, starting at a line boundary. */
30export function tailLines(text: string, cap = TRANSCRIPT_TAIL_CHARS): string {
31 if (text.length <= cap) return text
32 const cut = text.slice(text.length - cap)
33 const nl = cut.indexOf('\n')
34 return nl === -1 ? '' : cut.slice(nl + 1)
35}
36
37/** Transcript lines (JSONL) -> rows; a line that is not JSON is skipped. Only the tail is read. */
38export function parseRows(jsonl: string, cap = TRANSCRIPT_TAIL_CHARS): TranscriptRow[] {
39 const rows: TranscriptRow[] = []
40 for (const line of tailLines(jsonl, cap).split('\n')) {
41 if (!line.trim()) continue
42 try {
43 rows.push(JSON.parse(line) as TranscriptRow)
44 } catch {
45 // a half-written last line, or a non-row: not ours to repair
46 }
47 }
48 return rows
49}
50
51type Block = { type?: string; text?: string; name?: string; input?: Record<string, unknown> }
52
53/** The conversation as text: user and assistant text blocks, tool calls as one line each. */
54export function digestTranscript(rows: readonly TranscriptRow[], cap = DIGEST_CAP): Digest {
55 const lines: string[] = []
56 const prompts: string[] = []
57 const touched: string[] = []
58 for (const row of rows) {
59 if ((row.type !== 'user' && row.type !== 'assistant') || row.isSidechain || row.isMeta) continue
60 const role = row.message?.role ?? row.type
61 const content = row.message?.content
62 if (typeof content === 'string') {
63 if (role === 'user') prompts.push(content)
64 lines.push(`${role}: ${content}`)
65 continue
66 }
67 if (!Array.isArray(content)) continue
68 for (const block of content as Block[]) {
69 if (block.type === 'text' && block.text) {
70 if (role === 'user') prompts.push(block.text)
71 lines.push(`${role}: ${block.text}`)
72 } else if (block.type === 'tool_use') {
73 const target = toolTarget(block)
74 if (target) touched.push(target)
75 lines.push(`assistant used ${block.name ?? 'a tool'}${target ? `: ${target}` : ''}`)
76 }
77 // tool_result bodies are left out: large, and the assistant's text already reflects them
78 }
79 }
80 const joined = lines.join('\n')
81 const text = joined.length > cap ? `…\n${joined.slice(joined.length - cap)}` : joined
82 return {
83 text,
84 touched: dedupe(touched).slice(-TOUCHED_KEEP),
85 prompts: prompts.slice(-PROMPT_KEEP),
86 }
87}
88
89function toolTarget(block: Block): string | undefined {
90 const input = block.input ?? {}
91 const path = input.file_path ?? input.path ?? input.notebook_path
92 if (typeof path === 'string') return path
93 const command = input.command
94 if (typeof command === 'string') return command.length > COMMAND_PEEK ? `${command.slice(0, COMMAND_PEEK)}…` : command
95 return undefined
96}
97
98function dedupe(items: readonly string[]): string[] {
99 const seen = new Set<string>()
100 const out: string[] = []
101 for (const item of [...items].reverse()) {
102 if (seen.has(item)) continue
103 seen.add(item)
104 out.push(item)
105 }
106 return out.reverse()
107}
108
109export const HANDOFF_SYSTEM = `You write a handoff brief so a fresh Claude Code session can continue someone's work after the previous conversation was cleared.
110Write in plain Markdown with exactly these headings, each followed by short bullets or one line; say "none" where there is nothing:
111## Goal
112## Current state
113## Decisions made
114## Files and commands touched
115## Open items
116## Next step
117Be concrete: names of files, functions, flags, commands, errors. No preamble, no closing remarks, under 400 words.`
118
119export function handoffPrompt(digest: Digest): string {
120 const touched = digest.touched.length ? digest.touched.map(t => `- ${t}`).join('\n') : '- none recorded'
121 return `Tools touched (most recent last):\n${touched}\n\nConversation (oldest first, truncated from the start if long):\n${digest.text}`
122}
123
124/** When the model gave no brief: the person's last prompts, verbatim, and how to go back. */
125export function fallbackBrief(digest: Digest, previousSessionId: string): string {
126 const prompts = digest.prompts.length ? digest.prompts.map(p => `- ${firstLine(p)}`).join('\n') : '- none recorded'
127 const touched = digest.touched.length ? digest.touched.map(t => `- ${t}`).join('\n') : '- none recorded'
128 return `## Goal\nNot summarised (the brief could not be generated). The last prompts were:\n${prompts}\n\n## Files and commands touched\n${touched}\n\n## Next step\nAsk me what to continue with, or run \`claude --resume ${previousSessionId}\` to reopen the full conversation.`
129}
130
131/** The text the fresh session is sent: the brief, framed so the first turn stays short. */
132export function wrapForNewSession(brief: string, previousSessionId: string): string {
133 return `Handoff brief from the previous session (cleared to reset the prompt cache; its id was ${previousSessionId}, reopen it with \`claude --resume ${previousSessionId}\` if you need the full history).\n\nRead it, then acknowledge in one line and wait for my next instruction. Do not start any work yet.\n\n${brief.trim()}`
134}
135
136function firstLine(text: string): string {
137 const line = (text.split('\n')[0] ?? '').trim()
138 return line.length > 160 ? `${line.slice(0, 160)}…` : line
139}
140