SLOPSHOPPER

prompt-cache-control

Prompt-cache meter above the Claude Code prompt: how many tokens each request read from, wrote to and sent past the cache, a live countdown to the cache's…

newpanebandcommandtoaststatus
★ 1v0.2.0MITupdated 2026-10-08Nasrallah-Adel/claude-prompt-cache-control
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · prompt-cache-control
│ ┃ cache ✕ › fix the failing auth test and add an audit log call │ ┃ ⚡ PROMPT CACHE · 1h lifetime (Claude subscri │ ┃ ● prompt-cache-control: prompt-cache-control loaded: 1h cache (Claude │ ┃ ⏱ --:-- ⏺ Read(src/auth.ts) │ ┃ ⎿ Read 6 lines │ ┃ ○ no request yet: the first one writes the ⏺ Update(src/auth.ts) │ ┃ cache ⎿ Added 2 lines, removed 1 line │ ┃ ⏺ Bash(bun test) │ ┃ turn steps read wrote new hit ⎿ 3 pass, 1 fail │ ┃ no requests yet │ ┃ ● Done. refresh now rejects expired claims and logs an audit event. │ ┃ [ close ] │ ┃ [ Clear: write a handoff brief, /clear, and ✻ Worked for 42s · done 4:20 PM │ ┃ │ ┃ ■ read: served by the cache › /cache │ ┃ ■ wrote: new cache entry ⎿ prompt-cache-control: 1h cache (Claude subscription default) · n │ ┃ ■ new: sent uncached │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · cache
⚡ PROMPT CACHE · 1h lifetime (Claude subscription default) ⏱ --:-- ○ no request yet: the first one writes the cache turn steps read wrote new hit no requests yet [ close ] [ Clear: write a handoff brief, /clear, and send it to the n ■ read: served by the cache ■ wrote: new cache entry ■ new: sent uncached
README

prompt-cache-control

A prompt-cache meter above the Claude Code prompt. Every request Claude makes reports how much of its prompt the cache served, how much it wrote and how much went uncached; this mod keeps those numbers per request and per turn, counts down to the moment the cache lapses and tells you what to do about it: keep going, /compact or /clear.

cache ██████████ 98% read 80k · wrote 1k · new 300 ⏱ 3:41 5m · warm: keep going
cache ░░░░░░░░░░  0% read 0 · wrote 52k · new 300 ⏱ 4:58 5m · cache missed: model changed (…)
cache ██████████ 98% read 150k · wrote 1k · new 300 ⏱ 0:00 5m · expired: the next message rewrites 151k tokens. /compact first, or /clear if the task is done

/cache opens a pane: the time left with a solid bar that shrinks as the cache runs out (green, yellow below 40% of the lifetime, red from the warning threshold), a stacked read / wrote / new bar for the last request, and a colour-coded table with one row per turn. Bars are filled cells and columns have fixed widths with no-break spaces, so the terminal and the Desktop (HTML) pane render the same.

Clear button

The band and the /cache pane carry a Clear button (hotkey c once the band or pane has focus: ctrl+x tab for the band). Pressing it:

  1. reads this session's transcript (~/.claude/projects/<folder>/<session>.jsonl), keeps the user and assistant text and one line per tool call, and asks haiku for a brief with fixed headings: Goal, Current state, Decisions made, Files and commands touched, Open items, Next step;
  2. runs /clear, so the next request starts from an empty, cheap prefix;
  3. submits the brief to the fresh conversation as your own first prompt, framed to ask for a one-line acknowledgement and then wait.

If the model gives no brief, a fallback with your last prompts and a claude --resume <id> line is sent instead. If /clear or the submit is refused, the brief is put in the prompt box for you to send. The button is hidden while a turn is running. The old session id is in the brief, so claude --resume gets the full history back.

How the countdown works

From Anthropic's prompt caching documentation:

  • The cache lives 5 minutes by default, 1 hour when asked for.
  • Every request that reads the cache refreshes it at no extra cost, so a conversation that keeps talking keeps the 5-minute cache warm.
  • The lifetime is counted from the start of the request that wrote or read the entry; generation time counts against it.
  • A prompt is input_tokens (uncached remainder) + cache_read_input_tokens + cache_creation_input_tokens.
  • Writes cost 1.25x base input for 5 minutes and 2x for 1 hour; reads cost about 0.1x (less on some models). The expensive moment is an expired cache on a large context, which is when this mod suggests /compact.
  • /clear starts a new conversation in the same process, so the meter and the /cache table start over with it. A change in the prefix (model, effort or thinking settings, tool set, system prompt, CLAUDE.md) makes the next request write instead of read. The mod names the cause when it sees a miss: model changed, the cache had lapsed, or the prefix changed.

Which lifetime your account gets

The mod follows Claude Code's own rules (prompt caching: cache lifetime, Claude Code 2.1.242 or later). For the main conversation the TTL is the first match of:

#SourceResult
1the mod's ttl option (5m / 1h)what you set
2FORCE_PROMPT_CACHING_5M=15 minutes
3CLAUDE_CODE_PROMPT_CACHE_TTL5m or 1h
4the promptCacheTtl setting (local, project or user settings file)5m or 1h
5ENABLE_PROMPT_CACHING_1H=11 hour
6the account1 hour on a Claude subscription within its plan usage; 5 minutes on usage credits, an API key or a cloud provider

The account comes from the rate-limit windows the last response reported: a five_hour or seven_day window means a subscription, and one at 100% means requests now draw on usage credits. An API key or a cloud provider reports no such window, and before the first response nothing is known, so the mod starts from 5 minutes there. Managed settings are not readable from a mod.

On top of that the mod watches the traffic, which beats rows 2 to 6: a request that hits the cache more than 5 minutes after the previous one proves the 1-hour lifetime (a later miss does not undo it, since a changed prefix looks the same), and a miss 5 to 60 minutes after the previous request, with the same model and a prompt that did not shrink, says the entry lapsed, so 5 minutes (a later hit overrules it). That covers what the mod cannot see: managed settings, a gateway that rewrites the TTL, or a subscription that ran out of plan usage mid-session. The pane header names the source in use.

Why the mod infers instead of reading it: the API names the TTL of each write (cache_creation.ephemeral_5m_input_tokens / ephemeral_1h_input_tokens) and Claude Code's status line exposes it as prompt_cache.ttl, but the mod API passes on only the four token counts. To check by hand, claude -p "hello" --output-format json and read usage.cache_creation.

Other switches read from the environment at session start:

VariableEffect on the meter
DISABLE_PROMPT_CACHING=1 (and _HAIKU, _SONNET, _OPUS)the band says caching is off for that model

What it hooks

  • turn.step: reads each main-loop request's usage (subagents have their own prefixes and are left out)
  • $.clock.every(1000): redraws the countdown, and only while its text changes, so an idle expired session costs nothing
  • ui.render on AbovePrompt (the band) and on Pane (/cache)
  • $.ui.toast: once per cache entry at the warning threshold (60 s by default) and again at 10, 3, 2 and 1 seconds left, for prompts of 20k tokens or more

Options

  ttl: string               "auto" | "5m" | "1h" (default auto)
  warnSeconds: number       countdown threshold for the yellow state and the toast (default 60)
  compactAtTokens: number   prompt size that makes an expired cache suggest /compact (default 100000)
  band: boolean             row above the prompt (default true)
  status: boolean           entry under the prompt, "cache 98% · 3:41" (default false)
  toast: boolean            toasts at the threshold, 10, 3, 2 and 1 s (default true)

The 100k compactAtTokens is a judgement, not a figure from the documentation: lower it if your model's cache writes are expensive for you.

Install

In a Claude Code terminal session (2.1.287 or later, where mods load by default):

/plugin install prompt-cache-control --marketplace Nasrallah-Adel/claude-prompt-cache-control

Answer y to add the marketplace, then pick the user scope (Enter) so it loads in every project. It is active at once in that session and in each session started after.

To try it from a clone instead, for one session with hot reload:

git clone https://github.com/Nasrallah-Adel/claude-prompt-cache-control.git
claude --plugin-dir ./claude-prompt-cache-control

It is written to .claude/skills/prompt-cache-control/, which Claude Code auto-loads as prompt-cache-control@skills-dir once the workspace trust prompt is accepted. For one session with hot reload: claude --plugin-dir .claude/skills/prompt-cache-control. claude plugin validate .claude/skills/prompt-cache-control prints every event it hooks and every $ call it makes; claude plugin test .claude/skills/prompt-cache-control runs its tests.

Options are read from user settings (~/.claude/settings.json, never project settings), --settings <file> or managed settings, under the plugin's full id:

{ "pluginConfigs": { "prompt-cache-control@skills-dir": { "options": { } } } }

Requirements. Mods are on by default in Claude Code 2.1.287+. Typed against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods

License

MIT. The cache meter started from Daniel Ávila's prompt-cache-control mod (davila7/claude-code-templates); the Clear button, the handoff flow and the band layout are this repo's. LICENSE carries both copyright lines, as MIT requires.

Source 3 files
hooks/prompt-cache-control.tsx 514 lines
1/**
2 * prompt-cache-control — Claude Mod
3 *
4 * A prompt-cache meter for Claude Code. Every main-loop request reports how
5 * many prompt tokens the cache served (`cache_read_input_tokens`), wrote
6 * (`cache_creation_input_tokens`) and sent uncached (`input_tokens`); this mod
7 * keeps those per request and per turn, counts down to the moment the cache
8 * lapses, and says what to do about it: keep going, /compact or /clear.
9 *
10 *   - `turn.step` reads each main-loop request's usage (subagents have their
11 *     own prefixes and are left out)
12 *   - `$.clock.every(1000)` redraws the countdown, and only while its text
13 *     changes: an idle, expired session costs nothing
14 *   - a row above the prompt (the AbovePrompt component), an optional status
15 *     line entry, and `/cache`, a pane with one row per turn
16 *
17 * The lifetime is counted from the start of the request that last wrote or read
18 * the cache, as Anthropic documents it. Which lifetime Claude Code asked for
19 * follows its documented rules (see decideTtl in ./cache.ts): FORCE_PROMPT_CACHING_5M,
20 * CLAUDE_CODE_PROMPT_CACHE_TTL, the promptCacheTtl setting, ENABLE_PROMPT_CACHING_1H,
21 * then the account (1 hour on a Claude subscription, 5 minutes otherwise). The
22 * API names the TTL of a write but the mod API passes on only the token counts,
23 * so the mod also watches the gaps between requests (a hit after more than 5
24 * minutes proves 1 hour; see observeTtl). `ttl: "5m" | "1h"` pins it.
25 *
26 * Needs Claude Code >= 2.1.287.
27 *
28 * Options (pluginConfigs["prompt-cache-control@skills-dir"].options):
29 *   ttl: "auto" | "5m" | "1h"   cache lifetime (default auto)
30 *   warnSeconds: number         countdown threshold for the warning (default 60)
31 *   compactAtTokens: number     prompt size that makes an expired cache suggest /compact (default 100000)
32 *   band: boolean               row above the prompt (default true)
33 *   status: boolean             entry under the prompt (default false)
34 *   toast: boolean              toasts near expiry: at warnSeconds, then 10, 3, 2 and 1 s (default true)
35 */
36import type { EngineInterface, Register } from 'claude-code'
37import {
38  advise,
39  COUNTDOWN_MARKS,
40  bar,
41  byTurn,
42  fit,
43  fmtClock,
44  fmtTokens,
45  hitRatio,
46  isCachingDisabled,
47  accountOf,
48  decideTtl,
49  observeTtl,
50  lifeColor,
51  lifeRatio,
52  nextToastMark,
53  positive,
54  promptTokens,
55  remainingMs,
56  rowRatio,
57  segments,
58} from './cache.ts'
59import type { Account, Advice, CacheEnv, Sample, Ttl } from './cache.ts'
60import { digestTranscript, fallbackBrief, HANDOFF_SYSTEM, handoffPrompt, parseRows, wrapForNewSession } from './handoff.ts'
61
62const HANDOFF_MODEL = 'haiku'
63const HANDOFF_MAX_TOKENS = 1200
64const HANDOFF_TIMEOUT_MS = 90_000
65
66const PANE = 'cache'
67const COMMAND = 'cache'
68const KEEP = 200
69// below this a lapsed cache costs too little to interrupt anyone about
70const TOAST_MIN_TOKENS = 20_000
71
72let samples: Sample[] = []
73let ttl: Ttl = '5m'
74let baseTtl: Ttl = '5m'
75let pinned = false
76let observed: Ttl | undefined
77let setting: unknown
78let account: Account = 'other'
79let ttlSource = 'default'
80let envSource = 'default'
81let env: CacheEnv = {}
82let timer: { cancel: () => void } | undefined
83let lastKey = ''
84let toastedFor = 0
85let toastLevel = Infinity
86let isPaneOpen = false
87
88type Policy = { warnMs: number; compactAtTokens: number }
89
90function current(policy: Policy, now: number) {
91  const last = samples[samples.length - 1]
92  const prev = samples[samples.length - 2]
93  const disabled = last ? isCachingDisabled(last.model, env) : isCachingDisabled('', env)
94  const advice: Advice = advise(last, prev, { ttl, ...policy }, now, disabled)
95  const left = last ? remainingMs(last, ttl, now) : 0
96  return { last, advice, left }
97}
98
99const COLOR: Record<Advice['kind'], string | undefined> = {
100  warm: 'green',
101  soon: 'yellow',
102  expired: 'red',
103  miss: 'red',
104  off: undefined,
105  cold: undefined,
106  uncached: undefined,
107}
108
109function shortLine(policy: Policy, now: number): string {
110  const { last, advice, left } = current(policy, now)
111  if (!last || advice.kind === 'off') return `cache: ${advice.text}`
112  const clock = left > 0 ? ` · ${fmtClock(left)}` : ''
113  return `cache ${Math.round(hitRatio(last) * 100)}%${clock}`
114}
115
116// the promptCacheTtl setting, from the settings files that can carry it (local over project over user)
117async function readSetting($: EngineInterface): Promise<unknown> {
118  const home = await $.env.get('HOME').catch(() => undefined)
119  const cwd = await $.session.cwd().catch(() => undefined)
120  const files = [cwd && `${cwd}/.claude/settings.local.json`, cwd && `${cwd}/.claude/settings.json`, home && `${home}/.claude/settings.json`]
121  for (const file of files) {
122    if (!file) continue
123    try {
124      const value = JSON.parse(await $.fs.read(file)).promptCacheTtl
125      if (value === '5m' || value === '1h') return value
126    } catch {
127      // missing or unreadable: the next file
128    }
129  }
130  return undefined
131}
132
133// The Clear button's progress text while it runs; undefined when idle.
134let clearing: string | undefined
135
136/** Read this session's transcript file, if the engine wrote one where it usually does. */
137async function readTranscript($: EngineInterface, sessionId: string): Promise<string | undefined> {
138  const home = await $.env.get('HOME').catch(() => undefined)
139  if (!home) return undefined
140  const projects = `${home}/.claude/projects`
141  const dirs = await $.fs.list(projects).catch(() => [])
142  for (const dir of dirs) {
143    if (dir.kind !== 'dir') continue
144    const path = `${projects}/${dir.name}/${sessionId}.jsonl`
145    if (await $.fs.exists(path).catch(() => false)) {
146      return $.fs.read(path).catch(() => undefined) as Promise<string | undefined>
147    }
148  }
149  return undefined
150}
151
152/** Clear: write a handoff brief from the transcript, /clear, then hand the brief to the fresh session. */
153async function clearAndHandoff($: EngineInterface): Promise<void> {
154  if (clearing) return
155  const step = (text: string | undefined) => {
156    clearing = text
157    $.ui.invalidate('ui.render')
158  }
159  step('writing handoff…')
160  $.ui.toast('cache: writing a handoff brief, then clearing')
161  const sessionId = await $.session.id().catch(() => 'unknown')
162  let brief: string
163  try {
164    const digest = digestTranscript(parseRows((await readTranscript($, sessionId)) ?? ''))
165    const r = await $.model.complete({
166      model: HANDOFF_MODEL,
167      system: HANDOFF_SYSTEM,
168      prompt: handoffPrompt(digest),
169      maxTokens: HANDOFF_MAX_TOKENS,
170      effort: 'low',
171      timeoutMs: HANDOFF_TIMEOUT_MS,
172    })
173    brief = r.isAnswered && r.text.trim() ? r.text : fallbackBrief(digest, sessionId)
174    if (!r.isAnswered) $.ui.log(`prompt-cache-control: handoff brief not generated (${r.reason}); sent the fallback`, { to: 'debug' })
175  } catch (err) {
176    brief = fallbackBrief(digestTranscript([]), sessionId)
177    $.ui.log(`prompt-cache-control: handoff failed (${String(err)}); sent the fallback`, { to: 'debug' })
178  }
179  const text = wrapForNewSession(brief, sessionId)
180  try {
181    step('clearing…')
182    await $.command.run({ command: 'clear' })
183    step('sending handoff…')
184    await $.prompt.submit({ text, asUser: true })
185    $.ui.toast('cache: cleared; the handoff brief is in the new conversation')
186  } catch (err) {
187    $.ui.log(`prompt-cache-control: clear or submit failed (${String(err)})`, { to: 'debug' })
188    const filled = await $.prompt.fill({ text }).catch(() => ({ isFilled: false }))
189    $.ui.toast(filled.isFilled ? 'cache: handoff is in the prompt box, press Enter to send it' : 'cache: could not clear; run /clear yourself')
190  } finally {
191    step(undefined)
192  }
193}
194
195export const register: Register = (on, options) => {
196  const policy: Policy = {
197    warnMs: positive(options.warnSeconds, 60) * 1000,
198    compactAtTokens: positive(options.compactAtTokens, 100_000),
199  }
200  const showBand = options.band !== false
201  const showStatus = options.status === true
202  const wantToast = options.toast !== false
203
204  on('session.start', async ($, e, next) => {
205    const r = await next(e)
206    samples = []
207    lastKey = ''
208    toastedFor = 0
209    const none = () => undefined
210    env = {
211      enable1h: await $.env.get('ENABLE_PROMPT_CACHING_1H').catch(none),
212      force5m: await $.env.get('FORCE_PROMPT_CACHING_5M').catch(none),
213      ttlVar: await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL').catch(none),
214      disableAll: await $.env.get('DISABLE_PROMPT_CACHING').catch(none),
215      disableHaiku: await $.env.get('DISABLE_PROMPT_CACHING_HAIKU').catch(none),
216      disableSonnet: await $.env.get('DISABLE_PROMPT_CACHING_SONNET').catch(none),
217      disableOpus: await $.env.get('DISABLE_PROMPT_CACHING_OPUS').catch(none),
218    }
219    pinned = options.ttl === '5m' || options.ttl === '1h'
220    observed = undefined
221    setting = await readSetting($)
222    account = accountOf((await $.session.usage().catch(() => undefined))?.rateLimits ?? [])
223    const choice = decideTtl(options.ttl, env, setting, account)
224    baseTtl = choice.ttl
225    ttl = baseTtl
226    envSource = choice.source
227    ttlSource = envSource
228
229    await $.command
230      .register({
231        name: COMMAND,
232        description: 'Prompt-cache usage per turn and the time left before it lapses (stop closes)',
233        argumentHint: '[stop]',
234        immediate: true,
235      })
236      .catch(err => $.ui.log(`prompt-cache-control: /${COMMAND} not registered: ${err}`))
237    $.ui.log(`prompt-cache-control loaded: ${ttl} cache (${ttlSource}), /${COMMAND} opens the table`, { to: 'debug' })
238
239    timer?.cancel()
240    timer = $.clock.every(1000, () => {
241      const now = Date.now()
242      const { last, advice, left } = current(policy, now)
243      const key = `${advice.kind}|${advice.text}|${left > 0 ? fmtClock(left) : ''}`
244      if (key !== lastKey) {
245        lastKey = key
246        if (showStatus) $.ui.status(shortLine(policy, now))
247        $.ui.invalidate('ui.render')
248      }
249      if (wantToast && last && left > 0 && promptTokens(last) >= TOAST_MIN_TOKENS) {
250        if (toastedFor !== last.startedAt) {
251          toastedFor = last.startedAt
252          toastLevel = Infinity
253        }
254        // the first toast comes at warnSeconds, then 10, 3, 2 and 1 seconds; a late tick skips to the newest one
255        const secs = Math.ceil(left / 1000)
256        const mark = nextToastMark(secs, policy.warnMs / 1000, toastLevel)
257        if (mark !== undefined) {
258          toastLevel = mark
259          const tail = secs <= COUNTDOWN_MARKS[0] ? 'send a message now' : `send a message to keep ${fmtTokens(promptTokens(last))} tokens warm`
260          $.ui.toast(`cache expires in ${secs >= 60 ? fmtClock(left) : `${secs}s`}: ${tail}`)
261        }
262      }
263    })
264    return r
265  })
266
267  on('session.end', async ($, e, next) => {
268    // /clear starts a new conversation in the same process: its cache is a new one
269    if (e.reason === 'clear') {
270      samples = []
271      lastKey = ''
272      toastedFor = 0
273      observed = undefined
274      ttl = baseTtl
275      ttlSource = envSource
276      $.ui.invalidate('ui.render')
277      return next(e)
278    }
279    timer?.cancel()
280    timer = undefined
281    return next(e)
282  })
283
284  // each main-loop request: what the cache did with it
285  on('turn.step', async function* ($, e, next) {
286    if (e.agentId) return yield* next(e)
287    const startedAt = Date.now()
288    const r = yield* next(e)
289    if (r.usage) {
290      samples.push({
291        turnId: e.turnId,
292        index: e.index,
293        model: r.usage.model || e.model,
294        startedAt,
295        read: r.usage.cache_read_input_tokens,
296        write: r.usage.cache_creation_input_tokens,
297        fresh: r.usage.input_tokens,
298        output: r.usage.output_tokens,
299      })
300      if (samples.length > KEEP) samples = samples.slice(-KEEP)
301      if (!pinned) {
302        // the account can change under a session: a subscription running out of plan usage moves to usage credits
303        account = accountOf((await $.session.usage().catch(() => undefined))?.rateLimits ?? [])
304        const choice = decideTtl(options.ttl, env, setting, account)
305        baseTtl = choice.ttl
306        envSource = choice.source
307        if (observed === undefined) {
308          ttl = baseTtl
309          ttlSource = envSource
310        }
311        const seen = observeTtl(samples[samples.length - 2], samples[samples.length - 1], observed)
312        if (seen !== observed) {
313          observed = seen
314          ttl = seen ?? baseTtl
315          ttlSource = `observed from request timing; ${envSource} said ${baseTtl}`
316          $.ui.log(`prompt-cache-control: cache lifetime is ${ttl} (${ttlSource})`, { to: 'debug' })
317        }
318      }
319      lastKey = ''
320      if (showStatus) $.ui.status(shortLine(policy, Date.now()))
321      $.ui.invalidate('ui.render')
322    }
323    return r
324  })
325
326  on('command.run', { command: COMMAND }, async ($, e) => {
327    if (e.args.trim().toLowerCase() === 'stop') {
328      await $.ui.close({ id: PANE }).catch(() => undefined)
329      isPaneOpen = false
330      return { text: 'cache table closed' }
331    }
332    isPaneOpen = true
333    await $.ui.open({ id: PANE, title: 'cache', focus: true })
334    $.ui.invalidate('ui.render')
335    const { advice } = current(policy, Date.now())
336    return { text: `${ttl} cache (${ttlSource}) · ${advice.text} · /${COMMAND} stop closes` }
337  })
338
339  on('ui.close', async ($, e, next) => {
340    if (e.id !== PANE) return next(e)
341    isPaneOpen = false
342    return next(e)
343  })
344
345  on('ui.press', async ($, e, next) => {
346    if (e.plugin !== $.plugin.name || e.requestId !== PANE) return next(e)
347    if (e.element === 'close') await $.ui.close({ id: PANE }).catch(() => undefined)
348    return next(e)
349  })
350
351  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
352    if (!showBand || e.props.hasSurvey || isPaneOpen) return next(e)
353    const { last, advice, left } = current(policy, Date.now())
354    if (!last && advice.kind !== 'off') return next(e)
355    const { Box, Text, Button } = $.ui.resolve(e)
356    const columns = e.viewport?.columns ?? 100
357    const color = COLOR[advice.kind]
358    // The band is one instance: whatever the plugins beneath draw (another mod's band) stays, under this line.
359    const below = await next(e)
360    const stack = (own: ReturnType<typeof Text>) => (
361      <Box flexDirection="column">
362        {own}
363        {below}
364      </Box>
365    )
366    // Clear: brief + /clear + handoff. Always drawn; mid-turn it only says to wait (the brief would go stale).
367    const busy = e.props.isWorking
368    const clearControl = clearing ? (
369      <Text color="yellow">{clearing}</Text>
370    ) : (
371      <Button
372        key="clear"
373        label="Clear"
374        hotkey="c"
375        onPress={() => (busy ? $.ui.toast('cache: wait for the turn to finish, then press Clear') : void clearAndHandoff($))}
376      />
377    )
378
379    if (!last) return stack(<Text dimColor>{fit(`cache: ${advice.text}`, columns)}</Text>)
380
381    const ratio = hitRatio(last)
382    const wide = columns >= 90
383    return stack(
384      <Box flexDirection="row" columnGap={1}>
385        <Text bold color={color}>{advice.kind === 'warm' ? '●' : advice.kind === 'soon' ? '▲' : advice.kind === 'off' || advice.kind === 'cold' || advice.kind === 'uncached' ? '○' : '✖'}</Text>
386        <Text bold color="cyan">cache</Text>
387        <Text color={color}>{bar(ratio, wide ? 10 : 6)}</Text>
388        <Text bold>{`${Math.round(ratio * 100)}%`}</Text>
389        {wide && <Text color="green">{`read ${fmtTokens(last.read)}`}</Text>}
390        {wide && <Text color="yellow">{`wrote ${fmtTokens(last.write)}`}</Text>}
391        {wide && <Text color="cyan">{`new ${fmtTokens(last.fresh)}`}</Text>}
392        {!wide && <Text dimColor>{`${fmtTokens(promptTokens(last))} tok`}</Text>}
393        {advice.kind !== 'uncached' && advice.kind !== 'off' && (
394          <Text bold color={left > 0 ? lifeColor(left, ttl, policy.warnMs) : 'red'}>{left > 0 ? `⏱ ${fmtClock(left)}` : '⏱ 0:00'}</Text>
395        )}
396        {clearControl}
397        <Text dimColor wrap="truncate-end">{`${ttl} · ${advice.text}`}</Text>
398      </Box>,
399    )
400  })
401
402  on('ui.render', { component: 'Pane' }, async ($, e, next) => {
403    if (e.requestId !== PANE) return next(e)
404    const { Box, Text, Button } = $.ui.resolve(e)
405    const width = Math.max(30, e.props.bodyColumns - 1)
406    // HTML collapses runs of spaces and trims a text's ends; a no-break space keeps them
407    const sp = (t: string) => (e.surface === 'terminal' ? t : t.replace(/ /g, ' '))
408    const now = Date.now()
409    const { last, advice, left } = current(policy, now)
410    const all = byTurn(samples)
411    const counting = !!last && advice.kind !== 'uncached' && advice.kind !== 'off'
412    // the countdown goes green, then yellow, then red as the cache runs out
413    const clockColor = counting ? lifeColor(left, ttl, policy.warnMs) : undefined
414    const stateColor = advice.kind === 'expired' || advice.kind === 'miss' ? 'red' : (clockColor ?? COLOR[advice.kind])
415    const hitColor = (pct: number) => (pct >= 80 ? 'green' : pct >= 40 ? 'yellow' : 'red')
416    // solid bars are filled Boxes, not block characters, so HTML draws no seams between cells
417    const solid = (key: string, parts: [number, string | undefined][]) => (
418      <Box key={key} flexDirection="row" height={1} flexShrink={0}>
419        {parts.map(([w, c], i) => (w > 0 ? <Box key={`${key}:${i}`} width={w} height={1} flexShrink={0} backgroundColor={c} /> : null))}
420      </Box>
421    )
422    const cell = (key: string, w: number, text: string, c?: string, bold = false) => (
423      <Box key={key} width={w} flexShrink={0} justifyContent="flex-end">
424        <Text color={c} bold={bold} dimColor={!c}>{sp(text)}</Text>
425      </Box>
426    )
427
428    const barW = Math.min(width, 40)
429    const life = lifeRatio(left, ttl)
430    const lifeFilled = Math.round(life * barW)
431    const [sr, sw, sn] = last ? segments(last.read, last.write, last.fresh, barW) : [0, 0, 0]
432    const rows = all.slice(-Math.max(3, (e.viewport?.rows ?? 24) - 16))
433    const icon = advice.kind === 'warm' ? '●' : advice.kind === 'soon' ? '▲' : advice.kind === 'expired' || advice.kind === 'miss' ? '✖' : '○'
434
435    return (
436      <Box flexDirection="column">
437        <Box key="title" flexDirection="row" columnGap={1}>
438          <Text bold color="cyan">{sp('⚡ PROMPT CACHE')}</Text>
439          <Text dimColor>{sp(`· ${ttl} lifetime (${ttlSource})`)}</Text>
440        </Box>
441
442        <Box key="clock" flexDirection="column" marginTop={1}>
443          <Text bold color={clockColor}>{sp(counting ? `⏱ ${left > 0 ? fmtClock(left) : '0:00'}` : '⏱ --:--')}</Text>
444          {counting ? (
445            <Box flexDirection="row" columnGap={1}>
446              {solid('life', [[lifeFilled, clockColor], [barW - lifeFilled, 'gray']])}
447              <Text dimColor>{sp(`${Math.round(life * 100)}%`)}</Text>
448            </Box>
449          ) : null}
450        </Box>
451
452        <Box key="advice" marginTop={1} flexDirection="column">
453          <Text bold color={stateColor}>{sp(`${icon} ${advice.text}`)}</Text>
454          {last ? <Text dimColor>{sp(fit(`${last.model} · prompt ${fmtTokens(promptTokens(last))} tokens`, width))}</Text> : null}
455        </Box>
456
457        {last ? (
458          <Box key="stack" flexDirection="column" marginTop={1}>
459            <Box flexDirection="row" columnGap={1}>
460              {solid('stack', [[sr, 'green'], [sw, 'yellow'], [sn, 'cyan']])}
461              <Text bold color={hitColor(Math.round(hitRatio(last) * 100))}>{sp(`${Math.round(hitRatio(last) * 100)}% hit`)}</Text>
462            </Box>
463            <Box flexDirection="row" columnGap={2}>
464              <Text color="green">{sp(`■ read ${fmtTokens(last.read)}`)}</Text>
465              <Text color="yellow">{sp(`■ wrote ${fmtTokens(last.write)}`)}</Text>
466              <Text color="cyan">{sp(`■ new ${fmtTokens(last.fresh)}`)}</Text>
467            </Box>
468          </Box>
469        ) : null}
470
471        <Box key="table" flexDirection="column" marginTop={1}>
472          <Box key="head" flexDirection="row" columnGap={1}>
473            {cell('h:turn', 4, 'turn', 'cyan', true)}
474            {cell('h:steps', 5, 'steps', 'cyan', true)}
475            {cell('h:read', 6, 'read', 'green', true)}
476            {cell('h:wrote', 6, 'wrote', 'yellow', true)}
477            {cell('h:new', 5, 'new', 'cyan', true)}
478            {cell('h:hit', 4, 'hit', 'magenta', true)}
479          </Box>
480          {rows.length === 0 ? <Text dimColor>{sp('no requests yet')}</Text> : null}
481          {rows.map((row, i) => {
482            const n = all.length - rows.length + i + 1
483            const pct = Math.round(rowRatio(row) * 100)
484            return (
485              <Box key={`t:${row.turnId}`} flexDirection="row" columnGap={1}>
486                {cell(`c:turn:${row.turnId}`, 4, String(n))}
487                {cell(`c:steps:${row.turnId}`, 5, String(row.steps))}
488                {cell(`c:read:${row.turnId}`, 6, fmtTokens(row.read), 'green')}
489                {cell(`c:wrote:${row.turnId}`, 6, fmtTokens(row.write), 'yellow')}
490                {cell(`c:new:${row.turnId}`, 5, fmtTokens(row.fresh), 'cyan')}
491                {cell(`c:hit:${row.turnId}`, 4, `${pct}%`, hitColor(pct), true)}
492              </Box>
493            )
494          })}
495        </Box>
496
497        <Box key="foot" marginTop={1} flexDirection="column">
498          <Button key="close" label="close" onPress={() => {}} />
499          {clearing ? (
500            <Text color="yellow">{sp(clearing)}</Text>
501          ) : (
502            <Button key="clear" label="Clear: write a handoff brief, /clear, and send it to the new conversation" hotkey="c" onPress={() => void clearAndHandoff($)} />
503          )}
504          <Box key="legend" marginTop={1} flexDirection="column">
505            <Text color="green">{sp('■ read: served by the cache')}</Text>
506            <Text color="yellow">{sp('■ wrote: new cache entry')}</Text>
507            <Text color="cyan">{sp('■ new: sent uncached')}</Text>
508          </Box>
509        </Box>
510      </Box>
511    )
512  })
513}
514
hooks/cache.ts 318 lines
1/**
2 * cache.ts — the pure half of prompt-cache-control: no `$`, no engine.
3 *
4 * What it models, from Anthropic's prompt-caching documentation:
5 *   - the cache lives 5 minutes by default, 1 hour when asked for; a read
6 *     refreshes the entry at no extra cost, and the lifetime is measured from
7 *     the START of the request that wrote or read it
8 *   - a request's prompt is `input_tokens` (uncached remainder) +
9 *     `cache_read_input_tokens` + `cache_creation_input_tokens`
10 *   - writes cost 1.25x base input for 5m and 2x for 1h; reads about 0.1x
11 *     (less on some models), so an expired cache on a large context is the
12 *     expensive moment
13 *   - a prefix change (model, effort/thinking settings, tool set, system
14 *     prompt) makes the next request write instead of read
15 *
16 * Claude Code's own switches (read from the environment):
17 *   ENABLE_PROMPT_CACHING_1H=1   ask for the 1-hour TTL
18 *   FORCE_PROMPT_CACHING_5M=1    force the 5-minute TTL, beating the above
19 *   DISABLE_PROMPT_CACHING=1     no caching; DISABLE_PROMPT_CACHING_{HAIKU,SONNET,OPUS}
20 *                                 turn it off for that model family only
21 */
22
23export type Ttl = '5m' | '1h'
24
25export type CacheEnv = {
26  enable1h?: string
27  force5m?: string
28  /** CLAUDE_CODE_PROMPT_CACHE_TTL: "5m" or "1h" for the main conversation */
29  ttlVar?: string
30  disableAll?: string
31  disableHaiku?: string
32  disableSonnet?: string
33  disableOpus?: string
34}
35
36/** One main-loop request, as the API reported it. */
37export type Sample = {
38  turnId: string
39  index: number
40  model: string
41  /** ms since the epoch when the request started: the cache's lifetime is counted from here */
42  startedAt: number
43  read: number
44  write: number
45  fresh: number
46  output: number
47}
48
49export type AdviceKind = 'off' | 'cold' | 'uncached' | 'warm' | 'soon' | 'expired' | 'miss'
50
51export type Advice = {
52  kind: AdviceKind
53  /** one sentence for the band */
54  text: string
55}
56
57export type Policy = {
58  ttl: Ttl
59  warnMs: number
60  compactAtTokens: number
61}
62
63export const isOn = (v: string | undefined) => v === '1' || v?.toLowerCase() === 'true'
64
65/** What the account is billed as, as far as the mod can tell. */
66export type Account = 'subscription' | 'credits' | 'other'
67
68export type TtlChoice = { ttl: Ttl; source: string }
69
70const asTtl = (v: unknown): Ttl | undefined => (v === '5m' || v === '1h' ? v : undefined)
71
72/**
73 * Which lifetime Claude Code asks for on the main conversation, in the order
74 * its documentation gives (https://code.claude.com/docs/en/prompt-caching,
75 * "Choose the TTL yourself"), after the mod's own `ttl` option:
76 *
77 *   FORCE_PROMPT_CACHING_5M, CLAUDE_CODE_PROMPT_CACHE_TTL, the promptCacheTtl
78 *   setting, ENABLE_PROMPT_CACHING_1H, then the default of the account: one
79 *   hour on a Claude subscription within its plan usage, five minutes on usage
80 *   credits, an API key or a cloud provider.
81 */
82export function decideTtl(option: unknown, env: CacheEnv, setting?: unknown, account?: Account): TtlChoice {
83  const pinned = asTtl(option)
84  if (pinned) return { ttl: pinned, source: 'the ttl option' }
85  if (isOn(env.force5m)) return { ttl: '5m', source: 'FORCE_PROMPT_CACHING_5M' }
86  const fromVar = asTtl(env.ttlVar)
87  if (fromVar) return { ttl: fromVar, source: 'CLAUDE_CODE_PROMPT_CACHE_TTL' }
88  const fromSetting = asTtl(setting)
89  if (fromSetting) return { ttl: fromSetting, source: 'the promptCacheTtl setting' }
90  if (isOn(env.enable1h)) return { ttl: '1h', source: 'ENABLE_PROMPT_CACHING_1H' }
91  if (account === 'subscription') return { ttl: '1h', source: 'Claude subscription default' }
92  if (account === 'credits') return { ttl: '5m', source: 'usage credits default' }
93  return { ttl: '5m', source: 'default' }
94}
95
96export const resolveTtl = (option: unknown, env: CacheEnv, setting?: unknown, account?: Account): Ttl =>
97  decideTtl(option, env, setting, account).ttl
98
99/**
100 * The account, from the rate-limit windows the last response reported: a
101 * five-hour or seven-day window means a Claude subscription, and one that is
102 * full means the next requests draw on usage credits. No window (an API key,
103 * a cloud provider, or no response yet) says nothing.
104 */
105export function accountOf(windows: readonly { kind: string; percentUsed: number }[]): Account {
106  const plan = windows.filter(w => w.kind === 'five_hour' || w.kind === 'seven_day')
107  if (plan.length === 0) return 'other'
108  return plan.some(w => w.percentUsed >= 100) ? 'credits' : 'subscription'
109}
110
111export function ttlMs(ttl: Ttl): number {
112  return ttl === '1h' ? 3_600_000 : 300_000
113}
114
115/** Caching switched off for this model by the environment. */
116export function isCachingDisabled(model: string, env: CacheEnv): boolean {
117  if (isOn(env.disableAll)) return true
118  const name = model.toLowerCase()
119  if (name.includes('haiku')) return isOn(env.disableHaiku)
120  if (name.includes('sonnet')) return isOn(env.disableSonnet)
121  if (name.includes('opus')) return isOn(env.disableOpus)
122  return false
123}
124
125export const promptTokens = (s: Sample) => s.read + s.write + s.fresh
126
127/** Share of the prompt the cache served, 0 to 1; 0 for an empty prompt. */
128export function hitRatio(s: Sample): number {
129  const total = promptTokens(s)
130  return total === 0 ? 0 : s.read / total
131}
132
133/** When the cache entry the sample touched lapses, ms since the epoch. */
134export const expiresAt = (s: Sample, ttl: Ttl) => s.startedAt + ttlMs(ttl)
135
136/** Zero for a request that read and wrote nothing: it created or refreshed no entry, so there is nothing to count down. */
137export function remainingMs(s: Sample, ttl: Ttl, now: number): number {
138  if (s.read + s.write === 0) return 0
139  return Math.max(0, expiresAt(s, ttl) - now)
140}
141
142/**
143 * Why a request that should have read the cache wrote it instead; undefined
144 * when it did not miss. A prompt that shrank is a /compact or /clear, not a
145 * miss, and the first request of a session has nothing to read.
146 */
147export function missReason(prev: Sample | undefined, cur: Sample, ttl: Ttl): string | undefined {
148  if (!prev) return undefined
149  const before = promptTokens(prev)
150  if (before === 0 || promptTokens(cur) < before * 0.7) return undefined
151  if (cur.read >= before * 0.5 || cur.write === 0) return undefined
152  if (cur.model !== prev.model) return `model changed (${prev.model} to ${cur.model})`
153  if (cur.startedAt - prev.startedAt > ttlMs(ttl)) return `the ${ttl} cache had lapsed`
154  return 'the prompt prefix changed (effort, tools, system prompt or CLAUDE.md)'
155}
156
157export function advise(last: Sample | undefined, prev: Sample | undefined, policy: Policy, now: number, disabled: boolean): Advice {
158  if (disabled) return { kind: 'off', text: 'prompt caching is off for this model (DISABLE_PROMPT_CACHING*)' }
159  if (!last) return { kind: 'cold', text: 'no request yet: the first one writes the cache' }
160  if (last.read + last.write === 0) {
161    return { kind: 'uncached', text: 'this request was not cached (prompt under the model minimum, or caching off)' }
162  }
163  const miss = missReason(prev, last, policy.ttl)
164  const left = remainingMs(last, policy.ttl, now)
165  const size = promptTokens(last)
166  if (left <= 0) {
167    const big = size >= policy.compactAtTokens
168    return {
169      kind: 'expired',
170      text: big
171        ? `expired: the next message rewrites ${fmtTokens(size)} tokens. /compact first, or /clear if the task is done`
172        : `expired: only ${fmtTokens(size)} tokens to rebuild, just keep going`,
173    }
174  }
175  if (left <= policy.warnMs) {
176    return { kind: 'soon', text: 'expires soon: any message refreshes it for free' }
177  }
178  if (miss) return { kind: 'miss', text: `cache missed: ${miss}` }
179  return { kind: 'warm', text: 'warm: keep going' }
180}
181
182export function fmtTokens(n: number): string {
183  if (n < 1000) return String(n)
184  if (n < 100_000) return `${(n / 1000).toFixed(1).replace(/\.0$/, '')}k`
185  if (n < 1_000_000) return `${Math.round(n / 1000)}k`
186  return `${(n / 1_000_000).toFixed(1).replace(/\.0$/, '')}M`
187}
188
189/** m:ss, or h:mm:ss from an hour up. */
190export function fmtClock(ms: number): string {
191  const total = Math.max(0, Math.ceil(ms / 1000))
192  const h = Math.floor(total / 3600)
193  const m = Math.floor((total % 3600) / 60)
194  const s = total % 60
195  const pad = (n: number) => String(n).padStart(2, '0')
196  return h > 0 ? `${h}:${pad(m)}:${pad(s)}` : `${m}:${pad(s)}`
197}
198
199export function bar(ratio: number, width: number): string {
200  const filled = Math.round(Math.min(1, Math.max(0, ratio)) * width)
201  return '█'.repeat(filled) + '░'.repeat(width - filled)
202}
203
204export type TurnRow = {
205  turnId: string
206  steps: number
207  read: number
208  write: number
209  fresh: number
210  output: number
211}
212
213/** Samples grouped by turn, oldest first, each turn's requests summed. */
214export function byTurn(samples: readonly Sample[]): TurnRow[] {
215  const rows: TurnRow[] = []
216  for (const s of samples) {
217    let row = rows[rows.length - 1]
218    if (!row || row.turnId !== s.turnId) {
219      row = { turnId: s.turnId, steps: 0, read: 0, write: 0, fresh: 0, output: 0 }
220      rows.push(row)
221    }
222    row.steps += 1
223    row.read += s.read
224    row.write += s.write
225    row.fresh += s.fresh
226    row.output += s.output
227  }
228  return rows
229}
230
231export const rowRatio = (r: TurnRow) => {
232  const total = r.read + r.write + r.fresh
233  return total === 0 ? 0 : r.read / total
234}
235
236export function fit(text: string, width: number): string {
237  return text.length <= width ? text : `${text.slice(0, Math.max(0, width - 1))}…`
238}
239
240export function positive(v: unknown, fallback: number): number {
241  return typeof v === 'number' && Number.isFinite(v) && v > 0 ? v : fallback
242}
243
244/** Share of the cache lifetime left, 0 to 1. */
245export function lifeRatio(leftMs: number, ttl: Ttl): number {
246  return Math.min(1, Math.max(0, leftMs / ttlMs(ttl)))
247}
248
249/**
250 * Widths of the three stacked-bar segments (read, wrote, new) over `width`
251 * cells: proportional, each non-empty part at least one cell, summing to width.
252 */
253export function segments(read: number, write: number, fresh: number, width: number): [number, number, number] {
254  const total = read + write + fresh
255  if (total === 0 || width <= 0) return [0, 0, 0]
256  const parts = [read, write, fresh]
257  const cells = parts.map(p => (p > 0 ? Math.max(1, Math.round((p / total) * width)) : 0))
258  let over = cells.reduce((a, b) => a + b, 0) - width
259  while (over !== 0) {
260    const i = over > 0 ? cells.indexOf(Math.max(...cells)) : parts.indexOf(Math.max(...parts))
261    cells[i] += over > 0 ? -1 : 1
262    over += over > 0 ? -1 : 1
263  }
264  return [cells[0], cells[1], cells[2]]
265}
266
267/** Seconds left at which a toast counts down after the one at the warning threshold. */
268export const COUNTDOWN_MARKS = [10, 3, 2, 1]
269
270/**
271 * The toast mark to fire now, or undefined. `level` is the mark last fired for
272 * this cache entry (Infinity before any); a late tick skips straight to the
273 * newest mark crossed, so a stalled clock never replays old ones.
274 */
275export function nextToastMark(secsLeft: number, warnSecs: number, level: number): number | undefined {
276  const marks = [warnSecs, ...COUNTDOWN_MARKS].filter(m => m <= warnSecs)
277  const due = marks.filter(m => secsLeft <= m && m < level)
278  return due.length ? Math.min(...due) : undefined
279}
280
281export type LifeColor = 'green' | 'yellow' | 'red'
282
283/** Countdown colour: green while there is plenty, yellow below 40% of the lifetime, red from the warning threshold down. */
284export function lifeColor(leftMs: number, ttl: Ttl, warnMs: number): LifeColor {
285  if (leftMs <= warnMs) return 'red'
286  return leftMs / ttlMs(ttl) <= 0.4 ? 'yellow' : 'green'
287}
288
289// requests are timed from their start, so a little slack keeps a hit that
290// landed just inside the lifetime from reading as proof of the longer one
291const SLACK_MS = 10_000
292
293/**
294 * What the traffic says about the cache lifetime, given the request before and
295 * `known`, what earlier requests already showed.
296 *
297 *   - a hit (the cache served at least half of the previous prompt) more than
298 *     5 minutes after the previous request began proves the 1-hour lifetime,
299 *     and nothing later undoes it: a miss afterwards is more likely a changed
300 *     prefix than a lapse
301 *   - a miss with the same model and a prompt that did not shrink, 5 minutes to
302 *     an hour after the previous request, says the entry lapsed: 5 minutes
303 *     (weaker: a changed prefix looks the same, so a later hit overrules it)
304 *
305 * Needed because the API names the TTL of a write (`cache_creation.ephemeral_*`)
306 * but Claude Code's mod API passes on only the four token counts.
307 */
308export function observeTtl(prev: Sample | undefined, cur: Sample, known: Ttl | undefined): Ttl | undefined {
309  if (!prev || prev.read + prev.write === 0 || cur.model !== prev.model) return known
310  const gap = cur.startedAt - prev.startedAt
311  const before = promptTokens(prev)
312  if (gap <= ttlMs('5m') + SLACK_MS) return known
313  if (cur.read >= before * 0.5) return '1h'
314  if (known === '1h') return known
315  const lapsed = cur.write > 0 && promptTokens(cur) >= before * 0.7 && gap < ttlMs('1h') + SLACK_MS
316  return lapsed ? '5m' : known
317}
318
hooks/handoff.ts 140 lines
1// Pure helpers for the Clear button: turn a transcript into a digest a cheap
2// model can brief from, and build the text the fresh session receives.
3// No `$` calls here, so every function tests on its own.
4
5export type TranscriptRow = {
6  type?: string
7  isSidechain?: boolean
8  isMeta?: boolean
9  message?: { role?: string; content?: unknown }
10}
11
12export type Digest = {
13  /** Turns in order, "user: ..." / "assistant: ...", trimmed to the cap from the start. */
14  text: string
15  /** Files and commands the tools touched, most recent last, de-duplicated. */
16  touched: string[]
17  /** The person's prompts, verbatim, oldest first. */
18  prompts: string[]
19}
20
21export const DIGEST_CAP = 60_000
22const PROMPT_KEEP = 8
23const TOUCHED_KEEP = 20
24const COMMAND_PEEK = 80
25
26/** How much of a transcript's tail is parsed: a long session's file runs to many MB. */
27export const TRANSCRIPT_TAIL_CHARS = 3_000_000
28
29/** The last `cap` characters of a JSONL text, starting at a line boundary. */
30export function tailLines(text: string, cap = TRANSCRIPT_TAIL_CHARS): string {
31  if (text.length <= cap) return text
32  const cut = text.slice(text.length - cap)
33  const nl = cut.indexOf('\n')
34  return nl === -1 ? '' : cut.slice(nl + 1)
35}
36
37/** Transcript lines (JSONL) -> rows; a line that is not JSON is skipped. Only the tail is read. */
38export function parseRows(jsonl: string, cap = TRANSCRIPT_TAIL_CHARS): TranscriptRow[] {
39  const rows: TranscriptRow[] = []
40  for (const line of tailLines(jsonl, cap).split('\n')) {
41    if (!line.trim()) continue
42    try {
43      rows.push(JSON.parse(line) as TranscriptRow)
44    } catch {
45      // a half-written last line, or a non-row: not ours to repair
46    }
47  }
48  return rows
49}
50
51type Block = { type?: string; text?: string; name?: string; input?: Record<string, unknown> }
52
53/** The conversation as text: user and assistant text blocks, tool calls as one line each. */
54export function digestTranscript(rows: readonly TranscriptRow[], cap = DIGEST_CAP): Digest {
55  const lines: string[] = []
56  const prompts: string[] = []
57  const touched: string[] = []
58  for (const row of rows) {
59    if ((row.type !== 'user' && row.type !== 'assistant') || row.isSidechain || row.isMeta) continue
60    const role = row.message?.role ?? row.type
61    const content = row.message?.content
62    if (typeof content === 'string') {
63      if (role === 'user') prompts.push(content)
64      lines.push(`${role}: ${content}`)
65      continue
66    }
67    if (!Array.isArray(content)) continue
68    for (const block of content as Block[]) {
69      if (block.type === 'text' && block.text) {
70        if (role === 'user') prompts.push(block.text)
71        lines.push(`${role}: ${block.text}`)
72      } else if (block.type === 'tool_use') {
73        const target = toolTarget(block)
74        if (target) touched.push(target)
75        lines.push(`assistant used ${block.name ?? 'a tool'}${target ? `: ${target}` : ''}`)
76      }
77      // tool_result bodies are left out: large, and the assistant's text already reflects them
78    }
79  }
80  const joined = lines.join('\n')
81  const text = joined.length > cap ? `…\n${joined.slice(joined.length - cap)}` : joined
82  return {
83    text,
84    touched: dedupe(touched).slice(-TOUCHED_KEEP),
85    prompts: prompts.slice(-PROMPT_KEEP),
86  }
87}
88
89function toolTarget(block: Block): string | undefined {
90  const input = block.input ?? {}
91  const path = input.file_path ?? input.path ?? input.notebook_path
92  if (typeof path === 'string') return path
93  const command = input.command
94  if (typeof command === 'string') return command.length > COMMAND_PEEK ? `${command.slice(0, COMMAND_PEEK)}…` : command
95  return undefined
96}
97
98function dedupe(items: readonly string[]): string[] {
99  const seen = new Set<string>()
100  const out: string[] = []
101  for (const item of [...items].reverse()) {
102    if (seen.has(item)) continue
103    seen.add(item)
104    out.push(item)
105  }
106  return out.reverse()
107}
108
109export const HANDOFF_SYSTEM = `You write a handoff brief so a fresh Claude Code session can continue someone's work after the previous conversation was cleared.
110Write in plain Markdown with exactly these headings, each followed by short bullets or one line; say "none" where there is nothing:
111## Goal
112## Current state
113## Decisions made
114## Files and commands touched
115## Open items
116## Next step
117Be concrete: names of files, functions, flags, commands, errors. No preamble, no closing remarks, under 400 words.`
118
119export function handoffPrompt(digest: Digest): string {
120  const touched = digest.touched.length ? digest.touched.map(t => `- ${t}`).join('\n') : '- none recorded'
121  return `Tools touched (most recent last):\n${touched}\n\nConversation (oldest first, truncated from the start if long):\n${digest.text}`
122}
123
124/** When the model gave no brief: the person's last prompts, verbatim, and how to go back. */
125export function fallbackBrief(digest: Digest, previousSessionId: string): string {
126  const prompts = digest.prompts.length ? digest.prompts.map(p => `- ${firstLine(p)}`).join('\n') : '- none recorded'
127  const touched = digest.touched.length ? digest.touched.map(t => `- ${t}`).join('\n') : '- none recorded'
128  return `## Goal\nNot summarised (the brief could not be generated). The last prompts were:\n${prompts}\n\n## Files and commands touched\n${touched}\n\n## Next step\nAsk me what to continue with, or run \`claude --resume ${previousSessionId}\` to reopen the full conversation.`
129}
130
131/** The text the fresh session is sent: the brief, framed so the first turn stays short. */
132export function wrapForNewSession(brief: string, previousSessionId: string): string {
133  return `Handoff brief from the previous session (cleared to reset the prompt cache; its id was ${previousSessionId}, reopen it with \`claude --resume ${previousSessionId}\` if you need the full history).\n\nRead it, then acknowledge in one line and wait for my next instruction. Do not start any work yet.\n\n${brief.trim()}`
134}
135
136function firstLine(text: string): string {
137  const line = (text.split('\n')[0] ?? '').trim()
138  return line.length > 160 ? `${line.slice(0, 160)}…` : line
139}
140