When an idle session's prompt cache is about to expire, summarizes it while the cache is still warm; when you come back, pick A (keep the full session, pay the…

A Claude Code mod for when you walk away from a long session. Just before the idle prompt cache expires, it summarizes the conversation while the cache is still warm. When you come back to a cold cache, a menu asks how to continue:

Pick with the arrow keys and Enter, or press the option's number.
If you close the menu with Esc, the same choice stays above the prompt:

There, press 1 or 2, or type a or b and press Enter. Anything else you type before choosing is held, shown under the choices, and sent once you pick.
You need Claude Code 2.1.287 or newer.
git clone https://github.com/cjavdev/mods.git ~/mods
claude --plugin-dir ~/mods/cache-saver
Or add ~/mods/cache-saver to CLAUDE_CODE_PLUGIN_DIRS in the env block of ~/.claude/settings.json. It works well alongside cache-shot-clock.
The summary is a $.model.fork of the main thread. It's sent while the cache is warm, so its prefix is a cache read (0.1× base input), plus the summary's output tokens. One fork per idle stretch. Sessions smaller than minTokens are skipped, because re-caching them is cheap anyway.
The fork's cache read also refreshes the cache entry, so the cache stays warm for one more TTL after the summary is built. The choice only appears once it is actually cold. If you come back before then, just keep working: the summary is dropped when the next response arrives.
/compactThe engine skips a plugin's own hooks on any compaction that plugin starts directly. A compaction cache-saver started that way would go to Claude Code's summarizer, which reads the whole cold transcript. That is the cost this mod exists to avoid. So when you pick the summary, cache-saver runs /compact as a queued command, the way you would by typing it. That compaction reaches cache-saver's session.compact hook, which answers with the prepared summary, and Claude Code's summarizer never runs.
A /compact that you type yourself, or one with instructions (/compact focus on X), still goes to Claude Code as usual.
| Option | Default | |
|---|---|---|
enabled | true | Turns the mod off without unloading it. |
ttl | auto | auto reads the TTL from the transcript; 5m or 1h forces it. |
leadSeconds | 45 | How long before the idle cache expires to build the summary. |
minTokens | 30000 | Sessions smaller than this get no summary. |
testTtlSeconds | 0 | For testing: treat the cache as expiring after this many seconds. 0 is off. |
With a 1-hour cache, the choice shows up two hours after you go idle. testTtlSeconds runs the whole cycle on a short clock instead. The real cache stays warm, so the summary is still cheap:
claude --settings '{"pluginConfigs":{"cache-saver":{"options":{"testTtlSeconds":40,"leadSeconds":20}},"cache-shot-clock":{"options":{"testTtlSeconds":40}}}}'
Send one prompt and wait. The summary is built about 20 seconds after the reply, and the choice appears about 45 seconds after that. The session needs at least minTokens of context, which a fresh session with a few connectors already has.
claude plugin validate ~/mods/cache-saver
claude plugin test ~/mods/cache-saver
The tests drive a whole idle stretch on a mocked clock. They cover the summary built 45 seconds before expiry, a single fork per stretch, the choice appearing only once the cache is cold, a typed prompt held and then sent, the choice typed as a or b, /compact answered with the summary without the engine's summarizer, and the short test TTL.
hooks/register.tsx 411 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { Cache, Offer, Ttl } from '../types'
5import { TTL_MS, compact, fmtAgo, fmtClock, prefixOf, ttlFromTranscript } from './clock'
6
7const cache = atom({ plugin: 'cache-saver', key: 'cache' } as const, null)
8const offer = atom({ plugin: 'cache-saver', key: 'offer' } as const, null)
9const now = atom({ plugin: 'cache-saver', key: 'now' } as const, 0)
10
11const PREVIEW = 'cache-saver-summary'
12const TAIL_BYTES = '262144'
13// The fork has to start before the entry lapses to read it; past this little
14// time left it would pay the full write it is meant to save.
15const MIN_LEFT_MS = 3000
16// Where a prompt comes from the person (the terminal, the desktop app, Remote
17// Control), as opposed to a notification, a peer or a plugin.
18const PERSON = new Set(['composer', 'sdk', 'bridge'])
19
20const SUMMARY_PROMPT = `The prompt cache for this conversation is about to expire. Write a summary that lets you continue this work in a fresh context with nothing else to go on. Do not call tools and do not continue the task; reply with the summary alone, in Markdown, under these headings:
21
221. Primary request and intent: everything the user has asked for, in detail.
232. Key technical concepts: technologies, frameworks and decisions in play.
243. Files and code: every file examined, changed or created, why it matters, and the important snippets verbatim.
254. Errors and fixes: what went wrong and how it was resolved, including user feedback.
265. Problem solving: what is solved and what is still being worked out.
276. All user messages: every non-tool-result user message, verbatim or nearly so.
287. Pending tasks: what was asked for and is not done.
298. Current work: precisely what was happening just before this summary.
309. Next step: the next step in line with the user's latest request, quoting it where it helps.
31
32Leave this request for a summary out of it: it is not part of the work.`
33
34const wrapSummary = (summary: string) =>
35 `This session is being continued from an earlier conversation that was summarized to avoid re-writing its expired prompt cache. The summary below covers everything before this point.\n\n${summary}`
36
37type Usage = Parameters<typeof prefixOf>[0]
38
39const mem = {
40 forced: null as Ttl | null,
41 ttl: '5m' as Ttl,
42 leadMs: 45_000,
43 // For testing: treat the cache as expiring after this long; 0 is off.
44 testMs: 0,
45 minTokens: 30_000,
46 isWorking: false,
47 // The step a summary was already tried for: one try per idle stretch.
48 triedStep: -1,
49 // The step the menu was already raised for: once per cold cache.
50 askedStep: -1,
51}
52
53const cacheLeft = (c: Cache, at: number) => c.lastAt + (mem.testMs || TTL_MS[c.ttl]) - at
54
55async function learnTtl($: EngineInterface, ttl: Ttl) {
56 if (mem.forced) return
57 mem.ttl = ttl
58 await update($, cache, prev => (prev && prev.ttl !== ttl ? { ...prev, ttl } : prev))
59}
60
61// A main-thread response: the cache is fresh, and any summary is stale.
62async function touch($: EngineInterface, usage: Usage) {
63 const at = await $.clock.now()
64 const prev = await read($, cache)
65 if (prev && mem.ttl === '5m' && usage.cache_read_input_tokens > 0 && at - prev.lastAt > TTL_MS['5m'] + 15_000) {
66 await learnTtl($, '1h')
67 }
68 const next: Cache = { lastAt: at, prefixTokens: prefixOf(usage), ttl: mem.ttl, step: (prev?.step ?? 0) + 1 }
69 await update($, cache, () => next)
70 await update($, offer, o => (o && o.status !== 'compacting' && o.held === null ? null : o))
71}
72
73async function forget($: EngineInterface) {
74 await update($, cache, () => null)
75 await update($, offer, () => null)
76}
77
78// Summarize over the still-warm prefix: a fork reads it from the cache at a
79// tenth of the input price, where the same tokens cold cost a full write.
80async function prepare($: EngineInterface, c: Cache) {
81 mem.triedStep = c.step
82 await update($, offer, () => ({
83 status: 'preparing',
84 step: c.step,
85 summary: '',
86 summaryTokens: 0,
87 fullTokens: c.prefixTokens,
88 held: null,
89 }))
90
91 const reply = await $.model.fork({ prompt: SUMMARY_PROMPT }).catch(
92 (err: unknown) => ({ isAnswered: false as const, reason: String(err) }),
93 )
94 if (!reply.isAnswered || reply.text.trim() === '') {
95 $.ui.log(`cache-saver: no summary (${reply.isAnswered ? 'empty reply' : reply.reason})`, { to: 'debug' })
96 await update($, offer, o => (o?.step === c.step && o.status === 'preparing' ? null : o))
97 return
98 }
99
100 // The fork's read refreshed the entry: it lives a full TTL from now.
101 if (reply.usage.cache_read_input_tokens > 0) {
102 const at = await $.clock.now()
103 await update($, cache, prev => (prev && prev.step === c.step ? { ...prev, lastAt: at } : prev))
104 }
105 await update($, offer, o =>
106 o?.step === c.step && o.status === 'preparing'
107 ? { ...o, status: 'ready', summary: reply.text.trim(), summaryTokens: reply.usage.output_tokens }
108 : o,
109 )
110}
111
112// A: keep the full conversation; the next turn re-writes the cache.
113async function keepFull($: EngineInterface) {
114 const o = await read($, offer)
115 if (!o || (o.status !== 'ready' && o.status !== 'armed')) return
116 if (o.status === 'armed') await disarm($)
117 await update($, offer, () => null)
118 void $.ui.close({ id: PREVIEW }).catch(() => undefined)
119 if (o.held) await $.prompt.submit({ text: o.held })
120}
121
122// B: swap the conversation for the summary. The engine skips a plugin's own
123// hooks on a compaction its hook started directly, so the mod runs `/compact`
124// the way the person would: queued as a command, it reaches the
125// session.compact hook below, which answers with the summary built earlier.
126async function useSummary($: EngineInterface) {
127 const o = await read($, offer)
128 if (!o || o.status !== 'ready') return
129 void $.ui.close({ id: PREVIEW }).catch(() => undefined)
130 await update($, offer, () => ({ ...o, status: 'armed' }))
131
132 const hasRun = await $.command.run({ command: 'compact' }).then(
133 () => true,
134 () => false,
135 )
136 if (hasRun) {
137 // Still armed: the compaction went past our hook, and the session is
138 // compacted all the same.
139 const after = await read($, offer)
140 if (after?.status === 'armed') await forget($)
141 return
142 }
143
144 // The command could not be run for the person: leave it for their Enter.
145 const filled = await $.prompt.fill({ text: '/compact' })
146 if (!filled.isFilled) $.ui.toast('cache-saver: run /compact to continue from the summary')
147}
148
149// Back out of B: the box is the person's again.
150async function disarm($: EngineInterface) {
151 await update($, offer, o => (o?.status === 'armed' ? { ...o, status: 'ready' } : o))
152 const box = await $.prompt.read()
153 if (box.text.trim() === '/compact') await $.prompt.fill({ text: '' })
154}
155
156// The choice as the engine's own question dialog: it takes the keyboard by
157// itself (arrows and Enter, or the option's number). Dismissed with Esc, the
158// band's buttons stay up as the way back to it.
159async function ask($: EngineInterface, o: Offer) {
160 const saved = Math.round((1 - o.summaryTokens / Math.max(1, o.fullTokens)) * 100)
161 const keep = `Keep full session: re-cache ${compact(o.fullTokens)} tokens`
162 const summary = `Continue from summary: ${compact(o.summaryTokens)} tokens (${saved}% smaller)`
163 const look = 'Preview the summary first'
164
165 const answer = await $.ui
166 .ask('Your prompt cache went cold. How do you want to continue?', {
167 header: 'Cache cold',
168 options: [keep, summary, look],
169 })
170 .catch(() => null)
171
172 if (answer === keep) await keepFull($)
173 else if (answer === summary) await useSummary($)
174 else if (answer === look) await preview($)
175 // Typed under "Other": it is the next prompt, sent once they choose.
176 else if (answer) await update($, offer, prev => (prev ? { ...prev, held: prev.held ? `${prev.held}\n\n${answer}` : answer } : prev))
177}
178
179async function preview($: EngineInterface) {
180 await $.ui.open({ id: PREVIEW, title: 'Session summary (option B)' })
181}
182
183export const register: Register = (on, options) => {
184 // Off in /config: hook nothing at all.
185 if (options.enabled === false) return
186
187 mem.forced = options.ttl === '5m' || options.ttl === '1h' ? options.ttl : null
188 mem.ttl = mem.forced ?? '5m'
189 mem.leadMs = Math.max(5, Number(options.leadSeconds ?? 45)) * 1000
190 mem.minTokens = Math.max(0, Number(options.minTokens ?? 30_000))
191 mem.testMs = Math.max(0, Number(options.testTtlSeconds ?? 0)) * 1000
192
193 on('session.start', async ($, e, next) => {
194 const result = await next(e)
195 const c = await read($, cache)
196 if (c && !mem.forced) mem.ttl = c.ttl
197 // A summary left half-built by a reload will never finish.
198 await update($, offer, o => (o?.status === 'ready' || o?.status === 'armed' ? { ...o, status: 'ready' } : null))
199
200 $.clock.every(1000, async () => {
201 const at = await $.clock.now()
202 await update($, now, () => at)
203
204 const c = await read($, cache)
205 if (!c || mem.isWorking) return
206 const left = cacheLeft(c, at)
207
208 // Cold with a summary ready: raise the menu, once.
209 const o = await read($, offer)
210 if (left <= 0 && o?.status === 'ready' && o.step === c.step && mem.askedStep !== c.step) {
211 mem.askedStep = c.step
212 void ask($, o)
213 return
214 }
215
216 if (c.prefixTokens < mem.minTokens || mem.triedStep === c.step) return
217 if (left <= mem.leadMs && left > MIN_LEFT_MS) void prepare($, c)
218 })
219 return result
220 })
221
222 on('turn.start', async ($, e, next) => {
223 mem.isWorking = true
224 return next(e)
225 })
226
227 on('turn.complete', async ($, e, next) => {
228 if (e.agentId === undefined) mem.isWorking = false
229 return next(e)
230 })
231
232 on('turn.step', async function* ($, e, next) {
233 const result = yield* next(e)
234 if (e.agentId === undefined && result.usage) await touch($, result.usage)
235 return result
236 })
237
238 on('classic.Stop', async ($, e, next) => {
239 const result = await next(e)
240 if (!mem.forced) {
241 const tail = await $.process
242 .run(['tail', '-c', TAIL_BYTES, e.transcript_path], { timeoutMs: 5000 })
243 .catch(() => null)
244 const found = tail?.exitCode === 0 ? ttlFromTranscript(tail.stdout) : null
245 if (found) await learnTtl($, found)
246 }
247 return result
248 })
249
250 on('classic.PostModelSwitch', async ($, e, next) => {
251 const result = await next(e)
252 await forget($)
253 if (!mem.forced) mem.ttl = e.cache_ttl
254 return result
255 })
256
257 // B armed: the person's /compact gets the summary built over the warm
258 // cache, so no summarizer request re-reads the cold transcript. Any other
259 // compaction leaves a new prefix with nothing cached for it.
260 on('session.compact', async ($, e, next) => {
261 if (e.agentId !== undefined || e.trigger === 'precompute') return next(e)
262 const o = await read($, offer)
263 if (o?.status === 'armed' && (e.trigger === 'manual' || e.trigger === 'plugin') && !e.instructions) {
264 await forget($)
265 $.ui.toast(`Continuing from the summary: ${compact(o.summaryTokens)} tokens instead of ${compact(o.fullTokens)}`)
266 if (o.held) {
267 const held = o.held
268 $.clock.after(0, () => void $.prompt.submit({ text: held }))
269 }
270 return {
271 messages: [{ role: 'user', text: wrapSummary(o.summary), toolUses: [] }],
272 tokensBefore: o.fullTokens,
273 tokensAfter: o.summaryTokens,
274 }
275 }
276 const result = await next(e)
277 if (result.messages) await forget($)
278 return result
279 })
280
281 on('session.end', async ($, e, next) => {
282 if (e.reason === 'clear') await forget($)
283 return next(e)
284 })
285
286 // Typed before choosing: hold it, and send it once A or B is picked.
287 on('prompt.submit', async ($, e, next) => {
288 const o = await read($, offer)
289 const c = await read($, cache)
290 const isCold = c !== null && cacheLeft(c, await $.clock.now()) <= 0
291 const isChoosing = o !== null && (o.status === 'ready' || o.status === 'armed')
292 if (!o || !isChoosing || !isCold || !PERSON.has(e.origin.kind) || e.text.trimStart().startsWith('/')) return next(e)
293
294 // The choice typed instead of pressed: a or 1, b or 2. Acted on from a
295 // timer, since a command cannot be run from inside the hook a prompt waits on.
296 const typed = e.text.trim().toLowerCase()
297 if (o.status === 'ready' && (typed === 'a' || typed === '1')) {
298 $.clock.after(0, () => void keepFull($))
299 return { drop: 'cache-saver: keeping the full session' }
300 }
301 if (o.status === 'ready' && (typed === 'b' || typed === '2')) {
302 $.clock.after(0, () => void useSummary($))
303 return { drop: 'cache-saver: continuing from the summary' }
304 }
305
306 const held = o.held ? `${o.held}\n\n${e.text}` : e.text
307 await update($, offer, prev => (prev ? { ...prev, held } : prev))
308 return {
309 drop:
310 o.status === 'armed'
311 ? 'cache-saver: held until /compact runs; it sends after'
312 : 'cache-saver: held until you pick A (keep full session) or B (continue from summary) above',
313 }
314 })
315
316 on('ui.render', { component: 'Pane', requestId: PREVIEW }, async ($, e) => {
317 const { Box, Markdown, Text } = $.ui.resolve(e)
318 const o = await read($, offer)
319 if (!o || o.status !== 'ready') {
320 return (
321 <Box>
322 <Text dimColor>No summary waiting.</Text>
323 </Box>
324 )
325 }
326 return (
327 <Box flexDirection="column">
328 <Text dimColor>{`${compact(o.summaryTokens)} tokens, in place of ${compact(o.fullTokens)}`}</Text>
329 <Markdown text={o.summary.slice(0, 10_000)} />
330 </Box>
331 )
332 })
333
334 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
335 const below = await next(e)
336 const o = await read($, offer)
337 const c = await read($, cache)
338 if (e.props.hasSurvey || !o || !c) return below
339
340 const { Box, Button, Text } = $.ui.resolve(e)
341 const at = (await read($, now)) ?? Date.now()
342 const left = cacheLeft(c, at)
343 let mine
344
345 if (o.status === 'preparing') {
346 mine = (
347 <Box key="cache-saver">
348 <Text dimColor>{`cache-saver: summarizing ${compact(o.fullTokens)} tokens while the cache is warm…`}</Text>
349 </Box>
350 )
351 } else if (o.status === 'compacting') {
352 mine = (
353 <Box key="cache-saver">
354 <Text dimColor>cache-saver: switching to the summary…</Text>
355 </Box>
356 )
357 } else if (o.status === 'armed') {
358 mine = (
359 <Box key="cache-saver" flexDirection="column">
360 <Text>
361 <Text bold>Press Enter</Text>
362 <Text dimColor>{` to run /compact and continue from the ${compact(o.summaryTokens)}-token summary`}</Text>
363 </Text>
364 <Box>
365 <Button key="keep-full" hotkey="1" plain dimColor label="Keep full session instead" onPress={() => keepFull($)} />
366 </Box>
367 </Box>
368 )
369 } else if (left > 0) {
370 mine = (
371 <Box key="cache-saver">
372 <Text dimColor>{`cache-saver: summary ready (${compact(o.summaryTokens)} tokens), offered if the cache goes cold in ${fmtClock(left)}`}</Text>
373 </Box>
374 )
375 } else if (!e.props.isWorking) {
376 const saved = Math.round((1 - o.summaryTokens / Math.max(1, o.fullTokens)) * 100)
377 mine = (
378 <Box key="cache-saver" flexDirection="column">
379 <Text>
380 <Text bold>Prompt cache went cold </Text>
381 <Text dimColor>{`${fmtAgo(-left)} ago. How do you want to continue?`}</Text>
382 </Text>
383 <Box>
384 <Button key="keep-full" hotkey="1" variant="secondary" label={`A · Keep full session`} onPress={() => keepFull($)} />
385 <Text dimColor>{` re-cache ${compact(o.fullTokens)} tokens`}</Text>
386 </Box>
387 <Box>
388 <Button key="use-summary" hotkey="2" variant="primary" label={`B · Continue from summary`} onPress={() => useSummary($)} />
389 <Text dimColor>{` ${compact(o.summaryTokens)} tokens (${saved}% smaller)`}</Text>
390 </Box>
391 <Box>
392 <Button key="preview" hotkey="3" plain dimColor label="Preview summary" onPress={() => preview($)} />
393 <Text dimColor>{' · type A or B and press Enter, or press 1 or 2'}</Text>
394 </Box>
395 {o.held !== null && <Text dimColor>{`Held: “${o.held.slice(0, 60)}${o.held.length > 60 ? '…' : ''}” sends after you choose`}</Text>}
396 </Box>
397 )
398 }
399
400 if (!mine) return below
401 return below ? (
402 <Box flexDirection="column">
403 {below}
404 {mine}
405 </Box>
406 ) : (
407 mine
408 )
409 })
410}
411hooks/clock.ts 58 lines1// Pure helpers: no `$`, so tests can call them directly.
2
3import type { Ttl } from '../types'
4
5export const TTL_MS: Record<Ttl, number> = { '5m': 5 * 60_000, '1h': 60 * 60_000 }
6
7type Usage = {
8 input_tokens: number
9 output_tokens: number
10 cache_read_input_tokens: number
11 cache_creation_input_tokens: number
12}
13
14// What the next request re-sends: everything this one was answered over plus
15// what it generated. On a cold cache all of it is written afresh.
16export const prefixOf = (u: Usage): number =>
17 u.input_tokens + u.cache_read_input_tokens + u.cache_creation_input_tokens + u.output_tokens
18
19// The TTL the newest cache write used, from the tail of a session transcript
20// (JSONL). Each assistant row's usage carries
21// "cache_creation":{"ephemeral_1h_input_tokens":N,"ephemeral_5m_input_tokens":M};
22// a row that wrote nothing says nothing, so the last row that wrote decides.
23export const ttlFromTranscript = (tail: string): Ttl | null => {
24 let found: Ttl | null = null
25 for (const m of tail.matchAll(/"cache_creation":\{([^}]*)\}/g)) {
26 const body = m[1] ?? ''
27 const oneHour = Number(/"ephemeral_1h_input_tokens":(\d+)/.exec(body)?.[1] ?? 0)
28 const fiveMin = Number(/"ephemeral_5m_input_tokens":(\d+)/.exec(body)?.[1] ?? 0)
29 if (oneHour > 0) found = '1h'
30 else if (fiveMin > 0) found = '5m'
31 }
32 return found
33}
34
35// m:ss (a full hour reads 60:00), or h:mm:ss past it.
36export const fmtClock = (ms: number): string => {
37 const total = Math.max(0, Math.ceil(ms / 1000))
38 const h = total > 3600 ? Math.floor(total / 3600) : 0
39 const m = Math.floor((total - h * 3600) / 60)
40 const s = String(total % 60).padStart(2, '0')
41 return h > 0 ? `${h}:${String(m).padStart(2, '0')}:${s}` : `${m}:${s}`
42}
43
44// "45s", "3m", "1h 5m": how long the cache has been cold.
45export const fmtAgo = (ms: number): string => {
46 const s = Math.max(0, Math.floor(ms / 1000))
47 if (s < 60) return `${s}s`
48 const m = Math.floor(s / 60)
49 return m < 60 ? `${m}m` : `${Math.floor(m / 60)}h ${m % 60}m`
50}
51
52export const compact = (tokens: number): string =>
53 tokens >= 1_000_000
54 ? `${+(tokens / 1_000_000).toFixed(1)}M`
55 : tokens >= 1000
56 ? `${Math.round(tokens / 1000)}k`
57 : `${tokens}`
58types/index.d.ts 32 lines1export type Ttl = '5m' | '1h'
2
3// The main thread's cache as last seen.
4export type Cache = {
5 // When the main prefix was last read or written (ms since epoch).
6 lastAt: number
7 // Tokens the next request re-sends: what a cold cache writes afresh.
8 prefixTokens: number
9 ttl: Ttl
10 // Counts main-thread responses, so one idle stretch gets one summary.
11 step: number
12}
13
14// The A/B choice: built while the cache is warm, offered once it is cold.
15export type Offer = {
16 // armed: B was picked and `/compact` waits in the prompt box for Enter.
17 status: 'preparing' | 'ready' | 'armed' | 'compacting'
18 // The `Cache.step` it summarizes; a newer response makes it stale.
19 step: number
20 summary: string
21 summaryTokens: number
22 fullTokens: number
23 // A prompt typed before choosing, sent once the choice is made.
24 held: string | null
25}
26
27declare module 'claude-code' {
28 interface PluginState {
29 'cache-saver': { cache: Cache | null; offer: Offer | null; now: number }
30 }
31}
32