Compacts an idle session just before its prompt cache goes cold, or keeps the cache warm while you're away, so your next prompt doesn't re-read the whole…

Compacts an idle session just before its prompt cache goes cold, or keeps the cache warm while you're away.
It is separate from Claude Code's own auto-compact, which compacts when the context window fills up. warm-compact acts on time instead: it compacts when you've been away long enough for the cache to lapse, even when the window is far from full.
Claude Code caches the conversation for 5 minutes or an hour after each request, depending on your account. If you come back after that, your next prompt re-reads the whole conversation uncached, which is the most expensive way to read it. warm-compact compacts the session a minute before the cache lapses, while the compaction itself can still read the conversation from the cache. Your next prompt then starts from a short summary.
/model switch there's nothing left to keep warm.For the 30 seconds before it compacts, the status line under the prompt counts down: ⇊ compacting in 23s, before the cache goes cold · type to hold.
When you come back, a line in the transcript tells you what happened, for example Compacted at 14:32 (120k → 4k tokens), just before the prompt cache went cold. That line is for you only: it isn't sent to the model. If the compaction found the cache already gone, the line says so, because then it saved nothing.
Whether your account caches for 5 minutes or an hour is learned from the requests themselves and remembered between sessions. Until a request shows otherwise, it assumes an hour, as a Claude subscription has.
Headless runs (claude -p) are left alone.
Two chips sit side by side on one row under the prompt, below the mode line, a row they share with the chips of any other mod that puts its own there:
⏸ manual mode on · ? for shortcuts
on Warm compact off Keep warm
Click Warm compact to turn it off for this session, and again to turn it back on. Turning it off during the countdown cancels it.
The same from the prompt:
/warm-compact flips it
/warm-compact off turns it off
/warm-compact on turns it on
/warm-compact stats shows what it saved
A new session, or a /clear, starts with it on.
A compaction keeps your session cheap, but it trades the conversation for a summary. Keep warm keeps the whole conversation instead: a minute before the cache lapses, it renews it with one tiny request that re-reads the conversation from the cache. That request costs about a tenth of the conversation's size in tokens (a twentieth on Opus 5.5) and starts the cache's clock again. Coming back to a cold cache would cost twice its size.
It's off by default, because it spends tokens while you're away. Turn it on with the Keep warm chip, or from the prompt:
/keep-warm flips it
/keep-warm on turns it on
/keep-warm off turns it off
keepWarmRenewals option changes how many times.Kept the prompt cache warm at 14:32 (1 of 2)., is for you. Text left in the prompt box doesn't hold it off.Renewals count in /warm-compact stats too: each costs its request, and the last one before you come back is credited with the re-read it spared.
/warm-compact stats adds up what the compactions saved, for this session, the last 7 days and all time:
compactions renewals came back net saved
This session 1 0 1 +222k
Last 7 days 5 3 6 +1.4M
All time 9 5 11 +2.6M
Savings are counted in input tokens, the unit the API prices everything against: writing a token to an hour-long cache costs 2 of them (1.25 for a 5-minute cache), reading one from the cache 0.1 (0.05 on Opus 5.5, 0.025 on Fable 5.1), and generating one 5.
On a Claude subscription you don't pay per token, but the same tokens count toward your usage limits.
In /config, or in ~/.claude/settings.json under pluginConfigs:
| Option | Default | What it does |
|---|---|---|
minTokens | 50000 | The smallest conversation it compacts, in tokens |
leadSeconds | 60 | How long before the cache lapses a compaction or a renewal starts (15 at least) |
keepWarmRenewals | 2 | With keep warm on, how many times the cache is renewed in one pause before warm compact takes over |
claude plugin marketplace add ziedgithub/claude-code-mods && claude plugin install warm-compact@zied-mods
Then restart Claude Code.
See the repository README for running it from a clone.
hooks/register.tsx 388 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { CacheTtl, CompactionRecord, CompactionUsage } from '../types'
5import { CHIP_ROW_KEY, splitChipRow } from './chips'
6import { DEFAULT_TTL, TTL_MS, hitPercent, inferRenewTtl, inferTtl, isCacheTtl, plan, warnText } from './plan'
7import { MAX_RECORDS, clockTime, compactionHit, noticeText, recordsOf, statsText } from './savings'
8
9const TICK_MS = 1000
10const CACHE_TTL_KEY = 'cacheTtl'
11const RECORDS_KEY = 'compactions'
12const RENEW_TTL_KEY = 'renewTtl'
13// The models a renewal could not read the conversation from the cache on, each with the
14// engine version that failed: tried again once the engine changes
15const RENEW_MISSES_KEY = 'renewMisses'
16// The fork that renews the cache asks for as little as it can
17const RENEW_PROMPT = 'Reply with the single word: ok'
18// A renewal that read less of its request from the cache than this missed the entry
19const RENEW_HIT_FROM = 90
20const DEFAULT_MIN_TOKENS = 50_000
21const DEFAULT_LEAD_SECONDS = 60
22const DEFAULT_MAX_RENEWALS = 2
23// Never so close to the lapse that the compaction's request could miss the cache
24const MIN_LEAD_MS = 15_000
25const DONE_TEXT = 'Compacted before the prompt cache went cold'
26const COMMAND = 'warm-compact'
27const KEEP_WARM_COMMAND = 'keep-warm'
28const COMPACT_LABEL = 'Warm compact'
29const KEEP_WARM_LABEL = 'Keep warm'
30const ON_COLOR = '#238636'
31const OFF_COLOR = '#6e7681'
32
33const last = atom({ plugin: 'warm-compact', key: 'last' } as const, null)
34const isTurnRunning = atom({ plugin: 'warm-compact', key: 'isTurnRunning' } as const, false)
35const isEnabled = atom({ plugin: 'warm-compact', key: 'isEnabled' } as const, true)
36const isKeepWarm = atom({ plugin: 'warm-compact', key: 'isKeepWarm' } as const, false)
37const kept = atom({ plugin: 'warm-compact', key: 'kept' } as const, null)
38const pendingRecord = atom({ plugin: 'warm-compact', key: 'pendingRecord' } as const, null)
39
40let cacheTtl: CacheTtl = DEFAULT_TTL
41let shownStatus: string | undefined
42let isCompacting = false
43// The renewal or compaction last tried, named by the request and the renewals since: one
44// try each, whatever came of it
45let triedFor: string | null = null
46let maxRenewals = DEFAULT_MAX_RENEWALS
47// Until a read after a renewal shows otherwise, it holds the entry as long as a request does
48let renewTtl: CacheTtl | null = null
49let renewMisses: Record<string, string> = {}
50let engineVersion = ''
51// Each reason keep warm stood down is said once a session
52const saidOnce = new Set<string>()
53let minTokens = DEFAULT_MIN_TOKENS
54let leadMs = DEFAULT_LEAD_SECONDS * 1000
55
56const numberOr = (value: unknown, fallback: number): number =>
57 typeof value === 'number' && Number.isFinite(value) && value >= 0 ? value : fallback
58
59const showStatus = ($: EngineInterface, text: string | undefined): void => {
60 if (text === shownStatus) return
61 shownStatus = text
62 $.ui.status(text)
63}
64
65// Passes the response on as it streams, seeing each chunk on the way
66async function* tap<T, R>(stream: AsyncGenerator<T, R>, see: (chunk: T) => Promise<void>): AsyncGenerator<T, R> {
67 for (;;) {
68 const step = await stream.next()
69 if (step.done === true) return step.value
70 await see(step.value)
71 yield step.value
72 }
73}
74
75// The engine keeps its TTL to itself: a request after a long enough pause tells it, and it
76// is kept between sessions since it holds for the account
77const learnTtl = async ($: EngineInterface, ttl: CacheTtl | null): Promise<void> => {
78 if (ttl === null || ttl === cacheTtl) return
79 cacheTtl = ttl
80 await $.store.set(CACHE_TTL_KEY, ttl)
81}
82
83// A compaction leaves a new prefix, which the cache has none of yet
84const forgetRequest = async ($: EngineInterface): Promise<void> => {
85 await update($, last, prev => (prev === null ? prev : { ...prev, at: null }))
86 await update($, kept, () => null)
87}
88
89// A fork re-sends the main thread's last request with one more message after it: it reads
90// the conversation from the cache, which renews the entry, and leaves the transcript as
91// it was. Resolves when the request started and what it cost, or null when there was
92// nothing to fork
93const renewCache = async ($: EngineInterface): Promise<{ at: number; usage: CompactionUsage } | null> => {
94 const at = await $.clock.now()
95 const result = await $.model.fork({ prompt: RENEW_PROMPT })
96 if (!result.isAnswered && result.reason === 'nothing-to-fork') return null
97 return { at, usage: result.usage }
98}
99
100// A renewal that missed the entry wrote the conversation to the cache again itself, so a
101// compaction right after reads it from there. One missing after an earlier renewal shows
102// that renewal held five minutes; a first one, that this model's renewals miss
103const renew = async ($: EngineInterface, model: string, isCompactOn: boolean, tokens: number): Promise<void> => {
104 isCompacting = true
105 try {
106 const renewal = await renewCache($)
107 if (renewal === null) return
108 const previous = await read($, kept)
109 const { usage } = renewal
110 const record: CompactionRecord = {
111 kind: 'renewal',
112 at: renewal.at,
113 sessionId: await $.session.id(),
114 model,
115 ttl: cacheTtl,
116 before: usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens,
117 after: 0,
118 usage,
119 isReturned: false,
120 }
121 await addRecord($, record)
122 const hit = compactionHit(usage) ?? 0
123 if (previous !== null) await learnRenewTtl($, inferRenewTtl(previous.at, hit, renewal.at))
124 if (hit < RENEW_HIT_FROM) {
125 if (previous === null) await noteRenewMiss($, model)
126 $.ui.log(
127 `Keep warm: the renewal at ${clockTime(renewal.at)} found the prompt cache gone${isCompactOn ? ', so warm compact takes over' : ''}.`,
128 )
129 if (isCompactOn) await compact($, model, tokens)
130 return
131 }
132 const count = (previous?.count ?? 0) + 1
133 await update($, kept, () => ({ at: renewal.at, count }))
134 await update($, pendingRecord, () => renewal.at)
135 $.ui.log(`Kept the prompt cache warm at ${clockTime(renewal.at)} (${count} of ${maxRenewals}).`)
136 } finally {
137 isCompacting = false
138 }
139}
140
141const addRecord = async ($: EngineInterface, record: CompactionRecord): Promise<void> => {
142 const records = recordsOf(await $.store.get(RECORDS_KEY))
143 await $.store.set(RECORDS_KEY, [...records, record].slice(-MAX_RECORDS))
144}
145
146// The session going on after a compaction, or while a renewal still holds the cache, is
147// when the re-read it spared would have been paid
148const markReturned = async ($: EngineInterface): Promise<void> => {
149 const at = await read($, pendingRecord)
150 if (at === null) return
151 await update($, pendingRecord, () => null)
152 const records = recordsOf(await $.store.get(RECORDS_KEY))
153 const now = await $.clock.now()
154 const isKept = (r: CompactionRecord): boolean => r.kind !== 'renewal' || now < r.at + TTL_MS[renewTtl ?? r.ttl]
155 await $.store.set(RECORDS_KEY, records.map(r => (r.at === at && isKept(r) ? { ...r, isReturned: true } : r)))
156}
157
158const sayOnce = ($: EngineInterface, text: string): void => {
159 if (saidOnce.has(text)) return
160 saidOnce.add(text)
161 $.ui.log(text)
162}
163
164const learnRenewTtl = async ($: EngineInterface, ttl: CacheTtl | null): Promise<void> => {
165 if (ttl === null || ttl === renewTtl) return
166 renewTtl = ttl
167 await $.store.set(RENEW_TTL_KEY, ttl)
168}
169
170const noteRenewMiss = async ($: EngineInterface, model: string): Promise<void> => {
171 renewMisses = { ...renewMisses, [model]: engineVersion }
172 await $.store.set(RENEW_MISSES_KEY, renewMisses)
173}
174
175// Why keep warm cannot help here, or null when it can: a renewal that holds five minutes
176// would cost more than a compaction to cover an hour-long cache, and a model whose renewals
177// missed would pay for the whole conversation each time
178const keepWarmBlock = (model: string): string | null => {
179 if (cacheTtl === '1h' && renewTtl === '5m') return 'Keep warm: a renewal holds this account\'s cache for 5 minutes only, so warm compact takes over.'
180 if (renewMisses[model] === engineVersion) return `Keep warm: renewals can't read the conversation from the cache on ${model} with this Claude Code version, so warm compact takes over.`
181 return null
182}
183
184// A plugin's own compaction skips its own hooks, so the request is forgotten here. The toast
185// is for someone watching; the transcript line is for whoever comes back
186const compact = async ($: EngineInterface, model: string, tokens: number): Promise<void> => {
187 isCompacting = true
188 showStatus($, undefined)
189 try {
190 const result = await $.session.compact()
191 if (result.skip !== undefined) return
192 await forgetRequest($)
193 const record: CompactionRecord = {
194 at: await $.clock.now(),
195 sessionId: await $.session.id(),
196 model,
197 ttl: cacheTtl,
198 before: result.tokensBefore ?? tokens,
199 after: result.tokensAfter ?? result.usage?.output_tokens ?? 0,
200 usage: result.usage ?? null,
201 isReturned: false,
202 }
203 await addRecord($, record)
204 await update($, pendingRecord, () => record.at)
205 $.ui.toast(DONE_TEXT)
206 $.ui.log(noticeText(record, result.tokensBefore !== undefined && result.tokensAfter !== undefined))
207 } catch {
208 // A turn began meanwhile, and renews the cache by itself
209 } finally {
210 isCompacting = false
211 }
212}
213
214// Padded alike, so the labels after the chips line up
215export const chipLabel = (isOn: boolean): string => ` ${(isOn ? 'on' : 'off').padEnd(3)} `
216
217// `/warm-compact` flips it, `/warm-compact on` and `off` set it
218export const switchedTo = (args: string, isOn: boolean): boolean | null => {
219 const word = args.trim().toLowerCase()
220 if (word === '') return !isOn
221 if (word === 'on' || word === 'off') return word === 'on'
222 return null
223}
224
225const turnText = (name: string, isOn: boolean): string => `${name} is ${isOn ? 'on' : 'off'} for this session.`
226
227// Turned off mid-countdown, the countdown goes at once
228const turn = async ($: EngineInterface, isOn: boolean): Promise<void> => {
229 await update($, isEnabled, () => isOn)
230 if (!isOn) showStatus($, undefined)
231}
232
233const turnKeepWarm = async ($: EngineInterface, isOn: boolean): Promise<void> => {
234 await update($, isKeepWarm, () => isOn)
235}
236
237const tick = async ($: EngineInterface): Promise<void> => {
238 if (isCompacting) return
239 const isCompactOn = await read($, isEnabled)
240 let isKeepWarmOn = await read($, isKeepWarm)
241 if (!isCompactOn && !isKeepWarmOn) return
242 const request = await read($, last)
243 const renewals = await read($, kept)
244 const { context } = await $.session.usage()
245 const model = await $.session.model()
246 const block = isKeepWarmOn ? keepWarmBlock(model) : null
247 if (block !== null) isKeepWarmOn = false
248 let next = plan(request, {
249 now: await $.clock.now(),
250 ttl: cacheTtl,
251 leadMs,
252 model,
253 tokens: context.tokens,
254 minTokens,
255 isTurnRunning: await read($, isTurnRunning),
256 isCompactOn,
257 isKeepWarmOn,
258 maxRenewals,
259 kept: renewals,
260 renewTtl: renewTtl ?? cacheTtl,
261 })
262 const attempt = `${request?.at ?? ''}:${renewals?.count ?? 0}`
263 if (next.kind !== 'idle' && triedFor === attempt) next = { kind: 'idle' }
264 // A draft in the prompt box means the person is back, and holds a compaction off; a
265 // renewal keeps the cache for them all the same
266 if ((next.kind === 'warn' || next.kind === 'compact') && (await $.prompt.read()).text.trim() !== '') next = { kind: 'idle' }
267 if (next.kind === 'warn') showStatus($, warnText(next.inMs))
268 else showStatus($, undefined)
269 if (next.kind === 'idle' || next.kind === 'warn') return
270 triedFor = attempt
271 if (block !== null) sayOnce($, block)
272 if (next.kind === 'renew') await renew($, model, isCompactOn, context.tokens ?? 0)
273 else await compact($, model, context.tokens ?? 0)
274}
275
276export const register: Register = (on, options) => {
277 minTokens = numberOr(options.minTokens, DEFAULT_MIN_TOKENS)
278 leadMs = Math.max(MIN_LEAD_MS, numberOr(options.leadSeconds, DEFAULT_LEAD_SECONDS) * 1000)
279 maxRenewals = Math.round(numberOr(options.keepWarmRenewals, DEFAULT_MAX_RENEWALS))
280
281 on('session.start', async ($, e, next) => {
282 const result = await next(e)
283 if (!e.isInteractive) return result
284 const stored = await $.store.get(CACHE_TTL_KEY)
285 if (isCacheTtl(stored)) cacheTtl = stored
286 const storedRenewTtl = await $.store.get(RENEW_TTL_KEY)
287 if (isCacheTtl(storedRenewTtl)) renewTtl = storedRenewTtl
288 const misses = await $.store.get(RENEW_MISSES_KEY)
289 if (typeof misses === 'object' && misses !== null) renewMisses = misses as Record<string, string>
290 engineVersion = (await $.session.version()).version
291 await $.command.register({ name: COMMAND, description: 'Compact before the prompt cache goes cold: on or off for this session (nothing flips it), or stats for what it saved' })
292 await $.command.register({ name: KEEP_WARM_COMMAND, description: 'Keep the prompt cache warm while you are away, instead of compacting at once: on or off for this session (nothing flips it)' })
293 $.clock.every(TICK_MS, () => void tick($))
294 return result
295 })
296
297 on('turn.start', async ($, e, next) => {
298 await update($, isTurnRunning, () => true)
299 await markReturned($)
300 return next(e)
301 })
302
303 on('turn.complete', async ($, e, next) => {
304 if (e.agentId === undefined) await update($, isTurnRunning, () => false)
305 return next(e)
306 })
307
308 // Each main-thread request renews the cache; the first after a message says how much of
309 // the conversation was still cached when it was sent
310 on('turn.step', async function* ($, e, next) {
311 if (e.agentId !== undefined) return yield* next(e)
312 const model = await $.session.model()
313 const requestAt = await $.clock.now()
314 return yield* tap(next(e), async chunk => {
315 if (chunk.kind !== 'stop' || chunk.usage === null) return
316 const hit = hitPercent(chunk.usage)
317 const prev = await read($, last)
318 if (e.index === 0 && hit !== null && prev !== null) await learnTtl($, inferTtl(prev, model, hit, requestAt))
319 const renewed = await read($, kept)
320 if (e.index === 0 && hit !== null && renewed !== null && cacheTtl === '1h') await learnRenewTtl($, inferRenewTtl(renewed.at, hit, requestAt))
321 await update($, last, () => ({ at: requestAt, model }))
322 await update($, kept, () => null)
323 })
324 })
325
326 // Anyone's compaction of the main conversation that stands changes the prefix; one
327 // computed ahead or vetoed leaves it, and a subagent's leaves the main one
328 on('session.compact', async ($, e, next) => {
329 const result = await next(e)
330 if (e.trigger !== 'precompute' && e.agentId === undefined && result.skip === undefined) await forgetRequest($)
331 return result
332 })
333
334 on('command.run', { command: COMMAND }, async ($, e) => {
335 if (e.args.trim().toLowerCase() === 'stats') {
336 return { text: statsText(recordsOf(await $.store.get(RECORDS_KEY)), await $.clock.now(), await $.session.id()) }
337 }
338 const isOn = switchedTo(e.args, await read($, isEnabled))
339 if (isOn === null) return { text: `Usage: /${COMMAND} [on|off|stats]` }
340 await turn($, isOn)
341 return { text: turnText('Warm compact', isOn) }
342 })
343
344 on('command.run', { command: KEEP_WARM_COMMAND }, async ($, e) => {
345 const isOn = switchedTo(e.args, await read($, isKeepWarm))
346 if (isOn === null) return { text: `Usage: /${KEEP_WARM_COMMAND} [on|off]` }
347 await turnKeepWarm($, isOn)
348 return { text: turnText('Keep warm', isOn) }
349 })
350
351 // The chips go on the row of chips under the hint line, which the engine still draws: the
352 // one another mod's chips already started, or a row of their own
353 on('ui.render', { component: 'PromptHint' }, async ($, e, next) => {
354 const { above, chips } = splitChipRow(await next(e))
355 const { Box, Text, Button } = $.ui.resolve(e)
356 const isCompactOn = await read($, isEnabled)
357 const isKeepWarmOn = await read($, isKeepWarm)
358 return (
359 <Box flexDirection="column">
360 {above}
361 <Box key={CHIP_ROW_KEY} columnGap={2} flexWrap="wrap">
362 {chips}
363 <Box gap={1}>
364 <Box backgroundColor={isCompactOn ? ON_COLOR : OFF_COLOR}>
365 <Button key="toggle" label={chipLabel(isCompactOn)} plain onPress={() => void turn($, !isCompactOn)} />
366 </Box>
367 <Text dimColor>{COMPACT_LABEL}</Text>
368 </Box>
369 <Box gap={1}>
370 <Box backgroundColor={isKeepWarmOn ? ON_COLOR : OFF_COLOR}>
371 <Button key="keep-warm" label={chipLabel(isKeepWarmOn)} plain onPress={() => void turnKeepWarm($, !isKeepWarmOn)} />
372 </Box>
373 <Text dimColor>{KEEP_WARM_LABEL}</Text>
374 </Box>
375 </Box>
376 </Box>
377 )
378 })
379
380 on('session.end', async ($, e, next) => {
381 if (e.reason === 'clear') {
382 await update($, last, () => null)
383 await update($, kept, () => null)
384 }
385 return next(e)
386 })
387}
388hooks/chips.ts 20 lines1import type { RenderElement, RenderNode } from 'claude-code'
2
3// The mods that put on/off chips under the hint line share one row of them, whichever are
4// loaded: each leaves its chips in a Box under this key, and the one drawn over it adds its
5// own to that row instead of starting another. The same file in each such mod
6export const CHIP_ROW_KEY = 'mod-chip-row'
7
8// What the plugins beneath and the engine drew, as the rows above the chip row and the chips
9// already in it: a column ending in that row, or else the drawing whole with no chips yet
10export const splitChipRow = (beneath: RenderElement): { above: RenderNode[]; chips: RenderNode[] } => {
11 if (beneath.type === 'Box' && beneath.props?.flexDirection === 'column') {
12 const rows = beneath.children ?? []
13 const last = rows.at(-1)
14 if (typeof last === 'object' && last.type === 'Box' && last.props?.key === CHIP_ROW_KEY) {
15 return { above: rows.slice(0, -1), chips: last.children ?? [] }
16 }
17 }
18 return { above: [beneath], chips: [] }
19}
20hooks/plan.ts 102 lines1import type { CacheTtl, Kept, LastRequest } from '../types'
2
3const MINUTE_MS = 60_000
4const SECOND_MS = 1000
5
6export const TTL_MS: Record<CacheTtl, number> = { '5m': 5 * MINUTE_MS, '1h': 60 * MINUTE_MS }
7// What the engine picks for a main thread that may cache for an hour, which a session on a
8// Claude subscription does, until a request shows otherwise
9export const DEFAULT_TTL: CacheTtl = '1h'
10// Past a TTL by this much, a request's hit or miss is the entry's lapse and not a race
11const TTL_MARGIN_MS = 30_000
12const WARM_FROM = 50
13const COLD_BELOW = 10
14// The countdown shows this long before the compaction starts
15export const WARN_MS = 30 * SECOND_MS
16// A compaction started this close to the lapse may reach the API after it
17export const LATE_MS = 5 * SECOND_MS
18
19type RequestUsage = { input_tokens: number; cache_read_input_tokens: number; cache_creation_input_tokens: number }
20
21// The share of the prompt the cache served, out of all the input the request was answered over
22export const hitPercent = (usage: RequestUsage): number | null => {
23 const total = usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens
24 return total === 0 ? null : Math.round((usage.cache_read_input_tokens / total) * 100)
25}
26
27export const isCacheTtl = (value: unknown): value is CacheTtl => value === '5m' || value === '1h'
28
29// A pause longer than five minutes and shorter than an hour tells the TTL apart: a cache
30// still read means an hour, one gone means five minutes. A miss can also come from a
31// prompt that changed meanwhile (an edited CLAUDE.md, a server connected), which the
32// next long pause with a hit corrects
33export const inferTtl = (prev: LastRequest, model: string, hit: number, requestAt: number): CacheTtl | null => {
34 if (prev.at === null || prev.model !== model) return null
35 const pause = requestAt - prev.at
36 if (pause < TTL_MS['5m'] + TTL_MARGIN_MS || pause > TTL_MS['1h'] - TTL_MARGIN_MS) return null
37 if (hit >= WARM_FROM) return '1h'
38 return hit < COLD_BELOW ? '5m' : null
39}
40
41// Past a renewal, a read this warm found the conversation, and one this cold found only
42// the tools and the system prompt, which other sessions keep cached
43const RENEWED_FROM = 90
44const LAPSED_BELOW = 50
45
46// A renewal reads the entry without the hour-long TTL a session's own requests ask for.
47// Whether it holds the entry an hour or five minutes shows in what a read between the two
48// after it finds
49export const inferRenewTtl = (renewedAt: number, hit: number, readAt: number): CacheTtl | null => {
50 const pause = readAt - renewedAt
51 if (pause < TTL_MS['5m'] + TTL_MARGIN_MS || pause > TTL_MS['1h'] - TTL_MARGIN_MS) return null
52 if (hit >= RENEWED_FROM) return '1h'
53 return hit < LAPSED_BELOW ? '5m' : null
54}
55
56export type Moment = {
57 now: number
58 ttl: CacheTtl
59 leadMs: number
60 // The session's model now: the cache entry is per model
61 model: string
62 // The conversation's size as the last response measured it
63 tokens: number | undefined
64 minTokens: number
65 isTurnRunning: boolean
66 isCompactOn: boolean
67 isKeepWarmOn: boolean
68 // The most renewals in one pause, after which a compaction takes over
69 maxRenewals: number
70 kept: Kept | null
71 // How long a renewal holds the entry, which may be less than a request's TTL
72 renewTtl: CacheTtl
73}
74
75export type Plan = { kind: 'idle' } | { kind: 'warn'; inMs: number } | { kind: 'compact' } | { kind: 'renew' }
76
77const IDLE: Plan = { kind: 'idle' }
78
79// Each request renews the entry, so it lapses one TTL after the last, or after the last
80// renewal keep warm made. Before it does, keep warm renews it again, up to `maxRenewals`
81// times in a pause, and then a compaction takes over. Either starts `leadMs` ahead, while
82// its own request can still read the conversation from the cache. A compaction is
83// announced for WARN_MS before it starts; a renewal changes nothing anyone sees. A turn
84// renews the entry by itself; a switched model has none of the conversation cached, and a
85// small conversation costs little to read cold
86export const plan = (last: LastRequest | null, m: Moment): Plan => {
87 if (last === null || last.at === null || m.isTurnRunning || last.model !== m.model) return IDLE
88 if (m.tokens === undefined || m.tokens < m.minTokens) return IDLE
89 const renewals = m.kept?.count ?? 0
90 const action = m.isKeepWarmOn && renewals < m.maxRenewals ? 'renew' : m.isCompactOn ? 'compact' : null
91 if (action === null) return IDLE
92 const ttlMs = m.kept === null ? TTL_MS[m.ttl] : TTL_MS[m.renewTtl]
93 const expiresAt = (m.kept?.at ?? last.at) + ttlMs
94 const startAt = expiresAt - Math.min(m.leadMs, ttlMs - WARN_MS)
95 if (m.now >= expiresAt - LATE_MS) return IDLE
96 if (m.now >= startAt) return { kind: action }
97 if (action === 'compact' && m.now >= startAt - WARN_MS) return { kind: 'warn', inMs: startAt - m.now }
98 return IDLE
99}
100
101export const warnText = (inMs: number): string => `⇊ compacting in ${Math.ceil(inMs / SECOND_MS)}s, before the cache goes cold · type to hold`
102hooks/savings.ts 147 lines1import type { CacheTtl, CompactionRecord, CompactionUsage } from '../types'
2
3const DAY_MS = 24 * 60 * 60_000
4const WEEK_MS = 7 * DAY_MS
5export const MAX_RECORDS = 1000
6// What the API charges, in multiples of a plain input token: a cache write at each TTL,
7// and output, on every current model
8const WRITE_COST: Record<CacheTtl, number> = { '5m': 1.25, '1h': 2 }
9const OUTPUT_COST = 5
10// A cache read is a tenth of an input token, less on the newest models
11const READ_COSTS: readonly [RegExp, number][] = [
12 [/fable-5-1|mythos-5-1/, 0.025],
13 [/opus-5-5/, 0.05],
14]
15const DEFAULT_READ_COST = 0.1
16// Below this share read from the cache, the compaction's request found the entry gone
17const WARM_FROM = 50
18
19export const readCost = (model: string): number => READ_COSTS.find(([pattern]) => pattern.test(model))?.[1] ?? DEFAULT_READ_COST
20
21const isNumber = (value: unknown): value is number => typeof value === 'number' && Number.isFinite(value)
22
23const isUsage = (value: unknown): value is CompactionUsage => {
24 if (typeof value !== 'object' || value === null) return false
25 const usage = value as Record<string, unknown>
26 return ['input_tokens', 'output_tokens', 'cache_read_input_tokens', 'cache_creation_input_tokens'].every(key => isNumber(usage[key]))
27}
28
29export const isRecord = (value: unknown): value is CompactionRecord => {
30 if (typeof value !== 'object' || value === null) return false
31 const record = value as Record<string, unknown>
32 return (
33 (record.kind === undefined || record.kind === 'renewal') &&
34 isNumber(record.at) &&
35 typeof record.sessionId === 'string' &&
36 typeof record.model === 'string' &&
37 (record.ttl === '5m' || record.ttl === '1h') &&
38 isNumber(record.before) &&
39 isNumber(record.after) &&
40 (record.usage === null || isUsage(record.usage)) &&
41 typeof record.isReturned === 'boolean'
42 )
43}
44
45export const recordsOf = (stored: unknown): CompactionRecord[] => (Array.isArray(stored) ? stored.filter(isRecord) : [])
46
47// What a compaction's request read from the cache, out of all it was answered over
48export const compactionHit = (usage: CompactionUsage | null): number | null => {
49 if (usage === null) return null
50 const total = usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens
51 return total === 0 ? null : Math.round((usage.cache_read_input_tokens / total) * 100)
52}
53
54export const isCacheMissed = (usage: CompactionUsage | null): boolean => {
55 const hit = compactionHit(usage)
56 return hit !== null && hit < WARM_FROM
57}
58
59export type Saving = { avoided: number; spent: number }
60
61// In input-token equivalents. Coming back to a cold cache writes the whole conversation
62// to it again. A renewal read the conversation instead, and wrote only its own tail, for
63// five minutes. After a compaction, the summary is written on the way back, and the
64// compaction itself read the conversation and wrote the summary. What was never gone back
65// to, or went cold anyway, spared nothing and paid for itself. Requests after the first
66// back read a smaller conversation too, which is left out, so the tally errs low
67export const savingOf = (record: CompactionRecord): Saving => {
68 const write = WRITE_COST[record.ttl]
69 const usage = record.usage
70 if (record.kind === 'renewal') {
71 const renewal =
72 usage === null
73 ? readCost(record.model) * record.before
74 : readCost(record.model) * usage.cache_read_input_tokens +
75 WRITE_COST['5m'] * usage.cache_creation_input_tokens +
76 usage.input_tokens +
77 OUTPUT_COST * usage.output_tokens
78 return { avoided: record.isReturned ? write * record.before : 0, spent: renewal }
79 }
80 const compaction =
81 usage === null
82 ? readCost(record.model) * record.before + OUTPUT_COST * record.after
83 : readCost(record.model) * usage.cache_read_input_tokens +
84 write * usage.cache_creation_input_tokens +
85 usage.input_tokens +
86 OUTPUT_COST * usage.output_tokens
87 if (!record.isReturned) return { avoided: 0, spent: compaction }
88 return { avoided: write * record.before, spent: compaction + write * record.after }
89}
90
91export const formatTokens = (tokens: number): string => {
92 const size = Math.abs(tokens)
93 if (size >= 1_000_000) return `${(tokens / 1_000_000).toFixed(1)}M`
94 if (size >= 1000) return `${Math.round(tokens / 1000)}k`
95 return String(Math.round(tokens))
96}
97
98const signed = (tokens: number): string => (tokens >= 0 ? `+${formatTokens(tokens)}` : `-${formatTokens(-tokens)}`)
99
100type Row = { label: string; records: CompactionRecord[] }
101
102const rowText = ({ label, records }: Row): string => {
103 const savings = records.map(savingOf)
104 const net = savings.reduce((sum, s) => sum + s.avoided - s.spent, 0)
105 const returned = records.filter(r => r.isReturned).length
106 const renewals = records.filter(r => r.kind === 'renewal').length
107 return [
108 label.padEnd(14),
109 String(records.length - renewals).padStart(11),
110 String(renewals).padStart(10),
111 String(returned).padStart(11),
112 (records.length === 0 ? '-' : signed(net)).padStart(11),
113 ].join('')
114}
115
116export const statsText = (records: readonly CompactionRecord[], now: number, sessionId: string): string => {
117 const rows: Row[] = [
118 { label: 'This session', records: records.filter(r => r.sessionId === sessionId) },
119 { label: 'Last 7 days', records: records.filter(r => now - r.at < WEEK_MS) },
120 { label: 'All time', records: [...records] },
121 ]
122 return [
123 'Savings, in input tokens (a cache write at 1h counts as 2, a cache read as 0.1 or less):',
124 '',
125 `${''.padEnd(14)}${'compactions'.padStart(11)}${'renewals'.padStart(10)}${'came back'.padStart(11)}${'net saved'.padStart(11)}`,
126 ...rows.map(rowText),
127 '',
128 'A compaction, or the last renewal before you came back, counts once its session goes on: the re-read it spared is set against what it cost. What was never gone back to counts its cost alone.',
129 ].join('\n')
130}
131
132const pad = (n: number): string => String(n).padStart(2, '0')
133
134export const clockTime = (at: number): string => {
135 const date = new Date(at)
136 return `${pad(date.getHours())}:${pad(date.getMinutes())}`
137}
138
139// The transcript's lasting line, for whoever comes back to the session
140export const noticeText = (record: CompactionRecord, hasSizes: boolean): string => {
141 const sizes = hasSizes ? ` (${formatTokens(record.before)} → ${formatTokens(record.after)} tokens)` : ''
142 if (isCacheMissed(record.usage)) {
143 return `Compacted at ${clockTime(record.at)}${sizes}, but the prompt cache had already lapsed, so this one saved nothing.`
144 }
145 return `Compacted at ${clockTime(record.at)}${sizes}, just before the prompt cache went cold. The conversation goes on from the summary.`
146}
147types/index.d.ts 52 lines1export type CacheTtl = '5m' | '1h'
2// The last main-thread request, which renewed the cache entry the next prompt would read.
3// `at`: when it was made, null once a compaction or a /clear changed the prefix.
4// `model`: the session's model then, the one the entry belongs to
5export type LastRequest = { at: number | null; model: string }
6// The renewals keep warm made since the last request: the latest's start, and how many
7export type Kept = { at: number; count: number }
8
9// What a compaction's own request was answered over, as the API reports it
10export type CompactionUsage = {
11 input_tokens: number
12 output_tokens: number
13 cache_read_input_tokens: number
14 cache_creation_input_tokens: number
15}
16
17// One compaction or cache renewal warm-compact made, kept across sessions to tally what
18// they saved. `kind`: absent for a compaction.
19// `before`, `after`: the conversation's size in tokens either side of a compaction, or the
20// size a renewal read (`after` 0).
21// `usage`: its request's, when the engine reported one.
22// `isReturned`: whether the session went on afterwards with what it kept, which is when the
23// re-read it spared would have been paid
24export type CompactionRecord = {
25 kind?: 'renewal'
26 at: number
27 sessionId: string
28 model: string
29 ttl: CacheTtl
30 before: number
31 after: number
32 usage: CompactionUsage | null
33 isReturned: boolean
34}
35
36declare module 'claude-code' {
37 interface PluginState {
38 'warm-compact': {
39 last: LastRequest | null
40 isTurnRunning: boolean
41 // Whether it compacts at all in this session: the footer chip and /warm-compact flip it
42 isEnabled: boolean
43 // Whether it renews the cache in compaction's place: the footer's second chip and
44 // /keep-warm flip it
45 isKeepWarm: boolean
46 kept: Kept | null
47 // When this session's latest compaction or renewal ran, until the session goes on
48 pendingRecord: number | null
49 }
50 }
51}
52