SLOPSHOPPER

warm-compact

Compacts an idle session just before its prompt cache goes cold, or keeps the cache warm while you're away, so your next prompt doesn't re-read the whole…

newspinnercommandtoaststatusmodel
v0.3.1MITupdated 2026-10-06ziedgithub/claude-code-mods/warm-compact
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · warm-compact
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /warm-compact ⎿ warm-compact: Warm compact is off for this session. ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ⟨Claude Code's own drawing⟩ off Warm compact on Keep warm

Draws

Prompt hint
⟨Claude Code's own drawing⟩ off
README

warm-compact

Compacts an idle session just before its prompt cache goes cold, or keeps the cache warm while you're away.

It is separate from Claude Code's own auto-compact, which compacts when the context window fills up. warm-compact acts on time instead: it compacts when you've been away long enough for the cache to lapse, even when the window is far from full.

Claude Code caches the conversation for 5 minutes or an hour after each request, depending on your account. If you come back after that, your next prompt re-reads the whole conversation uncached, which is the most expensive way to read it. warm-compact compacts the session a minute before the cache lapses, while the compaction itself can still read the conversation from the cache. Your next prompt then starts from a short summary.

When it compacts

  • Only when the session is idle. Nothing is running, and the prompt box is empty. Typing anything holds it off.
  • Only a large conversation. Below 50,000 tokens, reading the conversation uncached costs little, so it's left alone.
  • Once per pause. After a compaction, nothing happens again until your next prompt.
  • Same model. The cache belongs to one model, so after a /model switch there's nothing left to keep warm.

For the 30 seconds before it compacts, the status line under the prompt counts down: ⇊ compacting in 23s, before the cache goes cold · type to hold.

When you come back, a line in the transcript tells you what happened, for example Compacted at 14:32 (120k → 4k tokens), just before the prompt cache went cold. That line is for you only: it isn't sent to the model. If the compaction found the cache already gone, the line says so, because then it saved nothing.

Whether your account caches for 5 minutes or an hour is learned from the requests themselves and remembered between sessions. Until a request shows otherwise, it assumes an hour, as a Claude subscription has.

Headless runs (claude -p) are left alone.

Turning it off for a session

Two chips sit side by side on one row under the prompt, below the mode line, a row they share with the chips of any other mod that puts its own there:

⏸ manual mode on · ? for shortcuts
 on   Warm compact   off  Keep warm

Click Warm compact to turn it off for this session, and again to turn it back on. Turning it off during the countdown cancels it.

The same from the prompt:

/warm-compact        flips it
/warm-compact off    turns it off
/warm-compact on     turns it on
/warm-compact stats  shows what it saved

A new session, or a /clear, starts with it on.

Keep warm: renew the cache instead

A compaction keeps your session cheap, but it trades the conversation for a summary. Keep warm keeps the whole conversation instead: a minute before the cache lapses, it renews it with one tiny request that re-reads the conversation from the cache. That request costs about a tenth of the conversation's size in tokens (a twentieth on Opus 5.5) and starts the cache's clock again. Coming back to a cold cache would cost twice its size.

It's off by default, because it spends tokens while you're away. Turn it on with the Keep warm chip, or from the prompt:

/keep-warm        flips it
/keep-warm on     turns it on
/keep-warm off    turns it off
  • It renews the cache at most twice per pause (two hours on an hour-long cache), then warm compact takes over, if it's on. The keepWarmRenewals option changes how many times.
  • A renewal is invisible. It adds nothing to the conversation and the model never sees it; a line in the transcript, Kept the prompt cache warm at 14:32 (1 of 2)., is for you. Text left in the prompt box doesn't hold it off.
  • Every renewal is checked. If one finds the cache gone, warm compact takes over at once, while it's cheap to.
  • Some models can't be renewed. On Sonnet 5.5, with Claude Code 2.1.289, the renewal request differs slightly from the session's own and misses the cache. Keep warm notices the first miss and leaves that model to warm compact until Claude Code updates.
  • How long a renewal holds the cache is learned the first time you come back, or renew again, after one. If it turns out to be only 5 minutes on your account, keeping an hour-long cache warm would cost more than compacting, so keep warm stands down and says so.

Renewals count in /warm-compact stats too: each costs its request, and the last one before you come back is credited with the re-read it spared.

What it saved

/warm-compact stats adds up what the compactions saved, for this session, the last 7 days and all time:

              compactions  renewals  came back  net saved
This session            1         0          1      +222k
Last 7 days             5         3          6      +1.4M
All time                9         5         11      +2.6M

Savings are counted in input tokens, the unit the API prices everything against: writing a token to an hour-long cache costs 2 of them (1.25 for a 5-minute cache), reading one from the cache 0.1 (0.05 on Opus 5.5, 0.025 on Fable 5.1), and generating one 5.

  • Once you come back to a session, the compaction is credited with the re-read it spared you: the whole conversation written to the cache again. Against that it is charged its own request (reading the conversation from the cache and writing the summary) and writing the summary to the cache when you return.
  • A session you never come back to is charged the compaction alone.
  • Every request after your return reads a smaller conversation too. That saving is left out, so the tally errs low.

On a Claude subscription you don't pay per token, but the same tokens count toward your usage limits.

Options

In /config, or in ~/.claude/settings.json under pluginConfigs:

OptionDefaultWhat it does
minTokens50000The smallest conversation it compacts, in tokens
leadSeconds60How long before the cache lapses a compaction or a renewal starts (15 at least)
keepWarmRenewals2With keep warm on, how many times the cache is renewed in one pause before warm compact takes over

Install

claude plugin marketplace add ziedgithub/claude-code-mods && claude plugin install warm-compact@zied-mods

Then restart Claude Code.

See the repository README for running it from a clone.

Source 5 files
hooks/register.tsx 388 lines
1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { CacheTtl, CompactionRecord, CompactionUsage } from '../types'
5import { CHIP_ROW_KEY, splitChipRow } from './chips'
6import { DEFAULT_TTL, TTL_MS, hitPercent, inferRenewTtl, inferTtl, isCacheTtl, plan, warnText } from './plan'
7import { MAX_RECORDS, clockTime, compactionHit, noticeText, recordsOf, statsText } from './savings'
8
9const TICK_MS = 1000
10const CACHE_TTL_KEY = 'cacheTtl'
11const RECORDS_KEY = 'compactions'
12const RENEW_TTL_KEY = 'renewTtl'
13// The models a renewal could not read the conversation from the cache on, each with the
14// engine version that failed: tried again once the engine changes
15const RENEW_MISSES_KEY = 'renewMisses'
16// The fork that renews the cache asks for as little as it can
17const RENEW_PROMPT = 'Reply with the single word: ok'
18// A renewal that read less of its request from the cache than this missed the entry
19const RENEW_HIT_FROM = 90
20const DEFAULT_MIN_TOKENS = 50_000
21const DEFAULT_LEAD_SECONDS = 60
22const DEFAULT_MAX_RENEWALS = 2
23// Never so close to the lapse that the compaction's request could miss the cache
24const MIN_LEAD_MS = 15_000
25const DONE_TEXT = 'Compacted before the prompt cache went cold'
26const COMMAND = 'warm-compact'
27const KEEP_WARM_COMMAND = 'keep-warm'
28const COMPACT_LABEL = 'Warm compact'
29const KEEP_WARM_LABEL = 'Keep warm'
30const ON_COLOR = '#238636'
31const OFF_COLOR = '#6e7681'
32
33const last = atom({ plugin: 'warm-compact', key: 'last' } as const, null)
34const isTurnRunning = atom({ plugin: 'warm-compact', key: 'isTurnRunning' } as const, false)
35const isEnabled = atom({ plugin: 'warm-compact', key: 'isEnabled' } as const, true)
36const isKeepWarm = atom({ plugin: 'warm-compact', key: 'isKeepWarm' } as const, false)
37const kept = atom({ plugin: 'warm-compact', key: 'kept' } as const, null)
38const pendingRecord = atom({ plugin: 'warm-compact', key: 'pendingRecord' } as const, null)
39
40let cacheTtl: CacheTtl = DEFAULT_TTL
41let shownStatus: string | undefined
42let isCompacting = false
43// The renewal or compaction last tried, named by the request and the renewals since: one
44// try each, whatever came of it
45let triedFor: string | null = null
46let maxRenewals = DEFAULT_MAX_RENEWALS
47// Until a read after a renewal shows otherwise, it holds the entry as long as a request does
48let renewTtl: CacheTtl | null = null
49let renewMisses: Record<string, string> = {}
50let engineVersion = ''
51// Each reason keep warm stood down is said once a session
52const saidOnce = new Set<string>()
53let minTokens = DEFAULT_MIN_TOKENS
54let leadMs = DEFAULT_LEAD_SECONDS * 1000
55
56const numberOr = (value: unknown, fallback: number): number =>
57  typeof value === 'number' && Number.isFinite(value) && value >= 0 ? value : fallback
58
59const showStatus = ($: EngineInterface, text: string | undefined): void => {
60  if (text === shownStatus) return
61  shownStatus = text
62  $.ui.status(text)
63}
64
65// Passes the response on as it streams, seeing each chunk on the way
66async function* tap<T, R>(stream: AsyncGenerator<T, R>, see: (chunk: T) => Promise<void>): AsyncGenerator<T, R> {
67  for (;;) {
68    const step = await stream.next()
69    if (step.done === true) return step.value
70    await see(step.value)
71    yield step.value
72  }
73}
74
75// The engine keeps its TTL to itself: a request after a long enough pause tells it, and it
76// is kept between sessions since it holds for the account
77const learnTtl = async ($: EngineInterface, ttl: CacheTtl | null): Promise<void> => {
78  if (ttl === null || ttl === cacheTtl) return
79  cacheTtl = ttl
80  await $.store.set(CACHE_TTL_KEY, ttl)
81}
82
83// A compaction leaves a new prefix, which the cache has none of yet
84const forgetRequest = async ($: EngineInterface): Promise<void> => {
85  await update($, last, prev => (prev === null ? prev : { ...prev, at: null }))
86  await update($, kept, () => null)
87}
88
89// A fork re-sends the main thread's last request with one more message after it: it reads
90// the conversation from the cache, which renews the entry, and leaves the transcript as
91// it was. Resolves when the request started and what it cost, or null when there was
92// nothing to fork
93const renewCache = async ($: EngineInterface): Promise<{ at: number; usage: CompactionUsage } | null> => {
94  const at = await $.clock.now()
95  const result = await $.model.fork({ prompt: RENEW_PROMPT })
96  if (!result.isAnswered && result.reason === 'nothing-to-fork') return null
97  return { at, usage: result.usage }
98}
99
100// A renewal that missed the entry wrote the conversation to the cache again itself, so a
101// compaction right after reads it from there. One missing after an earlier renewal shows
102// that renewal held five minutes; a first one, that this model's renewals miss
103const renew = async ($: EngineInterface, model: string, isCompactOn: boolean, tokens: number): Promise<void> => {
104  isCompacting = true
105  try {
106    const renewal = await renewCache($)
107    if (renewal === null) return
108    const previous = await read($, kept)
109    const { usage } = renewal
110    const record: CompactionRecord = {
111      kind: 'renewal',
112      at: renewal.at,
113      sessionId: await $.session.id(),
114      model,
115      ttl: cacheTtl,
116      before: usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens,
117      after: 0,
118      usage,
119      isReturned: false,
120    }
121    await addRecord($, record)
122    const hit = compactionHit(usage) ?? 0
123    if (previous !== null) await learnRenewTtl($, inferRenewTtl(previous.at, hit, renewal.at))
124    if (hit < RENEW_HIT_FROM) {
125      if (previous === null) await noteRenewMiss($, model)
126      $.ui.log(
127        `Keep warm: the renewal at ${clockTime(renewal.at)} found the prompt cache gone${isCompactOn ? ', so warm compact takes over' : ''}.`,
128      )
129      if (isCompactOn) await compact($, model, tokens)
130      return
131    }
132    const count = (previous?.count ?? 0) + 1
133    await update($, kept, () => ({ at: renewal.at, count }))
134    await update($, pendingRecord, () => renewal.at)
135    $.ui.log(`Kept the prompt cache warm at ${clockTime(renewal.at)} (${count} of ${maxRenewals}).`)
136  } finally {
137    isCompacting = false
138  }
139}
140
141const addRecord = async ($: EngineInterface, record: CompactionRecord): Promise<void> => {
142  const records = recordsOf(await $.store.get(RECORDS_KEY))
143  await $.store.set(RECORDS_KEY, [...records, record].slice(-MAX_RECORDS))
144}
145
146// The session going on after a compaction, or while a renewal still holds the cache, is
147// when the re-read it spared would have been paid
148const markReturned = async ($: EngineInterface): Promise<void> => {
149  const at = await read($, pendingRecord)
150  if (at === null) return
151  await update($, pendingRecord, () => null)
152  const records = recordsOf(await $.store.get(RECORDS_KEY))
153  const now = await $.clock.now()
154  const isKept = (r: CompactionRecord): boolean => r.kind !== 'renewal' || now < r.at + TTL_MS[renewTtl ?? r.ttl]
155  await $.store.set(RECORDS_KEY, records.map(r => (r.at === at && isKept(r) ? { ...r, isReturned: true } : r)))
156}
157
158const sayOnce = ($: EngineInterface, text: string): void => {
159  if (saidOnce.has(text)) return
160  saidOnce.add(text)
161  $.ui.log(text)
162}
163
164const learnRenewTtl = async ($: EngineInterface, ttl: CacheTtl | null): Promise<void> => {
165  if (ttl === null || ttl === renewTtl) return
166  renewTtl = ttl
167  await $.store.set(RENEW_TTL_KEY, ttl)
168}
169
170const noteRenewMiss = async ($: EngineInterface, model: string): Promise<void> => {
171  renewMisses = { ...renewMisses, [model]: engineVersion }
172  await $.store.set(RENEW_MISSES_KEY, renewMisses)
173}
174
175// Why keep warm cannot help here, or null when it can: a renewal that holds five minutes
176// would cost more than a compaction to cover an hour-long cache, and a model whose renewals
177// missed would pay for the whole conversation each time
178const keepWarmBlock = (model: string): string | null => {
179  if (cacheTtl === '1h' && renewTtl === '5m') return 'Keep warm: a renewal holds this account\'s cache for 5 minutes only, so warm compact takes over.'
180  if (renewMisses[model] === engineVersion) return `Keep warm: renewals can't read the conversation from the cache on ${model} with this Claude Code version, so warm compact takes over.`
181  return null
182}
183
184// A plugin's own compaction skips its own hooks, so the request is forgotten here. The toast
185// is for someone watching; the transcript line is for whoever comes back
186const compact = async ($: EngineInterface, model: string, tokens: number): Promise<void> => {
187  isCompacting = true
188  showStatus($, undefined)
189  try {
190    const result = await $.session.compact()
191    if (result.skip !== undefined) return
192    await forgetRequest($)
193    const record: CompactionRecord = {
194      at: await $.clock.now(),
195      sessionId: await $.session.id(),
196      model,
197      ttl: cacheTtl,
198      before: result.tokensBefore ?? tokens,
199      after: result.tokensAfter ?? result.usage?.output_tokens ?? 0,
200      usage: result.usage ?? null,
201      isReturned: false,
202    }
203    await addRecord($, record)
204    await update($, pendingRecord, () => record.at)
205    $.ui.toast(DONE_TEXT)
206    $.ui.log(noticeText(record, result.tokensBefore !== undefined && result.tokensAfter !== undefined))
207  } catch {
208    // A turn began meanwhile, and renews the cache by itself
209  } finally {
210    isCompacting = false
211  }
212}
213
214// Padded alike, so the labels after the chips line up
215export const chipLabel = (isOn: boolean): string => ` ${(isOn ? 'on' : 'off').padEnd(3)} `
216
217// `/warm-compact` flips it, `/warm-compact on` and `off` set it
218export const switchedTo = (args: string, isOn: boolean): boolean | null => {
219  const word = args.trim().toLowerCase()
220  if (word === '') return !isOn
221  if (word === 'on' || word === 'off') return word === 'on'
222  return null
223}
224
225const turnText = (name: string, isOn: boolean): string => `${name} is ${isOn ? 'on' : 'off'} for this session.`
226
227// Turned off mid-countdown, the countdown goes at once
228const turn = async ($: EngineInterface, isOn: boolean): Promise<void> => {
229  await update($, isEnabled, () => isOn)
230  if (!isOn) showStatus($, undefined)
231}
232
233const turnKeepWarm = async ($: EngineInterface, isOn: boolean): Promise<void> => {
234  await update($, isKeepWarm, () => isOn)
235}
236
237const tick = async ($: EngineInterface): Promise<void> => {
238  if (isCompacting) return
239  const isCompactOn = await read($, isEnabled)
240  let isKeepWarmOn = await read($, isKeepWarm)
241  if (!isCompactOn && !isKeepWarmOn) return
242  const request = await read($, last)
243  const renewals = await read($, kept)
244  const { context } = await $.session.usage()
245  const model = await $.session.model()
246  const block = isKeepWarmOn ? keepWarmBlock(model) : null
247  if (block !== null) isKeepWarmOn = false
248  let next = plan(request, {
249    now: await $.clock.now(),
250    ttl: cacheTtl,
251    leadMs,
252    model,
253    tokens: context.tokens,
254    minTokens,
255    isTurnRunning: await read($, isTurnRunning),
256    isCompactOn,
257    isKeepWarmOn,
258    maxRenewals,
259    kept: renewals,
260    renewTtl: renewTtl ?? cacheTtl,
261  })
262  const attempt = `${request?.at ?? ''}:${renewals?.count ?? 0}`
263  if (next.kind !== 'idle' && triedFor === attempt) next = { kind: 'idle' }
264  // A draft in the prompt box means the person is back, and holds a compaction off; a
265  // renewal keeps the cache for them all the same
266  if ((next.kind === 'warn' || next.kind === 'compact') && (await $.prompt.read()).text.trim() !== '') next = { kind: 'idle' }
267  if (next.kind === 'warn') showStatus($, warnText(next.inMs))
268  else showStatus($, undefined)
269  if (next.kind === 'idle' || next.kind === 'warn') return
270  triedFor = attempt
271  if (block !== null) sayOnce($, block)
272  if (next.kind === 'renew') await renew($, model, isCompactOn, context.tokens ?? 0)
273  else await compact($, model, context.tokens ?? 0)
274}
275
276export const register: Register = (on, options) => {
277  minTokens = numberOr(options.minTokens, DEFAULT_MIN_TOKENS)
278  leadMs = Math.max(MIN_LEAD_MS, numberOr(options.leadSeconds, DEFAULT_LEAD_SECONDS) * 1000)
279  maxRenewals = Math.round(numberOr(options.keepWarmRenewals, DEFAULT_MAX_RENEWALS))
280
281  on('session.start', async ($, e, next) => {
282    const result = await next(e)
283    if (!e.isInteractive) return result
284    const stored = await $.store.get(CACHE_TTL_KEY)
285    if (isCacheTtl(stored)) cacheTtl = stored
286    const storedRenewTtl = await $.store.get(RENEW_TTL_KEY)
287    if (isCacheTtl(storedRenewTtl)) renewTtl = storedRenewTtl
288    const misses = await $.store.get(RENEW_MISSES_KEY)
289    if (typeof misses === 'object' && misses !== null) renewMisses = misses as Record<string, string>
290    engineVersion = (await $.session.version()).version
291    await $.command.register({ name: COMMAND, description: 'Compact before the prompt cache goes cold: on or off for this session (nothing flips it), or stats for what it saved' })
292    await $.command.register({ name: KEEP_WARM_COMMAND, description: 'Keep the prompt cache warm while you are away, instead of compacting at once: on or off for this session (nothing flips it)' })
293    $.clock.every(TICK_MS, () => void tick($))
294    return result
295  })
296
297  on('turn.start', async ($, e, next) => {
298    await update($, isTurnRunning, () => true)
299    await markReturned($)
300    return next(e)
301  })
302
303  on('turn.complete', async ($, e, next) => {
304    if (e.agentId === undefined) await update($, isTurnRunning, () => false)
305    return next(e)
306  })
307
308  // Each main-thread request renews the cache; the first after a message says how much of
309  // the conversation was still cached when it was sent
310  on('turn.step', async function* ($, e, next) {
311    if (e.agentId !== undefined) return yield* next(e)
312    const model = await $.session.model()
313    const requestAt = await $.clock.now()
314    return yield* tap(next(e), async chunk => {
315      if (chunk.kind !== 'stop' || chunk.usage === null) return
316      const hit = hitPercent(chunk.usage)
317      const prev = await read($, last)
318      if (e.index === 0 && hit !== null && prev !== null) await learnTtl($, inferTtl(prev, model, hit, requestAt))
319      const renewed = await read($, kept)
320      if (e.index === 0 && hit !== null && renewed !== null && cacheTtl === '1h') await learnRenewTtl($, inferRenewTtl(renewed.at, hit, requestAt))
321      await update($, last, () => ({ at: requestAt, model }))
322      await update($, kept, () => null)
323    })
324  })
325
326  // Anyone's compaction of the main conversation that stands changes the prefix; one
327  // computed ahead or vetoed leaves it, and a subagent's leaves the main one
328  on('session.compact', async ($, e, next) => {
329    const result = await next(e)
330    if (e.trigger !== 'precompute' && e.agentId === undefined && result.skip === undefined) await forgetRequest($)
331    return result
332  })
333
334  on('command.run', { command: COMMAND }, async ($, e) => {
335    if (e.args.trim().toLowerCase() === 'stats') {
336      return { text: statsText(recordsOf(await $.store.get(RECORDS_KEY)), await $.clock.now(), await $.session.id()) }
337    }
338    const isOn = switchedTo(e.args, await read($, isEnabled))
339    if (isOn === null) return { text: `Usage: /${COMMAND} [on|off|stats]` }
340    await turn($, isOn)
341    return { text: turnText('Warm compact', isOn) }
342  })
343
344  on('command.run', { command: KEEP_WARM_COMMAND }, async ($, e) => {
345    const isOn = switchedTo(e.args, await read($, isKeepWarm))
346    if (isOn === null) return { text: `Usage: /${KEEP_WARM_COMMAND} [on|off]` }
347    await turnKeepWarm($, isOn)
348    return { text: turnText('Keep warm', isOn) }
349  })
350
351  // The chips go on the row of chips under the hint line, which the engine still draws: the
352  // one another mod's chips already started, or a row of their own
353  on('ui.render', { component: 'PromptHint' }, async ($, e, next) => {
354    const { above, chips } = splitChipRow(await next(e))
355    const { Box, Text, Button } = $.ui.resolve(e)
356    const isCompactOn = await read($, isEnabled)
357    const isKeepWarmOn = await read($, isKeepWarm)
358    return (
359      <Box flexDirection="column">
360        {above}
361        <Box key={CHIP_ROW_KEY} columnGap={2} flexWrap="wrap">
362          {chips}
363          <Box gap={1}>
364            <Box backgroundColor={isCompactOn ? ON_COLOR : OFF_COLOR}>
365              <Button key="toggle" label={chipLabel(isCompactOn)} plain onPress={() => void turn($, !isCompactOn)} />
366            </Box>
367            <Text dimColor>{COMPACT_LABEL}</Text>
368          </Box>
369          <Box gap={1}>
370            <Box backgroundColor={isKeepWarmOn ? ON_COLOR : OFF_COLOR}>
371              <Button key="keep-warm" label={chipLabel(isKeepWarmOn)} plain onPress={() => void turnKeepWarm($, !isKeepWarmOn)} />
372            </Box>
373            <Text dimColor>{KEEP_WARM_LABEL}</Text>
374          </Box>
375        </Box>
376      </Box>
377    )
378  })
379
380  on('session.end', async ($, e, next) => {
381    if (e.reason === 'clear') {
382      await update($, last, () => null)
383      await update($, kept, () => null)
384    }
385    return next(e)
386  })
387}
388
hooks/chips.ts 20 lines
1import type { RenderElement, RenderNode } from 'claude-code'
2
3// The mods that put on/off chips under the hint line share one row of them, whichever are
4// loaded: each leaves its chips in a Box under this key, and the one drawn over it adds its
5// own to that row instead of starting another. The same file in each such mod
6export const CHIP_ROW_KEY = 'mod-chip-row'
7
8// What the plugins beneath and the engine drew, as the rows above the chip row and the chips
9// already in it: a column ending in that row, or else the drawing whole with no chips yet
10export const splitChipRow = (beneath: RenderElement): { above: RenderNode[]; chips: RenderNode[] } => {
11  if (beneath.type === 'Box' && beneath.props?.flexDirection === 'column') {
12    const rows = beneath.children ?? []
13    const last = rows.at(-1)
14    if (typeof last === 'object' && last.type === 'Box' && last.props?.key === CHIP_ROW_KEY) {
15      return { above: rows.slice(0, -1), chips: last.children ?? [] }
16    }
17  }
18  return { above: [beneath], chips: [] }
19}
20
hooks/plan.ts 102 lines
1import type { CacheTtl, Kept, LastRequest } from '../types'
2
3const MINUTE_MS = 60_000
4const SECOND_MS = 1000
5
6export const TTL_MS: Record<CacheTtl, number> = { '5m': 5 * MINUTE_MS, '1h': 60 * MINUTE_MS }
7// What the engine picks for a main thread that may cache for an hour, which a session on a
8// Claude subscription does, until a request shows otherwise
9export const DEFAULT_TTL: CacheTtl = '1h'
10// Past a TTL by this much, a request's hit or miss is the entry's lapse and not a race
11const TTL_MARGIN_MS = 30_000
12const WARM_FROM = 50
13const COLD_BELOW = 10
14// The countdown shows this long before the compaction starts
15export const WARN_MS = 30 * SECOND_MS
16// A compaction started this close to the lapse may reach the API after it
17export const LATE_MS = 5 * SECOND_MS
18
19type RequestUsage = { input_tokens: number; cache_read_input_tokens: number; cache_creation_input_tokens: number }
20
21// The share of the prompt the cache served, out of all the input the request was answered over
22export const hitPercent = (usage: RequestUsage): number | null => {
23  const total = usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens
24  return total === 0 ? null : Math.round((usage.cache_read_input_tokens / total) * 100)
25}
26
27export const isCacheTtl = (value: unknown): value is CacheTtl => value === '5m' || value === '1h'
28
29// A pause longer than five minutes and shorter than an hour tells the TTL apart: a cache
30// still read means an hour, one gone means five minutes. A miss can also come from a
31// prompt that changed meanwhile (an edited CLAUDE.md, a server connected), which the
32// next long pause with a hit corrects
33export const inferTtl = (prev: LastRequest, model: string, hit: number, requestAt: number): CacheTtl | null => {
34  if (prev.at === null || prev.model !== model) return null
35  const pause = requestAt - prev.at
36  if (pause < TTL_MS['5m'] + TTL_MARGIN_MS || pause > TTL_MS['1h'] - TTL_MARGIN_MS) return null
37  if (hit >= WARM_FROM) return '1h'
38  return hit < COLD_BELOW ? '5m' : null
39}
40
41// Past a renewal, a read this warm found the conversation, and one this cold found only
42// the tools and the system prompt, which other sessions keep cached
43const RENEWED_FROM = 90
44const LAPSED_BELOW = 50
45
46// A renewal reads the entry without the hour-long TTL a session's own requests ask for.
47// Whether it holds the entry an hour or five minutes shows in what a read between the two
48// after it finds
49export const inferRenewTtl = (renewedAt: number, hit: number, readAt: number): CacheTtl | null => {
50  const pause = readAt - renewedAt
51  if (pause < TTL_MS['5m'] + TTL_MARGIN_MS || pause > TTL_MS['1h'] - TTL_MARGIN_MS) return null
52  if (hit >= RENEWED_FROM) return '1h'
53  return hit < LAPSED_BELOW ? '5m' : null
54}
55
56export type Moment = {
57  now: number
58  ttl: CacheTtl
59  leadMs: number
60  // The session's model now: the cache entry is per model
61  model: string
62  // The conversation's size as the last response measured it
63  tokens: number | undefined
64  minTokens: number
65  isTurnRunning: boolean
66  isCompactOn: boolean
67  isKeepWarmOn: boolean
68  // The most renewals in one pause, after which a compaction takes over
69  maxRenewals: number
70  kept: Kept | null
71  // How long a renewal holds the entry, which may be less than a request's TTL
72  renewTtl: CacheTtl
73}
74
75export type Plan = { kind: 'idle' } | { kind: 'warn'; inMs: number } | { kind: 'compact' } | { kind: 'renew' }
76
77const IDLE: Plan = { kind: 'idle' }
78
79// Each request renews the entry, so it lapses one TTL after the last, or after the last
80// renewal keep warm made. Before it does, keep warm renews it again, up to `maxRenewals`
81// times in a pause, and then a compaction takes over. Either starts `leadMs` ahead, while
82// its own request can still read the conversation from the cache. A compaction is
83// announced for WARN_MS before it starts; a renewal changes nothing anyone sees. A turn
84// renews the entry by itself; a switched model has none of the conversation cached, and a
85// small conversation costs little to read cold
86export const plan = (last: LastRequest | null, m: Moment): Plan => {
87  if (last === null || last.at === null || m.isTurnRunning || last.model !== m.model) return IDLE
88  if (m.tokens === undefined || m.tokens < m.minTokens) return IDLE
89  const renewals = m.kept?.count ?? 0
90  const action = m.isKeepWarmOn && renewals < m.maxRenewals ? 'renew' : m.isCompactOn ? 'compact' : null
91  if (action === null) return IDLE
92  const ttlMs = m.kept === null ? TTL_MS[m.ttl] : TTL_MS[m.renewTtl]
93  const expiresAt = (m.kept?.at ?? last.at) + ttlMs
94  const startAt = expiresAt - Math.min(m.leadMs, ttlMs - WARN_MS)
95  if (m.now >= expiresAt - LATE_MS) return IDLE
96  if (m.now >= startAt) return { kind: action }
97  if (action === 'compact' && m.now >= startAt - WARN_MS) return { kind: 'warn', inMs: startAt - m.now }
98  return IDLE
99}
100
101export const warnText = (inMs: number): string => `⇊ compacting in ${Math.ceil(inMs / SECOND_MS)}s, before the cache goes cold · type to hold`
102
hooks/savings.ts 147 lines
1import type { CacheTtl, CompactionRecord, CompactionUsage } from '../types'
2
3const DAY_MS = 24 * 60 * 60_000
4const WEEK_MS = 7 * DAY_MS
5export const MAX_RECORDS = 1000
6// What the API charges, in multiples of a plain input token: a cache write at each TTL,
7// and output, on every current model
8const WRITE_COST: Record<CacheTtl, number> = { '5m': 1.25, '1h': 2 }
9const OUTPUT_COST = 5
10// A cache read is a tenth of an input token, less on the newest models
11const READ_COSTS: readonly [RegExp, number][] = [
12  [/fable-5-1|mythos-5-1/, 0.025],
13  [/opus-5-5/, 0.05],
14]
15const DEFAULT_READ_COST = 0.1
16// Below this share read from the cache, the compaction's request found the entry gone
17const WARM_FROM = 50
18
19export const readCost = (model: string): number => READ_COSTS.find(([pattern]) => pattern.test(model))?.[1] ?? DEFAULT_READ_COST
20
21const isNumber = (value: unknown): value is number => typeof value === 'number' && Number.isFinite(value)
22
23const isUsage = (value: unknown): value is CompactionUsage => {
24  if (typeof value !== 'object' || value === null) return false
25  const usage = value as Record<string, unknown>
26  return ['input_tokens', 'output_tokens', 'cache_read_input_tokens', 'cache_creation_input_tokens'].every(key => isNumber(usage[key]))
27}
28
29export const isRecord = (value: unknown): value is CompactionRecord => {
30  if (typeof value !== 'object' || value === null) return false
31  const record = value as Record<string, unknown>
32  return (
33    (record.kind === undefined || record.kind === 'renewal') &&
34    isNumber(record.at) &&
35    typeof record.sessionId === 'string' &&
36    typeof record.model === 'string' &&
37    (record.ttl === '5m' || record.ttl === '1h') &&
38    isNumber(record.before) &&
39    isNumber(record.after) &&
40    (record.usage === null || isUsage(record.usage)) &&
41    typeof record.isReturned === 'boolean'
42  )
43}
44
45export const recordsOf = (stored: unknown): CompactionRecord[] => (Array.isArray(stored) ? stored.filter(isRecord) : [])
46
47// What a compaction's request read from the cache, out of all it was answered over
48export const compactionHit = (usage: CompactionUsage | null): number | null => {
49  if (usage === null) return null
50  const total = usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens
51  return total === 0 ? null : Math.round((usage.cache_read_input_tokens / total) * 100)
52}
53
54export const isCacheMissed = (usage: CompactionUsage | null): boolean => {
55  const hit = compactionHit(usage)
56  return hit !== null && hit < WARM_FROM
57}
58
59export type Saving = { avoided: number; spent: number }
60
61// In input-token equivalents. Coming back to a cold cache writes the whole conversation
62// to it again. A renewal read the conversation instead, and wrote only its own tail, for
63// five minutes. After a compaction, the summary is written on the way back, and the
64// compaction itself read the conversation and wrote the summary. What was never gone back
65// to, or went cold anyway, spared nothing and paid for itself. Requests after the first
66// back read a smaller conversation too, which is left out, so the tally errs low
67export const savingOf = (record: CompactionRecord): Saving => {
68  const write = WRITE_COST[record.ttl]
69  const usage = record.usage
70  if (record.kind === 'renewal') {
71    const renewal =
72      usage === null
73        ? readCost(record.model) * record.before
74        : readCost(record.model) * usage.cache_read_input_tokens +
75          WRITE_COST['5m'] * usage.cache_creation_input_tokens +
76          usage.input_tokens +
77          OUTPUT_COST * usage.output_tokens
78    return { avoided: record.isReturned ? write * record.before : 0, spent: renewal }
79  }
80  const compaction =
81    usage === null
82      ? readCost(record.model) * record.before + OUTPUT_COST * record.after
83      : readCost(record.model) * usage.cache_read_input_tokens +
84        write * usage.cache_creation_input_tokens +
85        usage.input_tokens +
86        OUTPUT_COST * usage.output_tokens
87  if (!record.isReturned) return { avoided: 0, spent: compaction }
88  return { avoided: write * record.before, spent: compaction + write * record.after }
89}
90
91export const formatTokens = (tokens: number): string => {
92  const size = Math.abs(tokens)
93  if (size >= 1_000_000) return `${(tokens / 1_000_000).toFixed(1)}M`
94  if (size >= 1000) return `${Math.round(tokens / 1000)}k`
95  return String(Math.round(tokens))
96}
97
98const signed = (tokens: number): string => (tokens >= 0 ? `+${formatTokens(tokens)}` : `-${formatTokens(-tokens)}`)
99
100type Row = { label: string; records: CompactionRecord[] }
101
102const rowText = ({ label, records }: Row): string => {
103  const savings = records.map(savingOf)
104  const net = savings.reduce((sum, s) => sum + s.avoided - s.spent, 0)
105  const returned = records.filter(r => r.isReturned).length
106  const renewals = records.filter(r => r.kind === 'renewal').length
107  return [
108    label.padEnd(14),
109    String(records.length - renewals).padStart(11),
110    String(renewals).padStart(10),
111    String(returned).padStart(11),
112    (records.length === 0 ? '-' : signed(net)).padStart(11),
113  ].join('')
114}
115
116export const statsText = (records: readonly CompactionRecord[], now: number, sessionId: string): string => {
117  const rows: Row[] = [
118    { label: 'This session', records: records.filter(r => r.sessionId === sessionId) },
119    { label: 'Last 7 days', records: records.filter(r => now - r.at < WEEK_MS) },
120    { label: 'All time', records: [...records] },
121  ]
122  return [
123    'Savings, in input tokens (a cache write at 1h counts as 2, a cache read as 0.1 or less):',
124    '',
125    `${''.padEnd(14)}${'compactions'.padStart(11)}${'renewals'.padStart(10)}${'came back'.padStart(11)}${'net saved'.padStart(11)}`,
126    ...rows.map(rowText),
127    '',
128    'A compaction, or the last renewal before you came back, counts once its session goes on: the re-read it spared is set against what it cost. What was never gone back to counts its cost alone.',
129  ].join('\n')
130}
131
132const pad = (n: number): string => String(n).padStart(2, '0')
133
134export const clockTime = (at: number): string => {
135  const date = new Date(at)
136  return `${pad(date.getHours())}:${pad(date.getMinutes())}`
137}
138
139// The transcript's lasting line, for whoever comes back to the session
140export const noticeText = (record: CompactionRecord, hasSizes: boolean): string => {
141  const sizes = hasSizes ? ` (${formatTokens(record.before)} → ${formatTokens(record.after)} tokens)` : ''
142  if (isCacheMissed(record.usage)) {
143    return `Compacted at ${clockTime(record.at)}${sizes}, but the prompt cache had already lapsed, so this one saved nothing.`
144  }
145  return `Compacted at ${clockTime(record.at)}${sizes}, just before the prompt cache went cold. The conversation goes on from the summary.`
146}
147
types/index.d.ts 52 lines
1export type CacheTtl = '5m' | '1h'
2// The last main-thread request, which renewed the cache entry the next prompt would read.
3// `at`: when it was made, null once a compaction or a /clear changed the prefix.
4// `model`: the session's model then, the one the entry belongs to
5export type LastRequest = { at: number | null; model: string }
6// The renewals keep warm made since the last request: the latest's start, and how many
7export type Kept = { at: number; count: number }
8
9// What a compaction's own request was answered over, as the API reports it
10export type CompactionUsage = {
11  input_tokens: number
12  output_tokens: number
13  cache_read_input_tokens: number
14  cache_creation_input_tokens: number
15}
16
17// One compaction or cache renewal warm-compact made, kept across sessions to tally what
18// they saved. `kind`: absent for a compaction.
19// `before`, `after`: the conversation's size in tokens either side of a compaction, or the
20// size a renewal read (`after` 0).
21// `usage`: its request's, when the engine reported one.
22// `isReturned`: whether the session went on afterwards with what it kept, which is when the
23// re-read it spared would have been paid
24export type CompactionRecord = {
25  kind?: 'renewal'
26  at: number
27  sessionId: string
28  model: string
29  ttl: CacheTtl
30  before: number
31  after: number
32  usage: CompactionUsage | null
33  isReturned: boolean
34}
35
36declare module 'claude-code' {
37  interface PluginState {
38    'warm-compact': {
39      last: LastRequest | null
40      isTurnRunning: boolean
41      // Whether it compacts at all in this session: the footer chip and /warm-compact flip it
42      isEnabled: boolean
43      // Whether it renews the cache in compaction's place: the footer's second chip and
44      // /keep-warm flip it
45      isKeepWarm: boolean
46      kept: Kept | null
47      // When this session's latest compaction or renewal ran, until the session goes on
48      pendingRecord: number | null
49    }
50  }
51}
52