SLOPSHOPPER

cache-watch

Shows how long the prompt cache stays warm and suggests compacting a large context at the right moment

newbandguardtoaststatusprompt
★ 2v0.1.0MITupdated 2026-10-04MichaelP17/claude-mods/cache-watch
A shopper browsing a rack in a slop shop
README

cache-watch

Shows how long the prompt cache stays warm and suggests compacting a large context at the moments where it pays off.

Every model request re-sends the whole conversation. The prompt cache makes that cheap, but it expires a fixed time after the last request — one hour in Claude Code, five minutes under usage overage. A request on an expired cache re-reads the entire context at the cache write price, which on a large context is the most expensive single request of a session. This mod keeps that moment visible and offers to shrink the context before it happens.

Status line

LineMeaning
cache 42m · ctx 250kWarm for another 42 minutes; the last request carried 250k tokens
cache cold · ctx 250kExpired; the next request rebuilds the cache over the whole context
cache rebuilds · ctx 90kA compaction just ran; the next request writes the cache for the smaller context

The line only appears when there is something to decide: from statusFromTokens of context on, or while the cache is about to expire. A small context on a cold cache shows nothing.

The timer starts over with every model request of the main conversation, including the steps Claude takes on its own between tool calls. Subagents have a cache of their own and do not count.

Suggestions

From largeContextTokens of context on, a band above the prompt offers Compact, Write handoff and Dismiss:

WhenWhy then
The cache expires in warnMinutes and you are idleCompacting on a warm cache reads the context at the cache read price
A turn committed or ran tests greenA unit of work ended; nothing in flight gets lost

A running turn is never interrupted.

Write handoff sends a prompt that asks Claude to write HANDOFF.md at the repository root: goal, state, decisions with their reasons, rejected approaches, open points and the exact next step.

When a conversation starts in a repository that has a HANDOFF.md, a toast names it and Claude is told where it lies, so asking it to continue is enough for the next session. Only the path is added to the context, not the file's content.

When you send a prompt on a cold cache with a large context, a dialog asks whether to compact first. The rebuild is paid either way; compacting first pays it on the smaller remainder.

Lifetime learned from the server

The server reports with every response how much of the prompt came from the cache. When a request finds nothing cached although the timer still ran, the mod switches to a five-minute lifetime; when a request is served from the cache after more than five minutes, it switches back. Prompts under 20k tokens teach nothing, since they may be too short to be cached at all.

Configuration

OptionDefaultMeaning
ttlMinutes60Cache lifetime Claude Code uses
warnMinutes10How early before expiry to suggest compacting
statusFromTokens300000Context size from which the status line shows
largeContextTokens300000Context size from which the mod suggests compacting

Change them in /config, or in ~/.claude/settings.json under pluginConfigs["cache-watch"].options.

Works with compact-shaper

Compact runs a normal compaction. With compact-shaper loaded, its summary follows the structure of a handoff.

Uninstall

Remove the mod from CLAUDE_CODE_PLUGIN_DIRS. Its options in pluginConfigs can be deleted.

Source 3 files
hooks/register.tsx 279 lines
1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register, Timer } from 'claude-code'
3
4import type { CacheReading } from '../types'
5import {
6  HANDOFF_FILE,
7  HANDOFF_PROMPT,
8  contextTokensOf,
9  formatRemaining,
10  formatTokens,
11  handoffContextOf,
12  isCold,
13  isMilestoneCommand,
14  learnTtl,
15  remainingMs,
16  statusText,
17} from './cache'
18
19const COMPACT_FIRST = 'Compact first'
20const SEND_AS_IS = 'Send as is'
21
22const INITIAL: CacheReading = { lastRequestAt: null, ttlMs: 60 * 60 * 1000, contextTokens: 0, isRebuildPending: false }
23
24const reading = atom({ plugin: 'cache-watch', key: 'reading' } as const, INITIAL)
25const suggestion = atom({ plugin: 'cache-watch', key: 'suggestion' } as const, null)
26const warnedFor = atom({ plugin: 'cache-watch', key: 'warnedFor' } as const, null)
27const isWorking = atom({ plugin: 'cache-watch', key: 'isWorking' } as const, false)
28
29// Set from the plugin's options when the module registers.
30let settings = { ttlMs: 60 * 60 * 1000, warnMs: 10 * 60 * 1000, largeContextTokens: 300_000, statusFromTokens: 300_000 }
31let hadMilestone = false
32let ticker: Timer | undefined
33
34async function refreshStatus($: EngineInterface): Promise<void> {
35  const now = await $.clock.now()
36  const visibility = { fromTokens: settings.statusFromTokens, warnMs: settings.warnMs }
37  $.ui.status(statusText(await read($, reading), now, visibility))
38}
39
40async function checkExpiring($: EngineInterface): Promise<void> {
41  if (await read($, isWorking)) {
42    return
43  }
44  const now = await $.clock.now()
45  const current = await read($, reading)
46  const remaining = remainingMs(current, now)
47  if (remaining === null || remaining === 0 || remaining > settings.warnMs || current.isRebuildPending) {
48    return
49  }
50  if (current.contextTokens < settings.largeContextTokens) {
51    return
52  }
53  if ((await read($, warnedFor)) === current.lastRequestAt) {
54    return
55  }
56  await update($, warnedFor, () => current.lastRequestAt)
57  await update($, suggestion, () => ({ kind: 'expiring', contextTokens: current.contextTokens }))
58  $.ui.toast(`Prompt cache expires in ${formatRemaining(remaining)} · context ${formatTokens(current.contextTokens)}`, {
59    timeoutMs: 10_000,
60  })
61}
62
63async function tick($: EngineInterface): Promise<void> {
64  await refreshStatus($)
65  await checkExpiring($)
66}
67
68async function markCompacted($: EngineInterface, tokensAfter: number | undefined): Promise<void> {
69  await update($, reading, current => ({
70    ...current,
71    contextTokens: tokensAfter ?? current.contextTokens,
72    isRebuildPending: true,
73  }))
74  await update($, suggestion, () => null)
75  await refreshStatus($)
76}
77
78async function compactNow($: EngineInterface): Promise<void> {
79  await update($, suggestion, () => null)
80  const result = await $.session.compact()
81  if (result.skip === undefined) {
82    await markCompacted($, result.tokensAfter)
83  }
84}
85
86// The handoff sits at the repository root, so a session started in a
87// subfolder still finds it.
88async function handoffPathOf($: EngineInterface): Promise<string | null> {
89  let directory = await $.session.cwd()
90  while (true) {
91    const candidate = `${directory}/${HANDOFF_FILE}`
92    if (await $.fs.exists(candidate)) {
93      return candidate
94    }
95    const parent = directory.slice(0, directory.lastIndexOf('/'))
96    if ((await $.fs.exists(`${directory}/.git`)) || parent.length === 0 || parent === directory) {
97      return null
98    }
99    directory = parent
100  }
101}
102
103export const register: Register = (on, options) => {
104  settings = {
105    ttlMs: Number(options.ttlMinutes ?? 60) * 60_000,
106    warnMs: Number(options.warnMinutes ?? 10) * 60_000,
107    largeContextTokens: Number(options.largeContextTokens ?? 300_000),
108    statusFromTokens: Number(options.statusFromTokens ?? 300_000),
109  }
110
111  on('session.start', async ($, e, next) => {
112    ticker?.cancel()
113    ticker = $.clock.every(15_000, () => {
114      void tick($)
115    })
116    await refreshStatus($)
117
118    return next(e)
119  })
120
121  on('prompt.context', async ($, e, next) => {
122    const result = await next(e)
123    const path = await handoffPathOf($)
124    if (path === null) {
125      return result
126    }
127    $.ui.toast(`Handoff ready: ${path}`, { timeoutMs: 10_000 })
128
129    return { ...result, blocks: [...result.blocks, { name: 'handoff', text: handoffContextOf(path) }] }
130  })
131
132  // Every model request of the main conversation reads the cache and starts
133  // its lifetime over. The lifetime counts from the request's start.
134  on('turn.step', async function* ($, e, next) {
135    if (e.agentId !== undefined) {
136      return yield* next(e)
137    }
138    const startedAt = await $.clock.now()
139    const result = yield* next(e)
140    const usage = result.usage
141    if (usage !== null) {
142      await update($, reading, current => ({
143        lastRequestAt: startedAt,
144        ttlMs: learnTtl(current, startedAt, usage, settings.ttlMs),
145        contextTokens: contextTokensOf(usage),
146        isRebuildPending: false,
147      }))
148      await refreshStatus($)
149    }
150
151    return result
152  })
153
154  on('turn.start', async ($, e, next) => {
155    hadMilestone = false
156    await update($, isWorking, () => true)
157
158    return next(e)
159  })
160
161  on('tool.call', { tool: 'Bash' }, async ($, e, next) => {
162    const ran = await next(e)
163    if (ran.deny === undefined && ran.isError !== true && typeof e.command === 'string' && isMilestoneCommand(e.command)) {
164      hadMilestone = true
165    }
166
167    return ran
168  })
169
170  on('turn.complete', async ($, e, next) => {
171    const result = await next(e)
172    await update($, isWorking, () => false)
173    const current = await read($, reading)
174    const now = await $.clock.now()
175    if (hadMilestone && current.contextTokens >= settings.largeContextTokens && !isCold(current, now)) {
176      await update($, suggestion, () => ({ kind: 'milestone', contextTokens: current.contextTokens }))
177    }
178    hadMilestone = false
179    await refreshStatus($)
180
181    return result
182  })
183
184  // On a cold cache the next request re-reads the whole context at the write
185  // price anyway. Compacting first means paying that price on the small
186  // remainder instead.
187  on('prompt.submit', async ($, e, next) => {
188    if (e.origin.kind !== 'composer' || e.turnId !== undefined || e.text.trimStart().startsWith('/')) {
189      return next(e)
190    }
191    await update($, suggestion, () => null)
192    const current = await read($, reading)
193    const now = await $.clock.now()
194    if (current.isRebuildPending || !isCold(current, now) || current.contextTokens < settings.largeContextTokens) {
195      return next(e)
196    }
197    let answer: string
198    try {
199      answer = await $.ui.ask(
200        `The prompt cache has expired and the context holds ${formatTokens(current.contextTokens)} tokens. Sending now re-reads all of it at the cache write price. Compact before sending?`,
201        { options: [COMPACT_FIRST, SEND_AS_IS], header: 'Cache' },
202      )
203    } catch {
204      answer = SEND_AS_IS
205    }
206    if (answer === COMPACT_FIRST) {
207      try {
208        await compactNow($)
209      } catch (error) {
210        $.ui.toast(`Compaction failed, sending as is: ${error instanceof Error ? error.message : String(error)}`)
211      }
212    }
213
214    return next(e)
215  })
216
217  // Catches /compact and automatic compactions. A compaction another plugin
218  // answers without passing on is not seen here; the next request's usage
219  // corrects the reading then.
220  on('session.compact', async ($, e, next) => {
221    const result = await next(e)
222    if (e.agentId === undefined && e.trigger !== 'precompute' && result.skip === undefined) {
223      await markCompacted($, result.tokensAfter)
224    }
225
226    return result
227  })
228
229  on('session.end', async ($, e, next) => {
230    if (e.reason === 'clear') {
231      await update($, reading, current => ({ ...INITIAL, ttlMs: current.ttlMs }))
232      await update($, suggestion, () => null)
233      await update($, warnedFor, () => null)
234      $.ui.status(undefined)
235    }
236
237    return next(e)
238  })
239
240  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
241    const current = await read($, suggestion)
242    if (current === null || e.props.hasSurvey || e.props.isWorking) {
243      return next(e)
244    }
245    // Another mod's band (status-band) may be drawn beneath; it stays under this
246    // one. An `engine` element means none was, and it cannot be nested.
247    const below = await next(e)
248    const beneath = below.type === 'engine' ? null : below
249    const { Box, Button, Text } = $.ui.resolve(e)
250    const context = formatTokens(current.contextTokens)
251    const message =
252      current.kind === 'expiring'
253        ? `Prompt cache expires soon · context ${context}. Compacting now keeps the next start cheap.`
254        : `Milestone reached · context ${context}. A good moment to compact.`
255
256    return (
257      <Box flexDirection="column">
258        <Text>{message}</Text>
259        <Box flexDirection="row">
260          <Button key="compact" label="Compact" hotkey="c" variant="primary" onPress={() => compactNow($)} />
261          <Text> </Text>
262          <Button
263            key="handoff"
264            label="Write handoff"
265            hotkey="h"
266            onPress={async () => {
267              await update($, suggestion, () => null)
268              await $.prompt.submit({ text: HANDOFF_PROMPT })
269            }}
270          />
271          <Text> </Text>
272          <Button key="dismiss" label="Dismiss" hotkey="d" role="dismiss" onPress={() => update($, suggestion, () => null)} />
273        </Box>
274        {beneath}
275      </Box>
276    )
277  })
278}
279
hooks/cache.ts 116 lines
1import type { CacheReading } from '../types'
2
3export const SHORT_TTL_MS = 5 * 60 * 1000
4
5// Below this prompt size a cache miss says nothing about the TTL: the prefix
6// may simply be too short to be cached at all.
7const MIN_LEARNING_TOKENS = 20_000
8
9export type Usage = {
10  input_tokens: number
11  output_tokens: number
12  cache_read_input_tokens: number
13  cache_creation_input_tokens: number
14}
15
16export function promptTokensOf(usage: Usage): number {
17  return usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens
18}
19
20export function contextTokensOf(usage: Usage): number {
21  return promptTokensOf(usage) + usage.output_tokens
22}
23
24export function remainingMs(reading: CacheReading, now: number): number | null {
25  if (reading.lastRequestAt === null) {
26    return null
27  }
28
29  return Math.max(0, reading.lastRequestAt + reading.ttlMs - now)
30}
31
32export function isCold(reading: CacheReading, now: number): boolean {
33  return remainingMs(reading, now) === 0
34}
35
36// The server is the only one who knows the real TTL. A request that found
37// nothing cached although the timer still ran means a shorter lifetime (usage
38// overage drops it to five minutes); one served after more than five minutes
39// means the configured lifetime holds again.
40export function learnTtl(reading: CacheReading, startedAt: number, usage: Usage, configuredTtlMs: number): number {
41  if (reading.lastRequestAt === null || reading.isRebuildPending) {
42    return reading.ttlMs
43  }
44  const prompt = promptTokensOf(usage)
45  if (prompt < MIN_LEARNING_TOKENS) {
46    return reading.ttlMs
47  }
48  const gap = startedAt - reading.lastRequestAt
49  const wasServed = usage.cache_read_input_tokens >= prompt * 0.5
50  if (!wasServed && gap > SHORT_TTL_MS && gap < reading.ttlMs) {
51    return SHORT_TTL_MS
52  }
53  if (wasServed && gap > reading.ttlMs) {
54    return configuredTtlMs
55  }
56
57  return reading.ttlMs
58}
59
60export function formatTokens(tokens: number): string {
61  if (tokens >= 1_000_000) {
62    return `${(tokens / 1_000_000).toFixed(1)}M`
63  }
64  if (tokens >= 1000) {
65    return `${Math.round(tokens / 1000)}k`
66  }
67
68  return String(tokens)
69}
70
71export function formatRemaining(ms: number): string {
72  return ms < 60_000 ? '<1m' : `${Math.floor(ms / 60_000)}m`
73}
74
75export type Visibility = { fromTokens: number; warnMs: number }
76
77// The line only earns its place when there is something to decide: a large
78// context, or a cache about to go cold.
79export function statusText(reading: CacheReading, now: number, visibility: Visibility): string | undefined {
80  const remaining = remainingMs(reading, now)
81  if (remaining === null) {
82    return undefined
83  }
84  const isExpiring = remaining > 0 && remaining <= visibility.warnMs && !reading.isRebuildPending
85  if (reading.contextTokens < visibility.fromTokens && !isExpiring) {
86    return undefined
87  }
88  const context = `ctx ${formatTokens(reading.contextTokens)}`
89  if (reading.isRebuildPending) {
90    return `cache rebuilds · ${context}`
91  }
92  if (remaining === 0) {
93    return `cache cold · ${context}`
94  }
95
96  return `cache ${formatRemaining(remaining)} · ${context}`
97}
98
99const COMMIT = /\bgit\s+(?:-C\s+\S+\s+)?commit\b/
100const TESTS = /\b(?:dotnet\s+test|(?:npm|pnpm|yarn|bun)\s+(?:run\s+)?test|pytest|go\s+test|cargo\s+test|vitest|jest|mvn\s+test|gradle\s+test|claude\s+plugin\s+test)\b/
101
102// A turn that committed or ran tests green ends a unit of work: compacting
103// there loses nothing that is still in flight.
104export function isMilestoneCommand(command: string): boolean {
105  return COMMIT.test(command) || TESTS.test(command)
106}
107
108export const HANDOFF_FILE = 'HANDOFF.md'
109
110export const HANDOFF_PROMPT =
111  'Write HANDOFF.md at the repository root so a new session can continue from it alone: date and branch, the goal, the current state, decisions and their reasons, approaches that were rejected and why, open points, the exact next step and which files to read first. Replace what an older handoff says. Leave out anything that can be re-read from the code.'
112
113export function handoffContextOf(path: string): string {
114  return `The previous session left a handoff in ${path}. When the person asks to continue without saying with what, read it first and check that branch and recent commits still match it.`
115}
116
types/index.d.ts 20 lines
1export type CacheReading = {
2  lastRequestAt: number | null
3  ttlMs: number
4  contextTokens: number
5  isRebuildPending: boolean
6}
7
8export type Suggestion = { kind: 'expiring' | 'milestone'; contextTokens: number } | null
9
10declare module 'claude-code' {
11  interface PluginState {
12    'cache-watch': {
13      reading: CacheReading
14      suggestion: Suggestion
15      warnedFor: number | null
16      isWorking: boolean
17    }
18  }
19}
20