Shows how long the prompt cache stays warm and suggests compacting a large context at the right moment

Shows how long the prompt cache stays warm and suggests compacting a large context at the moments where it pays off.
Every model request re-sends the whole conversation. The prompt cache makes that cheap, but it expires a fixed time after the last request — one hour in Claude Code, five minutes under usage overage. A request on an expired cache re-reads the entire context at the cache write price, which on a large context is the most expensive single request of a session. This mod keeps that moment visible and offers to shrink the context before it happens.
| Line | Meaning |
|---|---|
cache 42m · ctx 250k | Warm for another 42 minutes; the last request carried 250k tokens |
cache cold · ctx 250k | Expired; the next request rebuilds the cache over the whole context |
cache rebuilds · ctx 90k | A compaction just ran; the next request writes the cache for the smaller context |
The line only appears when there is something to decide: from statusFromTokens of context on, or while the cache is about to expire. A small context on a cold cache shows nothing.
The timer starts over with every model request of the main conversation, including the steps Claude takes on its own between tool calls. Subagents have a cache of their own and do not count.
From largeContextTokens of context on, a band above the prompt offers Compact, Write handoff and Dismiss:
| When | Why then |
|---|---|
The cache expires in warnMinutes and you are idle | Compacting on a warm cache reads the context at the cache read price |
| A turn committed or ran tests green | A unit of work ended; nothing in flight gets lost |
A running turn is never interrupted.
Write handoff sends a prompt that asks Claude to write HANDOFF.md at the repository root: goal, state, decisions with their reasons, rejected approaches, open points and the exact next step.
When a conversation starts in a repository that has a HANDOFF.md, a toast names it and Claude is told where it lies, so asking it to continue is enough for the next session. Only the path is added to the context, not the file's content.
When you send a prompt on a cold cache with a large context, a dialog asks whether to compact first. The rebuild is paid either way; compacting first pays it on the smaller remainder.
The server reports with every response how much of the prompt came from the cache. When a request finds nothing cached although the timer still ran, the mod switches to a five-minute lifetime; when a request is served from the cache after more than five minutes, it switches back. Prompts under 20k tokens teach nothing, since they may be too short to be cached at all.
| Option | Default | Meaning |
|---|---|---|
ttlMinutes | 60 | Cache lifetime Claude Code uses |
warnMinutes | 10 | How early before expiry to suggest compacting |
statusFromTokens | 300000 | Context size from which the status line shows |
largeContextTokens | 300000 | Context size from which the mod suggests compacting |
Change them in /config, or in ~/.claude/settings.json under pluginConfigs["cache-watch"].options.
Compact runs a normal compaction. With compact-shaper loaded, its summary follows the structure of a handoff.
Remove the mod from CLAUDE_CODE_PLUGIN_DIRS. Its options in pluginConfigs can be deleted.
hooks/register.tsx 279 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register, Timer } from 'claude-code'
3
4import type { CacheReading } from '../types'
5import {
6 HANDOFF_FILE,
7 HANDOFF_PROMPT,
8 contextTokensOf,
9 formatRemaining,
10 formatTokens,
11 handoffContextOf,
12 isCold,
13 isMilestoneCommand,
14 learnTtl,
15 remainingMs,
16 statusText,
17} from './cache'
18
19const COMPACT_FIRST = 'Compact first'
20const SEND_AS_IS = 'Send as is'
21
22const INITIAL: CacheReading = { lastRequestAt: null, ttlMs: 60 * 60 * 1000, contextTokens: 0, isRebuildPending: false }
23
24const reading = atom({ plugin: 'cache-watch', key: 'reading' } as const, INITIAL)
25const suggestion = atom({ plugin: 'cache-watch', key: 'suggestion' } as const, null)
26const warnedFor = atom({ plugin: 'cache-watch', key: 'warnedFor' } as const, null)
27const isWorking = atom({ plugin: 'cache-watch', key: 'isWorking' } as const, false)
28
29// Set from the plugin's options when the module registers.
30let settings = { ttlMs: 60 * 60 * 1000, warnMs: 10 * 60 * 1000, largeContextTokens: 300_000, statusFromTokens: 300_000 }
31let hadMilestone = false
32let ticker: Timer | undefined
33
34async function refreshStatus($: EngineInterface): Promise<void> {
35 const now = await $.clock.now()
36 const visibility = { fromTokens: settings.statusFromTokens, warnMs: settings.warnMs }
37 $.ui.status(statusText(await read($, reading), now, visibility))
38}
39
40async function checkExpiring($: EngineInterface): Promise<void> {
41 if (await read($, isWorking)) {
42 return
43 }
44 const now = await $.clock.now()
45 const current = await read($, reading)
46 const remaining = remainingMs(current, now)
47 if (remaining === null || remaining === 0 || remaining > settings.warnMs || current.isRebuildPending) {
48 return
49 }
50 if (current.contextTokens < settings.largeContextTokens) {
51 return
52 }
53 if ((await read($, warnedFor)) === current.lastRequestAt) {
54 return
55 }
56 await update($, warnedFor, () => current.lastRequestAt)
57 await update($, suggestion, () => ({ kind: 'expiring', contextTokens: current.contextTokens }))
58 $.ui.toast(`Prompt cache expires in ${formatRemaining(remaining)} · context ${formatTokens(current.contextTokens)}`, {
59 timeoutMs: 10_000,
60 })
61}
62
63async function tick($: EngineInterface): Promise<void> {
64 await refreshStatus($)
65 await checkExpiring($)
66}
67
68async function markCompacted($: EngineInterface, tokensAfter: number | undefined): Promise<void> {
69 await update($, reading, current => ({
70 ...current,
71 contextTokens: tokensAfter ?? current.contextTokens,
72 isRebuildPending: true,
73 }))
74 await update($, suggestion, () => null)
75 await refreshStatus($)
76}
77
78async function compactNow($: EngineInterface): Promise<void> {
79 await update($, suggestion, () => null)
80 const result = await $.session.compact()
81 if (result.skip === undefined) {
82 await markCompacted($, result.tokensAfter)
83 }
84}
85
86// The handoff sits at the repository root, so a session started in a
87// subfolder still finds it.
88async function handoffPathOf($: EngineInterface): Promise<string | null> {
89 let directory = await $.session.cwd()
90 while (true) {
91 const candidate = `${directory}/${HANDOFF_FILE}`
92 if (await $.fs.exists(candidate)) {
93 return candidate
94 }
95 const parent = directory.slice(0, directory.lastIndexOf('/'))
96 if ((await $.fs.exists(`${directory}/.git`)) || parent.length === 0 || parent === directory) {
97 return null
98 }
99 directory = parent
100 }
101}
102
103export const register: Register = (on, options) => {
104 settings = {
105 ttlMs: Number(options.ttlMinutes ?? 60) * 60_000,
106 warnMs: Number(options.warnMinutes ?? 10) * 60_000,
107 largeContextTokens: Number(options.largeContextTokens ?? 300_000),
108 statusFromTokens: Number(options.statusFromTokens ?? 300_000),
109 }
110
111 on('session.start', async ($, e, next) => {
112 ticker?.cancel()
113 ticker = $.clock.every(15_000, () => {
114 void tick($)
115 })
116 await refreshStatus($)
117
118 return next(e)
119 })
120
121 on('prompt.context', async ($, e, next) => {
122 const result = await next(e)
123 const path = await handoffPathOf($)
124 if (path === null) {
125 return result
126 }
127 $.ui.toast(`Handoff ready: ${path}`, { timeoutMs: 10_000 })
128
129 return { ...result, blocks: [...result.blocks, { name: 'handoff', text: handoffContextOf(path) }] }
130 })
131
132 // Every model request of the main conversation reads the cache and starts
133 // its lifetime over. The lifetime counts from the request's start.
134 on('turn.step', async function* ($, e, next) {
135 if (e.agentId !== undefined) {
136 return yield* next(e)
137 }
138 const startedAt = await $.clock.now()
139 const result = yield* next(e)
140 const usage = result.usage
141 if (usage !== null) {
142 await update($, reading, current => ({
143 lastRequestAt: startedAt,
144 ttlMs: learnTtl(current, startedAt, usage, settings.ttlMs),
145 contextTokens: contextTokensOf(usage),
146 isRebuildPending: false,
147 }))
148 await refreshStatus($)
149 }
150
151 return result
152 })
153
154 on('turn.start', async ($, e, next) => {
155 hadMilestone = false
156 await update($, isWorking, () => true)
157
158 return next(e)
159 })
160
161 on('tool.call', { tool: 'Bash' }, async ($, e, next) => {
162 const ran = await next(e)
163 if (ran.deny === undefined && ran.isError !== true && typeof e.command === 'string' && isMilestoneCommand(e.command)) {
164 hadMilestone = true
165 }
166
167 return ran
168 })
169
170 on('turn.complete', async ($, e, next) => {
171 const result = await next(e)
172 await update($, isWorking, () => false)
173 const current = await read($, reading)
174 const now = await $.clock.now()
175 if (hadMilestone && current.contextTokens >= settings.largeContextTokens && !isCold(current, now)) {
176 await update($, suggestion, () => ({ kind: 'milestone', contextTokens: current.contextTokens }))
177 }
178 hadMilestone = false
179 await refreshStatus($)
180
181 return result
182 })
183
184 // On a cold cache the next request re-reads the whole context at the write
185 // price anyway. Compacting first means paying that price on the small
186 // remainder instead.
187 on('prompt.submit', async ($, e, next) => {
188 if (e.origin.kind !== 'composer' || e.turnId !== undefined || e.text.trimStart().startsWith('/')) {
189 return next(e)
190 }
191 await update($, suggestion, () => null)
192 const current = await read($, reading)
193 const now = await $.clock.now()
194 if (current.isRebuildPending || !isCold(current, now) || current.contextTokens < settings.largeContextTokens) {
195 return next(e)
196 }
197 let answer: string
198 try {
199 answer = await $.ui.ask(
200 `The prompt cache has expired and the context holds ${formatTokens(current.contextTokens)} tokens. Sending now re-reads all of it at the cache write price. Compact before sending?`,
201 { options: [COMPACT_FIRST, SEND_AS_IS], header: 'Cache' },
202 )
203 } catch {
204 answer = SEND_AS_IS
205 }
206 if (answer === COMPACT_FIRST) {
207 try {
208 await compactNow($)
209 } catch (error) {
210 $.ui.toast(`Compaction failed, sending as is: ${error instanceof Error ? error.message : String(error)}`)
211 }
212 }
213
214 return next(e)
215 })
216
217 // Catches /compact and automatic compactions. A compaction another plugin
218 // answers without passing on is not seen here; the next request's usage
219 // corrects the reading then.
220 on('session.compact', async ($, e, next) => {
221 const result = await next(e)
222 if (e.agentId === undefined && e.trigger !== 'precompute' && result.skip === undefined) {
223 await markCompacted($, result.tokensAfter)
224 }
225
226 return result
227 })
228
229 on('session.end', async ($, e, next) => {
230 if (e.reason === 'clear') {
231 await update($, reading, current => ({ ...INITIAL, ttlMs: current.ttlMs }))
232 await update($, suggestion, () => null)
233 await update($, warnedFor, () => null)
234 $.ui.status(undefined)
235 }
236
237 return next(e)
238 })
239
240 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
241 const current = await read($, suggestion)
242 if (current === null || e.props.hasSurvey || e.props.isWorking) {
243 return next(e)
244 }
245 // Another mod's band (status-band) may be drawn beneath; it stays under this
246 // one. An `engine` element means none was, and it cannot be nested.
247 const below = await next(e)
248 const beneath = below.type === 'engine' ? null : below
249 const { Box, Button, Text } = $.ui.resolve(e)
250 const context = formatTokens(current.contextTokens)
251 const message =
252 current.kind === 'expiring'
253 ? `Prompt cache expires soon · context ${context}. Compacting now keeps the next start cheap.`
254 : `Milestone reached · context ${context}. A good moment to compact.`
255
256 return (
257 <Box flexDirection="column">
258 <Text>{message}</Text>
259 <Box flexDirection="row">
260 <Button key="compact" label="Compact" hotkey="c" variant="primary" onPress={() => compactNow($)} />
261 <Text> </Text>
262 <Button
263 key="handoff"
264 label="Write handoff"
265 hotkey="h"
266 onPress={async () => {
267 await update($, suggestion, () => null)
268 await $.prompt.submit({ text: HANDOFF_PROMPT })
269 }}
270 />
271 <Text> </Text>
272 <Button key="dismiss" label="Dismiss" hotkey="d" role="dismiss" onPress={() => update($, suggestion, () => null)} />
273 </Box>
274 {beneath}
275 </Box>
276 )
277 })
278}
279hooks/cache.ts 116 lines1import type { CacheReading } from '../types'
2
3export const SHORT_TTL_MS = 5 * 60 * 1000
4
5// Below this prompt size a cache miss says nothing about the TTL: the prefix
6// may simply be too short to be cached at all.
7const MIN_LEARNING_TOKENS = 20_000
8
9export type Usage = {
10 input_tokens: number
11 output_tokens: number
12 cache_read_input_tokens: number
13 cache_creation_input_tokens: number
14}
15
16export function promptTokensOf(usage: Usage): number {
17 return usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens
18}
19
20export function contextTokensOf(usage: Usage): number {
21 return promptTokensOf(usage) + usage.output_tokens
22}
23
24export function remainingMs(reading: CacheReading, now: number): number | null {
25 if (reading.lastRequestAt === null) {
26 return null
27 }
28
29 return Math.max(0, reading.lastRequestAt + reading.ttlMs - now)
30}
31
32export function isCold(reading: CacheReading, now: number): boolean {
33 return remainingMs(reading, now) === 0
34}
35
36// The server is the only one who knows the real TTL. A request that found
37// nothing cached although the timer still ran means a shorter lifetime (usage
38// overage drops it to five minutes); one served after more than five minutes
39// means the configured lifetime holds again.
40export function learnTtl(reading: CacheReading, startedAt: number, usage: Usage, configuredTtlMs: number): number {
41 if (reading.lastRequestAt === null || reading.isRebuildPending) {
42 return reading.ttlMs
43 }
44 const prompt = promptTokensOf(usage)
45 if (prompt < MIN_LEARNING_TOKENS) {
46 return reading.ttlMs
47 }
48 const gap = startedAt - reading.lastRequestAt
49 const wasServed = usage.cache_read_input_tokens >= prompt * 0.5
50 if (!wasServed && gap > SHORT_TTL_MS && gap < reading.ttlMs) {
51 return SHORT_TTL_MS
52 }
53 if (wasServed && gap > reading.ttlMs) {
54 return configuredTtlMs
55 }
56
57 return reading.ttlMs
58}
59
60export function formatTokens(tokens: number): string {
61 if (tokens >= 1_000_000) {
62 return `${(tokens / 1_000_000).toFixed(1)}M`
63 }
64 if (tokens >= 1000) {
65 return `${Math.round(tokens / 1000)}k`
66 }
67
68 return String(tokens)
69}
70
71export function formatRemaining(ms: number): string {
72 return ms < 60_000 ? '<1m' : `${Math.floor(ms / 60_000)}m`
73}
74
75export type Visibility = { fromTokens: number; warnMs: number }
76
77// The line only earns its place when there is something to decide: a large
78// context, or a cache about to go cold.
79export function statusText(reading: CacheReading, now: number, visibility: Visibility): string | undefined {
80 const remaining = remainingMs(reading, now)
81 if (remaining === null) {
82 return undefined
83 }
84 const isExpiring = remaining > 0 && remaining <= visibility.warnMs && !reading.isRebuildPending
85 if (reading.contextTokens < visibility.fromTokens && !isExpiring) {
86 return undefined
87 }
88 const context = `ctx ${formatTokens(reading.contextTokens)}`
89 if (reading.isRebuildPending) {
90 return `cache rebuilds · ${context}`
91 }
92 if (remaining === 0) {
93 return `cache cold · ${context}`
94 }
95
96 return `cache ${formatRemaining(remaining)} · ${context}`
97}
98
99const COMMIT = /\bgit\s+(?:-C\s+\S+\s+)?commit\b/
100const TESTS = /\b(?:dotnet\s+test|(?:npm|pnpm|yarn|bun)\s+(?:run\s+)?test|pytest|go\s+test|cargo\s+test|vitest|jest|mvn\s+test|gradle\s+test|claude\s+plugin\s+test)\b/
101
102// A turn that committed or ran tests green ends a unit of work: compacting
103// there loses nothing that is still in flight.
104export function isMilestoneCommand(command: string): boolean {
105 return COMMIT.test(command) || TESTS.test(command)
106}
107
108export const HANDOFF_FILE = 'HANDOFF.md'
109
110export const HANDOFF_PROMPT =
111 'Write HANDOFF.md at the repository root so a new session can continue from it alone: date and branch, the goal, the current state, decisions and their reasons, approaches that were rejected and why, open points, the exact next step and which files to read first. Replace what an older handoff says. Leave out anything that can be re-read from the code.'
112
113export function handoffContextOf(path: string): string {
114 return `The previous session left a handoff in ${path}. When the person asks to continue without saying with what, read it first and check that branch and recent commits still match it.`
115}
116types/index.d.ts 20 lines1export type CacheReading = {
2 lastRequestAt: number | null
3 ttlMs: number
4 contextTokens: number
5 isRebuildPending: boolean
6}
7
8export type Suggestion = { kind: 'expiring' | 'milestone'; contextTokens: number } | null
9
10declare module 'claude-code' {
11 interface PluginState {
12 'cache-watch': {
13 reading: CacheReading
14 suggestion: Suggestion
15 warnedFor: number | null
16 isWorking: boolean
17 }
18 }
19}
20