Live countdown above the prompt until your session's prompt cache expires, with an optional keep-alive that keeps it warm while you're away

A Claude Code mod that shows a live countdown until your session's prompt cache expires, and can keep the cache warm while you're away, so coming back to a long session doesn't re-cache the whole context from scratch.
⏳ cache 52:13 (TTL 1h) · ⟳ keep-alive 3h 41m · saved ~380k tok
Claude Code caches your conversation prefix on Anthropic's side. While the cache is warm, every new message reads the context from it cheaply. Once it expires (1 hour on a Claude subscription, 5 minutes on an API key), the next message writes the whole context to the cache again, and that write costs much more than a read.
| Price vs. one uncached input token | |
|---|---|
| Cache read | 0.05× (Opus 5.5), 0.1× (most models) |
| Cache write, 5-minute TTL | 1.25× |
| Cache write, 1-hour TTL | 2× |
With a 400k-token session that's the difference between ~20k and ~800k input-token equivalents for the first message after a break.
In a Claude Code terminal session:
/plugin install cache-timer --marketplace MarekBartczak/claude-cache-timer
Answer y to add the marketplace, pick a scope (user scope makes it load in every session), and confirm the options. The band appears above the prompt right away.
Requires a Claude Code build with function-hook plugins (mods). Developed and tested on 2.1.292. The mod API is early access and may change between releases.
A band above the prompt:
| Band | Meaning |
|---|---|
⏳ cache 52:13 (TTL 1h) | Time left until the cache expires, counted from the last request of the main conversation |
yellow ⏳ cache 4:59 | Less than 5 minutes left (1 minute on a 5m TTL), plus a one-time toast. Not shown while keep-alive covers it |
· ⟳ keep-alive 3h 41m | Keep-alive is active and will keep refreshing the cache for this long |
· saved ~380k tok (Σ 2.1M) | Net savings this session, and the running total across sessions |
❄ cache cold / ❄ cache expired 3 min ago | Nothing cached yet, or the cache is gone |
Subagent requests don't move the countdown, since they cache their own prefixes.
About 2 minutes before the cache would expire (1 minute on a 5m TTL), while no turn is running, the mod re-sends the conversation's last request with a one-word prompt through $.model.fork. That request reads the cache, which resets its timer. Nothing is added to your transcript or context.
/cache-ping refreshes the cache on demand and tells you how many tokens it read from the cache. Use it once to confirm keep-alive works on your setup.On Opus 5.5 an hourly ping costs about 1/40 of a re-cache, so keep-alive pays off if you come back within roughly 40 hours. The window guards against sessions you never return to: there, every ping is pure cost.
Counted in uncached-input-token equivalents, using API price ratios:
So the number can go negative when pings didn't pay off. On a subscription you're billed in usage limits, not tokens. Anthropic doesn't publish the exact weighting, so treat the figure as an estimate of direction and scale.
In /config:
| Setting | Values | Default |
|---|---|---|
| Prompt cache TTL | 1h, 5m | 1h |
| Cache keep-alive | off, 1h, 2h, 4h, 8h | 4h |
The engine doesn't expose the TTL to mods, so the mod works it out in this order:
/model).CLAUDE_CODE_PROMPT_CACHE_TTL, if set.git clone https://github.com/MarekBartczak/claude-cache-timer
cd claude-cache-timer
claude plugin validate .
claude plugin test .
claude --plugin-dir . # run a session with your working copy
To load your working copy in every session, add its path to CLAUDE_CODE_PLUGIN_DIRS in the env block of ~/.claude/settings.json. In an interactive session, saving a file reloads the mod.
hooks/register.tsx 315 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, ModelUsage, Register } from 'claude-code'
3
4import type { CacheTtl, Line, Tone } from '../types'
5
6const MINUTE = 60_000
7const HOUR = 60 * MINUTE
8const TTL_MS: Record<CacheTtl, number> = { '5m': 5 * MINUTE, '1h': HOUR }
9const KEEP_ALIVE_MS: Record<string, number> = { off: 0, '1h': HOUR, '2h': 2 * HOUR, '4h': 4 * HOUR, '8h': 8 * HOUR }
10const PING_PROMPT = 'Keep-alive ping for the prompt cache. Reply with the single word: ok'
11
12// API price ratios against one uncached input token; savings are counted in these units.
13const WRITE_WEIGHT: Record<CacheTtl, number> = { '5m': 1.25, '1h': 2 }
14const OUTPUT_WEIGHT = 5
15const readWeight = (model: string): number =>
16 /fable-5-1|mythos-5-1/.test(model) ? 0.025 : /opus-5-5/.test(model) ? 0.05 : 0.1
17
18const lastHitAt = atom({ plugin: 'cache-timer', key: 'lastHitAt' } as const, null)
19const lastModel = atom({ plugin: 'cache-timer', key: 'lastModel' } as const, null)
20const learnedTtl = atom({ plugin: 'cache-timer', key: 'learnedTtl' } as const, null)
21const warnedFor = atom({ plugin: 'cache-timer', key: 'warnedFor' } as const, null)
22const lastActivityAt = atom({ plugin: 'cache-timer', key: 'lastActivityAt' } as const, null)
23const lastRealHitAt = atom({ plugin: 'cache-timer', key: 'lastRealHitAt' } as const, null)
24const pingsSinceReal = atom({ plugin: 'cache-timer', key: 'pingsSinceReal' } as const, 0)
25const savedSession = atom({ plugin: 'cache-timer', key: 'savedSession' } as const, 0)
26const line = atom({ plugin: 'cache-timer', key: 'line' } as const, null)
27
28const COLORS: Record<Tone, string> = { ok: 'subtle', warn: 'warning', cold: 'inactive' }
29
30type Config = { assumedTtl: CacheTtl; keepAliveMs: number }
31
32const formatLeft = (ms: number): string => {
33 const seconds = Math.max(0, Math.ceil(ms / 1000))
34
35 return `${Math.floor(seconds / 60)}:${String(seconds % 60).padStart(2, '0')}`
36}
37
38const formatSpan = (ms: number): string => {
39 const minutes = Math.max(0, Math.floor(ms / MINUTE))
40
41 return minutes >= 60 ? `${Math.floor(minutes / 60)}h ${minutes % 60}m` : `${minutes}m`
42}
43
44const formatTokens = (n: number): string => {
45 const abs = Math.abs(n)
46 const text =
47 abs >= 1_000_000 ? `${(abs / 1_000_000).toFixed(1)}M` : abs >= 1000 ? `${(abs / 1000).toFixed(abs < 10_000 ? 1 : 0)}k` : `${Math.round(abs)}`
48
49 return n < 0 ? `-${text}` : text
50}
51
52const isTtl = (value: unknown): value is CacheTtl => value === '5m' || value === '1h'
53
54let shown: string | undefined
55let timer: { cancel: () => void } | undefined
56let isWorking = false
57let isPinging = false
58let pingedFor: number | undefined
59let savedTotal = 0
60
61const show = async ($: EngineInterface, text: string, tone: Tone) => {
62 if (text !== shown) {
63 shown = text
64 const next: Line = { text, tone }
65 await update($, line, () => next)
66 }
67}
68
69const credit = async ($: EngineInterface, delta: number) => {
70 await update($, savedSession, n => n + delta)
71 savedTotal += delta
72 await $.store.set('savedTotal', savedTotal)
73}
74
75// What a request cost, in uncached-input-token units.
76const costOf = (model: string, ttl: CacheTtl, usage: ModelUsage): number =>
77 usage.input_tokens +
78 usage.cache_read_input_tokens * readWeight(model) +
79 usage.cache_creation_input_tokens * WRITE_WEIGHT[ttl] +
80 usage.output_tokens * OUTPUT_WEIGHT
81
82// Re-sends the main thread's last request (nothing joins the transcript): a cache read
83// refreshes the entry's timer for a fraction of what re-writing it costs.
84const ping = async ($: EngineInterface, ttl: CacheTtl): Promise<string> => {
85 const at = await read($, lastHitAt)
86 isPinging = true
87 pingedFor = at ?? undefined
88
89 try {
90 const sentAt = await $.clock.now()
91 const reply = await $.model.fork({ prompt: PING_PROMPT })
92
93 if (!reply.isAnswered) {
94 return `Keep-alive failed: ${reply.reason}`
95 }
96
97 const cached = reply.usage.cache_read_input_tokens
98 await credit($, -costOf((await read($, lastModel)) ?? '', ttl, reply.usage))
99
100 if (cached === 0) {
101 return 'Keep-alive found the cache already gone; not refreshing further'
102 }
103
104 await update($, lastHitAt, () => sentAt)
105 await update($, pingsSinceReal, n => n + 1)
106
107 return `⟳ Cache refreshed (${Math.round(cached / 1000)}k tokens read from cache)`
108 } finally {
109 isPinging = false
110 }
111}
112
113const savedSuffix = async ($: EngineInterface): Promise<string> => {
114 const session = await read($, savedSession)
115
116 if (session === 0 && savedTotal === 0) {
117 return ''
118 }
119
120 const total = savedTotal !== session ? ` (Σ ${formatTokens(savedTotal)})` : ''
121
122 return ` · saved ~${formatTokens(session)} tok${total}`
123}
124
125const tick = async ($: EngineInterface, config: Config) => {
126 const at = await read($, lastHitAt)
127 const saved = await savedSuffix($)
128
129 if (at === null) {
130 await show($, `❄ cache cold${saved}`, 'cold')
131
132 return
133 }
134
135 const ttl = (await read($, learnedTtl)) ?? config.assumedTtl
136 const now = await $.clock.now()
137 const left = at + TTL_MS[ttl] - now
138
139 if (left <= 0) {
140 await show($, `❄ cache expired ${Math.floor(-left / MINUTE)} min ago (TTL ${ttl})${saved}`, 'cold')
141
142 return
143 }
144
145 const activeUntil = ((await read($, lastActivityAt)) ?? now) + config.keepAliveMs
146 const isKeptAlive = config.keepAliveMs > 0 && now < activeUntil
147 const lead = Math.min(2 * MINUTE, TTL_MS[ttl] / 5)
148
149 if (isKeptAlive && left <= lead && !isWorking && !isPinging && pingedFor !== at) {
150 void ping($, ttl).then(async text => {
151 $.ui.toast(text, { timeoutMs: 6000 })
152 await tick($, config)
153 })
154 }
155
156 const warnAt = Math.min(5 * MINUTE, TTL_MS[ttl] / 5)
157 const isWarning = left <= warnAt && !isKeptAlive
158
159 if (isWarning && (await read($, warnedFor)) !== at) {
160 await update($, warnedFor, () => at)
161 $.ui.toast(`Session cache expires in ${formatLeft(left)}`, { timeoutMs: 8000 })
162 }
163
164 const suffix = isKeptAlive ? ` · ⟳ keep-alive ${formatSpan(activeUntil - now)}` : ''
165 await show($, `⏳ cache ${formatLeft(left)} (TTL ${ttl})${suffix}${saved}`, isWarning ? 'warn' : 'ok')
166}
167
168// The first real request after the cache would have lapsed without keep-alive pings: what it
169// read from the cache would otherwise have been written again.
170const creditRescue = async ($: EngineInterface, config: Config, sentAt: number, model: string, usage: ModelUsage) => {
171 const realAt = await read($, lastRealHitAt)
172 const ttl = (await read($, learnedTtl)) ?? config.assumedTtl
173 const isRescue = realAt !== null && (await read($, pingsSinceReal)) > 0 && sentAt > realAt + TTL_MS[ttl]
174
175 if (isRescue && usage.cache_read_input_tokens > 0) {
176 await credit($, usage.cache_read_input_tokens * (WRITE_WEIGHT[ttl] - readWeight(model)))
177 }
178
179 await update($, lastRealHitAt, () => sentAt)
180 await update($, pingsSinceReal, () => 0)
181}
182
183// A request after a gap between 5m and 1h tells which TTL the cache really has.
184const learn = async ($: EngineInterface, sentAt: number, model: string, usage: ModelUsage) => {
185 const prev = await read($, lastHitAt)
186
187 if (prev === null || (await read($, lastModel)) !== model) {
188 return
189 }
190
191 const gap = sentAt - prev
192 const isTelling = gap > TTL_MS['5m'] + 30_000 && gap < TTL_MS['1h'] - MINUTE
193
194 if (isTelling && usage.cache_read_input_tokens > 0) {
195 await update($, learnedTtl, () => '1h' as const)
196 } else if (isTelling && usage.cache_creation_input_tokens > 0) {
197 await update($, learnedTtl, () => '5m' as const)
198 }
199}
200
201export const register: Register = (on, options) => {
202 const config: Config = {
203 assumedTtl: isTtl(options.ttl) ? options.ttl : '1h',
204 keepAliveMs: KEEP_ALIVE_MS[String(options.keepAlive)] ?? 0,
205 }
206
207 on('session.start', async ($, e, next) => {
208 const fromEnv = await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL')
209 config.assumedTtl = isTtl(fromEnv) ? fromEnv : config.assumedTtl
210 shown = undefined
211 savedTotal = Number((await $.store.get('savedTotal')) ?? 0) || 0
212 $.ui.status(undefined)
213
214 if ((await read($, lastActivityAt)) === null) {
215 const now = await $.clock.now()
216 await update($, lastActivityAt, () => now)
217 }
218
219 await $.command.register({ name: 'cache-ping', description: 'Refresh the session prompt cache now (keep-alive)' })
220 timer?.cancel()
221 timer = $.clock.every(1000, () => void tick($, config))
222 await tick($, config)
223
224 return next(e)
225 })
226
227 on('command.run', { command: 'cache-ping' }, async $ => {
228 const at = await read($, lastHitAt)
229 const ttl = (await read($, learnedTtl)) ?? config.assumedTtl
230
231 if (at === null || at + TTL_MS[ttl] <= (await $.clock.now())) {
232 return { text: 'No warm cache to refresh.' }
233 }
234
235 const text = await ping($, ttl)
236 await tick($, config)
237
238 return { text }
239 })
240
241 // What the person does counts as activity; the keep-alive window runs from it.
242 on('prompt.submit', async ($, e, next) => {
243 const now = await $.clock.now()
244 await update($, lastActivityAt, () => now)
245
246 return next(e)
247 }).catch(($, e, next) => next(e))
248
249 on('turn.start', async ($, e, next) => {
250 isWorking = true
251
252 return next(e)
253 })
254
255 on('turn.complete', async ($, e, next) => {
256 isWorking = false
257
258 return next(e)
259 })
260
261 // Only the main thread's requests: subagents cache their own prefixes.
262 on('turn.step', async function* ($, e, next) {
263 if (e.agentId !== undefined) {
264 return yield* next(e)
265 }
266
267 const sentAt = await $.clock.now()
268 const response = yield* next(e)
269
270 if (response.usage !== null) {
271 await creditRescue($, config, sentAt, e.model, response.usage)
272 await learn($, sentAt, e.model, response.usage)
273 await update($, lastHitAt, () => sentAt)
274 await update($, lastModel, () => e.model)
275 await tick($, config)
276 }
277
278 return response
279 })
280
281 // The engine knows the TTL here; the new model's cache starts cold.
282 on('classic.PostModelSwitch', async ($, e, next) => {
283 await update($, learnedTtl, () => e.cache_ttl)
284 await update($, lastHitAt, () => null)
285 await update($, lastModel, () => e.to_model)
286 await tick($, config)
287
288 return next(e)
289 }).catch(($, e, next) => next(e))
290
291 on('classic.SessionStart', async ($, e, next) => {
292 if (e.seconds_since_last_response !== undefined) {
293 const at = (await $.clock.now()) - e.seconds_since_last_response * 1000
294 await update($, lastHitAt, () => at)
295 await update($, lastModel, () => e.model ?? null)
296 await tick($, config)
297 }
298
299 return next(e)
300 }).catch(($, e, next) => next(e))
301
302 // Our own band above the prompt: the status line would prefix it with a warning icon.
303 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
304 const current = await read($, line)
305
306 if (e.props.hasSurvey || current === null) {
307 return next(e)
308 }
309
310 const { Text } = $.ui.resolve(e)
311
312 return <Text color={COLORS[current.tone]}>{current.text}</Text>
313 })
314}
315types/index.d.ts 20 lines1export type CacheTtl = '5m' | '1h'
2export type Tone = 'ok' | 'warn' | 'cold'
3export type Line = { text: string; tone: Tone }
4
5declare module 'claude-code' {
6 interface PluginState {
7 'cache-timer': {
8 lastHitAt: number | null
9 lastModel: string | null
10 learnedTtl: CacheTtl | null
11 warnedFor: number | null
12 lastActivityAt: number | null
13 lastRealHitAt: number | null
14 pingsSinceReal: number
15 savedSession: number
16 line: Line | null
17 }
18 }
19}
20