Prompt-cache indicator plus an opt-in, bounded keepalive (/warm) that keeps an idle Claude Code session's cache warm, so you don't pay to re-cache the whole…

See whether your prompt cache is warm, and keep it warm while you step away. That way the first message after a break doesn't re-send and re-cache the whole conversation.
<img src="docs/indicator.svg" alt="cache ● 1h ████░░ 38m left · hit 91% · keepalive 1h42m ↻2 — cache ○ cold · next message re-caches ~240k tokens">
| Part | What it does | Setup |
|---|---|---|
Plugin (plugins/cache-warmer) | A cache indicator under the prompt, plus /warm, a bounded keepalive. | Install the plugin; nothing else. |
Status line segment (statusline/cache-segment.sh, optional) | Claude Code's own exact cache figures in colour: hit ratio, misses, last miss cause. It works standalone or after your existing status line. | One line in settings.json. |
Claude Code caches your conversation prefix. The cache lasts 1 hour on a Claude subscription (within included usage) and 5 minutes on API keys, gateways, Bedrock, Vertex and Foundry, unless you set promptCacheTtl.
When the cache expires, your next message re-writes the whole context at the cache-write rate. That rate is 2× the input price for 1h, and 1.25× for 5m. A cache read costs 0.1×, and every read restarts the timer.
So one keepalive ping that reads the cache costs about 1/20 of the cold re-write it prevents on a 1h cache, and about 1/12 on a 5m cache. It pays off when you come back. It's wasted when you don't, which is why every window here has an end.
/plugin marketplace add Peek-Everything/claude-cache-warmer
/plugin install cache-warmer@claude-cache-warmer
To try it from a local clone instead: claude --plugin-dir ./plugins/cache-warmer.
Requires a Claude Code release with the plugin hooks API. It's tested on 2.1.289; earlier releases are untested. That API is early access and may change between releases.
| Command | Effect |
|---|---|
/warm / /warm 4h / /warm 90m | Keep this session warm for the auto window (default 2h) or the given time. The maximum is 8h. |
/warm off / /warm on | Stop for this session / undo that. |
/warm auto on / /warm auto off | Keep every session warm for the auto window after each reply. Applies to 1h caches only. Remembered across sessions. |
/warm status | Cache state, TTL (and where it came from), keepalive state. |
/warm test | Send one ping now and report what the cache served (read should be large and wrote small). |
Settings. The defaults below apply as soon as it's installed. To change them, run /plugin configure cache-warmer@claude-cache-warmer (or edit pluginConfigs in settings.json). The installer's note that options are "not yet set" just means the defaults are in use.
| Option | Default | ||
|---|---|---|---|
auto | false | Auto keepalive after each reply. `/warm auto on\ | off` overrides it. |
autoWindowMinutes | 120 | How long after the last reply auto keeps pinging. Maximum 480. | |
indicator | true | Show the cache indicator under the prompt. |
$.model.fork), sent about 5 minutes before a 1h cache expires, or about 1 minute before a 5m cache does. The API serves the same prefix from cache and the timer restarts.ANTHROPIC_BASE_URL, no credentials, no settings changes.claude -p and Agent SDK hosts it registers nothing and sends nothing.CACHE_WARMER_DISABLE=1~/.claude/cache-warmer-off exists/warm still allows it and tells you the rate first.FORCE_PROMPT_CACHING_5M, CLAUDE_CODE_PROMPT_CACHE_TTL, promptCacheTtl, or ENABLE_PROMPT_CACHING_1H. Authoritative.~1h / ~5m in the indicator.The plugin infers cache state from response timings and token counts. Claude Code also sends its own prompt_cache figures to status line commands. Those include misses and last_miss_cause, which the plugin can't see.
"statusLine": {
"type": "command",
"command": "/path/to/statusline/cache-segment.sh /path/to/your-existing-statusline.sh",
"refreshInterval": 30
}
Leave out the second path if you have no status line. Your existing output is printed unchanged, and the cache segment goes on its own line after it. Requires bash, jq and awk.
auto.The plugin reads a few environment variables and settings to work out the cache TTL. It stores timing and token counts, never prompt text. Its only network use is Claude Code's own request over your session. SECURITY.md has the full list.
/plugin uninstall cache-warmer@claude-cache-warmer
/plugin marketplace remove claude-cache-warmer
If you added the status line segment, restore your previous statusLine.command.
Working on the code, by hand or with an AI agent? See AGENTS.md for the layout, the engine's rules, the invariants and the release steps.
claude plugin validate plugins/cache-warmer
claude plugin test plugins/cache-warmer # 21 tests, mocked clock
statusline/test.sh # 9 tests
CI runs the status line tests and the marketplace check. It does not run the plugin checks: GitHub-hosted runners currently reject the early-access hooks module. Run the two claude plugin commands locally before tagging a release.
Ideas were weighed against cache-tax (fork-based keepalive, MIT), cachebeat, claude-code-cache-keepalive and clodex. This plugin avoids transcript turns, proxies and credential access, and adds TTL detection and an indicator.
MIT
hooks/register.ts 352 lines1import type { EngineInterface, PluginOptions, Register } from 'claude-code'
2
3// cache-warmer: a prompt-cache indicator plus a bounded keepalive for INTERACTIVE
4// Claude Code sessions.
5//
6// Indicator (under the prompt): cache ● 1h ████░░ 38m left · hit 91% · keepalive 1h42m ↻2
7// cache ○ cold · next message re-caches ~82k tokens
8//
9// Keepalive: one $.model.fork (tool-less, over the session's own transcript, so the
10// API serves the exact main-thread prefix from its cache) shortly before the cache
11// TTL lapses. Nothing is appended to the transcript; no proxy, no env changes, no
12// credentials are touched.
13//
14// /warm [2h|90m] keep this session warm for a window (max 8h)
15// /warm off | on stop for this session / undo that
16// /warm auto on|off keep every session warm after each reply (persists; 1h TTL only)
17// /warm status details
18// /warm test one ping now, report what the cache served
19//
20// Never pings: in a non-interactive session (claude -p, SDK hosts); outside a
21// window; while a turn runs; once the cache is already cold; when
22// CACHE_WARMER_DISABLE=1 or ~/.claude/cache-warmer-off exists. A ping that reads
23// nothing, or writes more than 10% of what it read, stops it for the session.
24
25const MIN = 60_000
26const HOUR = 60 * MIN
27const MAX_WINDOW = 8 * HOUR
28const TICK = 30_000
29const STALE = 24 * HOUR
30const PING_PROMPT = 'Reply with the single word: ok'
31
32type TtlSource = 'config' | 'learned' | 'guess'
33type Saved = {
34 until: number // keepalive window end (0 = none)
35 last: number // when the main thread last got a response
36 pings: number
37 optOut: boolean
38 ctx: number // context tokens the next request re-sends
39 read: number // cache-read tokens over the session's own turns
40 total: number // all input tokens over the session's own turns
41 seen: number
42}
43const EMPTY: Saved = { until: 0, last: 0, pings: 0, optOut: false, ctx: 0, read: 0, total: 0, seen: 0 }
44
45// Module state: a reload resets these; the session record comes back from $.store.
46let isInteractive = false
47let isPinging = false
48let sid = ''
49let s: Saved = { ...EMPTY }
50let ttl = HOUR
51let ttlSource: TtlSource = 'guess'
52let autoOverride: boolean | undefined
53let opts = { auto: false, autoWindow: 2 * HOUR, indicator: true }
54let submittedAt = 0
55const running = new Set<string>()
56
57// ---- pure helpers ----------------------------------------------------------
58
59const fmt = (ms: number) => {
60 const m = Math.max(0, Math.round(ms / MIN))
61 return m >= 60 ? `${Math.floor(m / 60)}h${String(m % 60).padStart(2, '0')}m` : `${m}m`
62}
63const fmtLeft = (ms: number) => (ms >= MIN ? `${Math.floor(ms / MIN)}m` : `${Math.max(0, Math.floor(ms / 1000))}s`)
64const parseWindow = (arg: string): number | undefined => {
65 const m = /^(\d+)\s*(m|h)$/i.exec(arg.trim())
66 if (!m) return undefined
67 return Number(m[1]) * (m[2]?.toLowerCase() === 'h' ? HOUR : MIN)
68}
69const lead = () => Math.min(5 * MIN, Math.floor(ttl / 5))
70const isAuto = () => autoOverride ?? opts.auto
71const ttlLabel = () => `${ttlSource === 'guess' ? '~' : ''}${ttl >= HOUR ? '1h' : '5m'}`
72const bar = (frac: number) => {
73 const n = Math.max(0, Math.min(6, Math.round(frac * 6)))
74 return '█'.repeat(n) + '░'.repeat(6 - n)
75}
76
77function readOptions(o: PluginOptions) {
78 const win = typeof o.autoWindowMinutes === 'number' ? o.autoWindowMinutes : 120
79 opts = {
80 auto: o.auto === true,
81 autoWindow: Math.min(MAX_WINDOW, Math.max(10, win) * MIN),
82 indicator: o.indicator !== false,
83 }
84}
85
86function indicatorText(now: number): string | undefined {
87 const keep = s.until > now ? `keepalive ${fmt(s.until - now)} ↻${s.pings}` : ''
88 if (!opts.indicator || s.last === 0) return keep || undefined
89 const left = s.last + ttl - now
90 const hit = s.total > 0 ? ` · hit ${Math.round((s.read / s.total) * 100)}%` : ''
91 const tail = keep ? ` · ${keep}` : ''
92 if (left > 0) return `cache ● ${ttlLabel()} ${bar(left / ttl)} ${fmtLeft(left)} left${hit}${tail}`
93 const k = Math.round(s.ctx / 1000)
94 const recache = s.ctx > 0 ? ` · next message re-caches ~${k > 0 ? `${k}k` : '<1k'} tokens` : ''
95 return `cache ○ cold${recache}${tail}`
96}
97
98function describe(now: number): string {
99 const lines = [indicatorText(now) ?? 'cache: no response yet in this session']
100 const src = { config: 'from your settings', learned: 'learned from this account', guess: 'guessed; it is learned after an idle gap' }[ttlSource]
101 lines.push(`cache TTL ${ttl >= HOUR ? '1h' : '5m'} (${src})`)
102 lines.push(`keepalive: ${s.until > now ? `on, ${fmt(s.until - now)} left, ${s.pings} ping(s)` : 'idle'}` +
103 ` · auto ${isAuto() ? 'on' : 'off'}${s.optOut ? ' · off for this session' : ''}`)
104 if (ttl < HOUR) lines.push('note: a 5m cache needs a ping about every 4 minutes; auto never runs on 5m, /warm does.')
105 return lines.join('\n')
106}
107
108// ---- helpers that use $ --------------------------------------------------------
109
110async function save($: EngineInterface) {
111 s = { ...s, seen: await $.clock.now() }
112 await $.store.set(`s:${sid}`, s)
113}
114
115async function show($: EngineInterface) {
116 $.ui.status(indicatorText(await $.clock.now()))
117}
118
119async function resolveTtl($: EngineInterface) {
120 const settings = (await $.settings.read().catch(() => undefined)) as Record<string, unknown> | undefined
121 const configured = (await $.env.get('FORCE_PROMPT_CACHING_5M')) === '1' ? '5m'
122 : (await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL')) ?? (settings?.promptCacheTtl as string | undefined)
123 ?? ((await $.env.get('ENABLE_PROMPT_CACHING_1H')) === '1' ? '1h' : undefined)
124 if (configured === '5m' || configured === '1h') {
125 ttl = configured === '1h' ? HOUR : 5 * MIN
126 ttlSource = 'config'
127 return
128 }
129 const learned = await $.store.get('ttl')
130 if (learned === '5m' || learned === '1h') {
131 ttl = learned === '1h' ? HOUR : 5 * MIN
132 ttlSource = 'learned'
133 return
134 }
135 // Subscriptions get 1h by default; API keys, gateways and cloud providers get 5m.
136 const isNonSub = [
137 await $.env.get('ANTHROPIC_API_KEY'), await $.env.get('ANTHROPIC_AUTH_TOKEN'),
138 await $.env.get('ANTHROPIC_BASE_URL'), await $.env.get('CLAUDE_CODE_USE_BEDROCK'),
139 await $.env.get('CLAUDE_CODE_USE_VERTEX'), await $.env.get('CLAUDE_CODE_USE_FOUNDRY'),
140 ].some(Boolean)
141 ttl = isNonSub ? 5 * MIN : HOUR
142 ttlSource = 'guess'
143}
144
145// After an idle gap that only a 1h cache survives, the first request's cache WRITES
146// tell which TTL this account really gets. Writes (not reads) are used because a
147// multi-step turn's reads include its own later steps.
148async function learnTtl($: EngineInterface, gap: number, wrote: number) {
149 if (ttlSource === 'config' || s.ctx < 10_000 || gap < 6 * MIN || gap > 54 * MIN) return
150 if (wrote <= 0.2 * s.ctx) {
151 await $.store.set('ttl', '1h')
152 await $.store.delete('ttl5m-votes')
153 ttl = HOUR
154 ttlSource = 'learned'
155 } else if (wrote >= 0.8 * s.ctx) {
156 const votes = Number((await $.store.get('ttl5m-votes')) ?? 0) + 1 // one miss can have other causes
157 await $.store.set('ttl5m-votes', votes)
158 if (votes >= 2) {
159 await $.store.set('ttl', '5m')
160 ttl = 5 * MIN
161 ttlSource = 'learned'
162 }
163 }
164}
165
166async function blockedReason($: EngineInterface): Promise<string | undefined> {
167 if ((await $.env.get('CACHE_WARMER_DISABLE')) === '1') return 'CACHE_WARMER_DISABLE=1'
168 const home = await $.env.get('HOME')
169 if (home && (await $.fs.stat(`${home}/.claude/cache-warmer-off`).then(() => true, () => false)))
170 return '~/.claude/cache-warmer-off exists'
171 return undefined
172}
173
174async function autoArm($: EngineInterface, now: number) {
175 if (!isAuto() || s.optOut || ttl < HOUR) return
176 if (await blockedReason($)) return
177 s = { ...s, until: Math.max(s.until, now + opts.autoWindow) }
178}
179
180async function stopWindow($: EngineInterface, why: string, optOut: boolean) {
181 s = { ...s, until: 0, optOut: s.optOut || optOut }
182 await save($)
183 await show($)
184 $.ui.toast(`cache-warmer stopped${optOut ? ' for this session' : ''}: ${why}`)
185}
186
187async function ping($: EngineInterface, now: number, isManual: boolean): Promise<string> {
188 const blocked = await blockedReason($)
189 if (blocked) return `cache-warmer: skipped, ${blocked}`
190 isPinging = true
191 try {
192 const r = await $.model.fork({ prompt: PING_PROMPT })
193 if (!r.isAnswered) {
194 // nothing-to-fork (after /clear) and aborted are not failures of the cache.
195 if (!isManual && r.reason === 'api-error') await stopWindow($, `ping failed (${r.reason})`, false)
196 return `cache-warmer ping: no reply (${r.reason})`
197 }
198 const { cache_read_input_tokens: read, cache_creation_input_tokens: wrote, output_tokens: out } = r.usage
199 const report = `cache read ${read.toLocaleString()} · wrote ${wrote.toLocaleString()} · output ${out}`
200 if (read === 0 || wrote > read * 0.1) {
201 await stopWindow($, `ping missed the cache (${report})`, true)
202 return `cache-warmer ping MISSED the cache: ${report}. Stopped for this session.`
203 }
204 s = { ...s, last: await $.clock.now(), pings: s.pings + (isManual ? 0 : 1) }
205 await save($)
206 await show($)
207 return `cache-warmer ping ok: ${report}`
208 } finally {
209 isPinging = false
210 }
211}
212
213async function tick($: EngineInterface) {
214 const now = await $.clock.now()
215 if (s.until > 0 && now >= s.until) {
216 s = { ...s, until: 0 }
217 await save($)
218 }
219 await show($)
220 if (s.until <= now || running.size > 0 || isPinging || s.last === 0) return
221 const expiresAt = s.last + ttl
222 if (now >= expiresAt) return // already cold: a ping would only pay a full write
223 if (now < expiresAt - lead()) return
224 await ping($, now, false)
225}
226
227async function prune($: EngineInterface, now: number) {
228 for (const key of await $.store.keys()) {
229 if (!key.startsWith('s:') || key === `s:${sid}`) continue
230 const rec = (await $.store.get(key)) as Partial<Saved> | undefined
231 if (!rec || (rec.seen ?? 0) < now - STALE) await $.store.delete(key)
232 }
233}
234
235// ---- hooks ---------------------------------------------------------------------
236
237export const register: Register = (on, options) => {
238 // Fresh state per load, whatever a previous load of this module left behind.
239 isInteractive = false
240 isPinging = false
241 sid = ''
242 s = { ...EMPTY }
243 ttl = HOUR
244 ttlSource = 'guess'
245 autoOverride = undefined
246 submittedAt = 0
247 running.clear()
248 readOptions(options)
249
250 on('session.start', async ($, e, next) => {
251 const r = await next(e)
252 isInteractive = e.isInteractive
253 if (!isInteractive) return r // headless and SDK sessions: nothing registered, nothing sent
254
255 sid = await $.session.id()
256 s = { ...EMPTY, ...((await $.store.get(`s:${sid}`)) as Partial<Saved> | undefined) }
257 const stored = await $.store.get('auto')
258 autoOverride = typeof stored === 'boolean' ? stored : undefined
259 await resolveTtl($)
260 await prune($, await $.clock.now())
261 await $.command.register({
262 name: 'warm',
263 description: 'Prompt-cache keepalive: /warm [2h|off|on|auto on|auto off|status|test]',
264 argumentHint: '[2h|off|on|auto on|auto off|status|test]',
265 })
266 $.clock.every(TICK, () => void tick($))
267 await show($)
268 return r
269 })
270
271 // A real turn refreshes the cache itself; never ping over one. The first turn after
272 // an idle gap also marks when that gap ended (for TTL learning).
273 on('turn.start', async ($, e, next) => {
274 const r = await next(e)
275 if (isInteractive) {
276 if (running.size === 0) submittedAt = await $.clock.now()
277 running.add(e.turnId)
278 }
279 return r
280 })
281
282 on('turn.complete', async ($, e, next) => {
283 const r = await next(e)
284 running.delete(e.turnId)
285 if (!isInteractive || e.agentId !== undefined) return r
286 const now = await $.clock.now()
287 const u = e.usage
288 if (u) {
289 if (submittedAt > 0 && s.last > 0) await learnTtl($, submittedAt - s.last, u.cache_creation_input_tokens)
290 s = {
291 ...s,
292 read: s.read + u.cache_read_input_tokens,
293 total: s.total + u.cache_read_input_tokens + u.cache_creation_input_tokens + u.input_tokens,
294 }
295 }
296 const ctx = (await $.session.usage().catch(() => undefined))?.context.tokens
297 s = { ...s, last: now, ctx: ctx ?? s.ctx }
298 submittedAt = 0
299 await autoArm($, now)
300 await save($)
301 await show($)
302 return r
303 })
304
305 on('command.run', { command: 'warm' }, async ($, e, next) => {
306 if (!isInteractive) return next(e)
307 const arg = e.args.trim().toLowerCase().replace(/\s+/g, ' ')
308 const now = await $.clock.now()
309 const done = async (text: string) => {
310 await save($)
311 await show($)
312 return { text }
313 }
314
315 if (arg === 'auto on' || arg === 'auto off') {
316 autoOverride = arg === 'auto on'
317 await $.store.set('auto', autoOverride)
318 if (!autoOverride) s = { ...s, until: 0 }
319 else if (s.last > 0) await autoArm($, now)
320 return done(autoOverride
321 ? `cache-warmer auto on: sessions stay warm up to ${fmt(opts.autoWindow)} after each reply${ttl < HOUR ? ' (not on this 5m cache)' : ''}.`
322 : 'cache-warmer auto off: use /warm 2h to keep a session warm by hand.')
323 }
324 if (arg === 'off') {
325 s = { ...s, until: 0, optOut: true }
326 return done('cache-warmer off for this session (/warm on to undo).')
327 }
328 if (arg === 'on') {
329 s = { ...s, optOut: false }
330 if (s.last > 0) await autoArm($, now)
331 return done(describe(now))
332 }
333 if (arg === 'status' || arg === 'help') return { text: describe(now) }
334 if (arg === 'test') return { text: await ping($, now, true) }
335
336 const span = arg === '' ? opts.autoWindow : parseWindow(arg)
337 if (span === undefined) return { text: 'usage: /warm [2h|90m|off|on|auto on|auto off|status|test]' }
338 const blocked = await blockedReason($)
339 if (blocked) return { text: `cache-warmer not armed: ${blocked}` }
340 const capped = Math.min(span, MAX_WINDOW)
341 s = { ...s, until: now + capped, optOut: false }
342 const perHour = Math.round(HOUR / (ttl - lead()))
343 return done(`cache-warmer on for ${fmt(capped)}${capped < span ? ' (capped at 8h)' : ''}` +
344 ` · cache TTL ${ttlLabel()} · about ${perHour} ping${perHour === 1 ? '' : 's'}/hour while idle.`)
345 })
346
347 on('session.end', async ($, e, next) => {
348 if (isInteractive && sid) await $.store.delete(`s:${sid}`)
349 return next(e)
350 })
351}
352