Prompt-cache meter above the Claude Code prompt: how many tokens each request read from, wrote to and sent past the cache, a live countdown to the cache's…

A prompt-cache meter above the Claude Code prompt. Every request Claude makes reports how much of its prompt the cache served, how much it wrote and how much went uncached; this mod keeps those numbers per request and per turn, counts down to the moment the cache lapses and tells you what to do about it: keep going, /compact or /clear.
cache ██████████ 98% read 80k · wrote 1k · new 300 ⏱ 3:41 5m · warm: keep going
cache ░░░░░░░░░░ 0% read 0 · wrote 52k · new 300 ⏱ 4:58 5m · cache missed: model changed (…)
cache ██████████ 98% read 150k · wrote 1k · new 300 ⏱ 0:00 5m · expired: the next message rewrites 151k tokens. /compact first, or /clear if the task is done
/cache opens a pane: the time left with a solid bar that shrinks as the cache runs out (green, yellow below 40% of the lifetime, red from the warning threshold), a stacked read / wrote / new bar for the last request, and a colour-coded table with one row per turn. Bars are filled cells and columns have fixed widths with no-break spaces, so the terminal and the Desktop (HTML) pane render the same.
From Anthropic's prompt caching documentation:
input_tokens (uncached remainder) + cache_read_input_tokens + cache_creation_input_tokens./compact./clear starts a new conversation in the same process, so the meter and the /cache table start over with it. A change in the prefix (model, effort or thinking settings, tool set, system prompt, CLAUDE.md) makes the next request write instead of read. The mod names the cause when it sees a miss: model changed, the cache had lapsed, or the prefix changed.The mod follows Claude Code's own rules (prompt caching: cache lifetime, Claude Code 2.1.242 or later). For the main conversation the TTL is the first match of:
| # | Source | Result |
|---|---|---|
| 1 | the mod's ttl option (5m / 1h) | what you set |
| 2 | FORCE_PROMPT_CACHING_5M=1 | 5 minutes |
| 3 | CLAUDE_CODE_PROMPT_CACHE_TTL | 5m or 1h |
| 4 | the promptCacheTtl setting (local, project or user settings file) | 5m or 1h |
| 5 | ENABLE_PROMPT_CACHING_1H=1 | 1 hour |
| 6 | the account | 1 hour on a Claude subscription within its plan usage; 5 minutes on usage credits, an API key or a cloud provider |
The account comes from the rate-limit windows the last response reported: a five_hour or seven_day window means a subscription, and one at 100% means requests now draw on usage credits. An API key or a cloud provider reports no such window, and before the first response nothing is known, so the mod starts from 5 minutes there. Managed settings are not readable from a mod.
On top of that the mod watches the traffic, which beats rows 2 to 6: a request that hits the cache more than 5 minutes after the previous one proves the 1-hour lifetime (a later miss does not undo it, since a changed prefix looks the same), and a miss 5 to 60 minutes after the previous request, with the same model and a prompt that did not shrink, says the entry lapsed, so 5 minutes (a later hit overrules it). That covers what the mod cannot see: managed settings, a gateway that rewrites the TTL, or a subscription that ran out of plan usage mid-session. The pane header names the source in use.
Why the mod infers instead of reading it: the API names the TTL of each write (cache_creation.ephemeral_5m_input_tokens / ephemeral_1h_input_tokens) and Claude Code's status line exposes it as prompt_cache.ttl, but the mod API passes on only the four token counts. To check by hand, claude -p "hello" --output-format json and read usage.cache_creation.
Other switches read from the environment at session start:
| Variable | Effect on the meter |
|---|---|
DISABLE_PROMPT_CACHING=1 (and _HAIKU, _SONNET, _OPUS) | the band says caching is off for that model |
turn.step: reads each main-loop request's usage (subagents have their own prefixes and are left out)$.clock.every(1000): redraws the countdown, and only while its text changes, so an idle expired session costs nothingui.render on AbovePrompt (the band) and on Pane (/cache)$.ui.toast: once per cache entry at the warning threshold (60 s by default) and again at 10, 3, 2 and 1 seconds left, for prompts of 20k tokens or more ttl: string "auto" | "5m" | "1h" (default auto)
warnSeconds: number countdown threshold for the yellow state and the toast (default 60)
compactAtTokens: number prompt size that makes an expired cache suggest /compact (default 100000)
band: boolean row above the prompt (default true)
status: boolean entry under the prompt, "cache 98% · 3:41" (default false)
toast: boolean toasts at the threshold, 10, 3, 2 and 1 s (default true)
The 100k compactAtTokens is a judgement, not a figure from the documentation: lower it if your model's cache writes are expensive for you.
From the claudemods marketplace:
claude plugin marketplace add jakerains/claudemods
claude plugin install prompt-cache-control@claudemods
/cache opens the table, /cache stop closes it, /cache off and /cache on hide or show the bar above the prompt (remembered).
Options are read from user settings (~/.claude/settings.json) under the plugin's full id:
{ "pluginConfigs": { "prompt-cache-control@claudemods": { "options": { } } } }
$.state, so /reload-plugins and hot reloads keep the history and countdown (the original reset them).noUncheckedIndexedAccess).Early access. Mods need Claude Code 2.1.259+ with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1; the $ API may change between releases. Typed against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods
hooks/prompt-cache-control.tsx 497 lines1/**
2 * prompt-cache-control — Claude Mod (EARLY ACCESS)
3 *
4 * A prompt-cache meter for Claude Code. Every main-loop request reports how
5 * many prompt tokens the cache served (`cache_read_input_tokens`), wrote
6 * (`cache_creation_input_tokens`) and sent uncached (`input_tokens`); this mod
7 * keeps those per request and per turn, counts down to the moment the cache
8 * lapses, and says what to do about it: keep going, /compact or /clear.
9 *
10 * - `turn.step` reads each main-loop request's usage (subagents have their
11 * own prefixes and are left out)
12 * - `$.clock.every(1000)` redraws the countdown, and only while its text
13 * changes: an idle, expired session costs nothing
14 * - a row above the prompt (the AbovePrompt component), an optional status
15 * line entry, and `/cache`, a pane with one row per turn
16 *
17 * The lifetime is counted from the start of the request that last wrote or read
18 * the cache, as Anthropic documents it. Which lifetime Claude Code asked for
19 * follows its documented rules (see decideTtl in ./cache.ts): FORCE_PROMPT_CACHING_5M,
20 * CLAUDE_CODE_PROMPT_CACHE_TTL, the promptCacheTtl setting, ENABLE_PROMPT_CACHING_1H,
21 * then the account (1 hour on a Claude subscription, 5 minutes otherwise). The
22 * API names the TTL of a write but the mod API passes on only the token counts,
23 * so the mod also watches the gaps between requests (a hit after more than 5
24 * minutes proves 1 hour; see observeTtl). `ttl: "5m" | "1h"` pins it.
25 *
26 * Needs CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 (Claude Code >= 2.1.259).
27 *
28 * What the band and the pane draw from lives in $.state (the Meter atom), so a
29 * hot reload or /reload-plugins keeps the history and the countdown. The band
30 * composes with other mods' rows above the prompt and sits last, on the input.
31 *
32 * Options (pluginConfigs["prompt-cache-control@claudemods"].options):
33 * ttl: "auto" | "5m" | "1h" cache lifetime (default auto)
34 * warnSeconds: number countdown threshold for the warning (default 60)
35 * compactAtTokens: number prompt size that makes an expired cache suggest /compact (default 100000)
36 * band: boolean row above the prompt (default true)
37 * status: boolean entry under the prompt (default false)
38 * toast: boolean toasts near expiry: at warnSeconds, then 10, 3, 2 and 1 s (default true)
39 */
40import { atom, read, update } from 'claude-code'
41import type { EngineInterface, Register } from 'claude-code'
42import {
43 advise,
44 COUNTDOWN_MARKS,
45 bar,
46 byTurn,
47 fit,
48 fmtClock,
49 fmtTokens,
50 hitRatio,
51 isCachingDisabled,
52 accountOf,
53 decideTtl,
54 observeTtl,
55 lifeColor,
56 lifeRatio,
57 nextToastMark,
58 positive,
59 promptTokens,
60 remainingMs,
61 rowRatio,
62 segments,
63} from './cache.ts'
64import type { Advice, CacheEnv } from './cache.ts'
65import type { Meter } from '../types'
66
67const PANE = 'cache'
68const COMMAND = 'cache'
69const KEEP = 200
70// below this a lapsed cache costs too little to interrupt anyone about
71const TOAST_MIN_TOKENS = 20_000
72
73const EMPTY: Meter = {
74 samples: [],
75 ttl: '5m',
76 baseTtl: '5m',
77 pinned: false,
78 observed: null,
79 account: 'other',
80 ttlSource: 'default',
81 envSource: 'default',
82 env: {},
83 setting: null,
84}
85
86// Held by the host: a reload starts this module over but keeps these.
87const meter = atom({ plugin: 'prompt-cache-control', key: 'meter' } as const, EMPTY)
88const paneOpen = atom({ plugin: 'prompt-cache-control', key: 'paneOpen' } as const, false)
89const barOn = atom({ plugin: 'prompt-cache-control', key: 'barOn' } as const, true)
90
91// The ticker's own bookkeeping; a reload restarts the ticker with it.
92let timer: { cancel: () => void } | undefined
93let lastKey = ''
94let toastedFor = 0
95let toastLevel = Infinity
96
97type Policy = { warnMs: number; compactAtTokens: number }
98
99function current(m: Meter, policy: Policy, now: number) {
100 const last = m.samples[m.samples.length - 1]
101 const prev = m.samples[m.samples.length - 2]
102 const disabled = last ? isCachingDisabled(last.model, m.env) : isCachingDisabled('', m.env)
103 const advice: Advice = advise(last, prev, { ttl: m.ttl, ...policy }, now, disabled)
104 const left = last ? remainingMs(last, m.ttl, now) : 0
105 return { last, advice, left }
106}
107
108const COLOR: Record<Advice['kind'], string | undefined> = {
109 warm: 'green',
110 soon: 'yellow',
111 expired: 'red',
112 miss: 'red',
113 off: undefined,
114 cold: undefined,
115 uncached: undefined,
116}
117
118function shortLine(m: Meter, policy: Policy, now: number): string {
119 const { last, advice, left } = current(m, policy, now)
120 if (!last || advice.kind === 'off') return `cache: ${advice.text}`
121 const clock = left > 0 ? ` · ${fmtClock(left)}` : ''
122 return `cache ${Math.round(hitRatio(last) * 100)}%${clock}`
123}
124
125// the promptCacheTtl setting, from the settings files that can carry it (local over project over user)
126async function readSetting($: EngineInterface): Promise<string | null> {
127 const home = await $.env.get('HOME').catch(() => undefined)
128 const cwd = await $.session.cwd().catch(() => undefined)
129 const files = [cwd && `${cwd}/.claude/settings.local.json`, cwd && `${cwd}/.claude/settings.json`, home && `${home}/.claude/settings.json`]
130 for (const file of files) {
131 if (!file) continue
132 try {
133 const value = JSON.parse(await $.fs.read(file)).promptCacheTtl
134 if (value === '5m' || value === '1h') return value
135 } catch {
136 // missing or unreadable: the next file
137 }
138 }
139 return null
140}
141
142export const register: Register = (on, options) => {
143 const policy: Policy = {
144 warnMs: positive(options.warnSeconds, 60) * 1000,
145 compactAtTokens: positive(options.compactAtTokens, 100_000),
146 }
147 const showBand = options.band !== false
148 const showStatus = options.status === true
149 const wantToast = options.toast !== false
150
151 on('session.start', async ($, e, next) => {
152 const r = await next(e)
153 lastKey = ''
154 toastedFor = 0
155 // A pinned status outlives a reload: with the option off, take down any
156 // entry an earlier load (or the option, before it was turned off) left.
157 if (!showStatus) $.ui.status(undefined)
158 const none = () => undefined
159 const env: CacheEnv = {
160 enable1h: await $.env.get('ENABLE_PROMPT_CACHING_1H').catch(none),
161 force5m: await $.env.get('FORCE_PROMPT_CACHING_5M').catch(none),
162 ttlVar: await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL').catch(none),
163 disableAll: await $.env.get('DISABLE_PROMPT_CACHING').catch(none),
164 disableHaiku: await $.env.get('DISABLE_PROMPT_CACHING_HAIKU').catch(none),
165 disableSonnet: await $.env.get('DISABLE_PROMPT_CACHING_SONNET').catch(none),
166 disableOpus: await $.env.get('DISABLE_PROMPT_CACHING_OPUS').catch(none),
167 }
168 const pinned = options.ttl === '5m' || options.ttl === '1h'
169 const setting = await readSetting($)
170 const account = accountOf((await $.session.usage().catch(() => undefined))?.rateLimits ?? [])
171 const choice = decideTtl(options.ttl, env, setting ?? undefined, account)
172 // A reload fires session.start again: the requests seen so far, and any
173 // lifetime they proved, stay. A new session's state starts empty anyway.
174 const m = await update($, meter, m => {
175 const observed = pinned ? null : m.observed
176 return {
177 ...m,
178 env,
179 pinned,
180 setting,
181 account,
182 observed,
183 baseTtl: choice.ttl,
184 envSource: choice.source,
185 ttl: observed ?? choice.ttl,
186 ttlSource: observed ? m.ttlSource : choice.source,
187 }
188 })
189
190 await $.command
191 .register({
192 name: COMMAND,
193 description: 'Prompt-cache table per turn (stop closes it); on / off shows or hides the bar above the prompt',
194 argumentHint: '[on|off|stop]',
195 immediate: true,
196 })
197 .catch(err => $.ui.log(`prompt-cache-control: /${COMMAND} not registered: ${err}`))
198 $.ui.log(`prompt-cache-control loaded: ${m.ttl} cache (${m.ttlSource}), /${COMMAND} opens the table`, { to: 'debug' })
199
200 // The bar's on/off is the person's last /cache on|off, else the `band` option.
201 const kept = await $.store.get('bar').catch(() => undefined)
202 await update($, barOn, () => (typeof kept === 'boolean' ? kept : showBand))
203
204 timer?.cancel()
205 timer = $.clock.every(1000, async () => {
206 const now = Date.now()
207 const m = await read($, meter)
208 const { last, advice, left } = current(m, policy, now)
209 const key = `${advice.kind}|${advice.text}|${left > 0 ? fmtClock(left) : ''}`
210 if (key !== lastKey) {
211 lastKey = key
212 if (showStatus) $.ui.status(shortLine(m, policy, now))
213 $.ui.invalidate('ui.render')
214 }
215 if (wantToast && last && left > 0 && promptTokens(last) >= TOAST_MIN_TOKENS) {
216 if (toastedFor !== last.startedAt) {
217 toastedFor = last.startedAt
218 toastLevel = Infinity
219 }
220 // the first toast comes at warnSeconds, then 10, 3, 2 and 1 seconds; a late tick skips to the newest one
221 const secs = Math.ceil(left / 1000)
222 const mark = nextToastMark(secs, policy.warnMs / 1000, toastLevel)
223 if (mark !== undefined) {
224 toastLevel = mark
225 const tail = secs <= COUNTDOWN_MARKS[0]! ? 'send a message now' : `send a message to keep ${fmtTokens(promptTokens(last))} tokens warm`
226 $.ui.toast(`cache expires in ${secs >= 60 ? fmtClock(left) : `${secs}s`}: ${tail}`)
227 }
228 }
229 })
230 return r
231 })
232
233 on('session.end', async ($, e, next) => {
234 // /clear starts a new conversation in the same process: its cache is a new one
235 if (e.reason === 'clear') {
236 lastKey = ''
237 toastedFor = 0
238 await update($, meter, m => ({ ...m, samples: [], observed: null, ttl: m.baseTtl, ttlSource: m.envSource }))
239 $.ui.invalidate('ui.render')
240 return next(e)
241 }
242 timer?.cancel()
243 timer = undefined
244 return next(e)
245 })
246
247 // each main-loop request: what the cache did with it
248 on('turn.step', async function* ($, e, next) {
249 if (e.agentId) return yield* next(e)
250 const startedAt = Date.now()
251 const r = yield* next(e)
252 if (r.usage) {
253 const usage = r.usage
254 const before = await read($, meter)
255 // the account can change under a session: a subscription running out of plan usage moves to usage credits
256 const account = before.pinned
257 ? before.account
258 : accountOf((await $.session.usage().catch(() => undefined))?.rateLimits ?? [])
259 const m = await update($, meter, m => {
260 const samples = [
261 ...m.samples,
262 {
263 turnId: e.turnId,
264 index: e.index,
265 model: usage.model || e.model,
266 startedAt,
267 read: usage.cache_read_input_tokens,
268 write: usage.cache_creation_input_tokens,
269 fresh: usage.input_tokens,
270 output: usage.output_tokens,
271 },
272 ].slice(-KEEP)
273 if (m.pinned) return { ...m, samples }
274 const choice = decideTtl(options.ttl, m.env, m.setting ?? undefined, account)
275 const moved = { ...m, samples, account, baseTtl: choice.ttl, envSource: choice.source }
276 const known = m.observed ?? undefined
277 const seen = observeTtl(samples[samples.length - 2], samples[samples.length - 1]!, known)
278 if (seen !== known) {
279 return {
280 ...moved,
281 observed: seen ?? null,
282 ttl: seen ?? choice.ttl,
283 ttlSource: `observed from request timing; ${choice.source} said ${choice.ttl}`,
284 }
285 }
286 return m.observed ? moved : { ...moved, ttl: choice.ttl, ttlSource: choice.source }
287 })
288 if (m.observed !== before.observed) {
289 $.ui.log(`prompt-cache-control: cache lifetime is ${m.ttl} (${m.ttlSource})`, { to: 'debug' })
290 }
291 lastKey = ''
292 if (showStatus) $.ui.status(shortLine(m, policy, Date.now()))
293 $.ui.invalidate('ui.render')
294 }
295 return r
296 })
297
298 on('command.run', { command: COMMAND }, async ($, e) => {
299 const word = e.args.trim().toLowerCase()
300 // Turning the bar off draws nothing above the prompt, so the engine has no
301 // band to collapse and no "panel hidden" line to show.
302 if (word === 'on' || word === 'off') {
303 const isOn = word === 'on'
304 await $.store.set('bar', isOn)
305 await update($, barOn, () => isOn)
306 return { text: isOn ? 'cache bar on, above the prompt' : `cache bar off · /${COMMAND} on brings it back, /${COMMAND} opens the table` }
307 }
308 if (word === 'stop') {
309 await $.ui.close({ id: PANE }).catch(() => undefined)
310 await update($, paneOpen, () => false)
311 return { text: 'cache table closed' }
312 }
313 await update($, paneOpen, () => true)
314 await $.ui.open({ id: PANE, title: 'cache', focus: true })
315 $.ui.invalidate('ui.render')
316 const m = await read($, meter)
317 const { advice } = current(m, policy, Date.now())
318 return { text: `${m.ttl} cache (${m.ttlSource}) · ${advice.text} · /${COMMAND} stop closes` }
319 })
320
321 on('ui.close', async ($, e, next) => {
322 if (e.id !== PANE) return next(e)
323 await update($, paneOpen, () => false)
324 return next(e)
325 })
326
327 on('ui.press', async ($, e, next) => {
328 if (e.plugin !== $.plugin.name || e.requestId !== PANE) return next(e)
329 if (e.element === 'close') await $.ui.close({ id: PANE }).catch(() => undefined)
330 return next(e)
331 })
332
333 // Composes: other mods (where-we-are, context-gauge) draw in this band too.
334 // Theirs first, the meter last, right on the input.
335 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
336 const others = await next(e)
337 if (!(await read($, barOn)) || e.props.hasSurvey || (await read($, paneOpen))) return others
338 const m = await read($, meter)
339 const { last, advice, left } = current(m, policy, Date.now())
340 if (!last && advice.kind !== 'off') return others
341 const { Box, Text } = $.ui.resolve(e)
342 const columns = e.props.bodyColumns ?? e.viewport?.columns ?? 100
343 const color = COLOR[advice.kind]
344 const ttl = m.ttl
345
346 if (!last) {
347 return (
348 <Box flexDirection="column">
349 {others}
350 <Text dimColor>{fit(`cache: ${advice.text}`, columns)}</Text>
351 </Box>
352 )
353 }
354
355 const ratio = hitRatio(last)
356 const wide = columns >= 90
357 const meterRow = (
358 <Box flexDirection="row" columnGap={1}>
359 <Text bold color={color}>{advice.kind === 'warm' ? '●' : advice.kind === 'soon' ? '▲' : advice.kind === 'off' || advice.kind === 'cold' || advice.kind === 'uncached' ? '○' : '✖'}</Text>
360 <Text bold color="cyan">cache</Text>
361 <Text color={color}>{bar(ratio, wide ? 10 : 6)}</Text>
362 <Text bold>{`${Math.round(ratio * 100)}%`}</Text>
363 {wide ? (
364 // a row of its own: a fragment here stacks its children as a column
365 <Box flexDirection="row" columnGap={1}>
366 <Text color="green">{`read ${fmtTokens(last.read)}`}</Text>
367 <Text color="yellow">{`wrote ${fmtTokens(last.write)}`}</Text>
368 <Text color="cyan">{`new ${fmtTokens(last.fresh)}`}</Text>
369 </Box>
370 ) : (
371 <Text dimColor>{`${fmtTokens(promptTokens(last))} tok`}</Text>
372 )}
373 {advice.kind !== 'uncached' && advice.kind !== 'off' && (
374 <Text bold color={left > 0 ? lifeColor(left, ttl, policy.warnMs) : 'red'}>{left > 0 ? `⏱ ${fmtClock(left)}` : '⏱ 0:00'}</Text>
375 )}
376 <Text dimColor wrap="truncate-end">{`${ttl} · ${advice.text}`}</Text>
377 </Box>
378 )
379
380 return (
381 <Box flexDirection="column">
382 {others}
383 {meterRow}
384 </Box>
385 )
386 })
387
388 on('ui.render', { component: 'Pane' }, async ($, e, next) => {
389 if (e.requestId !== PANE) return next(e)
390 const { Box, Text, Button } = $.ui.resolve(e)
391 const width = Math.max(30, e.props.bodyColumns - 1)
392 // HTML collapses runs of spaces and trims a text's ends; a no-break space keeps them
393 const sp = (t: string) => (e.surface === 'terminal' ? t : t.replace(/ /g, ' '))
394 const now = Date.now()
395 const m = await read($, meter)
396 const { ttl, ttlSource } = m
397 const { last, advice, left } = current(m, policy, now)
398 const all = byTurn(m.samples)
399 const counting = !!last && advice.kind !== 'uncached' && advice.kind !== 'off'
400 // the countdown goes green, then yellow, then red as the cache runs out
401 const clockColor = counting ? lifeColor(left, ttl, policy.warnMs) : undefined
402 const stateColor = advice.kind === 'expired' || advice.kind === 'miss' ? 'red' : (clockColor ?? COLOR[advice.kind])
403 const hitColor = (pct: number) => (pct >= 80 ? 'green' : pct >= 40 ? 'yellow' : 'red')
404 // solid bars are filled Boxes, not block characters, so HTML draws no seams between cells
405 const solid = (key: string, parts: [number, string | undefined][]) => (
406 <Box key={key} flexDirection="row" height={1} flexShrink={0}>
407 {parts.map(([w, c], i) => (w > 0 ? <Box key={`${key}:${i}`} width={w} height={1} flexShrink={0} backgroundColor={c} /> : null))}
408 </Box>
409 )
410 const cell = (key: string, w: number, text: string, c?: string, bold = false) => (
411 <Box key={key} width={w} flexShrink={0} justifyContent="flex-end">
412 <Text color={c} bold={bold} dimColor={!c}>{sp(text)}</Text>
413 </Box>
414 )
415
416 const barW = Math.min(width, 40)
417 const life = lifeRatio(left, ttl)
418 const lifeFilled = Math.round(life * barW)
419 const [sr, sw, sn] = last ? segments(last.read, last.write, last.fresh, barW) : [0, 0, 0]
420 const rows = all.slice(-Math.max(3, (e.viewport?.rows ?? 24) - 16))
421 const icon = advice.kind === 'warm' ? '●' : advice.kind === 'soon' ? '▲' : advice.kind === 'expired' || advice.kind === 'miss' ? '✖' : '○'
422
423 return (
424 <Box flexDirection="column">
425 <Box key="title" flexDirection="row" columnGap={1}>
426 <Text bold color="cyan">{sp('⚡ PROMPT CACHE')}</Text>
427 <Text dimColor>{sp(`· ${ttl} lifetime (${ttlSource})`)}</Text>
428 </Box>
429
430 <Box key="clock" flexDirection="column" marginTop={1}>
431 <Text bold color={clockColor}>{sp(counting ? `⏱ ${left > 0 ? fmtClock(left) : '0:00'}` : '⏱ --:--')}</Text>
432 {counting ? (
433 <Box flexDirection="row" columnGap={1}>
434 {solid('life', [[lifeFilled, clockColor], [barW - lifeFilled, 'gray']])}
435 <Text dimColor>{sp(`${Math.round(life * 100)}%`)}</Text>
436 </Box>
437 ) : null}
438 </Box>
439
440 <Box key="advice" marginTop={1} flexDirection="column">
441 <Text bold color={stateColor}>{sp(`${icon} ${advice.text}`)}</Text>
442 {last ? <Text dimColor>{sp(fit(`${last.model} · prompt ${fmtTokens(promptTokens(last))} tokens`, width))}</Text> : null}
443 </Box>
444
445 {last ? (
446 <Box key="stack" flexDirection="column" marginTop={1}>
447 <Box flexDirection="row" columnGap={1}>
448 {solid('stack', [[sr, 'green'], [sw, 'yellow'], [sn, 'cyan']])}
449 <Text bold color={hitColor(Math.round(hitRatio(last) * 100))}>{sp(`${Math.round(hitRatio(last) * 100)}% hit`)}</Text>
450 </Box>
451 <Box flexDirection="row" columnGap={2}>
452 <Text color="green">{sp(`■ read ${fmtTokens(last.read)}`)}</Text>
453 <Text color="yellow">{sp(`■ wrote ${fmtTokens(last.write)}`)}</Text>
454 <Text color="cyan">{sp(`■ new ${fmtTokens(last.fresh)}`)}</Text>
455 </Box>
456 </Box>
457 ) : null}
458
459 <Box key="table" flexDirection="column" marginTop={1}>
460 <Box key="head" flexDirection="row" columnGap={1}>
461 {cell('h:turn', 4, 'turn', 'cyan', true)}
462 {cell('h:steps', 5, 'steps', 'cyan', true)}
463 {cell('h:read', 6, 'read', 'green', true)}
464 {cell('h:wrote', 6, 'wrote', 'yellow', true)}
465 {cell('h:new', 5, 'new', 'cyan', true)}
466 {cell('h:hit', 4, 'hit', 'magenta', true)}
467 </Box>
468 {rows.length === 0 ? <Text dimColor>{sp('no requests yet')}</Text> : null}
469 {rows.map((row, i) => {
470 const n = all.length - rows.length + i + 1
471 const pct = Math.round(rowRatio(row) * 100)
472 return (
473 <Box key={`t:${row.turnId}`} flexDirection="row" columnGap={1}>
474 {cell(`c:turn:${row.turnId}`, 4, String(n))}
475 {cell(`c:steps:${row.turnId}`, 5, String(row.steps))}
476 {cell(`c:read:${row.turnId}`, 6, fmtTokens(row.read), 'green')}
477 {cell(`c:wrote:${row.turnId}`, 6, fmtTokens(row.write), 'yellow')}
478 {cell(`c:new:${row.turnId}`, 5, fmtTokens(row.fresh), 'cyan')}
479 {cell(`c:hit:${row.turnId}`, 4, `${pct}%`, hitColor(pct), true)}
480 </Box>
481 )
482 })}
483 </Box>
484
485 <Box key="foot" marginTop={1} flexDirection="column">
486 <Button key="close" label="close" onPress={() => {}} />
487 <Box key="legend" marginTop={1} flexDirection="column">
488 <Text color="green">{sp('■ read: served by the cache')}</Text>
489 <Text color="yellow">{sp('■ wrote: new cache entry')}</Text>
490 <Text color="cyan">{sp('■ new: sent uncached')}</Text>
491 </Box>
492 </Box>
493 </Box>
494 )
495 })
496}
497hooks/cache.ts 293 lines1/**
2 * cache.ts — the pure half of prompt-cache-control: no `$`, no engine.
3 *
4 * What it models, from Anthropic's prompt-caching documentation:
5 * - the cache lives 5 minutes by default, 1 hour when asked for; a read
6 * refreshes the entry at no extra cost, and the lifetime is measured from
7 * the START of the request that wrote or read it
8 * - a request's prompt is `input_tokens` (uncached remainder) +
9 * `cache_read_input_tokens` + `cache_creation_input_tokens`
10 * - writes cost 1.25x base input for 5m and 2x for 1h; reads about 0.1x
11 * (less on some models), so an expired cache on a large context is the
12 * expensive moment
13 * - a prefix change (model, effort/thinking settings, tool set, system
14 * prompt) makes the next request write instead of read
15 *
16 * Claude Code's own switches (read from the environment):
17 * ENABLE_PROMPT_CACHING_1H=1 ask for the 1-hour TTL
18 * FORCE_PROMPT_CACHING_5M=1 force the 5-minute TTL, beating the above
19 * DISABLE_PROMPT_CACHING=1 no caching; DISABLE_PROMPT_CACHING_{HAIKU,SONNET,OPUS}
20 * turn it off for that model family only
21 */
22
23import type { Account, CacheEnv, Sample, Ttl } from '../types'
24
25export type { Account, CacheEnv, Sample, Ttl }
26
27export type AdviceKind = 'off' | 'cold' | 'uncached' | 'warm' | 'soon' | 'expired' | 'miss'
28
29export type Advice = {
30 kind: AdviceKind
31 /** one sentence for the band */
32 text: string
33}
34
35export type Policy = {
36 ttl: Ttl
37 warnMs: number
38 compactAtTokens: number
39}
40
41export const isOn = (v: string | undefined) => v === '1' || v?.toLowerCase() === 'true'
42
43export type TtlChoice = { ttl: Ttl; source: string }
44
45const asTtl = (v: unknown): Ttl | undefined => (v === '5m' || v === '1h' ? v : undefined)
46
47/**
48 * Which lifetime Claude Code asks for on the main conversation, in the order
49 * its documentation gives (https://code.claude.com/docs/en/prompt-caching,
50 * "Choose the TTL yourself"), after the mod's own `ttl` option:
51 *
52 * FORCE_PROMPT_CACHING_5M, CLAUDE_CODE_PROMPT_CACHE_TTL, the promptCacheTtl
53 * setting, ENABLE_PROMPT_CACHING_1H, then the default of the account: one
54 * hour on a Claude subscription within its plan usage, five minutes on usage
55 * credits, an API key or a cloud provider.
56 */
57export function decideTtl(option: unknown, env: CacheEnv, setting?: unknown, account?: Account): TtlChoice {
58 const pinned = asTtl(option)
59 if (pinned) return { ttl: pinned, source: 'the ttl option' }
60 if (isOn(env.force5m)) return { ttl: '5m', source: 'FORCE_PROMPT_CACHING_5M' }
61 const fromVar = asTtl(env.ttlVar)
62 if (fromVar) return { ttl: fromVar, source: 'CLAUDE_CODE_PROMPT_CACHE_TTL' }
63 const fromSetting = asTtl(setting)
64 if (fromSetting) return { ttl: fromSetting, source: 'the promptCacheTtl setting' }
65 if (isOn(env.enable1h)) return { ttl: '1h', source: 'ENABLE_PROMPT_CACHING_1H' }
66 if (account === 'subscription') return { ttl: '1h', source: 'Claude subscription default' }
67 if (account === 'credits') return { ttl: '5m', source: 'usage credits default' }
68 return { ttl: '5m', source: 'default' }
69}
70
71export const resolveTtl = (option: unknown, env: CacheEnv, setting?: unknown, account?: Account): Ttl =>
72 decideTtl(option, env, setting, account).ttl
73
74/**
75 * The account, from the rate-limit windows the last response reported: a
76 * five-hour or seven-day window means a Claude subscription, and one that is
77 * full means the next requests draw on usage credits. No window (an API key,
78 * a cloud provider, or no response yet) says nothing.
79 */
80export function accountOf(windows: readonly { kind: string; percentUsed: number }[]): Account {
81 const plan = windows.filter(w => w.kind === 'five_hour' || w.kind === 'seven_day')
82 if (plan.length === 0) return 'other'
83 return plan.some(w => w.percentUsed >= 100) ? 'credits' : 'subscription'
84}
85
86export function ttlMs(ttl: Ttl): number {
87 return ttl === '1h' ? 3_600_000 : 300_000
88}
89
90/** Caching switched off for this model by the environment. */
91export function isCachingDisabled(model: string, env: CacheEnv): boolean {
92 if (isOn(env.disableAll)) return true
93 const name = model.toLowerCase()
94 if (name.includes('haiku')) return isOn(env.disableHaiku)
95 if (name.includes('sonnet')) return isOn(env.disableSonnet)
96 if (name.includes('opus')) return isOn(env.disableOpus)
97 return false
98}
99
100export const promptTokens = (s: Sample) => s.read + s.write + s.fresh
101
102/** Share of the prompt the cache served, 0 to 1; 0 for an empty prompt. */
103export function hitRatio(s: Sample): number {
104 const total = promptTokens(s)
105 return total === 0 ? 0 : s.read / total
106}
107
108/** When the cache entry the sample touched lapses, ms since the epoch. */
109export const expiresAt = (s: Sample, ttl: Ttl) => s.startedAt + ttlMs(ttl)
110
111/** Zero for a request that read and wrote nothing: it created or refreshed no entry, so there is nothing to count down. */
112export function remainingMs(s: Sample, ttl: Ttl, now: number): number {
113 if (s.read + s.write === 0) return 0
114 return Math.max(0, expiresAt(s, ttl) - now)
115}
116
117/**
118 * Why a request that should have read the cache wrote it instead; undefined
119 * when it did not miss. A prompt that shrank is a /compact or /clear, not a
120 * miss, and the first request of a session has nothing to read.
121 */
122export function missReason(prev: Sample | undefined, cur: Sample, ttl: Ttl): string | undefined {
123 if (!prev) return undefined
124 const before = promptTokens(prev)
125 if (before === 0 || promptTokens(cur) < before * 0.7) return undefined
126 if (cur.read >= before * 0.5 || cur.write === 0) return undefined
127 if (cur.model !== prev.model) return `model changed (${prev.model} to ${cur.model})`
128 if (cur.startedAt - prev.startedAt > ttlMs(ttl)) return `the ${ttl} cache had lapsed`
129 return 'the prompt prefix changed (effort, tools, system prompt or CLAUDE.md)'
130}
131
132export function advise(last: Sample | undefined, prev: Sample | undefined, policy: Policy, now: number, disabled: boolean): Advice {
133 if (disabled) return { kind: 'off', text: 'prompt caching is off for this model (DISABLE_PROMPT_CACHING*)' }
134 if (!last) return { kind: 'cold', text: 'no request yet: the first one writes the cache' }
135 if (last.read + last.write === 0) {
136 return { kind: 'uncached', text: 'this request was not cached (prompt under the model minimum, or caching off)' }
137 }
138 const miss = missReason(prev, last, policy.ttl)
139 const left = remainingMs(last, policy.ttl, now)
140 const size = promptTokens(last)
141 if (left <= 0) {
142 const big = size >= policy.compactAtTokens
143 return {
144 kind: 'expired',
145 text: big
146 ? `expired: the next message rewrites ${fmtTokens(size)} tokens. /compact first, or /clear if the task is done`
147 : `expired: only ${fmtTokens(size)} tokens to rebuild, just keep going`,
148 }
149 }
150 if (left <= policy.warnMs) {
151 return { kind: 'soon', text: 'expires soon: any message refreshes it for free' }
152 }
153 if (miss) return { kind: 'miss', text: `cache missed: ${miss}` }
154 return { kind: 'warm', text: 'warm: keep going' }
155}
156
157export function fmtTokens(n: number): string {
158 if (n < 1000) return String(n)
159 if (n < 100_000) return `${(n / 1000).toFixed(1).replace(/\.0$/, '')}k`
160 if (n < 1_000_000) return `${Math.round(n / 1000)}k`
161 return `${(n / 1_000_000).toFixed(1).replace(/\.0$/, '')}M`
162}
163
164/** m:ss, or h:mm:ss from an hour up. */
165export function fmtClock(ms: number): string {
166 const total = Math.max(0, Math.ceil(ms / 1000))
167 const h = Math.floor(total / 3600)
168 const m = Math.floor((total % 3600) / 60)
169 const s = total % 60
170 const pad = (n: number) => String(n).padStart(2, '0')
171 return h > 0 ? `${h}:${pad(m)}:${pad(s)}` : `${m}:${pad(s)}`
172}
173
174export function bar(ratio: number, width: number): string {
175 const filled = Math.round(Math.min(1, Math.max(0, ratio)) * width)
176 return '█'.repeat(filled) + '░'.repeat(width - filled)
177}
178
179export type TurnRow = {
180 turnId: string
181 steps: number
182 read: number
183 write: number
184 fresh: number
185 output: number
186}
187
188/** Samples grouped by turn, oldest first, each turn's requests summed. */
189export function byTurn(samples: readonly Sample[]): TurnRow[] {
190 const rows: TurnRow[] = []
191 for (const s of samples) {
192 let row = rows[rows.length - 1]
193 if (!row || row.turnId !== s.turnId) {
194 row = { turnId: s.turnId, steps: 0, read: 0, write: 0, fresh: 0, output: 0 }
195 rows.push(row)
196 }
197 row.steps += 1
198 row.read += s.read
199 row.write += s.write
200 row.fresh += s.fresh
201 row.output += s.output
202 }
203 return rows
204}
205
206export const rowRatio = (r: TurnRow) => {
207 const total = r.read + r.write + r.fresh
208 return total === 0 ? 0 : r.read / total
209}
210
211export function fit(text: string, width: number): string {
212 return text.length <= width ? text : `${text.slice(0, Math.max(0, width - 1))}…`
213}
214
215export function positive(v: unknown, fallback: number): number {
216 return typeof v === 'number' && Number.isFinite(v) && v > 0 ? v : fallback
217}
218
219/** Share of the cache lifetime left, 0 to 1. */
220export function lifeRatio(leftMs: number, ttl: Ttl): number {
221 return Math.min(1, Math.max(0, leftMs / ttlMs(ttl)))
222}
223
224/**
225 * Widths of the three stacked-bar segments (read, wrote, new) over `width`
226 * cells: proportional, each non-empty part at least one cell, summing to width.
227 */
228export function segments(read: number, write: number, fresh: number, width: number): [number, number, number] {
229 const total = read + write + fresh
230 if (total === 0 || width <= 0) return [0, 0, 0]
231 const parts = [read, write, fresh]
232 const cells = parts.map(p => (p > 0 ? Math.max(1, Math.round((p / total) * width)) : 0))
233 let over = cells.reduce((a, b) => a + b, 0) - width
234 while (over !== 0) {
235 const i = over > 0 ? cells.indexOf(Math.max(...cells)) : parts.indexOf(Math.max(...parts))
236 cells[i]! += over > 0 ? -1 : 1
237 over += over > 0 ? -1 : 1
238 }
239 return [cells[0]!, cells[1]!, cells[2]!]
240}
241
242/** Seconds left at which a toast counts down after the one at the warning threshold. */
243export const COUNTDOWN_MARKS = [10, 3, 2, 1]
244
245/**
246 * The toast mark to fire now, or undefined. `level` is the mark last fired for
247 * this cache entry (Infinity before any); a late tick skips straight to the
248 * newest mark crossed, so a stalled clock never replays old ones.
249 */
250export function nextToastMark(secsLeft: number, warnSecs: number, level: number): number | undefined {
251 const marks = [warnSecs, ...COUNTDOWN_MARKS].filter(m => m <= warnSecs)
252 const due = marks.filter(m => secsLeft <= m && m < level)
253 return due.length ? Math.min(...due) : undefined
254}
255
256export type LifeColor = 'green' | 'yellow' | 'red'
257
258/** Countdown colour: green while there is plenty, yellow below 40% of the lifetime, red from the warning threshold down. */
259export function lifeColor(leftMs: number, ttl: Ttl, warnMs: number): LifeColor {
260 if (leftMs <= warnMs) return 'red'
261 return leftMs / ttlMs(ttl) <= 0.4 ? 'yellow' : 'green'
262}
263
264// requests are timed from their start, so a little slack keeps a hit that
265// landed just inside the lifetime from reading as proof of the longer one
266const SLACK_MS = 10_000
267
268/**
269 * What the traffic says about the cache lifetime, given the request before and
270 * `known`, what earlier requests already showed.
271 *
272 * - a hit (the cache served at least half of the previous prompt) more than
273 * 5 minutes after the previous request began proves the 1-hour lifetime,
274 * and nothing later undoes it: a miss afterwards is more likely a changed
275 * prefix than a lapse
276 * - a miss with the same model and a prompt that did not shrink, 5 minutes to
277 * an hour after the previous request, says the entry lapsed: 5 minutes
278 * (weaker: a changed prefix looks the same, so a later hit overrules it)
279 *
280 * Needed because the API names the TTL of a write (`cache_creation.ephemeral_*`)
281 * but Claude Code's mod API passes on only the four token counts.
282 */
283export function observeTtl(prev: Sample | undefined, cur: Sample, known: Ttl | undefined): Ttl | undefined {
284 if (!prev || prev.read + prev.write === 0 || cur.model !== prev.model) return known
285 const gap = cur.startedAt - prev.startedAt
286 const before = promptTokens(prev)
287 if (gap <= ttlMs('5m') + SLACK_MS) return known
288 if (cur.read >= before * 0.5) return '1h'
289 if (known === '1h') return known
290 const lapsed = cur.write > 0 && promptTokens(cur) >= before * 0.7 && gap < ttlMs('1h') + SLACK_MS
291 return lapsed ? '5m' : known
292}
293types/index.d.ts 60 lines1export type Ttl = '5m' | '1h'
2
3export type CacheEnv = {
4 enable1h?: string
5 force5m?: string
6 /** CLAUDE_CODE_PROMPT_CACHE_TTL: "5m" or "1h" for the main conversation */
7 ttlVar?: string
8 disableAll?: string
9 disableHaiku?: string
10 disableSonnet?: string
11 disableOpus?: string
12}
13
14/** One main-loop request, as the API reported it. */
15export type Sample = {
16 turnId: string
17 index: number
18 model: string
19 /** ms since the epoch when the request started: the cache's lifetime is counted from here */
20 startedAt: number
21 read: number
22 write: number
23 fresh: number
24 output: number
25}
26
27/** What the account is billed as, as far as the mod can tell. */
28export type Account = 'subscription' | 'credits' | 'other'
29
30/** Everything the band and the pane draw from, held by the host so a reload keeps it. */
31export type Meter = {
32 samples: Sample[]
33 /** The lifetime in use: the observed one when traffic proved it, else `baseTtl`. */
34 ttl: Ttl
35 /** The lifetime Claude Code's rules give (option, environment, setting, account). */
36 baseTtl: Ttl
37 /** Set by the `ttl` option; nothing overrides it. */
38 pinned: boolean
39 /** The lifetime request timing proved, or null. */
40 observed: Ttl | null
41 account: Account
42 ttlSource: string
43 envSource: string
44 env: CacheEnv
45 /** The promptCacheTtl setting, as read at session start. */
46 setting: string | null
47}
48
49declare module 'claude-code' {
50 interface PluginState {
51 'prompt-cache-control': {
52 meter: Meter
53 // The band steps aside while /cache is open.
54 paneOpen: boolean
55 // The bar above the prompt: /cache on, /cache off (kept in $.store).
56 barOn: boolean
57 }
58 }
59}
60