SLOPSHOPPER

prompt-cache-control

Prompt-cache meter above the Claude Code prompt: how many tokens each request read from, wrote to and sent past the cache, a live countdown to the cache's…

newpanebandcommandtoaststatus
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · prompt-cache-control
│ ┃ cache ✕ › fix the failing auth test and add an audit log call │ ┃ ⚡ PROMPT CACHE · 1h lifetime (Claude subscri │ ┃ ● prompt-cache-control: prompt-cache-control loaded: 1h cache (Claude │ ┃ ⏱ --:-- ⏺ Read(src/auth.ts) │ ┃ ⎿ Read 6 lines │ ┃ ○ no request yet: the first one writes the ⏺ Update(src/auth.ts) │ ┃ cache ⎿ Added 2 lines, removed 1 line │ ┃ ⏺ Bash(bun test) │ ┃ turn steps read wrote new hit ⎿ 3 pass, 1 fail │ ┃ no requests yet │ ┃ ● Done. refresh now rejects expired claims and logs an audit event. │ ┃ [ close ] │ ┃ ✻ Worked for 42s · done 4:20 PM │ ┃ ■ read: served by the cache │ ┃ ■ wrote: new cache entry › /cache │ ┃ ■ new: sent uncached ⎿ prompt-cache-control: 1h cache (Claude subscription default) · n │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · cache
⚡ PROMPT CACHE · 1h lifetime (Claude subscription default) ⏱ --:-- ○ no request yet: the first one writes the cache turn steps read wrote new hit no requests yet [ close ] ■ read: served by the cache ■ wrote: new cache entry ■ new: sent uncached
README

prompt-cache-control (Claude Cache Control)

A prompt-cache meter above the Claude Code prompt. Every request Claude makes reports how much of its prompt the cache served, how much it wrote and how much went uncached; this mod keeps those numbers per request and per turn, counts down to the moment the cache lapses and tells you what to do about it: keep going, /compact or /clear.

cache ██████████ 98% read 80k · wrote 1k · new 300 ⏱ 3:41 5m · warm: keep going
cache ░░░░░░░░░░  0% read 0 · wrote 52k · new 300 ⏱ 4:58 5m · cache missed: model changed (…)
cache ██████████ 98% read 150k · wrote 1k · new 300 ⏱ 0:00 5m · expired: the next message rewrites 151k tokens. /compact first, or /clear if the task is done

/cache opens a pane: the time left with a solid bar that shrinks as the cache runs out (green, yellow below 40% of the lifetime, red from the warning threshold), a stacked read / wrote / new bar for the last request, and a colour-coded table with one row per turn. Bars are filled cells and columns have fixed widths with no-break spaces, so the terminal and the Desktop (HTML) pane render the same.

How the countdown works

From Anthropic's prompt caching documentation:

  • The cache lives 5 minutes by default, 1 hour when asked for.
  • Every request that reads the cache refreshes it at no extra cost, so a conversation that keeps talking keeps the 5-minute cache warm.
  • The lifetime is counted from the start of the request that wrote or read the entry; generation time counts against it.
  • A prompt is input_tokens (uncached remainder) + cache_read_input_tokens + cache_creation_input_tokens.
  • Writes cost 1.25x base input for 5 minutes and 2x for 1 hour; reads cost about 0.1x (less on some models). The expensive moment is an expired cache on a large context, which is when this mod suggests /compact.
  • /clear starts a new conversation in the same process, so the meter and the /cache table start over with it. A change in the prefix (model, effort or thinking settings, tool set, system prompt, CLAUDE.md) makes the next request write instead of read. The mod names the cause when it sees a miss: model changed, the cache had lapsed, or the prefix changed.

Which lifetime your account gets

The mod follows Claude Code's own rules (prompt caching: cache lifetime, Claude Code 2.1.242 or later). For the main conversation the TTL is the first match of:

#SourceResult
1the mod's ttl option (5m / 1h)what you set
2FORCE_PROMPT_CACHING_5M=15 minutes
3CLAUDE_CODE_PROMPT_CACHE_TTL5m or 1h
4the promptCacheTtl setting (local, project or user settings file)5m or 1h
5ENABLE_PROMPT_CACHING_1H=11 hour
6the account1 hour on a Claude subscription within its plan usage; 5 minutes on usage credits, an API key or a cloud provider

The account comes from the rate-limit windows the last response reported: a five_hour or seven_day window means a subscription, and one at 100% means requests now draw on usage credits. An API key or a cloud provider reports no such window, and before the first response nothing is known, so the mod starts from 5 minutes there. Managed settings are not readable from a mod.

On top of that the mod watches the traffic, which beats rows 2 to 6: a request that hits the cache more than 5 minutes after the previous one proves the 1-hour lifetime (a later miss does not undo it, since a changed prefix looks the same), and a miss 5 to 60 minutes after the previous request, with the same model and a prompt that did not shrink, says the entry lapsed, so 5 minutes (a later hit overrules it). That covers what the mod cannot see: managed settings, a gateway that rewrites the TTL, or a subscription that ran out of plan usage mid-session. The pane header names the source in use.

Why the mod infers instead of reading it: the API names the TTL of each write (cache_creation.ephemeral_5m_input_tokens / ephemeral_1h_input_tokens) and Claude Code's status line exposes it as prompt_cache.ttl, but the mod API passes on only the four token counts. To check by hand, claude -p "hello" --output-format json and read usage.cache_creation.

Other switches read from the environment at session start:

VariableEffect on the meter
DISABLE_PROMPT_CACHING=1 (and _HAIKU, _SONNET, _OPUS)the band says caching is off for that model

What it hooks

  • turn.step: reads each main-loop request's usage (subagents have their own prefixes and are left out)
  • $.clock.every(1000): redraws the countdown, and only while its text changes, so an idle expired session costs nothing
  • ui.render on AbovePrompt (the band) and on Pane (/cache)
  • $.ui.toast: once per cache entry at the warning threshold (60 s by default) and again at 10, 3, 2 and 1 seconds left, for prompts of 20k tokens or more

Options

  ttl: string               "auto" | "5m" | "1h" (default auto)
  warnSeconds: number       countdown threshold for the yellow state and the toast (default 60)
  compactAtTokens: number   prompt size that makes an expired cache suggest /compact (default 100000)
  band: boolean             row above the prompt (default true)
  status: boolean           entry under the prompt, "cache 98% · 3:41" (default false)
  toast: boolean            toasts at the threshold, 10, 3, 2 and 1 s (default true)

The 100k compactAtTokens is a judgement, not a figure from the documentation: lower it if your model's cache writes are expensive for you.

Install

From the claudemods marketplace:

claude plugin marketplace add jakerains/claudemods
claude plugin install prompt-cache-control@claudemods

/cache opens the table, /cache stop closes it, /cache off and /cache on hide or show the bar above the prompt (remembered).

Options are read from user settings (~/.claude/settings.json) under the plugin's full id:

{ "pluginConfigs": { "prompt-cache-control@claudemods": { "options": { } } } }

Changes from the claude-code-templates original

  • What the band and pane draw from lives in $.state, so /reload-plugins and hot reloads keep the history and countdown (the original reset them).
  • The band composes with other mods' rows above the prompt (where-we-are, context-gauge) instead of replacing them, and sits last, on the input.
  • Strict type-check fixes (noUncheckedIndexedAccess).

Early access. Mods need Claude Code 2.1.259+ with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1; the $ API may change between releases. Typed against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods

Source 3 files
hooks/prompt-cache-control.tsx 497 lines
1/**
2 * prompt-cache-control — Claude Mod (EARLY ACCESS)
3 *
4 * A prompt-cache meter for Claude Code. Every main-loop request reports how
5 * many prompt tokens the cache served (`cache_read_input_tokens`), wrote
6 * (`cache_creation_input_tokens`) and sent uncached (`input_tokens`); this mod
7 * keeps those per request and per turn, counts down to the moment the cache
8 * lapses, and says what to do about it: keep going, /compact or /clear.
9 *
10 *   - `turn.step` reads each main-loop request's usage (subagents have their
11 *     own prefixes and are left out)
12 *   - `$.clock.every(1000)` redraws the countdown, and only while its text
13 *     changes: an idle, expired session costs nothing
14 *   - a row above the prompt (the AbovePrompt component), an optional status
15 *     line entry, and `/cache`, a pane with one row per turn
16 *
17 * The lifetime is counted from the start of the request that last wrote or read
18 * the cache, as Anthropic documents it. Which lifetime Claude Code asked for
19 * follows its documented rules (see decideTtl in ./cache.ts): FORCE_PROMPT_CACHING_5M,
20 * CLAUDE_CODE_PROMPT_CACHE_TTL, the promptCacheTtl setting, ENABLE_PROMPT_CACHING_1H,
21 * then the account (1 hour on a Claude subscription, 5 minutes otherwise). The
22 * API names the TTL of a write but the mod API passes on only the token counts,
23 * so the mod also watches the gaps between requests (a hit after more than 5
24 * minutes proves 1 hour; see observeTtl). `ttl: "5m" | "1h"` pins it.
25 *
26 * Needs CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 (Claude Code >= 2.1.259).
27 *
28 * What the band and the pane draw from lives in $.state (the Meter atom), so a
29 * hot reload or /reload-plugins keeps the history and the countdown. The band
30 * composes with other mods' rows above the prompt and sits last, on the input.
31 *
32 * Options (pluginConfigs["prompt-cache-control@claudemods"].options):
33 *   ttl: "auto" | "5m" | "1h"   cache lifetime (default auto)
34 *   warnSeconds: number         countdown threshold for the warning (default 60)
35 *   compactAtTokens: number     prompt size that makes an expired cache suggest /compact (default 100000)
36 *   band: boolean               row above the prompt (default true)
37 *   status: boolean             entry under the prompt (default false)
38 *   toast: boolean              toasts near expiry: at warnSeconds, then 10, 3, 2 and 1 s (default true)
39 */
40import { atom, read, update } from 'claude-code'
41import type { EngineInterface, Register } from 'claude-code'
42import {
43  advise,
44  COUNTDOWN_MARKS,
45  bar,
46  byTurn,
47  fit,
48  fmtClock,
49  fmtTokens,
50  hitRatio,
51  isCachingDisabled,
52  accountOf,
53  decideTtl,
54  observeTtl,
55  lifeColor,
56  lifeRatio,
57  nextToastMark,
58  positive,
59  promptTokens,
60  remainingMs,
61  rowRatio,
62  segments,
63} from './cache.ts'
64import type { Advice, CacheEnv } from './cache.ts'
65import type { Meter } from '../types'
66
67const PANE = 'cache'
68const COMMAND = 'cache'
69const KEEP = 200
70// below this a lapsed cache costs too little to interrupt anyone about
71const TOAST_MIN_TOKENS = 20_000
72
73const EMPTY: Meter = {
74  samples: [],
75  ttl: '5m',
76  baseTtl: '5m',
77  pinned: false,
78  observed: null,
79  account: 'other',
80  ttlSource: 'default',
81  envSource: 'default',
82  env: {},
83  setting: null,
84}
85
86// Held by the host: a reload starts this module over but keeps these.
87const meter = atom({ plugin: 'prompt-cache-control', key: 'meter' } as const, EMPTY)
88const paneOpen = atom({ plugin: 'prompt-cache-control', key: 'paneOpen' } as const, false)
89const barOn = atom({ plugin: 'prompt-cache-control', key: 'barOn' } as const, true)
90
91// The ticker's own bookkeeping; a reload restarts the ticker with it.
92let timer: { cancel: () => void } | undefined
93let lastKey = ''
94let toastedFor = 0
95let toastLevel = Infinity
96
97type Policy = { warnMs: number; compactAtTokens: number }
98
99function current(m: Meter, policy: Policy, now: number) {
100  const last = m.samples[m.samples.length - 1]
101  const prev = m.samples[m.samples.length - 2]
102  const disabled = last ? isCachingDisabled(last.model, m.env) : isCachingDisabled('', m.env)
103  const advice: Advice = advise(last, prev, { ttl: m.ttl, ...policy }, now, disabled)
104  const left = last ? remainingMs(last, m.ttl, now) : 0
105  return { last, advice, left }
106}
107
108const COLOR: Record<Advice['kind'], string | undefined> = {
109  warm: 'green',
110  soon: 'yellow',
111  expired: 'red',
112  miss: 'red',
113  off: undefined,
114  cold: undefined,
115  uncached: undefined,
116}
117
118function shortLine(m: Meter, policy: Policy, now: number): string {
119  const { last, advice, left } = current(m, policy, now)
120  if (!last || advice.kind === 'off') return `cache: ${advice.text}`
121  const clock = left > 0 ? ` · ${fmtClock(left)}` : ''
122  return `cache ${Math.round(hitRatio(last) * 100)}%${clock}`
123}
124
125// the promptCacheTtl setting, from the settings files that can carry it (local over project over user)
126async function readSetting($: EngineInterface): Promise<string | null> {
127  const home = await $.env.get('HOME').catch(() => undefined)
128  const cwd = await $.session.cwd().catch(() => undefined)
129  const files = [cwd && `${cwd}/.claude/settings.local.json`, cwd && `${cwd}/.claude/settings.json`, home && `${home}/.claude/settings.json`]
130  for (const file of files) {
131    if (!file) continue
132    try {
133      const value = JSON.parse(await $.fs.read(file)).promptCacheTtl
134      if (value === '5m' || value === '1h') return value
135    } catch {
136      // missing or unreadable: the next file
137    }
138  }
139  return null
140}
141
142export const register: Register = (on, options) => {
143  const policy: Policy = {
144    warnMs: positive(options.warnSeconds, 60) * 1000,
145    compactAtTokens: positive(options.compactAtTokens, 100_000),
146  }
147  const showBand = options.band !== false
148  const showStatus = options.status === true
149  const wantToast = options.toast !== false
150
151  on('session.start', async ($, e, next) => {
152    const r = await next(e)
153    lastKey = ''
154    toastedFor = 0
155    // A pinned status outlives a reload: with the option off, take down any
156    // entry an earlier load (or the option, before it was turned off) left.
157    if (!showStatus) $.ui.status(undefined)
158    const none = () => undefined
159    const env: CacheEnv = {
160      enable1h: await $.env.get('ENABLE_PROMPT_CACHING_1H').catch(none),
161      force5m: await $.env.get('FORCE_PROMPT_CACHING_5M').catch(none),
162      ttlVar: await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL').catch(none),
163      disableAll: await $.env.get('DISABLE_PROMPT_CACHING').catch(none),
164      disableHaiku: await $.env.get('DISABLE_PROMPT_CACHING_HAIKU').catch(none),
165      disableSonnet: await $.env.get('DISABLE_PROMPT_CACHING_SONNET').catch(none),
166      disableOpus: await $.env.get('DISABLE_PROMPT_CACHING_OPUS').catch(none),
167    }
168    const pinned = options.ttl === '5m' || options.ttl === '1h'
169    const setting = await readSetting($)
170    const account = accountOf((await $.session.usage().catch(() => undefined))?.rateLimits ?? [])
171    const choice = decideTtl(options.ttl, env, setting ?? undefined, account)
172    // A reload fires session.start again: the requests seen so far, and any
173    // lifetime they proved, stay. A new session's state starts empty anyway.
174    const m = await update($, meter, m => {
175      const observed = pinned ? null : m.observed
176      return {
177        ...m,
178        env,
179        pinned,
180        setting,
181        account,
182        observed,
183        baseTtl: choice.ttl,
184        envSource: choice.source,
185        ttl: observed ?? choice.ttl,
186        ttlSource: observed ? m.ttlSource : choice.source,
187      }
188    })
189
190    await $.command
191      .register({
192        name: COMMAND,
193        description: 'Prompt-cache table per turn (stop closes it); on / off shows or hides the bar above the prompt',
194        argumentHint: '[on|off|stop]',
195        immediate: true,
196      })
197      .catch(err => $.ui.log(`prompt-cache-control: /${COMMAND} not registered: ${err}`))
198    $.ui.log(`prompt-cache-control loaded: ${m.ttl} cache (${m.ttlSource}), /${COMMAND} opens the table`, { to: 'debug' })
199
200    // The bar's on/off is the person's last /cache on|off, else the `band` option.
201    const kept = await $.store.get('bar').catch(() => undefined)
202    await update($, barOn, () => (typeof kept === 'boolean' ? kept : showBand))
203
204    timer?.cancel()
205    timer = $.clock.every(1000, async () => {
206      const now = Date.now()
207      const m = await read($, meter)
208      const { last, advice, left } = current(m, policy, now)
209      const key = `${advice.kind}|${advice.text}|${left > 0 ? fmtClock(left) : ''}`
210      if (key !== lastKey) {
211        lastKey = key
212        if (showStatus) $.ui.status(shortLine(m, policy, now))
213        $.ui.invalidate('ui.render')
214      }
215      if (wantToast && last && left > 0 && promptTokens(last) >= TOAST_MIN_TOKENS) {
216        if (toastedFor !== last.startedAt) {
217          toastedFor = last.startedAt
218          toastLevel = Infinity
219        }
220        // the first toast comes at warnSeconds, then 10, 3, 2 and 1 seconds; a late tick skips to the newest one
221        const secs = Math.ceil(left / 1000)
222        const mark = nextToastMark(secs, policy.warnMs / 1000, toastLevel)
223        if (mark !== undefined) {
224          toastLevel = mark
225          const tail = secs <= COUNTDOWN_MARKS[0]! ? 'send a message now' : `send a message to keep ${fmtTokens(promptTokens(last))} tokens warm`
226          $.ui.toast(`cache expires in ${secs >= 60 ? fmtClock(left) : `${secs}s`}: ${tail}`)
227        }
228      }
229    })
230    return r
231  })
232
233  on('session.end', async ($, e, next) => {
234    // /clear starts a new conversation in the same process: its cache is a new one
235    if (e.reason === 'clear') {
236      lastKey = ''
237      toastedFor = 0
238      await update($, meter, m => ({ ...m, samples: [], observed: null, ttl: m.baseTtl, ttlSource: m.envSource }))
239      $.ui.invalidate('ui.render')
240      return next(e)
241    }
242    timer?.cancel()
243    timer = undefined
244    return next(e)
245  })
246
247  // each main-loop request: what the cache did with it
248  on('turn.step', async function* ($, e, next) {
249    if (e.agentId) return yield* next(e)
250    const startedAt = Date.now()
251    const r = yield* next(e)
252    if (r.usage) {
253      const usage = r.usage
254      const before = await read($, meter)
255      // the account can change under a session: a subscription running out of plan usage moves to usage credits
256      const account = before.pinned
257        ? before.account
258        : accountOf((await $.session.usage().catch(() => undefined))?.rateLimits ?? [])
259      const m = await update($, meter, m => {
260        const samples = [
261          ...m.samples,
262          {
263            turnId: e.turnId,
264            index: e.index,
265            model: usage.model || e.model,
266            startedAt,
267            read: usage.cache_read_input_tokens,
268            write: usage.cache_creation_input_tokens,
269            fresh: usage.input_tokens,
270            output: usage.output_tokens,
271          },
272        ].slice(-KEEP)
273        if (m.pinned) return { ...m, samples }
274        const choice = decideTtl(options.ttl, m.env, m.setting ?? undefined, account)
275        const moved = { ...m, samples, account, baseTtl: choice.ttl, envSource: choice.source }
276        const known = m.observed ?? undefined
277        const seen = observeTtl(samples[samples.length - 2], samples[samples.length - 1]!, known)
278        if (seen !== known) {
279          return {
280            ...moved,
281            observed: seen ?? null,
282            ttl: seen ?? choice.ttl,
283            ttlSource: `observed from request timing; ${choice.source} said ${choice.ttl}`,
284          }
285        }
286        return m.observed ? moved : { ...moved, ttl: choice.ttl, ttlSource: choice.source }
287      })
288      if (m.observed !== before.observed) {
289        $.ui.log(`prompt-cache-control: cache lifetime is ${m.ttl} (${m.ttlSource})`, { to: 'debug' })
290      }
291      lastKey = ''
292      if (showStatus) $.ui.status(shortLine(m, policy, Date.now()))
293      $.ui.invalidate('ui.render')
294    }
295    return r
296  })
297
298  on('command.run', { command: COMMAND }, async ($, e) => {
299    const word = e.args.trim().toLowerCase()
300    // Turning the bar off draws nothing above the prompt, so the engine has no
301    // band to collapse and no "panel hidden" line to show.
302    if (word === 'on' || word === 'off') {
303      const isOn = word === 'on'
304      await $.store.set('bar', isOn)
305      await update($, barOn, () => isOn)
306      return { text: isOn ? 'cache bar on, above the prompt' : `cache bar off · /${COMMAND} on brings it back, /${COMMAND} opens the table` }
307    }
308    if (word === 'stop') {
309      await $.ui.close({ id: PANE }).catch(() => undefined)
310      await update($, paneOpen, () => false)
311      return { text: 'cache table closed' }
312    }
313    await update($, paneOpen, () => true)
314    await $.ui.open({ id: PANE, title: 'cache', focus: true })
315    $.ui.invalidate('ui.render')
316    const m = await read($, meter)
317    const { advice } = current(m, policy, Date.now())
318    return { text: `${m.ttl} cache (${m.ttlSource}) · ${advice.text} · /${COMMAND} stop closes` }
319  })
320
321  on('ui.close', async ($, e, next) => {
322    if (e.id !== PANE) return next(e)
323    await update($, paneOpen, () => false)
324    return next(e)
325  })
326
327  on('ui.press', async ($, e, next) => {
328    if (e.plugin !== $.plugin.name || e.requestId !== PANE) return next(e)
329    if (e.element === 'close') await $.ui.close({ id: PANE }).catch(() => undefined)
330    return next(e)
331  })
332
333  // Composes: other mods (where-we-are, context-gauge) draw in this band too.
334  // Theirs first, the meter last, right on the input.
335  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
336    const others = await next(e)
337    if (!(await read($, barOn)) || e.props.hasSurvey || (await read($, paneOpen))) return others
338    const m = await read($, meter)
339    const { last, advice, left } = current(m, policy, Date.now())
340    if (!last && advice.kind !== 'off') return others
341    const { Box, Text } = $.ui.resolve(e)
342    const columns = e.props.bodyColumns ?? e.viewport?.columns ?? 100
343    const color = COLOR[advice.kind]
344    const ttl = m.ttl
345
346    if (!last) {
347      return (
348        <Box flexDirection="column">
349          {others}
350          <Text dimColor>{fit(`cache: ${advice.text}`, columns)}</Text>
351        </Box>
352      )
353    }
354
355    const ratio = hitRatio(last)
356    const wide = columns >= 90
357    const meterRow = (
358      <Box flexDirection="row" columnGap={1}>
359        <Text bold color={color}>{advice.kind === 'warm' ? '●' : advice.kind === 'soon' ? '▲' : advice.kind === 'off' || advice.kind === 'cold' || advice.kind === 'uncached' ? '○' : '✖'}</Text>
360        <Text bold color="cyan">cache</Text>
361        <Text color={color}>{bar(ratio, wide ? 10 : 6)}</Text>
362        <Text bold>{`${Math.round(ratio * 100)}%`}</Text>
363        {wide ? (
364          // a row of its own: a fragment here stacks its children as a column
365          <Box flexDirection="row" columnGap={1}>
366            <Text color="green">{`read ${fmtTokens(last.read)}`}</Text>
367            <Text color="yellow">{`wrote ${fmtTokens(last.write)}`}</Text>
368            <Text color="cyan">{`new ${fmtTokens(last.fresh)}`}</Text>
369          </Box>
370        ) : (
371          <Text dimColor>{`${fmtTokens(promptTokens(last))} tok`}</Text>
372        )}
373        {advice.kind !== 'uncached' && advice.kind !== 'off' && (
374          <Text bold color={left > 0 ? lifeColor(left, ttl, policy.warnMs) : 'red'}>{left > 0 ? `⏱ ${fmtClock(left)}` : '⏱ 0:00'}</Text>
375        )}
376        <Text dimColor wrap="truncate-end">{`${ttl} · ${advice.text}`}</Text>
377      </Box>
378    )
379
380    return (
381      <Box flexDirection="column">
382        {others}
383        {meterRow}
384      </Box>
385    )
386  })
387
388  on('ui.render', { component: 'Pane' }, async ($, e, next) => {
389    if (e.requestId !== PANE) return next(e)
390    const { Box, Text, Button } = $.ui.resolve(e)
391    const width = Math.max(30, e.props.bodyColumns - 1)
392    // HTML collapses runs of spaces and trims a text's ends; a no-break space keeps them
393    const sp = (t: string) => (e.surface === 'terminal' ? t : t.replace(/ /g, ' '))
394    const now = Date.now()
395    const m = await read($, meter)
396    const { ttl, ttlSource } = m
397    const { last, advice, left } = current(m, policy, now)
398    const all = byTurn(m.samples)
399    const counting = !!last && advice.kind !== 'uncached' && advice.kind !== 'off'
400    // the countdown goes green, then yellow, then red as the cache runs out
401    const clockColor = counting ? lifeColor(left, ttl, policy.warnMs) : undefined
402    const stateColor = advice.kind === 'expired' || advice.kind === 'miss' ? 'red' : (clockColor ?? COLOR[advice.kind])
403    const hitColor = (pct: number) => (pct >= 80 ? 'green' : pct >= 40 ? 'yellow' : 'red')
404    // solid bars are filled Boxes, not block characters, so HTML draws no seams between cells
405    const solid = (key: string, parts: [number, string | undefined][]) => (
406      <Box key={key} flexDirection="row" height={1} flexShrink={0}>
407        {parts.map(([w, c], i) => (w > 0 ? <Box key={`${key}:${i}`} width={w} height={1} flexShrink={0} backgroundColor={c} /> : null))}
408      </Box>
409    )
410    const cell = (key: string, w: number, text: string, c?: string, bold = false) => (
411      <Box key={key} width={w} flexShrink={0} justifyContent="flex-end">
412        <Text color={c} bold={bold} dimColor={!c}>{sp(text)}</Text>
413      </Box>
414    )
415
416    const barW = Math.min(width, 40)
417    const life = lifeRatio(left, ttl)
418    const lifeFilled = Math.round(life * barW)
419    const [sr, sw, sn] = last ? segments(last.read, last.write, last.fresh, barW) : [0, 0, 0]
420    const rows = all.slice(-Math.max(3, (e.viewport?.rows ?? 24) - 16))
421    const icon = advice.kind === 'warm' ? '●' : advice.kind === 'soon' ? '▲' : advice.kind === 'expired' || advice.kind === 'miss' ? '✖' : '○'
422
423    return (
424      <Box flexDirection="column">
425        <Box key="title" flexDirection="row" columnGap={1}>
426          <Text bold color="cyan">{sp('⚡ PROMPT CACHE')}</Text>
427          <Text dimColor>{sp(`· ${ttl} lifetime (${ttlSource})`)}</Text>
428        </Box>
429
430        <Box key="clock" flexDirection="column" marginTop={1}>
431          <Text bold color={clockColor}>{sp(counting ? `⏱ ${left > 0 ? fmtClock(left) : '0:00'}` : '⏱ --:--')}</Text>
432          {counting ? (
433            <Box flexDirection="row" columnGap={1}>
434              {solid('life', [[lifeFilled, clockColor], [barW - lifeFilled, 'gray']])}
435              <Text dimColor>{sp(`${Math.round(life * 100)}%`)}</Text>
436            </Box>
437          ) : null}
438        </Box>
439
440        <Box key="advice" marginTop={1} flexDirection="column">
441          <Text bold color={stateColor}>{sp(`${icon} ${advice.text}`)}</Text>
442          {last ? <Text dimColor>{sp(fit(`${last.model} · prompt ${fmtTokens(promptTokens(last))} tokens`, width))}</Text> : null}
443        </Box>
444
445        {last ? (
446          <Box key="stack" flexDirection="column" marginTop={1}>
447            <Box flexDirection="row" columnGap={1}>
448              {solid('stack', [[sr, 'green'], [sw, 'yellow'], [sn, 'cyan']])}
449              <Text bold color={hitColor(Math.round(hitRatio(last) * 100))}>{sp(`${Math.round(hitRatio(last) * 100)}% hit`)}</Text>
450            </Box>
451            <Box flexDirection="row" columnGap={2}>
452              <Text color="green">{sp(`■ read ${fmtTokens(last.read)}`)}</Text>
453              <Text color="yellow">{sp(`■ wrote ${fmtTokens(last.write)}`)}</Text>
454              <Text color="cyan">{sp(`■ new ${fmtTokens(last.fresh)}`)}</Text>
455            </Box>
456          </Box>
457        ) : null}
458
459        <Box key="table" flexDirection="column" marginTop={1}>
460          <Box key="head" flexDirection="row" columnGap={1}>
461            {cell('h:turn', 4, 'turn', 'cyan', true)}
462            {cell('h:steps', 5, 'steps', 'cyan', true)}
463            {cell('h:read', 6, 'read', 'green', true)}
464            {cell('h:wrote', 6, 'wrote', 'yellow', true)}
465            {cell('h:new', 5, 'new', 'cyan', true)}
466            {cell('h:hit', 4, 'hit', 'magenta', true)}
467          </Box>
468          {rows.length === 0 ? <Text dimColor>{sp('no requests yet')}</Text> : null}
469          {rows.map((row, i) => {
470            const n = all.length - rows.length + i + 1
471            const pct = Math.round(rowRatio(row) * 100)
472            return (
473              <Box key={`t:${row.turnId}`} flexDirection="row" columnGap={1}>
474                {cell(`c:turn:${row.turnId}`, 4, String(n))}
475                {cell(`c:steps:${row.turnId}`, 5, String(row.steps))}
476                {cell(`c:read:${row.turnId}`, 6, fmtTokens(row.read), 'green')}
477                {cell(`c:wrote:${row.turnId}`, 6, fmtTokens(row.write), 'yellow')}
478                {cell(`c:new:${row.turnId}`, 5, fmtTokens(row.fresh), 'cyan')}
479                {cell(`c:hit:${row.turnId}`, 4, `${pct}%`, hitColor(pct), true)}
480              </Box>
481            )
482          })}
483        </Box>
484
485        <Box key="foot" marginTop={1} flexDirection="column">
486          <Button key="close" label="close" onPress={() => {}} />
487          <Box key="legend" marginTop={1} flexDirection="column">
488            <Text color="green">{sp('■ read: served by the cache')}</Text>
489            <Text color="yellow">{sp('■ wrote: new cache entry')}</Text>
490            <Text color="cyan">{sp('■ new: sent uncached')}</Text>
491          </Box>
492        </Box>
493      </Box>
494    )
495  })
496}
497
hooks/cache.ts 293 lines
1/**
2 * cache.ts — the pure half of prompt-cache-control: no `$`, no engine.
3 *
4 * What it models, from Anthropic's prompt-caching documentation:
5 *   - the cache lives 5 minutes by default, 1 hour when asked for; a read
6 *     refreshes the entry at no extra cost, and the lifetime is measured from
7 *     the START of the request that wrote or read it
8 *   - a request's prompt is `input_tokens` (uncached remainder) +
9 *     `cache_read_input_tokens` + `cache_creation_input_tokens`
10 *   - writes cost 1.25x base input for 5m and 2x for 1h; reads about 0.1x
11 *     (less on some models), so an expired cache on a large context is the
12 *     expensive moment
13 *   - a prefix change (model, effort/thinking settings, tool set, system
14 *     prompt) makes the next request write instead of read
15 *
16 * Claude Code's own switches (read from the environment):
17 *   ENABLE_PROMPT_CACHING_1H=1   ask for the 1-hour TTL
18 *   FORCE_PROMPT_CACHING_5M=1    force the 5-minute TTL, beating the above
19 *   DISABLE_PROMPT_CACHING=1     no caching; DISABLE_PROMPT_CACHING_{HAIKU,SONNET,OPUS}
20 *                                 turn it off for that model family only
21 */
22
23import type { Account, CacheEnv, Sample, Ttl } from '../types'
24
25export type { Account, CacheEnv, Sample, Ttl }
26
27export type AdviceKind = 'off' | 'cold' | 'uncached' | 'warm' | 'soon' | 'expired' | 'miss'
28
29export type Advice = {
30  kind: AdviceKind
31  /** one sentence for the band */
32  text: string
33}
34
35export type Policy = {
36  ttl: Ttl
37  warnMs: number
38  compactAtTokens: number
39}
40
41export const isOn = (v: string | undefined) => v === '1' || v?.toLowerCase() === 'true'
42
43export type TtlChoice = { ttl: Ttl; source: string }
44
45const asTtl = (v: unknown): Ttl | undefined => (v === '5m' || v === '1h' ? v : undefined)
46
47/**
48 * Which lifetime Claude Code asks for on the main conversation, in the order
49 * its documentation gives (https://code.claude.com/docs/en/prompt-caching,
50 * "Choose the TTL yourself"), after the mod's own `ttl` option:
51 *
52 *   FORCE_PROMPT_CACHING_5M, CLAUDE_CODE_PROMPT_CACHE_TTL, the promptCacheTtl
53 *   setting, ENABLE_PROMPT_CACHING_1H, then the default of the account: one
54 *   hour on a Claude subscription within its plan usage, five minutes on usage
55 *   credits, an API key or a cloud provider.
56 */
57export function decideTtl(option: unknown, env: CacheEnv, setting?: unknown, account?: Account): TtlChoice {
58  const pinned = asTtl(option)
59  if (pinned) return { ttl: pinned, source: 'the ttl option' }
60  if (isOn(env.force5m)) return { ttl: '5m', source: 'FORCE_PROMPT_CACHING_5M' }
61  const fromVar = asTtl(env.ttlVar)
62  if (fromVar) return { ttl: fromVar, source: 'CLAUDE_CODE_PROMPT_CACHE_TTL' }
63  const fromSetting = asTtl(setting)
64  if (fromSetting) return { ttl: fromSetting, source: 'the promptCacheTtl setting' }
65  if (isOn(env.enable1h)) return { ttl: '1h', source: 'ENABLE_PROMPT_CACHING_1H' }
66  if (account === 'subscription') return { ttl: '1h', source: 'Claude subscription default' }
67  if (account === 'credits') return { ttl: '5m', source: 'usage credits default' }
68  return { ttl: '5m', source: 'default' }
69}
70
71export const resolveTtl = (option: unknown, env: CacheEnv, setting?: unknown, account?: Account): Ttl =>
72  decideTtl(option, env, setting, account).ttl
73
74/**
75 * The account, from the rate-limit windows the last response reported: a
76 * five-hour or seven-day window means a Claude subscription, and one that is
77 * full means the next requests draw on usage credits. No window (an API key,
78 * a cloud provider, or no response yet) says nothing.
79 */
80export function accountOf(windows: readonly { kind: string; percentUsed: number }[]): Account {
81  const plan = windows.filter(w => w.kind === 'five_hour' || w.kind === 'seven_day')
82  if (plan.length === 0) return 'other'
83  return plan.some(w => w.percentUsed >= 100) ? 'credits' : 'subscription'
84}
85
86export function ttlMs(ttl: Ttl): number {
87  return ttl === '1h' ? 3_600_000 : 300_000
88}
89
90/** Caching switched off for this model by the environment. */
91export function isCachingDisabled(model: string, env: CacheEnv): boolean {
92  if (isOn(env.disableAll)) return true
93  const name = model.toLowerCase()
94  if (name.includes('haiku')) return isOn(env.disableHaiku)
95  if (name.includes('sonnet')) return isOn(env.disableSonnet)
96  if (name.includes('opus')) return isOn(env.disableOpus)
97  return false
98}
99
100export const promptTokens = (s: Sample) => s.read + s.write + s.fresh
101
102/** Share of the prompt the cache served, 0 to 1; 0 for an empty prompt. */
103export function hitRatio(s: Sample): number {
104  const total = promptTokens(s)
105  return total === 0 ? 0 : s.read / total
106}
107
108/** When the cache entry the sample touched lapses, ms since the epoch. */
109export const expiresAt = (s: Sample, ttl: Ttl) => s.startedAt + ttlMs(ttl)
110
111/** Zero for a request that read and wrote nothing: it created or refreshed no entry, so there is nothing to count down. */
112export function remainingMs(s: Sample, ttl: Ttl, now: number): number {
113  if (s.read + s.write === 0) return 0
114  return Math.max(0, expiresAt(s, ttl) - now)
115}
116
117/**
118 * Why a request that should have read the cache wrote it instead; undefined
119 * when it did not miss. A prompt that shrank is a /compact or /clear, not a
120 * miss, and the first request of a session has nothing to read.
121 */
122export function missReason(prev: Sample | undefined, cur: Sample, ttl: Ttl): string | undefined {
123  if (!prev) return undefined
124  const before = promptTokens(prev)
125  if (before === 0 || promptTokens(cur) < before * 0.7) return undefined
126  if (cur.read >= before * 0.5 || cur.write === 0) return undefined
127  if (cur.model !== prev.model) return `model changed (${prev.model} to ${cur.model})`
128  if (cur.startedAt - prev.startedAt > ttlMs(ttl)) return `the ${ttl} cache had lapsed`
129  return 'the prompt prefix changed (effort, tools, system prompt or CLAUDE.md)'
130}
131
132export function advise(last: Sample | undefined, prev: Sample | undefined, policy: Policy, now: number, disabled: boolean): Advice {
133  if (disabled) return { kind: 'off', text: 'prompt caching is off for this model (DISABLE_PROMPT_CACHING*)' }
134  if (!last) return { kind: 'cold', text: 'no request yet: the first one writes the cache' }
135  if (last.read + last.write === 0) {
136    return { kind: 'uncached', text: 'this request was not cached (prompt under the model minimum, or caching off)' }
137  }
138  const miss = missReason(prev, last, policy.ttl)
139  const left = remainingMs(last, policy.ttl, now)
140  const size = promptTokens(last)
141  if (left <= 0) {
142    const big = size >= policy.compactAtTokens
143    return {
144      kind: 'expired',
145      text: big
146        ? `expired: the next message rewrites ${fmtTokens(size)} tokens. /compact first, or /clear if the task is done`
147        : `expired: only ${fmtTokens(size)} tokens to rebuild, just keep going`,
148    }
149  }
150  if (left <= policy.warnMs) {
151    return { kind: 'soon', text: 'expires soon: any message refreshes it for free' }
152  }
153  if (miss) return { kind: 'miss', text: `cache missed: ${miss}` }
154  return { kind: 'warm', text: 'warm: keep going' }
155}
156
157export function fmtTokens(n: number): string {
158  if (n < 1000) return String(n)
159  if (n < 100_000) return `${(n / 1000).toFixed(1).replace(/\.0$/, '')}k`
160  if (n < 1_000_000) return `${Math.round(n / 1000)}k`
161  return `${(n / 1_000_000).toFixed(1).replace(/\.0$/, '')}M`
162}
163
164/** m:ss, or h:mm:ss from an hour up. */
165export function fmtClock(ms: number): string {
166  const total = Math.max(0, Math.ceil(ms / 1000))
167  const h = Math.floor(total / 3600)
168  const m = Math.floor((total % 3600) / 60)
169  const s = total % 60
170  const pad = (n: number) => String(n).padStart(2, '0')
171  return h > 0 ? `${h}:${pad(m)}:${pad(s)}` : `${m}:${pad(s)}`
172}
173
174export function bar(ratio: number, width: number): string {
175  const filled = Math.round(Math.min(1, Math.max(0, ratio)) * width)
176  return '█'.repeat(filled) + '░'.repeat(width - filled)
177}
178
179export type TurnRow = {
180  turnId: string
181  steps: number
182  read: number
183  write: number
184  fresh: number
185  output: number
186}
187
188/** Samples grouped by turn, oldest first, each turn's requests summed. */
189export function byTurn(samples: readonly Sample[]): TurnRow[] {
190  const rows: TurnRow[] = []
191  for (const s of samples) {
192    let row = rows[rows.length - 1]
193    if (!row || row.turnId !== s.turnId) {
194      row = { turnId: s.turnId, steps: 0, read: 0, write: 0, fresh: 0, output: 0 }
195      rows.push(row)
196    }
197    row.steps += 1
198    row.read += s.read
199    row.write += s.write
200    row.fresh += s.fresh
201    row.output += s.output
202  }
203  return rows
204}
205
206export const rowRatio = (r: TurnRow) => {
207  const total = r.read + r.write + r.fresh
208  return total === 0 ? 0 : r.read / total
209}
210
211export function fit(text: string, width: number): string {
212  return text.length <= width ? text : `${text.slice(0, Math.max(0, width - 1))}…`
213}
214
215export function positive(v: unknown, fallback: number): number {
216  return typeof v === 'number' && Number.isFinite(v) && v > 0 ? v : fallback
217}
218
219/** Share of the cache lifetime left, 0 to 1. */
220export function lifeRatio(leftMs: number, ttl: Ttl): number {
221  return Math.min(1, Math.max(0, leftMs / ttlMs(ttl)))
222}
223
224/**
225 * Widths of the three stacked-bar segments (read, wrote, new) over `width`
226 * cells: proportional, each non-empty part at least one cell, summing to width.
227 */
228export function segments(read: number, write: number, fresh: number, width: number): [number, number, number] {
229  const total = read + write + fresh
230  if (total === 0 || width <= 0) return [0, 0, 0]
231  const parts = [read, write, fresh]
232  const cells = parts.map(p => (p > 0 ? Math.max(1, Math.round((p / total) * width)) : 0))
233  let over = cells.reduce((a, b) => a + b, 0) - width
234  while (over !== 0) {
235    const i = over > 0 ? cells.indexOf(Math.max(...cells)) : parts.indexOf(Math.max(...parts))
236    cells[i]! += over > 0 ? -1 : 1
237    over += over > 0 ? -1 : 1
238  }
239  return [cells[0]!, cells[1]!, cells[2]!]
240}
241
242/** Seconds left at which a toast counts down after the one at the warning threshold. */
243export const COUNTDOWN_MARKS = [10, 3, 2, 1]
244
245/**
246 * The toast mark to fire now, or undefined. `level` is the mark last fired for
247 * this cache entry (Infinity before any); a late tick skips straight to the
248 * newest mark crossed, so a stalled clock never replays old ones.
249 */
250export function nextToastMark(secsLeft: number, warnSecs: number, level: number): number | undefined {
251  const marks = [warnSecs, ...COUNTDOWN_MARKS].filter(m => m <= warnSecs)
252  const due = marks.filter(m => secsLeft <= m && m < level)
253  return due.length ? Math.min(...due) : undefined
254}
255
256export type LifeColor = 'green' | 'yellow' | 'red'
257
258/** Countdown colour: green while there is plenty, yellow below 40% of the lifetime, red from the warning threshold down. */
259export function lifeColor(leftMs: number, ttl: Ttl, warnMs: number): LifeColor {
260  if (leftMs <= warnMs) return 'red'
261  return leftMs / ttlMs(ttl) <= 0.4 ? 'yellow' : 'green'
262}
263
264// requests are timed from their start, so a little slack keeps a hit that
265// landed just inside the lifetime from reading as proof of the longer one
266const SLACK_MS = 10_000
267
268/**
269 * What the traffic says about the cache lifetime, given the request before and
270 * `known`, what earlier requests already showed.
271 *
272 *   - a hit (the cache served at least half of the previous prompt) more than
273 *     5 minutes after the previous request began proves the 1-hour lifetime,
274 *     and nothing later undoes it: a miss afterwards is more likely a changed
275 *     prefix than a lapse
276 *   - a miss with the same model and a prompt that did not shrink, 5 minutes to
277 *     an hour after the previous request, says the entry lapsed: 5 minutes
278 *     (weaker: a changed prefix looks the same, so a later hit overrules it)
279 *
280 * Needed because the API names the TTL of a write (`cache_creation.ephemeral_*`)
281 * but Claude Code's mod API passes on only the four token counts.
282 */
283export function observeTtl(prev: Sample | undefined, cur: Sample, known: Ttl | undefined): Ttl | undefined {
284  if (!prev || prev.read + prev.write === 0 || cur.model !== prev.model) return known
285  const gap = cur.startedAt - prev.startedAt
286  const before = promptTokens(prev)
287  if (gap <= ttlMs('5m') + SLACK_MS) return known
288  if (cur.read >= before * 0.5) return '1h'
289  if (known === '1h') return known
290  const lapsed = cur.write > 0 && promptTokens(cur) >= before * 0.7 && gap < ttlMs('1h') + SLACK_MS
291  return lapsed ? '5m' : known
292}
293
types/index.d.ts 60 lines
1export type Ttl = '5m' | '1h'
2
3export type CacheEnv = {
4  enable1h?: string
5  force5m?: string
6  /** CLAUDE_CODE_PROMPT_CACHE_TTL: "5m" or "1h" for the main conversation */
7  ttlVar?: string
8  disableAll?: string
9  disableHaiku?: string
10  disableSonnet?: string
11  disableOpus?: string
12}
13
14/** One main-loop request, as the API reported it. */
15export type Sample = {
16  turnId: string
17  index: number
18  model: string
19  /** ms since the epoch when the request started: the cache's lifetime is counted from here */
20  startedAt: number
21  read: number
22  write: number
23  fresh: number
24  output: number
25}
26
27/** What the account is billed as, as far as the mod can tell. */
28export type Account = 'subscription' | 'credits' | 'other'
29
30/** Everything the band and the pane draw from, held by the host so a reload keeps it. */
31export type Meter = {
32  samples: Sample[]
33  /** The lifetime in use: the observed one when traffic proved it, else `baseTtl`. */
34  ttl: Ttl
35  /** The lifetime Claude Code's rules give (option, environment, setting, account). */
36  baseTtl: Ttl
37  /** Set by the `ttl` option; nothing overrides it. */
38  pinned: boolean
39  /** The lifetime request timing proved, or null. */
40  observed: Ttl | null
41  account: Account
42  ttlSource: string
43  envSource: string
44  env: CacheEnv
45  /** The promptCacheTtl setting, as read at session start. */
46  setting: string | null
47}
48
49declare module 'claude-code' {
50  interface PluginState {
51    'prompt-cache-control': {
52      meter: Meter
53      // The band steps aside while /cache is open.
54      paneOpen: boolean
55      // The bar above the prompt: /cache on, /cache off (kept in $.store).
56      barOn: boolean
57    }
58  }
59}
60