SLOPSHOPPER

cache-line

One-line prompt-cache meter above the prompt (hit rate, countdown, prompt size) that stacks with other mods instead of hiding them. /cache opens the per-turn…

newpanebandcommandtoaststatus
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · cache-line
│ ┃ cache ✕ › fix the failing auth test and add an audit log call │ ┃ ⚡ PROMPT CACHE · 1h lifetime (Claude subscri │ ┃ ● cache-line: cache-line loaded: 1h cache (Claude subscription defaul │ ┃ ⏱ --:-- ⏺ Read(src/auth.ts) │ ┃ ⎿ Read 6 lines │ ┃ ○ no request yet: the first one writes the ⏺ Update(src/auth.ts) │ ┃ cache ⎿ Added 2 lines, removed 1 line │ ┃ ⏺ Bash(bun test) │ ┃ turn steps read wrote new hit ⎿ 3 pass, 1 fail │ ┃ no requests yet │ ┃ ● Done. refresh now rejects expired claims and logs an audit event. │ ┃ [ close ] │ ┃ ✻ Worked for 42s · done 4:20 PM │ ┃ ■ read: served by the cache │ ┃ ■ wrote: new cache entry › /cache │ ┃ ■ new: sent uncached ⎿ cache-line: 1h cache (Claude subscription default) · no request │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · cache
⚡ PROMPT CACHE · 1h lifetime (Claude subscription default) ⏱ --:-- ○ no request yet: the first one writes the cache turn steps read wrote new hit no requests yet [ close ] ■ read: served by the cache ■ wrote: new cache entry ■ new: sent uncached
README

prompt-cache-control (Claude Cache Control)

A prompt-cache meter above the Claude Code prompt. Every request Claude makes reports how much of its prompt the cache served, how much it wrote and how much went uncached; this mod keeps those numbers per request and per turn, counts down to the moment the cache lapses and tells you what to do about it: keep going, /compact or /clear.

cache ██████████ 98% read 80k · wrote 1k · new 300 ⏱ 3:41 5m · warm: keep going
cache ░░░░░░░░░░  0% read 0 · wrote 52k · new 300 ⏱ 4:58 5m · cache missed: model changed (…)
cache ██████████ 98% read 150k · wrote 1k · new 300 ⏱ 0:00 5m · expired: the next message rewrites 151k tokens. /compact first, or /clear if the task is done

/cache opens a pane: the time left with a solid bar that shrinks as the cache runs out (green, yellow below 40% of the lifetime, red from the warning threshold), a stacked read / wrote / new bar for the last request, and a colour-coded table with one row per turn. Bars are filled cells and columns have fixed widths with no-break spaces, so the terminal and the Desktop (HTML) pane render the same.

How the countdown works

From Anthropic's prompt caching documentation:

  • The cache lives 5 minutes by default, 1 hour when asked for.
  • Every request that reads the cache refreshes it at no extra cost, so a conversation that keeps talking keeps the 5-minute cache warm.
  • The lifetime is counted from the start of the request that wrote or read the entry; generation time counts against it.
  • A prompt is input_tokens (uncached remainder) + cache_read_input_tokens + cache_creation_input_tokens.
  • Writes cost 1.25x base input for 5 minutes and 2x for 1 hour; reads cost about 0.1x (less on some models). The expensive moment is an expired cache on a large context, which is when this mod suggests /compact.
  • /clear starts a new conversation in the same process, so the meter and the /cache table start over with it. A change in the prefix (model, effort or thinking settings, tool set, system prompt, CLAUDE.md) makes the next request write instead of read. The mod names the cause when it sees a miss: model changed, the cache had lapsed, or the prefix changed.

Which lifetime your account gets

The mod follows Claude Code's own rules (prompt caching: cache lifetime, Claude Code 2.1.242 or later). For the main conversation the TTL is the first match of:

#SourceResult
1the mod's ttl option (5m / 1h)what you set
2FORCE_PROMPT_CACHING_5M=15 minutes
3CLAUDE_CODE_PROMPT_CACHE_TTL5m or 1h
4the promptCacheTtl setting (local, project or user settings file)5m or 1h
5ENABLE_PROMPT_CACHING_1H=11 hour
6the account1 hour on a Claude subscription within its plan usage; 5 minutes on usage credits, an API key or a cloud provider

The account comes from the rate-limit windows the last response reported: a five_hour or seven_day window means a subscription, and one at 100% means requests now draw on usage credits. An API key or a cloud provider reports no such window, and before the first response nothing is known, so the mod starts from 5 minutes there. Managed settings are not readable from a mod.

On top of that the mod watches the traffic, which beats rows 2 to 6: a request that hits the cache more than 5 minutes after the previous one proves the 1-hour lifetime (a later miss does not undo it, since a changed prefix looks the same), and a miss 5 to 60 minutes after the previous request, with the same model and a prompt that did not shrink, says the entry lapsed, so 5 minutes (a later hit overrules it). That covers what the mod cannot see: managed settings, a gateway that rewrites the TTL, or a subscription that ran out of plan usage mid-session. The pane header names the source in use.

Why the mod infers instead of reading it: the API names the TTL of each write (cache_creation.ephemeral_5m_input_tokens / ephemeral_1h_input_tokens) and Claude Code's status line exposes it as prompt_cache.ttl, but the mod API passes on only the four token counts. To check by hand, claude -p "hello" --output-format json and read usage.cache_creation.

Other switches read from the environment at session start:

VariableEffect on the meter
DISABLE_PROMPT_CACHING=1 (and _HAIKU, _SONNET, _OPUS)the band says caching is off for that model

What it hooks

  • turn.step: reads each main-loop request's usage (subagents have their own prefixes and are left out)
  • $.clock.every(1000): redraws the countdown, and only while its text changes, so an idle expired session costs nothing
  • ui.render on AbovePrompt (the band) and on Pane (/cache)
  • $.ui.toast: once per cache entry at the warning threshold (60 s by default) and again at 10, 3, 2 and 1 seconds left, for prompts of 20k tokens or more

Options

  ttl: string               "auto" | "5m" | "1h" (default auto)
  warnSeconds: number       countdown threshold for the yellow state and the toast (default 60)
  compactAtTokens: number   prompt size that makes an expired cache suggest /compact (default 100000)
  band: boolean             row above the prompt (default true)
  status: boolean           entry under the prompt, "cache 98% · 3:41" (default false)
  toast: boolean            toasts at the threshold, 10, 3, 2 and 1 s (default true)

The 100k compactAtTokens is a judgement, not a figure from the documentation: lower it if your model's cache writes are expensive for you.

Install

npx claude-code-templates@latest --mod observability/prompt-cache-control
claude

It is written to .claude/skills/prompt-cache-control/, which Claude Code auto-loads as prompt-cache-control@skills-dir once the workspace trust prompt is accepted. For one session with hot reload: claude --plugin-dir .claude/skills/prompt-cache-control. claude plugin validate .claude/skills/prompt-cache-control prints every event it hooks and every $ call it makes; claude plugin test .claude/skills/prompt-cache-control runs its tests.

Options are read from user settings (~/.claude/settings.json, never project settings), --settings <file> or managed settings, under the plugin's full id:

{ "pluginConfigs": { "prompt-cache-control@skills-dir": { "options": { } } } }

Requirements. Mods are on by default in Claude Code 2.1.287+. Typed against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods

Source 3 files
hooks/register.tsx 424 lines
1/**
2 * cache-line — one-line prompt-cache meter (fork of prompt-cache-control)
3 *
4 * A prompt-cache meter for Claude Code. Every main-loop request reports how
5 * many prompt tokens the cache served (`cache_read_input_tokens`), wrote
6 * (`cache_creation_input_tokens`) and sent uncached (`input_tokens`); this mod
7 * keeps those per request and per turn, counts down to the moment the cache
8 * lapses, and says what to do about it: keep going, /compact or /clear.
9 *
10 *   - `turn.step` reads each main-loop request's usage (subagents have their
11 *     own prefixes and are left out)
12 *   - `$.clock.every(1000)` redraws the countdown, and only while its text
13 *     changes: an idle, expired session costs nothing
14 *   - a row above the prompt (the AbovePrompt component), an optional status
15 *     line entry, and `/cache`, a pane with one row per turn
16 *
17 * The lifetime is counted from the start of the request that last wrote or read
18 * the cache, as Anthropic documents it. Which lifetime Claude Code asked for
19 * follows its documented rules (see decideTtl in ./cache.ts): FORCE_PROMPT_CACHING_5M,
20 * CLAUDE_CODE_PROMPT_CACHE_TTL, the promptCacheTtl setting, ENABLE_PROMPT_CACHING_1H,
21 * then the account (1 hour on a Claude subscription, 5 minutes otherwise). The
22 * API names the TTL of a write but the mod API passes on only the token counts,
23 * so the mod also watches the gaps between requests (a hit after more than 5
24 * minutes proves 1 hour; see observeTtl). `ttl: "5m" | "1h"` pins it.
25 *
26 * Needs Claude Code >= 2.1.287.
27 *
28 * Options (pluginConfigs["cache-line"].options):
29 *   ttl: "auto" | "5m" | "1h"   cache lifetime (default auto)
30 *   warnSeconds: number         countdown threshold for the warning (default 60)
31 *   compactAtTokens: number     prompt size that makes an expired cache suggest /compact (default 100000)
32 *   band: boolean               row above the prompt (default true)
33 *   status: boolean             entry under the prompt (default false)
34 *   toast: boolean              toasts near expiry: at warnSeconds, then 10, 3, 2 and 1 s (default true)
35 */
36import { atom, read, update } from 'claude-code'
37import type { EngineInterface, Register } from 'claude-code'
38import {
39  advise,
40  COUNTDOWN_MARKS,
41  bar,
42  byTurn,
43  fit,
44  fmtClock,
45  fmtTokens,
46  hitRatio,
47  isCachingDisabled,
48  accountOf,
49  decideTtl,
50  observeTtl,
51  lifeColor,
52  lifeRatio,
53  nextToastMark,
54  positive,
55  promptTokens,
56  remainingMs,
57  rowRatio,
58  segments,
59} from './cache.ts'
60import type { Account, Advice, CacheEnv, Sample, Ttl } from './cache.ts'
61
62// the session's requests live in $.state: a reload (each edit of this mod) keeps the last one instead of starting blank
63const saved = atom({ plugin: 'cache-line', key: 'samples' } as const, [])
64
65const PANE = 'cache'
66const COMMAND = 'cache'
67const KEEP = 200
68// below this a lapsed cache costs too little to interrupt anyone about
69const TOAST_MIN_TOKENS = 20_000
70
71let samples: Sample[] = []
72let ttl: Ttl = '5m'
73let baseTtl: Ttl = '5m'
74let pinned = false
75let observed: Ttl | undefined
76let setting: unknown
77let account: Account = 'other'
78let ttlSource = 'default'
79let envSource = 'default'
80let env: CacheEnv = {}
81let timer: { cancel: () => void } | undefined
82let lastKey = ''
83let toastedFor = 0
84let toastLevel = Infinity
85let isPaneOpen = false
86
87type Policy = { warnMs: number; compactAtTokens: number }
88
89function current(policy: Policy, now: number) {
90  const last = samples[samples.length - 1]
91  const prev = samples[samples.length - 2]
92  const disabled = last ? isCachingDisabled(last.model, env) : isCachingDisabled('', env)
93  const advice: Advice = advise(last, prev, { ttl, ...policy }, now, disabled)
94  const left = last ? remainingMs(last, ttl, now) : 0
95  return { last, advice, left }
96}
97
98const COLOR: Record<Advice['kind'], string | undefined> = {
99  warm: 'green',
100  soon: 'yellow',
101  expired: 'red',
102  miss: 'red',
103  off: undefined,
104  cold: undefined,
105  uncached: undefined,
106}
107
108function shortLine(policy: Policy, now: number): string {
109  const { last, advice, left } = current(policy, now)
110  if (!last || advice.kind === 'off') return `cache: ${advice.text}`
111  const clock = left > 0 ? ` · ${fmtClock(left)}` : ''
112  return `cache ${Math.round(hitRatio(last) * 100)}%${clock}`
113}
114
115// the promptCacheTtl setting, from the settings files that can carry it (local over project over user)
116async function readSetting($: EngineInterface): Promise<unknown> {
117  const home = await $.env.get('HOME').catch(() => undefined)
118  const cwd = await $.session.cwd().catch(() => undefined)
119  const files = [cwd && `${cwd}/.claude/settings.local.json`, cwd && `${cwd}/.claude/settings.json`, home && `${home}/.claude/settings.json`]
120  for (const file of files) {
121    if (!file) continue
122    try {
123      const value = JSON.parse(await $.fs.read(file)).promptCacheTtl
124      if (value === '5m' || value === '1h') return value
125    } catch {
126      // missing or unreadable: the next file
127    }
128  }
129  return undefined
130}
131
132export const register: Register = (on, options) => {
133  const policy: Policy = {
134    warnMs: positive(options.warnSeconds, 60) * 1000,
135    compactAtTokens: positive(options.compactAtTokens, 100_000),
136  }
137  const showBand = options.band !== false
138  const showStatus = options.status === true
139  const wantToast = options.toast !== false
140
141  on('session.start', async ($, e, next) => {
142    const r = await next(e)
143    samples = [...(await read($, saved).catch(() => []))]
144    lastKey = ''
145    toastedFor = 0
146    const none = () => undefined
147    env = {
148      enable1h: await $.env.get('ENABLE_PROMPT_CACHING_1H').catch(none),
149      force5m: await $.env.get('FORCE_PROMPT_CACHING_5M').catch(none),
150      ttlVar: await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL').catch(none),
151      disableAll: await $.env.get('DISABLE_PROMPT_CACHING').catch(none),
152      disableHaiku: await $.env.get('DISABLE_PROMPT_CACHING_HAIKU').catch(none),
153      disableSonnet: await $.env.get('DISABLE_PROMPT_CACHING_SONNET').catch(none),
154      disableOpus: await $.env.get('DISABLE_PROMPT_CACHING_OPUS').catch(none),
155    }
156    pinned = options.ttl === '5m' || options.ttl === '1h'
157    observed = undefined
158    setting = await readSetting($)
159    account = accountOf((await $.session.usage().catch(() => undefined))?.rateLimits ?? [])
160    const choice = decideTtl(options.ttl, env, setting, account)
161    baseTtl = choice.ttl
162    ttl = baseTtl
163    envSource = choice.source
164    ttlSource = envSource
165
166    await $.command
167      .register({
168        name: COMMAND,
169        description: 'Prompt-cache usage per turn and the time left before it lapses (stop closes)',
170        argumentHint: '[stop]',
171        immediate: true,
172      })
173      .catch(err => $.ui.log(`cache-line: /${COMMAND} not registered: ${err}`))
174    $.ui.log(`cache-line loaded: ${ttl} cache (${ttlSource}), /${COMMAND} opens the table`, { to: 'debug' })
175
176    timer?.cancel()
177    timer = $.clock.every(1000, () => {
178      const now = Date.now()
179      const { last, advice, left } = current(policy, now)
180      const key = `${advice.kind}|${advice.text}|${left > 0 ? fmtClock(left) : ''}`
181      if (key !== lastKey) {
182        lastKey = key
183        if (showStatus) $.ui.status(shortLine(policy, now))
184        $.ui.invalidate('ui.render')
185      }
186      if (wantToast && last && left > 0 && promptTokens(last) >= TOAST_MIN_TOKENS) {
187        if (toastedFor !== last.startedAt) {
188          toastedFor = last.startedAt
189          toastLevel = Infinity
190        }
191        // the first toast comes at warnSeconds, then 10, 3, 2 and 1 seconds; a late tick skips to the newest one
192        const secs = Math.ceil(left / 1000)
193        const mark = nextToastMark(secs, policy.warnMs / 1000, toastLevel)
194        if (mark !== undefined) {
195          toastLevel = mark
196          const tail = secs <= COUNTDOWN_MARKS[0] ? 'send a message now' : `send a message to keep ${fmtTokens(promptTokens(last))} tokens warm`
197          $.ui.toast(`cache expires in ${secs >= 60 ? fmtClock(left) : `${secs}s`}: ${tail}`)
198        }
199      }
200    })
201    return r
202  })
203
204  on('session.end', async ($, e, next) => {
205    // /clear starts a new conversation in the same process: its cache is a new one
206    if (e.reason === 'clear') {
207      samples = []
208      await update($, saved, () => []).catch(() => undefined)
209      lastKey = ''
210      toastedFor = 0
211      observed = undefined
212      ttl = baseTtl
213      ttlSource = envSource
214      $.ui.invalidate('ui.render')
215      return next(e)
216    }
217    timer?.cancel()
218    timer = undefined
219    return next(e)
220  })
221
222  // each main-loop request: what the cache did with it
223  on('turn.step', async function* ($, e, next) {
224    if (e.agentId) return yield* next(e)
225    const startedAt = Date.now()
226    const r = yield* next(e)
227    if (r.usage) {
228      samples.push({
229        turnId: e.turnId,
230        index: e.index,
231        model: r.usage.model || e.model,
232        startedAt,
233        read: r.usage.cache_read_input_tokens,
234        write: r.usage.cache_creation_input_tokens,
235        fresh: r.usage.input_tokens,
236        output: r.usage.output_tokens,
237      })
238      if (samples.length > KEEP) samples = samples.slice(-KEEP)
239      await update($, saved, () => samples).catch(() => undefined)
240      if (!pinned) {
241        // the account can change under a session: a subscription running out of plan usage moves to usage credits
242        account = accountOf((await $.session.usage().catch(() => undefined))?.rateLimits ?? [])
243        const choice = decideTtl(options.ttl, env, setting, account)
244        baseTtl = choice.ttl
245        envSource = choice.source
246        if (observed === undefined) {
247          ttl = baseTtl
248          ttlSource = envSource
249        }
250        const seen = observeTtl(samples[samples.length - 2], samples[samples.length - 1], observed)
251        if (seen !== observed) {
252          observed = seen
253          ttl = seen ?? baseTtl
254          ttlSource = `observed from request timing; ${envSource} said ${baseTtl}`
255          $.ui.log(`cache-line: cache lifetime is ${ttl} (${ttlSource})`, { to: 'debug' })
256        }
257      }
258      lastKey = ''
259      if (showStatus) $.ui.status(shortLine(policy, Date.now()))
260      $.ui.invalidate('ui.render')
261    }
262    return r
263  })
264
265  on('command.run', { command: COMMAND }, async ($, e) => {
266    if (e.args.trim().toLowerCase() === 'stop') {
267      await $.ui.close({ id: PANE }).catch(() => undefined)
268      isPaneOpen = false
269      return { text: 'cache table closed' }
270    }
271    isPaneOpen = true
272    await $.ui.open({ id: PANE, title: 'cache', focus: true })
273    $.ui.invalidate('ui.render')
274    const { advice } = current(policy, Date.now())
275    return { text: `${ttl} cache (${ttlSource}) · ${advice.text} · /${COMMAND} stop closes` }
276  })
277
278  on('ui.close', async ($, e, next) => {
279    if (e.id !== PANE) return next(e)
280    isPaneOpen = false
281    return next(e)
282  })
283
284  on('ui.press', async ($, e, next) => {
285    if (e.plugin !== $.plugin.name || e.requestId !== PANE) return next(e)
286    if (e.element === 'close') await $.ui.close({ id: PANE }).catch(() => undefined)
287    return next(e)
288  })
289
290  // one line, stacked under whatever the plugins beneath draw: never hides another mod's band
291  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
292    const below = await next(e)
293    if (!showBand || e.props.hasSurvey || isPaneOpen) return below
294    if (samples.length === 0) samples = [...(await read($, saved).catch(() => []))]
295    const { last, advice, left } = current(policy, Date.now())
296    const { Box, Text } = $.ui.resolve(e)
297    // a brand-new session before its first request: a bare dim marker
298    if (!last) {
299      const idle = <Text key="cache-line" dimColor>{advice.kind === 'off' ? '○ cache off' : '○ cache'}</Text>
300      return below ? <Box flexDirection="column">{below}{idle}</Box> : idle
301    }
302    const pct = Math.round(hitRatio(last) * 100)
303    const counting = advice.kind !== 'uncached' && advice.kind !== 'off'
304    const icon = advice.kind === 'warm' ? '●' : advice.kind === 'soon' ? '▲' : advice.kind === 'expired' || advice.kind === 'miss' ? '✖' : '○'
305    const line = (
306      <Box key="cache-line" flexDirection="row" columnGap={1}>
307        <Text color={COLOR[advice.kind]}>{icon}</Text>
308        <Text dimColor>cache</Text>
309        <Text bold>{`${pct}%`}</Text>
310        {counting ? <Text color={left > 0 ? lifeColor(left, ttl, policy.warnMs) : 'red'}>{`⏱ ${left > 0 ? fmtClock(left) : '0:00'}`}</Text> : null}
311        <Text dimColor wrap="truncate-end">{`${fmtTokens(promptTokens(last))} tok`}</Text>
312      </Box>
313    )
314    return below ? <Box flexDirection="column">{below}{line}</Box> : line
315  })
316
317  on('ui.render', { component: 'Pane' }, async ($, e, next) => {
318    if (e.requestId !== PANE) return next(e)
319    const { Box, Text, Button } = $.ui.resolve(e)
320    const width = Math.max(30, e.props.bodyColumns - 1)
321    // HTML collapses runs of spaces and trims a text's ends; a no-break space keeps them
322    const sp = (t: string) => (e.surface === 'terminal' ? t : t.replace(/ /g, ' '))
323    const now = Date.now()
324    const { last, advice, left } = current(policy, now)
325    const all = byTurn(samples)
326    const counting = !!last && advice.kind !== 'uncached' && advice.kind !== 'off'
327    // the countdown goes green, then yellow, then red as the cache runs out
328    const clockColor = counting ? lifeColor(left, ttl, policy.warnMs) : undefined
329    const stateColor = advice.kind === 'expired' || advice.kind === 'miss' ? 'red' : (clockColor ?? COLOR[advice.kind])
330    const hitColor = (pct: number) => (pct >= 80 ? 'green' : pct >= 40 ? 'yellow' : 'red')
331    // solid bars are filled Boxes, not block characters, so HTML draws no seams between cells
332    const solid = (key: string, parts: [number, string | undefined][]) => (
333      <Box key={key} flexDirection="row" height={1} flexShrink={0}>
334        {parts.map(([w, c], i) => (w > 0 ? <Box key={`${key}:${i}`} width={w} height={1} flexShrink={0} backgroundColor={c} /> : null))}
335      </Box>
336    )
337    const cell = (key: string, w: number, text: string, c?: string, bold = false) => (
338      <Box key={key} width={w} flexShrink={0} justifyContent="flex-end">
339        <Text color={c} bold={bold} dimColor={!c}>{sp(text)}</Text>
340      </Box>
341    )
342
343    const barW = Math.min(width, 40)
344    const life = lifeRatio(left, ttl)
345    const lifeFilled = Math.round(life * barW)
346    const [sr, sw, sn] = last ? segments(last.read, last.write, last.fresh, barW) : [0, 0, 0]
347    const rows = all.slice(-Math.max(3, (e.viewport?.rows ?? 24) - 16))
348    const icon = advice.kind === 'warm' ? '●' : advice.kind === 'soon' ? '▲' : advice.kind === 'expired' || advice.kind === 'miss' ? '✖' : '○'
349
350    return (
351      <Box flexDirection="column">
352        <Box key="title" flexDirection="row" columnGap={1}>
353          <Text bold color="cyan">{sp('⚡ PROMPT CACHE')}</Text>
354          <Text dimColor>{sp(`· ${ttl} lifetime (${ttlSource})`)}</Text>
355        </Box>
356
357        <Box key="clock" flexDirection="column" marginTop={1}>
358          <Text bold color={clockColor}>{sp(counting ? `⏱ ${left > 0 ? fmtClock(left) : '0:00'}` : '⏱ --:--')}</Text>
359          {counting ? (
360            <Box flexDirection="row" columnGap={1}>
361              {solid('life', [[lifeFilled, clockColor], [barW - lifeFilled, 'gray']])}
362              <Text dimColor>{sp(`${Math.round(life * 100)}%`)}</Text>
363            </Box>
364          ) : null}
365        </Box>
366
367        <Box key="advice" marginTop={1} flexDirection="column">
368          <Text bold color={stateColor}>{sp(`${icon} ${advice.text}`)}</Text>
369          {last ? <Text dimColor>{sp(fit(`${last.model} · prompt ${fmtTokens(promptTokens(last))} tokens`, width))}</Text> : null}
370        </Box>
371
372        {last ? (
373          <Box key="stack" flexDirection="column" marginTop={1}>
374            <Box flexDirection="row" columnGap={1}>
375              {solid('stack', [[sr, 'green'], [sw, 'yellow'], [sn, 'cyan']])}
376              <Text bold color={hitColor(Math.round(hitRatio(last) * 100))}>{sp(`${Math.round(hitRatio(last) * 100)}% hit`)}</Text>
377            </Box>
378            <Box flexDirection="row" columnGap={2}>
379              <Text color="green">{sp(`■ read ${fmtTokens(last.read)}`)}</Text>
380              <Text color="yellow">{sp(`■ wrote ${fmtTokens(last.write)}`)}</Text>
381              <Text color="cyan">{sp(`■ new ${fmtTokens(last.fresh)}`)}</Text>
382            </Box>
383          </Box>
384        ) : null}
385
386        <Box key="table" flexDirection="column" marginTop={1}>
387          <Box key="head" flexDirection="row" columnGap={1}>
388            {cell('h:turn', 4, 'turn', 'cyan', true)}
389            {cell('h:steps', 5, 'steps', 'cyan', true)}
390            {cell('h:read', 6, 'read', 'green', true)}
391            {cell('h:wrote', 6, 'wrote', 'yellow', true)}
392            {cell('h:new', 5, 'new', 'cyan', true)}
393            {cell('h:hit', 4, 'hit', 'magenta', true)}
394          </Box>
395          {rows.length === 0 ? <Text dimColor>{sp('no requests yet')}</Text> : null}
396          {rows.map((row, i) => {
397            const n = all.length - rows.length + i + 1
398            const pct = Math.round(rowRatio(row) * 100)
399            return (
400              <Box key={`t:${row.turnId}`} flexDirection="row" columnGap={1}>
401                {cell(`c:turn:${row.turnId}`, 4, String(n))}
402                {cell(`c:steps:${row.turnId}`, 5, String(row.steps))}
403                {cell(`c:read:${row.turnId}`, 6, fmtTokens(row.read), 'green')}
404                {cell(`c:wrote:${row.turnId}`, 6, fmtTokens(row.write), 'yellow')}
405                {cell(`c:new:${row.turnId}`, 5, fmtTokens(row.fresh), 'cyan')}
406                {cell(`c:hit:${row.turnId}`, 4, `${pct}%`, hitColor(pct), true)}
407              </Box>
408            )
409          })}
410        </Box>
411
412        <Box key="foot" marginTop={1} flexDirection="column">
413          <Button key="close" label="close" onPress={() => {}} />
414          <Box key="legend" marginTop={1} flexDirection="column">
415            <Text color="green">{sp('■ read: served by the cache')}</Text>
416            <Text color="yellow">{sp('■ wrote: new cache entry')}</Text>
417            <Text color="cyan">{sp('■ new: sent uncached')}</Text>
418          </Box>
419        </Box>
420      </Box>
421    )
422  })
423}
424
hooks/cache.ts 318 lines
1/**
2 * cache.ts — the pure half of prompt-cache-control: no `$`, no engine.
3 *
4 * What it models, from Anthropic's prompt-caching documentation:
5 *   - the cache lives 5 minutes by default, 1 hour when asked for; a read
6 *     refreshes the entry at no extra cost, and the lifetime is measured from
7 *     the START of the request that wrote or read it
8 *   - a request's prompt is `input_tokens` (uncached remainder) +
9 *     `cache_read_input_tokens` + `cache_creation_input_tokens`
10 *   - writes cost 1.25x base input for 5m and 2x for 1h; reads about 0.1x
11 *     (less on some models), so an expired cache on a large context is the
12 *     expensive moment
13 *   - a prefix change (model, effort/thinking settings, tool set, system
14 *     prompt) makes the next request write instead of read
15 *
16 * Claude Code's own switches (read from the environment):
17 *   ENABLE_PROMPT_CACHING_1H=1   ask for the 1-hour TTL
18 *   FORCE_PROMPT_CACHING_5M=1    force the 5-minute TTL, beating the above
19 *   DISABLE_PROMPT_CACHING=1     no caching; DISABLE_PROMPT_CACHING_{HAIKU,SONNET,OPUS}
20 *                                 turn it off for that model family only
21 */
22
23export type Ttl = '5m' | '1h'
24
25export type CacheEnv = {
26  enable1h?: string
27  force5m?: string
28  /** CLAUDE_CODE_PROMPT_CACHE_TTL: "5m" or "1h" for the main conversation */
29  ttlVar?: string
30  disableAll?: string
31  disableHaiku?: string
32  disableSonnet?: string
33  disableOpus?: string
34}
35
36/** One main-loop request, as the API reported it. */
37export type Sample = {
38  turnId: string
39  index: number
40  model: string
41  /** ms since the epoch when the request started: the cache's lifetime is counted from here */
42  startedAt: number
43  read: number
44  write: number
45  fresh: number
46  output: number
47}
48
49export type AdviceKind = 'off' | 'cold' | 'uncached' | 'warm' | 'soon' | 'expired' | 'miss'
50
51export type Advice = {
52  kind: AdviceKind
53  /** one sentence for the band */
54  text: string
55}
56
57export type Policy = {
58  ttl: Ttl
59  warnMs: number
60  compactAtTokens: number
61}
62
63export const isOn = (v: string | undefined) => v === '1' || v?.toLowerCase() === 'true'
64
65/** What the account is billed as, as far as the mod can tell. */
66export type Account = 'subscription' | 'credits' | 'other'
67
68export type TtlChoice = { ttl: Ttl; source: string }
69
70const asTtl = (v: unknown): Ttl | undefined => (v === '5m' || v === '1h' ? v : undefined)
71
72/**
73 * Which lifetime Claude Code asks for on the main conversation, in the order
74 * its documentation gives (https://code.claude.com/docs/en/prompt-caching,
75 * "Choose the TTL yourself"), after the mod's own `ttl` option:
76 *
77 *   FORCE_PROMPT_CACHING_5M, CLAUDE_CODE_PROMPT_CACHE_TTL, the promptCacheTtl
78 *   setting, ENABLE_PROMPT_CACHING_1H, then the default of the account: one
79 *   hour on a Claude subscription within its plan usage, five minutes on usage
80 *   credits, an API key or a cloud provider.
81 */
82export function decideTtl(option: unknown, env: CacheEnv, setting?: unknown, account?: Account): TtlChoice {
83  const pinned = asTtl(option)
84  if (pinned) return { ttl: pinned, source: 'the ttl option' }
85  if (isOn(env.force5m)) return { ttl: '5m', source: 'FORCE_PROMPT_CACHING_5M' }
86  const fromVar = asTtl(env.ttlVar)
87  if (fromVar) return { ttl: fromVar, source: 'CLAUDE_CODE_PROMPT_CACHE_TTL' }
88  const fromSetting = asTtl(setting)
89  if (fromSetting) return { ttl: fromSetting, source: 'the promptCacheTtl setting' }
90  if (isOn(env.enable1h)) return { ttl: '1h', source: 'ENABLE_PROMPT_CACHING_1H' }
91  if (account === 'subscription') return { ttl: '1h', source: 'Claude subscription default' }
92  if (account === 'credits') return { ttl: '5m', source: 'usage credits default' }
93  return { ttl: '5m', source: 'default' }
94}
95
96export const resolveTtl = (option: unknown, env: CacheEnv, setting?: unknown, account?: Account): Ttl =>
97  decideTtl(option, env, setting, account).ttl
98
99/**
100 * The account, from the rate-limit windows the last response reported: a
101 * five-hour or seven-day window means a Claude subscription, and one that is
102 * full means the next requests draw on usage credits. No window (an API key,
103 * a cloud provider, or no response yet) says nothing.
104 */
105export function accountOf(windows: readonly { kind: string; percentUsed: number }[]): Account {
106  const plan = windows.filter(w => w.kind === 'five_hour' || w.kind === 'seven_day')
107  if (plan.length === 0) return 'other'
108  return plan.some(w => w.percentUsed >= 100) ? 'credits' : 'subscription'
109}
110
111export function ttlMs(ttl: Ttl): number {
112  return ttl === '1h' ? 3_600_000 : 300_000
113}
114
115/** Caching switched off for this model by the environment. */
116export function isCachingDisabled(model: string, env: CacheEnv): boolean {
117  if (isOn(env.disableAll)) return true
118  const name = model.toLowerCase()
119  if (name.includes('haiku')) return isOn(env.disableHaiku)
120  if (name.includes('sonnet')) return isOn(env.disableSonnet)
121  if (name.includes('opus')) return isOn(env.disableOpus)
122  return false
123}
124
125export const promptTokens = (s: Sample) => s.read + s.write + s.fresh
126
127/** Share of the prompt the cache served, 0 to 1; 0 for an empty prompt. */
128export function hitRatio(s: Sample): number {
129  const total = promptTokens(s)
130  return total === 0 ? 0 : s.read / total
131}
132
133/** When the cache entry the sample touched lapses, ms since the epoch. */
134export const expiresAt = (s: Sample, ttl: Ttl) => s.startedAt + ttlMs(ttl)
135
136/** Zero for a request that read and wrote nothing: it created or refreshed no entry, so there is nothing to count down. */
137export function remainingMs(s: Sample, ttl: Ttl, now: number): number {
138  if (s.read + s.write === 0) return 0
139  return Math.max(0, expiresAt(s, ttl) - now)
140}
141
142/**
143 * Why a request that should have read the cache wrote it instead; undefined
144 * when it did not miss. A prompt that shrank is a /compact or /clear, not a
145 * miss, and the first request of a session has nothing to read.
146 */
147export function missReason(prev: Sample | undefined, cur: Sample, ttl: Ttl): string | undefined {
148  if (!prev) return undefined
149  const before = promptTokens(prev)
150  if (before === 0 || promptTokens(cur) < before * 0.7) return undefined
151  if (cur.read >= before * 0.5 || cur.write === 0) return undefined
152  if (cur.model !== prev.model) return `model changed (${prev.model} to ${cur.model})`
153  if (cur.startedAt - prev.startedAt > ttlMs(ttl)) return `the ${ttl} cache had lapsed`
154  return 'the prompt prefix changed (effort, tools, system prompt or CLAUDE.md)'
155}
156
157export function advise(last: Sample | undefined, prev: Sample | undefined, policy: Policy, now: number, disabled: boolean): Advice {
158  if (disabled) return { kind: 'off', text: 'prompt caching is off for this model (DISABLE_PROMPT_CACHING*)' }
159  if (!last) return { kind: 'cold', text: 'no request yet: the first one writes the cache' }
160  if (last.read + last.write === 0) {
161    return { kind: 'uncached', text: 'this request was not cached (prompt under the model minimum, or caching off)' }
162  }
163  const miss = missReason(prev, last, policy.ttl)
164  const left = remainingMs(last, policy.ttl, now)
165  const size = promptTokens(last)
166  if (left <= 0) {
167    const big = size >= policy.compactAtTokens
168    return {
169      kind: 'expired',
170      text: big
171        ? `expired: the next message rewrites ${fmtTokens(size)} tokens. /compact first, or /clear if the task is done`
172        : `expired: only ${fmtTokens(size)} tokens to rebuild, just keep going`,
173    }
174  }
175  if (left <= policy.warnMs) {
176    return { kind: 'soon', text: 'expires soon: any message refreshes it for free' }
177  }
178  if (miss) return { kind: 'miss', text: `cache missed: ${miss}` }
179  return { kind: 'warm', text: 'warm: keep going' }
180}
181
182export function fmtTokens(n: number): string {
183  if (n < 1000) return String(n)
184  if (n < 100_000) return `${(n / 1000).toFixed(1).replace(/\.0$/, '')}k`
185  if (n < 1_000_000) return `${Math.round(n / 1000)}k`
186  return `${(n / 1_000_000).toFixed(1).replace(/\.0$/, '')}M`
187}
188
189/** m:ss, or h:mm:ss from an hour up. */
190export function fmtClock(ms: number): string {
191  const total = Math.max(0, Math.ceil(ms / 1000))
192  const h = Math.floor(total / 3600)
193  const m = Math.floor((total % 3600) / 60)
194  const s = total % 60
195  const pad = (n: number) => String(n).padStart(2, '0')
196  return h > 0 ? `${h}:${pad(m)}:${pad(s)}` : `${m}:${pad(s)}`
197}
198
199export function bar(ratio: number, width: number): string {
200  const filled = Math.round(Math.min(1, Math.max(0, ratio)) * width)
201  return '█'.repeat(filled) + '░'.repeat(width - filled)
202}
203
204export type TurnRow = {
205  turnId: string
206  steps: number
207  read: number
208  write: number
209  fresh: number
210  output: number
211}
212
213/** Samples grouped by turn, oldest first, each turn's requests summed. */
214export function byTurn(samples: readonly Sample[]): TurnRow[] {
215  const rows: TurnRow[] = []
216  for (const s of samples) {
217    let row = rows[rows.length - 1]
218    if (!row || row.turnId !== s.turnId) {
219      row = { turnId: s.turnId, steps: 0, read: 0, write: 0, fresh: 0, output: 0 }
220      rows.push(row)
221    }
222    row.steps += 1
223    row.read += s.read
224    row.write += s.write
225    row.fresh += s.fresh
226    row.output += s.output
227  }
228  return rows
229}
230
231export const rowRatio = (r: TurnRow) => {
232  const total = r.read + r.write + r.fresh
233  return total === 0 ? 0 : r.read / total
234}
235
236export function fit(text: string, width: number): string {
237  return text.length <= width ? text : `${text.slice(0, Math.max(0, width - 1))}…`
238}
239
240export function positive(v: unknown, fallback: number): number {
241  return typeof v === 'number' && Number.isFinite(v) && v > 0 ? v : fallback
242}
243
244/** Share of the cache lifetime left, 0 to 1. */
245export function lifeRatio(leftMs: number, ttl: Ttl): number {
246  return Math.min(1, Math.max(0, leftMs / ttlMs(ttl)))
247}
248
249/**
250 * Widths of the three stacked-bar segments (read, wrote, new) over `width`
251 * cells: proportional, each non-empty part at least one cell, summing to width.
252 */
253export function segments(read: number, write: number, fresh: number, width: number): [number, number, number] {
254  const total = read + write + fresh
255  if (total === 0 || width <= 0) return [0, 0, 0]
256  const parts = [read, write, fresh]
257  const cells = parts.map(p => (p > 0 ? Math.max(1, Math.round((p / total) * width)) : 0))
258  let over = cells.reduce((a, b) => a + b, 0) - width
259  while (over !== 0) {
260    const i = over > 0 ? cells.indexOf(Math.max(...cells)) : parts.indexOf(Math.max(...parts))
261    cells[i] += over > 0 ? -1 : 1
262    over += over > 0 ? -1 : 1
263  }
264  return [cells[0], cells[1], cells[2]]
265}
266
267/** Seconds left at which a toast counts down after the one at the warning threshold. */
268export const COUNTDOWN_MARKS = [10, 3, 2, 1]
269
270/**
271 * The toast mark to fire now, or undefined. `level` is the mark last fired for
272 * this cache entry (Infinity before any); a late tick skips straight to the
273 * newest mark crossed, so a stalled clock never replays old ones.
274 */
275export function nextToastMark(secsLeft: number, warnSecs: number, level: number): number | undefined {
276  const marks = [warnSecs, ...COUNTDOWN_MARKS].filter(m => m <= warnSecs)
277  const due = marks.filter(m => secsLeft <= m && m < level)
278  return due.length ? Math.min(...due) : undefined
279}
280
281export type LifeColor = 'green' | 'yellow' | 'red'
282
283/** Countdown colour: green while there is plenty, yellow below 40% of the lifetime, red from the warning threshold down. */
284export function lifeColor(leftMs: number, ttl: Ttl, warnMs: number): LifeColor {
285  if (leftMs <= warnMs) return 'red'
286  return leftMs / ttlMs(ttl) <= 0.4 ? 'yellow' : 'green'
287}
288
289// requests are timed from their start, so a little slack keeps a hit that
290// landed just inside the lifetime from reading as proof of the longer one
291const SLACK_MS = 10_000
292
293/**
294 * What the traffic says about the cache lifetime, given the request before and
295 * `known`, what earlier requests already showed.
296 *
297 *   - a hit (the cache served at least half of the previous prompt) more than
298 *     5 minutes after the previous request began proves the 1-hour lifetime,
299 *     and nothing later undoes it: a miss afterwards is more likely a changed
300 *     prefix than a lapse
301 *   - a miss with the same model and a prompt that did not shrink, 5 minutes to
302 *     an hour after the previous request, says the entry lapsed: 5 minutes
303 *     (weaker: a changed prefix looks the same, so a later hit overrules it)
304 *
305 * Needed because the API names the TTL of a write (`cache_creation.ephemeral_*`)
306 * but Claude Code's mod API passes on only the four token counts.
307 */
308export function observeTtl(prev: Sample | undefined, cur: Sample, known: Ttl | undefined): Ttl | undefined {
309  if (!prev || prev.read + prev.write === 0 || cur.model !== prev.model) return known
310  const gap = cur.startedAt - prev.startedAt
311  const before = promptTokens(prev)
312  if (gap <= ttlMs('5m') + SLACK_MS) return known
313  if (cur.read >= before * 0.5) return '1h'
314  if (known === '1h') return known
315  const lapsed = cur.write > 0 && promptTokens(cur) >= before * 0.7 && gap < ttlMs('1h') + SLACK_MS
316  return lapsed ? '5m' : known
317}
318
types/index.d.ts 18 lines
1/** One main-loop request as the cache saw it; kept in $.state so a reload keeps the last request. */
2export type SavedSample = {
3  turnId: string
4  index: number
5  model: string
6  startedAt: number
7  read: number
8  write: number
9  fresh: number
10  output: number
11}
12
13declare module 'claude-code' {
14  interface PluginState {
15    'cache-line': { samples: SavedSample[] }
16  }
17}
18