SLOPSHOPPER

prompt-cache-control

Prompt-cache meter above the Claude Code prompt: how many tokens each request read from, wrote to and sent past the cache, a live countdown to the cache's…

newpanebandcommandtoaststatus
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · prompt-cache-control
│ ┃ cache ✕ › fix the failing auth test and add an audit log call │ ┃ ⚡ PROMPT CACHE ○ no request yet: the first o │ ┃ cache ● prompt-cache-control: prompt-cache-control loaded: 1h cache (Claude │ ┃ ⏺ Read(src/auth.ts) │ ┃ CACHE ⎿ Read 6 lines │ ┃ session 8767:11:26 ⏺ Update(src/auth.ts) │ ┃ lifetime 1h (Claude subscription defaul ⎿ Added 2 lines, removed 1 line │ ┃ expires in --:-- ⏺ Bash(bun test) │ ┃ ⎿ 3 pass, 1 fail │ ┃ KEEPWARM │ ┃ window ● Done. refresh now rejects expired claims and logs an audit event. │ ┃ next ping after the next turn every 50m │ ┃ ✻ Worked for 42s · done 4:20 PM │ ┃ COST │ ┃ guard refuse once on a cold cache of › /cache │ ┃ 5h window 🧪 calibrating… · 1 active ses ⎿ prompt-cache-control: 1h cache (Claude subscription default) · n │ ┃ this session 0 cold writes · $0.00 │ ┃ │ ┃ TURNS │ ┃ no requests yet │ ┃ │ ┃ [ close ] │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⚠ prompt-cache-control: 5h 31%

Draws

Pane · cache
⚡ PROMPT CACHE ○ no request yet: the first one writes the ca CACHE session 8767:11:26 lifetime 1h (Claude subscription default) expires in --:-- KEEPWARM window 6h00m left [ next ping after the next turn every 50m COST guard refuse once on a cold cache of 50k+ tokens 5h window 🧪 calibrating… · 1 active session this session 0 cold writes · $0.00 TURNS no requests yet [ close ]
README

prompt-cache-control (Claude Cache Control)

A prompt-cache meter above the Claude Code prompt. Every request Claude makes reports how much of its prompt the cache served, how much it wrote and how much went uncached; this mod keeps those numbers per request and per turn, counts down to the moment the cache lapses and tells you what to do about it: keep going, /compact or /clear.

cache ██████████ 98% read 80k · wrote 1k · new 300 ⏱ 3:41 5m · warm: keep going
cache ░░░░░░░░░░  0% read 0 · wrote 52k · new 300 ⏱ 4:58 5m · cache missed: model changed (…)
cache ██████████ 98% read 150k · wrote 1k · new 300 ⏱ 0:00 5m · expired: the next message rewrites 151k tokens. /compact first, or /clear if the task is done

/cache opens a pane: the time left with a solid bar that shrinks as the cache runs out (green, yellow below 40% of the lifetime, red from the warning threshold), a stacked read / wrote / new bar for the last request, and a colour-coded table with one row per turn. Bars are filled cells and columns have fixed widths with no-break spaces, so the terminal and the Desktop (HTML) pane render the same.

How the countdown works

From Anthropic's prompt caching documentation:

  • The cache lives 5 minutes by default, 1 hour when asked for.
  • Every request that reads the cache refreshes it at no extra cost, so a conversation that keeps talking keeps the 5-minute cache warm.
  • The lifetime is counted from the start of the request that wrote or read the entry; generation time counts against it.
  • A prompt is input_tokens (uncached remainder) + cache_read_input_tokens + cache_creation_input_tokens.
  • Writes cost 1.25x base input for 5 minutes and 2x for 1 hour; reads cost about 0.1x (less on some models). The expensive moment is an expired cache on a large context, which is when this mod suggests /compact.
  • /clear starts a new conversation in the same process, so the meter and the /cache table start over with it. A change in the prefix (model, effort or thinking settings, tool set, system prompt, CLAUDE.md) makes the next request write instead of read. The mod names the cause when it sees a miss: model changed, the cache had lapsed, or the prefix changed.

Which lifetime your account gets

The mod follows Claude Code's own rules (prompt caching: cache lifetime, Claude Code 2.1.242 or later). For the main conversation the TTL is the first match of:

#SourceResult
1the mod's ttl option (5m / 1h)what you set
2FORCE_PROMPT_CACHING_5M=15 minutes
3CLAUDE_CODE_PROMPT_CACHE_TTL5m or 1h
4the promptCacheTtl setting (local, project or user settings file)5m or 1h
5ENABLE_PROMPT_CACHING_1H=11 hour
6the account1 hour on a Claude subscription within its plan usage; 5 minutes on usage credits, an API key or a cloud provider

The account comes from the rate-limit windows the last response reported: a five_hour or seven_day window means a subscription, and one at 100% means requests now draw on usage credits. An API key or a cloud provider reports no such window, and before the first response nothing is known, so the mod starts from 5 minutes there. Managed settings are not readable from a mod.

On top of that the mod watches the traffic, which beats rows 2 to 6: a request that hits the cache more than 5 minutes after the previous one proves the 1-hour lifetime (a later miss does not undo it, since a changed prefix looks the same), and a miss 5 to 60 minutes after the previous request, with the same model and a prompt that did not shrink, says the entry lapsed, so 5 minutes (a later hit overrules it). That covers what the mod cannot see: managed settings, a gateway that rewrites the TTL, or a subscription that ran out of plan usage mid-session. The pane header names the source in use.

Why the mod infers instead of reading it: the API names the TTL of each write (cache_creation.ephemeral_5m_input_tokens / ephemeral_1h_input_tokens) and Claude Code's status line exposes it as prompt_cache.ttl, but the mod API passes on only the four token counts. To check by hand, claude -p "hello" --output-format json and read usage.cache_creation.

Other switches read from the environment at session start:

VariableEffect on the meter
DISABLE_PROMPT_CACHING=1 (and _HAIKU, _SONNET, _OPUS)the band says caching is off for that model

What it hooks

  • turn.step: reads each main-loop request's usage (subagents have their own prefixes and are left out)
  • $.clock.every(1000): redraws the countdown, and only while its text changes, so an idle expired session costs nothing
  • ui.render on AbovePrompt (the band) and on Pane (/cache)
  • $.ui.toast: once per cache entry at the warning threshold (60 s by default) and again at 10, 3, 2 and 1 seconds left, for prompts of 20k tokens or more

Options

  ttl: string               "auto" | "5m" | "1h" (default auto)
  warnSeconds: number       countdown threshold for the yellow state and the toast (default 60)
  compactAtTokens: number   prompt size that makes an expired cache suggest /compact (default 100000)
  band: boolean             row above the prompt (default true)
  status: boolean           entry under the prompt, "cache 98% · 3:41" (default false)
  toast: boolean            toasts at the threshold, 10, 3, 2 and 1 s (default true)

The 100k compactAtTokens is a judgement, not a figure from the documentation: lower it if your model's cache writes are expensive for you.

Install

npx claude-code-templates@latest --mod observability/prompt-cache-control
claude

It is written to .claude/skills/prompt-cache-control/, which Claude Code auto-loads as prompt-cache-control@skills-dir once the workspace trust prompt is accepted. For one session with hot reload: claude --plugin-dir .claude/skills/prompt-cache-control. claude plugin validate .claude/skills/prompt-cache-control prints every event it hooks and every $ call it makes; claude plugin test .claude/skills/prompt-cache-control runs its tests.

Options are read from user settings (~/.claude/settings.json, never project settings), --settings <file> or managed settings, under the plugin's full id:

{ "pluginConfigs": { "prompt-cache-control@skills-dir": { "options": { } } } }

Requirements. Mods are on by default in Claude Code 2.1.287+. Typed against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods

Source 2 files
hooks/prompt-cache-control.tsx 1107 lines
1/**
2 * prompt-cache-control — Claude Mod
3 *
4 * A prompt-cache meter for Claude Code. Every main-loop request reports how
5 * many prompt tokens the cache served (`cache_read_input_tokens`), wrote
6 * (`cache_creation_input_tokens`) and sent uncached (`input_tokens`); this mod
7 * keeps those per request and per turn, counts down to the moment the cache
8 * lapses, and says what to do about it: keep going, /compact or /clear.
9 *
10 *   - `turn.step` reads each main-loop request's usage (subagents have their
11 *     own prefixes and are left out)
12 *   - `$.clock.every(1000)` redraws the countdown, and only while its text
13 *     changes: an idle, expired session costs nothing
14 *   - `/keepwarm` pings the cache after idle stretches (ported from cache-tax)
15 *   - a row above the prompt (the AbovePrompt component), an optional status
16 *     line entry, and `/cache`, a pane with one row per turn
17 *
18 * The lifetime is counted from the start of the request that last wrote or read
19 * the cache, as Anthropic documents it. Which lifetime Claude Code asked for
20 * follows its documented rules (see decideTtl in ./cache.ts): FORCE_PROMPT_CACHING_5M,
21 * CLAUDE_CODE_PROMPT_CACHE_TTL, the promptCacheTtl setting, ENABLE_PROMPT_CACHING_1H,
22 * then the account (1 hour on a Claude subscription, 5 minutes otherwise). The
23 * API names the TTL of a write but the mod API passes on only the token counts,
24 * so the mod also watches the gaps between requests (a hit after more than 5
25 * minutes proves 1 hour; see observeTtl). `ttl: "5m" | "1h"` pins it.
26 *
27 * Needs Claude Code >= 2.1.287.
28 *
29 * Options (pluginConfigs["prompt-cache-control@skills-dir"].options):
30 *   ttl: "auto" | "5m" | "1h"   cache lifetime (default auto)
31 *   warnSeconds: number         countdown threshold for the warning (default 60)
32 *   compactAtTokens: number     prompt size that makes an expired cache suggest /compact (default 100000)
33 *   band: boolean               row above the prompt (default true)
34 *   breakdown: boolean          read/wrote/new token counts in the band (default true)
35 *   status: boolean             entry under the prompt (default false)
36 *   toast: boolean              toasts near expiry: at warnSeconds, then 10, 3, 2 and 1 s (default true)
37 */
38import type { EngineInterface, Register, RenderChildren } from 'claude-code'
39import {
40  calibRate,
41  windowSpend,
42  type Spend,
43  guardVerdict,
44  isColdWrite,
45  advise,
46  ttlMs,
47  COUNTDOWN_MARKS,
48  bar,
49  byTurn,
50  fmtClock,
51  fmtCountdown,
52  fmtLimit,
53  fmtTokens,
54  hitRatio,
55  isCachingDisabled,
56  accountOf,
57  decideTtl,
58  observeTtl,
59  lifeColor,
60  lifeRatio,
61  nextToastMark,
62  positive,
63  promptTokens,
64  remainingMs,
65  rowRatio,
66  segments,
67} from './cache.ts'
68import type { Account, Advice, CacheEnv, GuardMode, Sample, Ttl } from './cache.ts'
69
70const PANE = 'cache'
71// the pane's label column
72const LABEL_W = 13
73const COMMAND = 'cache'
74const KEEP = 200
75// below this a lapsed cache costs too little to interrupt anyone about
76const TOAST_MIN_TOKENS = 20_000
77
78let samples: Sample[] = []
79let ttl: Ttl = '5m'
80let baseTtl: Ttl = '5m'
81let pinned = false
82let observed: Ttl | undefined
83let setting: unknown
84let account: Account = 'other'
85/** the 5-hour plan window as the last response reported it, for the line under the prompt */
86let fiveHour: { percentUsed: number; resetsAt?: string } | undefined
87let ttlSource = 'default'
88let envSource = 'default'
89let env: CacheEnv = {}
90let timer: { cancel: () => void } | undefined
91let lastKey = ''
92let toastedFor = 0
93let toastLevel = Infinity
94let isPaneOpen = false
95/** when the session began (its first launch if resumed, /clear restarts it); 0 until known */
96let sessionStartedAt = 0
97
98type Policy = { warnMs: number; compactAtTokens: number }
99
100function current(policy: Policy, now: number) {
101  const last = samples[samples.length - 1]
102  const prev = samples[samples.length - 2]
103  const disabled = last ? isCachingDisabled(last.model, env) : isCachingDisabled('', env)
104  const advice: Advice = advise(last, prev, { ttl, ...policy }, now, disabled)
105  const left = last ? remainingMs(last, ttl, now) : 0
106  return { last, advice, left }
107}
108
109const COLOR: Record<Advice['kind'], string | undefined> = {
110  warm: 'green',
111  soon: 'yellow',
112  expired: 'red',
113  miss: 'red',
114  off: undefined,
115  cold: undefined,
116  uncached: undefined,
117}
118
119/** a hit rate with one decimal, e.g. 99.5% */
120const pct1 = (ratio: number) => `${(ratio * 100).toFixed(1)}%`
121
122async function openPane($: EngineInterface) {
123  isPaneOpen = true
124  await $.ui.open({ id: PANE, title: 'cache', focus: true })
125  $.ui.invalidate('ui.render')
126}
127
128function shortLine(policy: Policy, now: number): string {
129  const { last, advice, left } = current(policy, now)
130  if (!last || advice.kind === 'off') return `cache: ${advice.text}`
131  const clock = left > 0 ? ` · ${fmtCountdown(left)}` : ''
132  return `cache ${pct1(hitRatio(last))}${clock}`
133}
134
135// the promptCacheTtl setting, from the settings files that can carry it (local over project over user)
136async function readSetting($: EngineInterface): Promise<unknown> {
137  const home = await $.env.get('HOME').catch(() => undefined)
138  const cwd = await $.session.cwd().catch(() => undefined)
139  const files = [cwd && `${cwd}/.claude/settings.local.json`, cwd && `${cwd}/.claude/settings.json`, home && `${home}/.claude/settings.json`]
140  for (const file of files) {
141    if (!file) continue
142    try {
143      const value = JSON.parse(await $.fs.read(file)).promptCacheTtl
144      if (value === '5m' || value === '1h') return value
145    } catch {
146      // missing or unreadable: the next file
147    }
148  }
149  return undefined
150}
151
152
153// ---- /keepwarm, ported from cache-tax 2.2.1 (MIT): a tool-less $.model.fork after each idle stretch re-reads
154// the cached prefix; each ping is recorded as a Sample, so the meter counts it
155const DEFAULT_WINDOW_MS = 6 * 60 * 60 * 1000
156const MIN_PING_MS = 60 * 1000
157const PING_PROMPT = 'Reply with the single word: warm'
158const KEY_DEADLINE = 'keepwarm.deadline'
159const KEY_EVERY = 'keepwarm.every'
160const KEY_ALWAYS = 'keepwarm.always'
161
162// $ per million tokens, [cache read, 1h cache write, output], list prices September 2026 (from cache-tax).
163const PRICES: Array<[string, number, number, number]> = [
164  ['fable-5-1', 0.25, 20, 50],
165  ['fable-5', 1, 20, 50],
166  ['opus-5', 0.5, 10, 25],
167  ['opus-4', 0.5, 10, 25],
168  ['sonnet-5', 0.2, 4, 10],
169  ['sonnet', 0.3, 6, 15],
170  ['haiku', 0.1, 2, 5],
171]
172
173type Ping = { read: number; usd: number | null }
174
175function parseDuration(text: string): number | null {
176  const m = /^(?:(\d+)h)?(?:(\d+)m)?$/.exec(text.trim())
177  if (!m || (m[1] === undefined && m[2] === undefined)) return null
178  return (Number(m[1] ?? 0) * 60 + Number(m[2] ?? 0)) * 60 * 1000
179}
180
181function fmtDuration(ms: number): string {
182  const total = Math.max(0, Math.round(ms / 60000))
183  const h = Math.floor(total / 60)
184  const m = total % 60
185  return h > 0 ? `${h}h${String(m).padStart(2, '0')}m` : `${m}m`
186}
187
188function priceOf(model: string): [number, number, number] | null {
189  const m = model.toLowerCase().replace(/[\s.]+/g, '-')
190  for (const [family, read, write, output] of PRICES) if (m.includes(family)) return [read, write, output]
191  return null
192}
193
194function fmtUsd(usd: number | null): string {
195  return usd == null ? 'n/a' : '$' + usd.toFixed(2)
196}
197
198/** 50m on a 1h cache, about 4m on a 5m one */
199const defaultEvery = (ttlMs: number) => Math.max(MIN_PING_MS, Math.round((ttlMs * 5) / 6))
200
201let sid = ''
202let deadline = 0
203let windowStart = 0 // when the current window was armed, for the pane's bar
204let every = 0 // 0: the default for the TTL in force
205let always = false
206// ponytail: in memory, a plugin reload or a restarted process lets autoKeepwarm start it again
207let userOff = false // /keepwarm off or the pane's stop: autoKeepwarm leaves it off until it is started by hand
208let compacted = false
209let stopped: string | null = null
210let lastPing: Ping | null = null
211/** keepwarm pings in the current window, counted from the samples (they survive a restart with them) */
212// ponytail: a restart keeps the last SAVED samples, so a longer run of pings than that counts short
213const pingsInWindow = () => samples.filter(x => x.turnId.startsWith('keepwarm-') && x.startedAt >= windowStart).length
214let pending: { cancel: () => void } | null = null
215
216const period = () => every || defaultEvery(ttlMs(ttl))
217const lastAt = () => samples[samples.length - 1]?.startedAt ?? 0
218const isCold = (now: number) => lastAt() > 0 && now - lastAt() >= ttlMs(ttl)
219
220/** the band's segment, undefined when keepwarm is off */
221function keepwarmStatus(now: number): { text: string; stopped: boolean } | undefined {
222  if (stopped) return { text: `keepwarm stopped: ${stopped}`, stopped: true }
223  if (!deadline || now >= deadline) return undefined
224  const next = !lastAt() || compacted ? 'waiting for a turn'
225    : isCold(now) ? 'cold'
226    : `ping ${fmtDuration(lastAt() + period() - now)}`
227  const pings = pingsInWindow()
228  const last = pings ? ` · ${pings} ping${pings === 1 ? '' : 's'}` : ''
229  return { text: `♨ keepwarm ${fmtDuration(deadline - now)} · ${next}${last}`, stopped: false }
230}
231
232/** keepwarm for the pane, as values rather than one line */
233function keepwarmInfo(now: number) {
234  const on = !!deadline && now < deadline
235  const next = !on ? undefined
236    : !lastAt() || compacted ? 'after the next turn'
237    : isCold(now) ? 'cold, after the next turn'
238    : `in ${fmtDuration(lastAt() + period() - now)}`
239  return {
240    on,
241    stopped,
242    always,
243    left: on ? deadline - now : 0,
244    window: on ? Math.max(1, deadline - windowStart) : 1,
245    next,
246    every: period(),
247    lastPing,
248  }
249}
250
251/** from the module's session.start (one hook per event per module) */
252async function keepwarmStart($: EngineInterface) {
253  disarm()
254  sid = await $.session.id()
255  const now = Date.now()
256  const saved = await $.store.get(keyOf(KEY_DEADLINE))
257  const savedEvery = await $.store.get(keyOf(KEY_EVERY))
258  deadline = typeof saved === 'number' && saved > now ? saved : 0
259  // a restored window's start is not stored: assume the default length
260  windowStart = deadline ? Math.min(now, deadline - DEFAULT_WINDOW_MS) : 0
261  if (!deadline) {
262    await $.store.delete(keyOf(KEY_DEADLINE))
263    await $.store.delete(keyOf(KEY_EVERY))
264  }
265  every = typeof savedEvery === 'number' && savedEvery >= MIN_PING_MS ? savedEvery : 0
266  always = (await $.store.get(KEY_ALWAYS)) === true
267  if (always) await startWindow($, DEFAULT_WINDOW_MS, 0)
268  await $.command.register({
269    name: 'keepwarm',
270    description: 'Keep the prompt cache warm: bare for 6h, a window such as 90m, always, off, or status',
271    argumentHint: '[6h | always | off | status]',
272    immediate: true,
273  })
274  await arm($)
275}
276
277/** from the module's session.end */
278async function keepwarmEnd($: EngineInterface, reason: string) {
279  // /clear starts a new conversation: nothing to keep warm until it has a turn
280  if (reason === 'clear') {
281    await stop($, null)
282    lastPing = null
283    userOff = false
284    misses = []
285    ackedAt = 0
286    coldWritePending = false
287  } else disarm()
288}
289
290const keyOf = (k: string) => `${k}:${sid}`
291
292const disarm = () => {
293  pending?.cancel()
294  pending = null
295}
296
297async function stop($: EngineInterface, why: string | null, forgetAlways = false) {
298  if (forgetAlways) userOff = true
299  deadline = 0
300  every = 0
301  stopped = why
302  disarm()
303  await $.store.delete(keyOf(KEY_DEADLINE))
304  await $.store.delete(keyOf(KEY_EVERY))
305  if (forgetAlways) {
306    always = false
307    await $.store.delete(KEY_ALWAYS)
308  }
309  $.ui.invalidate('ui.render')
310}
311
312async function arm($: EngineInterface) {
313  disarm()
314  if (!deadline) return
315  const now = Date.now()
316  if (now >= deadline) return stop($, null)
317  // a cold or compacted cache is not pinged: the next turn writes it anew, then pinging resumes
318  if (lastAt() && !compacted && !isCold(now)) {
319    const delay = Math.min(deadline - now, lastAt() + ttlMs(ttl) - now, Math.max(1000, lastAt() + period() - now))
320    pending = $.clock.after(delay, () => void ping($))
321  } else {
322    pending = $.clock.after(deadline - now, () => void arm($))
323  }
324  $.ui.invalidate('ui.render')
325}
326
327async function ping($: EngineInterface) {
328  pending = null
329  if (!deadline) return
330  const now = Date.now()
331  if (now >= deadline || isCold(now)) return arm($)
332  // a turn in the meantime re-armed the timer; this callback is stale
333  if (now - lastAt() < period() - 1000) return
334  const model = samples[samples.length - 1]?.model ?? ''
335  let reply
336  try {
337    reply = await $.model.fork({ prompt: PING_PROMPT })
338  } catch (err) {
339    return stop($, `the ping failed, ${err instanceof Error ? err.message : String(err)}`)
340  }
341  if (reply.isAnswered === false) {
342    const reason = reply.reason === 'nothing-to-fork' ? 'no conversation to warm yet'
343      : reply.reason === 'api-error' ? `the API call failed${reply.status === null ? '' : ` (${reply.status})`}`
344      : reply.reason === 'aborted' ? 'the ping was interrupted'
345      : 'the ping returned no text'
346    return stop($, reason)
347  }
348  const u = reply.usage
349  const price = priceOf(model)
350  const usd = price
351    ? (u.cache_read_input_tokens * price[0] + u.cache_creation_input_tokens * price[1] + (u.input_tokens * price[1]) / 2 + u.output_tokens * price[2]) / 1e6
352    : null
353  lastPing = { read: u.cache_read_input_tokens, usd }
354  if (usd != null) spentUsd += usd
355  await pushSample($, {
356    turnId: `keepwarm-${now}`,
357    index: 0,
358    model,
359    startedAt: now,
360    read: u.cache_read_input_tokens,
361    write: u.cache_creation_input_tokens,
362    fresh: u.input_tokens,
363    output: u.output_tokens,
364  })
365  $.ui.invalidate('ui.render')
366  // a warm ping reads the prefix and writes only its own message; a write of a tenth of the read or more means the prefix broke
367  const warm = u.cache_read_input_tokens > 0 && u.cache_creation_input_tokens < 0.1 * u.cache_read_input_tokens
368  if (!warm) return stop($, `the ping wrote ${fmtTokens(u.cache_creation_input_tokens)} tokens (${fmtUsd(usd)}), the cache was already gone`)
369  await arm($)
370}
371
372async function startWindow($: EngineInterface, windowMs: number, everyMs: number) {
373  every = everyMs
374  if (every) await $.store.set(keyOf(KEY_EVERY), every)
375  else await $.store.delete(keyOf(KEY_EVERY))
376  windowStart = Date.now()
377  deadline = windowStart + windowMs
378  userOff = false
379  stopped = null
380  await $.store.set(keyOf(KEY_DEADLINE), deadline)
381  await arm($)
382}
383
384function statusLine(now: number): string | undefined {
385  if (stopped) return `keepwarm stopped: ${stopped}`
386  if (!deadline || now >= deadline) return undefined
387  const next = !lastAt() ? ' · waiting for the first turn'
388    : compacted ? ' · waiting for the first turn after compaction'
389    : isCold(now) ? ` · cold now, first ping ${fmtDuration(period())} after the next turn`
390    : ` · ping in ${fmtDuration(lastAt() + period() - now)}`
391  const last = lastPing ? ` · last ping read ${fmtTokens(lastPing.read)} ${fmtUsd(lastPing.usd)}` : ''
392  return `keepwarm ${fmtDuration(deadline - now)} left${next}${last}`
393}
394
395function armedText(windowMs: number): string {
396  if (isCold(Date.now())) return `keepwarm on for ${fmtDuration(windowMs)}. The cache is cold now, so the first ping comes ${fmtDuration(period())} after the next turn`
397  return `keepwarm on for ${fmtDuration(windowMs)}, a ping ${fmtDuration(period())} after each idle stretch keeps the cache read, not re-written`
398}
399
400// ---- cold-cache price and guard, ported from cache-tax 2.2.1 (MIT)
401/** the last request's cache has lapsed: the next message writes the whole prefix again */
402function coldSince(now: number): number | undefined {
403  const last = samples[samples.length - 1]
404  if (!last || compacted || last.read + last.write === 0) return undefined
405  const at = last.startedAt + ttlMs(ttl)
406  return now >= at ? at : undefined
407}
408
409/** $ per million tokens written to the cache for `model` at the TTL in force: the 1h rate is 2x base, the 5m one 1.25x */
410function writeRate(model: string): number | null {
411  const price = priceOf(model)
412  return price ? (ttl === '1h' ? price[1] : price[1] * 0.625) : null
413}
414
415const AUTO_WARM_MS = 3 * 60 * 60 * 1000
416type Miss = { at: number; turnId: string; tokens: number; usd: number | null }
417/** this session's cold writes */
418let misses: Miss[] = []
419/** the guard let a message to a cold cache through: its turn pays a cold write */
420let coldWritePending = false
421
422/** the request of `turnId` that re-wrote the prefix: after the guard let it through, or one that wrote half the previous prompt or more */
423function coldWriteOf(turnId: string): Sample | undefined {
424  const i = samples.findIndex(x => x.turnId === turnId)
425  if (i < 0) return undefined
426  const prevTokens = i > 0 ? promptTokens(samples[i - 1]) : 0
427  const first = samples[i]
428  return isColdWrite(prevTokens, first.write, coldWritePending) ? first : undefined
429}
430
431
432/** what re-writing the last prompt costs, and what a warm turn would have read it for */
433function coldPrice(): { tokens: number; cold: number | null; warm: number | null } {
434  const last = samples[samples.length - 1]
435  const tokens = last ? promptTokens(last) : 0
436  const price = priceOf(last?.model ?? '')
437  const write = writeRate(last?.model ?? '')
438  return { tokens, cold: write == null ? null : (tokens * write) / 1e6, warm: price ? (tokens * price[0]) / 1e6 : null }
439}
440
441/** the band's and pane's segment while the cache is cold */
442function coldStatus(now: number): string | undefined {
443  if (coldSince(now) === undefined) return undefined
444  const { tokens, cold } = coldPrice()
445  return `❄ re-warm ${fmtTokens(tokens)} ≈ ${fmtUsd(cold)}`
446}
447
448function guardText(now: number): string {
449  const since = coldSince(now) ?? now
450  const { tokens, cold, warm } = coldPrice()
451  return `the prompt cache went cold ${fmtDuration(now - since)} ago. Sending this re-writes about ${fmtTokens(tokens)} tokens ≈ ${fmtUsd(cold)}` +
452    (warm == null ? '' : ` (a warm turn would have cost ${fmtUsd(warm)})`) + '.'
453}
454
455let ackedAt = 0
456
457// ---- experimental: list price as a share of the 5-hour plan window. The engine
458// reports the window's percentUsed but not its size, so this session's spend at
459// list price is set against how far the window moved. Each session writes its
460// spend in the window under `spend:<sid>` in the shared store and the estimate
461// sums them all; sessions without this mod (claude.ai, other machines) still
462// make it read high.
463const KEY_CALIB = 'calib.pctPerUsd'
464const KEY_SPEND = 'spend'
465/** this session's spend at list price, requests and keepwarm pings */
466let spentUsd = 0
467/** the window this session last wrote its spend for, and spentUsd when that window began for it */
468let winResetsAt: string | undefined
469let winBase = 0
470/** where the window and every session's spend stood when this window was first seen */
471let anchor: { resetsAt: string; pct: number; usd: number } | undefined
472/** the last estimate: percent of the 5-hour window per list-price dollar */
473let pctPerUsd: number | undefined
474/** sessions that wrote their spend in the last 10 minutes, this one included */
475let activeSessions = 0
476
477/** a call's list price: cache read, cache write at the TTL in force, uncached input at base (half the 1h write rate), output */
478function usdOf(u: { input_tokens: number; output_tokens: number; cache_read_input_tokens: number; cache_creation_input_tokens: number }, model: string): number | null {
479  const price = priceOf(model)
480  const rate = writeRate(model)
481  if (!price || rate == null) return null
482  return (u.cache_read_input_tokens * price[0] + u.cache_creation_input_tokens * rate + (u.input_tokens * price[1]) / 2 + u.output_tokens * price[2]) / 1e6
483}
484
485/** a turn's requests at list price, each at its own model's rates; null when a model's price is unknown */
486function turnUsd(turnId: string): number | null {
487  let sum = 0
488  for (const x of samples) {
489    if (x.turnId !== turnId) continue
490    const usd = usdOf({ input_tokens: x.fresh, output_tokens: x.output, cache_read_input_tokens: x.read, cache_creation_input_tokens: x.write }, x.model)
491    if (usd == null) return null
492    sum += usd
493  }
494  return sum
495}
496
497/** after a turn: move the anchor on a new window, otherwise update the estimate */
498async function calibrate($: EngineInterface) {
499  const w = (await $.session.usage().catch(() => undefined))?.rateLimits?.find(x => x.kind === 'five_hour')
500  if (!w || !w.resetsAt) return
501  const mine = keyOf(KEY_SPEND)
502  if (winResetsAt !== w.resetsAt) {
503    // a reload zeroes spentUsd: carry on from what this session wrote for the same window
504    const saved = winResetsAt === undefined ? ((await $.store.get(mine)) as Partial<Spend> | undefined) : undefined
505    winBase = saved?.resetsAt === w.resetsAt && typeof saved.usd === 'number' ? spentUsd - saved.usd : spentUsd
506    winResetsAt = w.resetsAt
507  }
508  const now = Date.now()
509  await $.store.set(mine, { resetsAt: w.resetsAt, usd: spentUsd - winBase, at: now } satisfies Spend)
510  // ponytail: one JSON file for every session, two writing at once can drop one write; the next turn rewrites it
511  const keys = (await $.store.keys()).filter(k => k.startsWith(`${KEY_SPEND}:`))
512  const all = await Promise.all(keys.map(async k => [k, await $.store.get(k)] as [string, unknown]))
513  const { usd, active, stale } = windowSpend(all, w.resetsAt, now)
514  activeSessions = active
515  for (const k of stale) await $.store.delete(k)
516  if (!anchor || anchor.resetsAt !== w.resetsAt || w.percentUsed < anchor.pct) {
517    anchor = { resetsAt: w.resetsAt, pct: w.percentUsed, usd }
518    return
519  }
520  const r = calibRate(anchor.pct, anchor.usd, w.percentUsed, usd)
521  if (r !== undefined) {
522    pctPerUsd = r
523    await $.store.set(KEY_CALIB, r)
524  }
525}
526
527/** recount the active sessions from the store, for /cache between turns */
528async function refreshActive($: EngineInterface) {
529  if (!winResetsAt) return
530  const keys = (await $.store.keys()).filter(k => k.startsWith(`${KEY_SPEND}:`))
531  const all = await Promise.all(keys.map(async k => [k, await $.store.get(k)] as [string, unknown]))
532  activeSessions = windowSpend(all, winResetsAt, Date.now()).active
533}
534
535const sessionsText = () => `${activeSessions} active session${activeSessions === 1 ? '' : 's'}`
536
537// the engine's process can stop under an idle session and start again on resume, which runs session.start
538// with the module's variables zeroed: keep the latest requests in the store, under `samples:<sid>`
539const KEY_SAMPLES = 'samples'
540const SAVED = 30
541const SAVED_FOR_MS = 7 * 24 * 60 * 60 * 1000
542
543async function pushSample($: EngineInterface, sample: Sample) {
544  samples.push(sample)
545  if (samples.length > KEEP) samples = samples.slice(-KEEP)
546  lastKey = ''
547  await $.store.set(keyOf(KEY_SAMPLES), { at: Date.now(), samples: samples.slice(-SAVED) }).catch(() => undefined)
548}
549
550/** this session's saved requests; drops other sessions' older than a week */
551async function restoreSamples($: EngineInterface): Promise<Sample[]> {
552  const now = Date.now()
553  let mine: Sample[] = []
554  for (const k of (await $.store.keys()).filter(k => k.startsWith(`${KEY_SAMPLES}:`))) {
555    const v = (await $.store.get(k)) as { at?: number; samples?: Sample[] } | undefined
556    if (k === keyOf(KEY_SAMPLES)) mine = Array.isArray(v?.samples) ? v.samples : []
557    else if (!v || typeof v.at !== 'number' || now - v.at > SAVED_FOR_MS) await $.store.delete(k)
558  }
559  return mine
560}
561
562export const register: Register = (on, options) => {
563  const policy: Policy = {
564    warnMs: positive(options.warnSeconds, 60) * 1000,
565    compactAtTokens: positive(options.compactAtTokens, 100_000),
566  }
567  const showBand = options.band !== false
568  const showBreakdown = options.breakdown !== false
569  const showStatus = options.status === true
570  const showLimit = options.limit !== false
571  const autoWarm = options.autoKeepwarm === true
572  /** the entry under the prompt: the cache's short line and/or the 5-hour window */
573  const statusText = (now: number) =>
574    [showStatus ? shortLine(policy, now) : undefined, showLimit ? fmtLimit(fiveHour, now) : undefined].filter(Boolean).join(' · ') || undefined
575  const wantToast = options.toast !== false
576  const guard: GuardMode = options.guard === 'warn' || options.guard === 'off' ? options.guard : 'refuse'
577
578  on('session.start', async ($, e, next) => {
579    const r = await next(e)
580    samples = []
581    lastKey = ''
582    toastedFor = 0
583    const none = () => undefined
584    env = {
585      enable1h: await $.env.get('ENABLE_PROMPT_CACHING_1H').catch(none),
586      force5m: await $.env.get('FORCE_PROMPT_CACHING_5M').catch(none),
587      ttlVar: await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL').catch(none),
588      disableAll: await $.env.get('DISABLE_PROMPT_CACHING').catch(none),
589      disableHaiku: await $.env.get('DISABLE_PROMPT_CACHING_HAIKU').catch(none),
590      disableSonnet: await $.env.get('DISABLE_PROMPT_CACHING_SONNET').catch(none),
591      disableOpus: await $.env.get('DISABLE_PROMPT_CACHING_OPUS').catch(none),
592    }
593    pinned = options.ttl === '5m' || options.ttl === '1h'
594    observed = undefined
595    setting = await readSetting($)
596    const usage = await $.session.usage().catch(() => undefined)
597    sessionStartedAt = usage?.startedAt ?? 0
598    const limits = usage?.rateLimits ?? []
599    account = accountOf(limits)
600    fiveHour = limits.find(x => x.kind === 'five_hour')
601    const choice = decideTtl(options.ttl, env, setting, account)
602    baseTtl = choice.ttl
603    ttl = baseTtl
604    envSource = choice.source
605    ttlSource = envSource
606
607    await $.command
608      .register({
609        name: COMMAND,
610        description: 'Prompt-cache usage per turn and the time left before it lapses (stop closes)',
611        argumentHint: '[stop]',
612        immediate: true,
613      })
614      .catch(err => $.ui.log(`prompt-cache-control: /${COMMAND} not registered: ${err}`))
615    $.ui.log(`prompt-cache-control loaded: ${ttl} cache (${ttlSource}), /${COMMAND} opens the table`, { to: 'debug' })
616
617    // the cache's part waits for the first request; this also clears what a previous load left
618    $.ui.status(showLimit ? fmtLimit(fiveHour, Date.now()) : undefined)
619    spentUsd = 0
620    winResetsAt = undefined
621    winBase = 0
622    activeSessions = 0
623    anchor = undefined
624    const savedRate = await $.store.get(KEY_CALIB).catch(() => undefined)
625    pctPerUsd = typeof savedRate === 'number' && savedRate > 0 ? savedRate : undefined
626    // a resumed session, or a restarted process under it, goes on from its saved requests
627    sid = await $.session.id().catch(() => '')
628    samples = sid ? await restoreSamples($).catch(() => []) : []
629    // keepwarm failing to start must not take the meter down with it
630    try {
631      await keepwarmStart($)
632    } catch (err) {
633      $.ui.log(`prompt-cache-control: /keepwarm not started: ${err}`, { to: 'debug' })
634    }
635    timer?.cancel()
636    timer = $.clock.every(1000, () => {
637      const now = Date.now()
638      const { last, advice, left } = current(policy, now)
639      // the band counts in minutes from 10 minutes up; the open pane counts seconds, so it redraws each one
640      const key = `${advice.kind}|${advice.text}|${left > 0 ? (isPaneOpen ? fmtClock(left) : fmtCountdown(left)) : ''}|${keepwarmStatus(now)?.text ?? ''}|${coldStatus(now) ?? ''}|${showLimit ? fmtLimit(fiveHour, now) ?? '' : ''}|${isPaneOpen && sessionStartedAt ? fmtClock(now - sessionStartedAt) : ''}`
641      if (key !== lastKey) {
642        lastKey = key
643        if (showStatus || showLimit) $.ui.status(statusText(now))
644        $.ui.invalidate('ui.render')
645      }
646      if (wantToast && last && left > 0 && promptTokens(last) >= TOAST_MIN_TOKENS) {
647        if (toastedFor !== last.startedAt) {
648          toastedFor = last.startedAt
649          toastLevel = Infinity
650        }
651        // the first toast comes at warnSeconds, then 10, 3, 2 and 1 seconds; a late tick skips to the newest one
652        const secs = Math.ceil(left / 1000)
653        const mark = nextToastMark(secs, policy.warnMs / 1000, toastLevel)
654        if (mark !== undefined) {
655          toastLevel = mark
656          const tail = secs <= COUNTDOWN_MARKS[0] ? 'send a message now' : `send a message to keep ${fmtTokens(promptTokens(last))} tokens warm`
657          $.ui.toast(`cache expires in ${secs >= 60 ? fmtClock(left) : `${secs}s`}: ${tail}`)
658        }
659      }
660    })
661    return r
662  })
663
664  on('session.end', async ($, e, next) => {
665    await keepwarmEnd($, e.reason)
666    // /clear starts a new conversation in the same process: its cache is a new one
667    if (e.reason === 'clear') {
668      samples = []
669      lastKey = ''
670      toastedFor = 0
671      observed = undefined
672      ttl = baseTtl
673      ttlSource = envSource
674      $.ui.invalidate('ui.render')
675      return next(e)
676    }
677    timer?.cancel()
678    timer = undefined
679    return next(e)
680  })
681
682  // the plan windows moved: the 5-hour one goes under the prompt
683  on('session.measure', async ($, e, next) => {
684    const w = e.rateLimits.find(x => x.kind === 'five_hour')
685    if (w) fiveHour = w
686    if (showLimit) $.ui.status(statusText(Date.now()))
687    return next(e)
688  })
689
690  // each main-loop request: what the cache did with it
691  on('turn.step', async function* ($, e, next) {
692    if (e.agentId) return yield* next(e)
693    const startedAt = Date.now()
694    const r = yield* next(e)
695    if (r.usage) {
696      await pushSample($, {
697        turnId: e.turnId,
698        index: e.index,
699        model: r.usage.model || e.model,
700        startedAt,
701        read: r.usage.cache_read_input_tokens,
702        write: r.usage.cache_creation_input_tokens,
703        fresh: r.usage.input_tokens,
704        output: r.usage.output_tokens,
705      })
706      // autoKeepwarm: a message opens a window when none is running, unless the person turned it off
707      if (autoWarm && !userOff && !(deadline > Date.now())) {
708        await startWindow($, DEFAULT_WINDOW_MS, every).catch(err => $.ui.log(`prompt-cache-control: autoKeepwarm not started: ${err}`, { to: 'debug' }))
709      }
710      if (!pinned) {
711        // the account can change under a session: a subscription running out of plan usage moves to usage credits
712        account = accountOf((await $.session.usage().catch(() => undefined))?.rateLimits ?? [])
713        const choice = decideTtl(options.ttl, env, setting, account)
714        baseTtl = choice.ttl
715        envSource = choice.source
716        if (observed === undefined) {
717          ttl = baseTtl
718          ttlSource = envSource
719        }
720        const seen = observeTtl(samples[samples.length - 2], samples[samples.length - 1], observed)
721        if (seen !== observed) {
722          observed = seen
723          ttl = seen ?? baseTtl
724          ttlSource = `observed from request timing; ${envSource} said ${baseTtl}`
725          $.ui.log(`prompt-cache-control: cache lifetime is ${ttl} (${ttlSource})`, { to: 'debug' })
726        }
727      }
728      lastKey = ''
729      if (showStatus || showLimit) $.ui.status(statusText(Date.now()))
730      $.ui.invalidate('ui.render')
731    }
732    return r
733  })
734
735  /** the pane as plain lines, for where nothing draws it (Remote Control) */
736  function textSummary(now: number): string {
737    const { last, advice, left } = current(policy, now)
738    const kwi = keepwarmInfo(now)
739    const { cold, warm } = coldPrice()
740    const isColdNow = coldSince(now) !== undefined
741    const price = priceOf(last?.model ?? '')
742    const rate = writeRate(last?.model ?? '')
743    const breakEven = price && rate ? Math.floor(rate / price[0]) : undefined
744    const paid = misses.reduce((a, m) => a + (m.usd ?? 0), 0)
745    const counting = !!last && advice.kind !== 'uncached' && advice.kind !== 'off'
746    const row = (label: string, value: string) => `${label.padEnd(LABEL_W)}${value}`
747    // the engine prefixes the plugin's name to the first line
748    const lines = [advice.text]
749    if (sessionStartedAt) lines.push(row('session', fmtClock(now - sessionStartedAt)))
750    if (counting) lines.push(row('expires in', left > 0 ? fmtClock(left) : '0:00'))
751    if (last) {
752      lines.push(row('last request', `${pct1(hitRatio(last))} hit`))
753    }
754    lines.push(row('keepwarm', kwi.stopped ? `stopped: ${kwi.stopped}`
755      : kwi.on ? `${fmtDuration(kwi.left)} left · next ping ${kwi.next ?? ''} · every ${fmtDuration(kwi.every)}`
756      : kwi.always ? 'off until next session (always)' : 'off (/keepwarm arms 6h)'))
757    if (kwi.lastPing) lines.push(row('last ping', `${fmtTokens(kwi.lastPing.read)} read · ${fmtUsd(kwi.lastPing.usd)}`))
758    if (breakEven !== undefined) lines.push(row('break-even', `${breakEven} pings = one cold write, ~${fmtDuration(breakEven * kwi.every)} idle`))
759    if (last && cold != null) {
760      lines.push(row('cold write', `${fmtUsd(cold)}${isColdNow ? ' · cold now: the next message pays it' : ''}` +
761        (warm != null ? ` (warm turn ${fmtUsd(warm)})` : '')))
762    }
763    lines.push(row('guard', guard === 'refuse' ? 'refuse once' : guard === 'warn' ? 'warn only' : 'off'))
764    if (pctPerUsd !== undefined && cold != null && warm != null) {
765      lines.push(row('5h window', `≈ ${(cold * pctPerUsd).toFixed(1)}% cold · ${(warm * pctPerUsd).toFixed(1)}% warm (experimental, ${sessionsText()})`))
766    }
767    lines.push(row('this session', `${misses.length} cold write${misses.length === 1 ? '' : 's'} · ${fmtUsd(paid)}`))
768    return lines.join('\n')
769  }
770
771  on('command.run', { command: COMMAND }, async ($, e) => {
772    const arg = e.args.trim().toLowerCase()
773    if (arg === 'stop') {
774      await $.ui.close({ id: PANE }).catch(() => undefined)
775      isPaneOpen = false
776      $.ui.invalidate('ui.render')
777      return { text: 'cache table closed' }
778    }
779    // nothing draws the pane under Remote Control: answer in text
780    await refreshActive($).catch(() => undefined)
781    const surfaces = await $.session.surfaces().catch(() => [])
782    if (arg === 'text' || surfaces.length === 0) return { text: textSummary(Date.now()) }
783    await openPane($)
784    const { advice } = current(policy, Date.now())
785    return { text: `${ttl} cache (${ttlSource}) · ${advice.text} · /${COMMAND} stop closes` }
786  })
787
788  on('ui.close', async ($, e, next) => {
789    if (e.id !== PANE) return next(e)
790    isPaneOpen = false
791    // the band hides while the pane is open: draw it again now, not at the next countdown tick
792    $.ui.invalidate('ui.render')
793    return next(e)
794  })
795
796  on('ui.press', async ($, e, next) => {
797    if (e.plugin !== $.plugin.name || e.requestId !== PANE) return next(e)
798    if (e.element === 'close') await $.ui.close({ id: PANE }).catch(() => undefined)
799    if (e.element === 'kw-start') await startWindow($, DEFAULT_WINDOW_MS, 0)
800    if (e.element === 'kw-stop') await stop($, null, true)
801    return next(e)
802  })
803
804  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
805    if (!showBand || e.props.hasSurvey || isPaneOpen) return next(e)
806    const { last, advice, left } = current(policy, Date.now())
807    const { Box, Text, Button } = $.ui.resolve(e)
808    const columns = e.viewport?.columns ?? 100
809    const color = COLOR[advice.kind]
810    // other mods' rows above the prompt come back from next(e). Ours goes on top, with a gap before theirs,
811    // so the order does not hang on which mod the engine loaded first (it differs by install source)
812    const rest = await next(e)
813    const band = (row: RenderChildren) => (
814      <Box flexDirection="column">
815        {row}
816        {rest ? <Box key="rest" flexDirection="column" marginTop={1}>{rest}</Box> : null}
817      </Box>
818    )
819
820    // no request yet (new session, /clear, a plugin reload): a placeholder row, so the band does not vanish
821    if (!last) {
822      return band(
823        <Box key="cache" flexDirection="row" columnGap={1}>
824          <Text dimColor wrap="truncate-end">{advice.kind === 'off' ? `cache: ${advice.text}` : '○ cache · waiting for the first request'}</Text>
825          <Button key="open" label="details" onPress={() => openPane($)} />
826        </Box>,
827      )
828    }
829
830    const ratio = hitRatio(last)
831    const wide = columns >= 90
832    const kw = keepwarmStatus(Date.now())
833    const cold = coldStatus(Date.now())
834    return band(
835      <Box key="cache" flexDirection="row" columnGap={1}>
836        <Text bold color={color}>{advice.kind === 'warm' ? '●' : advice.kind === 'soon' ? '▲' : advice.kind === 'off' || advice.kind === 'cold' || advice.kind === 'uncached' ? '○' : '✖'}</Text>
837        <Text bold color="cyan">cache</Text>
838        <Text color={color}>{bar(ratio, wide ? 10 : 6)}</Text>
839        <Text bold>{pct1(ratio)}</Text>
840        {!showBreakdown ? null : wide ? (
841          <>
842            <Text color="green">{`read ${fmtTokens(last.read)}`}</Text>
843            <Text color="yellow">{`wrote ${fmtTokens(last.write)}`}</Text>
844            <Text color="cyan">{`new ${fmtTokens(last.fresh)}`}</Text>
845          </>
846        ) : (
847          <Text dimColor>{`${fmtTokens(promptTokens(last))} tok`}</Text>
848        )}
849        {advice.kind !== 'uncached' && advice.kind !== 'off' && (
850          <Text bold color={left > 0 ? lifeColor(left, ttl, policy.warnMs) : 'red'}>{left > 0 ? `⏱ ${fmtCountdown(left)}` : '⏱ 0:00'}</Text>
851        )}
852        {kw && <Text color={kw.stopped ? 'red' : 'magenta'}>{kw.text}</Text>}
853        {cold && <Text color="cyan">{cold}</Text>}
854        <Button key="open" label="details" onPress={() => openPane($)} />
855        <Text dimColor wrap="truncate-end">{`${ttl} · ${advice.text}`}</Text>
856      </Box>,
857    )
858  })
859
860  on('ui.render', { component: 'Pane' }, async ($, e, next) => {
861    if (e.requestId !== PANE) return next(e)
862    const { Box, Text, Button } = $.ui.resolve(e)
863    const width = Math.max(30, e.props.bodyColumns - 1)
864    // HTML collapses runs of spaces and trims a text's ends; a no-break space keeps them
865    const sp = (t: string) => (e.surface === 'terminal' ? t : t.replace(/ /g, ' '))
866    const now = Date.now()
867    const { last, advice, left } = current(policy, now)
868    const kw = keepwarmStatus(now)
869    const all = byTurn(samples)
870    const counting = !!last && advice.kind !== 'uncached' && advice.kind !== 'off'
871    // the countdown goes green, then yellow, then red as the cache runs out
872    const clockColor = counting ? lifeColor(left, ttl, policy.warnMs) : undefined
873    const stateColor = advice.kind === 'expired' || advice.kind === 'miss' ? 'red' : (clockColor ?? COLOR[advice.kind])
874    const hitColor = (pct: number) => (pct >= 80 ? 'green' : pct >= 40 ? 'yellow' : 'red')
875    // solid bars are filled Boxes, not block characters, so HTML draws no seams between cells
876    const solid = (key: string, parts: [number, string | undefined][]) => (
877      <Box key={key} flexDirection="row" height={1} flexShrink={0}>
878        {parts.map(([w, c], i) => (w > 0 ? <Box key={`${key}:${i}`} width={w} height={1} flexShrink={0} backgroundColor={c} /> : null))}
879      </Box>
880    )
881    const cell = (key: string, w: number, text: string, c?: string, bold = false) => (
882      <Box key={key} width={w} flexShrink={0} justifyContent="flex-end">
883        <Text color={c} bold={bold} dimColor={!c}>{sp(text)}</Text>
884      </Box>
885    )
886
887    const barW = Math.max(8, Math.min(width - LABEL_W - 10, 32))
888    const life = lifeRatio(left, ttl)
889    const lifeFilled = Math.round(life * barW)
890    const [sr, sw, sn] = last ? segments(last.read, last.write, last.fresh, barW) : [0, 0, 0]
891    const rows = all.slice(-Math.max(3, (e.viewport?.rows ?? 24) - 30))
892    const icon = advice.kind === 'warm' ? '●' : advice.kind === 'soon' ? '▲' : advice.kind === 'expired' || advice.kind === 'miss' ? '✖' : '○'
893    const kwi = keepwarmInfo(now)
894    const { tokens, cold, warm } = coldPrice()
895    const isColdNow = coldSince(now) !== undefined
896    const price = priceOf(last?.model ?? '')
897    const rate = writeRate(last?.model ?? '')
898    const breakEven = price && rate ? Math.floor(rate / price[0]) : undefined
899    const paid = misses.reduce((a, m) => a + (m.usd ?? 0), 0)
900
901    // a section: a title, then label/value rows under it
902    const section = (key: string, title: string, children: RenderChildren[]) => (
903      <Box key={key} flexDirection="column" marginTop={1}>
904        <Text bold color="cyan">{sp(title)}</Text>
905        {children}
906      </Box>
907    )
908    // one row: the label dim in a fixed column, the value after it
909    const row = (key: string, label: string, ...value: RenderChildren[]) => (
910      <Box key={key} flexDirection="row" columnGap={1}>
911        <Box width={LABEL_W} flexShrink={0}>
912          <Text dimColor>{sp(label)}</Text>
913        </Box>
914        {value}
915      </Box>
916    )
917    const bar = (key: string, ratio: number, c: string) => {
918      const w = Math.max(0, Math.min(barW, Math.round(ratio * barW)))
919      return solid(key, [[w, c], [barW - w, 'gray']])
920    }
921
922    // same as /keepwarm and /keepwarm off; handled in ui.press like close
923    const kwButton = kwi.on
924      ? <Button key="kw-stop" label="stop" onPress={() => {}} />
925      : <Button key="kw-start" label={`start ${DEFAULT_WINDOW_MS / 3600000}h`} onPress={() => {}} />
926
927    return (
928      <Box flexDirection="column">
929        <Box key="title" flexDirection="row" columnGap={1}>
930          <Text bold color="cyan">{sp('⚡ PROMPT CACHE')}</Text>
931          <Text bold color={stateColor}>{sp(`${icon} ${advice.text}`)}</Text>
932        </Box>
933
934        {section('cache', 'CACHE', [
935          sessionStartedAt ? row('c:session', 'session', <Text key="v" bold>{sp(fmtClock(now - sessionStartedAt))}</Text>) : null,
936          row('c:ttl', 'lifetime', <Text key="v" bold>{sp(ttl)}</Text>, <Text key="s" dimColor>{sp(`(${ttlSource})`)}</Text>),
937          row('c:left', 'expires in',
938            counting ? solid('life', [[lifeFilled, clockColor], [barW - lifeFilled, 'gray']]) : null,
939            <Text key="v" bold color={clockColor}>{sp(counting ? (left > 0 ? fmtClock(left) : '0:00') : '--:--')}</Text>),
940          last ? row('c:model', 'model', <Text key="v" bold>{sp(last.model)}</Text>) : null,
941          last ? row('c:prompt', 'prompt', <Text key="v" bold>{sp(`${fmtTokens(promptTokens(last))} tokens`)}</Text>) : null,
942        ])}
943
944        {last ? section('req', 'LAST REQUEST', [
945          row('r:hit', 'hit rate', solid('stack', [[sr, 'green'], [sw, 'yellow'], [sn, 'cyan']]),
946            <Text key="v" bold color={hitColor(Math.round(hitRatio(last) * 100))}>{sp(pct1(hitRatio(last)))}</Text>),
947          row('r:read', 'read', <Text key="v" bold color="green">{sp(`■ ${fmtTokens(last.read)}`)}</Text>, <Text key="s" dimColor>{sp('served by the cache')}</Text>),
948          row('r:wrote', 'wrote', <Text key="v" bold color="yellow">{sp(`■ ${fmtTokens(last.write)}`)}</Text>, <Text key="s" dimColor>{sp('new cache entry')}</Text>),
949          row('r:new', 'new', <Text key="v" bold color="cyan">{sp(`■ ${fmtTokens(last.fresh)}`)}</Text>, <Text key="s" dimColor>{sp('sent uncached')}</Text>),
950        ]) : null}
951
952        {section('kw', 'KEEPWARM', [
953          kwi.stopped
954            ? row('k:state', 'status', <Text key="v" bold color="red">{sp(`stopped: ${kwi.stopped}`)}</Text>, kwButton)
955            : kwi.on
956              ? row('k:state', 'window', bar('kwbar', kwi.left / kwi.window, 'magenta'),
957                  <Text key="v" bold color="magenta">{sp(`${fmtDuration(kwi.left)} left`)}</Text>, kwButton)
958              : row('k:state', 'status', <Text key="v" bold>{sp(kwi.always ? 'off until next session (always)' : 'off')}</Text>, kwButton),
959          kwi.on ? row('k:next', 'next ping', <Text key="v" bold>{sp(kwi.next ?? '')}</Text>, <Text key="s" dimColor>{sp(`every ${fmtDuration(kwi.every)}`)}</Text>) : null,
960          kwi.lastPing ? row('k:last', 'last ping', <Text key="v" bold>{sp(`${fmtTokens(kwi.lastPing.read)} read · ${fmtUsd(kwi.lastPing.usd)}`)}</Text>) : null,
961          breakEven !== undefined ? row('k:be', 'break-even', <Text key="v" bold>{sp(`${breakEven} pings`)}</Text>,
962            <Text key="s" dimColor>{sp(`= one cold write, ~${fmtDuration(breakEven * kwi.every)} idle`)}</Text>) : null,
963        ])}
964
965        {section('cost', 'COST', [
966          last && cold != null ? row('$:cold', 'cold write', bar('coldbar', 1, isColdNow ? 'red' : 'yellow'),
967            <Text key="v" bold color={isColdNow ? 'red' : 'yellow'}>{sp(fmtUsd(cold))}</Text>,
968            <Text key="s" dimColor>{sp(isColdNow ? `cold now: the next message re-writes ${fmtTokens(tokens)}` : `to re-write ${fmtTokens(tokens)}`)}</Text>) : null,
969          last && warm != null && cold ? row('$:warm', 'warm turn', bar('warmbar', Math.max(warm / cold, 1 / barW), 'green'),
970            <Text key="v" bold color="green">{sp(fmtUsd(warm))}</Text>) : null,
971          row('$:guard', 'guard', <Text key="v" bold>{sp(guard === 'refuse' ? 'refuse once' : guard === 'warn' ? 'warn only' : 'off')}</Text>,
972            <Text key="s" dimColor>{sp('on a cold cache of 50k+ tokens')}</Text>),
973          pctPerUsd !== undefined && cold != null && warm != null
974            ? row('$:plan', '5h window', <Text key="v" bold color="magenta">{sp(`≈ ${(cold * pctPerUsd).toFixed(1)}% cold · ${(warm * pctPerUsd).toFixed(1)}% warm`)}</Text>,
975                <Text key="s" dimColor>{sp(`🧪 ${pctPerUsd.toFixed(2)}% per $ · ${sessionsText()}`)}</Text>)
976            : anchor
977              ? row('$:plan', '5h window', <Text key="v" dimColor>{sp(`🧪 calibrating… · ${sessionsText()}`)}</Text>)
978              : null,
979          row('$:session', 'this session', <Text key="v" bold color={misses.length ? 'red' : undefined}>{sp(`${misses.length} cold write${misses.length === 1 ? '' : 's'} · ${fmtUsd(paid)}`)}</Text>),
980        ])}
981
982        {section('turns', 'TURNS', [
983          rows.length === 0 ? <Text key="none" dimColor>{sp('no requests yet')}</Text> : <Box key="head" flexDirection="row" columnGap={1}>
984            {cell('h:turn', 4, 'turn', 'cyan', true)}
985            {cell('h:steps', 5, 'steps', 'cyan', true)}
986            {cell('h:read', 6, 'read', 'green', true)}
987            {cell('h:wrote', 6, 'wrote', 'yellow', true)}
988            {cell('h:new', 5, 'new', 'cyan', true)}
989            {cell('h:hit', 6, 'hit', 'magenta', true)}
990            {cell('h:usd', 6, 'cost', 'cyan', true)}
991            {pctPerUsd !== undefined ? cell('h:pct', 7, '5h% 🧪', 'magenta', true) : null}
992          </Box>,
993          ...rows.map((r, i) => {
994            const n = all.length - rows.length + i + 1
995            const pct = Math.round(rowRatio(r) * 100)
996            const usd = turnUsd(r.turnId)
997            return (
998              <Box key={`t:${r.turnId}`} flexDirection="row" columnGap={1}>
999                {r.turnId.startsWith('keepwarm-') ? cell(`c:turn:${r.turnId}`, 4, '♨')
1000                  : misses.some(m => m.turnId === r.turnId) ? cell(`c:turn:${r.turnId}`, 4, `❄${n}`, 'cyan', true)
1001                  : cell(`c:turn:${r.turnId}`, 4, String(n))}
1002                {cell(`c:steps:${r.turnId}`, 5, String(r.steps))}
1003                {cell(`c:read:${r.turnId}`, 6, fmtTokens(r.read), 'green')}
1004                {cell(`c:wrote:${r.turnId}`, 6, fmtTokens(r.write), 'yellow')}
1005                {cell(`c:new:${r.turnId}`, 5, fmtTokens(r.fresh), 'cyan')}
1006                {cell(`c:hit:${r.turnId}`, 6, pct1(rowRatio(r)), hitColor(pct), true)}
1007                {cell(`c:usd:${r.turnId}`, 6, fmtUsd(usd))}
1008                {pctPerUsd !== undefined ? cell(`c:pct:${r.turnId}`, 7, usd == null ? 'n/a' : `${(usd * pctPerUsd).toFixed(1)}%`, 'magenta') : null}
1009              </Box>
1010            )
1011          }),
1012        ])}
1013
1014        <Box key="foot" marginTop={1}>
1015          <Button key="close" label="close" onPress={() => {}} />
1016        </Box>
1017      </Box>
1018    )
1019  })
1020
1021  // a message to a cold cache: refused once with its price (guard: refuse), or sent with the price logged (warn)
1022  on('prompt.submit', async ($, e, next) => {
1023    if (e.origin.kind === 'plugin' || typeof e.text !== 'string' || e.text.trimStart().startsWith('/')) return next(e)
1024    const now = Date.now()
1025    const last = samples[samples.length - 1]
1026    if (!last) return next(e)
1027    const verdict = guardVerdict({ mode: guard, coldAt: coldSince(now), tokens: promptTokens(last), startedAt: last.startedAt, ackedAt })
1028    if (verdict === 'pass') return next(e)
1029    if (verdict === 'warn') {
1030      $.ui.log(`prompt-cache-control: ${guardText(now)} Sending anyway; keepwarm will hold the cache for ${fmtDuration(AUTO_WARM_MS)} once it lands.`)
1031      coldWritePending = true
1032      return next(e)
1033    }
1034    if (verdict === 'resend') {
1035      ackedAt = 0
1036      coldWritePending = true
1037      return next(e)
1038    }
1039    ackedAt = last.startedAt
1040    return { drop: `prompt-cache-control: ${guardText(now)} Send it again to pay it, and keepwarm will then hold the cache for ${fmtDuration(AUTO_WARM_MS)}. Or /clear and start from a note.` }
1041  })
1042
1043  on('session.compact', async ($, e, next) => {
1044    const r = await next(e)
1045    if (!e.agentId) {
1046      compacted = true
1047      await arm($)
1048    }
1049    return r
1050  })
1051
1052  on('turn.complete', async ($, e, next) => {
1053    const r = await next(e)
1054    if (e.agentId) return r
1055    compacted = false
1056    if (e.usage) {
1057      const usd = usdOf(e.usage, e.usage.model || (samples[samples.length - 1]?.model ?? ''))
1058      if (usd != null) spentUsd += usd
1059    }
1060    await calibrate($).catch(() => undefined)
1061    const paid = coldWriteOf(e.turnId)
1062    coldWritePending = false
1063    if (paid) {
1064      const rate = writeRate(paid.model)
1065      const usd = rate == null ? null : (paid.write * rate) / 1e6
1066      misses.push({ at: paid.startedAt, turnId: paid.turnId, tokens: paid.write, usd })
1067      // a cold write was just paid: keep it from being paid again today
1068      if (deadline < Date.now() + AUTO_WARM_MS) {
1069        await startWindow($, AUTO_WARM_MS, every)
1070        $.ui.log(`prompt-cache-control: cold write of ${fmtTokens(paid.write)} tokens paid (${fmtUsd(usd)}). Keeping the cache warm for ${fmtDuration(AUTO_WARM_MS)} so it is not paid again today; /keepwarm off to stop.`)
1071      }
1072    }
1073    await arm($)
1074    return r
1075  })
1076
1077  on('command.run', { command: 'keepwarm' }, async ($, e) => {
1078    const words = String(e.args ?? '').trim().split(/\s+/).filter(Boolean)
1079    if (words[0] === 'off') {
1080      const wasAlways = always
1081      await stop($, null, true)
1082      return { text: wasAlways ? 'keepwarm is off, and no longer arms itself at session start' : 'keepwarm is off' }
1083    }
1084    if (words[0] === 'always') {
1085      always = true
1086      await $.store.set(KEY_ALWAYS, true)
1087      await startWindow($, DEFAULT_WINDOW_MS, 0)
1088      return { text: `keepwarm always on: every session starts with a ${fmtDuration(DEFAULT_WINDOW_MS)} window; /keepwarm off turns it off for good` }
1089    }
1090    if (!words.length) {
1091      await startWindow($, DEFAULT_WINDOW_MS, 0)
1092      return { text: armedText(DEFAULT_WINDOW_MS) }
1093    }
1094    if (words[0] === 'status') return { text: statusLine(Date.now()) ?? 'keepwarm is off' }
1095    const windowMs = parseDuration(words[0])
1096    if (windowMs == null) return { text: 'keepwarm takes a window such as 6h or 90m, or always, off, or status' }
1097    let everyMs = 0
1098    if (words[1] === 'every') {
1099      const p = parseDuration(words[2] ?? '')
1100      if (p == null || p < MIN_PING_MS) return { text: 'every takes a period of at least 1m' }
1101      everyMs = p
1102    }
1103    await startWindow($, windowMs, everyMs)
1104    return { text: armedText(windowMs) }
1105  })
1106}
1107
hooks/cache.ts 386 lines
1/**
2 * cache.ts — the pure half of prompt-cache-control: no `$`, no engine.
3 *
4 * What it models, from Anthropic's prompt-caching documentation:
5 *   - the cache lives 5 minutes by default, 1 hour when asked for; a read
6 *     refreshes the entry at no extra cost, and the lifetime is measured from
7 *     the START of the request that wrote or read it
8 *   - a request's prompt is `input_tokens` (uncached remainder) +
9 *     `cache_read_input_tokens` + `cache_creation_input_tokens`
10 *   - writes cost 1.25x base input for 5m and 2x for 1h; reads about 0.1x
11 *     (less on some models), so an expired cache on a large context is the
12 *     expensive moment
13 *   - a prefix change (model, effort/thinking settings, tool set, system
14 *     prompt) makes the next request write instead of read
15 *
16 * Claude Code's own switches (read from the environment):
17 *   ENABLE_PROMPT_CACHING_1H=1   ask for the 1-hour TTL
18 *   FORCE_PROMPT_CACHING_5M=1    force the 5-minute TTL, beating the above
19 *   DISABLE_PROMPT_CACHING=1     no caching; DISABLE_PROMPT_CACHING_{HAIKU,SONNET,OPUS}
20 *                                 turn it off for that model family only
21 */
22
23export type Ttl = '5m' | '1h'
24
25export type CacheEnv = {
26  enable1h?: string
27  force5m?: string
28  /** CLAUDE_CODE_PROMPT_CACHE_TTL: "5m" or "1h" for the main conversation */
29  ttlVar?: string
30  disableAll?: string
31  disableHaiku?: string
32  disableSonnet?: string
33  disableOpus?: string
34}
35
36/** One main-loop request, as the API reported it. */
37export type Sample = {
38  turnId: string
39  index: number
40  model: string
41  /** ms since the epoch when the request started: the cache's lifetime is counted from here */
42  startedAt: number
43  read: number
44  write: number
45  fresh: number
46  output: number
47}
48
49export type AdviceKind = 'off' | 'cold' | 'uncached' | 'warm' | 'soon' | 'expired' | 'miss'
50
51export type Advice = {
52  kind: AdviceKind
53  /** one sentence for the band */
54  text: string
55}
56
57export type Policy = {
58  ttl: Ttl
59  warnMs: number
60  compactAtTokens: number
61}
62
63export const isOn = (v: string | undefined) => v === '1' || v?.toLowerCase() === 'true'
64
65/** What the account is billed as, as far as the mod can tell. */
66export type Account = 'subscription' | 'credits' | 'other'
67
68export type TtlChoice = { ttl: Ttl; source: string }
69
70const asTtl = (v: unknown): Ttl | undefined => (v === '5m' || v === '1h' ? v : undefined)
71
72/**
73 * Which lifetime Claude Code asks for on the main conversation, in the order
74 * its documentation gives (https://code.claude.com/docs/en/prompt-caching,
75 * "Choose the TTL yourself"), after the mod's own `ttl` option:
76 *
77 *   FORCE_PROMPT_CACHING_5M, CLAUDE_CODE_PROMPT_CACHE_TTL, the promptCacheTtl
78 *   setting, ENABLE_PROMPT_CACHING_1H, then the default of the account: one
79 *   hour on a Claude subscription within its plan usage, five minutes on usage
80 *   credits, an API key or a cloud provider.
81 */
82export function decideTtl(option: unknown, env: CacheEnv, setting?: unknown, account?: Account): TtlChoice {
83  const pinned = asTtl(option)
84  if (pinned) return { ttl: pinned, source: 'the ttl option' }
85  if (isOn(env.force5m)) return { ttl: '5m', source: 'FORCE_PROMPT_CACHING_5M' }
86  const fromVar = asTtl(env.ttlVar)
87  if (fromVar) return { ttl: fromVar, source: 'CLAUDE_CODE_PROMPT_CACHE_TTL' }
88  const fromSetting = asTtl(setting)
89  if (fromSetting) return { ttl: fromSetting, source: 'the promptCacheTtl setting' }
90  if (isOn(env.enable1h)) return { ttl: '1h', source: 'ENABLE_PROMPT_CACHING_1H' }
91  if (account === 'subscription') return { ttl: '1h', source: 'Claude subscription default' }
92  if (account === 'credits') return { ttl: '5m', source: 'usage credits default' }
93  return { ttl: '5m', source: 'default' }
94}
95
96export const resolveTtl = (option: unknown, env: CacheEnv, setting?: unknown, account?: Account): Ttl =>
97  decideTtl(option, env, setting, account).ttl
98
99/**
100 * The account, from the rate-limit windows the last response reported: a
101 * five-hour or seven-day window means a Claude subscription, and one that is
102 * full means the next requests draw on usage credits. No window (an API key,
103 * a cloud provider, or no response yet) says nothing.
104 */
105export function accountOf(windows: readonly { kind: string; percentUsed: number }[]): Account {
106  const plan = windows.filter(w => w.kind === 'five_hour' || w.kind === 'seven_day')
107  if (plan.length === 0) return 'other'
108  return plan.some(w => w.percentUsed >= 100) ? 'credits' : 'subscription'
109}
110
111export function ttlMs(ttl: Ttl): number {
112  return ttl === '1h' ? 3_600_000 : 300_000
113}
114
115/** Caching switched off for this model by the environment. */
116export function isCachingDisabled(model: string, env: CacheEnv): boolean {
117  if (isOn(env.disableAll)) return true
118  const name = model.toLowerCase()
119  if (name.includes('haiku')) return isOn(env.disableHaiku)
120  if (name.includes('sonnet')) return isOn(env.disableSonnet)
121  if (name.includes('opus')) return isOn(env.disableOpus)
122  return false
123}
124
125export const promptTokens = (s: Sample) => s.read + s.write + s.fresh
126
127/** Share of the prompt the cache served, 0 to 1; 0 for an empty prompt. */
128export function hitRatio(s: Sample): number {
129  const total = promptTokens(s)
130  return total === 0 ? 0 : s.read / total
131}
132
133/** When the cache entry the sample touched lapses, ms since the epoch. */
134export const expiresAt = (s: Sample, ttl: Ttl) => s.startedAt + ttlMs(ttl)
135
136/** Zero for a request that read and wrote nothing: it created or refreshed no entry, so there is nothing to count down. */
137export function remainingMs(s: Sample, ttl: Ttl, now: number): number {
138  if (s.read + s.write === 0) return 0
139  return Math.max(0, expiresAt(s, ttl) - now)
140}
141
142/**
143 * Why a request that should have read the cache wrote it instead; undefined
144 * when it did not miss. A prompt that shrank is a /compact or /clear, not a
145 * miss, and the first request of a session has nothing to read.
146 */
147export function missReason(prev: Sample | undefined, cur: Sample, ttl: Ttl): string | undefined {
148  if (!prev) return undefined
149  const before = promptTokens(prev)
150  if (before === 0 || promptTokens(cur) < before * 0.7) return undefined
151  if (cur.read >= before * 0.5 || cur.write === 0) return undefined
152  if (cur.model !== prev.model) return `model changed (${prev.model} to ${cur.model})`
153  if (cur.startedAt - prev.startedAt > ttlMs(ttl)) return `the ${ttl} cache had lapsed`
154  return 'the prompt prefix changed (effort, tools, system prompt or CLAUDE.md)'
155}
156
157export function advise(last: Sample | undefined, prev: Sample | undefined, policy: Policy, now: number, disabled: boolean): Advice {
158  if (disabled) return { kind: 'off', text: 'prompt caching is off for this model (DISABLE_PROMPT_CACHING*)' }
159  if (!last) return { kind: 'cold', text: 'no request yet: the first one writes the cache' }
160  if (last.read + last.write === 0) {
161    return { kind: 'uncached', text: 'this request was not cached (prompt under the model minimum, or caching off)' }
162  }
163  const miss = missReason(prev, last, policy.ttl)
164  const left = remainingMs(last, policy.ttl, now)
165  const size = promptTokens(last)
166  if (left <= 0) {
167    const big = size >= policy.compactAtTokens
168    return {
169      kind: 'expired',
170      text: big
171        ? `expired: the next message rewrites ${fmtTokens(size)} tokens. /compact first, or /clear if the task is done`
172        : `expired: only ${fmtTokens(size)} tokens to rebuild, just keep going`,
173    }
174  }
175  if (left <= policy.warnMs) {
176    return { kind: 'soon', text: 'expires soon: any message refreshes it for free' }
177  }
178  if (miss) return { kind: 'miss', text: `cache missed: ${miss}` }
179  return { kind: 'warm', text: 'warm: keep going' }
180}
181
182export function fmtTokens(n: number): string {
183  if (n < 1000) return String(n)
184  if (n < 100_000) return `${(n / 1000).toFixed(1).replace(/\.0$/, '')}k`
185  if (n < 1_000_000) return `${Math.round(n / 1000)}k`
186  return `${(n / 1_000_000).toFixed(1).replace(/\.0$/, '')}M`
187}
188
189/** m:ss, or h:mm:ss from an hour up. */
190export function fmtClock(ms: number): string {
191  const total = Math.max(0, Math.ceil(ms / 1000))
192  const h = Math.floor(total / 3600)
193  const m = Math.floor((total % 3600) / 60)
194  const s = total % 60
195  const pad = (n: number) => String(n).padStart(2, '0')
196  return h > 0 ? `${h}:${pad(m)}:${pad(s)}` : `${m}:${pad(s)}`
197}
198
199/** the band's countdown: whole minutes (rounded down) from 10 minutes up, so it does not tick; m:ss below */
200export function fmtCountdown(ms: number): string {
201  if (ms < 10 * 60_000) return fmtClock(ms)
202  const total = Math.floor(ms / 60_000)
203  const h = Math.floor(total / 60)
204  const m = total % 60
205  return h > 0 ? `${h}h${String(m).padStart(2, '0')}m` : `${m}m`
206}
207
208/** the 5-hour plan window under the prompt: `5h 42% · reset 2h13m`; undefined without a reading or once it has reset */
209export function fmtLimit(w: { percentUsed: number; resetsAt?: string } | undefined, now: number): string | undefined {
210  if (!w) return undefined
211  const at = w.resetsAt ? Date.parse(w.resetsAt) : NaN
212  if (Number.isNaN(at)) return `5h ${w.percentUsed}%`
213  if (at <= now) return undefined
214  const total = Math.ceil((at - now) / 60_000)
215  const h = Math.floor(total / 60)
216  const m = total % 60
217  return `5h ${w.percentUsed}% · reset ${h > 0 ? `${h}h${String(m).padStart(2, '0')}m` : `${m}m`}`
218}
219
220export function bar(ratio: number, width: number): string {
221  const filled = Math.round(Math.min(1, Math.max(0, ratio)) * width)
222  return '█'.repeat(filled) + '░'.repeat(width - filled)
223}
224
225export type TurnRow = {
226  turnId: string
227  steps: number
228  read: number
229  write: number
230  fresh: number
231  output: number
232}
233
234/** Samples grouped by turn, oldest first, each turn's requests summed. */
235export function byTurn(samples: readonly Sample[]): TurnRow[] {
236  const rows: TurnRow[] = []
237  for (const s of samples) {
238    let row = rows[rows.length - 1]
239    if (!row || row.turnId !== s.turnId) {
240      row = { turnId: s.turnId, steps: 0, read: 0, write: 0, fresh: 0, output: 0 }
241      rows.push(row)
242    }
243    row.steps += 1
244    row.read += s.read
245    row.write += s.write
246    row.fresh += s.fresh
247    row.output += s.output
248  }
249  return rows
250}
251
252export const rowRatio = (r: TurnRow) => {
253  const total = r.read + r.write + r.fresh
254  return total === 0 ? 0 : r.read / total
255}
256
257export function positive(v: unknown, fallback: number): number {
258  return typeof v === 'number' && Number.isFinite(v) && v > 0 ? v : fallback
259}
260
261/** Share of the cache lifetime left, 0 to 1. */
262export function lifeRatio(leftMs: number, ttl: Ttl): number {
263  return Math.min(1, Math.max(0, leftMs / ttlMs(ttl)))
264}
265
266/**
267 * Widths of the three stacked-bar segments (read, wrote, new) over `width`
268 * cells: proportional, each non-empty part at least one cell, summing to width.
269 */
270export function segments(read: number, write: number, fresh: number, width: number): [number, number, number] {
271  const total = read + write + fresh
272  if (total === 0 || width <= 0) return [0, 0, 0]
273  const parts = [read, write, fresh]
274  const cells = parts.map(p => (p > 0 ? Math.max(1, Math.round((p / total) * width)) : 0))
275  let over = cells.reduce((a, b) => a + b, 0) - width
276  while (over !== 0) {
277    const i = over > 0 ? cells.indexOf(Math.max(...cells)) : parts.indexOf(Math.max(...parts))
278    cells[i] += over > 0 ? -1 : 1
279    over += over > 0 ? -1 : 1
280  }
281  return [cells[0], cells[1], cells[2]]
282}
283
284/** Seconds left at which a toast counts down after the one at the warning threshold. */
285export const COUNTDOWN_MARKS = [10, 3, 2, 1]
286
287/**
288 * The toast mark to fire now, or undefined. `level` is the mark last fired for
289 * this cache entry (Infinity before any); a late tick skips straight to the
290 * newest mark crossed, so a stalled clock never replays old ones.
291 */
292export function nextToastMark(secsLeft: number, warnSecs: number, level: number): number | undefined {
293  const marks = [warnSecs, ...COUNTDOWN_MARKS].filter(m => m <= warnSecs)
294  const due = marks.filter(m => secsLeft <= m && m < level)
295  return due.length ? Math.min(...due) : undefined
296}
297
298export type LifeColor = 'green' | 'yellow' | 'red'
299
300/** Countdown colour: green while there is plenty, yellow below 40% of the lifetime, red from the warning threshold down. */
301export function lifeColor(leftMs: number, ttl: Ttl, warnMs: number): LifeColor {
302  if (leftMs <= warnMs) return 'red'
303  return leftMs / ttlMs(ttl) <= 0.4 ? 'yellow' : 'green'
304}
305
306// requests are timed from their start, so a little slack keeps a hit that
307// landed just inside the lifetime from reading as proof of the longer one
308const SLACK_MS = 10_000
309
310/**
311 * What the traffic says about the cache lifetime, given the request before and
312 * `known`, what earlier requests already showed.
313 *
314 *   - a hit (the cache served at least half of the previous prompt) more than
315 *     5 minutes after the previous request began proves the 1-hour lifetime,
316 *     and nothing later undoes it: a miss afterwards is more likely a changed
317 *     prefix than a lapse
318 *   - a miss with the same model and a prompt that did not shrink, 5 minutes to
319 *     an hour after the previous request, says the entry lapsed: 5 minutes
320 *     (weaker: a changed prefix looks the same, so a later hit overrules it)
321 *
322 * Needed because the API names the TTL of a write (`cache_creation.ephemeral_*`)
323 * but Claude Code's mod API passes on only the four token counts.
324 */
325export function observeTtl(prev: Sample | undefined, cur: Sample, known: Ttl | undefined): Ttl | undefined {
326  if (!prev || prev.read + prev.write === 0 || cur.model !== prev.model) return known
327  const gap = cur.startedAt - prev.startedAt
328  const before = promptTokens(prev)
329  if (gap <= ttlMs('5m') + SLACK_MS) return known
330  if (cur.read >= before * 0.5) return '1h'
331  if (known === '1h') return known
332  const lapsed = cur.write > 0 && promptTokens(cur) >= before * 0.7 && gap < ttlMs('1h') + SLACK_MS
333  return lapsed ? '5m' : known
334}
335
336export type GuardMode = 'refuse' | 'warn' | 'off'
337
338/**
339 * What to do with a message (cold-cache guard, from cache-tax): `coldAt` is when
340 * the last request's cache lapsed (undefined while warm), `startedAt` that
341 * request's start, `ackedAt` the request a previous refusal was for.
342 */
343export function guardVerdict(g: { mode: GuardMode; coldAt: number | undefined; tokens: number; startedAt: number; ackedAt: number }): 'pass' | 'warn' | 'drop' | 'resend' {
344  if (g.mode === 'off' || g.coldAt === undefined || g.tokens < 50_000) return 'pass'
345  if (g.mode === 'warn') return 'warn'
346  return g.ackedAt === g.startedAt ? 'resend' : 'drop'
347}
348
349/** a request re-wrote the prefix: the guard let a cold send through, or it wrote half a 20k+ previous prompt or more (from cache-tax) */
350export const isColdWrite = (prevTokens: number, write: number, guardLetThrough: boolean) =>
351  guardLetThrough || (prevTokens > 20_000 && write >= 0.5 * prevTokens)
352
353/**
354 * Experimental: percent of a plan window per list-price dollar, from how far
355 * the window moved (`pct0` to `pct`) while `usd0` to `usd` was spent. Undefined
356 * until the window moved half a percent, as it reports one decimal.
357 */
358export function calibRate(pct0: number, usd0: number, pct: number, usd: number): number | undefined {
359  const dp = pct - pct0
360  const du = usd - usd0
361  return dp >= 0.5 && du > 0 ? dp / du : undefined
362}
363
364/** one session's spend in a plan window, as each session writes it to the shared store */
365export type Spend = { resetsAt: string; usd: number; at: number }
366
367/**
368 * Every session's spend in the window that resets at `resetsAt`, summed; how
369 * many of them wrote within `activeMs`; and the keys left from other windows.
370 */
371export function windowSpend(entries: [string, unknown][], resetsAt: string, now: number, activeMs = 10 * 60_000) {
372  let usd = 0
373  let active = 0
374  const stale: string[] = []
375  for (const [key, v] of entries) {
376    const s = v as Partial<Spend> | undefined
377    if (!s || s.resetsAt !== resetsAt || typeof s.usd !== 'number') {
378      stale.push(key)
379      continue
380    }
381    usd += s.usd
382    if (typeof s.at === 'number' && now - s.at <= activeMs) active++
383  }
384  return { usd, active, stale }
385}
386