SLOPSHOPPER

cache-guard

Keeps the prompt cache warm while you are away and asks before a prompt that would re-cache a large conversation.

newcommandpromptmodelprocesstimer
v0.1.0MITupdated 2026-10-08justmytwospence/claude-cache-guard
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · cache-guard
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /cache-guard ⎿ cache-guard: No cached request in this session yet (or it was just compacted). ⎿ cache-guard: Keep-warm: idle (no cached request yet). ⎿ cache-guard: Warning: on, from $0.50 at API prices. ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts
README

claude-cache-guard

A Claude Code mod that keeps the prompt cache warm while you are away and asks before a prompt that would re-cache a large conversation. The Claude Code port of pi-cache-guard, with the same core (hooks/core.ts) and settings file.

What it does

Keeps the cache warm. A cached prefix lives for its TTL from the start of the last request that used it: 1 hour for the main conversation on a Claude subscription within plan usage, 5 minutes with an API key, a cloud provider or usage credits. The mod reads the tier off the session's own cache writes in the transcript. At 90% of the TTL with no request since, it calls $.model.fork, which re-sends the main thread's last request (same model, system prompt, tools and messages) with one short question after it, tools denied and its own tail never cached. The API serves the prefix from the cache, which restarts its clock.

Each refresh has to pay for itself by Pi's rule:

  • When: the expected saving must be at least $0.05. While you are idle that saving is 15% x miss cost - refresh cost; during a long tool run it is the whole miss cost.
  • Size: on the 1h tier of Opus 5.5 that works out to about 52k tokens of context or more.
  • For how long: refreshes stop 2 hours (1h tier) or 30 minutes (5m tier) after the last real request.
  • On a miss: a refresh that reads less than 80% of the prefix from the cache stops the keep-warm until the next turn.
  • In the transcript: each refresh logs a line such as kept the prompt cache warm (601k tokens read, ~$0.12 at API prices). The model never sees these lines.

Warns. A prompt you type while idle onto an expired cache asks first when the rewrite would cost at least warn.minCost (default $0.50 at API prices):

Prompt cache miss. The prompt cache expired 10m ago: this prompt re-caches 601k tokens (~$4.69 at API prices). What now?
1. Keep the prompt
2. New conversation (~$0)
3. Compact first (~$2.40)
4. Send anyway (~$4.81)

The options:

  • Keep the prompt. The default, so Enter or Esc spends nothing. It drops the submission and puts the text back in the prompt box.
  • New conversation. Runs /clear and sends the prompt as the new conversation's first.
  • Compact first. Asks what the summary should keep: the default summary, a summary focused on the held prompt (keep what it needs, drop the rest), or your own guidance typed under "Type something". It then compacts and sends the prompt onto the summary. If compaction is skipped or fails, the prompt goes back in the box.
  • Send anyway.

The headline cost is what the miss adds over a cache hit; the costs in the options are each path's total. Images attached to a held prompt are not resent by the New conversation and Compact paths. /cache-guard off stops asking in the session. Resumed and forked sessions are judged from the transcript, or from the SessionStart hook input (seconds_since_last_response, context_tokens) when the transcript is not written yet. Prompts from loops, peers, notifications, the SDK and -p are never held, and neither are slash commands.

Status. Claude Code reports the cache to the status line itself (prompt_cache.warm, expires_at, ttl, recache_tokens_if_cold). A mod cannot ship a status line, so the countdown belongs in your status line script, e.g. jq '.prompt_cache.expires_at'. The mod draws none.

Tells herdr. Inside a herdr pane, the mod sets the pane token cache with herdr pane report-metadata --source cache-guard. The token reads cold 601k once the cache has expired and the next prompt would re-cache at least the warning threshold. It is cleared while the cache is warm or small, and when the session ends. An expiry timer flips it on time. herdr's agents sidebar shows it with { token = "$cache" } in a [ui.sidebar.agents] row. Set "herdr": { "enabled": false } to turn it off.

Claude Code's own guards stay in place: /model and /effort ask while the cache is warm, and /resume offers to resume from a summary after a long break.

Commands

  • /cache-guard or /cache-guard status: the last request, TTL and its source, time left or how long since it expired, and the keep-warm state.
  • /cache-guard warm: a refresh now, whatever the schedule and the expected saving say. Use it to bridge a break, or to check that refreshes hit.
  • /cache-guard on, /cache-guard off: for this session.

Settings

~/.config/agents/cache-guard.json (shared with the other ports), then ~/.claude/cache-guard.json, then the project's .agents/cache-guard.json and .claude/cache-guard.json. Later files win; objects merge.

{
  "enabled": true,
  "warn": { "enabled": true, "minCost": 0.5 },
  "warm": { "enabled": true, "continuationProbability": 0.15, "minSavings": 0.05,
            "idleMinutes": { "5m": 30, "1h": 120 },
            "prompt": "Cache keep-alive. Do not use tools or think. Reply with exactly: ok" }
}

Install

It is a plugin directory with a hooks module. Load it with claude --plugin-dir <checkout>, or list it in CLAUDE_CODE_PLUGIN_DIRS. It needs Claude Code 2.1.287 or newer and was tested with 2.1.289.

Limits

  • First turn: a fork right after a session's first response reads only the system prompt and tools from the cache, not the messages. This is likely because the fork merges into the conversation's only user message. From the second response on, a refresh reads the whole prefix. Measured: 42k of 42k tokens, about $0.01 on Sonnet 5.5.
  • Plan usage: refreshes count against your plan like any other request. On the 1h tier they are rare: at most two per break with the defaults.
  • The clock: it is the API's guaranteed minimum, measured from each request's start; entries are deleted soon after, not exactly then. Changes to the system prompt or tools, model switches and compaction also invalidate the cache. Claude Code's prompt_cache.last_miss_cause reports those; this mod only judges time.
  • Prices: list prices of the Claude models, by model-id prefix (hooks/core.ts). Unknown models are priced as Opus 5.5.

Development

claude plugin validate .   # static analysis of the hooks module
claude plugin test         # tests/, no session or network
tsc -p .                   # after one load, which writes .claude-plugin/types/
Source 4 files
hooks/register.ts 482 lines
1// cache-guard: keeps Claude Code's prompt cache warm while you are away, and asks before a prompt
2// that would re-cache a large conversation.
3//
4// The clock: every main-loop model request (turn.step) reads or writes the cache, and the entry
5// lives for its TTL from that request's start. The TTL is the tier the session's own cache writes
6// use, read off the transcript (1h on a subscription within plan usage, 5m otherwise).
7//
8// Keep-warm: at 90% of the TTL with no request since, `$.model.fork` re-sends the main thread's
9// last request (same model, system prompt, tools and messages) with one short question after it,
10// tools denied and its own tail never cached, so the API serves the prefix from the cache and
11// restarts its clock. Pi's rule decides each refresh (expected saving at least $0.05, with a 15%
12// chance you come back before it lapses while idle), and refreshing stops 30 minutes (5m tier) or
13// 2 hours (1h tier) after the last real request. A refresh that misses stops it.
14//
15// Warning: a prompt typed while idle onto an expired cache, when the rewrite would cost at least
16// warn.minCost at API prices, asks first; keeping it puts the text back in the prompt box.
17//
18// The status line shows the cache itself (Claude Code's prompt_cache input), so this mod draws none.
19//
20// The host reads on(...) and $.noun.method(...) from source, so calls are spelled in full and
21// helpers that take $ are top-level functions in this file.
22import type { EngineInterface, On } from 'claude-code'
23
24import {
25  DEFAULT_SETTINGS,
26  NAME,
27  ONE_HOUR,
28  FIVE_MINUTES,
29  HERDR_SOURCE,
30  HERDR_TOKEN,
31  type Settings,
32  claudePrice,
33  decideWarm,
34  describeMiss,
35  formatClock,
36  formatCost,
37  formatDuration,
38  formatTokens,
39  choiceCosts,
40  compactionFocus,
41  herdrCacheValue,
42  mergeSettings,
43  missCost,
44  warmDeadline,
45  warmDelayMs,
46  worthWarning,
47} from './core'
48import { settingsFiles } from './settings'
49import { parseTail, transcriptPath } from './transcript'
50
51const COMMAND = NAME
52const KEEP = 'Keep the prompt'
53const FRESH = 'New conversation'
54const COMPACT = 'Compact first'
55const SEND = 'Send anyway'
56const SUMMARY_DEFAULT = 'Default summary'
57const SUMMARY_FOCUS = 'Focus on this prompt'
58/** How much of a transcript's end to read: enough for the last few responses. */
59const TAIL_BYTES = 400_000
60const TAIL_SCRIPT = 'f=$1; [ -f "$f" ] || f=$(ls "$2"/*/"$3" 2>/dev/null | head -n 1); [ -n "$f" ] && tail -c "$4" "$f"'
61
62interface Clock {
63  /** Start of the last request that read or wrote the cache (a real one or a refresh). */
64  lastAt: number
65  /** Start of the last real request. */
66  lastRealAt: number
67  /** The last request's prompt: what a refresh should read from the cache. */
68  promptTokens: number
69  /** Prompt plus reply: what the next real request re-sends. */
70  tokens: number
71  model: string
72}
73
74interface State {
75  settings: Settings
76  sessionOn: boolean
77  interactive: boolean
78  clock?: Clock
79  /** The transcript file, from the settings-hook SessionStart input. */
80  transcript?: string
81  ttl: { ms: number; source: string }
82  timer?: { cancel: () => void }
83  stepsInFlight: number
84  busy: boolean
85  warming: boolean
86  warms: number
87  /** Why no refresh is scheduled, for /cache-guard. */
88  stopped: string
89  /** Fires when the cache expires, to tell herdr. */
90  expiry?: { cancel: () => void }
91  /** The `cache` token last reported to herdr; null before the first report. */
92  herdrLast: string | undefined | null
93}
94
95export function register(on: On): void {
96  const s: State = {
97    settings: DEFAULT_SETTINGS,
98    sessionOn: true,
99    interactive: false,
100    ttl: { ms: FIVE_MINUTES, source: 'default' },
101    stepsInFlight: 0,
102    busy: false,
103    warming: false,
104    warms: 0,
105    stopped: 'waiting for the first response',
106    herdrLast: null,
107  }
108
109  on('session.start', async ($, e, next) => {
110    const result = await next(e)
111    s.interactive = e.isInteractive
112    s.settings = await loadSettings($)
113    await $.command.register({
114      name: COMMAND,
115      description: 'cache-guard: prompt cache status, warm (refresh now), on, or off (this session)',
116      argumentHint: '[status|warm|on|off]',
117      immediate: true,
118    })
119    // A resumed session: its last response and TTL come from the transcript.
120    await readTranscript($, s, true)
121    await syncHerdr($, s)
122    return result
123  })
124
125  on('session.end', async ($, e, next) => {
126    s.expiry?.cancel()
127    await reportHerdr($, s, undefined)
128    return next(e)
129  })
130
131  // The settings-hook SessionStart: its transcript_path, and on a resume or fork how long ago the
132  // last response was and how big it was. A forked session's transcript is only written with its
133  // first message, so this is the one record of the conversation's cache until then.
134  on('classic.SessionStart', async ($, e, next) => {
135    const result = await next(e)
136    s.transcript = e.transcript_path
137    const seconds = e.seconds_since_last_response
138    if (seconds !== undefined && e.context_tokens !== undefined && e.context_tokens > 0) {
139      const at = (await $.clock.now()) - seconds * 1000
140      if (!s.clock || s.clock.lastAt < at) {
141        s.clock = { lastAt: at, lastRealAt: at, promptTokens: e.context_tokens, tokens: e.context_tokens, model: e.model ?? (await $.session.model()) }
142      }
143    }
144    await syncHerdr($, s)
145    return result
146  })
147
148  on('turn.start', async ($, e, next) => {
149    s.busy = true
150    return next(e)
151  })
152
153  on('turn.complete', async ($, e, next) => {
154    const result = await next(e)
155    if (e.agentId === undefined) {
156      s.busy = false
157      await readTranscript($, s, false)
158      await schedule($, s)
159    }
160    return result
161  })
162
163  on('turn.step', async function* ($, e, next) {
164    if (e.agentId !== undefined) return yield* next(e)
165    const startedAt = await $.clock.now()
166    s.stepsInFlight++
167    s.timer?.cancel()
168    s.timer = undefined
169    try {
170      const result = yield* next(e)
171      const usage = result.usage
172      const prompt = usage ? usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens : 0
173      if (usage && prompt > 0) {
174        s.clock = { lastAt: startedAt, lastRealAt: startedAt, promptTokens: prompt, tokens: prompt + usage.output_tokens, model: usage.model }
175      }
176      return result
177    } finally {
178      s.stepsInFlight--
179      // A long tool run between steps can outlast a 5-minute TTL: keep warming then too.
180      if (s.stepsInFlight === 0) await schedule($, s)
181    }
182  })
183
184  on('prompt.submit', async ($, e, next) => {
185    if (!s.sessionOn || e.origin.kind !== 'composer' || e.turnId !== undefined) return next(e)
186    const text = e.text.trim()
187    if (!text || text.startsWith('/')) return next(e)
188    const miss = await assess($, s)
189    if (!miss) return next(e)
190    // Keep first, so a reflexive Enter (or Esc) spends nothing; then cheapest to dearest.
191    const labels = { keep: KEEP, fresh: `${FRESH} (~$0)`, compact: `${COMPACT} (~${formatCost(miss.costs.compact)})`, send: `${SEND} (~${formatCost(miss.costs.send)})` }
192    let answer: string | undefined
193    try {
194      answer = await $.ui.ask(`Prompt cache miss. ${miss.line} What now?`, { header: 'Cache', options: [labels.keep, labels.fresh, labels.compact, labels.send] })
195    } catch {
196      answer = undefined // dismissed: keep it
197    }
198    if (answer === labels.send) return next(e)
199    if (answer === labels.fresh) {
200      $.clock.after(0, () => {
201        void startFresh($, s, e.text)
202      })
203      return { drop: 'Starting a new conversation with your prompt.' }
204    }
205    if (answer === labels.compact) {
206      let how: string | undefined
207      try {
208        how = await $.ui.ask('What should the summary keep? (Type your own guidance under Other.)', { header: 'Compact', options: [SUMMARY_DEFAULT, SUMMARY_FOCUS] })
209      } catch {
210        how = undefined
211      }
212      if (how !== undefined) {
213        const instructions = how === SUMMARY_DEFAULT ? undefined : how === SUMMARY_FOCUS ? compactionFocus(e.text) : how
214        $.clock.after(0, () => {
215          void compactThenSend($, s, e.text, instructions)
216        })
217        return { drop: 'Compacting, then sending your prompt.' }
218      }
219    }
220    await $.prompt.fill({ text: e.text })
221    return { drop: `Kept in the prompt box. ${miss.line} /cache-guard off stops asking in this session.` }
222  })
223
224  on('command.run', { command: COMMAND }, async ($, e) => {
225    const arg = e.args.trim()
226    if (arg === 'on' || arg === 'off') {
227      s.sessionOn = arg === 'on'
228      if (s.sessionOn) await schedule($, s)
229      else {
230        cancel(s, 'turned off for this session')
231        await syncHerdr($, s)
232      }
233      return { text: `cache-guard ${arg} for this session` }
234    }
235    if (arg === 'warm') {
236      // A refresh now, whatever the schedule and the expected saving say (to bridge a break, or to check one works).
237      if (!s.clock) return { text: 'Nothing to keep warm yet.' }
238      await refresh($, s, true)
239      return { text: await status($, s) }
240    }
241    s.settings = await loadSettings($)
242    return { text: await status($, s) }
243  })
244}
245
246/** The shared settings files merged over the defaults; a missing file is skipped. */
247async function loadSettings($: EngineInterface): Promise<Settings> {
248  const home = (await $.env.get('HOME')) ?? ''
249  const xdg = await $.env.get('XDG_CONFIG_HOME')
250  const root = await $.session.root()
251  const texts: (string | undefined)[] = []
252  for (const file of settingsFiles(home, xdg, root)) {
253    texts.push(await $.fs.read(file).catch(() => undefined))
254  }
255  return mergeSettings(DEFAULT_SETTINGS, texts)
256}
257
258/**
259 * Reads the transcript's tail for the TTL the session's cache writes use and, when `seed` (a new or
260 * resumed session) or nothing was tracked yet, the last response as the clock.
261 */
262async function readTranscript($: EngineInterface, s: State, seed: boolean): Promise<void> {
263  try {
264    const home = (await $.env.get('HOME')) ?? ''
265    const configDir = (await $.env.get('CLAUDE_CONFIG_DIR')) || `${home}/.claude`
266    const id = await $.session.id()
267    // The project directory is the session root with every non-alphanumeric turned into '-', but
268    // of the path as Claude Code saw it (/private/tmp, not /tmp): try that, then look the id up.
269    const run = await $.process.run(
270      ['sh', '-c', TAIL_SCRIPT, 'cache-guard', s.transcript ?? transcriptPath(configDir, await $.session.root(), id), `${configDir}/projects`, `${id}.jsonl`, String(TAIL_BYTES)],
271      { timeoutMs: 5_000 },
272    )
273    if (run.exitCode !== 0) {
274      s.ttl = await configuredTtl($, s.ttl)
275      return
276    }
277    const tail = parseTail(run.stdout)
278    s.ttl = tail.ttlMs !== undefined ? { ms: tail.ttlMs, source: 'transcript' } : await configuredTtl($, s.ttl)
279    if (tail.compacted) {
280      s.clock = undefined
281      return
282    }
283    if ((seed || !s.clock) && tail.at !== undefined && tail.tokens !== undefined && tail.promptTokens !== undefined) {
284      s.clock = { lastAt: tail.at, lastRealAt: tail.at, promptTokens: tail.promptTokens, tokens: tail.tokens, model: tail.model ?? '' }
285    }
286  } catch {
287    // No transcript (a -p run with --no-session-persistence, a host without tail): keep what we have.
288  }
289}
290
291/** The TTL Claude Code is configured to request, when the transcript shows no write yet. */
292async function configuredTtl($: EngineInterface, current: { ms: number; source: string }): Promise<{ ms: number; source: string }> {
293  if (current.source === 'transcript') return current
294  if ((await $.env.get('FORCE_PROMPT_CACHING_5M')) === '1') return { ms: FIVE_MINUTES, source: 'FORCE_PROMPT_CACHING_5M' }
295  const env = (await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL'))?.trim()
296  if (env === '5m' || env === '1h') return { ms: env === '1h' ? ONE_HOUR : FIVE_MINUTES, source: 'CLAUDE_CODE_PROMPT_CACHE_TTL' }
297  const usage = await $.session.usage().catch(() => undefined)
298  // Rate-limit windows are reported to subscribers only; within plan usage the main conversation gets 1h.
299  if (usage?.rateLimits.some((r) => r.kind === 'five_hour' || r.kind === 'seven_day')) return { ms: ONE_HOUR, source: 'subscription' }
300  if ((await $.env.get('ENABLE_PROMPT_CACHING_1H')) === '1') return { ms: ONE_HOUR, source: 'ENABLE_PROMPT_CACHING_1H' }
301  // Before the first response reports rate limits: no API key or cloud provider means a Claude login.
302  const keyed = [
303    await $.env.get('ANTHROPIC_API_KEY'),
304    await $.env.get('ANTHROPIC_AUTH_TOKEN'),
305    await $.env.get('CLAUDE_CODE_USE_BEDROCK'),
306    await $.env.get('CLAUDE_CODE_USE_VERTEX'),
307    await $.env.get('CLAUDE_CODE_USE_FOUNDRY'),
308  ].some((value) => !!value)
309  return keyed ? { ms: FIVE_MINUTES, source: 'API key or cloud provider' } : { ms: ONE_HOUR, source: 'Claude login' }
310}
311
312function cancel(s: State, why: string): void {
313  s.timer?.cancel()
314  s.timer = undefined
315  s.stopped = why
316}
317
318/** Arms the next refresh (and the expiry report to herdr) after the cache was used. */
319async function schedule($: EngineInterface, s: State): Promise<void> {
320  await arm($, s)
321  await syncHerdr($, s)
322}
323
324/** Arms the next refresh at 90% of the TTL after the last cache use, inside the idle limit. */
325async function arm($: EngineInterface, s: State): Promise<void> {
326  s.timer?.cancel()
327  s.timer = undefined
328  const clock = s.clock
329  if (!clock) return cancel(s, 'no cached request yet')
330  if (!s.interactive) return cancel(s, 'not an interactive session')
331  if (!s.sessionOn || !s.settings.enabled || !s.settings.warm.enabled) return cancel(s, 'keep-warm is off')
332  const delay = warmDelayMs(s.ttl.ms)
333  if (delay === undefined) return cancel(s, 'TTL too short')
334  const due = clock.lastAt + delay
335  if (due > warmDeadline(clock.lastRealAt, s.ttl.ms, s.settings)) {
336    return cancel(s, `idle limit reached (${s.settings.warm.idleMinutes[s.ttl.ms >= ONE_HOUR ? '1h' : '5m']}m after the last request)`)
337  }
338  const now = await $.clock.now()
339  s.stopped = ''
340  s.timer = $.clock.after(Math.max(0, due - now), () => {
341    void refresh($, s)
342  })
343}
344
345/** One keep-warm request, when Pi's rule says it pays; reschedules after a hit, stops otherwise. */
346async function refresh($: EngineInterface, s: State, forced = false): Promise<void> {
347  if (!forced) s.timer = undefined
348  const clock = s.clock
349  if (!clock || s.stepsInFlight > 0 || s.warming) return
350  const now = await $.clock.now()
351  // A timer that fired late (sleep, a blocked loop) would pay a full write, not a refresh.
352  if (!forced && now >= clock.lastAt + s.ttl.ms - 1_000) return cancel(s, 'missed the refresh window (the cache expired first)')
353  const price = claudePrice(clock.model)
354  const decision = decideWarm(clock.promptTokens, price, s.ttl.ms, !s.busy, s.settings)
355  if (!forced && decision.action === 'stop') {
356    return cancel(s, `expected saving ${formatCost(decision.expectedSavings)} is under ${formatCost(s.settings.warm.minSavings)}`)
357  }
358  s.warming = true
359  try {
360    const result = await $.model.fork({ prompt: s.settings.warm.prompt })
361    const usage = 'usage' in result ? result.usage : undefined
362    const read = usage?.cache_read_input_tokens ?? 0
363    if (s.clock !== clock) return // a real request went out meanwhile and took over the clock
364    const cost = usage
365      ? (read * price.cacheRead + (usage.input_tokens + usage.cache_creation_input_tokens * 1.25) * price.input + usage.output_tokens * price.input * 5) / 1e6
366      : 0
367    if (read >= 0.8 * clock.promptTokens) {
368      s.clock = { ...clock, lastAt: now }
369      s.warms++
370      $.ui.log(`kept the prompt cache warm (${forced ? 'on request, ' : ''}${formatTokens(read)} tokens read, ~${formatCost(cost)} at API prices)`)
371      await schedule($, s)
372    } else {
373      const why = result.isAnswered || result.reason === 'empty-reply'
374        ? `the refresh read only ${formatTokens(read)} of ${formatTokens(clock.promptTokens)} tokens from the cache`
375        : `the refresh failed (${result.reason})`
376      $.ui.log(`stopped keeping the prompt cache warm: ${why}`)
377      cancel(s, why)
378    }
379  } catch (error) {
380    cancel(s, `the refresh failed (${error instanceof Error ? error.message : String(error)})`)
381  } finally {
382    s.warming = false
383  }
384}
385
386/** Why the next prompt misses the cache and what that costs, when it is worth asking about. */
387async function assess($: EngineInterface, s: State): Promise<{ line: string; costs: { send: number; compact: number } } | undefined> {
388  if (!s.settings.enabled || !s.settings.warn.enabled) return undefined
389  if (!s.clock) await readTranscript($, s, true)
390  const clock = s.clock
391  if (!clock) return undefined
392  const now = await $.clock.now()
393  const left = clock.lastAt + s.ttl.ms - now
394  if (left > 0) return undefined
395  const cost = missCost(clock.tokens, claudePrice(clock.model), s.ttl.ms)
396  if (!worthWarning(clock.tokens, cost, s.settings)) return undefined
397  return { line: describeMiss({ kind: 'expired', idleMs: -left }, clock.tokens, cost), costs: choiceCosts(clock.tokens, claudePrice(clock.model), s.ttl.ms) }
398}
399
400async function status($: EngineInterface, s: State): Promise<string> {
401  const now = await $.clock.now()
402  const clock = s.clock
403  const lines: string[] = []
404  if (!clock) {
405    lines.push('No cached request in this session yet (or it was just compacted).')
406  } else {
407    const left = clock.lastAt + s.ttl.ms - now
408    lines.push(`Last request: ${clock.model}, ${formatTokens(clock.tokens)} tokens, ${formatDuration(now - clock.lastRealAt)} ago` +
409      (clock.lastAt > clock.lastRealAt ? `; kept warm ${s.warms}x, last ${formatDuration(now - clock.lastAt)} ago.` : '.'))
410    lines.push(`TTL: ${s.ttl.ms >= ONE_HOUR ? '1h' : '5m'} (${s.ttl.source}). ${left > 0 ? `Warm, ${formatClock(left)} left.` : describeMiss({ kind: 'expired', idleMs: -left }, clock.tokens, missCost(clock.tokens, claudePrice(clock.model), s.ttl.ms))}`)
411  }
412  const on = s.sessionOn && s.settings.enabled
413  lines.push(`Keep-warm: ${on && s.settings.warm.enabled ? (s.timer ? 'next refresh scheduled' : `idle (${s.stopped || 'nothing to keep'})`) : 'off'}.`)
414  lines.push(`Warning: ${on && s.settings.warn.enabled ? `on, from ${formatCost(s.settings.warn.minCost)} at API prices` : 'off'}.`)
415  return lines.join('\n')
416}
417
418/**
419 * herdr's `cache` pane token: "cold 664k" once the cache has expired with a re-cache worth a
420 * warning, nothing while it is warm (an expiry timer re-checks then) or small. Inside herdr only.
421 */
422async function syncHerdr($: EngineInterface, s: State): Promise<void> {
423  s.expiry?.cancel()
424  s.expiry = undefined
425  const clock = s.clock
426  let value: string | undefined
427  if (clock && s.sessionOn && s.stepsInFlight === 0) {
428    const now = await $.clock.now()
429    const left = clock.lastAt + s.ttl.ms - now
430    if (left > 0) {
431      s.expiry = $.clock.after(left + 1_000, () => {
432        void syncHerdr($, s)
433      })
434    } else {
435      value = herdrCacheValue({ kind: 'expired', idleMs: -left }, clock.tokens, missCost(clock.tokens, claudePrice(clock.model), s.ttl.ms), s.settings)
436    }
437  }
438  await reportHerdr($, s, value)
439}
440
441/** One `herdr pane report-metadata` call when the token changes; any failure is ignored. */
442async function reportHerdr($: EngineInterface, s: State, value: string | undefined): Promise<void> {
443  if (value === s.herdrLast) return
444  if ((await $.env.get('HERDR_ENV')) !== '1') return
445  const pane = await $.env.get('HERDR_PANE_ID')
446  if (!pane) return
447  s.herdrLast = value
448  const change = value === undefined ? ['--clear-token', HERDR_TOKEN] : ['--token', `${HERDR_TOKEN}=${value}`, '--ttl-ms', '86400000']
449  await $.process.run(['herdr', 'pane', 'report-metadata', pane, '--source', HERDR_SOURCE, '--agent', 'claude', ...change], { timeoutMs: 3_000 }).catch(() => undefined)
450}
451
452/** /clear, then the held prompt as the new conversation's first. */
453async function startFresh($: EngineInterface, s: State, text: string): Promise<void> {
454  try {
455    await $.command.run({ command: 'clear' })
456    s.clock = undefined
457    s.timer?.cancel()
458    s.timer = undefined
459    await $.prompt.submit({ text, asUser: true })
460  } catch (error) {
461    $.ui.log(`could not start a new conversation (${error instanceof Error ? error.message : String(error)}); your prompt is back in the box`)
462    await $.prompt.fill({ text })
463  }
464}
465
466/** Compaction (with the chosen guidance), then the held prompt onto the summary. */
467async function compactThenSend($: EngineInterface, s: State, text: string, instructions: string | undefined): Promise<void> {
468  try {
469    const result = await $.session.compact(instructions ? { instructions } : {})
470    if ('skip' in result && typeof result.skip === 'string') {
471      $.ui.log(`compaction was skipped (${result.skip}); your prompt is back in the box`)
472      await $.prompt.fill({ text })
473      return
474    }
475    s.clock = undefined
476    await $.prompt.submit({ text, asUser: true })
477  } catch (error) {
478    $.ui.log(`compaction failed (${error instanceof Error ? error.message : String(error)}); your prompt is back in the box`)
479    await $.prompt.fill({ text })
480  }
481}
482
hooks/core.ts 276 lines
1// cache-guard core: the prompt-cache clock, the cost of a miss, and when a keep-warm request or a
2// warning pays. Shared verbatim by pi-cache-guard, claude-cache-guard, opencode-cache-guard and
3// codex-cache-guard; keep it free of harness imports so each port can copy this file as is.
4//
5// The model is Anthropic's: a cached prefix lives for its TTL from the start of the last request
6// that read or wrote it (the API guarantees that minimum and deletes soon after), and a request
7// after that re-writes the whole prefix. Prices are dollars per million tokens at API list rates;
8// on a subscription they stand for plan usage in the same proportions.
9
10export const NAME = "cache-guard";
11
12export interface Price {
13  /** Uncached input. */
14  input: number;
15  /** A cache read (hit or refresh). */
16  cacheRead: number;
17  /** A 5-minute cache write; 1.25x input when absent. A 1-hour write is always 2x input. */
18  cacheWrite?: number;
19}
20
21export interface Settings {
22  enabled: boolean;
23  warn: {
24    enabled: boolean;
25    /** Ask before a prompt whose re-cache costs at least this many dollars (when prices are known). */
26    minCost: number;
27    /** Without prices, ask when at least this many tokens would be re-sent uncached. */
28    minTokens: number;
29    /** Where a harness can only block, sending the same prompt again within this window sends it. */
30    confirmSeconds: number;
31    /** Providers that publish no TTL (OpenAI, Codex): warn after this long idle. */
32    idleMinutes: number;
33  };
34  warm: {
35    enabled: boolean;
36    /** Chance a real request arrives before the entry expires while idle (Pi's measured constant). */
37    continuationProbability: number;
38    /** A refresh is sent only when it is expected to save at least this many dollars. */
39    minSavings: number;
40    /** Stop refreshing this long after the last real request, per TTL tier. */
41    idleMinutes: { "5m": number; "1h": number };
42    /** What a keep-warm request asks, where the harness has to send a message. */
43    prompt: string;
44  };
45  herdr: {
46    /** Report the pane token `cache` to herdr (inside a herdr pane) so its agents sidebar can show doomed sessions. */
47    enabled: boolean;
48  };
49}
50
51export const DEFAULT_SETTINGS: Settings = {
52  enabled: true,
53  warn: { enabled: true, minCost: 0.5, minTokens: 100_000, confirmSeconds: 120, idleMinutes: 180 },
54  warm: {
55    enabled: true,
56    continuationProbability: 0.15,
57    minSavings: 0.05,
58    idleMinutes: { "5m": 30, "1h": 120 },
59    prompt: "Cache keep-alive. Do not use tools or think. Reply with exactly: ok",
60  },
61  herdr: { enabled: true },
62};
63
64/** The herdr pane token the ports report (`herdr pane report-metadata --source cache-guard`). */
65export const HERDR_TOKEN = "cache";
66export const HERDR_SOURCE = NAME;
67
68export const FIVE_MINUTES = 5 * 60_000;
69export const ONE_HOUR = 60 * 60_000;
70
71/** "5m" or "1h": the tier a TTL bills as (anything from an hour up writes at 2x). */
72export function tier(ttlMs: number): "5m" | "1h" {
73  return ttlMs >= ONE_HOUR ? "1h" : "5m";
74}
75
76/** Refresh at 90% of the TTL, keeping at least ten seconds of margin (Pi's rule). */
77export function warmDelayMs(ttlMs: number): number | undefined {
78  if (ttlMs <= 10_000) return undefined;
79  return Math.max(1, Math.floor(Math.min(ttlMs * 0.9, ttlMs - 10_000)));
80}
81
82/** Milliseconds the entry written or refreshed at `lastAt` has left; 0 once expired. */
83export function remainingMs(lastAt: number, ttlMs: number, now: number): number {
84  return Math.max(0, lastAt + ttlMs - now);
85}
86
87export function writePrice(price: Price, ttlMs: number): number {
88  return tier(ttlMs) === "1h" ? price.input * 2 : (price.cacheWrite ?? price.input * 1.25);
89}
90
91/** What a miss costs over a hit: the prefix written again instead of read. */
92export function missCost(tokens: number, price: Price, ttlMs: number): number {
93  return Math.max(0, (tokens * (writePrice(price, ttlMs) - price.cacheRead)) / 1e6);
94}
95
96/** What one refresh costs: the prefix read (plus a token or two, ignored). */
97export function warmCost(tokens: number, price: Price): number {
98  return (tokens * price.cacheRead) / 1e6;
99}
100
101export interface WarmDecision {
102  action: "warm" | "stop";
103  warmCost: number;
104  missCost: number;
105  expectedSavings: number;
106}
107
108/**
109 * Pi's rule: refresh when `p * missCost - warmCost` is at least `minSavings`, with p = 1 while the
110 * agent is still running (its next request is certain) and the idle constant otherwise.
111 */
112export function decideWarm(tokens: number, price: Price, ttlMs: number, idle: boolean, settings: Settings): WarmDecision {
113  const miss = missCost(tokens, price, ttlMs);
114  const warm = warmCost(tokens, price);
115  const p = idle ? settings.warm.continuationProbability : 1;
116  const expectedSavings = p * miss - warm;
117  return { action: expectedSavings >= settings.warm.minSavings ? "warm" : "stop", warmCost: warm, missCost: miss, expectedSavings };
118}
119
120/** The last moment a refresh may be sent for a cache last used by a real request at `lastRealAt`. */
121export function warmDeadline(lastRealAt: number, ttlMs: number, settings: Settings): number {
122  return lastRealAt + settings.warm.idleMinutes[tier(ttlMs)] * 60_000;
123}
124
125/** Whether a re-cache of `tokens` (costing `cost`, when prices are known) is worth asking about. */
126export function worthWarning(tokens: number, cost: number | undefined, settings: Settings): boolean {
127  if (!settings.enabled || !settings.warn.enabled) return false;
128  return cost !== undefined ? cost >= settings.warn.minCost : tokens >= settings.warn.minTokens;
129}
130
131export type ColdReason =
132  | { kind: "expired"; idleMs: number }
133  | { kind: "model"; from: string; to: string }
134  | { kind: "idle"; idleMs: number };
135
136/**
137 * The herdr `cache` token: `cold 664k` (or `cold? 180k` when only idle time suggests it) while the
138 * next prompt would re-cache at least the warning threshold, else undefined (clear the token).
139 * Warm and small caches report nothing, so the sidebar lists only the doomed sessions.
140 */
141export function herdrCacheValue(reason: ColdReason | undefined, tokens: number, cost: number | undefined, settings: Settings): string | undefined {
142  if (!reason || !settings.enabled || !settings.herdr.enabled) return undefined;
143  const big = cost !== undefined ? cost >= settings.warn.minCost : tokens >= settings.warn.minTokens;
144  if (!big) return undefined;
145  return `${reason.kind === "idle" ? "cold?" : "cold"} ${formatTokens(tokens)}`;
146}
147
148/**
149 * Rough costs, in dollars at list prices, of the two ways through a cold cache that keep the
150 * history: send the prompt and write the whole prefix to the cache again, or compact first, which
151 * reads it once uncached (plus a summary, not counted) and continues on a small context. Starting
152 * fresh costs about nothing.
153 */
154export function choiceCosts(tokens: number, price: Price, ttlMs: number): { send: number; compact: number } {
155  return { send: (tokens * writePrice(price, ttlMs)) / 1e6, compact: (tokens * price.input) / 1e6 };
156}
157
158/** Compaction guidance that keeps what the held prompt needs. */
159export function compactionFocus(prompt: string): string {
160  return `Keep what is needed to continue with the user's next request, quoted below, and drop the rest.\n\n${prompt.trim()}`;
161}
162
163/** One line saying why the next request misses and what that costs. */
164export function describeMiss(reason: ColdReason, tokens: number, cost: number | undefined): string {
165  const amount = `${formatTokens(tokens)} tokens${cost === undefined ? "" : ` (~${formatCost(cost)} at API prices)`}`;
166  switch (reason.kind) {
167    case "expired":
168      return `The prompt cache expired ${formatDuration(reason.idleMs)} ago: this prompt re-caches ${amount}.`;
169    case "idle":
170      return `Idle ${formatDuration(reason.idleMs)}: the prompt cache has probably expired, so this prompt may re-cache ${amount}.`;
171    case "model":
172      return `${reason.to} has no cache of this conversation (it was cached for ${reason.from}): this prompt re-caches ${amount}.`;
173  }
174}
175
176/** Remembers a blocked prompt so the same prompt sent again within the window goes through. */
177export class ConfirmMemo {
178  private pending?: { key: string; text: string; at: number };
179
180  /** True when `text` repeats the prompt blocked for `key` within `windowMs`; clears it either way. */
181  confirmed(key: string, text: string, now: number, windowMs: number): boolean {
182    const pending = this.pending;
183    this.pending = undefined;
184    return pending !== undefined && pending.key === key && pending.text === text.trim() && now - pending.at <= windowMs;
185  }
186
187  arm(key: string, text: string, now: number): void {
188    this.pending = { key, text: text.trim(), at: now };
189  }
190}
191
192/** "4:05" under an hour, "1h05m" above, "0:00" once gone. */
193export function formatClock(ms: number): string {
194  const seconds = Math.max(0, Math.ceil(ms / 1000));
195  if (seconds >= 3600) return `${Math.floor(seconds / 3600)}h${String(Math.floor((seconds % 3600) / 60)).padStart(2, "0")}m`;
196  return `${Math.floor(seconds / 60)}:${String(seconds % 60).padStart(2, "0")}`;
197}
198
199/** "45s", "12m", "3h20m", "2d4h". */
200export function formatDuration(ms: number): string {
201  const seconds = Math.max(0, Math.round(ms / 1000));
202  if (seconds < 60) return `${seconds}s`;
203  const minutes = Math.floor(seconds / 60);
204  if (minutes < 60) return `${minutes}m`;
205  const hours = Math.floor(minutes / 60);
206  if (hours < 48) return `${hours}h${minutes % 60 ? `${minutes % 60}m` : ""}`;
207  return `${Math.floor(hours / 24)}d${hours % 24 ? `${hours % 24}h` : ""}`;
208}
209
210export function formatTokens(n: number): string {
211  if (n >= 1e6) return `${(n / 1e6).toFixed(n >= 1e7 ? 0 : 1)}M`;
212  if (n >= 1e3) return `${Math.round(n / 1e3)}k`;
213  return String(Math.round(n));
214}
215
216export function formatCost(dollars: number): string {
217  return dollars >= 10 ? `$${dollars.toFixed(0)}` : `$${dollars.toFixed(2)}`;
218}
219
220/**
221 * List prices of the Claude models, for harnesses that do not carry a model catalog (Claude Code,
222 * Codex has no use for it). Longest prefix wins; unknown models price as Opus 5.5.
223 */
224const CLAUDE_PRICES: ReadonlyArray<readonly [string, Price]> = [
225  ["claude-fable-5-1", { input: 10, cacheRead: 0.25 }],
226  ["claude-fable-5", { input: 10, cacheRead: 1 }],
227  ["claude-mythos-5-1", { input: 10, cacheRead: 0.25 }],
228  ["claude-opus-5-5", { input: 4, cacheRead: 0.2 }],
229  ["claude-opus-5", { input: 5, cacheRead: 0.5 }],
230  ["claude-opus-4", { input: 5, cacheRead: 0.5 }],
231  ["claude-sonnet-5", { input: 2, cacheRead: 0.2 }],
232  ["claude-sonnet-4", { input: 3, cacheRead: 0.3 }],
233  ["claude-haiku-5", { input: 0.1, cacheRead: 0.01 }],
234  ["claude-haiku-4", { input: 1, cacheRead: 0.1 }],
235];
236
237export function claudePrice(model: string): Price {
238  const id = model.toLowerCase().replace(/^.*\//, "").replace(/\[.*$/, "");
239  let best: Price | undefined;
240  let length = 0;
241  for (const [prefix, price] of CLAUDE_PRICES) {
242    if (id.startsWith(prefix) && prefix.length > length) {
243      best = price;
244      length = prefix.length;
245    }
246  }
247  return best ?? { input: 4, cacheRead: 0.2 };
248}
249
250export function mergeSettings(base: Settings, texts: readonly (string | undefined)[]): Settings {
251  let merged: Record<string, unknown> = structuredClone(base) as unknown as Record<string, unknown>;
252  for (const text of texts) {
253    if (text === undefined) continue;
254    try {
255      const value: unknown = JSON.parse(text);
256      if (isRecord(value)) merged = merge(merged, value);
257    } catch {
258      // Invalid JSON: keep what the earlier files said.
259    }
260  }
261  return merged as unknown as Settings;
262}
263
264function merge(base: Record<string, unknown>, over: Record<string, unknown>): Record<string, unknown> {
265  const out: Record<string, unknown> = { ...base };
266  for (const [key, value] of Object.entries(over)) {
267    const current = out[key];
268    out[key] = isRecord(current) && isRecord(value) ? merge(current, value) : value;
269  }
270  return out;
271}
272
273function isRecord(value: unknown): value is Record<string, unknown> {
274  return typeof value === "object" && value !== null && !Array.isArray(value);
275}
276
hooks/settings.ts 17 lines
1// Settings shared with the pi, opencode and Codex ports: `~/.config/agents/cache-guard.json` (or
2// under $XDG_CONFIG_HOME) and `<project>/.agents/cache-guard.json`, with Claude Code's own
3// `~/.claude/cache-guard.json` and `<project>/.claude/cache-guard.json` as overrides. Later files
4// win; objects merge, other values replace.
5import { NAME } from './core'
6
7/** The settings files, lowest precedence first. */
8export function settingsFiles(home: string, xdgConfigHome: string | undefined, root: string): string[] {
9  const shared = xdgConfigHome || `${home}/.config`
10  return [
11    `${shared}/agents/${NAME}.json`,
12    `${home}/.claude/${NAME}.json`,
13    `${root}/.agents/${NAME}.json`,
14    `${root}/.claude/${NAME}.json`,
15  ]
16}
17
hooks/transcript.ts 74 lines
1// Reads the tail of a Claude Code transcript (~/.claude/projects/<slug>/<session>.jsonl): the last
2// main-conversation response, and which TTL the session's cache writes use. The usage rows carry
3// `cache_creation.ephemeral_1h_input_tokens` / `ephemeral_5m_input_tokens`, which nothing in the
4// mod API reports.
5
6export interface TranscriptTail {
7  /** When the last main-thread response was written (its end, so later than the request's start). */
8  at?: number;
9  /** Its prompt: input + cache reads + cache writes. */
10  promptTokens?: number;
11  /** Prompt plus the reply: what the next request re-sends. */
12  tokens?: number;
13  model?: string;
14  /** The tier of the most recent main-thread cache write in the tail. */
15  ttlMs?: number;
16  /** A compaction after the last response: the next request carries a fresh context. */
17  compacted?: boolean;
18}
19
20interface Usage {
21  input_tokens?: number;
22  output_tokens?: number;
23  cache_read_input_tokens?: number;
24  cache_creation_input_tokens?: number;
25  cache_creation?: { ephemeral_1h_input_tokens?: number; ephemeral_5m_input_tokens?: number };
26}
27
28/** The transcript file Claude Code writes for `sessionId` started in `root`. */
29export function transcriptPath(configDir: string, root: string, sessionId: string): string {
30  return `${configDir}/projects/${root.replace(/[^a-zA-Z0-9]/g, '-')}/${sessionId}.jsonl`
31}
32
33/** Parses JSONL text (possibly starting mid-line) from the end back. */
34export function parseTail(text: string): TranscriptTail {
35  const out: TranscriptTail = {}
36  const lines = text.split('\n')
37  for (let i = lines.length - 1; i >= 0; i--) {
38    const line = lines[i]
39    if (!line || line[0] !== '{') continue
40    let row: Record<string, any>
41    try {
42      row = JSON.parse(line)
43    } catch {
44      continue // the first, cut line
45    }
46    if (out.at === undefined && row.type === 'system' && row.subtype === 'compact_boundary') {
47      out.compacted = true
48      continue
49    }
50    if (row.type !== 'assistant' || row.isSidechain === true) continue
51    const message = row.message ?? {}
52    if (message.model === '<synthetic>' || row.isApiErrorMessage === true) continue
53    const usage: Usage | undefined = message.usage
54    if (!usage) continue
55    const prompt = (usage.input_tokens ?? 0) + (usage.cache_read_input_tokens ?? 0) + (usage.cache_creation_input_tokens ?? 0)
56    if (prompt <= 0) continue
57    if (out.at === undefined) {
58      const at = Date.parse(row.timestamp)
59      if (!Number.isFinite(at)) continue
60      out.at = at
61      out.promptTokens = prompt
62      out.tokens = prompt + (usage.output_tokens ?? 0)
63      out.model = String(message.model ?? '')
64    }
65    const oneHour = usage.cache_creation?.ephemeral_1h_input_tokens ?? 0
66    const fiveMinutes = usage.cache_creation?.ephemeral_5m_input_tokens ?? 0
67    if (oneHour > 0 || fiveMinutes > 0) {
68      out.ttlMs = oneHour >= fiveMinutes ? 3_600_000 : 300_000
69      break
70    }
71  }
72  return out
73}
74