Keeps the prompt cache warm while you are away and asks before a prompt that would re-cache a large conversation.

A Claude Code mod that keeps the prompt cache warm while you are away and asks before a prompt that would re-cache a large conversation. The Claude Code port of pi-cache-guard, with the same core (hooks/core.ts) and settings file.
Keeps the cache warm. A cached prefix lives for its TTL from the start of the last request that used it: 1 hour for the main conversation on a Claude subscription within plan usage, 5 minutes with an API key, a cloud provider or usage credits. The mod reads the tier off the session's own cache writes in the transcript. At 90% of the TTL with no request since, it calls $.model.fork, which re-sends the main thread's last request (same model, system prompt, tools and messages) with one short question after it, tools denied and its own tail never cached. The API serves the prefix from the cache, which restarts its clock.
Each refresh has to pay for itself by Pi's rule:
15% x miss cost - refresh cost; during a long tool run it is the whole miss cost.kept the prompt cache warm (601k tokens read, ~$0.12 at API prices). The model never sees these lines.Warns. A prompt you type while idle onto an expired cache asks first when the rewrite would cost at least warn.minCost (default $0.50 at API prices):
Prompt cache miss. The prompt cache expired 10m ago: this prompt re-caches 601k tokens (~$4.69 at API prices). What now?
1. Keep the prompt
2. New conversation (~$0)
3. Compact first (~$2.40)
4. Send anyway (~$4.81)
The options:
/clear and sends the prompt as the new conversation's first.The headline cost is what the miss adds over a cache hit; the costs in the options are each path's total. Images attached to a held prompt are not resent by the New conversation and Compact paths. /cache-guard off stops asking in the session. Resumed and forked sessions are judged from the transcript, or from the SessionStart hook input (seconds_since_last_response, context_tokens) when the transcript is not written yet. Prompts from loops, peers, notifications, the SDK and -p are never held, and neither are slash commands.
Status. Claude Code reports the cache to the status line itself (prompt_cache.warm, expires_at, ttl, recache_tokens_if_cold). A mod cannot ship a status line, so the countdown belongs in your status line script, e.g. jq '.prompt_cache.expires_at'. The mod draws none.
Tells herdr. Inside a herdr pane, the mod sets the pane token cache with herdr pane report-metadata --source cache-guard. The token reads cold 601k once the cache has expired and the next prompt would re-cache at least the warning threshold. It is cleared while the cache is warm or small, and when the session ends. An expiry timer flips it on time. herdr's agents sidebar shows it with { token = "$cache" } in a [ui.sidebar.agents] row. Set "herdr": { "enabled": false } to turn it off.
Claude Code's own guards stay in place: /model and /effort ask while the cache is warm, and /resume offers to resume from a summary after a long break.
/cache-guard or /cache-guard status: the last request, TTL and its source, time left or how long since it expired, and the keep-warm state./cache-guard warm: a refresh now, whatever the schedule and the expected saving say. Use it to bridge a break, or to check that refreshes hit./cache-guard on, /cache-guard off: for this session.~/.config/agents/cache-guard.json (shared with the other ports), then ~/.claude/cache-guard.json, then the project's .agents/cache-guard.json and .claude/cache-guard.json. Later files win; objects merge.
{
"enabled": true,
"warn": { "enabled": true, "minCost": 0.5 },
"warm": { "enabled": true, "continuationProbability": 0.15, "minSavings": 0.05,
"idleMinutes": { "5m": 30, "1h": 120 },
"prompt": "Cache keep-alive. Do not use tools or think. Reply with exactly: ok" }
}
It is a plugin directory with a hooks module. Load it with claude --plugin-dir <checkout>, or list it in CLAUDE_CODE_PLUGIN_DIRS. It needs Claude Code 2.1.287 or newer and was tested with 2.1.289.
prompt_cache.last_miss_cause reports those; this mod only judges time.hooks/core.ts). Unknown models are priced as Opus 5.5.claude plugin validate . # static analysis of the hooks module
claude plugin test # tests/, no session or network
tsc -p . # after one load, which writes .claude-plugin/types/hooks/register.ts 482 lines1// cache-guard: keeps Claude Code's prompt cache warm while you are away, and asks before a prompt
2// that would re-cache a large conversation.
3//
4// The clock: every main-loop model request (turn.step) reads or writes the cache, and the entry
5// lives for its TTL from that request's start. The TTL is the tier the session's own cache writes
6// use, read off the transcript (1h on a subscription within plan usage, 5m otherwise).
7//
8// Keep-warm: at 90% of the TTL with no request since, `$.model.fork` re-sends the main thread's
9// last request (same model, system prompt, tools and messages) with one short question after it,
10// tools denied and its own tail never cached, so the API serves the prefix from the cache and
11// restarts its clock. Pi's rule decides each refresh (expected saving at least $0.05, with a 15%
12// chance you come back before it lapses while idle), and refreshing stops 30 minutes (5m tier) or
13// 2 hours (1h tier) after the last real request. A refresh that misses stops it.
14//
15// Warning: a prompt typed while idle onto an expired cache, when the rewrite would cost at least
16// warn.minCost at API prices, asks first; keeping it puts the text back in the prompt box.
17//
18// The status line shows the cache itself (Claude Code's prompt_cache input), so this mod draws none.
19//
20// The host reads on(...) and $.noun.method(...) from source, so calls are spelled in full and
21// helpers that take $ are top-level functions in this file.
22import type { EngineInterface, On } from 'claude-code'
23
24import {
25 DEFAULT_SETTINGS,
26 NAME,
27 ONE_HOUR,
28 FIVE_MINUTES,
29 HERDR_SOURCE,
30 HERDR_TOKEN,
31 type Settings,
32 claudePrice,
33 decideWarm,
34 describeMiss,
35 formatClock,
36 formatCost,
37 formatDuration,
38 formatTokens,
39 choiceCosts,
40 compactionFocus,
41 herdrCacheValue,
42 mergeSettings,
43 missCost,
44 warmDeadline,
45 warmDelayMs,
46 worthWarning,
47} from './core'
48import { settingsFiles } from './settings'
49import { parseTail, transcriptPath } from './transcript'
50
51const COMMAND = NAME
52const KEEP = 'Keep the prompt'
53const FRESH = 'New conversation'
54const COMPACT = 'Compact first'
55const SEND = 'Send anyway'
56const SUMMARY_DEFAULT = 'Default summary'
57const SUMMARY_FOCUS = 'Focus on this prompt'
58/** How much of a transcript's end to read: enough for the last few responses. */
59const TAIL_BYTES = 400_000
60const TAIL_SCRIPT = 'f=$1; [ -f "$f" ] || f=$(ls "$2"/*/"$3" 2>/dev/null | head -n 1); [ -n "$f" ] && tail -c "$4" "$f"'
61
62interface Clock {
63 /** Start of the last request that read or wrote the cache (a real one or a refresh). */
64 lastAt: number
65 /** Start of the last real request. */
66 lastRealAt: number
67 /** The last request's prompt: what a refresh should read from the cache. */
68 promptTokens: number
69 /** Prompt plus reply: what the next real request re-sends. */
70 tokens: number
71 model: string
72}
73
74interface State {
75 settings: Settings
76 sessionOn: boolean
77 interactive: boolean
78 clock?: Clock
79 /** The transcript file, from the settings-hook SessionStart input. */
80 transcript?: string
81 ttl: { ms: number; source: string }
82 timer?: { cancel: () => void }
83 stepsInFlight: number
84 busy: boolean
85 warming: boolean
86 warms: number
87 /** Why no refresh is scheduled, for /cache-guard. */
88 stopped: string
89 /** Fires when the cache expires, to tell herdr. */
90 expiry?: { cancel: () => void }
91 /** The `cache` token last reported to herdr; null before the first report. */
92 herdrLast: string | undefined | null
93}
94
95export function register(on: On): void {
96 const s: State = {
97 settings: DEFAULT_SETTINGS,
98 sessionOn: true,
99 interactive: false,
100 ttl: { ms: FIVE_MINUTES, source: 'default' },
101 stepsInFlight: 0,
102 busy: false,
103 warming: false,
104 warms: 0,
105 stopped: 'waiting for the first response',
106 herdrLast: null,
107 }
108
109 on('session.start', async ($, e, next) => {
110 const result = await next(e)
111 s.interactive = e.isInteractive
112 s.settings = await loadSettings($)
113 await $.command.register({
114 name: COMMAND,
115 description: 'cache-guard: prompt cache status, warm (refresh now), on, or off (this session)',
116 argumentHint: '[status|warm|on|off]',
117 immediate: true,
118 })
119 // A resumed session: its last response and TTL come from the transcript.
120 await readTranscript($, s, true)
121 await syncHerdr($, s)
122 return result
123 })
124
125 on('session.end', async ($, e, next) => {
126 s.expiry?.cancel()
127 await reportHerdr($, s, undefined)
128 return next(e)
129 })
130
131 // The settings-hook SessionStart: its transcript_path, and on a resume or fork how long ago the
132 // last response was and how big it was. A forked session's transcript is only written with its
133 // first message, so this is the one record of the conversation's cache until then.
134 on('classic.SessionStart', async ($, e, next) => {
135 const result = await next(e)
136 s.transcript = e.transcript_path
137 const seconds = e.seconds_since_last_response
138 if (seconds !== undefined && e.context_tokens !== undefined && e.context_tokens > 0) {
139 const at = (await $.clock.now()) - seconds * 1000
140 if (!s.clock || s.clock.lastAt < at) {
141 s.clock = { lastAt: at, lastRealAt: at, promptTokens: e.context_tokens, tokens: e.context_tokens, model: e.model ?? (await $.session.model()) }
142 }
143 }
144 await syncHerdr($, s)
145 return result
146 })
147
148 on('turn.start', async ($, e, next) => {
149 s.busy = true
150 return next(e)
151 })
152
153 on('turn.complete', async ($, e, next) => {
154 const result = await next(e)
155 if (e.agentId === undefined) {
156 s.busy = false
157 await readTranscript($, s, false)
158 await schedule($, s)
159 }
160 return result
161 })
162
163 on('turn.step', async function* ($, e, next) {
164 if (e.agentId !== undefined) return yield* next(e)
165 const startedAt = await $.clock.now()
166 s.stepsInFlight++
167 s.timer?.cancel()
168 s.timer = undefined
169 try {
170 const result = yield* next(e)
171 const usage = result.usage
172 const prompt = usage ? usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens : 0
173 if (usage && prompt > 0) {
174 s.clock = { lastAt: startedAt, lastRealAt: startedAt, promptTokens: prompt, tokens: prompt + usage.output_tokens, model: usage.model }
175 }
176 return result
177 } finally {
178 s.stepsInFlight--
179 // A long tool run between steps can outlast a 5-minute TTL: keep warming then too.
180 if (s.stepsInFlight === 0) await schedule($, s)
181 }
182 })
183
184 on('prompt.submit', async ($, e, next) => {
185 if (!s.sessionOn || e.origin.kind !== 'composer' || e.turnId !== undefined) return next(e)
186 const text = e.text.trim()
187 if (!text || text.startsWith('/')) return next(e)
188 const miss = await assess($, s)
189 if (!miss) return next(e)
190 // Keep first, so a reflexive Enter (or Esc) spends nothing; then cheapest to dearest.
191 const labels = { keep: KEEP, fresh: `${FRESH} (~$0)`, compact: `${COMPACT} (~${formatCost(miss.costs.compact)})`, send: `${SEND} (~${formatCost(miss.costs.send)})` }
192 let answer: string | undefined
193 try {
194 answer = await $.ui.ask(`Prompt cache miss. ${miss.line} What now?`, { header: 'Cache', options: [labels.keep, labels.fresh, labels.compact, labels.send] })
195 } catch {
196 answer = undefined // dismissed: keep it
197 }
198 if (answer === labels.send) return next(e)
199 if (answer === labels.fresh) {
200 $.clock.after(0, () => {
201 void startFresh($, s, e.text)
202 })
203 return { drop: 'Starting a new conversation with your prompt.' }
204 }
205 if (answer === labels.compact) {
206 let how: string | undefined
207 try {
208 how = await $.ui.ask('What should the summary keep? (Type your own guidance under Other.)', { header: 'Compact', options: [SUMMARY_DEFAULT, SUMMARY_FOCUS] })
209 } catch {
210 how = undefined
211 }
212 if (how !== undefined) {
213 const instructions = how === SUMMARY_DEFAULT ? undefined : how === SUMMARY_FOCUS ? compactionFocus(e.text) : how
214 $.clock.after(0, () => {
215 void compactThenSend($, s, e.text, instructions)
216 })
217 return { drop: 'Compacting, then sending your prompt.' }
218 }
219 }
220 await $.prompt.fill({ text: e.text })
221 return { drop: `Kept in the prompt box. ${miss.line} /cache-guard off stops asking in this session.` }
222 })
223
224 on('command.run', { command: COMMAND }, async ($, e) => {
225 const arg = e.args.trim()
226 if (arg === 'on' || arg === 'off') {
227 s.sessionOn = arg === 'on'
228 if (s.sessionOn) await schedule($, s)
229 else {
230 cancel(s, 'turned off for this session')
231 await syncHerdr($, s)
232 }
233 return { text: `cache-guard ${arg} for this session` }
234 }
235 if (arg === 'warm') {
236 // A refresh now, whatever the schedule and the expected saving say (to bridge a break, or to check one works).
237 if (!s.clock) return { text: 'Nothing to keep warm yet.' }
238 await refresh($, s, true)
239 return { text: await status($, s) }
240 }
241 s.settings = await loadSettings($)
242 return { text: await status($, s) }
243 })
244}
245
246/** The shared settings files merged over the defaults; a missing file is skipped. */
247async function loadSettings($: EngineInterface): Promise<Settings> {
248 const home = (await $.env.get('HOME')) ?? ''
249 const xdg = await $.env.get('XDG_CONFIG_HOME')
250 const root = await $.session.root()
251 const texts: (string | undefined)[] = []
252 for (const file of settingsFiles(home, xdg, root)) {
253 texts.push(await $.fs.read(file).catch(() => undefined))
254 }
255 return mergeSettings(DEFAULT_SETTINGS, texts)
256}
257
258/**
259 * Reads the transcript's tail for the TTL the session's cache writes use and, when `seed` (a new or
260 * resumed session) or nothing was tracked yet, the last response as the clock.
261 */
262async function readTranscript($: EngineInterface, s: State, seed: boolean): Promise<void> {
263 try {
264 const home = (await $.env.get('HOME')) ?? ''
265 const configDir = (await $.env.get('CLAUDE_CONFIG_DIR')) || `${home}/.claude`
266 const id = await $.session.id()
267 // The project directory is the session root with every non-alphanumeric turned into '-', but
268 // of the path as Claude Code saw it (/private/tmp, not /tmp): try that, then look the id up.
269 const run = await $.process.run(
270 ['sh', '-c', TAIL_SCRIPT, 'cache-guard', s.transcript ?? transcriptPath(configDir, await $.session.root(), id), `${configDir}/projects`, `${id}.jsonl`, String(TAIL_BYTES)],
271 { timeoutMs: 5_000 },
272 )
273 if (run.exitCode !== 0) {
274 s.ttl = await configuredTtl($, s.ttl)
275 return
276 }
277 const tail = parseTail(run.stdout)
278 s.ttl = tail.ttlMs !== undefined ? { ms: tail.ttlMs, source: 'transcript' } : await configuredTtl($, s.ttl)
279 if (tail.compacted) {
280 s.clock = undefined
281 return
282 }
283 if ((seed || !s.clock) && tail.at !== undefined && tail.tokens !== undefined && tail.promptTokens !== undefined) {
284 s.clock = { lastAt: tail.at, lastRealAt: tail.at, promptTokens: tail.promptTokens, tokens: tail.tokens, model: tail.model ?? '' }
285 }
286 } catch {
287 // No transcript (a -p run with --no-session-persistence, a host without tail): keep what we have.
288 }
289}
290
291/** The TTL Claude Code is configured to request, when the transcript shows no write yet. */
292async function configuredTtl($: EngineInterface, current: { ms: number; source: string }): Promise<{ ms: number; source: string }> {
293 if (current.source === 'transcript') return current
294 if ((await $.env.get('FORCE_PROMPT_CACHING_5M')) === '1') return { ms: FIVE_MINUTES, source: 'FORCE_PROMPT_CACHING_5M' }
295 const env = (await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL'))?.trim()
296 if (env === '5m' || env === '1h') return { ms: env === '1h' ? ONE_HOUR : FIVE_MINUTES, source: 'CLAUDE_CODE_PROMPT_CACHE_TTL' }
297 const usage = await $.session.usage().catch(() => undefined)
298 // Rate-limit windows are reported to subscribers only; within plan usage the main conversation gets 1h.
299 if (usage?.rateLimits.some((r) => r.kind === 'five_hour' || r.kind === 'seven_day')) return { ms: ONE_HOUR, source: 'subscription' }
300 if ((await $.env.get('ENABLE_PROMPT_CACHING_1H')) === '1') return { ms: ONE_HOUR, source: 'ENABLE_PROMPT_CACHING_1H' }
301 // Before the first response reports rate limits: no API key or cloud provider means a Claude login.
302 const keyed = [
303 await $.env.get('ANTHROPIC_API_KEY'),
304 await $.env.get('ANTHROPIC_AUTH_TOKEN'),
305 await $.env.get('CLAUDE_CODE_USE_BEDROCK'),
306 await $.env.get('CLAUDE_CODE_USE_VERTEX'),
307 await $.env.get('CLAUDE_CODE_USE_FOUNDRY'),
308 ].some((value) => !!value)
309 return keyed ? { ms: FIVE_MINUTES, source: 'API key or cloud provider' } : { ms: ONE_HOUR, source: 'Claude login' }
310}
311
312function cancel(s: State, why: string): void {
313 s.timer?.cancel()
314 s.timer = undefined
315 s.stopped = why
316}
317
318/** Arms the next refresh (and the expiry report to herdr) after the cache was used. */
319async function schedule($: EngineInterface, s: State): Promise<void> {
320 await arm($, s)
321 await syncHerdr($, s)
322}
323
324/** Arms the next refresh at 90% of the TTL after the last cache use, inside the idle limit. */
325async function arm($: EngineInterface, s: State): Promise<void> {
326 s.timer?.cancel()
327 s.timer = undefined
328 const clock = s.clock
329 if (!clock) return cancel(s, 'no cached request yet')
330 if (!s.interactive) return cancel(s, 'not an interactive session')
331 if (!s.sessionOn || !s.settings.enabled || !s.settings.warm.enabled) return cancel(s, 'keep-warm is off')
332 const delay = warmDelayMs(s.ttl.ms)
333 if (delay === undefined) return cancel(s, 'TTL too short')
334 const due = clock.lastAt + delay
335 if (due > warmDeadline(clock.lastRealAt, s.ttl.ms, s.settings)) {
336 return cancel(s, `idle limit reached (${s.settings.warm.idleMinutes[s.ttl.ms >= ONE_HOUR ? '1h' : '5m']}m after the last request)`)
337 }
338 const now = await $.clock.now()
339 s.stopped = ''
340 s.timer = $.clock.after(Math.max(0, due - now), () => {
341 void refresh($, s)
342 })
343}
344
345/** One keep-warm request, when Pi's rule says it pays; reschedules after a hit, stops otherwise. */
346async function refresh($: EngineInterface, s: State, forced = false): Promise<void> {
347 if (!forced) s.timer = undefined
348 const clock = s.clock
349 if (!clock || s.stepsInFlight > 0 || s.warming) return
350 const now = await $.clock.now()
351 // A timer that fired late (sleep, a blocked loop) would pay a full write, not a refresh.
352 if (!forced && now >= clock.lastAt + s.ttl.ms - 1_000) return cancel(s, 'missed the refresh window (the cache expired first)')
353 const price = claudePrice(clock.model)
354 const decision = decideWarm(clock.promptTokens, price, s.ttl.ms, !s.busy, s.settings)
355 if (!forced && decision.action === 'stop') {
356 return cancel(s, `expected saving ${formatCost(decision.expectedSavings)} is under ${formatCost(s.settings.warm.minSavings)}`)
357 }
358 s.warming = true
359 try {
360 const result = await $.model.fork({ prompt: s.settings.warm.prompt })
361 const usage = 'usage' in result ? result.usage : undefined
362 const read = usage?.cache_read_input_tokens ?? 0
363 if (s.clock !== clock) return // a real request went out meanwhile and took over the clock
364 const cost = usage
365 ? (read * price.cacheRead + (usage.input_tokens + usage.cache_creation_input_tokens * 1.25) * price.input + usage.output_tokens * price.input * 5) / 1e6
366 : 0
367 if (read >= 0.8 * clock.promptTokens) {
368 s.clock = { ...clock, lastAt: now }
369 s.warms++
370 $.ui.log(`kept the prompt cache warm (${forced ? 'on request, ' : ''}${formatTokens(read)} tokens read, ~${formatCost(cost)} at API prices)`)
371 await schedule($, s)
372 } else {
373 const why = result.isAnswered || result.reason === 'empty-reply'
374 ? `the refresh read only ${formatTokens(read)} of ${formatTokens(clock.promptTokens)} tokens from the cache`
375 : `the refresh failed (${result.reason})`
376 $.ui.log(`stopped keeping the prompt cache warm: ${why}`)
377 cancel(s, why)
378 }
379 } catch (error) {
380 cancel(s, `the refresh failed (${error instanceof Error ? error.message : String(error)})`)
381 } finally {
382 s.warming = false
383 }
384}
385
386/** Why the next prompt misses the cache and what that costs, when it is worth asking about. */
387async function assess($: EngineInterface, s: State): Promise<{ line: string; costs: { send: number; compact: number } } | undefined> {
388 if (!s.settings.enabled || !s.settings.warn.enabled) return undefined
389 if (!s.clock) await readTranscript($, s, true)
390 const clock = s.clock
391 if (!clock) return undefined
392 const now = await $.clock.now()
393 const left = clock.lastAt + s.ttl.ms - now
394 if (left > 0) return undefined
395 const cost = missCost(clock.tokens, claudePrice(clock.model), s.ttl.ms)
396 if (!worthWarning(clock.tokens, cost, s.settings)) return undefined
397 return { line: describeMiss({ kind: 'expired', idleMs: -left }, clock.tokens, cost), costs: choiceCosts(clock.tokens, claudePrice(clock.model), s.ttl.ms) }
398}
399
400async function status($: EngineInterface, s: State): Promise<string> {
401 const now = await $.clock.now()
402 const clock = s.clock
403 const lines: string[] = []
404 if (!clock) {
405 lines.push('No cached request in this session yet (or it was just compacted).')
406 } else {
407 const left = clock.lastAt + s.ttl.ms - now
408 lines.push(`Last request: ${clock.model}, ${formatTokens(clock.tokens)} tokens, ${formatDuration(now - clock.lastRealAt)} ago` +
409 (clock.lastAt > clock.lastRealAt ? `; kept warm ${s.warms}x, last ${formatDuration(now - clock.lastAt)} ago.` : '.'))
410 lines.push(`TTL: ${s.ttl.ms >= ONE_HOUR ? '1h' : '5m'} (${s.ttl.source}). ${left > 0 ? `Warm, ${formatClock(left)} left.` : describeMiss({ kind: 'expired', idleMs: -left }, clock.tokens, missCost(clock.tokens, claudePrice(clock.model), s.ttl.ms))}`)
411 }
412 const on = s.sessionOn && s.settings.enabled
413 lines.push(`Keep-warm: ${on && s.settings.warm.enabled ? (s.timer ? 'next refresh scheduled' : `idle (${s.stopped || 'nothing to keep'})`) : 'off'}.`)
414 lines.push(`Warning: ${on && s.settings.warn.enabled ? `on, from ${formatCost(s.settings.warn.minCost)} at API prices` : 'off'}.`)
415 return lines.join('\n')
416}
417
418/**
419 * herdr's `cache` pane token: "cold 664k" once the cache has expired with a re-cache worth a
420 * warning, nothing while it is warm (an expiry timer re-checks then) or small. Inside herdr only.
421 */
422async function syncHerdr($: EngineInterface, s: State): Promise<void> {
423 s.expiry?.cancel()
424 s.expiry = undefined
425 const clock = s.clock
426 let value: string | undefined
427 if (clock && s.sessionOn && s.stepsInFlight === 0) {
428 const now = await $.clock.now()
429 const left = clock.lastAt + s.ttl.ms - now
430 if (left > 0) {
431 s.expiry = $.clock.after(left + 1_000, () => {
432 void syncHerdr($, s)
433 })
434 } else {
435 value = herdrCacheValue({ kind: 'expired', idleMs: -left }, clock.tokens, missCost(clock.tokens, claudePrice(clock.model), s.ttl.ms), s.settings)
436 }
437 }
438 await reportHerdr($, s, value)
439}
440
441/** One `herdr pane report-metadata` call when the token changes; any failure is ignored. */
442async function reportHerdr($: EngineInterface, s: State, value: string | undefined): Promise<void> {
443 if (value === s.herdrLast) return
444 if ((await $.env.get('HERDR_ENV')) !== '1') return
445 const pane = await $.env.get('HERDR_PANE_ID')
446 if (!pane) return
447 s.herdrLast = value
448 const change = value === undefined ? ['--clear-token', HERDR_TOKEN] : ['--token', `${HERDR_TOKEN}=${value}`, '--ttl-ms', '86400000']
449 await $.process.run(['herdr', 'pane', 'report-metadata', pane, '--source', HERDR_SOURCE, '--agent', 'claude', ...change], { timeoutMs: 3_000 }).catch(() => undefined)
450}
451
452/** /clear, then the held prompt as the new conversation's first. */
453async function startFresh($: EngineInterface, s: State, text: string): Promise<void> {
454 try {
455 await $.command.run({ command: 'clear' })
456 s.clock = undefined
457 s.timer?.cancel()
458 s.timer = undefined
459 await $.prompt.submit({ text, asUser: true })
460 } catch (error) {
461 $.ui.log(`could not start a new conversation (${error instanceof Error ? error.message : String(error)}); your prompt is back in the box`)
462 await $.prompt.fill({ text })
463 }
464}
465
466/** Compaction (with the chosen guidance), then the held prompt onto the summary. */
467async function compactThenSend($: EngineInterface, s: State, text: string, instructions: string | undefined): Promise<void> {
468 try {
469 const result = await $.session.compact(instructions ? { instructions } : {})
470 if ('skip' in result && typeof result.skip === 'string') {
471 $.ui.log(`compaction was skipped (${result.skip}); your prompt is back in the box`)
472 await $.prompt.fill({ text })
473 return
474 }
475 s.clock = undefined
476 await $.prompt.submit({ text, asUser: true })
477 } catch (error) {
478 $.ui.log(`compaction failed (${error instanceof Error ? error.message : String(error)}); your prompt is back in the box`)
479 await $.prompt.fill({ text })
480 }
481}
482hooks/core.ts 276 lines1// cache-guard core: the prompt-cache clock, the cost of a miss, and when a keep-warm request or a
2// warning pays. Shared verbatim by pi-cache-guard, claude-cache-guard, opencode-cache-guard and
3// codex-cache-guard; keep it free of harness imports so each port can copy this file as is.
4//
5// The model is Anthropic's: a cached prefix lives for its TTL from the start of the last request
6// that read or wrote it (the API guarantees that minimum and deletes soon after), and a request
7// after that re-writes the whole prefix. Prices are dollars per million tokens at API list rates;
8// on a subscription they stand for plan usage in the same proportions.
9
10export const NAME = "cache-guard";
11
12export interface Price {
13 /** Uncached input. */
14 input: number;
15 /** A cache read (hit or refresh). */
16 cacheRead: number;
17 /** A 5-minute cache write; 1.25x input when absent. A 1-hour write is always 2x input. */
18 cacheWrite?: number;
19}
20
21export interface Settings {
22 enabled: boolean;
23 warn: {
24 enabled: boolean;
25 /** Ask before a prompt whose re-cache costs at least this many dollars (when prices are known). */
26 minCost: number;
27 /** Without prices, ask when at least this many tokens would be re-sent uncached. */
28 minTokens: number;
29 /** Where a harness can only block, sending the same prompt again within this window sends it. */
30 confirmSeconds: number;
31 /** Providers that publish no TTL (OpenAI, Codex): warn after this long idle. */
32 idleMinutes: number;
33 };
34 warm: {
35 enabled: boolean;
36 /** Chance a real request arrives before the entry expires while idle (Pi's measured constant). */
37 continuationProbability: number;
38 /** A refresh is sent only when it is expected to save at least this many dollars. */
39 minSavings: number;
40 /** Stop refreshing this long after the last real request, per TTL tier. */
41 idleMinutes: { "5m": number; "1h": number };
42 /** What a keep-warm request asks, where the harness has to send a message. */
43 prompt: string;
44 };
45 herdr: {
46 /** Report the pane token `cache` to herdr (inside a herdr pane) so its agents sidebar can show doomed sessions. */
47 enabled: boolean;
48 };
49}
50
51export const DEFAULT_SETTINGS: Settings = {
52 enabled: true,
53 warn: { enabled: true, minCost: 0.5, minTokens: 100_000, confirmSeconds: 120, idleMinutes: 180 },
54 warm: {
55 enabled: true,
56 continuationProbability: 0.15,
57 minSavings: 0.05,
58 idleMinutes: { "5m": 30, "1h": 120 },
59 prompt: "Cache keep-alive. Do not use tools or think. Reply with exactly: ok",
60 },
61 herdr: { enabled: true },
62};
63
64/** The herdr pane token the ports report (`herdr pane report-metadata --source cache-guard`). */
65export const HERDR_TOKEN = "cache";
66export const HERDR_SOURCE = NAME;
67
68export const FIVE_MINUTES = 5 * 60_000;
69export const ONE_HOUR = 60 * 60_000;
70
71/** "5m" or "1h": the tier a TTL bills as (anything from an hour up writes at 2x). */
72export function tier(ttlMs: number): "5m" | "1h" {
73 return ttlMs >= ONE_HOUR ? "1h" : "5m";
74}
75
76/** Refresh at 90% of the TTL, keeping at least ten seconds of margin (Pi's rule). */
77export function warmDelayMs(ttlMs: number): number | undefined {
78 if (ttlMs <= 10_000) return undefined;
79 return Math.max(1, Math.floor(Math.min(ttlMs * 0.9, ttlMs - 10_000)));
80}
81
82/** Milliseconds the entry written or refreshed at `lastAt` has left; 0 once expired. */
83export function remainingMs(lastAt: number, ttlMs: number, now: number): number {
84 return Math.max(0, lastAt + ttlMs - now);
85}
86
87export function writePrice(price: Price, ttlMs: number): number {
88 return tier(ttlMs) === "1h" ? price.input * 2 : (price.cacheWrite ?? price.input * 1.25);
89}
90
91/** What a miss costs over a hit: the prefix written again instead of read. */
92export function missCost(tokens: number, price: Price, ttlMs: number): number {
93 return Math.max(0, (tokens * (writePrice(price, ttlMs) - price.cacheRead)) / 1e6);
94}
95
96/** What one refresh costs: the prefix read (plus a token or two, ignored). */
97export function warmCost(tokens: number, price: Price): number {
98 return (tokens * price.cacheRead) / 1e6;
99}
100
101export interface WarmDecision {
102 action: "warm" | "stop";
103 warmCost: number;
104 missCost: number;
105 expectedSavings: number;
106}
107
108/**
109 * Pi's rule: refresh when `p * missCost - warmCost` is at least `minSavings`, with p = 1 while the
110 * agent is still running (its next request is certain) and the idle constant otherwise.
111 */
112export function decideWarm(tokens: number, price: Price, ttlMs: number, idle: boolean, settings: Settings): WarmDecision {
113 const miss = missCost(tokens, price, ttlMs);
114 const warm = warmCost(tokens, price);
115 const p = idle ? settings.warm.continuationProbability : 1;
116 const expectedSavings = p * miss - warm;
117 return { action: expectedSavings >= settings.warm.minSavings ? "warm" : "stop", warmCost: warm, missCost: miss, expectedSavings };
118}
119
120/** The last moment a refresh may be sent for a cache last used by a real request at `lastRealAt`. */
121export function warmDeadline(lastRealAt: number, ttlMs: number, settings: Settings): number {
122 return lastRealAt + settings.warm.idleMinutes[tier(ttlMs)] * 60_000;
123}
124
125/** Whether a re-cache of `tokens` (costing `cost`, when prices are known) is worth asking about. */
126export function worthWarning(tokens: number, cost: number | undefined, settings: Settings): boolean {
127 if (!settings.enabled || !settings.warn.enabled) return false;
128 return cost !== undefined ? cost >= settings.warn.minCost : tokens >= settings.warn.minTokens;
129}
130
131export type ColdReason =
132 | { kind: "expired"; idleMs: number }
133 | { kind: "model"; from: string; to: string }
134 | { kind: "idle"; idleMs: number };
135
136/**
137 * The herdr `cache` token: `cold 664k` (or `cold? 180k` when only idle time suggests it) while the
138 * next prompt would re-cache at least the warning threshold, else undefined (clear the token).
139 * Warm and small caches report nothing, so the sidebar lists only the doomed sessions.
140 */
141export function herdrCacheValue(reason: ColdReason | undefined, tokens: number, cost: number | undefined, settings: Settings): string | undefined {
142 if (!reason || !settings.enabled || !settings.herdr.enabled) return undefined;
143 const big = cost !== undefined ? cost >= settings.warn.minCost : tokens >= settings.warn.minTokens;
144 if (!big) return undefined;
145 return `${reason.kind === "idle" ? "cold?" : "cold"} ${formatTokens(tokens)}`;
146}
147
148/**
149 * Rough costs, in dollars at list prices, of the two ways through a cold cache that keep the
150 * history: send the prompt and write the whole prefix to the cache again, or compact first, which
151 * reads it once uncached (plus a summary, not counted) and continues on a small context. Starting
152 * fresh costs about nothing.
153 */
154export function choiceCosts(tokens: number, price: Price, ttlMs: number): { send: number; compact: number } {
155 return { send: (tokens * writePrice(price, ttlMs)) / 1e6, compact: (tokens * price.input) / 1e6 };
156}
157
158/** Compaction guidance that keeps what the held prompt needs. */
159export function compactionFocus(prompt: string): string {
160 return `Keep what is needed to continue with the user's next request, quoted below, and drop the rest.\n\n${prompt.trim()}`;
161}
162
163/** One line saying why the next request misses and what that costs. */
164export function describeMiss(reason: ColdReason, tokens: number, cost: number | undefined): string {
165 const amount = `${formatTokens(tokens)} tokens${cost === undefined ? "" : ` (~${formatCost(cost)} at API prices)`}`;
166 switch (reason.kind) {
167 case "expired":
168 return `The prompt cache expired ${formatDuration(reason.idleMs)} ago: this prompt re-caches ${amount}.`;
169 case "idle":
170 return `Idle ${formatDuration(reason.idleMs)}: the prompt cache has probably expired, so this prompt may re-cache ${amount}.`;
171 case "model":
172 return `${reason.to} has no cache of this conversation (it was cached for ${reason.from}): this prompt re-caches ${amount}.`;
173 }
174}
175
176/** Remembers a blocked prompt so the same prompt sent again within the window goes through. */
177export class ConfirmMemo {
178 private pending?: { key: string; text: string; at: number };
179
180 /** True when `text` repeats the prompt blocked for `key` within `windowMs`; clears it either way. */
181 confirmed(key: string, text: string, now: number, windowMs: number): boolean {
182 const pending = this.pending;
183 this.pending = undefined;
184 return pending !== undefined && pending.key === key && pending.text === text.trim() && now - pending.at <= windowMs;
185 }
186
187 arm(key: string, text: string, now: number): void {
188 this.pending = { key, text: text.trim(), at: now };
189 }
190}
191
192/** "4:05" under an hour, "1h05m" above, "0:00" once gone. */
193export function formatClock(ms: number): string {
194 const seconds = Math.max(0, Math.ceil(ms / 1000));
195 if (seconds >= 3600) return `${Math.floor(seconds / 3600)}h${String(Math.floor((seconds % 3600) / 60)).padStart(2, "0")}m`;
196 return `${Math.floor(seconds / 60)}:${String(seconds % 60).padStart(2, "0")}`;
197}
198
199/** "45s", "12m", "3h20m", "2d4h". */
200export function formatDuration(ms: number): string {
201 const seconds = Math.max(0, Math.round(ms / 1000));
202 if (seconds < 60) return `${seconds}s`;
203 const minutes = Math.floor(seconds / 60);
204 if (minutes < 60) return `${minutes}m`;
205 const hours = Math.floor(minutes / 60);
206 if (hours < 48) return `${hours}h${minutes % 60 ? `${minutes % 60}m` : ""}`;
207 return `${Math.floor(hours / 24)}d${hours % 24 ? `${hours % 24}h` : ""}`;
208}
209
210export function formatTokens(n: number): string {
211 if (n >= 1e6) return `${(n / 1e6).toFixed(n >= 1e7 ? 0 : 1)}M`;
212 if (n >= 1e3) return `${Math.round(n / 1e3)}k`;
213 return String(Math.round(n));
214}
215
216export function formatCost(dollars: number): string {
217 return dollars >= 10 ? `$${dollars.toFixed(0)}` : `$${dollars.toFixed(2)}`;
218}
219
220/**
221 * List prices of the Claude models, for harnesses that do not carry a model catalog (Claude Code,
222 * Codex has no use for it). Longest prefix wins; unknown models price as Opus 5.5.
223 */
224const CLAUDE_PRICES: ReadonlyArray<readonly [string, Price]> = [
225 ["claude-fable-5-1", { input: 10, cacheRead: 0.25 }],
226 ["claude-fable-5", { input: 10, cacheRead: 1 }],
227 ["claude-mythos-5-1", { input: 10, cacheRead: 0.25 }],
228 ["claude-opus-5-5", { input: 4, cacheRead: 0.2 }],
229 ["claude-opus-5", { input: 5, cacheRead: 0.5 }],
230 ["claude-opus-4", { input: 5, cacheRead: 0.5 }],
231 ["claude-sonnet-5", { input: 2, cacheRead: 0.2 }],
232 ["claude-sonnet-4", { input: 3, cacheRead: 0.3 }],
233 ["claude-haiku-5", { input: 0.1, cacheRead: 0.01 }],
234 ["claude-haiku-4", { input: 1, cacheRead: 0.1 }],
235];
236
237export function claudePrice(model: string): Price {
238 const id = model.toLowerCase().replace(/^.*\//, "").replace(/\[.*$/, "");
239 let best: Price | undefined;
240 let length = 0;
241 for (const [prefix, price] of CLAUDE_PRICES) {
242 if (id.startsWith(prefix) && prefix.length > length) {
243 best = price;
244 length = prefix.length;
245 }
246 }
247 return best ?? { input: 4, cacheRead: 0.2 };
248}
249
250export function mergeSettings(base: Settings, texts: readonly (string | undefined)[]): Settings {
251 let merged: Record<string, unknown> = structuredClone(base) as unknown as Record<string, unknown>;
252 for (const text of texts) {
253 if (text === undefined) continue;
254 try {
255 const value: unknown = JSON.parse(text);
256 if (isRecord(value)) merged = merge(merged, value);
257 } catch {
258 // Invalid JSON: keep what the earlier files said.
259 }
260 }
261 return merged as unknown as Settings;
262}
263
264function merge(base: Record<string, unknown>, over: Record<string, unknown>): Record<string, unknown> {
265 const out: Record<string, unknown> = { ...base };
266 for (const [key, value] of Object.entries(over)) {
267 const current = out[key];
268 out[key] = isRecord(current) && isRecord(value) ? merge(current, value) : value;
269 }
270 return out;
271}
272
273function isRecord(value: unknown): value is Record<string, unknown> {
274 return typeof value === "object" && value !== null && !Array.isArray(value);
275}
276hooks/settings.ts 17 lines1// Settings shared with the pi, opencode and Codex ports: `~/.config/agents/cache-guard.json` (or
2// under $XDG_CONFIG_HOME) and `<project>/.agents/cache-guard.json`, with Claude Code's own
3// `~/.claude/cache-guard.json` and `<project>/.claude/cache-guard.json` as overrides. Later files
4// win; objects merge, other values replace.
5import { NAME } from './core'
6
7/** The settings files, lowest precedence first. */
8export function settingsFiles(home: string, xdgConfigHome: string | undefined, root: string): string[] {
9 const shared = xdgConfigHome || `${home}/.config`
10 return [
11 `${shared}/agents/${NAME}.json`,
12 `${home}/.claude/${NAME}.json`,
13 `${root}/.agents/${NAME}.json`,
14 `${root}/.claude/${NAME}.json`,
15 ]
16}
17hooks/transcript.ts 74 lines1// Reads the tail of a Claude Code transcript (~/.claude/projects/<slug>/<session>.jsonl): the last
2// main-conversation response, and which TTL the session's cache writes use. The usage rows carry
3// `cache_creation.ephemeral_1h_input_tokens` / `ephemeral_5m_input_tokens`, which nothing in the
4// mod API reports.
5
6export interface TranscriptTail {
7 /** When the last main-thread response was written (its end, so later than the request's start). */
8 at?: number;
9 /** Its prompt: input + cache reads + cache writes. */
10 promptTokens?: number;
11 /** Prompt plus the reply: what the next request re-sends. */
12 tokens?: number;
13 model?: string;
14 /** The tier of the most recent main-thread cache write in the tail. */
15 ttlMs?: number;
16 /** A compaction after the last response: the next request carries a fresh context. */
17 compacted?: boolean;
18}
19
20interface Usage {
21 input_tokens?: number;
22 output_tokens?: number;
23 cache_read_input_tokens?: number;
24 cache_creation_input_tokens?: number;
25 cache_creation?: { ephemeral_1h_input_tokens?: number; ephemeral_5m_input_tokens?: number };
26}
27
28/** The transcript file Claude Code writes for `sessionId` started in `root`. */
29export function transcriptPath(configDir: string, root: string, sessionId: string): string {
30 return `${configDir}/projects/${root.replace(/[^a-zA-Z0-9]/g, '-')}/${sessionId}.jsonl`
31}
32
33/** Parses JSONL text (possibly starting mid-line) from the end back. */
34export function parseTail(text: string): TranscriptTail {
35 const out: TranscriptTail = {}
36 const lines = text.split('\n')
37 for (let i = lines.length - 1; i >= 0; i--) {
38 const line = lines[i]
39 if (!line || line[0] !== '{') continue
40 let row: Record<string, any>
41 try {
42 row = JSON.parse(line)
43 } catch {
44 continue // the first, cut line
45 }
46 if (out.at === undefined && row.type === 'system' && row.subtype === 'compact_boundary') {
47 out.compacted = true
48 continue
49 }
50 if (row.type !== 'assistant' || row.isSidechain === true) continue
51 const message = row.message ?? {}
52 if (message.model === '<synthetic>' || row.isApiErrorMessage === true) continue
53 const usage: Usage | undefined = message.usage
54 if (!usage) continue
55 const prompt = (usage.input_tokens ?? 0) + (usage.cache_read_input_tokens ?? 0) + (usage.cache_creation_input_tokens ?? 0)
56 if (prompt <= 0) continue
57 if (out.at === undefined) {
58 const at = Date.parse(row.timestamp)
59 if (!Number.isFinite(at)) continue
60 out.at = at
61 out.promptTokens = prompt
62 out.tokens = prompt + (usage.output_tokens ?? 0)
63 out.model = String(message.model ?? '')
64 }
65 const oneHour = usage.cache_creation?.ephemeral_1h_input_tokens ?? 0
66 const fiveMinutes = usage.cache_creation?.ephemeral_5m_input_tokens ?? 0
67 if (oneHour > 0 || fiveMinutes > 0) {
68 out.ttlMs = oneHour >= fiveMinutes ? 3_600_000 : 300_000
69 break
70 }
71 }
72 return out
73}
74