SLOPSHOPPER

cache-warm

Keeps the 1-hour prompt cache warm for a window you set, or with no end in every session under always, with one cache-shared fork per idle stretch and one…

newcommandstatuspromptmodeltimer
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · cache-warm
› fix the failing auth test and add an audit log call ● cache-warm: the last request time of this session was not read: /Users/dev/.claude/projects/-work-app/preview-session.jsonl does not exist ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /cache-warm ⎿ cache-warm: on for 6h, a ping 50m after each idle stretch keeps the cache read, not re-written ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⚠ cache-warm: 6h left · ping in 50m · last turn read 91k $0.09 (01:53)
README

cache-warm

Claude Code keeps your conversation in a prompt cache for one hour. Step away for longer, and the next message re-writes the whole context at the cache-write rate: on Claude Fable 5.1 that is $4.00 for 200k tokens, against $0.05 to read the same tokens from a warm cache. This mod keeps the cache warm for a window you choose, and shows you what a cold cache would cost.

It follows the behaviour of the cache-tax mod by Karan Bansal (karanb192/claude-code-mods), without its send guard. The code is new.

What it does

Keeps the cache warm. /cache-warm arms a six-hour window. Inside it, 50 minutes after the main loop's last model request, the mod sends one tool-less $.model.fork over the session's own transcript. The server answers it from the cache, and that refreshes the hour. Every new request pushes the ping later, so a session you are actively using sends no ping at all. The first request of a session, and the first after /clear or a compaction, sets the ping as well, so even a tool call that runs past the hour inside the first turn gets its ping in the middle of the turn (measured on 2.1.281: the ping went out while sleep 110 ran, and the turn ended normally).

Runs with no end under always. /cache-warm always is not a window: the ping goes out every 50 minutes for as long as the session lives, and only /cache-warm off ends it. The switch is one global key in the mod's own $.store, so every later session of every project starts the same loop at its start and after /clear.

  • Each turn's end keeps the last request's time in $.store under the session's id, and each ping keeps its own read there too. A module loaded into a running conversation (/reload-plugins, an update) therefore pings on time from the later of the two; a session kept warm by pings alone has a last turn older than its cache. When both are more than an hour old, the cache is gone, and the loop waits for the next turn instead of paying for a cold ping.
  • The session's transcript is not used for this, because /reload-plugins writes a line of its own there and the file's last write would then look like a request (measured: a reload 90 seconds after the last turn would have set the ping 90 seconds late). Only a session with no time kept yet, one that ran an older version, reads the last write of its transcript (~/.claude/projects/<directory>/<session id>.jsonl, under CLAUDE_CONFIG_DIR when it is set) once. The transcript is looked up under the directory the session started in, never the one a shell cd moved to, and one it cannot find is named in a transcript line.
  • A ping that finds the cache gone does not end this loop. The write that ping paid for is the new cache: the mod says so in a transcript line, counts the write in the session's tally and keeps going. A warm ping reads the context at the read rate, about $0.05 for 200k tokens, so an idle day of pings costs about $1.40.

Stops when the cache is gone. This applies to a window with an end, not to always. A warm ping reads the context and writes only its own few tokens. When a ping reads nothing, or writes a tenth of what it read or more, the cache was already gone and the ping itself paid for the write, so the mod stops and shows why. It also stops when the API answers the fork with an error (the line names its status and kind), and when the fork is cut before it replies. When the engine has nothing to fork, as in a resumed process before its first reply, the window does not stop: it waits for the next reply, whose turn arms the ping again, and the line says until when the cache holds. A reply without text still read the cache, so it counts as a ping. Under always such a failure stops the loop for that turn only: the next turn starts it again, so the session never holds the switch while nothing runs.

Arms itself after a paid cold write. When a turn re-writes at least half of a context larger than 20k tokens, the mod counts that cold write and arms a six-hour window, unless a longer one is already armed. Under always no six-hour window is armed, because the endless loop already keeps that cache.

Shows the state. /cache-status prints the model, warm or cold, the context size, the cold price, the window, the break-even and this session's cold writes.

A message you send to a cold cache is never stopped or delayed. A resumed session whose cache has lapsed gets one line with the price of its first message. Claude Code dates the cache from the transcript's last reply, which a ping never writes, so that line is left out while the last ping the mod kept for the session read the cache within the hour.

Sends a keep-warm message after a resume. A resumed process cannot fork before its own first reply: $.model.fork answers nothing-to-fork (measured with a headless claude --resume). So no ping can go out, and a session you closed and opened again 30 minutes later would lose its cache at the hour unless you wrote something. When an interactive session is resumed while a window or always runs, its context is 50k tokens or more and its cache still holds, the mod therefore sends one message three seconds after the resume, through its own /cache-warm:send command:

/cache-warm:send This message was sent by the cache-warm plugin, not by the person. The session was resumed, and a resumed session can keep its prompt cache warm only after a reply. Do not run a tool or continue a task. Reply with the single word: warm

This is a real turn: it reads the cache, the model answers one word, the pair stays in the conversation, and its end arms the ping again. Nothing is sent when the cache is already gone (your next message pays the same write anyway), in a -p run, or when you sent a message within those three seconds. Other mods see it like any other turn: task-poke may send its continue prompt after it while tasks are open, and desk-notify shows its turn-end notification. Measured on 2.1.283 in a resumed interactive session of 116k tokens: the message went out at the resume, the model answered warm, and the next ping forked and read 117k tokens. A resume can break part of the prefix on its own (that session re-wrote 42k of the 116k; another, resumed 40 minutes after its last ping, 4k); the keep-warm turn pays that write at the resume instead of your first message. The three seconds count from the later of the two start hooks, classic.SessionStart and session.start, because they settle in no fixed order: with every mod of the marketplace loaded, session.start settled four seconds after classic.SessionStart (measured on 2.1.285), and a wait started by the first alone found a session that had not started and sent nothing.

Commands

/cache-warm keep warm for six hours /cache-warm 90m keep warm for a window of your own (also 2h30m) /cache-warm always keep the cache warm with no end, in every session of every project /cache-warm 6h every 2m ping every two minutes; a test setting, floor 1m, forgotten after this window /cache-warm status the status line text /cache-warm off stop, forget the window, and turn always off /cache-status the card /cache-warm:send <text> the keep-warm message the mod sends after a resume; its body is the text alone

What it shows

A status line under the prompt while a window is armed or after a stop:

cache-warm: 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42) cache-warm: stopped: the ping read 0 and wrote 180k tokens ($3.60), the cache was already gone

The last part names the last request that read the cache, a ping or a main-loop turn, whichever came later: last ping read 200k $0.05 or last turn read 250k $0.07, never both. A turn's figures cover the whole turn, every request summed. The time in brackets is when that answer came, in local time; one from an earlier day carries its day and month, as (22 Sep 23:10). The record is kept in $.store under the session's id, so /reload-plugins or an update shows it again at once.

While a window runs, an interactive session redraws the line every minute, so the time left and the time to the next ping count down between turns and pings, and the last part stays on the line. The last minute before a ping reads ping now, because minutes are rounded and the line is drawn once a minute; while the ping's fork is out it reads pinging…. When the engine has nothing to fork, the line says so instead of counting down:

cache-warm: always · no ping before the next reply · cache holds until 19:16

One stream entry per ping attempt while the sidebar is open. It is also kept in the sidebar's log file (~/.claude/sidebar/<project>-<date>.log), so you can read back later whether a ping went out and what it did; with the sidebar closed the same text is a transcript line:

ping sent · read 901k · wrote 0 · $0.45 ping found the cache gone · read 0 · wrote 180k · $3.60 ping not sent: the conversation has no reply to fork yet; the ping waits for the next reply, and the cache holds until 19:16 ping failed: the ping failed, the API answered 529 (overloaded) keep-warm message sent: the session was resumed and its cache holds until 19:16; a resumed session pings only after a reply

With the sidebar open, the status line moves there as a cache window section that stays for the session and is rewritten at every change, and the status line stays clear. Only the time left (or always) is coloured: yellow when the window ends sooner than one ping period, green while it holds, faint while it waits for the first turn. The ping details after it are faint, and a stopped window shows its stopped: front in red with the reason in the default colour. Without the sidebar, the status line is drawn as above.

A stop reason stays for one turn. At the next turn the section shows the idle line instead: faint, except for a paid N cold writes paid $X, which is yellow. That way the pane shows a measurement of now, not the last sentence of a window that ended. The reason stays in the transcript, and the status line is empty while no window runs:

cache window off · 2 cold writes paid $6.30 · context 315k tokens

A window that runs out of time is armed again by your next message, as long as the one that ended and with the same ping period, and the idle line says so while it waits:

cache window off · 6h again at your next message · 1 cold write paid $4.01 · context 201k tokens

cache-warm: the 6h window ran out; this message arms another one. /cache-warm off stops it.

Only a window that ran out of time comes back. One that a ping stopped does not: the cache is already gone there, and the cold write of your next message arms its own 6h window. /cache-warm off forgets a window waiting to come back.

Under the window the section holds a second, faint line: the last transcript line, shortened, with the cost of a cold write in yellow. The window line says how long the cache is kept; the second line says what the mod did last:

cache window 6h left · ping in 50m cold write 201k tokens paid ($4.01)

The card of /cache-status:

claude-fable-5-1 state warm, 42m left context 200,502 tokens cold cost $4.01 to re-write it (warm turn $0.05) keep warm on, 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42) break-even up to 80 pings at the read rate cost one cold write, about 2d 18h of idle at one ping per 50m session 1 cold write paid, $4.01

One transcript line, not sent to the model, when a cold write arms the window or a resume starts cold. While the watcher is off (no window, always off), the line goes to the sidebar's stream instead and writes nothing to the transcript; a closed sidebar drops it. The line a running watcher writes is unchanged.

Prices

The table in hooks/pricing.ts holds the cache-read, 1-hour cache-write and output rates of every model on the Anthropic pricing page, read in September 2026. A model id takes the first family it contains, so claude-opus-4-1 is priced as Opus 4.1 ($1.50 / $30 / $75) and claude-opus-4-8 as Opus 4.8 ($0.50 / $10 / $25). claude-opus-5-5 also contains opus-5, so its own row comes first: Opus 5.5 is $0.20 / $8 / $20, cheaper than Opus 5. Sonnet 5.5 has its own row at the Sonnet 5 rates, $0.20 / $4 / $10. A ping is priced in full: the cache read, its cache write, its uncached input at the base rate (half the 1-hour write rate) and its output. An unknown model shows n/a.

Fast mode bills Opus 5.5, Opus 5 and Opus 4.8 at their own base rates ($8 and $10 input), with the cache multipliers applied on top. The mod prices Opus 5.5 at $0.40 / $16 / $40 and Opus 5 and 4.8 at $1 / $20 / $50 while the fastMode setting, which /fast writes, is on. It reads the settings at the session's start and at the end of each main-loop turn, so a /fast counts from the next turn. With fastModePerSessionOptIn set to true, every session starts with fast mode off, so the standard rates apply. Any other model keeps its standard rates, and the card mentions fast mode only when the rates changed:

claude-opus-5-5 · fast mode rates (the fastMode setting)

On a subscription the dollars are a yardstick, not your bill. How a cache read counts against the 5-hour and weekly limits is not documented.

Install

claude plugin marketplace add KilimcininKorOglu/claude-code-mods claude plugin install cache-warm@kilimcininkoroglu-mods

Function hooks are early access. Claude Code 2.1.288 and later load them by default, so there is nothing to switch on.

To load it from a local checkout for one session:

claude --plugin-dir plugins/cache-warm

After installing

  1. Restart Claude Code.
  2. Disable every other keep-warm mod, for example claude plugin disable cache-tax@claude-code-mods. Two keep-warm mods in one session send two pings per idle stretch.
  3. Check once that a ping reads your cache, as "Prove it on your own session" below describes.
  4. To keep the cache warm with no end, run /cache-warm always once. The switch is global: every later session of every project starts the loop by itself, and /cache-warm off ends it for good. Without it, a window is armed only by /cache-warm or after a paid cold write.

What it can reach

Validated with claude plugin validate on Claude Code 2.1.288:

❯ ./register.ts hooks: session.start, classic.SessionStart, prompt.submit, command.run{command=cache-warm}, command.run{command=cache-status}, turn.step, turn.complete, session.compact ❯ ./register.ts calls: $.clock.after (via arm, scheduleKeepWarm), $.clock.every, $.clock.now, $.command.register (via registerCommands), $.command.run (via keepWarmAfterResume), $.env.get (via seedFromTranscript), $.fs.exists (via seedFromTranscript), $.fs.stat (via seedFromTranscript), $.model.fork (via forkPing), $.prompt.submit (via keepWarmAfterResume), $.session.id, $.session.model, $.session.root (via seedFromTranscript), $.session.usage, $.settings.read (via readFast), $.sidebar.set (via logEvent, toSidebar, toStream), $.store.delete (via prune, pruneRequests, startEndless, startWindow, stop), $.store.get, $.store.keys (via prune, pruneRequests), $.store.set (via afterTurn, keepLastRead, startWindow, warmCommand), $.ui.log (via logEvent, seedFromTranscript, toStream), $.ui.status (via showStatusAt) ❯ ./register.ts env writes: nothing ❯ ./register.ts env reads: CLAUDE_CONFIG_DIR, HOME

Reach L2: it drives Claude.

  1. Reads: the time of each main-loop model request; the token counts and model id of each turn and of each ping; the live context size; the origin of each message, to arm a window again; the resume fields Claude Code computes for settings hooks; the session id and model; the last write time of the session's own transcript file, once when the module loads into a running conversation; the fastMode and fastModePerSessionOptIn settings, at the session's start and at each turn's end; its own $.store. It never reads a prompt's text, a file's content or a tool result.
  2. Runs: one $.model.fork per idle stretch while a window or the always loop runs, 50 minutes after the last request unless the test setting is used (floor 1 minute); never while off; a ping that found the cache gone ends a window with an end, and under always the loop carries on; after a resume of an interactive session whose cache still holds, one keep-warm message through /cache-warm:send (a plugin prompt when the engine refuses the command), which is a real turn
  3. Sends: the fork, an API request over the session's own transcript with a fixed one-line prompt, and after a resume the fixed keep-warm message as a turn of the conversation
  4. Persists: in $.store, the window end, the ping period and the last main-loop request's time and the last ping or turn read (tokens, cost, time) under this session's id, and the global always switch, which the endless loop needs no window key beside; this session's ended window is deleted at stop and at its next start, another session's window one week after it ended, another session's request time and last read once they are an hour old; the cold-write tally lives in memory and ends with the session
  5. Hostile input: the only text it parses is the argument of /cache-warm, matched against a duration pattern and three words; the fork's prompt is a constant, so nothing crafted can reach it

Prove it on your own session

The mock-clock tests prove the timer and the scoring, not that a fork reads the main cache. One ping proves that. In a warm session:

Reply with one word: ready /cache-warm 1h every 1m

After a minute the status line should read last ping read <close to your context> $.... A stopped: the ping read ... line means the fork did not share the cache, and the mod has already stopped. /cache-warm off ends the test.

Limits

  • The 50-minute ping assumes the 1-hour cache tier, which the main conversation uses.
  • A warm ping only proves the cache was warm at that moment. A model or effort switch, an edited CLAUDE.md or a changed tool list breaks the prefix no matter the time, and your next message pays.
  • A ping's output cannot be capped; a model at high effort may think before it answers. The status line prices what the ping really billed.
  • The resume logic is covered by hook tests that raise classic.SessionStart with the resume fields, and by a live check of a resumed interactive session in tmux. Whether a resume keeps the whole prefix is out of the mod's hands: it re-wrote 42k of 116k tokens in one measured session and 4k in another.
  • The cold-write tally is per session and lives in memory. /clear empties it.
  • Fast mode is read from the fastMode setting, the saved preference, not from the request. The mod does not see Claude Code fall back to standard speed within a session (a fast mode rate-limit cooldown, usage credits that ran out, an organization that turned fast mode off). Those turns bill standard rates while the mod prices them as fast.
  • Whether a ping, a $.model.fork, runs at fast speed while the session does has not been measured; the mod prices it at the session's rates.

Development

make install # eslint, typescript-eslint, typescript make lint # complexity limit 10, the build fails above it make typecheck # needs .claude/types/ from /plugin-types make validate make test # claude plugin test

Source 3 files
hooks/register.ts 665 lines
1import type { EngineInterface, ModelForkResult, Register, TurnUsage } from 'claude-code'
2import { responseUsd, writeUsd, type Usage } from './pricing.ts'
3import {
4  AUTO_WARM_MS,
5  KEEP_WARM_TEXT,
6  MIN_PING_MS,
7  PING_AFTER_MS,
8  card,
9  coldPingText,
10  coldWriteShort,
11  eventShort,
12  fmtDuration,
13  fmtTok,
14  fmtUsd,
15  freshState,
16  hasWindow,
17  idleText,
18  isColdWrite,
19  isOver,
20  isWarmPing,
21  keepWarmLine,
22  noReplyText,
23  parseWarmArgs,
24  pingLine,
25  pingRecordOf,
26  priceNow,
27  resetForClear,
28  seedFromResume,
29  statusText,
30  transcriptPath,
31  TTL_MS,
32  unsentText,
33  wantsKeepWarm,
34  windowLine,
35  type Line,
36  type PingRecord,
37  type State,
38} from './warm.ts'
39
40const PING_PROMPT = 'Reply with the single word: warm'
41/** The mod's own markdown command, whose body is its arguments alone, so the model reads the text as it is written. */
42const SEND_COMMAND = 'cache-warm:send'
43/** How long after a resume the keep-warm message waits, so the session is idle and a quick prompt of the person's goes first. */
44const KEEP_WARM_AFTER_MS = 3000
45const KEY_ALWAYS = 'always'
46const DEADLINE = 'deadline:'
47const EVERY = 'every:'
48const REQUEST = 'request:'
49const LAST = 'last:'
50
51/** How often a running window's line is drawn again, so its minutes count down between turns and pings. */
52const REDRAW_MS = 60_000
53
54// The window and its ping period belong to the session that armed them, so a
55// second session never inherits them and cannot turn them off. The always
56// switch is global.
57function deadlineKey(s: State): string {
58  return DEADLINE + s.sid
59}
60
61function everyKey(s: State): string {
62  return EVERY + s.sid
63}
64
65/** The session's last main-loop request, kept for a module loaded into the running conversation. */
66function requestKey(s: State): string {
67  return REQUEST + s.sid
68}
69
70/** The session's last ping or turn read, kept so a reloaded module draws it again. */
71function lastKey(s: State): string {
72  return LAST + s.sid
73}
74
75/** Records the last request that read the cache, in memory and in the store. */
76async function keepLastRead($: EngineInterface, s: State, record: PingRecord): Promise<void> {
77  s.lastRead = record
78  await $.store.set(lastKey(s), record)
79}
80
81function errorText(err: unknown): string {
82  return err instanceof Error ? err.message : String(err)
83}
84
85function disarm(s: State): void {
86  s.pending?.cancel()
87  s.pending = null
88}
89
90/** The section this mod owns in the shared sidebar. */
91const SECTION = { consumer: 'cache-warm', key: 'window' }
92
93/**
94 * Writes one event line. While the watcher runs (a window or always) it goes to the transcript as
95 * before. While it is off, the sidebar's stream takes it instead and a closed sidebar drops it, so
96 * an off watcher writes nothing to the transcript; the short form still stands under the section.
97 */
98async function logEvent($: EngineInterface, s: State, text: string, short: Line): Promise<void> {
99  s.event = short
100  if (hasWindow(s) || (await readAlways($, s))) {
101    $.ui.log(text)
102    return
103  }
104  try {
105    if (await $.sidebar.set({ ...SECTION, key: 'event', title: 'event', lines: [{ text }], until: 'stream' })) return
106  } catch {
107    // The sidebar mod is not installed; the line is dropped while the watcher is off.
108  }
109}
110
111/**
112 * One entry in the sidebar's stream, which the sidebar keeps in its log file, so every ping attempt can be
113 * read back later; the transcript line while the pane is closed or the sidebar mod is missing.
114 */
115async function toStream($: EngineInterface, line: Line): Promise<void> {
116  try {
117    if (await $.sidebar.set({ ...SECTION, key: 'ping', title: 'ping', lines: [line], until: 'stream' })) return
118  } catch {
119    // The sidebar mod is not installed; the log line carries the entry.
120  }
121  $.ui.log(line.text)
122}
123
124/**
125 * Writes the window's state into the shared sidebar and answers whether it took it. The section also
126 * carries the last transcript line under the state, faint: the first line says how long the window
127 * holds, the second what the mod last did. A closed sidebar, and a sidebar mod that is not installed,
128 * both answer false, so the status line is drawn instead.
129 */
130async function toSidebar($: EngineInterface, s: State, line: Line): Promise<boolean> {
131  try {
132    const lines = [line, ...(s.event === undefined ? [] : [s.event])]
133    return await $.sidebar.set({ ...SECTION, title: 'cache window', lines, until: 'session', order: 20 })
134  } catch {
135    // The sidebar mod is not installed.
136    return false
137  }
138}
139
140/**
141 * Draws the window's state. The sidebar keeps a line either way: the window's own line, or the faint
142 * idle line while none runs, so the section is never a snapshot of a window that ended. The status line
143 * is unchanged: with no window and no stop reason it carries nothing.
144 */
145async function showStatusAt($: EngineInterface, s: State, now: number): Promise<void> {
146  const taken = await toSidebar($, s, windowLine(s, now))
147  $.ui.status(taken ? undefined : statusText(s, now))
148}
149
150async function showStatus($: EngineInterface, s: State): Promise<void> {
151  await showStatusAt($, s, await $.clock.now())
152}
153
154/**
155 * Deletes every session's ended window. A resumed session reads an ended window
156 * as no window (`restore`), so keeping one only grows the store; a window still
157 * running is left alone.
158 */
159async function prune($: EngineInterface, now: number): Promise<void> {
160  for (const key of await $.store.keys()) {
161    if (!key.startsWith(DEADLINE)) continue
162    const deadline = await $.store.get(key)
163    if (typeof deadline === 'number' && deadline > now) continue
164    await $.store.delete(key)
165    await $.store.delete(EVERY + key.slice(DEADLINE.length))
166  }
167}
168
169/** The time a stored request or read record was made, or undefined for a value of another shape. */
170function storedAt(key: string, value: unknown): number | undefined {
171  if (key.startsWith(REQUEST)) return typeof value === 'number' ? value : undefined
172  return pingRecordOf(value)?.at
173}
174
175/**
176 * Deletes every other session's last request time and last read that are older than the cache, which is
177 * gone by then. This session's own stay: the request's age is what tells the seed the cache is gone, and
178 * the next turn rewrites both.
179 */
180async function pruneRequests($: EngineInterface, s: State, now: number): Promise<void> {
181  for (const key of await $.store.keys()) {
182    if (!(key.startsWith(REQUEST) || key.startsWith(LAST)) || key === requestKey(s) || key === lastKey(s)) continue
183    const at = storedAt(key, await $.store.get(key))
184    if (at === undefined || now - at >= TTL_MS) await $.store.delete(key)
185  }
186}
187
188/**
189 * Reads the always switch from the store and answers it. Every session of every project shares the
190 * switch, so it is read where it acts, and /cache-warm always or off in another session applies here.
191 */
192async function readAlways($: EngineInterface, s: State): Promise<boolean> {
193  s.always = (await $.store.get(KEY_ALWAYS)) === true
194  return s.always
195}
196
197async function restore($: EngineInterface, s: State, now: number): Promise<void> {
198  const deadline = await $.store.get(deadlineKey(s))
199  const every = await $.store.get(everyKey(s))
200  s.deadline = typeof deadline === 'number' && deadline > now ? deadline : 0
201  s.every = typeof every === 'number' && every >= MIN_PING_MS ? every : PING_AFTER_MS
202  await readAlways($, s)
203  s.lastRead = pingRecordOf(await $.store.get(lastKey(s)))
204}
205
206async function stop($: EngineInterface, s: State, why: string | null, forgetAlways = false): Promise<void> {
207  s.deadline = 0
208  s.endless = false
209  s.every = PING_AFTER_MS
210  s.stopped = why
211  disarm(s)
212  await $.store.delete(deadlineKey(s))
213  await $.store.delete(everyKey(s))
214  if (forgetAlways) {
215    s.always = false
216    await $.store.delete(KEY_ALWAYS)
217  }
218  await showStatus($, s)
219}
220
221/**
222 * Ends this session's endless loop once another session turned always off, before it pays for another
223 * ping; answers whether nothing runs any more. A window this session armed itself is its own, and runs on.
224 */
225async function endedElsewhere($: EngineInterface, s: State): Promise<boolean> {
226  if (!s.endless || (await readAlways($, s))) return false
227  s.endless = false
228  if (hasWindow(s)) return false
229  await stop($, s, null)
230  return true
231}
232
233/**
234 * Follows the always switch as the store holds it now: another session's /cache-warm off ends this
235 * session's endless loop, and its /cache-warm always starts one here, unless a window this session armed
236 * itself runs.
237 */
238async function followAlways($: EngineInterface, s: State): Promise<void> {
239  if (await endedElsewhere($, s)) return
240  if ((await readAlways($, s)) && !hasWindow(s)) await startEndless($, s)
241}
242
243/** The minute's redraw of a running window, which first ends a loop another session turned off. */
244async function redraw($: EngineInterface, s: State): Promise<void> {
245  if (!(await endedElsewhere($, s))) await showStatus($, s)
246}
247
248/** Schedules the next ping one period after the last request. */
249async function arm($: EngineInterface, s: State): Promise<void> {
250  disarm(s)
251  const now = await $.clock.now()
252  if (isOver(s, now)) {
253    // A window that ran out of time is armed again by the next message, as long as this one was. A
254    // window the ping stopped is not: the cache is gone there, and the cold write of the next message
255    // arms its own window. The endless loop never reaches this.
256    const again = { window: s.window, every: s.every }
257    await stop($, s, null)
258    s.renew = again
259    return showStatusAt($, s, now)
260  }
261  if (hasWindow(s) && s.lastRequestAt && !s.compacted) {
262    const delay = Math.max(1000, s.lastRequestAt + s.every - now)
263    s.pending = $.clock.after(delay, () => { void runPing($, s) })
264  }
265  // A session with no window draws too, so the pane follows the state instead of holding the last one.
266  await showStatusAt($, s, now)
267}
268
269/**
270 * Records the ping's answer and schedules the next one. A ping that found the cache gone stops a window
271 * with an end, because the write it paid for is the last thing that window wanted. The endless loop of
272 * `always` carries on instead: that write is the new cache, and the next ping keeps it.
273 */
274async function settlePing($: EngineInterface, s: State, usage: Usage, now: number): Promise<void> {
275  const price = priceNow(s)
276  const usd = price ? responseUsd(usage, price) : null
277  await keepLastRead($, s, { kind: 'ping', read: usage.cache_read_input_tokens, write: usage.cache_creation_input_tokens, usd, at: now })
278  await toStream($, pingLine({ kind: isWarmPing(usage) ? 'sent' : 'cold', usage, usd }))
279  if (!isWarmPing(usage)) {
280    if (!s.endless) return stop($, s, coldPingText(usage, usd))
281    const write = usage.cache_creation_input_tokens
282    s.coldWrites.push({ tokens: write, usd })
283    await logEvent($, s, `the ping found the cache gone and re-wrote ${fmtTok(write)} tokens (${fmtUsd(usd)}); always keeps the loop running. /cache-warm off stops it.`, coldWriteShort(write, usd))
284  }
285  s.lastRequestAt = now
286  await arm($, s)
287}
288
289/** Sends the ping's fork while the line reads `pinging…`; a throw comes back as its Error. */
290async function forkPing($: EngineInterface, s: State): Promise<ModelForkResult | Error> {
291  s.pinging = true
292  await showStatus($, s)
293  try {
294    return await $.model.fork({ prompt: PING_PROMPT })
295  } catch (err) {
296    return err instanceof Error ? err : new Error(String(err))
297  } finally {
298    s.pinging = false
299  }
300}
301
302/**
303 * A ping the engine could not send because this process has no reply to fork (a resume, or a module
304 * loaded into one): the window waits for the next reply, whose turn arms the ping again, and says until
305 * when the cache holds.
306 */
307async function waitForReply($: EngineInterface, s: State, now: number): Promise<void> {
308  s.waitingReply = true
309  await toStream($, pingLine({ kind: 'unsent', reason: noReplyText(s, now) }))
310  await showStatusAt($, s, now)
311}
312
313/** A ping that failed: one stream entry, and the window stops with the reason. */
314async function pingFailed($: EngineInterface, s: State, why: string): Promise<void> {
315  await toStream($, pingLine({ kind: 'failed', reason: why }))
316  await stop($, s, why)
317}
318
319async function ping($: EngineInterface, s: State): Promise<void> {
320  s.pending = null
321  if (!hasWindow(s) || (await endedElsewhere($, s))) return
322  const now = await $.clock.now()
323  // The window ended, or a request since the timer was set moved the ping later.
324  if (isOver(s, now) || now - s.lastRequestAt < s.every - 1000) return arm($, s)
325  // Known before any fork: a resumed process has no reply to fork (measured: `nothing-to-fork`).
326  if (s.waitingReply) return showStatusAt($, s, now)
327  await settleReply($, s, await forkPing($, s), now)
328}
329
330/** What the fork answered, scored: a ping, a wait for the next reply, or a failure that stops the window. */
331async function settleReply($: EngineInterface, s: State, reply: ModelForkResult | Error, now: number): Promise<void> {
332  if (reply instanceof Error) return pingFailed($, s, `the ping failed, ${errorText(reply)}`)
333  if (!reply.isAnswered && reply.reason === 'nothing-to-fork') return waitForReply($, s, now)
334  // A reply without text still read the cache, so it is scored as a ping; every other unanswered fork stops.
335  if (!reply.isAnswered && reply.reason !== 'empty-reply') return pingFailed($, s, unsentText(reply))
336  await settlePing($, s, reply.usage, now)
337}
338
339async function runPing($: EngineInterface, s: State): Promise<void> {
340  try {
341    await ping($, s)
342  } catch (err) {
343    await stop($, s, `the ping step failed, ${errorText(err)}`)
344  }
345}
346
347async function startWindow($: EngineInterface, s: State, windowMs: number, every: number): Promise<void> {
348  s.every = every
349  s.window = windowMs
350  s.renew = null
351  s.deadline = await $.clock.now() + windowMs
352  s.stopped = null
353  await $.store.set(deadlineKey(s), s.deadline)
354  if (every === PING_AFTER_MS) await $.store.delete(everyKey(s))
355  else await $.store.set(everyKey(s), every)
356  await arm($, s)
357}
358
359/**
360 * Starts the endless loop of `always`: a ping every period until `/cache-warm off`, with no end time.
361 * It writes no per-session deadline, so the switch stays one global key that every session of every
362 * project reads at its start.
363 */
364async function startEndless($: EngineInterface, s: State): Promise<void> {
365  s.every = PING_AFTER_MS
366  s.window = 0
367  s.renew = null
368  s.deadline = 0
369  s.endless = true
370  s.stopped = null
371  await $.store.delete(deadlineKey(s))
372  await $.store.delete(everyKey(s))
373  await arm($, s)
374}
375
376async function registerCommands($: EngineInterface): Promise<void> {
377  await $.command.register({
378    name: 'cache-warm',
379    description: 'Keep the prompt cache warm: bare for 6h, a window such as 90m, always, off, or status (cache-warm)',
380    argumentHint: '[6h | 90m | always | off | status]',
381    immediate: true,
382  })
383  await $.command.register({
384    name: 'cache-status',
385    description: 'Prompt cache state, cold price and this session\'s cold writes (cache-warm)',
386    immediate: true,
387  })
388}
389
390async function warmCommand($: EngineInterface, s: State, args: string): Promise<string> {
391  const command = parseWarmArgs(args)
392  switch (command.kind) {
393    case 'error':
394      return command.text
395    case 'status':
396      await followAlways($, s)
397      return statusText(s, await $.clock.now()) ?? idleText(s)
398    case 'off': {
399      const wasAlways = await readAlways($, s)
400      s.renew = null
401      await stop($, s, null, true)
402      return wasAlways ? 'off, and no longer starts itself in any session' : 'off'
403    }
404    case 'always':
405      s.always = true
406      await $.store.set(KEY_ALWAYS, true)
407      await startEndless($, s)
408      return `always on: a ping every ${fmtDuration(PING_AFTER_MS)} with no end, in this session and in every later session of every project; /cache-warm off turns it off for good`
409    case 'arm':
410      await startWindow($, s, command.window, command.every)
411      return `on for ${fmtDuration(command.window)}, a ping ${fmtDuration(command.every)} after each idle stretch keeps the cache read, not re-written`
412  }
413}
414
415async function clearSession($: EngineInterface, s: State): Promise<void> {
416  await stop($, s, null)
417  resetForClear(s)
418  s.sid = await $.session.id()
419  if (await readAlways($, s)) await startEndless($, s)
420}
421
422/** The origins of a message the person sent themselves, which is what arms a window again. */
423const USER_ORIGINS: readonly string[] = ['composer', 'bridge', 'sdk']
424
425/**
426 * Arms the window again with the person's next message, as long as the one that ran out of time was and
427 * with the same ping period. A message whose origin the engine does not name arms nothing, so a plugin's
428 * own prompt never renews a window the person let end.
429 */
430async function renewWindow($: EngineInterface, s: State, kind: string | undefined): Promise<void> {
431  const again = s.renew
432  if (again === null || s.deadline || kind === undefined || !USER_ORIGINS.includes(kind)) return
433  // The line is written before the window starts, so the sidebar's redraw already carries it.
434  const text = `the ${fmtDuration(again.window)} window ran out; this message arms another one. /cache-warm off stops it.`
435  await logEvent($, s, text, eventShort(`window armed again for ${fmtDuration(again.window)}`))
436  await startWindow($, s, again.window, again.every)
437}
438
439/** Scores a turn that re-wrote the context, and keeps the cache warm after it. */
440async function measure($: EngineInterface, s: State, u: TurnUsage, now: number): Promise<void> {
441  if (u.model) s.model = u.model
442  // A turn's usage sums its requests, so this is what the whole turn read and cost.
443  const price = priceNow(s)
444  await keepLastRead($, s, { kind: 'turn', read: u.cache_read_input_tokens, write: u.cache_creation_input_tokens, usd: price ? responseUsd(u, price) : null, at: now })
445  const previous = s.ctx
446  const write = u.cache_creation_input_tokens
447  // A turn's usage sums its responses, so a ten-step turn counts the context ten
448  // times; the engine's live window figure is the context, the sum only a fallback.
449  const live = (await $.session.usage()).context.tokens
450  s.ctx = live && live > 0 ? live : u.input_tokens + u.cache_read_input_tokens + write
451  if (!isColdWrite(previous, write)) return
452  const usd = writeUsd(write, priceNow(s))
453  s.coldWrites.push({ tokens: write, usd })
454  // The endless loop already keeps this cache; a window with an end would only shorten it.
455  if (s.endless || s.deadline >= now + AUTO_WARM_MS) return
456  // The line is written before the window starts, so the sidebar's redraw already carries it.
457  await logEvent($, s, `cold write of ${fmtTok(write)} tokens paid (${fmtUsd(usd)}). Keeping the cache warm for ${fmtDuration(AUTO_WARM_MS)}; /cache-warm off stops it.`, coldWriteShort(write, usd))
458  await startWindow($, s, AUTO_WARM_MS, s.every)
459}
460
461/**
462 * Whether the session bills fast mode rates, as far as the settings say: `/fast` writes `fastMode`, and
463 * `fastModePerSessionOptIn` starts every session with fast mode off whatever `fastMode` holds. Read at
464 * each turn, because `/fast` can change it mid-session.
465 */
466async function readFast($: EngineInterface, s: State): Promise<void> {
467  const settings = await $.settings.read() as { fastMode?: unknown; fastModePerSessionOptIn?: unknown }
468  s.fast = settings.fastMode === true && settings.fastModePerSessionOptIn !== true
469}
470
471async function afterTurn($: EngineInterface, s: State, durationMs: number, usage: TurnUsage | undefined): Promise<void> {
472  const now = await $.clock.now()
473  await readFast($, s)
474  // turn.step stamps each request; when no step of this turn did, the turn's end is the floor.
475  if (now - s.lastRequestAt > durationMs) s.lastRequestAt = now
476  await $.store.set(requestKey(s), s.lastRequestAt)
477  s.compacted = false
478  s.waitingReply = false
479  // The stop reason and the line under it belong to the window that ended: one turn later the pane
480  // carries the idle line instead, and the reason stays in the transcript.
481  if (s.stopped) {
482    s.stopped = null
483    s.event = undefined
484  }
485  // `always` runs until /cache-warm off, so a ping that failed stopped the loop for this turn alone:
486  // the next turn starts it again. Without this the session keeps `always` stored while running no
487  // loop, and the next cold write arms a 6h window in its place. A window the person armed by hand
488  // holds, because it is a window of its own. The switch is read from the store, so another session's
489  // /cache-warm always or off applies here too.
490  await followAlways($, s)
491  if (usage) await measure($, s, usage, now)
492  await arm($, s)
493}
494
495/**
496 * Stamps a main-loop request. The request writes the cache a compaction dropped, and it sets the ping
497 * when none is pending: a turn's end sets it otherwise, so the first turn of a session, or the first
498 * after /clear or a compaction, would run one long tool call past the hour with no ping. A pending ping
499 * reads the newest request when it fires and moves itself later, so a later step sets nothing.
500 */
501async function stampRequest($: EngineInterface, s: State): Promise<void> {
502  s.lastRequestAt = await $.clock.now()
503  s.compacted = false
504  // This request's reply is one the engine can fork, and the ping it arms comes long after it.
505  s.waitingReply = false
506  if (s.pending === null && hasWindow(s)) await arm($, s)
507}
508
509/**
510 * The time of the last request of a conversation this module did not see: a reloaded module starts with
511 * no request time, and `always` would wait for the first turn to arm its ping. Each turn's end keeps that
512 * time in the store, and each ping keeps its own read (`restore` put it in `s.lastRead`); the later of the
513 * two counts, because a session kept warm by pings alone has a last turn older than the cache it holds.
514 * Only a session with neither kept (one that ran an older version) reads the last write of its
515 * transcript, which a reload's own line moves to the reload. Only a cache that is still warm is taken, so
516 * a reload never pays for a cold ping the next message would pay anyway.
517 */
518async function seedLastRequest($: EngineInterface, s: State, now: number): Promise<void> {
519  const kept = await $.store.get(requestKey(s))
520  const latest = Math.max(typeof kept === 'number' ? kept : 0, s.lastRead?.at ?? 0)
521  if (latest === 0) return seedFromTranscript($, s, now)
522  if (now - latest < TTL_MS) s.lastRequestAt = latest
523}
524
525/**
526 * Sends the keep-warm message after a resume. The resumed process has no reply to fork, so no ping can
527 * go before one comes, and the cache would lapse an hour after the last request while the person wrote
528 * nothing. The message is a real turn: it reads the cache, and its end arms the ping again. It runs from
529 * a timer through the mod's own markdown command; a run the engine refuses goes out as a plugin prompt.
530 */
531async function keepWarmAfterResume($: EngineInterface, s: State): Promise<void> {
532  await followAlways($, s)
533  const now = await $.clock.now()
534  if (!wantsKeepWarm(s, now)) return
535  await toStream($, keepWarmLine(s, now))
536  try {
537    await $.command.run({ command: SEND_COMMAND, args: KEEP_WARM_TEXT })
538  } catch (err) {
539    await toStream($, pingLine({ kind: 'unsent', reason: `the send command did not run, the keep-warm message goes out as a plugin prompt: ${errorText(err)}` }))
540    const res = await $.prompt.submit({ text: KEEP_WARM_TEXT })
541    if (res.drop !== undefined) await toStream($, pingLine({ kind: 'failed', reason: `the keep-warm message was dropped: ${res.drop}` }))
542  }
543}
544
545/**
546 * Starts the keep-warm wait once both start hooks ran: classic.SessionStart names the resume, and
547 * session.start reads the switch, the window and whether the session is interactive. The two settle in no
548 * fixed order: with every mod of the marketplace loaded, session.start settled four seconds after
549 * classic.SessionStart (measured on 2.1.285), so a wait started by classic.SessionStart alone read a session
550 * that had not started and sent nothing.
551 */
552function scheduleKeepWarm($: EngineInterface, s: State): void {
553  if (!s.started || !s.keepWarmDue) return
554  s.keepWarmDue = false
555  $.clock.after(KEEP_WARM_AFTER_MS, () => { void keepWarmAfterResume($, s) })
556}
557
558/**
559 * The transcript lies under the directory the session started in, which a shell `cd` does not move
560 * (measured: a module reloaded after `cd sub` looked under `sub` and found nothing).
561 */
562async function seedFromTranscript($: EngineInterface, s: State, now: number): Promise<void> {
563  const configDir = (await $.env.get('CLAUDE_CONFIG_DIR')) || `${(await $.env.get('HOME')) ?? ''}/.claude`
564  const path = transcriptPath(configDir, await $.session.root(), s.sid)
565  try {
566    if (!(await $.fs.exists(path))) {
567      $.ui.log(`the last request time of this session was not read: ${path} does not exist`)
568      return
569    }
570    const { mtimeMs } = await $.fs.stat(path)
571    if (now - mtimeMs < TTL_MS) s.lastRequestAt = mtimeMs
572  } catch (err) {
573    $.ui.log(`the last request time of this session was not read from ${path}: ${errorText(err)}`)
574  }
575}
576
577export const register: Register = on => {
578  const s = freshState()
579
580  on('session.start', async ($, e, next) => {
581    const r = await next(e)
582    s.sid = await $.session.id()
583    s.interactive = e.isInteractive
584    const now = await $.clock.now()
585    await prune($, now)
586    await pruneRequests($, s, now)
587    await restore($, s, now)
588    await readFast($, s)
589    // A reload gets no classic.SessionStart, and a ping can come before the first turn names the model.
590    s.model ??= await $.session.model()
591    const live = (await $.session.usage()).context.tokens
592    if (live) s.ctx = live
593    // A loaded conversation this module has not seen a request of: a reload, or an update mid-session.
594    if (live && !s.lastRequestAt) await seedLastRequest($, s, now)
595    // The always switch is one global key, so every session of every project starts the endless loop,
596    // whatever the last window of this session left behind.
597    if (s.always) await startEndless($, s)
598    await registerCommands($)
599    // A -p run draws nothing, so only an interactive session redraws on a timer, and only while a window runs.
600    if (e.isInteractive) $.clock.every(REDRAW_MS, () => { if (hasWindow(s)) void redraw($, s) })
601    await showStatus($, s)
602    s.started = true
603    scheduleKeepWarm($, s)
604    return r
605  })
606
607  // /clear arrives only through the classic seam, and a resume brings the
608  // fields Claude Code computes for settings hooks.
609  on('classic.SessionStart', async ($, e, next) => {
610    const r = await next(e)
611    if (e.agent_id !== undefined) return r
612    if (e.source === 'clear') {
613      await clearSession($, s)
614      return r
615    }
616    // The last ping this session kept, also when this seam runs before session.start restored it.
617    s.sid ||= await $.session.id()
618    s.lastRead ??= pingRecordOf(await $.store.get(lastKey(s)))
619    const line = seedFromResume(s, e, await $.clock.now())
620    s.model ??= await $.session.model()
621    if (line) await logEvent($, s, line, eventShort(line))
622    // The conditions are read when the timer fires, which scheduleKeepWarm starts after session.start too.
623    if (e.source === 'resume') {
624      s.keepWarmDue = true
625      scheduleKeepWarm($, s)
626    }
627    return r
628  })
629
630  // Only the origin is read; the prompt text passes through untouched.
631  on('prompt.submit', async ($, e, next) => {
632    const r = await next(e)
633    await renewWindow($, s, (e.origin as { kind?: string } | undefined)?.kind)
634    return r
635  })
636
637  on('command.run', { command: 'cache-warm' }, async ($, e) => ({ text: await warmCommand($, s, String(e.args ?? '')) }))
638
639  on('command.run', { command: 'cache-status' }, async $ => {
640    await followAlways($, s)
641    return { text: card(s, await $.clock.now()) }
642  })
643
644  on('turn.step', async function* ($, e, next) {
645    if (!e.agentId) await stampRequest($, s)
646    yield* next(e)
647  })
648
649  on('turn.complete', async ($, e, next) => {
650    const r = await next(e)
651    if (!e.agentId) await afterTurn($, s, e.durationMs, e.usage)
652    return r
653  })
654
655  on('session.compact', async ($, e, next) => {
656    const r = await next(e)
657    if (e.agentId) return r
658    s.compacted = true
659    s.ctx = 0
660    disarm(s)
661    await showStatus($, s)
662    return r
663  })
664}
665
hooks/pricing.ts 92 lines
1/** List prices in dollars per million tokens. */
2export interface Price {
3  /** Cache hits and refreshes. */
4  read: number
5  /** 1-hour cache writes; the base input rate is half of it. */
6  write: number
7  output: number
8}
9
10/** The token counts an API response bills. */
11export interface Usage {
12  input_tokens: number
13  output_tokens: number
14  cache_read_input_tokens: number
15  cache_creation_input_tokens: number
16}
17
18/**
19 * Model pricing from platform.claude.com/docs/en/about-claude/pricing, read in
20 * September 2026. A model id takes the first row whose family it contains, so
21 * the specific families come before the general ones.
22 */
23const PRICES: ReadonlyArray<readonly [family: string, price: Price]> = [
24  ['fable-5-1', { read: 0.25, write: 20, output: 50 }],
25  ['mythos-5-1', { read: 0.25, write: 20, output: 50 }],
26  ['fable-5', { read: 1, write: 20, output: 50 }],
27  ['mythos-5', { read: 1, write: 20, output: 50 }],
28  ['opus-5-5', { read: 0.2, write: 8, output: 20 }],
29  ['opus-5', { read: 0.5, write: 10, output: 25 }],
30  ['opus-4-5', { read: 0.5, write: 10, output: 25 }],
31  ['opus-4-6', { read: 0.5, write: 10, output: 25 }],
32  ['opus-4-7', { read: 0.5, write: 10, output: 25 }],
33  ['opus-4-8', { read: 0.5, write: 10, output: 25 }],
34  ['opus-4', { read: 1.5, write: 30, output: 75 }],
35  ['sonnet-5-5', { read: 0.2, write: 4, output: 10 }],
36  ['sonnet-5', { read: 0.2, write: 4, output: 10 }],
37  ['sonnet', { read: 0.3, write: 6, output: 15 }],
38  ['haiku-4-5', { read: 0.1, write: 2, output: 5 }],
39  ['haiku', { read: 0.08, write: 1.6, output: 4 }],
40]
41
42/**
43 * Fast mode rates, from the same page: the fast base input and output, with the cache multipliers
44 * applied on top of the fast base input (2x for the 1-hour write, 0.05x read on Opus 5.5, 0.1x on
45 * the others). Only these models take fast mode; any other model bills its standard rates.
46 */
47const FAST_PRICES: ReadonlyArray<readonly [family: string, price: Price]> = [
48  ['opus-5-5', { read: 0.4, write: 16, output: 40 }],
49  ['opus-5', { read: 1, write: 20, output: 50 }],
50  ['opus-4-8', { read: 1, write: 20, output: 50 }],
51]
52
53const MTOK = 1e6
54
55function rowOf(rows: ReadonlyArray<readonly [string, Price]>, id: string): Price | undefined {
56  return rows.find(([family]) => id.includes(family))?.[1]
57}
58
59const idOf = (model: string | null): string => (model ?? '').toLowerCase().replace(/[\s.]+/g, '-')
60
61/** Whether fast mode bills this model at its own rates. */
62export function hasFastRate(model: string | null): boolean {
63  return rowOf(FAST_PRICES, idOf(model)) !== undefined
64}
65
66export function priceOf(model: string | null, fast = false): Price | null {
67  const id = idOf(model)
68  return (fast ? rowOf(FAST_PRICES, id) : undefined) ?? rowOf(PRICES, id) ?? null
69}
70
71/** What writing this many tokens to the 1-hour cache costs. */
72export function writeUsd(tokens: number, price: Price | null): number | null {
73  return price ? tokens * price.write / MTOK : null
74}
75
76/** What reading this many tokens from the cache costs. */
77export function readUsd(tokens: number, price: Price | null): number | null {
78  return price ? tokens * price.read / MTOK : null
79}
80
81/** Everything one response bills: the cache read, the cache write, the uncached input at the base rate and the output. */
82export function responseUsd(u: Usage, price: Price): number {
83  const base = price.write / 2
84  return (u.cache_read_input_tokens * price.read + u.cache_creation_input_tokens * price.write +
85    u.input_tokens * base + u.output_tokens * price.output) / MTOK
86}
87
88/** How many reads of a context cost as much as one write of it; the upper bound of pings worth sending. */
89export function breakEvenPings(price: Price | null): number | null {
90  return price ? Math.floor(price.write / price.read) : null
91}
92
hooks/warm.ts 445 lines
1import { breakEvenPings, hasFastRate, priceOf, readUsd, writeUsd, type Price, type Usage } from './pricing.ts'
2
3const MIN = 60 * 1000
4const HOUR = 60 * MIN
5
6/** The 1-hour cache tier the main conversation uses. */
7export const TTL_MS = HOUR
8/** The ping lands this long after the last request, well inside the hour. */
9export const PING_AFTER_MS = 50 * MIN
10/** The floor of the `every` test setting. */
11export const MIN_PING_MS = MIN
12/** The window armed after a paid cold write. */
13export const AUTO_WARM_MS = 6 * HOUR
14/** The window of a bare /cache-warm and of `always`. */
15export const DEFAULT_WINDOW_MS = 6 * HOUR
16/** Below this context a cold resume is not worth a line. */
17export const BIG_TOKENS = 50_000
18/** A turn is a cold write only when the context before it was at least this large. */
19const COLD_WRITE_MIN_CONTEXT = 20_000
20/** A warm ping writes only its own message; a write of this share of the read or more means the prefix broke. */
21const WARM_WRITE_RATIO = 0.1
22
23/** The last request that read the cache: a ping, or a main-loop turn, whichever came last. */
24export interface PingRecord {
25  kind: 'ping' | 'turn'
26  read: number
27  write: number
28  usd: number | null
29  /** When the answer came, in ms since the epoch. */
30  at: number
31}
32
33const isCount = (v: unknown): v is number => typeof v === 'number' && Number.isFinite(v) && v >= 0
34
35/** A record kept in the store, or null when the stored value is of another shape. */
36export function pingRecordOf(value: unknown): PingRecord | null {
37  if (typeof value !== 'object' || value === null) return null
38  const r = value as Record<string, unknown>
39  if (r.kind !== 'ping' && r.kind !== 'turn') return null
40  if (!isCount(r.read) || !isCount(r.write) || !isCount(r.at)) return null
41  if (r.usd !== null && !isCount(r.usd)) return null
42  return { kind: r.kind, read: r.read, write: r.write, usd: r.usd, at: r.at }
43}
44
45export interface ColdWrite {
46  tokens: number
47  usd: number | null
48}
49
50export interface State {
51  sid: string
52  /** When the keep-warm window ends; 0 means off. `endless` runs with no deadline at all. */
53  deadline: number
54  /** The `always` loop: a ping every period until /cache-warm off, with no end time. */
55  endless: boolean
56  /** How long the running window was armed for, so one that runs out can start again as long. */
57  window: number
58  /** The window the next message of the person starts again; null when none ran out. */
59  renew: { window: number; every: number } | null
60  every: number
61  always: boolean
62  lastRequestAt: number
63  model: string | null
64  /** The `fastMode` setting is on and not reset per session, so a model with fast rates bills them. */
65  fast: boolean
66  ctx: number
67  compacted: boolean
68  coldWrites: ColdWrite[]
69  pending: { cancel: () => void } | null
70  lastRead: PingRecord | null
71  stopped: string | null
72  /** The short form of the last transcript line, drawn faint under the window line in the sidebar. */
73  event?: Line
74  /** A ping's fork is out; the line says so until its answer comes. */
75  pinging: boolean
76  /**
77   * The engine has no reply of this process to fork, as after a resume (measured: `nothing-to-fork`
78   * until the first reply), so no ping can go before the next reply.
79   */
80  waitingReply: boolean
81  /** Only an interactive session sends the keep-warm message after a resume. */
82  interactive: boolean
83  /** When this process resumed the conversation; 0 when it did not. */
84  resumedAt: number
85  /** session.start has read the switch, the window and whether the session is interactive. */
86  started: boolean
87  /** A resume waits to send the keep-warm message until session.start has run too. */
88  keepWarmDue: boolean
89}
90
91/** The rates the session bills now: its model, at fast mode rates while the setting says so. */
92export function priceNow(s: State): Price | null {
93  return priceOf(s.model, s.fast)
94}
95
96export function freshState(): State {
97  return {
98    sid: '', deadline: 0, endless: false, window: 0, renew: null, every: PING_AFTER_MS, always: false, lastRequestAt: 0, model: null, fast: false, ctx: 0,
99    compacted: false, coldWrites: [], pending: null, lastRead: null, stopped: null, pinging: false, waitingReply: false, interactive: false, resumedAt: 0,
100    started: false, keepWarmDue: false,
101  }
102}
103
104export function parseDuration(text: string): number | null {
105  const m = /^(?:(\d+)h)?(?:(\d+)m)?$/.exec(text.trim())
106  if (!m || (m[1] === undefined && m[2] === undefined)) return null
107  return (Number(m[1] ?? 0) * 60 + Number(m[2] ?? 0)) * MIN
108}
109
110export function fmtDuration(ms: number): string {
111  const total = Math.max(0, Math.round(ms / MIN))
112  const h = Math.floor(total / 60)
113  const m = total % 60
114  if (h >= 48) return `${Math.floor(h / 24)}d ${h % 24}h`
115  if (h === 0) return `${m}m`
116  return m > 0 ? `${h}h ${m}m` : `${h}h`
117}
118
119export function fmtUsd(usd: number | null): string {
120  if (usd == null) return 'n/a'
121  return '$' + (usd >= 100 ? usd.toFixed(0) : usd.toFixed(2))
122}
123
124export function fmtTok(n: number): string {
125  if (n >= 1e6) return (n / 1e6).toFixed(1) + 'M'
126  return n >= 1000 ? Math.round(n / 1000) + 'k' : String(n)
127}
128
129function fmtCount(n: number): string {
130  return n.toLocaleString('en-US')
131}
132
133export type WarmCommand =
134  | { kind: 'arm'; window: number; every: number }
135  | { kind: 'always' | 'off' | 'status' }
136  | { kind: 'error'; text: string }
137
138const USAGE = 'expects a window such as 6h or 90m, or always, off, or status'
139
140function parseEvery(words: readonly string[]): number | null {
141  if (words[0] === undefined) return PING_AFTER_MS
142  if (words[0] !== 'every' || words.length !== 2) return null
143  const period = parseDuration(words[1] ?? '')
144  return period != null && period >= MIN_PING_MS ? period : null
145}
146
147/** Reads the argument of /cache-warm. */
148export function parseWarmArgs(args: string): WarmCommand {
149  const words = args.trim().split(/\s+/).filter(Boolean)
150  const [first, ...rest] = words
151  if (first === undefined) return { kind: 'arm', window: DEFAULT_WINDOW_MS, every: PING_AFTER_MS }
152  if ((first === 'always' || first === 'off' || first === 'status') && rest.length === 0) return { kind: first }
153  const window = parseDuration(first)
154  if (window == null || window === 0) return { kind: 'error', text: USAGE }
155  const every = parseEvery(rest)
156  if (every == null) return { kind: 'error', text: 'every takes a period of at least 1m, as in 6h every 2m' }
157  return { kind: 'arm', window, every }
158}
159
160/** True while the mod keeps the cache warm: a window with an end, or the endless `always` loop. */
161export function hasWindow(s: State): boolean {
162  return s.endless || s.deadline > 0
163}
164
165/** True when a window with an end has reached it. The endless loop never does. */
166export function isOver(s: State, now: number): boolean {
167  return !s.endless && s.deadline > 0 && now >= s.deadline
168}
169
170export function isCold(s: State, now: number): boolean {
171  return s.lastRequestAt > 0 && !s.compacted && now - s.lastRequestAt >= TTL_MS
172}
173
174export function isWarmPing(u: Usage): boolean {
175  return u.cache_read_input_tokens > 0 && u.cache_creation_input_tokens < WARM_WRITE_RATIO * u.cache_read_input_tokens
176}
177
178/** A turn re-wrote the context when it wrote at least half of a sizeable context. */
179export function isColdWrite(previousContext: number, write: number): boolean {
180  return previousContext > COLD_WRITE_MIN_CONTEXT && write >= 0.5 * previousContext
181}
182
183export function coldPingText(u: Usage, usd: number | null): string {
184  return `the ping read ${fmtTok(u.cache_read_input_tokens)} and wrote ${fmtTok(u.cache_creation_input_tokens)} tokens (${fmtUsd(usd)}), the cache was already gone`
185}
186
187/** A fork result with no reply to score: nothing to fork yet, an API error, or a call cut before its reply. */
188export type Unsent = { reason: 'nothing-to-fork' } | { reason: 'api-error'; status: number | null; error: string } | { reason: 'aborted' }
189
190export function unsentText(r: Unsent): string {
191  if (r.reason === 'nothing-to-fork') return 'the engine did not send the ping; the conversation has no reply to fork yet'
192  if (r.reason === 'aborted') return 'the ping was cut before a reply came'
193  return `the ping failed, the API answered ${r.status ?? 'nothing'} (${r.error})`
194}
195
196const MONTHS = ['Jan', 'Feb', 'Mar', 'Apr', 'May', 'Jun', 'Jul', 'Aug', 'Sep', 'Oct', 'Nov', 'Dec']
197
198/** A local clock time as `05:42`, with the day and month in front when it is not today: `22 Sep 23:10`. */
199export function clockText(at: number, now: number): string {
200  const date = new Date(at)
201  const time = `${String(date.getHours()).padStart(2, '0')}:${String(date.getMinutes()).padStart(2, '0')}`
202  return date.toDateString() === new Date(now).toDateString() ? time : `${date.getDate()} ${MONTHS[date.getMonth()]} ${time}`
203}
204
205/** How the sidebar colours a line or a part of one. */
206export type Tone = 'ok' | 'warn' | 'error' | 'dim'
207export type Part = { text: string; kind?: Tone }
208/** A line; `parts` colour pieces of it, and `text` holds the whole line for a sidebar that draws no parts. */
209export type Line = { text: string; kind?: Tone; parts?: Part[] }
210
211const part = (text: string, kind: Tone | undefined): Part => (kind === undefined ? { text } : { text, kind })
212
213const joined = (parts: Part[]): string => parts.map(p => p.text).join('')
214
215/** A line made of parts, its `text` their texts joined. */
216export const partsLine = (parts: Part[]): Line => ({ text: joined(parts), parts })
217
218/** When the cache lapses: an hour after the last request that read it. */
219function cacheEnd(s: State, now: number): string {
220  const end = s.lastRequestAt + TTL_MS
221  return end > now ? `cache holds until ${clockText(end, now)}` : `cache ended at ${clockText(end, now)}`
222}
223
224/**
225 * What the window waits for next: its own fork, a reply the engine can fork, or the ping's time. The
226 * last minute reads `ping now`, because minutes round and the line is drawn once a minute.
227 */
228function nextText(s: State, now: number): string {
229  if (s.pinging) return ' · pinging…'
230  if (!s.lastRequestAt || s.compacted) return ' · waiting for the first turn'
231  if (s.waitingReply) return ` · no ping before the next reply · ${cacheEnd(s, now)}`
232  const left = s.lastRequestAt + s.every - now
233  return left < MIN ? ' · ping now' : ` · ping in ${fmtDuration(left)}`
234}
235
236/** Why a ping could not go: this process has no reply to fork yet. */
237export function noReplyText(s: State, now: number): string {
238  return `the conversation has no reply to fork yet; the ping waits for the next reply, and the ${cacheEnd(s, now)}`
239}
240
241/** What one ping attempt did: sent, found the cache gone, not sent, or failed. */
242export type PingOutcome = { kind: 'sent' | 'cold'; usage: Usage; usd: number | null } | { kind: 'unsent' | 'failed'; reason: string }
243
244/** One stream entry per ping attempt; only the verdict is coloured. */
245export function pingLine(o: PingOutcome): Line {
246  if ('reason' in o) return partsLine([o.kind === 'unsent' ? part('ping not sent', 'warn') : part('ping failed', 'error'), part(`: ${o.reason}`, 'dim')])
247  const head = o.kind === 'sent' ? part('ping sent', 'ok') : part('ping found the cache gone', 'error')
248  return partsLine([head, part(` · read ${fmtTok(o.usage.cache_read_input_tokens)} · wrote ${fmtTok(o.usage.cache_creation_input_tokens)} · ${fmtUsd(o.usd)}`, 'dim')])
249}
250
251/** The message a resumed session sends to keep its cache; the model reads it as the person's prompt. */
252export const KEEP_WARM_TEXT =
253  'This message was sent by the cache-warm plugin, not by the person. The session was resumed, and a resumed session can keep its prompt cache warm only after a reply. Do not run a tool or continue a task. Reply with the single word: warm'
254
255/**
256 * Whether a resumed session sends the keep-warm message now: an interactive session under a window or
257 * `always`, with a context worth keeping, whose cache still holds and which sent no request since.
258 */
259export function wantsKeepWarm(s: State, now: number): boolean {
260  if (!s.interactive || !hasWindow(s) || s.compacted || s.ctx < BIG_TOKENS || s.resumedAt === 0) return false
261  return s.lastRequestAt > 0 && s.lastRequestAt < s.resumedAt && now - s.lastRequestAt < TTL_MS
262}
263
264/** The stream entry of the keep-warm message. */
265export function keepWarmLine(s: State, now: number): Line {
266  return partsLine([part('keep-warm message sent', 'ok'), part(`: the session was resumed and its ${cacheEnd(s, now)}; a resumed session pings only after a reply`, 'dim')])
267}
268
269/**
270 * The window's state in parts; undefined while no window runs. Only the time left (or `always`) takes
271 * the window's colour, the ping details are faint, and a stop shows its `stopped:` front red.
272 */
273export function statusParts(s: State, now: number): Part[] | undefined {
274  if (s.stopped) return [part('stopped:', 'error'), part(` ${s.stopped}`, undefined)]
275  if (!hasWindow(s)) return undefined
276  const next = nextText(s, now)
277  const last = s.lastRead
278  const ping = last ? [part(` · last ${last.kind} read ${fmtTok(last.read)} ${fmtUsd(last.usd)} (${clockText(last.at, now)})`, 'dim')] : []
279  const left = s.endless ? 'always' : `${fmtDuration(s.deadline - now)} left`
280  return [part(left, statusTone(s, now)), part(next, 'dim'), ...ping]
281}
282
283/** The status line; undefined clears it. The engine puts the mod name in front. */
284export function statusText(s: State, now: number): string | undefined {
285  const parts = statusParts(s, now)
286  return parts === undefined ? undefined : joined(parts)
287}
288
289/**
290 * The sidebar line while no window runs, in parts: what this session paid for cold writes and how large
291 * the context is. It replaces the stop reason at the next turn, so the pane holds a measurement of now
292 * instead of one sentence of the window that ended. The transcript keeps the reason. The line is faint
293 * but for a paid cold write, which is yellow.
294 */
295export function idleParts(s: State): Part[] {
296  const count = s.coldWrites.length
297  const paid = s.coldWrites.reduce((sum, w) => sum + (w.usd ?? 0), 0)
298  const writes = count === 0 ? part('no cold write', 'dim') : part(`${count} cold write${count === 1 ? '' : 's'} paid ${fmtUsd(paid)}`, 'warn')
299  const context = s.ctx > 0 ? [part(` · context ${fmtTok(s.ctx)} tokens`, 'dim')] : []
300  const again = s.renew ? ` · ${fmtDuration(s.renew.window)} again at your next message` : ''
301  return [part(`off${again} · `, 'dim'), writes, ...context]
302}
303
304export function idleText(s: State): string {
305  return joined(idleParts(s))
306}
307
308/** The sidebar's first line: the window's state, or the idle line while none runs. */
309export function windowLine(s: State, now: number): Line {
310  return partsLine(statusParts(s, now) ?? idleParts(s))
311}
312
313/**
314 * The colour of that line in the sidebar: red for a window the mod stopped, yellow while the window
315 * ends within one ping period (no further ping renews it), green while it holds, faint before the
316 * first turn, when there is nothing to keep warm yet.
317 */
318export function statusTone(s: State, now: number): 'ok' | 'warn' | 'error' | 'dim' {
319  if (s.stopped) return 'error'
320  if (!s.lastRequestAt || s.compacted) return 'dim'
321  // No ping can keep the cache before the next reply.
322  if (s.waitingReply) return 'warn'
323  if (s.endless) return 'ok'
324  return s.deadline - now <= s.every ? 'warn' : 'ok'
325}
326
327/** How many characters of the last event the sidebar's second line holds. */
328const MAX_EVENT = 120
329
330/** A transcript line as the sidebar's second line: faint, and cut, because the pane holds one row for it. */
331export function eventShort(text: string): Line {
332  return { text: text.length > MAX_EVENT ? `${text.slice(0, MAX_EVENT - 1)}…` : text, kind: 'dim' }
333}
334
335/** The cold write as the sidebar's second line: what it cost, without the instruction the line carries; the cost yellow. */
336export function coldWriteShort(tokens: number, usd: number | null): Line {
337  return partsLine([part(`cold write ${fmtTok(tokens)} tokens paid `, 'dim'), part(`(${fmtUsd(usd)})`, 'warn')])
338}
339
340export type ResumeFields = {
341  source: string
342  model?: string
343  context_tokens?: number
344  seconds_since_last_response?: number
345  prompt_cache_likely_expired?: boolean
346  estimated_cache_write_usd?: number
347}
348
349/**
350 * The last request of a resumed conversation: its last reply, or a later ping this mod kept. The resumed
351 * process has no reply of its own to fork until the first one comes.
352 */
353function seedResumeClock(s: State, e: ResumeFields, now: number): void {
354  if (typeof e.seconds_since_last_response === 'number') s.lastRequestAt = now - e.seconds_since_last_response * 1000
355  if (s.lastRead && s.lastRead.at > s.lastRequestAt) s.lastRequestAt = s.lastRead.at
356  s.resumedAt = now
357  s.waitingReply = true
358}
359
360/** Whether a request this mod kept read the cache within its lifetime. */
361function readRecently(s: State, now: number): boolean {
362  return s.lastRead !== null && now - s.lastRead.at < TTL_MS
363}
364
365/**
366 * Applies the fields Claude Code computes for a resumed session; returns the line to log, if any. Claude
367 * Code dates the cache from the transcript's last reply, which a ping never writes, so a ping this mod
368 * kept (`s.lastRead`) within the cache's lifetime means the cache was warm when the session closed.
369 */
370export function seedFromResume(s: State, e: ResumeFields, now: number): string | null {
371  if (e.source !== 'resume' && e.source !== 'fork') return null
372  if (typeof e.context_tokens === 'number' && e.context_tokens > 0) s.ctx = e.context_tokens
373  seedResumeClock(s, e, now)
374  if (typeof e.model === 'string') s.model = e.model
375  s.compacted = false
376  if (e.prompt_cache_likely_expired !== true || readRecently(s, now) || s.ctx < BIG_TOKENS) return null
377  const usd = typeof e.estimated_cache_write_usd === 'number' ? e.estimated_cache_write_usd : writeUsd(s.ctx, priceNow(s))
378  return `the cache expired while the session was closed. The first message will re-write ${fmtCount(s.ctx)} tokens, about ${fmtUsd(usd)}.`
379}
380
381/**
382 * Where the engine keeps a session's transcript: under `projects/`, in a directory named after the
383 * session's start directory with every character but a letter or a digit turned into `-` (measured on
384 * 2.1.280).
385 */
386export function transcriptPath(configDir: string, cwd: string, sid: string): string {
387  return `${configDir}/projects/${cwd.replace(/[^A-Za-z0-9]/g, '-')}/${sid}.jsonl`
388}
389
390/** /clear starts a new conversation in the same process; nothing measured before it still applies. */
391export function resetForClear(s: State): void {
392  s.pending?.cancel()
393  s.pending = null
394  s.ctx = 0
395  s.lastRequestAt = 0
396  s.compacted = false
397  s.coldWrites = []
398  s.lastRead = null
399  s.stopped = null
400  s.event = undefined
401  s.renew = null
402  s.waitingReply = false
403  s.resumedAt = 0
404}
405
406function stateLine(s: State, now: number): string {
407  if (s.compacted) return 'reset by compaction, waiting for the first turn'
408  if (!s.lastRequestAt) return 'no request yet this session'
409  if (isCold(s, now)) return `COLD, last request ${fmtDuration(now - s.lastRequestAt)} ago`
410  return `warm, ${fmtDuration(s.lastRequestAt + TTL_MS - now)} left`
411}
412
413function warmLine(s: State, now: number): string {
414  const always = s.always ? ' (always)' : ''
415  if (hasWindow(s)) return `on, ${statusText(s, now) ?? ''}${always}`
416  if (s.stopped) return `stopped, ${s.stopped}${always}`
417  if (s.renew) return `off, ${fmtDuration(s.renew.window)} again at your next message${always}`
418  if (s.always) return `off until the next session start or /clear, which start the endless loop again (always)`
419  return `off (/cache-warm arms it for ${fmtDuration(DEFAULT_WINDOW_MS)})`
420}
421
422function breakEvenLine(s: State): string | null {
423  const pings = breakEvenPings(priceNow(s))
424  if (pings == null || s.ctx <= 0) return null
425  return `up to ${pings} pings at the read rate cost one cold write, about ${fmtDuration(pings * s.every)} of idle at one ping per ${fmtDuration(s.every)}`
426}
427
428/** The /cache-status card. */
429export function card(s: State, now: number): string {
430  const price = priceNow(s)
431  const paid = s.coldWrites.reduce((sum, w) => sum + (w.usd ?? 0), 0)
432  const count = s.coldWrites.length
433  const breakEven = breakEvenLine(s)
434  const fast = s.fast && hasFastRate(s.model) ? ' · fast mode rates (the fastMode setting)' : ''
435  return [
436    `${s.model ?? 'model not seen yet'}${fast}`,
437    `state       ${stateLine(s, now)}`,
438    `context     ${fmtCount(s.ctx)} tokens`,
439    `cold cost   ${fmtUsd(writeUsd(s.ctx, price))} to re-write it (warm turn ${fmtUsd(readUsd(s.ctx, price))})`,
440    `keep warm   ${warmLine(s, now)}`,
441    ...(breakEven ? [`break-even  ${breakEven}`] : []),
442    `session     ${count} cold write${count === 1 ? '' : 's'} paid, ${fmtUsd(paid)}`,
443  ].join('\n')
444}
445