SLOPSHOPPER

Cache Keep-Warm

Shows how long your session's prompt cache stays warm, and can keep it warm in the background while you are away

newbandguardcommandtoastmodel
★ 4v1.0.8MITupdated 2026-10-08andreichiritescu/claude-code-cache-keep-warm
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · cache-keep-warm
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /keepwarm ⎿ cache-keep-warm: Keep-warm off: the cache expires normally. ⎿ cache-keep-warm: This session's cache lasts 1-hour (subscription). ○ ▱▱▱▱▱▱▱▱▱▱ nothing cached yet [ Keep warm ] ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Band
○ ▱▱▱▱▱▱▱▱▱▱ nothing cached yet [ Keep warm ]
README

Cache Keep-Warm

Stop paying to re-cache your whole Claude Code conversation after a break.

A small mod for Claude Code that shows how long your session's prompt cache stays warm, and keeps it warm while you are away.

License: MIT Claude Code Works in

<img src="docs/images/screenshot.png" alt="Cache Keep-Warm in the Claude Code desktop app: a green countdown line above the prompt reading '57 min left (21:38) · next ping 21:33 · 1 ping ✓ · 460k cached', and the session list with a green dot in front of the sessions being kept warm" width="100%">

Why you would want it

Claude Code keeps your conversation in a prompt cache, so every message only pays full price for what is new. But the cache expires when you stop typing for a while: after one hour on a Claude subscription, after five minutes on an API key.

Come back after that, and your next message has to write the whole conversation into the cache again. For a long session that is the most expensive message of the day.

Cache Keep-Warm fixes that:

  • ⏱️ See it. A line above the prompt counts down until the cache expires.
  • ♨️ Keep it. One click, and the cache stays warm while you are away: a tiny background read just before it would expire.
  • 🟢 Know where. In the desktop app, every session being kept warm gets a green dot in the session list.

<img src="docs/images/sessions.svg" alt="The desktop app's session list, with a green dot in front of the two sessions that keep-warm is on in" width="420">

What it saves

Example: a long conversation of 500k tokens on Claude Opus 5.5, at current API list prices.

Cost
Your next message after the cache expired (re-writes everything)≈ $4.00
One keep-warm ping (reads it from the cache)≈ $0.10
Keeping it warm through an 8-hour night (8 pings)≈ $0.80

On a subscription you do not pay per token, but the same proportions apply to your usage limits.

Install

The easy way: let Claude do it

Copy this into any Claude Code session:

Install the Cache Keep-Warm mod from https://github.com/andreichiritescu/claude-code-cache-keep-warm for me:
1. Run: claude plugin marketplace add https://github.com/andreichiritescu/claude-code-cache-keep-warm.git
2. Run: claude plugin install cache-keep-warm@cache-keep-warm
3. If a step fails because my Claude Code is too old, tell me to update it (the mod needs Claude Code 2.1.287 or newer, or the desktop app with 2.1.286 or newer) and stop.
4. When it is installed, tell me to run /reload-plugins so this session loads it, and that I can turn keep-warm on with the Keep warm button on the new line above the prompt, or with /keepwarm on.

Or with one command in a session

Type this at the Claude Code prompt (Claude Code 2.1.275 or newer):

/plugin install cache-keep-warm --marketplace https://github.com/andreichiritescu/claude-code-cache-keep-warm.git

Or in your terminal

claude plugin marketplace add https://github.com/andreichiritescu/claude-code-cache-keep-warm.git
claude plugin install cache-keep-warm@cache-keep-warm

Then start a new session, or run /reload-plugins in an open one. The line appears above the prompt, and the countdown starts with your first message. If Claude Code says an option is not set yet, that is the Mark the session title setting: it is on by default.

Needs Claude Code 2.1.287 or newer in a terminal, or the Claude desktop app with Claude Code 2.1.286 or newer. Type /status to check your version.

Use it

Click Keep warm on the line, or type a command:

CommandWhat it does
/keepwarm onKeep this session warm
/keepwarm 3hKeep it warm for 3 hours
/keepwarm until 18:00Keep it warm until 18:00
/keepwarm offStop (or click Stop)
/keepwarmShow the state and the last pings

Keep-warm is off by default and you switch it per session: only the sessions you want are kept warm. It stops by itself after about 18 hours (20 pings in a row) unless you gave it a time, because past that point one re-write is cheaper than more pings.

Reading the line

You seeIt means
◕ and ▰▰▰▰▰▰▰▱▱▱How much of the cache's life is left. Green while keep-warm is on, grey when it is off, and ○ once the cache has expired.
nothing cached yetA new conversation, or one you just cleared or compacted: nothing of it is cached until your next message, and the countdown starts with that.
40 min left (14:12)Time left, and the clock time the cache expires if nothing happens.
next ping 14:07When keep-warm will read the cache next. (shows in chat) means this ping is a short visible message, because the session hasn't answered since the app restarted.
3 pings ✓Pings so far, and whether the last one found the cache warm (✓) or cold (✗).
conversation nearly full: keep-warm waits for youAfter a restart the next ping has to be the visible one, and Claude Code re-sends your project's instruction files (CLAUDE.md, the files it imports, MEMORY.md) with it. Here those would not fit next to the conversation: the ping would fail or make Claude Code compact the conversation. So keep-warm waits for your own next message.
ping failed: …The last ping never reached the model, with the reason Claude Code gave. "Prompt is too long" means the conversation is over the model's limit: run /compact in that session. The cache then expires normally.
keep-warm starts after your next messageKeep-warm is on but doesn't know your cache length yet. That only happens before the first answer after you install the mod.
773k cachedThe size of your conversation, which is what your next message would re-write if the cache went cold.

How it works

flowchart LR
    A["You stop typing"] --> B{"55 minutes later:<br/>still away?"}
    B -- yes --> C["Keep-warm reads the<br/>conversation from the cache"]
    C --> D["The cache is good<br/>for another hour"]
    D --> B
    B -- "you are back" --> E["Your next message<br/>reads from the cache:<br/>cheap"]
  • It never wakes a cold cache. The ping goes out 5 minutes before the cache would expire. If the cache has already expired, keep-warm waits for your next message instead of paying to rebuild it.
  • Nothing is added to your conversation. The ping sends your conversation once more, outside the chat, with a one-line question, and throws the answer away. One exception: after the app restarts, Claude Code can repeat a conversation only once it has answered again. So in a session you haven't touched since the restart, the first ping is a short message in the chat, which Claude answers with "ok". The line says so in advance: "next ping 14:07 (shows in chat)".
  • It works while Claude is waiting for you, for example on a question it asked you or a permission prompt.
  • It is light. The line itself never calls the model: one quick local check every 15 seconds, redrawn at most once a minute. The only thing that uses your plan or credits is the ping, about once every 55 minutes, in sessions where keep-warm is on.
  1. Your settings, if you set one: promptCacheTtl, CLAUDE_CODE_PROMPT_CACHE_TTL, FORCE_PROMPT_CACHING_5M or ENABLE_PROMPT_CACHING_1H.
  2. Otherwise your plan. A Claude subscription within its usage limits gets one hour. Once you are over the limit (paying with usage credits), or on an API key or a cloud provider, it is five minutes.
  3. Then it checks what really happens. Every request shows whether it read the conversation from the cache or had to re-write it. If the cache went cold sooner than expected, keep-warm pauses and tells you, and resumes once the cache is seen lasting again. So if Claude Code ever changes how long the cache lasts, the mod follows.

Claude Code reports your plan only with an answer, so after a restart the mod uses the length your plan showed last time until the session answers again.

Keep-warm only pings a cache that lasts 30 minutes or more: on the five-minute cache, pinging every few minutes would cost more than it saves.

SettingDefaultWhat it does
Mark the session titleOnWhile keep-warm is on, put 🟢 in front of the session's title in the desktop app's session list

Change it in /config, or with /plugin configure cache-keep-warm@cache-keep-warm.

In the desktop app, renaming a title you typed yourself asks for your approval each time; titles the app generated change without asking.

I use an API key. Does keep-warm work for me? Your cache lasts five minutes by default, so keep-warm does not ping. If you take longer breaks, set promptCacheTtl to "1h" in your Claude Code settings: the cache then lasts an hour and keep-warm works. One-hour cache writes cost more than five-minute ones, so this pays off only if you do take breaks.

Will it ping forever if I forget it? No. It stops by itself after 20 pings in a row (about 18 hours), or at the time you gave it with /keepwarm 3h or /keepwarm until 18:00.

Does it work in VS Code? The keep-warm pings work there, but the line is only drawn in the terminal and in the desktop app's Code tab.

Where does the line not appear? In VS Code's chat panel, in claude -p runs and in cloud sessions.

How do I remove it? claude plugin uninstall cache-keep-warm@cache-keep-warm

What it runs, reads and sends

A mod runs inside Claude Code with your permissions, so here is everything this one does. It has no server, makes no network calls of its own, and runs no shell commands. This section is also its privacy policy: the mod collects nothing about you, and nothing reaches its author.

What leaves your computer

Only the keep-warm ping, and only while keep-warm is on: about 5 minutes before the cache would expire, so about once every 55 minutes.

The mod asks Claude Code to send this session's last request again, with one line added: Keep-warm ping. Reply with just "ok". (the $.model.fork call in hooks/register.tsx). It goes to Anthropic, through Claude Code's own connection, on your own account, like any message you send. That request reads your conversation from the cache, which is what keeps it warm. The answer is thrown away, and nothing is added to your conversation.

After an app restart, Claude Code can repeat a session's request this way only once the session has answered again. Until then the ping is a short message the mod sends into the chat instead (the $.prompt.submit call): Keep-warm ping from the Cache Keep-Warm mod. Reply with exactly "ok" and nothing else. Use no tools. Claude answers "ok", and from then on the pings are invisible again. Same destination, same account, and the same rule: only in the last minutes before the cache would expire, never once it is cold.

The mod itself never reads what your conversation says: from each request it only takes the usage figures (how many tokens were read from the cache or written to it).

The tools it calls itself

Only on the current session. Both are the Claude desktop app's own session tools, and nothing they do leaves your computer:

ToolWhat it doesWhen
mcp__ccd_session_mgmt__get_sessionReads the session's titleWhen the session starts, when you switch keep-warm on or off, and every 5 minutes
mcp__ccd_session_mgmt__set_session_titlePuts 🟢 in front of the title, or takes it offOnly when the mark has to change

With the setting Mark the session title off, the mod never adds the mark; it still reads the title at those moments, to take off a mark left from before. In a terminal these tools do not exist: the call fails quietly and nothing changes.

What it reads

  • Your Claude Code settings, for promptCacheTtl only, and three environment variables: FORCE_PROMPT_CACHING_5M, CLAUDE_CODE_PROMPT_CACHE_TTL and ENABLE_PROMPT_CACHING_1H, to know how long your cache lasts. Claude Code hands a mod the settings as one object: this mod looks at promptCacheTtl and nothing else, and keeps or sends no other setting, credential or API key.
  • This session's figures from Claude Code: the conversation's size, the cache tokens of each request, and whether your plan's usage limits are reported (which tells a subscription from an API key). After a restart, while the next ping would be the visible one, also the free space left in the context window and the size of the instruction files (their token counts only, never their text), to check the visible ping fits.
  • When you reopen a session, what Claude Code tells mods at that moment: how long ago the session's last answer came and how big the conversation is. That lets the countdown carry on before your first message.
  • The clock.

It checks these every 15 seconds, on your computer, and sends none of them anywhere. It does not read your files.

What it keeps

Per session, in Claude Code's plugin storage on your computer: the time of the last request, whether keep-warm is on, and the stop time you gave it. A session untouched for a week is forgotten. Shared by your sessions, it also keeps the last cache length your plan showed (one hour or five minutes), so a session you reopen knows it before its first answer.

The events it hooks

EventWhat the mod does with it
Session startAdds the /keepwarm command and starts its 15-second local check
A session is reopened or cleared (Claude Code's SessionStart hook)Reads how long ago the last answer came and the conversation's size, so the countdown carries on; after /clear, starts the countdown over. Changes nothing.
A compaction (/compact or automatic)Checks only whose it is and whether it went through: when your conversation itself was compacted, starts the countdown over. Ignores a subagent's compaction of its own work and a summary prepared ahead of time. Never reads or changes the summary or your messages.
Each request of the main conversationNotes the time and how much was read from the cache. Changes nothing. Subagents' requests are ignored.
A turn endsNotes that an answer arrived (which helps tell a subscription from an API key)
Claude's AskUserQuestion toolShows "waiting for your answer" while the question is open. Never changes the question or your answer.
The /keepwarm commandAnswers its own command, and no other
The line above the promptDraws the countdown line

To check all of this before you install it, clone the repository and run:

claude plugin validate ./claude-code-cache-keep-warm/.claude-plugin/plugin.json

The hooks:, calls: and env reads: lines list every event the mod handles, everything it asks Claude Code to do, and every environment variable it reads.


MIT License · A community mod, not made by Anthropic · Issues and ideas welcome

Source 2 files
hooks/register.tsx 722 lines
1import type { EngineInterface, ModelUsage, Register, SessionUsage } from 'claude-code'
2
3import type { KeepWarmPing, KeepWarmSaved, KeepWarmTtl } from '../types'
4
5const HOUR_MS = 60 * 60_000
6const FIVE_MIN_MS = 5 * 60_000
7// The cache lengths Claude Code offers today; a value seen elsewhere is still accepted.
8const KNOWN_TTLS_MS = [FIVE_MIN_MS, HOUR_MS]
9// Keep-warm only for a cache this long or longer: on a short cache, pinging every few
10// minutes costs more than the single re-write it saves after any real break.
11const KEEP_WARM_MIN_TTL_MS = 30 * 60_000
12// A ping goes out this long before the cache would expire (55 min into a 1-hour cache), and
13// never in its last minute: a cold cache is never pinged. A later ping means fewer of them
14// (one per ~55 min instead of one per 45), and the lead still leaves four minutes of retries
15// if the API is busy. The cache's life runs from when a request reaches the server, a moment
16// after the time recorded here, so the real deadline is slightly later than computed.
17const PING_LEAD_MS = 5 * 60_000
18const PING_MARGIN_MS = 60_000
19// After a failed ping, wait this long before trying again inside the window.
20const PING_RETRY_MS = 30_000
21// Without an explicit stop time, keep-warm stops itself after this many pings in a row
22// (~15 h): past ~30 h idle, re-writing the cache once is cheaper than pinging.
23const MAX_PINGS = 20
24// One check every 15 s: local reads only (clock, session figures, settings), no model call.
25// The line is redrawn only when its text changes, at most once a minute.
26const TICK_MS = 15_000
27const PING_PENDING_MS = 5 * 60_000
28// A request reading at least this share of the conversation from cache found it still warm.
29const HIT_SHARE = 0.5
30// A re-write after a shorter gap is a change to the prompt (compaction, a model switch),
31// not an expiry, so it says nothing about the cache's length.
32const EVIDENCE_GAP_MS = FIVE_MIN_MS + 30_000
33// A reopened session reports how long ago its last answer came. Its cache's life began when
34// that last request was sent, a little earlier: counting from this much before the answer keeps
35// the ping early rather than late (a late ping would re-write a cache that already went cold).
36const RESUME_MARGIN_MS = 3 * 60_000
37// With keep-warm off, the band names what is at stake once this little time is left.
38const EXPIRING_MS = 10 * 60_000
39// Blocks that empty over the cache's life: filled blocks green while keep-warm is on (grey
40// otherwise), outlined blocks for the time already gone.
41const BAR_CELLS = 10
42const BAR_FULL = '▰'
43const BAR_EMPTY = '▱'
44// Space between the details, in the surface's own units (columns on the terminal).
45const DETAIL_GAP = 2
46// The circle in front empties with the cache's life, a quarter at a time; hollow once cold.
47// The full one is the "medium black circle" in text style: the plain ● is drawn smaller than
48// the quarter circles in the desktop app's font (picked by the user from a test row).
49const CIRCLE_BY_QUARTER = ['◔', '◑', '◕', '⚫︎']
50const CIRCLE_COLD = '○'
51// While keep-warm is on, the session's title in the app's list starts with this mark. The app
52// asks before changing a title you typed yourself; titles it generated change without asking.
53const TITLE_MARK = '🟢 '
54// The userConfig key in plugin.json that switches the title mark off.
55const TITLE_MARK_OPTION = 'title_mark'
56const TITLE_SYNC_MS = 5 * 60_000
57const GET_SESSION_TOOL = 'mcp__ccd_session_mgmt__get_session'
58const SET_TITLE_TOOL = 'mcp__ccd_session_mgmt__set_session_title'
59const TITLE_FIELD = /"title"\s*:\s*"((?:[^"\\]|\\.)*)"/
60const HISTORY_SHOWN = 5
61const PING_TEXT = 'Keep-warm ping. Reply with just "ok".'
62const COMMAND = 'keepwarm'
63const STORE_PREFIX = 'session:'
64const STORE_MAX_AGE_MS = 7 * 24 * HOUR_MS
65// Shared by every session: the cache length the plan showed last. Claude Code reports the plan
66// only with a response, so a session reopened after a relaunch would otherwise not know its
67// cache length (and could not keep it warm) until its first answer.
68const PLAN_TTL_KEY = 'plan-ttl'
69// Rate-limit windows that only a Claude subscription reports; an API key or a cloud provider reports none.
70const PLAN_WINDOWS = ['five_hour', 'seven_day']
71const PLAN_LIMIT_PERCENT = 100
72const TTL_VALUE = /^(\d+)\s*(s|m|h)$/
73const DURATION = /^(\d+(?:\.\d+)?)\s*(h|m|min)$/
74const UNTIL = /^until\s+(\d{1,2}):(\d{2})$/
75const WAITS_TEXT = 'keep-warm starts after your next message'
76// After an app restart Claude Code can repeat (fork) a conversation only once it has answered in
77// this run. Until then the ping is a real message instead: it shows in the chat, Claude answers
78// "ok", and from then on the pings are invisible again. Same window, so never on a cold cache.
79const VISIBLE_PING_TEXT = 'Keep-warm ping from the Cache Keep-Warm mod. Reply with exactly "ok" and nothing else. Use no tools.'
80const SHOWS_IN_CHAT = ' (shows in chat)'
81// The visible ping is a real turn, and after a restart Claude Code re-sends the project's
82// instruction files (CLAUDE.md, the files it imports, MEMORY.md) with it: 130k tokens in a big
83// project. If the conversation, those files and this much room for the ping and its reply do not
84// fit in the window's free space (before the compaction reserve), the turn would fail ("Prompt is
85// too long") or be compacted, which a ping must never cause: keep-warm then waits for the
86// person's own message instead.
87const VISIBLE_PING_ROOM = 30_000
88// Only where Claude Code gives no breakdown of the window: wait above this share of it.
89const VISIBLE_PING_MAX_FILL = 0.75
90const NEARLY_FULL_TEXT = 'conversation nearly full: keep-warm waits for you'
91
92const LAST = { plugin: 'cache-keep-warm', key: 'lastRequestAt' } as const
93const KEEP_WARM = { plugin: 'cache-keep-warm', key: 'keepWarm' } as const
94const STOP_AT = { plugin: 'cache-keep-warm', key: 'stopAt' } as const
95const NOW = { plugin: 'cache-keep-warm', key: 'now' } as const
96const PINGS = { plugin: 'cache-keep-warm', key: 'pings' } as const
97const TTL = { plugin: 'cache-keep-warm', key: 'ttl' } as const
98const CONTEXT = { plugin: 'cache-keep-warm', key: 'contextTokens' } as const
99const WAITING = { plugin: 'cache-keep-warm', key: 'isWaiting' } as const
100const SAW_RESPONSE = { plugin: 'cache-keep-warm', key: 'sawResponse' } as const
101const SHORT_AFTER = { plugin: 'cache-keep-warm', key: 'shortAfterMs' } as const
102const WAITS = { plugin: 'cache-keep-warm', key: 'waitsForMessage' } as const
103const RESTORED = { plugin: 'cache-keep-warm', key: 'restored' } as const
104const PING_ERROR = { plugin: 'cache-keep-warm', key: 'pingError' } as const
105const NEARLY_FULL = { plugin: 'cache-keep-warm', key: 'nearlyFull' } as const
106
107// The module's own variables start over on a reload; what must survive one is in $.state.
108// `titleMark` is the user's "Mark the session title" setting; changing it in /config reloads
109// the module, so it is read once in register(). `planTtl` mirrors the shared PLAN_TTL_KEY;
110// `savedTtl` is the length this session's own cache had when the store last saw it.
111// `visiblePingAt`: when a visible ping was submitted, until its request is seen.
112// `pingTurnId`: the visible ping's turn, which is no sign of use when it ends.
113// `roomChecked`: whether this process has worked out if the visible ping fits.
114// `isFirstRequest`: the first request after a start says nothing about the cache's length.
115const live: {
116  storeKey: string
117  pingPendingUntil: number
118  shownKey: string
119  titleMark: boolean
120  planTtl: KeepWarmTtl | undefined
121  savedTtl: KeepWarmTtl | undefined
122  isFirstRequest: boolean
123  visiblePingAt: number
124  pingTurnId: string
125  roomChecked: boolean
126} = {
127  storeKey: '',
128  pingPendingUntil: 0,
129  shownKey: '',
130  titleMark: true,
131  planTtl: undefined,
132  savedTtl: undefined,
133  isFirstRequest: true,
134  visiblePingAt: 0,
135  pingTurnId: '',
136  roomChecked: false,
137}
138
139function minutes(ms: number) {
140  return Math.max(0, Math.floor(ms / 60_000))
141}
142
143function clockTime(ms: number) {
144  const d = new Date(ms)
145  return `${String(d.getHours()).padStart(2, '0')}:${String(d.getMinutes()).padStart(2, '0')}`
146}
147
148function tokens(n: number) {
149  return n >= 1_000_000 ? `${(n / 1_000_000).toFixed(1)}M` : `${Math.round(n / 1000)}k`
150}
151
152function timeLeft(ms: number) {
153  return ms < 60_000 ? 'under a minute left' : `${minutes(ms)} min left`
154}
155
156function cacheLength(ms: number) {
157  return ms >= HOUR_MS && ms % HOUR_MS === 0 ? `${ms / HOUR_MS}-hour` : `${minutes(ms)}-minute`
158}
159
160// "5m", "1h", "30m", "90s": any length, so a cache length Claude Code adds later still reads.
161function parseTtl(value: unknown) {
162  if (typeof value !== 'string') return 0
163  const m = TTL_VALUE.exec(value.trim())
164  if (!m) return 0
165  return Number(m[1]) * (m[2] === 'h' ? HOUR_MS : m[2] === 'm' ? 60_000 : 1000)
166}
167
168// Keeps what the plan showed for the next session that opens before its first answer; written
169// only when it changes, though this runs on every check.
170async function rememberPlan($: EngineInterface, ttl: NonNullable<KeepWarmTtl>): Promise<KeepWarmTtl> {
171  if (live.planTtl?.ms !== ttl.ms || live.planTtl?.reason !== ttl.reason) {
172    live.planTtl = ttl
173    await $.store.set(PLAN_TTL_KEY, ttl)
174  }
175  return ttl
176}
177
178// The same precedence Claude Code itself applies, then the plan: a subscription within its
179// limits gets the hour, one over its limit (paying usage credits) and an API key get 5 minutes.
180async function detectTtl($: EngineInterface, usage: SessionUsage): Promise<KeepWarmTtl> {
181  if ((await $.env.get('FORCE_PROMPT_CACHING_5M')) === '1') return { ms: FIVE_MIN_MS, reason: 'setting' }
182  const fromEnv = parseTtl(await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL'))
183  if (fromEnv) return { ms: fromEnv, reason: 'setting' }
184  const settings = await $.settings.read()
185  const fromSettings = parseTtl(settings['promptCacheTtl'])
186  if (fromSettings) return { ms: fromSettings, reason: 'setting' }
187  if ((await $.env.get('ENABLE_PROMPT_CACHING_1H')) === '1') return { ms: HOUR_MS, reason: 'setting' }
188
189  const plan = usage.rateLimits.filter(w => PLAN_WINDOWS.includes(w.kind))
190  if (plan.length === 0) {
191    // The windows arrive with the first response, so before one an empty list proves nothing:
192    // until then, go by the length this session's cache had when last saved (what its last
193    // request really got), else by what the plan showed last time in any session.
194    const { value: sawResponse = false } = await $.state.get(SAW_RESPONSE)
195    if (!sawResponse) return live.savedTtl ?? live.planTtl ?? null
196    return rememberPlan($, { ms: FIVE_MIN_MS, reason: 'api' })
197  }
198  return rememberPlan(
199    $,
200    plan.some(w => w.percentUsed >= PLAN_LIMIT_PERCENT) ? { ms: FIVE_MIN_MS, reason: 'overage' } : { ms: HOUR_MS, reason: 'subscription' },
201  )
202}
203
204// What the cache actually did beats what the settings say: if it went cold after a gap it
205// should have survived, use the longest known length shorter than that gap.
206function correctTtl(base: KeepWarmTtl, shortAfterMs: number): KeepWarmTtl {
207  if (!base || shortAfterMs <= 0 || shortAfterMs >= base.ms - PING_MARGIN_MS) return base
208  const shorter = KNOWN_TTLS_MS.filter(ms => ms < shortAfterMs)
209  return { ms: shorter.length ? Math.max(...shorter) : FIVE_MIN_MS, reason: 'observed' }
210}
211
212// Each request after an idle gap shows whether the cache outlived that gap: a read clears an
213// earlier "went cold too soon"; a re-write after a gap the cache should have survived records one.
214async function observe($: EngineInterface, gapMs: number, usage: ModelUsage) {
215  const total = usage.cache_read_input_tokens + usage.cache_creation_input_tokens
216  if (total === 0 || gapMs < EVIDENCE_GAP_MS) return
217  if (usage.cache_read_input_tokens >= HIT_SHARE * total) {
218    await $.state.set(SHORT_AFTER, 0)
219    return
220  }
221  const { value: ttl = null } = await $.state.get(TTL)
222  const { value: shortAfter = 0 } = await $.state.get(SHORT_AFTER)
223  const expected = ttl?.reason === 'observed' ? HOUR_MS : (ttl?.ms ?? HOUR_MS)
224  if (gapMs < expected - PING_MARGIN_MS) await $.state.set(SHORT_AFTER, shortAfter > 0 ? Math.min(shortAfter, gapMs) : gapMs)
225}
226
227async function save($: EngineInterface) {
228  if (!live.storeKey) return
229  const { value: lastRequestAt = 0 } = await $.state.get(LAST)
230  const { value: keepWarm = false } = await $.state.get(KEEP_WARM)
231  const { value: stopAt = 0 } = await $.state.get(STOP_AT)
232  const { value: ttl = null } = await $.state.get(TTL)
233  const saved: KeepWarmSaved = { lastRequestAt, keepWarm, stopAt, ttl, savedAt: await $.clock.now() }
234  await $.store.set(live.storeKey, saved)
235}
236
237async function setKeepWarm($: EngineInterface, isOn: boolean, stopAt: number) {
238  await $.state.set(KEEP_WARM, isOn)
239  await $.state.set(STOP_AT, isOn ? stopAt : 0)
240  await $.state.set(PINGS, [])
241  await $.state.set(PING_ERROR, '')
242  await save($)
243  await syncTitle($)
244}
245
246// The conversation starts over (cleared, or replaced by its summary): nothing of it is cached
247// until the next request writes it.
248async function startOver($: EngineInterface) {
249  await $.state.set(LAST, 0)
250  await $.state.set(PINGS, [])
251  await $.state.set(SAW_RESPONSE, false)
252  await $.state.set(CONTEXT, 0)
253  await $.state.set(PING_ERROR, '')
254  await $.state.set(NOW, await $.clock.now())
255  live.roomChecked = false
256  const { value: restored = false } = await $.state.get(RESTORED)
257  if (restored) await save($)
258}
259
260// Puts the mark on this session's title in the app's list while keep-warm is on, and takes it
261// off otherwise (or always, with the setting off). Runs on every switch, at start (which also
262// clears a mark a crash left behind) and every few minutes (a title the app regenerated loses
263// the mark). Outside the desktop app there are no session tools and nothing is marked.
264async function syncTitle($: EngineInterface) {
265  const { value: isOn = false } = await $.state.get(KEEP_WARM)
266  const { value: ttl = null } = await $.state.get(TTL)
267  const wantsMark = live.titleMark && isOn && !(ttl && ttl.ms < KEEP_WARM_MIN_TTL_MS)
268  try {
269    const found = await $.tool.call({ tool: GET_SESSION_TOOL, session_id: 'self' })
270    if ('deny' in found && found.deny) return
271    const match = TITLE_FIELD.exec(found.text ?? '')
272    if (!match) return
273    const title: string = JSON.parse(`"${match[1]}"`)
274    const hasMark = title.startsWith(TITLE_MARK)
275    if (wantsMark === hasMark) return
276    const base = hasMark ? title.slice(TITLE_MARK.length) : title
277    await $.tool.call({ tool: SET_TITLE_TOOL, session_id: 'self', title: wantsMark ? TITLE_MARK + base : base })
278  } catch {
279    // No session tools here (a terminal session), or the app declined the rename: leave the title.
280  }
281}
282
283// The ping re-sends the main thread's last request outside the conversation with one short
284// question after it: it reads the conversation from cache (restarting its life), adds no
285// message, and runs even while the session waits on a question or a permission prompt.
286async function ping($: EngineInterface, startedAt: number, lastRequestAt: number) {
287  const r = await $.model.fork({ prompt: PING_TEXT })
288  if (!r.isAnswered && r.reason === 'nothing-to-fork') {
289    // Claude Code has no request of this conversation to repeat yet: send the visible one.
290    await $.state.set(WAITS, true)
291    const { value: contextTokens = 0 } = await $.state.get(CONTEXT)
292    const { context } = await $.session.usage()
293    if (!(await checkRoom($, context.window, contextTokens))) await visiblePing($, startedAt)
294    return
295  }
296  if (!r.isAnswered && r.reason === 'api-error') {
297    // A busy or failing API may answer on a retry; a request the API refuses (the conversation
298    // over the model's limit, for one) gets the same answer every time, so stop and say so.
299    const status = r.status ?? 0
300    if (status >= 400 && status < 500 && status !== 429) {
301      await $.state.set(PING_ERROR, `HTTP ${status}`)
302      $.ui.toast(`Keep-warm ping failed (HTTP ${status}); the cache will expire normally.`)
303      return
304    }
305    $.ui.toast(`Keep-warm ping failed (HTTP ${status || '?'}); retrying in ${PING_RETRY_MS / 1000} s.`)
306    live.pingPendingUntil = startedAt + PING_RETRY_MS
307    return
308  }
309  const { cache_read_input_tokens: read, cache_creation_input_tokens: written } = r.usage
310  const isHit = read > 0 && read >= HIT_SHARE * (read + written)
311  const done: KeepWarmPing = { at: startedAt, isHit, read }
312  const { value: pings = [] } = await $.state.get(PINGS)
313  await $.state.set(PINGS, [...pings, done].slice(-MAX_PINGS))
314  // A ping that found the cache cold proves the cache is shorter than assumed, which pauses
315  // keep-warm (another cold ping would only pay for a full re-write again).
316  await observe($, startedAt - lastRequestAt, r.usage)
317  await $.state.set(LAST, startedAt)
318  await save($)
319}
320
321// Whether the visible ping would not fit: the conversation, the instruction files Claude Code
322// re-sends after a restart, and the ping with its reply must all fit in the window's free space.
323// Claude Code estimates the breakdown locally (no request), from the same figures as /context.
324async function visiblePingWontFit($: EngineInterface, window: number, contextTokens: number) {
325  try {
326    const { breakdown } = (await $.session.usage({ breakdown: 'summary' })).context
327    if (breakdown) {
328      const free = breakdown.categories.filter(c => c.kind === 'free').reduce((n, c) => n + c.tokens, 0)
329      const instructions = breakdown.memoryFiles.reduce((n, f) => n + f.tokens, 0)
330      return free < instructions + VISIBLE_PING_ROOM
331    }
332  } catch {
333    // No breakdown here: go by the share of the window below.
334  }
335  return window > 0 && contextTokens >= VISIBLE_PING_MAX_FILL * window
336}
337
338// Works out once per start, and only while the next ping would be the visible one, whether it
339// fits; the line shows the answer before the ping is due.
340async function checkRoom($: EngineInterface, window: number, contextTokens: number) {
341  live.roomChecked = true
342  const wontFit = await visiblePingWontFit($, window, contextTokens)
343  const { value: was = false } = await $.state.get(NEARLY_FULL)
344  if (wontFit !== was) await $.state.set(NEARLY_FULL, wontFit)
345  return wontFit
346}
347
348// The fallback ping: a real prompt, so it shows in the chat with Claude's "ok". Its request reads
349// the conversation from cache like the invisible ping, and the turn.step hook records it.
350async function visiblePing($: EngineInterface, startedAt: number) {
351  live.visiblePingAt = startedAt
352  const r = await $.prompt.submit({ text: VISIBLE_PING_TEXT })
353  if (typeof r.drop === 'string') {
354    live.visiblePingAt = 0
355    $.ui.toast(`Keep-warm ping not sent: ${r.drop}`)
356  }
357}
358
359async function tick($: EngineInterface) {
360  const t = await $.clock.now()
361  const usage = await $.session.usage()
362  const { value: shortAfter = 0 } = await $.state.get(SHORT_AFTER)
363  const ttl = correctTtl(await detectTtl($, usage), shortAfter)
364  const { value: last = 0 } = await $.state.get(LAST)
365
366  // Redraw only when what the band shows would change: each minute, and whenever the cache
367  // length or the context size moves. A reopened session reports its size before its first
368  // answer, while the engine has none yet: keep that one until the engine has its own.
369  const { value: knownContext = 0 } = await $.state.get(CONTEXT)
370  const contextTokens = usage.context.tokens ?? knownContext
371  const shownKey = `${Math.floor(t / 60_000)}|${ttl?.ms ?? 0}|${ttl?.reason ?? ''}|${contextTokens}|${last}|${shortAfter}`
372  if (shownKey !== live.shownKey) {
373    live.shownKey = shownKey
374    await $.state.set(TTL, ttl)
375    await $.state.set(CONTEXT, contextTokens)
376    await $.state.set(NOW, t)
377  }
378
379  const { value: isOn = false } = await $.state.get(KEEP_WARM)
380  if (!isOn) return
381  // Invisible once the conversation has answered in this run; before that, the visible ping,
382  // whose fit is worked out once so the line can say in advance when it has to wait.
383  const { value: sawResponse = false } = await $.state.get(SAW_RESPONSE)
384  const { value: waits = false } = await $.state.get(WAITS)
385  const needsVisible = !sawResponse || waits
386  if (needsVisible && !live.roomChecked) await checkRoom($, usage.context.window, contextTokens)
387  const { value: stopAt = 0 } = await $.state.get(STOP_AT)
388  if (stopAt > 0 && t >= stopAt) {
389    await setKeepWarm($, false, 0)
390    $.ui.toast(`Keep-warm stopped at ${clockTime(stopAt)} as asked; the cache will expire normally.`)
391    return
392  }
393  const idle = t - last
394  const isPingWindow =
395    ttl !== null &&
396    ttl.ms >= KEEP_WARM_MIN_TTL_MS &&
397    last > 0 &&
398    idle >= ttl.ms - PING_LEAD_MS &&
399    idle < ttl.ms - PING_MARGIN_MS
400  if (!isPingWindow || t < live.pingPendingUntil) return
401
402  const { value: pings = [] } = await $.state.get(PINGS)
403  if (stopAt === 0 && pings.length >= MAX_PINGS) {
404    await setKeepWarm($, false, 0)
405    $.ui.toast(`Keep-warm stopped itself after ${MAX_PINGS} pings in a row; the cache will expire.`)
406    return
407  }
408  live.pingPendingUntil = t + PING_PENDING_MS
409  const { value: nearlyFull = false } = await $.state.get(NEARLY_FULL)
410  if (!needsVisible) await ping($, t, last)
411  else if (!nearlyFull) await visiblePing($, t)
412}
413
414// "3h", "90m", "until 18:00" -> when keep-warm stops; 0 = no stop time; -1 = not understood.
415function parseStop(text: string, now: number) {
416  if (text === '') return 0
417  const duration = DURATION.exec(text)
418  if (duration) return now + Number(duration[1]) * (duration[2] === 'h' ? HOUR_MS : 60_000)
419  const until = UNTIL.exec(text)
420  if (!until) return -1
421  const at = new Date(now)
422  at.setHours(Number(until[1]), Number(until[2]), 0, 0)
423  return at.getTime() > now ? at.getTime() : at.getTime() + 24 * HOUR_MS
424}
425
426export const register: Register = (on, options) => {
427  live.titleMark = options[TITLE_MARK_OPTION] !== false
428
429  on('session.start', async ($, e, next) => {
430    await $.command.register({
431      name: COMMAND,
432      description: 'Keep this session\'s prompt cache warm: on [3h | until 18:00], off, or status',
433      argumentHint: '[on [3h|until 18:00] | off]',
434      immediate: true,
435    })
436
437    live.storeKey = STORE_PREFIX + (await $.session.id())
438    live.isFirstRequest = true
439    const plan = (await $.store.get(PLAN_TTL_KEY)) as KeepWarmTtl | undefined
440    if (plan && typeof plan.ms === 'number') live.planTtl = plan
441
442    // After an app relaunch the session's values are gone; the store keeps them per session id.
443    // A reopened session may already have its last answer's time from Claude Code (the
444    // classic.SessionStart hook below, which can run before or after this one): keep the later.
445    const { value: restored = false } = await $.state.get(RESTORED)
446    if (!restored) {
447      const saved = (await $.store.get(live.storeKey)) as KeepWarmSaved | undefined
448      if (saved) {
449        if (saved.ttl && typeof saved.ttl.ms === 'number') live.savedTtl = saved.ttl
450        const { value: last = 0 } = await $.state.get(LAST)
451        if (saved.lastRequestAt > last) await $.state.set(LAST, saved.lastRequestAt)
452        await $.state.set(KEEP_WARM, saved.keepWarm)
453        await $.state.set(STOP_AT, saved.stopAt ?? 0)
454      }
455      await $.state.set(RESTORED, true)
456      await save($)
457    }
458
459    // Forget sessions untouched for a week so the store stays small. Aged by the last save: a
460    // session /compact or /clear just reset has no last request, yet is in use (counting from 0
461    // deleted its entry, keep-warm switch included). An old entry with no date at all is kept
462    // while its keep-warm is on.
463    const t = await $.clock.now()
464    for (const key of await $.store.keys()) {
465      if (!key.startsWith(STORE_PREFIX) || key === live.storeKey) continue
466      const old = (await $.store.get(key)) as KeepWarmSaved | undefined
467      const at = old ? (old.savedAt ?? old.lastRequestAt) : 0
468      if (!old || (at > 0 ? t - at > STORE_MAX_AGE_MS : !old.keepWarm)) await $.store.delete(key)
469    }
470
471    await tick($)
472    $.clock.every(TICK_MS, () => void tick($))
473    void syncTitle($)
474    $.clock.every(TITLE_SYNC_MS, () => void syncTitle($))
475
476    return next(e)
477  })
478
479  // A session reopened (an app relaunch, a resume): Claude Code says how long ago its last
480  // answer came and how big the conversation is, so the countdown carries on before the first
481  // message. After /clear the conversation starts over: no cache to count down yet. Its
482  // 'compact' is not used: it also comes when a subagent compacts its own conversation, with
483  // nothing saying so (no agent_id), which reset the main countdown. session.compact says whose.
484  on('classic.SessionStart', async ($, e, next) => {
485    if (e.source === 'clear') await startOver($)
486    else if ((e.source === 'resume' || e.source === 'fork') && typeof e.seconds_since_last_response === 'number') {
487      const t = await $.clock.now()
488      const lastAt = t - e.seconds_since_last_response * 1000 - RESUME_MARGIN_MS
489      const { value: last = 0 } = await $.state.get(LAST)
490      if (lastAt > last) await $.state.set(LAST, lastAt)
491      if (e.context_tokens) await $.state.set(CONTEXT, e.context_tokens)
492      await $.state.set(NOW, t)
493      // At a relaunch this can run before session.start has restored the saved values: saving
494      // now would write the defaults (keep-warm off) over them. session.start saves after it.
495      const { value: restored = false } = await $.state.get(RESTORED)
496      if (restored) await save($)
497    }
498    return next(e)
499  })
500
501  // A compaction of the main conversation (/compact, the threshold, a plugin) replaces it with its
502  // summary, so the countdown starts over. A subagent's own (agentId) is not this cache, a
503  // precomputed one is not applied yet, and a skipped one changed nothing. Read only: the
504  // compaction passes on unchanged.
505  on('session.compact', async ($, e, next) => {
506    const result = await next(e)
507    if (!e.agentId && e.trigger !== 'precompute' && result.skip === undefined) {
508      try {
509        await startOver($)
510      } catch {
511        // Never fail a compaction over the countdown.
512      }
513    }
514    return result
515  })
516
517  // Each model request of the main conversation (not a subagent's: its cache is its own):
518  // it restarts the cache's life, and its usage shows whether the cache outlived the gap.
519  on('turn.step', async function* ($, e, next) {
520    if (e.agentId) return yield* next(e)
521    const t = await $.clock.now()
522    const { value: previous = 0 } = await $.state.get(LAST)
523    live.pingPendingUntil = 0
524    await $.state.set(LAST, t)
525    await $.state.set(NOW, t)
526    await save($)
527    const result = yield* next(e)
528    // A request that never reached the model (an error, the conversation over its limit) did not
529    // touch the cache, so its life still runs from the request before.
530    const u = result?.usage
531    if (!u || u.input_tokens + u.cache_read_input_tokens + u.cache_creation_input_tokens === 0) {
532      await $.state.set(LAST, previous)
533      await save($)
534      return result
535    }
536    await $.state.set(WAITS, false)
537    await $.state.set(PING_ERROR, '')
538    if (live.visiblePingAt > 0) {
539      const { cache_read_input_tokens: read, cache_creation_input_tokens: written } = u
540      const done: KeepWarmPing = { at: live.visiblePingAt, isHit: read > 0 && read >= HIT_SHARE * (read + written), read }
541      const { value: pings = [] } = await $.state.get(PINGS)
542      await $.state.set(PINGS, [...pings, done].slice(-MAX_PINGS))
543      live.visiblePingAt = 0
544      live.pingTurnId = e.turnId
545    }
546    // Not the first request after a start: a restart can change the prompt itself (a new
547    // Claude Code version), which re-writes the cache whatever the gap, so it proves nothing.
548    const isFirst = live.isFirstRequest
549    live.isFirstRequest = false
550    if (previous > 0 && !isFirst) await observe($, t - previous, u)
551    return result
552  })
553
554  // A turn with real requests behind it: the conversation has answered in this run (so the next
555  // ping can be invisible, and an empty plan list now means an API key), and unless it was the
556  // mod's own visible ping, the conversation is in use, yours or a turn it ran by itself (an
557  // agent's report, another session's message): the pings in a row start over. A turn that died
558  // on an error has no usage; if it was the visible ping, the line says why it failed.
559  on('turn.complete', async ($, e, next) => {
560    if (!e.agentId) {
561      const isPingTurn = e.turnId === live.pingTurnId
562      if (isPingTurn) live.pingTurnId = ''
563      if (e.usage) {
564        await $.state.set(SAW_RESPONSE, true)
565        if (!isPingTurn) await $.state.set(PINGS, [])
566      } else if (live.visiblePingAt > 0) {
567        live.visiblePingAt = 0
568        if (e.reason === 'error') {
569          const why = e.answer.trim() || 'an API error'
570          await $.state.set(PING_ERROR, why)
571          $.ui.toast(`Keep-warm ping failed: ${why}. The cache will expire normally.`)
572        }
573      }
574    }
575    return next(e)
576  })
577
578  // While Claude's question waits for an answer the turn stays open; the band says so.
579  on('tool.call', { tool: 'AskUserQuestion' }, async ($, e, next) => {
580    await $.state.set(WAITING, true)
581    try {
582      return await next(e)
583    } finally {
584      await $.state.set(WAITING, false)
585    }
586  })
587
588  on('command.run', { command: COMMAND }, async ($, e) => {
589    const t = await $.clock.now()
590    const words = e.args.trim().toLowerCase()
591    const [first = ''] = words.split(/\s+/)
592    if (first === 'off') {
593      await setKeepWarm($, false, 0)
594    } else if (words !== '' && words !== 'status') {
595      const stopAt = parseStop(first === 'on' ? words.slice(2).trim() : words, t)
596      if (stopAt < 0) return { text: 'Not understood. Try: /keepwarm on, /keepwarm 3h, /keepwarm until 18:00, /keepwarm off' }
597      await setKeepWarm($, true, stopAt)
598    }
599
600    const { value: isOn = false } = await $.state.get(KEEP_WARM)
601    const { value: stopAt = 0 } = await $.state.get(STOP_AT)
602    const { value: ttl = null } = await $.state.get(TTL)
603    const { value: pings = [] } = await $.state.get(PINGS)
604    const { value: waits = false } = await $.state.get(WAITS)
605    const { value: sawResponse = false } = await $.state.get(SAW_RESPONSE)
606    const lines = [
607      isOn
608        ? `Keep-warm on${stopAt ? ` until ${clockTime(stopAt)}` : ` (stops itself after ${MAX_PINGS} pings in a row)`}: about ${minutes(PING_LEAD_MS)} min before the cache would expire, it re-reads the conversation from cache in the background. A cold cache is never pinged.`
609        : 'Keep-warm off: the cache expires normally.',
610    ]
611    if (isOn && (waits || !sawResponse)) lines.push('Claude Code can repeat this conversation invisibly only once it has answered since the app started, so the next ping is a short message in the chat.')
612    if (ttl) lines.push(`This session's cache lasts ${cacheLength(ttl.ms)} (${ttl.reason === 'observed' ? 'seen expiring sooner than expected' : ttl.reason}).`)
613    if (ttl && ttl.ms < KEEP_WARM_MIN_TTL_MS) lines.push(`Keep-warm only pings a cache of ${minutes(KEEP_WARM_MIN_TTL_MS)} min or longer.`)
614    for (const p of pings.slice(-HISTORY_SHOWN)) {
615      lines.push(`${clockTime(p.at)} · ${p.isHit ? `read ${tokens(p.read)} from cache` : 'found the cache cold'}`)
616    }
617    return { text: lines.join('\n') }
618  })
619
620  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
621    if (e.props.hasSurvey) return next(e)
622
623    const { Box, Button, Text } = $.ui.resolve(e)
624    const { value: last = 0 } = await $.state.get(LAST)
625    const { value: now = 0 } = await $.state.get(NOW)
626    const { value: isOn = false } = await $.state.get(KEEP_WARM)
627    const { value: stopAt = 0 } = await $.state.get(STOP_AT)
628    const { value: pings = [] } = await $.state.get(PINGS)
629    const { value: ttl = null } = await $.state.get(TTL)
630    const { value: contextTokens = 0 } = await $.state.get(CONTEXT)
631    const { value: isWaiting = false } = await $.state.get(WAITING)
632    const { value: shortAfter = 0 } = await $.state.get(SHORT_AFTER)
633    const { value: waits = false } = await $.state.get(WAITS)
634    const { value: sawResponse = false } = await $.state.get(SAW_RESPONSE)
635    const { value: pingError = '' } = await $.state.get(PING_ERROR)
636    const { value: nearlyFull = false } = await $.state.get(NEARLY_FULL)
637    // The next ping has to be the visible one but the conversation is too full for it.
638    const isBlocked = (!sawResponse || waits) && nearlyFull
639
640    const t = Math.max(now, last)
641    const ttlMs = ttl?.ms ?? HOUR_MS
642    const canKeepWarm = ttl !== null && ttl.ms >= KEEP_WARM_MIN_TTL_MS
643    // Green only while keep-warm is on and able to ping; everything else stays default grey.
644    const isGreen = isOn && canKeepWarm && !pingError && !isBlocked
645    const expiresAt = last + ttlMs
646    const remaining = expiresAt - t
647    const isWarm = last > 0 && remaining > 0
648    const filled = isWarm ? Math.max(1, Math.ceil((remaining / ttlMs) * BAR_CELLS)) : 0
649    const quarters = CIRCLE_BY_QUARTER.length
650    const circle = isWarm
651      ? CIRCLE_BY_QUARTER[Math.min(quarters, Math.max(1, Math.ceil((remaining / ttlMs) * quarters))) - 1]
652      : CIRCLE_COLD
653
654    // The first part (time left) is drawn bold; the dot and the button already say whether
655    // keep-warm is on, so the text no longer repeats it.
656    const main = last === 0 ? 'nothing cached yet' : isWarm ? timeLeft(remaining) : `cold since ${clockTime(expiresAt)}`
657    // "58 min left (14:12)": the clock time the cache expires, right after the time left.
658    const expiresText = isWarm ? ` (${clockTime(expiresAt)})` : ''
659    const parts: string[] = []
660    if (isWaiting && isWarm) parts.push('waiting for your answer')
661
662    if (ttl && !canKeepWarm) {
663      if (ttl.reason === 'observed') parts.push(`cache expired after ${minutes(shortAfter)} min, sooner than expected`, 'keep-warm paused')
664      else if (ttl.reason === 'overage') parts.push(`${cacheLength(ttl.ms)} cache while over your plan limit`, 'keep-warm paused')
665      else parts.push(`${cacheLength(ttl.ms)} cache (${ttl.reason === 'api' ? 'API key' : 'your settings'})`, 'keep-warm needs the 1-hour cache')
666    } else if (isOn) {
667      const lastPing = pings[pings.length - 1]
668      if (stopAt) parts.push(`keep-warm stops ${clockTime(stopAt)}`)
669      // With the cache length unknown no ping can go out; before the conversation has answered in
670      // this run the next one is the visible ping.
671      if (pingError) parts.push(`ping failed: ${pingError}`)
672      else if (!ttl) parts.push(WAITS_TEXT)
673      else if (isBlocked) parts.push(NEARLY_FULL_TEXT)
674      else if (isWarm) parts.push(`next ping ${clockTime(last + ttlMs - PING_LEAD_MS)}${sawResponse && !waits ? '' : SHOWS_IN_CHAT}`)
675      if (lastPing) parts.push(`${pings.length} ping${pings.length === 1 ? '' : 's'} ${lastPing.isHit ? '✓' : '✗'}`)
676    }
677
678    if (contextTokens > 0 && last > 0) {
679      if (!isWarm) parts.push(`your next message re-writes ${tokens(contextTokens)}`, ...(isGreen ? ['keep-warm restarts after it'] : []))
680      else if (!isGreen && remaining < EXPIRING_MS) parts.push(`after ${clockTime(expiresAt)} your next message re-writes ${tokens(contextTokens)}`)
681      else parts.push(`${tokens(contextTokens)} cached`)
682    }
683
684    // No toggle where keep-warm can never ping (an API key, or a short cache set on purpose).
685    const canToggle = isOn || !ttl || canKeepWarm || ttl.reason === 'overage' || ttl.reason === 'observed'
686
687    // Each detail is its own element and the spacing comes from the layout (column gaps), which
688    // no surface squeezes the way it squeezes runs of spaces inside one text. In a narrow window
689    // nothing wraps: the time and the button keep their size and the details are cut at the edge.
690    return (
691      <Box alignItems="center" flexWrap="nowrap">
692        <Box flexGrow={1} flexShrink={1} minWidth={0} overflow="hidden" flexWrap="nowrap" alignItems="center" columnGap={DETAIL_GAP}>
693          <Box flexShrink={0}>
694            {isGreen ? <Text color="green" bold>{circle}</Text> : <Text dimColor>{circle}</Text>}
695          </Box>
696          <Box flexShrink={0}>
697            {isGreen ? <Text color="green">{BAR_FULL.repeat(filled)}</Text> : <Text dimColor>{BAR_FULL.repeat(filled)}</Text>}
698            <Text dimColor>{BAR_EMPTY.repeat(BAR_CELLS - filled)}</Text>
699          </Box>
700          <Box flexShrink={0}>
701            <Text bold>{main}</Text>
702            <Text dimColor>{expiresText}</Text>
703          </Box>
704          {parts.flatMap(p => [
705            <Box flexShrink={0}><Text dimColor>·</Text></Box>,
706            <Box flexShrink={0}><Text dimColor>{p}</Text></Box>,
707          ])}
708        </Box>
709        {canToggle && (
710          <Box flexShrink={0}>
711            <Button
712              key="toggle"
713              label={isOn ? 'Stop' : 'Keep warm'}
714              onPress={() => setKeepWarm($, !isOn, 0)}
715            />
716          </Box>
717        )}
718      </Box>
719    )
720  })
721}
722
types/index.d.ts 55 lines
1/**
2 * What the mod remembers per session id in `$.store`, so a relaunch keeps it. `ttl` is the cache
3 * length in effect when it was saved (absent before 1.0.3); `savedAt` when it was saved, which
4 * ages the entry for the weekly clean-up (absent before 1.0.4).
5 */
6export type KeepWarmSaved = { lastRequestAt: number; keepWarm: boolean; stopAt: number; ttl?: KeepWarmTtl; savedAt?: number }
7
8/** One keep-warm ping: when it started, whether the cache still held the conversation, how much it read. */
9export type KeepWarmPing = { at: number; isHit: boolean; read: number }
10
11/**
12 * How long the main conversation's cache lives, and why: from a setting, from the plan, or
13 * corrected by what the cache actually did. Null until it can be told. The last one the plan
14 * showed is also kept in `$.store` under `plan-ttl`, shared by every session.
15 */
16export type KeepWarmTtl = { ms: number; reason: 'subscription' | 'api' | 'overage' | 'setting' | 'observed' } | null
17
18declare module 'claude-code' {
19  interface PluginState {
20    'cache-keep-warm': {
21      /** When the main conversation's last model request started (ms since epoch); 0 = unknown. */
22      lastRequestAt: number
23      /** Whether this session keeps its cache warm. */
24      keepWarm: boolean
25      /** When keep-warm switches itself off (ms since epoch); 0 = after the ping limit. */
26      stopAt: number
27      /** The clock as the band last showed it. */
28      now: number
29      /** Pings since keep-warm was last switched on, oldest first. */
30      pings: KeepWarmPing[]
31      /** The cache lifetime as last worked out. */
32      ttl: KeepWarmTtl
33      /** Tokens the last request sent: what a cold cache would make the next message re-write. */
34      contextTokens: number
35      /** True while Claude waits on an answer to its question. */
36      isWaiting: boolean
37      /**
38       * True once this process has seen a main-thread response: the ping can then be invisible (a
39       * fork), and an empty plan list then means an API key. Reset by /clear.
40       */
41      sawResponse: boolean
42      /** The idle gap after which the cache was found cold although it should have lasted; 0 = none seen. */
43      shortAfterMs: number
44      /** True when the invisible ping found nothing to repeat, so the visible one is used, until the next request. */
45      waitsForMessage: boolean
46      /** True once this process has restored the session's saved values from the store. */
47      restored: boolean
48      /** Why the last ping did not reach the model (Claude Code's words or the HTTP status); '' = none. */
49      pingError: string
50      /** True while the conversation fills too much of the context window for a visible ping. */
51      nearlyFull: boolean
52    }
53  }
54}
55