Shows how long your session's prompt cache stays warm, and can keep it warm in the background while you are away

Stop paying to re-cache your whole Claude Code conversation after a break.
A small mod for Claude Code that shows how long your session's prompt cache stays warm, and keeps it warm while you are away.
<img src="docs/images/screenshot.png" alt="Cache Keep-Warm in the Claude Code desktop app: a green countdown line above the prompt reading '57 min left (21:38) · next ping 21:33 · 1 ping ✓ · 460k cached', and the session list with a green dot in front of the sessions being kept warm" width="100%">
Claude Code keeps your conversation in a prompt cache, so every message only pays full price for what is new. But the cache expires when you stop typing for a while: after one hour on a Claude subscription, after five minutes on an API key.
Come back after that, and your next message has to write the whole conversation into the cache again. For a long session that is the most expensive message of the day.
Cache Keep-Warm fixes that:
<img src="docs/images/sessions.svg" alt="The desktop app's session list, with a green dot in front of the two sessions that keep-warm is on in" width="420">
Example: a long conversation of 500k tokens on Claude Opus 5.5, at current API list prices.
| Cost | |
|---|---|
| Your next message after the cache expired (re-writes everything) | ≈ $4.00 |
| One keep-warm ping (reads it from the cache) | ≈ $0.10 |
| Keeping it warm through an 8-hour night (8 pings) | ≈ $0.80 |
On a subscription you do not pay per token, but the same proportions apply to your usage limits.
Copy this into any Claude Code session:
Install the Cache Keep-Warm mod from https://github.com/andreichiritescu/claude-code-cache-keep-warm for me:
1. Run: claude plugin marketplace add https://github.com/andreichiritescu/claude-code-cache-keep-warm.git
2. Run: claude plugin install cache-keep-warm@cache-keep-warm
3. If a step fails because my Claude Code is too old, tell me to update it (the mod needs Claude Code 2.1.287 or newer, or the desktop app with 2.1.286 or newer) and stop.
4. When it is installed, tell me to run /reload-plugins so this session loads it, and that I can turn keep-warm on with the Keep warm button on the new line above the prompt, or with /keepwarm on.
Type this at the Claude Code prompt (Claude Code 2.1.275 or newer):
/plugin install cache-keep-warm --marketplace https://github.com/andreichiritescu/claude-code-cache-keep-warm.git
claude plugin marketplace add https://github.com/andreichiritescu/claude-code-cache-keep-warm.git
claude plugin install cache-keep-warm@cache-keep-warm
Then start a new session, or run /reload-plugins in an open one. The line appears above the prompt, and the countdown starts with your first message. If Claude Code says an option is not set yet, that is the Mark the session title setting: it is on by default.
Needs Claude Code 2.1.287 or newer in a terminal, or the Claude desktop app with Claude Code 2.1.286 or newer. Type
/statusto check your version.
Click Keep warm on the line, or type a command:
| Command | What it does |
|---|---|
/keepwarm on | Keep this session warm |
/keepwarm 3h | Keep it warm for 3 hours |
/keepwarm until 18:00 | Keep it warm until 18:00 |
/keepwarm off | Stop (or click Stop) |
/keepwarm | Show the state and the last pings |
Keep-warm is off by default and you switch it per session: only the sessions you want are kept warm. It stops by itself after about 18 hours (20 pings in a row) unless you gave it a time, because past that point one re-write is cheaper than more pings.
| You see | It means |
|---|---|
| ◕ and ▰▰▰▰▰▰▰▱▱▱ | How much of the cache's life is left. Green while keep-warm is on, grey when it is off, and ○ once the cache has expired. |
| nothing cached yet | A new conversation, or one you just cleared or compacted: nothing of it is cached until your next message, and the countdown starts with that. |
| 40 min left (14:12) | Time left, and the clock time the cache expires if nothing happens. |
| next ping 14:07 | When keep-warm will read the cache next. (shows in chat) means this ping is a short visible message, because the session hasn't answered since the app restarted. |
| 3 pings ✓ | Pings so far, and whether the last one found the cache warm (✓) or cold (✗). |
| conversation nearly full: keep-warm waits for you | After a restart the next ping has to be the visible one, and Claude Code re-sends your project's instruction files (CLAUDE.md, the files it imports, MEMORY.md) with it. Here those would not fit next to the conversation: the ping would fail or make Claude Code compact the conversation. So keep-warm waits for your own next message. |
| ping failed: … | The last ping never reached the model, with the reason Claude Code gave. "Prompt is too long" means the conversation is over the model's limit: run /compact in that session. The cache then expires normally. |
| keep-warm starts after your next message | Keep-warm is on but doesn't know your cache length yet. That only happens before the first answer after you install the mod. |
| 773k cached | The size of your conversation, which is what your next message would re-write if the cache went cold. |
flowchart LR
A["You stop typing"] --> B{"55 minutes later:<br/>still away?"}
B -- yes --> C["Keep-warm reads the<br/>conversation from the cache"]
C --> D["The cache is good<br/>for another hour"]
D --> B
B -- "you are back" --> E["Your next message<br/>reads from the cache:<br/>cheap"]
promptCacheTtl, CLAUDE_CODE_PROMPT_CACHE_TTL, FORCE_PROMPT_CACHING_5M or ENABLE_PROMPT_CACHING_1H.Claude Code reports your plan only with an answer, so after a restart the mod uses the length your plan showed last time until the session answers again.
Keep-warm only pings a cache that lasts 30 minutes or more: on the five-minute cache, pinging every few minutes would cost more than it saves.
| Setting | Default | What it does |
|---|---|---|
| Mark the session title | On | While keep-warm is on, put 🟢 in front of the session's title in the desktop app's session list |
Change it in /config, or with /plugin configure cache-keep-warm@cache-keep-warm.
In the desktop app, renaming a title you typed yourself asks for your approval each time; titles the app generated change without asking.
I use an API key. Does keep-warm work for me? Your cache lasts five minutes by default, so keep-warm does not ping. If you take longer breaks, set promptCacheTtl to "1h" in your Claude Code settings: the cache then lasts an hour and keep-warm works. One-hour cache writes cost more than five-minute ones, so this pays off only if you do take breaks.
Will it ping forever if I forget it? No. It stops by itself after 20 pings in a row (about 18 hours), or at the time you gave it with /keepwarm 3h or /keepwarm until 18:00.
Does it work in VS Code? The keep-warm pings work there, but the line is only drawn in the terminal and in the desktop app's Code tab.
Where does the line not appear? In VS Code's chat panel, in claude -p runs and in cloud sessions.
How do I remove it? claude plugin uninstall cache-keep-warm@cache-keep-warm
A mod runs inside Claude Code with your permissions, so here is everything this one does. It has no server, makes no network calls of its own, and runs no shell commands. This section is also its privacy policy: the mod collects nothing about you, and nothing reaches its author.
Only the keep-warm ping, and only while keep-warm is on: about 5 minutes before the cache would expire, so about once every 55 minutes.
The mod asks Claude Code to send this session's last request again, with one line added: Keep-warm ping. Reply with just "ok". (the $.model.fork call in hooks/register.tsx). It goes to Anthropic, through Claude Code's own connection, on your own account, like any message you send. That request reads your conversation from the cache, which is what keeps it warm. The answer is thrown away, and nothing is added to your conversation.
After an app restart, Claude Code can repeat a session's request this way only once the session has answered again. Until then the ping is a short message the mod sends into the chat instead (the $.prompt.submit call): Keep-warm ping from the Cache Keep-Warm mod. Reply with exactly "ok" and nothing else. Use no tools. Claude answers "ok", and from then on the pings are invisible again. Same destination, same account, and the same rule: only in the last minutes before the cache would expire, never once it is cold.
The mod itself never reads what your conversation says: from each request it only takes the usage figures (how many tokens were read from the cache or written to it).
Only on the current session. Both are the Claude desktop app's own session tools, and nothing they do leaves your computer:
| Tool | What it does | When |
|---|---|---|
mcp__ccd_session_mgmt__get_session | Reads the session's title | When the session starts, when you switch keep-warm on or off, and every 5 minutes |
mcp__ccd_session_mgmt__set_session_title | Puts 🟢 in front of the title, or takes it off | Only when the mark has to change |
With the setting Mark the session title off, the mod never adds the mark; it still reads the title at those moments, to take off a mark left from before. In a terminal these tools do not exist: the call fails quietly and nothing changes.
promptCacheTtl only, and three environment variables: FORCE_PROMPT_CACHING_5M, CLAUDE_CODE_PROMPT_CACHE_TTL and ENABLE_PROMPT_CACHING_1H, to know how long your cache lasts. Claude Code hands a mod the settings as one object: this mod looks at promptCacheTtl and nothing else, and keeps or sends no other setting, credential or API key.It checks these every 15 seconds, on your computer, and sends none of them anywhere. It does not read your files.
Per session, in Claude Code's plugin storage on your computer: the time of the last request, whether keep-warm is on, and the stop time you gave it. A session untouched for a week is forgotten. Shared by your sessions, it also keeps the last cache length your plan showed (one hour or five minutes), so a session you reopen knows it before its first answer.
| Event | What the mod does with it |
|---|---|
| Session start | Adds the /keepwarm command and starts its 15-second local check |
A session is reopened or cleared (Claude Code's SessionStart hook) | Reads how long ago the last answer came and the conversation's size, so the countdown carries on; after /clear, starts the countdown over. Changes nothing. |
A compaction (/compact or automatic) | Checks only whose it is and whether it went through: when your conversation itself was compacted, starts the countdown over. Ignores a subagent's compaction of its own work and a summary prepared ahead of time. Never reads or changes the summary or your messages. |
| Each request of the main conversation | Notes the time and how much was read from the cache. Changes nothing. Subagents' requests are ignored. |
| A turn ends | Notes that an answer arrived (which helps tell a subscription from an API key) |
Claude's AskUserQuestion tool | Shows "waiting for your answer" while the question is open. Never changes the question or your answer. |
The /keepwarm command | Answers its own command, and no other |
| The line above the prompt | Draws the countdown line |
To check all of this before you install it, clone the repository and run:
claude plugin validate ./claude-code-cache-keep-warm/.claude-plugin/plugin.json
The hooks:, calls: and env reads: lines list every event the mod handles, everything it asks Claude Code to do, and every environment variable it reads.
MIT License · A community mod, not made by Anthropic · Issues and ideas welcome
hooks/register.tsx 722 lines1import type { EngineInterface, ModelUsage, Register, SessionUsage } from 'claude-code'
2
3import type { KeepWarmPing, KeepWarmSaved, KeepWarmTtl } from '../types'
4
5const HOUR_MS = 60 * 60_000
6const FIVE_MIN_MS = 5 * 60_000
7// The cache lengths Claude Code offers today; a value seen elsewhere is still accepted.
8const KNOWN_TTLS_MS = [FIVE_MIN_MS, HOUR_MS]
9// Keep-warm only for a cache this long or longer: on a short cache, pinging every few
10// minutes costs more than the single re-write it saves after any real break.
11const KEEP_WARM_MIN_TTL_MS = 30 * 60_000
12// A ping goes out this long before the cache would expire (55 min into a 1-hour cache), and
13// never in its last minute: a cold cache is never pinged. A later ping means fewer of them
14// (one per ~55 min instead of one per 45), and the lead still leaves four minutes of retries
15// if the API is busy. The cache's life runs from when a request reaches the server, a moment
16// after the time recorded here, so the real deadline is slightly later than computed.
17const PING_LEAD_MS = 5 * 60_000
18const PING_MARGIN_MS = 60_000
19// After a failed ping, wait this long before trying again inside the window.
20const PING_RETRY_MS = 30_000
21// Without an explicit stop time, keep-warm stops itself after this many pings in a row
22// (~15 h): past ~30 h idle, re-writing the cache once is cheaper than pinging.
23const MAX_PINGS = 20
24// One check every 15 s: local reads only (clock, session figures, settings), no model call.
25// The line is redrawn only when its text changes, at most once a minute.
26const TICK_MS = 15_000
27const PING_PENDING_MS = 5 * 60_000
28// A request reading at least this share of the conversation from cache found it still warm.
29const HIT_SHARE = 0.5
30// A re-write after a shorter gap is a change to the prompt (compaction, a model switch),
31// not an expiry, so it says nothing about the cache's length.
32const EVIDENCE_GAP_MS = FIVE_MIN_MS + 30_000
33// A reopened session reports how long ago its last answer came. Its cache's life began when
34// that last request was sent, a little earlier: counting from this much before the answer keeps
35// the ping early rather than late (a late ping would re-write a cache that already went cold).
36const RESUME_MARGIN_MS = 3 * 60_000
37// With keep-warm off, the band names what is at stake once this little time is left.
38const EXPIRING_MS = 10 * 60_000
39// Blocks that empty over the cache's life: filled blocks green while keep-warm is on (grey
40// otherwise), outlined blocks for the time already gone.
41const BAR_CELLS = 10
42const BAR_FULL = '▰'
43const BAR_EMPTY = '▱'
44// Space between the details, in the surface's own units (columns on the terminal).
45const DETAIL_GAP = 2
46// The circle in front empties with the cache's life, a quarter at a time; hollow once cold.
47// The full one is the "medium black circle" in text style: the plain ● is drawn smaller than
48// the quarter circles in the desktop app's font (picked by the user from a test row).
49const CIRCLE_BY_QUARTER = ['◔', '◑', '◕', '⚫︎']
50const CIRCLE_COLD = '○'
51// While keep-warm is on, the session's title in the app's list starts with this mark. The app
52// asks before changing a title you typed yourself; titles it generated change without asking.
53const TITLE_MARK = '🟢 '
54// The userConfig key in plugin.json that switches the title mark off.
55const TITLE_MARK_OPTION = 'title_mark'
56const TITLE_SYNC_MS = 5 * 60_000
57const GET_SESSION_TOOL = 'mcp__ccd_session_mgmt__get_session'
58const SET_TITLE_TOOL = 'mcp__ccd_session_mgmt__set_session_title'
59const TITLE_FIELD = /"title"\s*:\s*"((?:[^"\\]|\\.)*)"/
60const HISTORY_SHOWN = 5
61const PING_TEXT = 'Keep-warm ping. Reply with just "ok".'
62const COMMAND = 'keepwarm'
63const STORE_PREFIX = 'session:'
64const STORE_MAX_AGE_MS = 7 * 24 * HOUR_MS
65// Shared by every session: the cache length the plan showed last. Claude Code reports the plan
66// only with a response, so a session reopened after a relaunch would otherwise not know its
67// cache length (and could not keep it warm) until its first answer.
68const PLAN_TTL_KEY = 'plan-ttl'
69// Rate-limit windows that only a Claude subscription reports; an API key or a cloud provider reports none.
70const PLAN_WINDOWS = ['five_hour', 'seven_day']
71const PLAN_LIMIT_PERCENT = 100
72const TTL_VALUE = /^(\d+)\s*(s|m|h)$/
73const DURATION = /^(\d+(?:\.\d+)?)\s*(h|m|min)$/
74const UNTIL = /^until\s+(\d{1,2}):(\d{2})$/
75const WAITS_TEXT = 'keep-warm starts after your next message'
76// After an app restart Claude Code can repeat (fork) a conversation only once it has answered in
77// this run. Until then the ping is a real message instead: it shows in the chat, Claude answers
78// "ok", and from then on the pings are invisible again. Same window, so never on a cold cache.
79const VISIBLE_PING_TEXT = 'Keep-warm ping from the Cache Keep-Warm mod. Reply with exactly "ok" and nothing else. Use no tools.'
80const SHOWS_IN_CHAT = ' (shows in chat)'
81// The visible ping is a real turn, and after a restart Claude Code re-sends the project's
82// instruction files (CLAUDE.md, the files it imports, MEMORY.md) with it: 130k tokens in a big
83// project. If the conversation, those files and this much room for the ping and its reply do not
84// fit in the window's free space (before the compaction reserve), the turn would fail ("Prompt is
85// too long") or be compacted, which a ping must never cause: keep-warm then waits for the
86// person's own message instead.
87const VISIBLE_PING_ROOM = 30_000
88// Only where Claude Code gives no breakdown of the window: wait above this share of it.
89const VISIBLE_PING_MAX_FILL = 0.75
90const NEARLY_FULL_TEXT = 'conversation nearly full: keep-warm waits for you'
91
92const LAST = { plugin: 'cache-keep-warm', key: 'lastRequestAt' } as const
93const KEEP_WARM = { plugin: 'cache-keep-warm', key: 'keepWarm' } as const
94const STOP_AT = { plugin: 'cache-keep-warm', key: 'stopAt' } as const
95const NOW = { plugin: 'cache-keep-warm', key: 'now' } as const
96const PINGS = { plugin: 'cache-keep-warm', key: 'pings' } as const
97const TTL = { plugin: 'cache-keep-warm', key: 'ttl' } as const
98const CONTEXT = { plugin: 'cache-keep-warm', key: 'contextTokens' } as const
99const WAITING = { plugin: 'cache-keep-warm', key: 'isWaiting' } as const
100const SAW_RESPONSE = { plugin: 'cache-keep-warm', key: 'sawResponse' } as const
101const SHORT_AFTER = { plugin: 'cache-keep-warm', key: 'shortAfterMs' } as const
102const WAITS = { plugin: 'cache-keep-warm', key: 'waitsForMessage' } as const
103const RESTORED = { plugin: 'cache-keep-warm', key: 'restored' } as const
104const PING_ERROR = { plugin: 'cache-keep-warm', key: 'pingError' } as const
105const NEARLY_FULL = { plugin: 'cache-keep-warm', key: 'nearlyFull' } as const
106
107// The module's own variables start over on a reload; what must survive one is in $.state.
108// `titleMark` is the user's "Mark the session title" setting; changing it in /config reloads
109// the module, so it is read once in register(). `planTtl` mirrors the shared PLAN_TTL_KEY;
110// `savedTtl` is the length this session's own cache had when the store last saw it.
111// `visiblePingAt`: when a visible ping was submitted, until its request is seen.
112// `pingTurnId`: the visible ping's turn, which is no sign of use when it ends.
113// `roomChecked`: whether this process has worked out if the visible ping fits.
114// `isFirstRequest`: the first request after a start says nothing about the cache's length.
115const live: {
116 storeKey: string
117 pingPendingUntil: number
118 shownKey: string
119 titleMark: boolean
120 planTtl: KeepWarmTtl | undefined
121 savedTtl: KeepWarmTtl | undefined
122 isFirstRequest: boolean
123 visiblePingAt: number
124 pingTurnId: string
125 roomChecked: boolean
126} = {
127 storeKey: '',
128 pingPendingUntil: 0,
129 shownKey: '',
130 titleMark: true,
131 planTtl: undefined,
132 savedTtl: undefined,
133 isFirstRequest: true,
134 visiblePingAt: 0,
135 pingTurnId: '',
136 roomChecked: false,
137}
138
139function minutes(ms: number) {
140 return Math.max(0, Math.floor(ms / 60_000))
141}
142
143function clockTime(ms: number) {
144 const d = new Date(ms)
145 return `${String(d.getHours()).padStart(2, '0')}:${String(d.getMinutes()).padStart(2, '0')}`
146}
147
148function tokens(n: number) {
149 return n >= 1_000_000 ? `${(n / 1_000_000).toFixed(1)}M` : `${Math.round(n / 1000)}k`
150}
151
152function timeLeft(ms: number) {
153 return ms < 60_000 ? 'under a minute left' : `${minutes(ms)} min left`
154}
155
156function cacheLength(ms: number) {
157 return ms >= HOUR_MS && ms % HOUR_MS === 0 ? `${ms / HOUR_MS}-hour` : `${minutes(ms)}-minute`
158}
159
160// "5m", "1h", "30m", "90s": any length, so a cache length Claude Code adds later still reads.
161function parseTtl(value: unknown) {
162 if (typeof value !== 'string') return 0
163 const m = TTL_VALUE.exec(value.trim())
164 if (!m) return 0
165 return Number(m[1]) * (m[2] === 'h' ? HOUR_MS : m[2] === 'm' ? 60_000 : 1000)
166}
167
168// Keeps what the plan showed for the next session that opens before its first answer; written
169// only when it changes, though this runs on every check.
170async function rememberPlan($: EngineInterface, ttl: NonNullable<KeepWarmTtl>): Promise<KeepWarmTtl> {
171 if (live.planTtl?.ms !== ttl.ms || live.planTtl?.reason !== ttl.reason) {
172 live.planTtl = ttl
173 await $.store.set(PLAN_TTL_KEY, ttl)
174 }
175 return ttl
176}
177
178// The same precedence Claude Code itself applies, then the plan: a subscription within its
179// limits gets the hour, one over its limit (paying usage credits) and an API key get 5 minutes.
180async function detectTtl($: EngineInterface, usage: SessionUsage): Promise<KeepWarmTtl> {
181 if ((await $.env.get('FORCE_PROMPT_CACHING_5M')) === '1') return { ms: FIVE_MIN_MS, reason: 'setting' }
182 const fromEnv = parseTtl(await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL'))
183 if (fromEnv) return { ms: fromEnv, reason: 'setting' }
184 const settings = await $.settings.read()
185 const fromSettings = parseTtl(settings['promptCacheTtl'])
186 if (fromSettings) return { ms: fromSettings, reason: 'setting' }
187 if ((await $.env.get('ENABLE_PROMPT_CACHING_1H')) === '1') return { ms: HOUR_MS, reason: 'setting' }
188
189 const plan = usage.rateLimits.filter(w => PLAN_WINDOWS.includes(w.kind))
190 if (plan.length === 0) {
191 // The windows arrive with the first response, so before one an empty list proves nothing:
192 // until then, go by the length this session's cache had when last saved (what its last
193 // request really got), else by what the plan showed last time in any session.
194 const { value: sawResponse = false } = await $.state.get(SAW_RESPONSE)
195 if (!sawResponse) return live.savedTtl ?? live.planTtl ?? null
196 return rememberPlan($, { ms: FIVE_MIN_MS, reason: 'api' })
197 }
198 return rememberPlan(
199 $,
200 plan.some(w => w.percentUsed >= PLAN_LIMIT_PERCENT) ? { ms: FIVE_MIN_MS, reason: 'overage' } : { ms: HOUR_MS, reason: 'subscription' },
201 )
202}
203
204// What the cache actually did beats what the settings say: if it went cold after a gap it
205// should have survived, use the longest known length shorter than that gap.
206function correctTtl(base: KeepWarmTtl, shortAfterMs: number): KeepWarmTtl {
207 if (!base || shortAfterMs <= 0 || shortAfterMs >= base.ms - PING_MARGIN_MS) return base
208 const shorter = KNOWN_TTLS_MS.filter(ms => ms < shortAfterMs)
209 return { ms: shorter.length ? Math.max(...shorter) : FIVE_MIN_MS, reason: 'observed' }
210}
211
212// Each request after an idle gap shows whether the cache outlived that gap: a read clears an
213// earlier "went cold too soon"; a re-write after a gap the cache should have survived records one.
214async function observe($: EngineInterface, gapMs: number, usage: ModelUsage) {
215 const total = usage.cache_read_input_tokens + usage.cache_creation_input_tokens
216 if (total === 0 || gapMs < EVIDENCE_GAP_MS) return
217 if (usage.cache_read_input_tokens >= HIT_SHARE * total) {
218 await $.state.set(SHORT_AFTER, 0)
219 return
220 }
221 const { value: ttl = null } = await $.state.get(TTL)
222 const { value: shortAfter = 0 } = await $.state.get(SHORT_AFTER)
223 const expected = ttl?.reason === 'observed' ? HOUR_MS : (ttl?.ms ?? HOUR_MS)
224 if (gapMs < expected - PING_MARGIN_MS) await $.state.set(SHORT_AFTER, shortAfter > 0 ? Math.min(shortAfter, gapMs) : gapMs)
225}
226
227async function save($: EngineInterface) {
228 if (!live.storeKey) return
229 const { value: lastRequestAt = 0 } = await $.state.get(LAST)
230 const { value: keepWarm = false } = await $.state.get(KEEP_WARM)
231 const { value: stopAt = 0 } = await $.state.get(STOP_AT)
232 const { value: ttl = null } = await $.state.get(TTL)
233 const saved: KeepWarmSaved = { lastRequestAt, keepWarm, stopAt, ttl, savedAt: await $.clock.now() }
234 await $.store.set(live.storeKey, saved)
235}
236
237async function setKeepWarm($: EngineInterface, isOn: boolean, stopAt: number) {
238 await $.state.set(KEEP_WARM, isOn)
239 await $.state.set(STOP_AT, isOn ? stopAt : 0)
240 await $.state.set(PINGS, [])
241 await $.state.set(PING_ERROR, '')
242 await save($)
243 await syncTitle($)
244}
245
246// The conversation starts over (cleared, or replaced by its summary): nothing of it is cached
247// until the next request writes it.
248async function startOver($: EngineInterface) {
249 await $.state.set(LAST, 0)
250 await $.state.set(PINGS, [])
251 await $.state.set(SAW_RESPONSE, false)
252 await $.state.set(CONTEXT, 0)
253 await $.state.set(PING_ERROR, '')
254 await $.state.set(NOW, await $.clock.now())
255 live.roomChecked = false
256 const { value: restored = false } = await $.state.get(RESTORED)
257 if (restored) await save($)
258}
259
260// Puts the mark on this session's title in the app's list while keep-warm is on, and takes it
261// off otherwise (or always, with the setting off). Runs on every switch, at start (which also
262// clears a mark a crash left behind) and every few minutes (a title the app regenerated loses
263// the mark). Outside the desktop app there are no session tools and nothing is marked.
264async function syncTitle($: EngineInterface) {
265 const { value: isOn = false } = await $.state.get(KEEP_WARM)
266 const { value: ttl = null } = await $.state.get(TTL)
267 const wantsMark = live.titleMark && isOn && !(ttl && ttl.ms < KEEP_WARM_MIN_TTL_MS)
268 try {
269 const found = await $.tool.call({ tool: GET_SESSION_TOOL, session_id: 'self' })
270 if ('deny' in found && found.deny) return
271 const match = TITLE_FIELD.exec(found.text ?? '')
272 if (!match) return
273 const title: string = JSON.parse(`"${match[1]}"`)
274 const hasMark = title.startsWith(TITLE_MARK)
275 if (wantsMark === hasMark) return
276 const base = hasMark ? title.slice(TITLE_MARK.length) : title
277 await $.tool.call({ tool: SET_TITLE_TOOL, session_id: 'self', title: wantsMark ? TITLE_MARK + base : base })
278 } catch {
279 // No session tools here (a terminal session), or the app declined the rename: leave the title.
280 }
281}
282
283// The ping re-sends the main thread's last request outside the conversation with one short
284// question after it: it reads the conversation from cache (restarting its life), adds no
285// message, and runs even while the session waits on a question or a permission prompt.
286async function ping($: EngineInterface, startedAt: number, lastRequestAt: number) {
287 const r = await $.model.fork({ prompt: PING_TEXT })
288 if (!r.isAnswered && r.reason === 'nothing-to-fork') {
289 // Claude Code has no request of this conversation to repeat yet: send the visible one.
290 await $.state.set(WAITS, true)
291 const { value: contextTokens = 0 } = await $.state.get(CONTEXT)
292 const { context } = await $.session.usage()
293 if (!(await checkRoom($, context.window, contextTokens))) await visiblePing($, startedAt)
294 return
295 }
296 if (!r.isAnswered && r.reason === 'api-error') {
297 // A busy or failing API may answer on a retry; a request the API refuses (the conversation
298 // over the model's limit, for one) gets the same answer every time, so stop and say so.
299 const status = r.status ?? 0
300 if (status >= 400 && status < 500 && status !== 429) {
301 await $.state.set(PING_ERROR, `HTTP ${status}`)
302 $.ui.toast(`Keep-warm ping failed (HTTP ${status}); the cache will expire normally.`)
303 return
304 }
305 $.ui.toast(`Keep-warm ping failed (HTTP ${status || '?'}); retrying in ${PING_RETRY_MS / 1000} s.`)
306 live.pingPendingUntil = startedAt + PING_RETRY_MS
307 return
308 }
309 const { cache_read_input_tokens: read, cache_creation_input_tokens: written } = r.usage
310 const isHit = read > 0 && read >= HIT_SHARE * (read + written)
311 const done: KeepWarmPing = { at: startedAt, isHit, read }
312 const { value: pings = [] } = await $.state.get(PINGS)
313 await $.state.set(PINGS, [...pings, done].slice(-MAX_PINGS))
314 // A ping that found the cache cold proves the cache is shorter than assumed, which pauses
315 // keep-warm (another cold ping would only pay for a full re-write again).
316 await observe($, startedAt - lastRequestAt, r.usage)
317 await $.state.set(LAST, startedAt)
318 await save($)
319}
320
321// Whether the visible ping would not fit: the conversation, the instruction files Claude Code
322// re-sends after a restart, and the ping with its reply must all fit in the window's free space.
323// Claude Code estimates the breakdown locally (no request), from the same figures as /context.
324async function visiblePingWontFit($: EngineInterface, window: number, contextTokens: number) {
325 try {
326 const { breakdown } = (await $.session.usage({ breakdown: 'summary' })).context
327 if (breakdown) {
328 const free = breakdown.categories.filter(c => c.kind === 'free').reduce((n, c) => n + c.tokens, 0)
329 const instructions = breakdown.memoryFiles.reduce((n, f) => n + f.tokens, 0)
330 return free < instructions + VISIBLE_PING_ROOM
331 }
332 } catch {
333 // No breakdown here: go by the share of the window below.
334 }
335 return window > 0 && contextTokens >= VISIBLE_PING_MAX_FILL * window
336}
337
338// Works out once per start, and only while the next ping would be the visible one, whether it
339// fits; the line shows the answer before the ping is due.
340async function checkRoom($: EngineInterface, window: number, contextTokens: number) {
341 live.roomChecked = true
342 const wontFit = await visiblePingWontFit($, window, contextTokens)
343 const { value: was = false } = await $.state.get(NEARLY_FULL)
344 if (wontFit !== was) await $.state.set(NEARLY_FULL, wontFit)
345 return wontFit
346}
347
348// The fallback ping: a real prompt, so it shows in the chat with Claude's "ok". Its request reads
349// the conversation from cache like the invisible ping, and the turn.step hook records it.
350async function visiblePing($: EngineInterface, startedAt: number) {
351 live.visiblePingAt = startedAt
352 const r = await $.prompt.submit({ text: VISIBLE_PING_TEXT })
353 if (typeof r.drop === 'string') {
354 live.visiblePingAt = 0
355 $.ui.toast(`Keep-warm ping not sent: ${r.drop}`)
356 }
357}
358
359async function tick($: EngineInterface) {
360 const t = await $.clock.now()
361 const usage = await $.session.usage()
362 const { value: shortAfter = 0 } = await $.state.get(SHORT_AFTER)
363 const ttl = correctTtl(await detectTtl($, usage), shortAfter)
364 const { value: last = 0 } = await $.state.get(LAST)
365
366 // Redraw only when what the band shows would change: each minute, and whenever the cache
367 // length or the context size moves. A reopened session reports its size before its first
368 // answer, while the engine has none yet: keep that one until the engine has its own.
369 const { value: knownContext = 0 } = await $.state.get(CONTEXT)
370 const contextTokens = usage.context.tokens ?? knownContext
371 const shownKey = `${Math.floor(t / 60_000)}|${ttl?.ms ?? 0}|${ttl?.reason ?? ''}|${contextTokens}|${last}|${shortAfter}`
372 if (shownKey !== live.shownKey) {
373 live.shownKey = shownKey
374 await $.state.set(TTL, ttl)
375 await $.state.set(CONTEXT, contextTokens)
376 await $.state.set(NOW, t)
377 }
378
379 const { value: isOn = false } = await $.state.get(KEEP_WARM)
380 if (!isOn) return
381 // Invisible once the conversation has answered in this run; before that, the visible ping,
382 // whose fit is worked out once so the line can say in advance when it has to wait.
383 const { value: sawResponse = false } = await $.state.get(SAW_RESPONSE)
384 const { value: waits = false } = await $.state.get(WAITS)
385 const needsVisible = !sawResponse || waits
386 if (needsVisible && !live.roomChecked) await checkRoom($, usage.context.window, contextTokens)
387 const { value: stopAt = 0 } = await $.state.get(STOP_AT)
388 if (stopAt > 0 && t >= stopAt) {
389 await setKeepWarm($, false, 0)
390 $.ui.toast(`Keep-warm stopped at ${clockTime(stopAt)} as asked; the cache will expire normally.`)
391 return
392 }
393 const idle = t - last
394 const isPingWindow =
395 ttl !== null &&
396 ttl.ms >= KEEP_WARM_MIN_TTL_MS &&
397 last > 0 &&
398 idle >= ttl.ms - PING_LEAD_MS &&
399 idle < ttl.ms - PING_MARGIN_MS
400 if (!isPingWindow || t < live.pingPendingUntil) return
401
402 const { value: pings = [] } = await $.state.get(PINGS)
403 if (stopAt === 0 && pings.length >= MAX_PINGS) {
404 await setKeepWarm($, false, 0)
405 $.ui.toast(`Keep-warm stopped itself after ${MAX_PINGS} pings in a row; the cache will expire.`)
406 return
407 }
408 live.pingPendingUntil = t + PING_PENDING_MS
409 const { value: nearlyFull = false } = await $.state.get(NEARLY_FULL)
410 if (!needsVisible) await ping($, t, last)
411 else if (!nearlyFull) await visiblePing($, t)
412}
413
414// "3h", "90m", "until 18:00" -> when keep-warm stops; 0 = no stop time; -1 = not understood.
415function parseStop(text: string, now: number) {
416 if (text === '') return 0
417 const duration = DURATION.exec(text)
418 if (duration) return now + Number(duration[1]) * (duration[2] === 'h' ? HOUR_MS : 60_000)
419 const until = UNTIL.exec(text)
420 if (!until) return -1
421 const at = new Date(now)
422 at.setHours(Number(until[1]), Number(until[2]), 0, 0)
423 return at.getTime() > now ? at.getTime() : at.getTime() + 24 * HOUR_MS
424}
425
426export const register: Register = (on, options) => {
427 live.titleMark = options[TITLE_MARK_OPTION] !== false
428
429 on('session.start', async ($, e, next) => {
430 await $.command.register({
431 name: COMMAND,
432 description: 'Keep this session\'s prompt cache warm: on [3h | until 18:00], off, or status',
433 argumentHint: '[on [3h|until 18:00] | off]',
434 immediate: true,
435 })
436
437 live.storeKey = STORE_PREFIX + (await $.session.id())
438 live.isFirstRequest = true
439 const plan = (await $.store.get(PLAN_TTL_KEY)) as KeepWarmTtl | undefined
440 if (plan && typeof plan.ms === 'number') live.planTtl = plan
441
442 // After an app relaunch the session's values are gone; the store keeps them per session id.
443 // A reopened session may already have its last answer's time from Claude Code (the
444 // classic.SessionStart hook below, which can run before or after this one): keep the later.
445 const { value: restored = false } = await $.state.get(RESTORED)
446 if (!restored) {
447 const saved = (await $.store.get(live.storeKey)) as KeepWarmSaved | undefined
448 if (saved) {
449 if (saved.ttl && typeof saved.ttl.ms === 'number') live.savedTtl = saved.ttl
450 const { value: last = 0 } = await $.state.get(LAST)
451 if (saved.lastRequestAt > last) await $.state.set(LAST, saved.lastRequestAt)
452 await $.state.set(KEEP_WARM, saved.keepWarm)
453 await $.state.set(STOP_AT, saved.stopAt ?? 0)
454 }
455 await $.state.set(RESTORED, true)
456 await save($)
457 }
458
459 // Forget sessions untouched for a week so the store stays small. Aged by the last save: a
460 // session /compact or /clear just reset has no last request, yet is in use (counting from 0
461 // deleted its entry, keep-warm switch included). An old entry with no date at all is kept
462 // while its keep-warm is on.
463 const t = await $.clock.now()
464 for (const key of await $.store.keys()) {
465 if (!key.startsWith(STORE_PREFIX) || key === live.storeKey) continue
466 const old = (await $.store.get(key)) as KeepWarmSaved | undefined
467 const at = old ? (old.savedAt ?? old.lastRequestAt) : 0
468 if (!old || (at > 0 ? t - at > STORE_MAX_AGE_MS : !old.keepWarm)) await $.store.delete(key)
469 }
470
471 await tick($)
472 $.clock.every(TICK_MS, () => void tick($))
473 void syncTitle($)
474 $.clock.every(TITLE_SYNC_MS, () => void syncTitle($))
475
476 return next(e)
477 })
478
479 // A session reopened (an app relaunch, a resume): Claude Code says how long ago its last
480 // answer came and how big the conversation is, so the countdown carries on before the first
481 // message. After /clear the conversation starts over: no cache to count down yet. Its
482 // 'compact' is not used: it also comes when a subagent compacts its own conversation, with
483 // nothing saying so (no agent_id), which reset the main countdown. session.compact says whose.
484 on('classic.SessionStart', async ($, e, next) => {
485 if (e.source === 'clear') await startOver($)
486 else if ((e.source === 'resume' || e.source === 'fork') && typeof e.seconds_since_last_response === 'number') {
487 const t = await $.clock.now()
488 const lastAt = t - e.seconds_since_last_response * 1000 - RESUME_MARGIN_MS
489 const { value: last = 0 } = await $.state.get(LAST)
490 if (lastAt > last) await $.state.set(LAST, lastAt)
491 if (e.context_tokens) await $.state.set(CONTEXT, e.context_tokens)
492 await $.state.set(NOW, t)
493 // At a relaunch this can run before session.start has restored the saved values: saving
494 // now would write the defaults (keep-warm off) over them. session.start saves after it.
495 const { value: restored = false } = await $.state.get(RESTORED)
496 if (restored) await save($)
497 }
498 return next(e)
499 })
500
501 // A compaction of the main conversation (/compact, the threshold, a plugin) replaces it with its
502 // summary, so the countdown starts over. A subagent's own (agentId) is not this cache, a
503 // precomputed one is not applied yet, and a skipped one changed nothing. Read only: the
504 // compaction passes on unchanged.
505 on('session.compact', async ($, e, next) => {
506 const result = await next(e)
507 if (!e.agentId && e.trigger !== 'precompute' && result.skip === undefined) {
508 try {
509 await startOver($)
510 } catch {
511 // Never fail a compaction over the countdown.
512 }
513 }
514 return result
515 })
516
517 // Each model request of the main conversation (not a subagent's: its cache is its own):
518 // it restarts the cache's life, and its usage shows whether the cache outlived the gap.
519 on('turn.step', async function* ($, e, next) {
520 if (e.agentId) return yield* next(e)
521 const t = await $.clock.now()
522 const { value: previous = 0 } = await $.state.get(LAST)
523 live.pingPendingUntil = 0
524 await $.state.set(LAST, t)
525 await $.state.set(NOW, t)
526 await save($)
527 const result = yield* next(e)
528 // A request that never reached the model (an error, the conversation over its limit) did not
529 // touch the cache, so its life still runs from the request before.
530 const u = result?.usage
531 if (!u || u.input_tokens + u.cache_read_input_tokens + u.cache_creation_input_tokens === 0) {
532 await $.state.set(LAST, previous)
533 await save($)
534 return result
535 }
536 await $.state.set(WAITS, false)
537 await $.state.set(PING_ERROR, '')
538 if (live.visiblePingAt > 0) {
539 const { cache_read_input_tokens: read, cache_creation_input_tokens: written } = u
540 const done: KeepWarmPing = { at: live.visiblePingAt, isHit: read > 0 && read >= HIT_SHARE * (read + written), read }
541 const { value: pings = [] } = await $.state.get(PINGS)
542 await $.state.set(PINGS, [...pings, done].slice(-MAX_PINGS))
543 live.visiblePingAt = 0
544 live.pingTurnId = e.turnId
545 }
546 // Not the first request after a start: a restart can change the prompt itself (a new
547 // Claude Code version), which re-writes the cache whatever the gap, so it proves nothing.
548 const isFirst = live.isFirstRequest
549 live.isFirstRequest = false
550 if (previous > 0 && !isFirst) await observe($, t - previous, u)
551 return result
552 })
553
554 // A turn with real requests behind it: the conversation has answered in this run (so the next
555 // ping can be invisible, and an empty plan list now means an API key), and unless it was the
556 // mod's own visible ping, the conversation is in use, yours or a turn it ran by itself (an
557 // agent's report, another session's message): the pings in a row start over. A turn that died
558 // on an error has no usage; if it was the visible ping, the line says why it failed.
559 on('turn.complete', async ($, e, next) => {
560 if (!e.agentId) {
561 const isPingTurn = e.turnId === live.pingTurnId
562 if (isPingTurn) live.pingTurnId = ''
563 if (e.usage) {
564 await $.state.set(SAW_RESPONSE, true)
565 if (!isPingTurn) await $.state.set(PINGS, [])
566 } else if (live.visiblePingAt > 0) {
567 live.visiblePingAt = 0
568 if (e.reason === 'error') {
569 const why = e.answer.trim() || 'an API error'
570 await $.state.set(PING_ERROR, why)
571 $.ui.toast(`Keep-warm ping failed: ${why}. The cache will expire normally.`)
572 }
573 }
574 }
575 return next(e)
576 })
577
578 // While Claude's question waits for an answer the turn stays open; the band says so.
579 on('tool.call', { tool: 'AskUserQuestion' }, async ($, e, next) => {
580 await $.state.set(WAITING, true)
581 try {
582 return await next(e)
583 } finally {
584 await $.state.set(WAITING, false)
585 }
586 })
587
588 on('command.run', { command: COMMAND }, async ($, e) => {
589 const t = await $.clock.now()
590 const words = e.args.trim().toLowerCase()
591 const [first = ''] = words.split(/\s+/)
592 if (first === 'off') {
593 await setKeepWarm($, false, 0)
594 } else if (words !== '' && words !== 'status') {
595 const stopAt = parseStop(first === 'on' ? words.slice(2).trim() : words, t)
596 if (stopAt < 0) return { text: 'Not understood. Try: /keepwarm on, /keepwarm 3h, /keepwarm until 18:00, /keepwarm off' }
597 await setKeepWarm($, true, stopAt)
598 }
599
600 const { value: isOn = false } = await $.state.get(KEEP_WARM)
601 const { value: stopAt = 0 } = await $.state.get(STOP_AT)
602 const { value: ttl = null } = await $.state.get(TTL)
603 const { value: pings = [] } = await $.state.get(PINGS)
604 const { value: waits = false } = await $.state.get(WAITS)
605 const { value: sawResponse = false } = await $.state.get(SAW_RESPONSE)
606 const lines = [
607 isOn
608 ? `Keep-warm on${stopAt ? ` until ${clockTime(stopAt)}` : ` (stops itself after ${MAX_PINGS} pings in a row)`}: about ${minutes(PING_LEAD_MS)} min before the cache would expire, it re-reads the conversation from cache in the background. A cold cache is never pinged.`
609 : 'Keep-warm off: the cache expires normally.',
610 ]
611 if (isOn && (waits || !sawResponse)) lines.push('Claude Code can repeat this conversation invisibly only once it has answered since the app started, so the next ping is a short message in the chat.')
612 if (ttl) lines.push(`This session's cache lasts ${cacheLength(ttl.ms)} (${ttl.reason === 'observed' ? 'seen expiring sooner than expected' : ttl.reason}).`)
613 if (ttl && ttl.ms < KEEP_WARM_MIN_TTL_MS) lines.push(`Keep-warm only pings a cache of ${minutes(KEEP_WARM_MIN_TTL_MS)} min or longer.`)
614 for (const p of pings.slice(-HISTORY_SHOWN)) {
615 lines.push(`${clockTime(p.at)} · ${p.isHit ? `read ${tokens(p.read)} from cache` : 'found the cache cold'}`)
616 }
617 return { text: lines.join('\n') }
618 })
619
620 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
621 if (e.props.hasSurvey) return next(e)
622
623 const { Box, Button, Text } = $.ui.resolve(e)
624 const { value: last = 0 } = await $.state.get(LAST)
625 const { value: now = 0 } = await $.state.get(NOW)
626 const { value: isOn = false } = await $.state.get(KEEP_WARM)
627 const { value: stopAt = 0 } = await $.state.get(STOP_AT)
628 const { value: pings = [] } = await $.state.get(PINGS)
629 const { value: ttl = null } = await $.state.get(TTL)
630 const { value: contextTokens = 0 } = await $.state.get(CONTEXT)
631 const { value: isWaiting = false } = await $.state.get(WAITING)
632 const { value: shortAfter = 0 } = await $.state.get(SHORT_AFTER)
633 const { value: waits = false } = await $.state.get(WAITS)
634 const { value: sawResponse = false } = await $.state.get(SAW_RESPONSE)
635 const { value: pingError = '' } = await $.state.get(PING_ERROR)
636 const { value: nearlyFull = false } = await $.state.get(NEARLY_FULL)
637 // The next ping has to be the visible one but the conversation is too full for it.
638 const isBlocked = (!sawResponse || waits) && nearlyFull
639
640 const t = Math.max(now, last)
641 const ttlMs = ttl?.ms ?? HOUR_MS
642 const canKeepWarm = ttl !== null && ttl.ms >= KEEP_WARM_MIN_TTL_MS
643 // Green only while keep-warm is on and able to ping; everything else stays default grey.
644 const isGreen = isOn && canKeepWarm && !pingError && !isBlocked
645 const expiresAt = last + ttlMs
646 const remaining = expiresAt - t
647 const isWarm = last > 0 && remaining > 0
648 const filled = isWarm ? Math.max(1, Math.ceil((remaining / ttlMs) * BAR_CELLS)) : 0
649 const quarters = CIRCLE_BY_QUARTER.length
650 const circle = isWarm
651 ? CIRCLE_BY_QUARTER[Math.min(quarters, Math.max(1, Math.ceil((remaining / ttlMs) * quarters))) - 1]
652 : CIRCLE_COLD
653
654 // The first part (time left) is drawn bold; the dot and the button already say whether
655 // keep-warm is on, so the text no longer repeats it.
656 const main = last === 0 ? 'nothing cached yet' : isWarm ? timeLeft(remaining) : `cold since ${clockTime(expiresAt)}`
657 // "58 min left (14:12)": the clock time the cache expires, right after the time left.
658 const expiresText = isWarm ? ` (${clockTime(expiresAt)})` : ''
659 const parts: string[] = []
660 if (isWaiting && isWarm) parts.push('waiting for your answer')
661
662 if (ttl && !canKeepWarm) {
663 if (ttl.reason === 'observed') parts.push(`cache expired after ${minutes(shortAfter)} min, sooner than expected`, 'keep-warm paused')
664 else if (ttl.reason === 'overage') parts.push(`${cacheLength(ttl.ms)} cache while over your plan limit`, 'keep-warm paused')
665 else parts.push(`${cacheLength(ttl.ms)} cache (${ttl.reason === 'api' ? 'API key' : 'your settings'})`, 'keep-warm needs the 1-hour cache')
666 } else if (isOn) {
667 const lastPing = pings[pings.length - 1]
668 if (stopAt) parts.push(`keep-warm stops ${clockTime(stopAt)}`)
669 // With the cache length unknown no ping can go out; before the conversation has answered in
670 // this run the next one is the visible ping.
671 if (pingError) parts.push(`ping failed: ${pingError}`)
672 else if (!ttl) parts.push(WAITS_TEXT)
673 else if (isBlocked) parts.push(NEARLY_FULL_TEXT)
674 else if (isWarm) parts.push(`next ping ${clockTime(last + ttlMs - PING_LEAD_MS)}${sawResponse && !waits ? '' : SHOWS_IN_CHAT}`)
675 if (lastPing) parts.push(`${pings.length} ping${pings.length === 1 ? '' : 's'} ${lastPing.isHit ? '✓' : '✗'}`)
676 }
677
678 if (contextTokens > 0 && last > 0) {
679 if (!isWarm) parts.push(`your next message re-writes ${tokens(contextTokens)}`, ...(isGreen ? ['keep-warm restarts after it'] : []))
680 else if (!isGreen && remaining < EXPIRING_MS) parts.push(`after ${clockTime(expiresAt)} your next message re-writes ${tokens(contextTokens)}`)
681 else parts.push(`${tokens(contextTokens)} cached`)
682 }
683
684 // No toggle where keep-warm can never ping (an API key, or a short cache set on purpose).
685 const canToggle = isOn || !ttl || canKeepWarm || ttl.reason === 'overage' || ttl.reason === 'observed'
686
687 // Each detail is its own element and the spacing comes from the layout (column gaps), which
688 // no surface squeezes the way it squeezes runs of spaces inside one text. In a narrow window
689 // nothing wraps: the time and the button keep their size and the details are cut at the edge.
690 return (
691 <Box alignItems="center" flexWrap="nowrap">
692 <Box flexGrow={1} flexShrink={1} minWidth={0} overflow="hidden" flexWrap="nowrap" alignItems="center" columnGap={DETAIL_GAP}>
693 <Box flexShrink={0}>
694 {isGreen ? <Text color="green" bold>{circle}</Text> : <Text dimColor>{circle}</Text>}
695 </Box>
696 <Box flexShrink={0}>
697 {isGreen ? <Text color="green">{BAR_FULL.repeat(filled)}</Text> : <Text dimColor>{BAR_FULL.repeat(filled)}</Text>}
698 <Text dimColor>{BAR_EMPTY.repeat(BAR_CELLS - filled)}</Text>
699 </Box>
700 <Box flexShrink={0}>
701 <Text bold>{main}</Text>
702 <Text dimColor>{expiresText}</Text>
703 </Box>
704 {parts.flatMap(p => [
705 <Box flexShrink={0}><Text dimColor>·</Text></Box>,
706 <Box flexShrink={0}><Text dimColor>{p}</Text></Box>,
707 ])}
708 </Box>
709 {canToggle && (
710 <Box flexShrink={0}>
711 <Button
712 key="toggle"
713 label={isOn ? 'Stop' : 'Keep warm'}
714 onPress={() => setKeepWarm($, !isOn, 0)}
715 />
716 </Box>
717 )}
718 </Box>
719 )
720 })
721}
722types/index.d.ts 55 lines1/**
2 * What the mod remembers per session id in `$.store`, so a relaunch keeps it. `ttl` is the cache
3 * length in effect when it was saved (absent before 1.0.3); `savedAt` when it was saved, which
4 * ages the entry for the weekly clean-up (absent before 1.0.4).
5 */
6export type KeepWarmSaved = { lastRequestAt: number; keepWarm: boolean; stopAt: number; ttl?: KeepWarmTtl; savedAt?: number }
7
8/** One keep-warm ping: when it started, whether the cache still held the conversation, how much it read. */
9export type KeepWarmPing = { at: number; isHit: boolean; read: number }
10
11/**
12 * How long the main conversation's cache lives, and why: from a setting, from the plan, or
13 * corrected by what the cache actually did. Null until it can be told. The last one the plan
14 * showed is also kept in `$.store` under `plan-ttl`, shared by every session.
15 */
16export type KeepWarmTtl = { ms: number; reason: 'subscription' | 'api' | 'overage' | 'setting' | 'observed' } | null
17
18declare module 'claude-code' {
19 interface PluginState {
20 'cache-keep-warm': {
21 /** When the main conversation's last model request started (ms since epoch); 0 = unknown. */
22 lastRequestAt: number
23 /** Whether this session keeps its cache warm. */
24 keepWarm: boolean
25 /** When keep-warm switches itself off (ms since epoch); 0 = after the ping limit. */
26 stopAt: number
27 /** The clock as the band last showed it. */
28 now: number
29 /** Pings since keep-warm was last switched on, oldest first. */
30 pings: KeepWarmPing[]
31 /** The cache lifetime as last worked out. */
32 ttl: KeepWarmTtl
33 /** Tokens the last request sent: what a cold cache would make the next message re-write. */
34 contextTokens: number
35 /** True while Claude waits on an answer to its question. */
36 isWaiting: boolean
37 /**
38 * True once this process has seen a main-thread response: the ping can then be invisible (a
39 * fork), and an empty plan list then means an API key. Reset by /clear.
40 */
41 sawResponse: boolean
42 /** The idle gap after which the cache was found cold although it should have lasted; 0 = none seen. */
43 shortAfterMs: number
44 /** True when the invisible ping found nothing to repeat, so the visible one is used, until the next request. */
45 waitsForMessage: boolean
46 /** True once this process has restored the session's saved values from the store. */
47 restored: boolean
48 /** Why the last ping did not reach the model (Claude Code's words or the HTTP status); '' = none. */
49 pingError: string
50 /** True while the conversation fills too much of the context window for a visible ping. */
51 nearlyFull: boolean
52 }
53 }
54}
55