Keeps the prompt cache of an idle Claude Code session warm: just before the cache lapses it submits a short service prompt, so your next message reads the…

Мод для Claude Code, который не дает кэшу промпта истечь, пока сессия простаивает. За пару минут до истечения кэша мод отправляет в чат короткое служебное сообщение. Модель читает контекст из кэша, и срок жизни кэша начинается заново.
Claude Code кэширует начало промпта: системные инструкции, файлы, историю разговора. На подписке Claude кэш живет 1 час, на API-ключе 5 минут, и каждый запрос, который его читает, продлевает срок. По документации Anthropic чтение из кэша стоит 0,1 цены обычного ввода, запись в кэш на 1 час стоит 2x, на 5 минут 1,25x.
Пример: контекст 150 тыс. токенов, вы ушли на обед. Без мода первое сообщение после перерыва заново запишет весь контекст, это 300 тыс. токенов в пересчете на обычный ввод. Пинг мода прочитает его из кэша за 15 тыс. и короткий ответ «ok». Один пинг примерно в 20 раз дешевле перезаписи.
Пинг выгоден, только если вы вернетесь. Поэтому мод выключен по умолчанию и делает не больше 3 пингов подряд, потом кэш истекает сам.
Нужен Claude Code 2.1.259 или новее: моды работают на функциональных хуках, это ранний доступ. Мод проверен в Claude Code 2.1.287, в настольном приложении флаг не понадобился. Если мод не загрузился, задайте переменную CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1.
macOS, Linux, Git Bash:
git clone --depth 1 https://github.com/mupamuc/cache-keepwarm ~/.claude/skills/cache-keepwarm
Windows PowerShell:
git clone --depth 1 https://github.com/mupamuc/cache-keepwarm "$HOME\.claude\skills\cache-keepwarm"
Мод подключится в новой сессии. Обновление: git pull в этой папке. Удаление: удалить папку.
Над полем ввода мод рисует одну строку:
● cache 99% · 56m [x] keepwarm 12:58 · 0/3
Слева состояние кэша: цвет точки (зеленый, желтый перед пингом, красный когда кэш истек), доля промпта из кэша в последнем запросе и сколько кэшу осталось жить. Справа галочка: щелчок включает и выключает мод. При включенном моде рядом видно время следующего пинга и сколько пингов подряд уже было.
То же самое делают команды:
| Команда | Что делает |
|---|---|
/keepwarm on | включает мод, выбор сохраняется между сессиями |
/keepwarm off | выключает |
/keepwarm status | показывает, когда будет пинг и сколько пингов осталось |
Перед пингом приходит уведомление, сам пинг виден в чате как сообщение от мода.
Мод отправляет пинг, только если выполнены все условия:
leadSeconds, но кэш еще жив (пинг позже чем за 15 секунд до истечения может не успеть);minTokens токенов;maxPings пингов.Мод запоминает время последнего запроса вместе с id сессии. После перезапуска приложения или --resume той же сессии отсчет продолжается, ждать вашего сообщения не нужно.
Ваше сообщение, уведомление фоновой задачи или сообщение другой сессии сбрасывает счетчик пингов. Запросы субагентов мод не учитывает: у них свой кэш.
| Параметр | По умолчанию | Смысл |
|---|---|---|
enabled | false | включен ли мод, если /keepwarm еще ни разу не вызывали |
ttl | 1h | срок кэша: 1h на подписке, 5m на API-ключе, у облачного провайдера, при оплате по использованию |
leadSeconds | 120 | за сколько секунд до истечения слать пинг, не больше половины срока |
maxPings | 3 | сколько пингов подряд без вашего сообщения |
minTokens | 20000 | кэш меньше этого дешево пересоздать, его мод не трогает |
message | пусто | свой текст пинга, пусто значит встроенный |
band | true | строка над полем ввода с галочкой |
status | false | короткая запись в строке состояния |
В терминальной версии параметры меняются через /config. В настольном приложении их задают в ~/.claude/settings.json:
{
"pluginConfigs": {
"cache-keepwarm@skills-dir": {
"options": { "ttl": "5m", "maxPings": 2 }
}
}
}
04.10.2026, Claude Code 2.1.286, настольное приложение для Windows, подписка Claude, контекст около 190 тыс. токенов. Мод отправил три пинга подряд примерно раз в час. Модель на каждый ответила одним «ok» без вызова инструментов. После третьего пинга мод остановился (3/3 wait). Доля попаданий в кэш после пингов была 100 %, то есть контекст ни разу не записывался заново.
Восстановление отсчета после перезапуска приложения в живой сессии пока не проверено.
maxPings пингов потрачены зря.Чтобы видеть состояние кэша, поставьте мод prompt-cache-control из каталога claude-code-templates. Он показывает над полем ввода долю попаданий в кэш и отсчет до истечения. Сам он ничего не отправляет, только показывает и напоминает. cache-keepwarm делает то, о чем тот напоминает.
npx claude-code-templates@latest --mod observability/prompt-cache-control
Команда ставит мод в папку текущего проекта. Чтобы он работал во всех проектах, перенесите папку .claude/skills/prompt-cache-control в ~/.claude/skills/.
cache-keepwarm is a Claude Code mod that keeps the prompt cache of an idle session warm. Shortly before the cache lapses (1 hour on a Claude subscription, 5 minutes on an API key) it submits a short service prompt; the model reads the context from the cache at 0.1x the input price, and the cache lifetime restarts. Without it, your next message would write the whole context again at 2x (1-hour cache) or 1.25x (5-minute cache).
Off by default. /keepwarm on|off|status. It pings only while the session is idle, only inside the last leadSeconds before expiry, never after the cache has lapsed, at most maxPings (default 3) times in a row; your own prompt resets the count.
Install: git clone --depth 1 https://github.com/mupamuc/cache-keepwarm ~/.claude/skills/cache-keepwarm, then start a new session. Requires Claude Code 2.1.259+ (function hooks, early access).
Companion mod for watching the cache: prompt-cache-control by davila7 and contributors, MIT.
MIT, см. LICENSE.
hooks/keepwarm.tsx 280 lines1/**
2 * cache-keepwarm - Claude Code mod (function hooks, early access)
3 *
4 * Keeps the prompt cache of an idle session warm. Every main-loop request
5 * reads or writes the cache and restarts its lifetime (1 hour on a Claude
6 * subscription, 5 minutes on an API key). When the session has been idle long
7 * enough that the cache is about to lapse, this mod submits a short service
8 * prompt; the model reads the whole prefix from the cache (about 10 % of the
9 * input price) and the lifetime starts over. Without it the next prompt would
10 * write the whole prefix again (1.25x or 2x the input price).
11 *
12 * - one row above the prompt: cache hit rate, time left, and a keepwarm
13 * checkbox; `/keepwarm on|off|status` does the same from the prompt
14 * - off by default; the choice is kept across sessions
15 * - pings only while idle, only inside the last `leadSeconds` before expiry,
16 * never after the cache has already lapsed
17 * - at most `maxPings` pings in a row; any prompt that is not ours resets it
18 * - skips caches smaller than `minTokens`
19 *
20 * Options (pluginConfigs["cache-keepwarm@skills-dir"].options):
21 * enabled: boolean keep the cache warm when the checkbox was never used (default false)
22 * ttl: "1h" | "5m" cache lifetime the session runs with (default 1h)
23 * leadSeconds: number ping this long before expiry (default 120)
24 * maxPings: number pings in a row without your own prompt (default 3)
25 * minTokens: number smallest cached prompt worth keeping (default 20000)
26 * message: string the service prompt the model receives
27 * band: boolean the row above the prompt (default true)
28 * status: boolean a short entry in the status line (default false)
29 */
30import type { EngineInterface, Register } from 'claude-code'
31
32const COMMAND = 'keepwarm'
33const STORE_KEY = 'enabled'
34const LAST_KEY = 'last'
35const TICK_MS = 5_000
36// a ping needs a few seconds to reach the API; closer to expiry it may miss
37const SAFETY_MS = 15_000
38const DEFAULT_MESSAGE =
39 'Service ping from the cache-keepwarm mod to keep the prompt cache warm. ' +
40 'Reply with the single word "ok" and do nothing else: no tools, no summary.'
41
42// read: tokens the cache served; tokens: the whole prompt (read + written + uncached)
43type Last = { at: number; tokens: number; read: number }
44type SavedLast = Last & { session: string }
45type Config = {
46 enabled: boolean
47 ttlMs: number
48 leadMs: number
49 maxPings: number
50 minTokens: number
51 message: string
52 band: boolean
53 status: boolean
54}
55
56let cfg: Config = {
57 enabled: false,
58 ttlMs: 60 * 60_000,
59 leadMs: 120_000,
60 maxPings: 3,
61 minTokens: 20_000,
62 message: DEFAULT_MESSAGE,
63 band: true,
64 status: false,
65}
66let last: Last | undefined
67let sessionId = ''
68let busy = false
69let pending = false
70let pings = 0
71let switchedOn: boolean | undefined
72let timer: { cancel: () => void } | undefined
73let drawnKey = ''
74
75function num(value: unknown, fallback: number): number {
76 return typeof value === 'number' && Number.isFinite(value) && value >= 0 ? value : fallback
77}
78
79function fmtTokens(n: number): string {
80 return n >= 1000 ? `${Math.round(n / 1000)}k` : String(n)
81}
82
83function fmtTime(ms: number): string {
84 const d = new Date(ms)
85 return `${String(d.getHours()).padStart(2, '0')}:${String(d.getMinutes()).padStart(2, '0')}`
86}
87
88// minutes while there are minutes, seconds at the end: the row redraws only when this text changes
89function fmtLeft(ms: number): string {
90 return ms >= 60_000 ? `${Math.floor(ms / 60_000)}m` : `${Math.ceil(ms / 1000)}s`
91}
92
93function isOn(): boolean {
94 return switchedOn ?? cfg.enabled
95}
96
97// short on purpose: the status line and the checkbox already name the plugin
98function describe(now: number): string {
99 if (!last) return 'on'
100 if (last.tokens < cfg.minTokens) return `skip <${fmtTokens(cfg.minTokens)}`
101 if (now >= last.at + cfg.ttlMs) return 'lapsed'
102 if (pings >= cfg.maxPings) return `${pings}/${cfg.maxPings} wait`
103 return `${fmtTime(last.at + cfg.ttlMs - cfg.leadMs)} · ${pings}/${cfg.maxPings}`
104}
105
106// the /keepwarm reply, read once, can say it in words
107function explain(now: number): string {
108 if (!isOn()) return 'off'
109 if (!last) return 'on, waiting for the first request'
110 if (last.tokens < cfg.minTokens) return `on, ${fmtTokens(last.tokens)} cached is below ${fmtTokens(cfg.minTokens)}: no ping`
111 if (now >= last.at + cfg.ttlMs) return 'on, the cache already lapsed: no ping'
112 if (pings >= cfg.maxPings) return `on, ${pings}/${cfg.maxPings} pings used, waiting for your prompt`
113 return `on, next ping at ${fmtTime(last.at + cfg.ttlMs - cfg.leadMs)}, ${pings}/${cfg.maxPings} used`
114}
115
116function cacheText(now: number): string {
117 if (!last) return 'cache: no request yet'
118 const left = last.at + cfg.ttlMs - now
119 const hit = last.tokens > 0 ? Math.round((last.read / last.tokens) * 100) : 0
120 return `cache ${hit}% · ${left > 0 ? fmtLeft(left) : 'lapsed'}`
121}
122
123function cacheColor(now: number): string | undefined {
124 if (!last) return undefined
125 const left = last.at + cfg.ttlMs - now
126 return left <= 0 ? 'red' : left <= cfg.leadMs ? 'yellow' : 'green'
127}
128
129function refresh($: EngineInterface) {
130 const now = Date.now()
131 if (cfg.status) $.ui.status(isOn() ? describe(now) : undefined)
132 const key = `${isOn()}|${cacheText(now)}|${describe(now)}`
133 if (key !== drawnKey) {
134 drawnKey = key
135 $.ui.invalidate('ui.render')
136 }
137}
138
139async function setOn($: EngineInterface, on: boolean) {
140 switchedOn = on
141 pings = 0
142 await $.store.set(STORE_KEY, on)
143 refresh($)
144}
145
146async function maybePing($: EngineInterface) {
147 if (!isOn() || busy || pending || !last) return
148 if (last.tokens < cfg.minTokens || pings >= cfg.maxPings) return
149 const now = Date.now()
150 const expires = last.at + cfg.ttlMs
151 if (now < expires - cfg.leadMs || now > expires - SAFETY_MS) return
152 pending = true
153 pings += 1
154 $.ui.toast(`keepwarm: ping ${pings}/${cfg.maxPings} keeps ${fmtTokens(last.tokens)} tokens cached`)
155 try {
156 await $.prompt.submit({ text: cfg.message })
157 } catch (err) {
158 pending = false
159 $.ui.log(`cache-keepwarm: ping not sent: ${err}`)
160 }
161 refresh($)
162}
163
164export const register: Register = (on, options) => {
165 const ttlMs = options.ttl === '5m' ? 5 * 60_000 : 60 * 60_000
166 cfg = {
167 enabled: options.enabled === true,
168 ttlMs,
169 leadMs: Math.min(num(options.leadSeconds, 120) * 1000, ttlMs / 2),
170 maxPings: Math.floor(num(options.maxPings, 3)),
171 minTokens: num(options.minTokens, 20_000),
172 message: typeof options.message === 'string' && options.message.trim() ? options.message : DEFAULT_MESSAGE,
173 band: options.band !== false,
174 status: options.status === true,
175 }
176
177 on('session.start', async ($, e, next) => {
178 const r = await next(e)
179 last = undefined
180 busy = false
181 pending = false
182 pings = 0
183 drawnKey = ''
184 const stored = await $.store.get(STORE_KEY).catch(() => undefined)
185 switchedOn = typeof stored === 'boolean' ? stored : undefined
186 // an app restart or a resume reloads the mod: the same session keeps its cache, so keep counting from it
187 sessionId = await $.session.id().catch(() => '')
188 const saved = (await $.store.get(LAST_KEY).catch(() => undefined)) as SavedLast | undefined
189 if (saved && sessionId && saved.session === sessionId && typeof saved.at === 'number' && typeof saved.tokens === 'number') {
190 last = { at: saved.at, tokens: saved.tokens, read: typeof saved.read === 'number' ? saved.read : 0 }
191 }
192 await $.command
193 .register({
194 name: COMMAND,
195 description: 'Keep the prompt cache warm while idle: on, off or status',
196 argumentHint: '[on|off|status]',
197 immediate: true,
198 })
199 .catch(err => $.ui.log(`cache-keepwarm: /${COMMAND} not registered: ${err}`))
200 timer?.cancel()
201 timer = $.clock.every(TICK_MS, () => {
202 refresh($)
203 void maybePing($)
204 })
205 refresh($)
206 return r
207 })
208
209 on('session.end', async ($, e, next) => {
210 last = undefined
211 pending = false
212 pings = 0
213 if (e.reason !== 'clear') {
214 timer?.cancel()
215 timer = undefined
216 }
217 return next(e)
218 })
219
220 // any prompt that reaches this hook is not ours: $.prompt.submit skips the calling hook
221 on('prompt.submit', async ($, e, next) => {
222 pings = 0
223 return next(e)
224 })
225
226 on('turn.start', async ($, e, next) => {
227 busy = true
228 return next(e)
229 })
230
231 on('turn.step', async function* ($, e, next) {
232 if (e.agentId) return yield* next(e)
233 const at = Date.now()
234 const r = yield* next(e)
235 const u = r.usage
236 if (u && u.cache_read_input_tokens + u.cache_creation_input_tokens > 0) {
237 last = {
238 at,
239 tokens: u.cache_read_input_tokens + u.cache_creation_input_tokens + u.input_tokens,
240 read: u.cache_read_input_tokens,
241 }
242 pending = false
243 if (sessionId) void $.store.set(LAST_KEY, { session: sessionId, ...last }).catch(() => undefined)
244 refresh($)
245 }
246 return r
247 })
248
249 on('turn.complete', async ($, e, next) => {
250 const r = await next(e)
251 if (!e.agentId) {
252 busy = false
253 pending = false
254 refresh($)
255 }
256 return r
257 })
258
259 on('command.run', { command: COMMAND }, async ($, e) => {
260 const arg = e.args.trim().toLowerCase()
261 if (arg === 'on' || arg === 'off') await setOn($, arg === 'on')
262 return { text: explain(Date.now()) }
263 })
264
265 // one row: cache state on the left, the keepwarm checkbox on the right
266 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
267 if (!cfg.band || e.props.hasSurvey) return next(e)
268 const { Box, Text, Button } = $.ui.resolve(e)
269 const now = Date.now()
270 const label = isOn() ? `[x] keepwarm ${describe(now)}` : '[ ] keepwarm'
271 return (
272 <Box flexDirection="row" columnGap={1}>
273 <Text color={cacheColor(now)}>●</Text>
274 <Text dimColor>{cacheText(now)}</Text>
275 <Button key="toggle" plain label={label} dimColor={!isOn()} onPress={() => setOn($, !isOn())} />
276 </Box>
277 )
278 })
279}
280