SLOPSHOPPER

cache-keepwarm

Keeps the prompt cache of an idle Claude Code session warm: just before the cache lapses it submits a short service prompt, so your next message reads the…

newbandcommandtoaststatusprompt
v0.1.0MITupdated 2026-10-04mupamuc/cache-keepwarm
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · cache-keepwarm
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /keepwarm ⎿ cache-keepwarm: off ● cache: no request yet [ ] keepwarm ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Band
● cache: no request yet [ ] keepwarm
README

cache-keepwarm

Мод для Claude Code, который не дает кэшу промпта истечь, пока сессия простаивает. За пару минут до истечения кэша мод отправляет в чат короткое служебное сообщение. Модель читает контекст из кэша, и срок жизни кэша начинается заново.

English below

Зачем

Claude Code кэширует начало промпта: системные инструкции, файлы, историю разговора. На подписке Claude кэш живет 1 час, на API-ключе 5 минут, и каждый запрос, который его читает, продлевает срок. По документации Anthropic чтение из кэша стоит 0,1 цены обычного ввода, запись в кэш на 1 час стоит 2x, на 5 минут 1,25x.

Пример: контекст 150 тыс. токенов, вы ушли на обед. Без мода первое сообщение после перерыва заново запишет весь контекст, это 300 тыс. токенов в пересчете на обычный ввод. Пинг мода прочитает его из кэша за 15 тыс. и короткий ответ «ok». Один пинг примерно в 20 раз дешевле перезаписи.

Пинг выгоден, только если вы вернетесь. Поэтому мод выключен по умолчанию и делает не больше 3 пингов подряд, потом кэш истекает сам.

Установка

Нужен Claude Code 2.1.259 или новее: моды работают на функциональных хуках, это ранний доступ. Мод проверен в Claude Code 2.1.287, в настольном приложении флаг не понадобился. Если мод не загрузился, задайте переменную CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1.

macOS, Linux, Git Bash:

git clone --depth 1 https://github.com/mupamuc/cache-keepwarm ~/.claude/skills/cache-keepwarm

Windows PowerShell:

git clone --depth 1 https://github.com/mupamuc/cache-keepwarm "$HOME\.claude\skills\cache-keepwarm"

Мод подключится в новой сессии. Обновление: git pull в этой папке. Удаление: удалить папку.

Как пользоваться

Над полем ввода мод рисует одну строку:

● cache 99% · 56m  [x] keepwarm 12:58 · 0/3

Слева состояние кэша: цвет точки (зеленый, желтый перед пингом, красный когда кэш истек), доля промпта из кэша в последнем запросе и сколько кэшу осталось жить. Справа галочка: щелчок включает и выключает мод. При включенном моде рядом видно время следующего пинга и сколько пингов подряд уже было.

То же самое делают команды:

КомандаЧто делает
/keepwarm onвключает мод, выбор сохраняется между сессиями
/keepwarm offвыключает
/keepwarm statusпоказывает, когда будет пинг и сколько пингов осталось

Перед пингом приходит уведомление, сам пинг виден в чате как сообщение от мода.

Правила пинга

Мод отправляет пинг, только если выполнены все условия:

  • мод включен;
  • сессия простаивает, модель ничего не делает;
  • до истечения кэша осталось меньше leadSeconds, но кэш еще жив (пинг позже чем за 15 секунд до истечения может не успеть);
  • в кэше не меньше minTokens токенов;
  • после вашего последнего сообщения было меньше maxPings пингов.

Мод запоминает время последнего запроса вместе с id сессии. После перезапуска приложения или --resume той же сессии отсчет продолжается, ждать вашего сообщения не нужно.

Ваше сообщение, уведомление фоновой задачи или сообщение другой сессии сбрасывает счетчик пингов. Запросы субагентов мод не учитывает: у них свой кэш.

Настройки

ПараметрПо умолчаниюСмысл
enabledfalseвключен ли мод, если /keepwarm еще ни разу не вызывали
ttl1hсрок кэша: 1h на подписке, 5m на API-ключе, у облачного провайдера, при оплате по использованию
leadSeconds120за сколько секунд до истечения слать пинг, не больше половины срока
maxPings3сколько пингов подряд без вашего сообщения
minTokens20000кэш меньше этого дешево пересоздать, его мод не трогает
messageпустосвой текст пинга, пусто значит встроенный
bandtrueстрока над полем ввода с галочкой
statusfalseкороткая запись в строке состояния

В терминальной версии параметры меняются через /config. В настольном приложении их задают в ~/.claude/settings.json:

{
  "pluginConfigs": {
    "cache-keepwarm@skills-dir": {
      "options": { "ttl": "5m", "maxPings": 2 }
    }
  }
}

Проверка

04.10.2026, Claude Code 2.1.286, настольное приложение для Windows, подписка Claude, контекст около 190 тыс. токенов. Мод отправил три пинга подряд примерно раз в час. Модель на каждый ответила одним «ok» без вызова инструментов. После третьего пинга мод остановился (3/3 wait). Доля попаданий в кэш после пингов была 100 %, то есть контекст ни разу не записывался заново.

Восстановление отсчета после перезапуска приложения в живой сессии пока не проверено.

Ограничения

  • Мод работает, пока сессия открыта. Закрытое приложение или спящий компьютер пинг не отправят.
  • Пинг и ответ «ok» остаются в истории разговора, по несколько десятков токенов на каждый пинг.
  • Мод работает только в Claude Code: моды используют API функциональных хуков Claude Code.
  • Мод не знает, когда вы вернетесь. Если не вернетесь, maxPings пингов потрачены зря.
  • API модов в раннем доступе и может измениться между версиями Claude Code.

Вместе с prompt-cache-control

Чтобы видеть состояние кэша, поставьте мод prompt-cache-control из каталога claude-code-templates. Он показывает над полем ввода долю попаданий в кэш и отсчет до истечения. Сам он ничего не отправляет, только показывает и напоминает. cache-keepwarm делает то, о чем тот напоминает.

npx claude-code-templates@latest --mod observability/prompt-cache-control

Команда ставит мод в папку текущего проекта. Чтобы он работал во всех проектах, перенесите папку .claude/skills/prompt-cache-control в ~/.claude/skills/.

Авторы и благодарности

English

cache-keepwarm is a Claude Code mod that keeps the prompt cache of an idle session warm. Shortly before the cache lapses (1 hour on a Claude subscription, 5 minutes on an API key) it submits a short service prompt; the model reads the context from the cache at 0.1x the input price, and the cache lifetime restarts. Without it, your next message would write the whole context again at 2x (1-hour cache) or 1.25x (5-minute cache).

Off by default. /keepwarm on|off|status. It pings only while the session is idle, only inside the last leadSeconds before expiry, never after the cache has lapsed, at most maxPings (default 3) times in a row; your own prompt resets the count.

Install: git clone --depth 1 https://github.com/mupamuc/cache-keepwarm ~/.claude/skills/cache-keepwarm, then start a new session. Requires Claude Code 2.1.259+ (function hooks, early access).

Companion mod for watching the cache: prompt-cache-control by davila7 and contributors, MIT.

Лицензия

MIT, см. LICENSE.

Source 1 files
hooks/keepwarm.tsx 280 lines
1/**
2 * cache-keepwarm - Claude Code mod (function hooks, early access)
3 *
4 * Keeps the prompt cache of an idle session warm. Every main-loop request
5 * reads or writes the cache and restarts its lifetime (1 hour on a Claude
6 * subscription, 5 minutes on an API key). When the session has been idle long
7 * enough that the cache is about to lapse, this mod submits a short service
8 * prompt; the model reads the whole prefix from the cache (about 10 % of the
9 * input price) and the lifetime starts over. Without it the next prompt would
10 * write the whole prefix again (1.25x or 2x the input price).
11 *
12 *   - one row above the prompt: cache hit rate, time left, and a keepwarm
13 *     checkbox; `/keepwarm on|off|status` does the same from the prompt
14 *   - off by default; the choice is kept across sessions
15 *   - pings only while idle, only inside the last `leadSeconds` before expiry,
16 *     never after the cache has already lapsed
17 *   - at most `maxPings` pings in a row; any prompt that is not ours resets it
18 *   - skips caches smaller than `minTokens`
19 *
20 * Options (pluginConfigs["cache-keepwarm@skills-dir"].options):
21 *   enabled: boolean     keep the cache warm when the checkbox was never used (default false)
22 *   ttl: "1h" | "5m"     cache lifetime the session runs with (default 1h)
23 *   leadSeconds: number  ping this long before expiry (default 120)
24 *   maxPings: number     pings in a row without your own prompt (default 3)
25 *   minTokens: number    smallest cached prompt worth keeping (default 20000)
26 *   message: string      the service prompt the model receives
27 *   band: boolean        the row above the prompt (default true)
28 *   status: boolean      a short entry in the status line (default false)
29 */
30import type { EngineInterface, Register } from 'claude-code'
31
32const COMMAND = 'keepwarm'
33const STORE_KEY = 'enabled'
34const LAST_KEY = 'last'
35const TICK_MS = 5_000
36// a ping needs a few seconds to reach the API; closer to expiry it may miss
37const SAFETY_MS = 15_000
38const DEFAULT_MESSAGE =
39  'Service ping from the cache-keepwarm mod to keep the prompt cache warm. ' +
40  'Reply with the single word "ok" and do nothing else: no tools, no summary.'
41
42// read: tokens the cache served; tokens: the whole prompt (read + written + uncached)
43type Last = { at: number; tokens: number; read: number }
44type SavedLast = Last & { session: string }
45type Config = {
46  enabled: boolean
47  ttlMs: number
48  leadMs: number
49  maxPings: number
50  minTokens: number
51  message: string
52  band: boolean
53  status: boolean
54}
55
56let cfg: Config = {
57  enabled: false,
58  ttlMs: 60 * 60_000,
59  leadMs: 120_000,
60  maxPings: 3,
61  minTokens: 20_000,
62  message: DEFAULT_MESSAGE,
63  band: true,
64  status: false,
65}
66let last: Last | undefined
67let sessionId = ''
68let busy = false
69let pending = false
70let pings = 0
71let switchedOn: boolean | undefined
72let timer: { cancel: () => void } | undefined
73let drawnKey = ''
74
75function num(value: unknown, fallback: number): number {
76  return typeof value === 'number' && Number.isFinite(value) && value >= 0 ? value : fallback
77}
78
79function fmtTokens(n: number): string {
80  return n >= 1000 ? `${Math.round(n / 1000)}k` : String(n)
81}
82
83function fmtTime(ms: number): string {
84  const d = new Date(ms)
85  return `${String(d.getHours()).padStart(2, '0')}:${String(d.getMinutes()).padStart(2, '0')}`
86}
87
88// minutes while there are minutes, seconds at the end: the row redraws only when this text changes
89function fmtLeft(ms: number): string {
90  return ms >= 60_000 ? `${Math.floor(ms / 60_000)}m` : `${Math.ceil(ms / 1000)}s`
91}
92
93function isOn(): boolean {
94  return switchedOn ?? cfg.enabled
95}
96
97// short on purpose: the status line and the checkbox already name the plugin
98function describe(now: number): string {
99  if (!last) return 'on'
100  if (last.tokens < cfg.minTokens) return `skip <${fmtTokens(cfg.minTokens)}`
101  if (now >= last.at + cfg.ttlMs) return 'lapsed'
102  if (pings >= cfg.maxPings) return `${pings}/${cfg.maxPings} wait`
103  return `${fmtTime(last.at + cfg.ttlMs - cfg.leadMs)} · ${pings}/${cfg.maxPings}`
104}
105
106// the /keepwarm reply, read once, can say it in words
107function explain(now: number): string {
108  if (!isOn()) return 'off'
109  if (!last) return 'on, waiting for the first request'
110  if (last.tokens < cfg.minTokens) return `on, ${fmtTokens(last.tokens)} cached is below ${fmtTokens(cfg.minTokens)}: no ping`
111  if (now >= last.at + cfg.ttlMs) return 'on, the cache already lapsed: no ping'
112  if (pings >= cfg.maxPings) return `on, ${pings}/${cfg.maxPings} pings used, waiting for your prompt`
113  return `on, next ping at ${fmtTime(last.at + cfg.ttlMs - cfg.leadMs)}, ${pings}/${cfg.maxPings} used`
114}
115
116function cacheText(now: number): string {
117  if (!last) return 'cache: no request yet'
118  const left = last.at + cfg.ttlMs - now
119  const hit = last.tokens > 0 ? Math.round((last.read / last.tokens) * 100) : 0
120  return `cache ${hit}% · ${left > 0 ? fmtLeft(left) : 'lapsed'}`
121}
122
123function cacheColor(now: number): string | undefined {
124  if (!last) return undefined
125  const left = last.at + cfg.ttlMs - now
126  return left <= 0 ? 'red' : left <= cfg.leadMs ? 'yellow' : 'green'
127}
128
129function refresh($: EngineInterface) {
130  const now = Date.now()
131  if (cfg.status) $.ui.status(isOn() ? describe(now) : undefined)
132  const key = `${isOn()}|${cacheText(now)}|${describe(now)}`
133  if (key !== drawnKey) {
134    drawnKey = key
135    $.ui.invalidate('ui.render')
136  }
137}
138
139async function setOn($: EngineInterface, on: boolean) {
140  switchedOn = on
141  pings = 0
142  await $.store.set(STORE_KEY, on)
143  refresh($)
144}
145
146async function maybePing($: EngineInterface) {
147  if (!isOn() || busy || pending || !last) return
148  if (last.tokens < cfg.minTokens || pings >= cfg.maxPings) return
149  const now = Date.now()
150  const expires = last.at + cfg.ttlMs
151  if (now < expires - cfg.leadMs || now > expires - SAFETY_MS) return
152  pending = true
153  pings += 1
154  $.ui.toast(`keepwarm: ping ${pings}/${cfg.maxPings} keeps ${fmtTokens(last.tokens)} tokens cached`)
155  try {
156    await $.prompt.submit({ text: cfg.message })
157  } catch (err) {
158    pending = false
159    $.ui.log(`cache-keepwarm: ping not sent: ${err}`)
160  }
161  refresh($)
162}
163
164export const register: Register = (on, options) => {
165  const ttlMs = options.ttl === '5m' ? 5 * 60_000 : 60 * 60_000
166  cfg = {
167    enabled: options.enabled === true,
168    ttlMs,
169    leadMs: Math.min(num(options.leadSeconds, 120) * 1000, ttlMs / 2),
170    maxPings: Math.floor(num(options.maxPings, 3)),
171    minTokens: num(options.minTokens, 20_000),
172    message: typeof options.message === 'string' && options.message.trim() ? options.message : DEFAULT_MESSAGE,
173    band: options.band !== false,
174    status: options.status === true,
175  }
176
177  on('session.start', async ($, e, next) => {
178    const r = await next(e)
179    last = undefined
180    busy = false
181    pending = false
182    pings = 0
183    drawnKey = ''
184    const stored = await $.store.get(STORE_KEY).catch(() => undefined)
185    switchedOn = typeof stored === 'boolean' ? stored : undefined
186    // an app restart or a resume reloads the mod: the same session keeps its cache, so keep counting from it
187    sessionId = await $.session.id().catch(() => '')
188    const saved = (await $.store.get(LAST_KEY).catch(() => undefined)) as SavedLast | undefined
189    if (saved && sessionId && saved.session === sessionId && typeof saved.at === 'number' && typeof saved.tokens === 'number') {
190      last = { at: saved.at, tokens: saved.tokens, read: typeof saved.read === 'number' ? saved.read : 0 }
191    }
192    await $.command
193      .register({
194        name: COMMAND,
195        description: 'Keep the prompt cache warm while idle: on, off or status',
196        argumentHint: '[on|off|status]',
197        immediate: true,
198      })
199      .catch(err => $.ui.log(`cache-keepwarm: /${COMMAND} not registered: ${err}`))
200    timer?.cancel()
201    timer = $.clock.every(TICK_MS, () => {
202      refresh($)
203      void maybePing($)
204    })
205    refresh($)
206    return r
207  })
208
209  on('session.end', async ($, e, next) => {
210    last = undefined
211    pending = false
212    pings = 0
213    if (e.reason !== 'clear') {
214      timer?.cancel()
215      timer = undefined
216    }
217    return next(e)
218  })
219
220  // any prompt that reaches this hook is not ours: $.prompt.submit skips the calling hook
221  on('prompt.submit', async ($, e, next) => {
222    pings = 0
223    return next(e)
224  })
225
226  on('turn.start', async ($, e, next) => {
227    busy = true
228    return next(e)
229  })
230
231  on('turn.step', async function* ($, e, next) {
232    if (e.agentId) return yield* next(e)
233    const at = Date.now()
234    const r = yield* next(e)
235    const u = r.usage
236    if (u && u.cache_read_input_tokens + u.cache_creation_input_tokens > 0) {
237      last = {
238        at,
239        tokens: u.cache_read_input_tokens + u.cache_creation_input_tokens + u.input_tokens,
240        read: u.cache_read_input_tokens,
241      }
242      pending = false
243      if (sessionId) void $.store.set(LAST_KEY, { session: sessionId, ...last }).catch(() => undefined)
244      refresh($)
245    }
246    return r
247  })
248
249  on('turn.complete', async ($, e, next) => {
250    const r = await next(e)
251    if (!e.agentId) {
252      busy = false
253      pending = false
254      refresh($)
255    }
256    return r
257  })
258
259  on('command.run', { command: COMMAND }, async ($, e) => {
260    const arg = e.args.trim().toLowerCase()
261    if (arg === 'on' || arg === 'off') await setOn($, arg === 'on')
262    return { text: explain(Date.now()) }
263  })
264
265  // one row: cache state on the left, the keepwarm checkbox on the right
266  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
267    if (!cfg.band || e.props.hasSurvey) return next(e)
268    const { Box, Text, Button } = $.ui.resolve(e)
269    const now = Date.now()
270    const label = isOn() ? `[x] keepwarm ${describe(now)}` : '[ ] keepwarm'
271    return (
272      <Box flexDirection="row" columnGap={1}>
273        <Text color={cacheColor(now)}>●</Text>
274        <Text dimColor>{cacheText(now)}</Text>
275        <Button key="toggle" plain label={label} dimColor={!isOn()} onPress={() => setOn($, !isOn())} />
276      </Box>
277    )
278  })
279}
280