Cache Keeper by Nate Herk, cut down to token economics: cache stats, a dollar-based cold-send guard, keep warm with a break-even, and handoff

A Claude Code mod for token economics, for subscription and API users alike. It is a fork of Cache Keeper by Nate Herk, from nateherkai/claude-code-mods (MIT), cut down to four jobs: show what the cache costs, set it up, keep it warm, and hand a big chat off to a fresh one.
All of Cache Keeper's original work is Nate Herk's: the warm/cold cache band, /keepwarm, the cold-send guard, TTL detection, /handoff and the session-handoff skill, plus the shared pricing and fmt helpers. Go check out his repo and videos. This fork changes and adds the things listed below. It leaves out Cache Keeper's /board and recording mode for now.
/cache./cache big ($1 by default). Token Keeper works out what that means in tokens for the current model and TTL, for example $1 ≈ 125k tokens on Opus 5.5 with the 1h TTL. Above it, the cold-send guard asks before a cold send, a line in the chat warns before the cache goes cold, and the band shows the keep-warm button./keepwarm runs for 30 minutes on the 5-minute TTL and for 4 hours on the 1-hour TTL by default. A rewrite costs as much as (cache-write price / cache-read price) pings, whatever the context size, so pinging pays off only up to a point, for example about 1h 28m on Opus 5.5 with the 5-minute TTL. Ask for longer and you get a warning that suggests /handoff or /clear instead./cache ttl sets the real TTL, and checks it. /cache ttl 5 or 60 sets CLAUDE_CODE_PROMPT_CACHE_TTL for the running Claude Code, so the next request already writes 5-minute or 1-hour cache entries; auto gives the variable back. Every cache write reports which TTL it used, so the cache window is measured rather than guessed, and a line warns once if Claude Code writes another TTL than the one you set (FORCE_PROMPT_CACHING_5M wins over it). Switching while the cache is warm makes the next message rewrite the part of the context cached under the other TTL, so /cache ttl asks first and names the price; once the cache is cold it switches without asking, because that rewrite happens anyway./keepwarm compares the two for the time you asked for and says when the handoff is cheaper, for example at 300k tokens on Opus 5.5 with the 5-minute TTL after about 14 minutes. It also says what each later turn saves, which is usually worth more: at 300k tokens each turn re-reads $0.06 of context, after a handoff about $0.005. The fresh chat's start is measured from /context, and the handoff's size is the average of your last five. Past /cache handoff tokens (140k by default, a common rule of thumb for where a model starts to lose track of a long context) a notice suggests a handoff or a /compact, which costs about the same, and asks whether to hand off now, be reminded after another 100k tokens, or get no more reminders. Not counted: files the fresh chat reads again because the handoff left them out, so a handoff pays off most when much of the old context is done with./handoff before the cache goes cold.≥ instead of being left out without a word.Token Keeper runs wherever Claude Code runs, but each app draws a different part of it. Claude Code decides which parts each app gets, not the mod.

The band above the prompt shows the cache state (warm, cooling, cold, kept warm), the context size, what a rewrite costs, rate limits, cold restarts and cache breaks. While a big cache is cooling, it shows a keep warm button you press with 1.

The Desktop app draws the band too, with its buttons as native buttons. The footer also gets a short label, but only when there is something to act on.

The VS Code extension does not draw the band, the footer label or the hint line under the prompt; Claude Code gives mods no place there. What still works:
/cache, /keepwarm and /handoff, with their full text output./cache ttl while the cache is warm, and the handoff reminder.What you miss in VS Code:
/cache: it shows everything the band does, including the context against the window, the effort and what a rewrite costs.1. Type /keepwarm instead./handoff, type /handoff continue to clear the chat and continue with the handoff.If you want the band, run claude in VS Code's integrated terminal instead of the extension's panel.
Needs Claude Code 2.1.287 or later, with mods turned on for your account.
claude plugin marketplace add LamAnhNguyenHTW/token-keeper
claude plugin install token-keeper@token-keeper
Don't install it alongside Cache Keeper: both register /cache, /keepwarm and /handoff. If a name is already taken, Token Keeper falls back to /tk-cache and so on.
| Command | What it does | ||
|---|---|---|---|
/cache | Status and settings | ||
| `/cache ttl 5\ | 60\ | auto` | Sets Claude Code's cache TTL for the main conversation (CLAUDE_CODE_PROMPT_CACHE_TTL, needs Claude Code v2.1.242+). auto leaves it to Claude Code. Asks first while the cache is warm, since the switch rewrites the context once |
| `/cache guard on\ | off` | Ask before a cold send of a big context | |
| `/cache big $1\ | default` | When a context counts as big: what a rewrite costs. $1 by default | |
| `/cache handoff 140k\ | off\ | default` | Suggest a handoff once the context passes this size, then every 100k more. 140k by default |
| `/cache alerts on\ | off` | Warnings in the chat | |
| `/keepwarm [30m\ | 4h\ | off]` | Keep the cache warm. 30m by default on the 5-minute TTL, 4h on the 1-hour TTL |
/handoff [continue] | Hand the chat off to a fresh one |
Everything stays local, under ~/.claude/mods-data/token-keeper/. The mod makes no network calls of its own. The only extra requests are the keep-warm pings, and those run only while /keepwarm is on.
MIT. See LICENSE. The original copyright belongs to Nate Herk.
hooks/register.mjs 1083 lines1// Token Keeper (based on Cache Keeper by Nate Herk): watches this session's
2// prompt cache, keeps a big cache warm on request, asks before a cold send,
3// and hands a big chat off to a fresh one.
4//
5// Why: each request re-reads the whole context. From the cache that costs about
6// a tenth of normal input. Once the cache expires (1 hour idle on a subscription,
7// 5 minutes on the default API TTL), the next message writes the whole context
8// again at 1.25x to 2x input. A cache read also restarts the timer, so a tiny
9// ping before expiry costs a fraction of a rewrite.
10
11import { rewriteCost, requestCost, totalInput, cachedShare, priceFor, writeRate } from './pricing.mjs'
12import { tokens, usd, minutes, clock, clip, basename } from './fmt.mjs'
13
14const MIN = 60000
15const TICK_EVERY = 30000
16const TODAY_EVERY = 10 * MIN
17const MAX_READ = 4 * 1024 * 1024 // $.fs.read refuses larger files
18// The default cold-send threshold: below a rewrite of about $1 a dialog
19// bothers more than the rewrite costs (about 125k tokens on Opus 5.5, 1h TTL)
20const BIG_USD = 1
21// What a handoff costs besides one read of the context: its output, and a fresh
22// chat's start (system prompt, tools, memory files). Both are measured; these
23// stand in until they are
24const HANDOFF_OUT = 3000
25const FRESH_START = 20000
26// The default context size from which a handoff is suggested (a rule of thumb:
27// past about 140k a model loses track more often), and how much more the
28// context grows before the next reminder
29const HANDOFF_AT = 140000
30const HANDOFF_STEP = 100000
31
32// This session
33const S = {
34 id: '',
35 cwd: '',
36 model: '',
37 lastActivity: 0, // last main-loop request or keep-warm ping that touched the cache
38 ctx: 0,
39 lastTotal: 0, // the last main-loop request's prompt size, what its cache entry holds
40 window: 0,
41 costUsd: 0,
42 ttlMin: 60,
43 ttlSource: 'default',
44 ttlWritten: 0, // the TTL of the last main-loop cache write, in minutes (0: none since /cache ttl)
45 ttlWarned: false, // said once that Claude Code does not write the TTL /cache ttl set
46 coldRestarts: [],
47 working: false,
48 turnId: '', // the main loop's running (or last) turn
49 keepWarm: false,
50 keepWarmUntil: 0,
51 pings: 0,
52 pingUsd: 0,
53 pingsSinceTurn: 0, // keep-warm pings since the last main-loop request
54 rateLimits: [],
55 cacheBreaks: [], // rewrites while the cache was still warm
56 effort: '',
57 todayUsd: null, // every session's spend since local midnight
58 todayPartial: false, // a transcript was too big to read, so todayUsd is a lower bound
59 startTokens: 0, // what a fresh chat starts with, as /context counts it (0: not measured)
60 handoffOuts: [], // output tokens of the last few handoffs, kept across sessions
61 handoffNextAt: 0, // the size of the next handoff reminder (0: /cache handoff)
62}
63
64const settings = { bigUsd: BIG_USD, guard: true, ttlMin: 0, alerts: true, handoffAt: HANDOFF_AT }
65const names = { cache: 'cache', keepwarm: 'keepwarm', handoff: 'handoff' }
66
67let now = 0
68let home = ''
69let warnedFor = 0 // the lastActivity the cooling warning was shown for
70let justCompacted = false
71let ttlSwitched = false // /cache ttl changed while the cache was warm: the next request rewrites
72let handoffPending = false
73let envPrior // CLAUDE_CODE_PROMPT_CACHE_TTL before /cache ttl set it
74
75// The handoff flow. When /session-handoff runs (the band's button, /handoff, a
76// typed command, or Claude calling the skill), the answer of the turn it starts
77// is captured and saved to a file. "clear and continue" then runs /clear and
78// sends the handoff as the fresh chat's first prompt.
79const HANDOFF_SKILL = /(^|:)session-handoff$/
80const H = {
81 armed: false,
82 outTokens: 0, // output of the handoff turn so far
83 notTurn: '', // a turn already running when the handoff was queued: not the handoff's
84 askAfter: false, // /handoff asks "clear and continue?" once the handoff is in
85 text: '',
86 path: '',
87 continuing: false,
88}
89
90function armHandoff(insideTurn) {
91 H.armed = true
92 H.outTokens = 0
93 H.notTurn = insideTurn ? '' : S.working ? S.turnId : ''
94}
95
96// The handoff without any preamble Claude put before its heading
97function handoffBody(answer) {
98 const text = String(answer || '').trim()
99 const i = text.indexOf('# Session Handoff')
100 return i > 0 ? text.slice(i) : text
101}
102
103function fileStamp(ms) {
104 const d = new Date(ms)
105 const two = (n) => String(n).padStart(2, '0')
106 return `${d.getFullYear()}-${two(d.getMonth() + 1)}-${two(d.getDate())}-${two(d.getHours())}${two(d.getMinutes())}`
107}
108
109function dataDir() {
110 return `${home || '.'}/.claude/mods-data/token-keeper`.replace(/\\/g, '/')
111}
112
113async function saveHandoff($, text) {
114 const project = basename(S.cwd).replace(/[^A-Za-z0-9-]+/g, '-') || 'session'
115 const path = `${dataDir()}/handoffs/${fileStamp(now)}-${project}-${S.id.slice(0, 8)}.md`
116 await $.fs.write(path, text + '\n')
117 return path
118}
119
120async function captureHandoff($, e) {
121 H.armed = false
122 H.notTurn = ''
123 const askAfter = H.askAfter
124 H.askAfter = false
125 const body = e.reason === 'answer' ? handoffBody(e.answer) : ''
126 if (body.length < 200) {
127 warn($, 'The handoff turn ended without a handoff, so nothing was saved.')
128 return
129 }
130 H.text = body
131 if (H.outTokens > 0) {
132 S.handoffOuts = [...S.handoffOuts, H.outTokens].slice(-5)
133 await $.store.set('handoffOuts', S.handoffOuts).catch(() => {})
134 }
135 try {
136 H.path = await saveHandoff($, body)
137 } catch {
138 H.path = ''
139 }
140 const saved = H.path ? `Handoff saved to ${H.path}.` : 'Handoff captured (the backup file could not be written).'
141 note($, `${saved} Press c on the band or type /${names.handoff} continue to clear and continue.`)
142 // Off the hook: the turn is ending, and a dialog would hold it open
143 if (askAfter) $.clock.after(300, () => offerContinue($).catch(() => {}))
144}
145
146async function offerContinue($) {
147 if (!H.text) return
148 let answer = ''
149 try {
150 answer = await $.ui.ask('Handoff saved. Clear this chat and continue with it in a fresh context?', ['Clear and continue', 'Keep this chat'])
151 } catch {
152 return // dismissed: the band button and /handoff continue still work
153 }
154 if (answer === 'Clear and continue') await clearAndContinue($)
155}
156
157function continuationPrompt() {
158 const where = H.path ? ` (saved at ${H.path})` : ''
159 return `Handoff from my previous session${where}:\n\n${H.text}`
160}
161
162async function clearAndContinue($) {
163 if (H.continuing) return
164 if (!H.text) {
165 note($, `No handoff ready. Press h on the band or type /${names.handoff} first.`)
166 return
167 }
168 if (S.working) {
169 note($, 'Claude is still working. Clear and continue once this turn ends.')
170 return
171 }
172 const text = continuationPrompt()
173 const path = H.path
174 H.continuing = true
175 $.ui.invalidate('ui.render')
176 try {
177 await $.command.run({ command: 'clear', args: '' })
178 } catch (err) {
179 H.continuing = false
180 $.ui.invalidate('ui.render')
181 warn($, 'Could not run /clear: ' + clip(String((err && err.message) || err), 100) + (path ? ` The handoff is saved at ${path}.` : ''))
182 return
183 }
184 H.text = ''
185 H.path = ''
186 try {
187 await $.prompt.submit({ text, asUser: true })
188 } catch {
189 // Not sent: leave it in the prompt box for one Enter
190 let filled = false
191 try {
192 filled = !!(await $.prompt.fill({ text })).isFilled
193 } catch {
194 filled = false
195 }
196 if (filled) note($, 'Cleared. The handoff is in the prompt box: press Enter to send it.')
197 else warn($, `Cleared, but the handoff could not be sent.${path ? ` It's saved at ${path}.` : ''}`)
198 } finally {
199 H.continuing = false
200 $.ui.invalidate('ui.render')
201 }
202}
203
204// Runs the /session-handoff skill as if you typed it. The engine queues it
205// until Claude finishes the current turn.
206async function runHandoff($) {
207 if (handoffPending) {
208 note($, 'Session handoff is already queued.')
209 return
210 }
211 try {
212 const commands = await $.command.list()
213 const cmd = commands.find((c) => c.name === 'session-handoff') || commands.find((c) => HANDOFF_SKILL.test(c.name))
214 if (!cmd) {
215 H.askAfter = false
216 warn($, 'No /session-handoff skill in this session.')
217 return
218 }
219 handoffPending = true
220 armHandoff(false)
221 $.ui.invalidate('ui.render')
222 note($, S.working ? 'Session handoff queued: it runs when Claude finishes this turn.' : 'Running /session-handoff.')
223 await $.command.run({ command: cmd.name, args: '' })
224 } catch (err) {
225 H.armed = false
226 H.askAfter = false
227 warn($, 'Could not start /session-handoff: ' + clip(String((err && err.message) || err), 100))
228 } finally {
229 handoffPending = false
230 $.ui.invalidate('ui.render')
231 }
232}
233
234// What the cache really uses: the last write's TTL, else /cache ttl, else the estimate
235function ttlMin() {
236 return S.ttlWritten || settings.ttlMin || S.ttlMin
237}
238
239// The session's spend as Claude Code counts it, or null when it does not say
240async function sessionUsd($) {
241 try {
242 const cost = (await $.session.usage()).cost
243 return cost ? cost.usd : null
244 } catch {
245 return null
246 }
247}
248
249// The TTL a cache write used: the usage's 1h/5m split when it has one; else the
250// price that matches what Claude Code charged for the request, since its cost
251// counts the TTL really written (0: too small to tell, or another request in between)
252function writeTtl(u, written, charged) {
253 if (u.cache_creation && written > 0) return (u.cache_creation.ephemeral_1h_input_tokens || 0) * 2 >= written ? 60 : 5
254 if (charged === null || written < 1000) return 0
255 const p5 = requestCost(u, u.model || S.model, 5)
256 const p60 = requestCost(u, u.model || S.model, 60)
257 const near = Math.abs(charged - p60) < Math.abs(charged - p5) ? 60 : 5
258 return Math.abs(charged - (near === 60 ? p60 : p5)) < Math.abs(p60 - p5) / 4 ? near : 0
259}
260
261function ttlEnv(min) {
262 return min === 5 ? '5m' : min === 60 ? '1h' : undefined
263}
264
265// Points Claude Code's own cache TTL at the /cache ttl choice; auto gives the
266// variable back as it was before Token Keeper set it
267async function applyTtl($) {
268 // a refusal must not break /cache: the next write's TTL check reports it
269 await $.env.set('CLAUDE_CODE_PROMPT_CACHE_TTL', settings.ttlMin ? ttlEnv(settings.ttlMin) : envPrior).catch(() => {})
270}
271
272function ttlSourceText() {
273 if (!settings.ttlMin) return S.ttlSource
274 if (!S.ttlWritten) return 'set by you'
275 if (S.ttlWritten === settings.ttlMin) return 'set by you · confirmed'
276 return `⚠️ you set ${settings.ttlMin} min, but Claude Code writes ${S.ttlWritten} min`
277}
278
279function ttlName(ttl) {
280 return ttl >= 60 ? '1h TTL' : '5-min TTL'
281}
282
283// Plan limit windows as the API reports them: five_hour, seven_day, spend_limit
284const LIMIT_LABEL = { five_hour: '5h', seven_day: 'week', spend_limit: 'spend' }
285
286function limitsText(limits) {
287 return limits.map((l) => `${LIMIT_LABEL[l.kind] || l.kind} ${Math.round(l.percentUsed)}%`).join(' · ')
288}
289
290function limitTone(limits, extra) {
291 const top = Math.max(...limits.map((l) => l.percentUsed || 0))
292 if (top >= 95) return { ...extra, color: 'red', bold: true }
293 if (top >= 80) return { ...extra, color: 'yellow' }
294 return { ...extra, dimColor: true }
295}
296
297// 'claude-opus-5-5' -> 'Opus 5.5'
298function modelName(model) {
299 const [family, ...ver] = priceFor(model).id.split('-')
300 return family[0].toUpperCase() + family.slice(1) + ' ' + ver.join('.')
301}
302
303function msLeft(lastActivity, ttl) {
304 if (!lastActivity) return null
305 return ttl * MIN - (now - lastActivity)
306}
307
308function cacheState() {
309 const left = msLeft(S.lastActivity, ttlMin())
310 if (left === null) return { kind: 'unknown', left: 0 }
311 if (S.keepWarm) return { kind: 'kept', left }
312 if (left <= 0) return { kind: 'cold', left }
313 // Cooling: from when a keep-warm ping would go out (5 minutes before on the 1h
314 // TTL, 90s on the 5-minute one), so a 5-minute cache is not cooling right away
315 const ttl = ttlMin()
316 if (left <= Math.min(5 * MIN, ttl * MIN - pingEvery(ttl))) return { kind: 'cooling', left }
317 return { kind: 'warm', left }
318}
319
320// Big: rewriting this context costs at least the threshold, on this model and TTL
321function isBig() {
322 return rewriteCost(S.ctx, S.model, ttlMin()) >= settings.bigUsd
323}
324
325// "$1.00 (default) ≈ 125k tokens on Opus 5.5, 1h TTL"
326function bigText() {
327 const ttl = ttlMin()
328 const atTokens = settings.bigUsd / (writeRate(S.model, ttl) / 1e6)
329 return `${usd(settings.bigUsd)}${settings.bigUsd === BIG_USD ? ' (default)' : ''} ≈ ${tokens(atTokens)} tokens on ${modelName(S.model)}, ${ttlName(ttl)}`
330}
331
332// Keep warm pings this long after the last read, and every ping after it
333function pingEvery(ttl) {
334 return ttl * MIN - (ttl >= 60 ? 8 * MIN : 90000)
335}
336
337// How long pinging stays cheaper than one rewrite: a rewrite costs as much as
338// (write price / read price) pings, whatever the context size
339function breakEvenMs() {
340 const ttl = ttlMin()
341 return (writeRate(S.model, ttl) / priceFor(S.model).read) * pingEvery(ttl)
342}
343
344function defaultKeepWarmMs() {
345 return ttlMin() >= 60 ? 4 * 60 * MIN : 30 * MIN
346}
347
348// "30m", "30min", "2h", "1.5" (hours); capped at 24 hours
349function parseDuration(text) {
350 const m = String(text || '').match(/^(\d+(?:\.\d+)?)\s*(m|min|h)?$/)
351 if (!m || !(Number(m[1]) > 0)) return null
352 const ms = Number(m[1]) * (m[2] === 'm' || m[2] === 'min' ? MIN : 60 * MIN)
353 return Math.min(24 * 60 * MIN, ms)
354}
355
356// "300k", "1.5M", "300000"
357function parseTokens(text) {
358 const m = String(text || '').match(/^(\d+(?:\.\d+)?)\s*(k|m)?$/)
359 if (!m || !(Number(m[1]) > 0)) return null
360 return Math.round(Number(m[1]) * (m[2] === 'k' ? 1e3 : m[2] === 'm' ? 1e6 : 1))
361}
362
363// What a fresh chat starts with: every /context row but the messages. A local
364// estimate, no request
365async function measureStart($) {
366 try {
367 const u = await $.session.usage({ breakdown: 'summary' })
368 const rows = (u.context && u.context.breakdown && u.context.breakdown.categories) || []
369 const sum = rows.filter((r) => r.kind === 'used' && r.name !== 'Messages').reduce((a, r) => a + (r.tokens || 0), 0)
370 if (sum > 0) S.startTokens = sum
371 } catch {
372 // no breakdown: keep the last figure or the default
373 }
374}
375
376// When the context passes the handoff size: a notice, then a choice to hand
377// off now, be reminded HANDOFF_STEP later, or turn the reminders off
378function handoffStep($) {
379 if (!settings.alerts || !settings.handoffAt) return
380 const at = S.handoffNextAt || settings.handoffAt
381 if (S.ctx < at) return
382 let next = at + HANDOFF_STEP
383 while (next <= S.ctx) next += HANDOFF_STEP
384 S.handoffNextAt = next
385 const c = keepWarmVsHandoff(0)
386 // Quality is the reason; the saving is named only when the fresh chat is clearly smaller
387 const saves = c.perTurn >= 0.01 ? `, and every turn after it re-reads about ${tokens(c.fresh)} instead of ${tokens(S.ctx)}: ${usd(c.perTurn)} less per turn` : ''
388 note($, `This chat is at ${tokens(S.ctx)} tokens, and a long context gets less reliable. A /${names.handoff} or /compact (about ${usd(c.handoff)}, either) carries on in a fresh chat${saves}.`)
389 // Off the hook: the turn is ending, and a dialog would hold it open
390 $.clock.after(300, () => offerHandoff($, next).catch(() => {}))
391}
392
393async function offerHandoff($, next) {
394 const remind = `Remind me after another ${tokens(HANDOFF_STEP)} (${tokens(next)})`
395 let answer = ''
396 try {
397 answer = await $.ui.ask(`This chat is at ${tokens(S.ctx)} tokens. Hand it off to a fresh chat?`, ['Handoff now', remind, 'No more reminders'])
398 } catch {
399 return // dismissed: remind at the next step
400 }
401 if (answer === 'Handoff now') {
402 H.askAfter = true
403 await runHandoff($)
404 } else if (answer === 'No more reminders') {
405 settings.handoffAt = 0
406 await $.store.set('settings', settings)
407 note($, `Handoff reminders off. /${names.cache} handoff default turns them back on.`)
408 }
409}
410
411function parseUsd(text) {
412 const m = String(text || '').match(/^\$?(\d+(?:\.\d+)?)$/)
413 return m && Number(m[1]) > 0 ? Number(m[1]) : null
414}
415
416// Lines in the chat (plain text, drawn dim): ⚠️ a warning, ℹ️ a notice
417function warn($, text) {
418 $.ui.log('⚠️ ' + text)
419}
420
421function note($, text) {
422 $.ui.log('ℹ️ ' + text)
423}
424
425// Bold in a command's output, which draws markdown; plain in a chat line
426function bolder(md) {
427 return md ? (s) => `**${s}**` : (s) => s
428}
429
430// Keep warm for ms vs a handoff now, in dollars. A handoff reads the context
431// once, writes the handoff, and the fresh chat writes its start; then no pings
432// The handoff's output: the average of the last few, or the estimate
433function handoffOut() {
434 return S.handoffOuts.length ? S.handoffOuts.reduce((a, n) => a + n, 0) / S.handoffOuts.length : HANDOFF_OUT
435}
436
437// Keep warm for ms vs a handoff now, in dollars, and what each turn after the
438// handoff saves on re-reading the context (a /compact costs about the same)
439function keepWarmVsHandoff(ms) {
440 const ttl = ttlMin()
441 const p = priceFor(S.model)
442 const ping = S.ctx * p.read / 1e6
443 const start = S.startTokens || FRESH_START
444 const fresh = start + handoffOut()
445 return {
446 ping,
447 fresh,
448 keepWarm: Math.floor(ms / pingEvery(ttl)) * ping, // a ping every pingEvery, none at the start
449 handoff: ping + handoffOut() * p.output / 1e6 + start * writeRate(S.model, ttl) / 1e6,
450 perTurn: Math.max(0, S.ctx - fresh) * p.read / 1e6,
451 }
452}
453
454// Starts keeping the cache warm and says so, with ⚠️ past the break-even
455function startKeepWarm(ms, isDefault, md) {
456 // A new run counts its own pings; extending a running one keeps counting
457 if (!S.keepWarm) {
458 S.pings = 0
459 S.pingUsd = 0
460 }
461 S.keepWarm = true
462 S.keepWarmUntil = now + ms
463 const b = bolder(md)
464 const ttl = ttlMin()
465 const be = breakEvenMs()
466 const how = `${b(minutes(ms))}${isDefault ? `, the default on the ${ttlName(ttl)}` : ''}, until ${clock(S.keepWarmUntil)}`
467 const on = `On ${modelName(S.model)} with the ${ttlName(ttl)}`
468 if (ms > be) return `⚠️ Keeping this cache warm for ${how}. ${on} that only pays off up to ${b('~' + minutes(be))}: after that, one rewrite (${usd(rewriteCost(S.ctx, S.model, ttl))}) is cheaper. Better: ${b(`/${names.keepwarm} ${ttl >= 60 ? '4h' : '30m'}`)} (the default), ${b('/' + names.handoff)} now while the cache is warm, or ${b('/clear')}.`
469 const c = keepWarmVsHandoff(ms)
470 const vs = c.handoff < c.keepWarm ? ` A ${b('/' + names.handoff)} now is cheaper: about ${usd(c.handoff)} vs ${usd(c.keepWarm)} for ${minutes(ms)} of pings. After it, each turn re-reads about ${tokens(c.fresh)} instead of ${tokens(S.ctx)}, about ${usd(c.perTurn)} less.` : ''
471 return `ℹ️ Keeping this cache warm for ${how}. ${on} it pays off up to ~${minutes(be)}.${vs}`
472}
473
474// Stops keeping the cache warm and says so ('' when it was off), ⚠️ for a problem
475function stopKeepWarm(why, problem) {
476 if (!S.keepWarm) return ''
477 S.keepWarm = false
478 return `${problem ? '⚠️' : 'ℹ️'} Keep warm off (${why}). ${S.pings} ping(s), ${usd(S.pingUsd)}.`
479}
480
481function log($, text) {
482 if (text) $.ui.log(text)
483}
484
485async function keepWarmStep($) {
486 if (!S.keepWarm) return
487 if (now >= S.keepWarmUntil) return log($, stopKeepWarm('time limit reached'))
488 if (S.working || !S.lastActivity) return
489 const ttl = ttlMin()
490 if (now - S.lastActivity < pingEvery(ttl)) return
491 // Past the TTL the cache is gone (keep warm started late, or the machine slept):
492 // a ping would only write it again
493 if (now - S.lastActivity >= ttl * MIN) return log($, stopKeepWarm(`the cache already went cold ${minutes(now - S.lastActivity - ttl * MIN)} ago, so a ping would only write it again`, true))
494 let reply = null
495 try {
496 reply = await $.model.fork({ prompt: 'token-keeper keep-alive ping. Reply with only: ok' })
497 } catch {
498 reply = null
499 }
500 if (!reply || !reply.usage) {
501 return log($, stopKeepWarm('the ping got no answer, so the cache may already be cold', true))
502 }
503 const u = reply.usage
504 const cost = requestCost(u, S.model, ttl)
505 S.pings += 1
506 S.pingUsd += cost
507 addToday(cost)
508 // The first ping after a turn may write the turn's tail (its prefix ends after the
509 // reply, the main thread's entry before it); a later ping that writes much missed
510 // the cache: stop paying for it
511 const first = S.pingsSinceTurn === 0
512 S.pingsSinceTurn += 1
513 if (!first && (u.cache_creation_input_tokens || 0) > 0.1 * Math.max(1, u.cache_read_input_tokens || 0)) {
514 return log($, stopKeepWarm(`the ping wrote ${tokens(u.cache_creation_input_tokens)} tokens instead of reading the cache`, true))
515 }
516 S.lastActivity = now
517}
518
519// Shortly before a big cache goes cold: keep it warm, or hand off while it's warm
520function warnStep($) {
521 if (!settings.alerts || S.keepWarm || !isBig()) return
522 const st = cacheState()
523 if (st.kind !== 'cooling' || warnedFor === S.lastActivity) return
524 warnedFor = S.lastActivity
525 const cost = usd(rewriteCost(S.ctx, S.model, ttlMin()))
526 warn($, `This ${tokens(S.ctx)}-token cache goes cold in ${minutes(st.left)}. Rewriting it then costs about ${cost}. Back soon: /${names.keepwarm} (or press 1). Done for now: /${names.handoff} while it's still warm.`)
527}
528
529async function tick($) {
530 now = await $.clock.now()
531 await keepWarmStep($)
532 warnStep($)
533 if (now - todayAt > TODAY_EVERY) await refreshToday($).catch(() => {})
534 $.ui.invalidate('ui.render')
535}
536
537// Today's spend at API list prices, every session's (subagents too). A cache
538// file per day keeps each transcript's costs, so only transcripts that changed
539// are read again. This session's own transcripts are read once, then its
540// requests are counted live. Each response counts once across files.
541let todayAt = 0
542let todayDate = ''
543let todayFiles = {} // path -> { mtimeMs, costs: { responseKey: usd } } or { mtimeMs, tooBig: true }
544let othersUsd = 0
545let othersPartial = false
546const own = { scanned: false, base: 0, live: 0, partial: false }
547
548function localDate(ms) {
549 const d = new Date(ms)
550 return `${d.getFullYear()}-${d.getMonth() + 1}-${d.getDate()}`
551}
552
553function setToday() {
554 S.todayUsd = othersUsd + own.base + own.live
555 S.todayPartial = othersPartial || own.partial
556}
557
558// A request of this session: counted live once its transcripts were read
559function addToday(cost) {
560 if (!own.scanned) return // the first read finds it in the transcript
561 own.live += cost
562 setToday()
563}
564
565async function readCosts($, path, mtimeMs, size, date) {
566 if (size > MAX_READ) return { mtimeMs, tooBig: true }
567 const costs = {}
568 const text = await $.fs.read(path).catch(() => null)
569 if (text === null) return { mtimeMs, tooBig: true }
570 for (const line of text.split('\n')) {
571 if (!line.includes('"usage"')) continue
572 let entry
573 try { entry = JSON.parse(line) } catch { continue }
574 const msg = entry.message
575 if (!msg || !msg.usage || !entry.timestamp || msg.model === '<synthetic>') continue
576 if (localDate(Date.parse(entry.timestamp)) !== date) continue
577 const cc = msg.usage.cache_creation || {}
578 const w1h = cc.ephemeral_1h_input_tokens || 0
579 const w5m = cc.ephemeral_5m_input_tokens ?? Math.max(0, (msg.usage.cache_creation_input_tokens || 0) - w1h)
580 costs[`${msg.id}:${entry.requestId}`] = requestCost({ ...msg.usage, cache_creation_input_tokens: w1h }, msg.model, 60) + requestCost({ cache_creation_input_tokens: w5m }, msg.model, 5)
581 }
582 return { mtimeMs, costs }
583}
584
585async function refreshToday($) {
586 todayAt = now
587 const date = localDate(now)
588 const cachePath = `${dataDir()}/today/${date}.json`
589 if (date !== todayDate) {
590 todayDate = date
591 todayFiles = {}
592 try {
593 todayFiles = JSON.parse(await $.fs.read(cachePath)).files || {}
594 } catch {
595 // no cache file yet today
596 }
597 own.scanned = false
598 }
599 const midnight = new Date(now).setHours(0, 0, 0, 0)
600 const root = ((await $.env.get('CLAUDE_CONFIG_DIR')) || home + '/.claude') + '/projects'
601 const files = []
602 const walk = async (dir) => {
603 for (const f of await $.fs.list(dir).catch(() => [])) {
604 const p = dir + '/' + f.name
605 if (f.kind === 'dir') await walk(p)
606 else if (f.name.endsWith('.jsonl') && f.mtimeMs >= midnight) files.push({ p, mtimeMs: f.mtimeMs, size: f.size })
607 }
608 }
609 await walk(root)
610 const scanOwn = !own.scanned
611 if (scanOwn) Object.assign(own, { base: 0, live: 0, partial: false })
612 const seen = new Map()
613 let partial = false
614 let changed = false
615 for (const { p, mtimeMs, size } of files) {
616 const mine = !!S.id && p.includes(S.id) // the transcript and its subagents' folder
617 if (mine && !scanOwn) continue
618 let entry = todayFiles[p]
619 if (!entry || entry.mtimeMs !== mtimeMs) {
620 entry = await readCosts($, p, mtimeMs, size, date)
621 todayFiles[p] = entry
622 changed = true
623 }
624 if (entry.tooBig) {
625 if (mine) own.partial = true
626 else partial = true
627 } else if (mine) {
628 own.base += Object.values(entry.costs).reduce((a, v) => a + v, 0)
629 } else {
630 for (const [k, v] of Object.entries(entry.costs)) seen.set(k, v)
631 }
632 }
633 own.scanned = true
634 othersUsd = [...seen.values()].reduce((a, v) => a + v, 0)
635 othersPartial = partial
636 setToday()
637 if (changed) await $.fs.write(cachePath, JSON.stringify({ files: todayFiles })).catch(() => {})
638}
639
640function todayText() {
641 return (S.todayPartial ? '≥ ' : '') + usd(S.todayUsd)
642}
643
644async function loadSettings($) {
645 const saved = await $.store.get('settings')
646 if (!saved || typeof saved !== 'object') return
647 for (const k of Object.keys(settings)) if (saved[k] !== undefined) settings[k] = saved[k]
648}
649
650function statusText() {
651 const st = cacheState()
652 const ttl = ttlMin()
653 const rewrite = rewriteCost(S.ctx, S.model, ttl)
654 const c = names.cache
655 // Markdown, as a command's output draws: a bold label per line, the band's icons.
656 // Everything the band shows is here too: VS Code draws no band, only this.
657 const lines = []
658 lines.push(`**Token Keeper** · cache window **${ttl} min** (${ttlSourceText()})`)
659 lines.push('')
660 const pct = S.window ? Math.floor((S.ctx / S.window) * 100) : null
661 const full = pct >= 80 ? '⚠️ ' : ''
662 lines.push(`- **Context:** ${full}**${tokens(S.ctx)}**${pct === null ? '' : ` / ${tokens(S.window)} (${pct}%)`} tokens on **${modelName(S.model)}**${S.effort ? ` · ${S.effort}` : ''}`)
663 if (st.kind === 'unknown') lines.push('- – **Cache:** no request yet this session, so its state is unknown')
664 if (st.kind === 'warm') lines.push(`- ● **Cache:** warm for about **${minutes(st.left)}** more · a rewrite would cost about ${usd(rewrite)}`)
665 if (st.kind === 'kept') lines.push(`- ● **Cache:** kept warm (see below) · a rewrite would cost about ${usd(rewrite)}`)
666 if (st.kind === 'cooling') lines.push(`- ◐ **Cache:** ⚠️ cools in **${minutes(st.left)}**, then the next message rewrites it for about **${usd(rewrite)}**`)
667 if (st.kind === 'cold') lines.push(`- ○ **Cache:** ⚠️ cold for ${minutes(-st.left)}, the next message rewrites it for about **${usd(rewrite)}**`)
668 if (H.text) lines.push(`- ✅ **Handoff ready:** \`/${names.handoff} continue\` clears this chat and carries on with it`)
669 if (S.keepWarm) lines.push(`- ◆ **Keep warm:** on until **${clock(S.keepWarmUntil)}**, ${S.pings} ping(s), ${usd(S.pingUsd)} so far`)
670 else lines.push(`- ◆ **Keep warm:** ${minutes(defaultKeepWarmMs())} by default on the ${ttlName(ttl)}, pays off up to **~${minutes(breakEvenMs())}** on ${modelName(S.model)}`)
671 if (S.rateLimits.length) lines.push(`- **Plan limits:** ${S.rateLimits.some((l) => (l.percentUsed || 0) >= 80) ? '⚠️ ' : ''}${limitsText(S.rateLimits)}`)
672 lines.push(`- **Cost:** session **${usd(S.costUsd)}**${S.todayUsd === null ? '' : ` · today **${todayText()}**`} (API list prices)`)
673 let restarts = `- **Cold restarts:** ${S.coldRestarts.length} (${usd(S.coldRestarts.reduce((a, r) => a + r.usd, 0))})`
674 if (S.cacheBreaks.length) restarts += ` · ⚠️ **Cache breaks:** ${S.cacheBreaks.length} (${usd(S.cacheBreaks.reduce((a, r) => a + r.usd, 0))}), last: ${S.cacheBreaks[S.cacheBreaks.length - 1].cause}`
675 lines.push(restarts)
676 lines.push(`- **Cold-send guard:** ${settings.guard ? 'on' : 'off'}, asks above **${bigText()}** · rewriting now ≈ ${usd(rewrite)}, ${isBig() ? 'above' : 'below'} that`)
677 lines.push(`- **Handoff hint:** ${settings.handoffAt ? `from **${tokens(settings.handoffAt)}** tokens${settings.handoffAt === HANDOFF_AT ? ' (default)' : ''}, then every ${tokens(HANDOFF_STEP)} more` : 'off'} · a handoff costs about **${usd(keepWarmVsHandoff(0).handoff)}** now (fresh chat ${tokens(S.startTokens || FRESH_START)}${S.startTokens ? '' : ' estimated'}, handoff ${tokens(handoffOut())}${S.handoffOuts.length ? `, average of the last ${S.handoffOuts.length}` : ' estimated'})`)
678 lines.push(`- **Alerts:** ${settings.alerts ? 'on' : 'off'}`)
679 lines.push(`- **Settings:** \`/${c} ttl 5|60|auto\` · \`/${c} guard on|off\` · \`/${c} big $1|default\` · \`/${c} handoff 140k|off|default\` · \`/${c} alerts on|off\``)
680 return lines.join('\n')
681}
682
683async function registerCommand($, name, description, argumentHint, immediate) {
684 const spec = immediate ? { name, description, argumentHint, immediate: true } : { name, description, argumentHint }
685 try {
686 await $.command.register(spec)
687 return name
688 } catch {
689 try {
690 await $.command.register({ ...spec, name: 'tk-' + name })
691 return 'tk-' + name
692 } catch {
693 return null
694 }
695 }
696}
697
698// The guard's "Compact first, then send", after its prompt.submit hook dropped
699// the message. Whatever fails, the message goes back in the prompt box unsent.
700async function compactThenSend($, held) {
701 let why = ''
702 try {
703 const r = await $.session.compact()
704 if (r && r.skip) why = r.skip
705 } catch (err) {
706 why = String((err && err.message) || err || 'it was refused')
707 }
708 if (why) {
709 log($, `guard: compact failed: ${why}`)
710 await $.prompt.fill({ text: held.text })
711 return warn($, `Not sent: compacting didn't work (${why}). Your message is back in the prompt box.`)
712 }
713 // Pasted images and @file mentions can't be sent again by a plugin
714 if (held.attached) {
715 await $.prompt.fill({ text: held.text })
716 return note($, 'Compacted. Your message is back in the prompt box: attach any images again, then send it.')
717 }
718 await $.prompt.submit({ text: held.text, asUser: true })
719}
720
721export function register(on) {
722 on('session.start', async ($, e, next) => {
723 now = await $.clock.now()
724 home = (await $.env.get('USERPROFILE')) || (await $.env.get('HOME')) || ''
725 S.id = await $.session.id()
726 S.cwd = await $.session.cwd()
727 S.model = await $.session.model()
728 await loadSettings($)
729 // Remember the variable as found, unless it is the value /cache ttl set (a reload)
730 const prior = await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL')
731 envPrior = settings.ttlMin && prior === ttlEnv(settings.ttlMin) ? undefined : prior
732 if (settings.ttlMin) await applyTtl($)
733 const outs = await $.store.get('handoffOuts')
734 S.handoffOuts = Array.isArray(outs) ? outs.filter((n) => n > 0).slice(-5) : []
735 measureStart($).catch(() => {})
736 names.cache = (await registerCommand($, 'cache', 'Token Keeper status and settings', '[ttl 5|60|auto] [guard on|off] [big $1|default] [handoff 140k|off|default] [alerts on|off]')) || names.cache
737 names.keepwarm = (await registerCommand($, 'keepwarm', 'Keep this session\'s prompt cache warm (default 30m on the 5-min TTL, 4h on the 1h TTL), or /keepwarm off', '[30m|4h|off]', true)) || names.keepwarm
738 names.handoff = (await registerCommand($, 'handoff', 'Session handoff, then clear this chat and continue with it (/handoff continue)', '[continue]')) || names.handoff
739 $.clock.every(TICK_EVERY, () => tick($).catch(() => {}))
740 refreshToday($).catch(() => {})
741 return next(e)
742 })
743
744 // /clear, /resume and /branch start from an unknown cache. The process goes
745 // on under a new session id, and no session.start fires for it.
746 on('classic.SessionStart', { source: ['clear', 'resume', 'fork'] }, async ($, e, next) => {
747 S.lastActivity = 0
748 S.keepWarm = false
749 H.armed = false
750 try {
751 S.id = (await $.session.id()) || S.id
752 } catch {
753 // keep the old id
754 }
755 // The old transcript is another session's now: read today's costs afresh
756 own.scanned = false
757 todayAt = 0
758 return next(e)
759 })
760
761 on('prompt.submit', async ($, e, next) => {
762 now = await $.clock.now()
763 // A new message in this chat makes a waiting handoff out of date (its file stays)
764 if (H.text && !H.continuing && e.origin && e.origin.kind === 'composer' && !String(e.text || '').trim().startsWith('/')) {
765 H.text = ''
766 H.path = ''
767 $.ui.invalidate('ui.render')
768 }
769 const ttl = ttlMin()
770 const left = msLeft(S.lastActivity, ttl)
771 const isCold = left !== null && left <= 0
772 const fromUser = e.origin && (e.origin.kind === 'composer' || e.origin.kind === 'bridge')
773 if (!settings.guard || !fromUser || !isCold || !isBig() || e.turnId) return next(e)
774 const cost = usd(rewriteCost(S.ctx, S.model, ttl))
775 let answer = 'Send anyway'
776 try {
777 answer = await $.ui.ask(
778 `Cache went cold ${minutes(-left)} ago. Sending now rewrites ${tokens(S.ctx)} tokens of context (about ${cost}). What should happen?`,
779 ['Send anyway', 'Compact first, then send', 'Cancel'],
780 )
781 } catch {
782 // nobody to ask (a -p run, or the dialog was dismissed): send as typed
783 return next(e)
784 }
785 if (answer === 'Compact first, then send') {
786 // Claude Code refuses compact() from a prompt.submit hook (it would
787 // compact under the turn this hook holds): hold the message, compact
788 // once this hook has returned, then send it
789 // A plugin's prompt expands no pasted images and no @file mentions
790 const attached = !!(e.attachments && e.attachments.length) || /(^|\s)@\S/.test(e.text)
791 const held = { text: e.text, attached }
792 $.clock.after(0, () => { compactThenSend($, held).catch(() => {}) })
793 return { drop: 'Held by token-keeper: compacting first, then sending it' }
794 }
795 if (answer === 'Send anyway') return next(e)
796 // A handoff now would rewrite the cache too: it only saves while warm
797 warn($, `Not sent. Next time, /${names.handoff} before the cache goes cold skips the rewrite.`)
798 return { drop: 'Cancelled by token-keeper before a cold cache rewrite' }
799 })
800
801 on('turn.start', async ($, e, next) => {
802 S.working = true
803 if (e.turnId) S.turnId = e.turnId
804 return next(e)
805 })
806
807 // Each request: what it cost today, and for the main loop, did it read the
808 // cache or rewrite it?
809 on('turn.step', async function* ($, e, next) {
810 const startedAt = await $.clock.now()
811 const usdBefore = e.agentId ? null : await sessionUsd($)
812 const result = yield* next(e)
813 if (!result || !result.usage) return result
814 const u = result.usage
815 addToday(requestCost(u, u.model || S.model, ttlMin()))
816 if (e.agentId) return result
817 now = await $.clock.now()
818 const total = totalInput(u)
819 const written = u.cache_creation_input_tokens || 0
820 // Each write's TTL is the cache window, measured
821 const usdAfter = usdBefore === null ? null : await sessionUsd($)
822 const seen = writeTtl(u, written, usdAfter === null ? null : usdAfter - usdBefore)
823 if (seen) {
824 S.ttlWritten = seen
825 if (!settings.ttlMin) {
826 S.ttlMin = S.ttlWritten
827 S.ttlSource = 'measured'
828 } else if (S.ttlWritten !== settings.ttlMin && !S.ttlWarned) {
829 S.ttlWarned = true
830 if (settings.alerts) warn($, `You set the ${settings.ttlMin}-min cache TTL, but Claude Code still writes ${S.ttlWritten}-min entries. FORCE_PROMPT_CACHING_5M overrides it, and setting it needs Claude Code v2.1.242 or later.`)
831 }
832 }
833 const gap = S.lastActivity ? startedAt - S.lastActivity : 0
834 const prevModel = S.model
835 S.model = u.model || S.model
836 const afterCompact = justCompacted
837 justCompacted = false
838 const afterTtlSwitch = ttlSwitched
839 ttlSwitched = false
840 // What the cache lost: the part of the previous context this request did not read.
841 // New content (a file read, a tool result) is written too, but that is no loss
842 // (against the last request's own size: S.ctx may already count the new content)
843 const lost = Math.max(0, S.lastTotal - (u.cache_read_input_tokens || 0))
844 const lostMuch = S.lastActivity && S.lastTotal > 30000 && lost > 0.2 * S.lastTotal
845 if (afterCompact) {
846 // the first request after a compaction writes the new, shorter context: expected
847 } else if (afterTtlSwitch) {
848 // a /cache ttl switch while warm: the rewrite is expected, so say what it really cost
849 if (written > 0) note($, `Switching to the ${ttlName(ttlMin())} rewrote ${tokens(written)} tokens for about ${usd(rewriteCost(written, S.model, ttlMin()))}.`)
850 } else if (lostMuch && gap < 4.5 * MIN) {
851 // Lost although the cache was still warm: something changed the prompt prefix
852 const cause = prevModel && u.model && prevModel !== u.model ? `the model changed (${priceFor(prevModel).id} → ${priceFor(u.model).id})` : 'the prompt prefix changed (CLAUDE.md, MCP tools, settings, effort, or system prompt)'
853 const cost = rewriteCost(lost, S.model, ttlMin())
854 S.cacheBreaks.push({ at: now, tokens: lost, usd: cost, cause })
855 if (settings.alerts) warn($, `Cache broken while warm: ${cause}. Rewrote ${tokens(lost)} tokens for about ${usd(cost)}.`)
856 } else if (lostMuch) {
857 // Cold: a fifth or more of the previous context was not read after a pause (the
858 // system prompt part often stays warm through other sessions).
859 // A rewrite after 5-60 idle minutes means this session runs on the 5-minute TTL
860 if (!settings.ttlMin && gap > 5.5 * MIN && gap < S.ttlMin * MIN) {
861 S.ttlMin = 5
862 S.ttlSource = 'measured'
863 }
864 S.coldRestarts.push({ at: now, tokens: lost, usd: rewriteCost(lost, S.model, ttlMin()), gapMin: Math.round(gap / MIN) })
865 } else if (S.lastActivity && total > 30000 && gap > 5.5 * MIN && cachedShare(u) > 0.8) {
866 // A hit after more than 5 idle minutes proves the 1-hour TTL
867 if (!settings.ttlMin) {
868 S.ttlMin = 60
869 S.ttlSource = 'measured'
870 }
871 }
872 if (H.armed) H.outTokens += u.output_tokens || 0
873 S.ctx = total
874 S.lastTotal = total
875 S.lastActivity = now
876 S.pingsSinceTurn = 0
877 if (e.effort) S.effort = String(e.effort)
878 return result
879 })
880
881 on('turn.complete', async ($, e, next) => {
882 const r = await next(e)
883 if (e.agentId) return r
884 now = await $.clock.now()
885 S.working = false
886 try {
887 const usage = await $.session.usage()
888 if (usage.context && usage.context.tokens) S.ctx = usage.context.tokens
889 if (usage.context && usage.context.window) S.window = usage.context.window
890 if (usage.cost) S.costUsd = usage.cost.usd
891 if (Array.isArray(usage.rateLimits) && usage.rateLimits.length) S.rateLimits = usage.rateLimits
892 } catch {
893 // usage unavailable: keep the per-request figures
894 }
895 if (H.armed && e.turnId !== H.notTurn) await captureHandoff($, e)
896 else handoffStep($)
897 $.ui.invalidate('ui.render')
898 return r
899 })
900
901 on('session.compact', async ($, e, next) => {
902 const r = await next(e)
903 justCompacted = true
904 return r
905 })
906
907 on('session.measure', async ($, e, next) => {
908 if (e.context && e.context.tokens) S.ctx = e.context.tokens
909 if (e.context && e.context.window) S.window = e.context.window
910 if (e.cost) S.costUsd = e.cost.usd
911 if (Array.isArray(e.rateLimits) && e.rateLimits.length) S.rateLimits = e.rateLimits
912 return next(e)
913 })
914
915 // Claude calling the handoff skill itself: this turn's answer is the handoff
916 on('tool.call', { tool: 'Skill' }, async ($, e, next) => {
917 if (!e.agentId && HANDOFF_SKILL.test(String(e.skill || ''))) armHandoff(true)
918 return next(e)
919 })
920
921 // /session-handoff from anywhere (typed, the band, /handoff): watch for its answer
922 on('command.run', async ($, e, next) => {
923 if (HANDOFF_SKILL.test(String(e.command || ''))) armHandoff(false)
924 return next(e)
925 })
926
927 // /handoff runs the handoff and then asks to clear and continue (the band's
928 // buttons draw only in the terminal); /handoff continue does the second half.
929 // Both run off the hook: a command started inside a hook the session is
930 // waiting on is refused.
931 on('command.run', { command: ['handoff', 'tk-handoff'] }, async ($, e) => {
932 now = await $.clock.now()
933 const arg = String(e.args || '').trim().toLowerCase()
934 if (arg === 'continue' || arg === 'go') {
935 if (!H.text) return { text: `No handoff ready yet. Type /${names.handoff} to make one.` }
936 $.clock.after(50, () => clearAndContinue($).catch(() => {}))
937 return { text: 'Clearing this chat and continuing with the handoff.' }
938 }
939 H.askAfter = true
940 $.clock.after(50, () => runHandoff($).catch(() => {}))
941 return { text: 'Running the session handoff. When it finishes, you can clear this chat and continue with it.' }
942 })
943
944 on('command.run', { command: ['cache', 'tk-cache'] }, async ($, e) => {
945 now = await $.clock.now()
946 const [key, value] = String(e.args || '').trim().toLowerCase().split(/\s+/)
947 if (key === 'ttl') {
948 const want = value === '5' ? 5 : value === '60' ? 60 : 0
949 // While the cache is warm, the next request rewrites what was cached under
950 // the old TTL; once it is cold that rewrite happens anyway, so no question
951 const st = cacheState()
952 if (want && want !== ttlMin() && ['warm', 'cooling', 'kept'].includes(st.kind) && S.ctx > 30000) {
953 const cost = usd(rewriteCost(S.ctx, S.model, want))
954 const keep = `Keep the ${ttlName(ttlMin())}`
955 let answer = keep
956 try {
957 answer = await $.ui.ask(`The cache is warm. Switching to the ${ttlName(want)} makes your next message rewrite up to ${tokens(S.ctx)} tokens (about ${cost}). Once the cache is cold, switching costs nothing extra. Switch now?`, ['Switch now', keep])
958 } catch {
959 // dismissed, or nobody to ask: keep the TTL
960 }
961 if (answer !== 'Switch now') return { text: `Cache TTL unchanged: ${ttlName(ttlMin())}. Switching now would rewrite up to ${tokens(S.ctx)} tokens (about ${cost}); after the cache goes cold it is free.` }
962 }
963 if (want !== settings.ttlMin && ['warm', 'cooling', 'kept'].includes(st.kind)) ttlSwitched = true
964 settings.ttlMin = want
965 S.ttlWritten = 0
966 S.ttlWarned = false
967 await applyTtl($)
968 } else if (key === 'guard' || key === 'alerts') {
969 settings[key] = value !== 'off'
970 } else if (key === 'big') {
971 const n = value === 'default' ? BIG_USD : parseUsd(value)
972 if (!n) return { text: `Usage: \`/${names.cache} big $1\` (a dollar amount), or \`/${names.cache} big default\`.` }
973 settings.bigUsd = n
974 } else if (key === 'handoff') {
975 const n = value === 'default' ? HANDOFF_AT : value === 'off' ? 0 : parseTokens(value)
976 if (n === null) return { text: `Usage: \`/${names.cache} handoff 140k\` (a context size in tokens), \`off\`, or \`default\`.` }
977 settings.handoffAt = n
978 S.handoffNextAt = 0
979 }
980 if (key) {
981 await $.store.set('settings', settings)
982 $.ui.invalidate('ui.render')
983 }
984 return { text: statusText() }
985 })
986
987 on('command.run', { command: ['keepwarm', 'tk-keepwarm'] }, async ($, e) => {
988 now = await $.clock.now()
989 const arg = String(e.args || '').trim().toLowerCase()
990 let text
991 if (arg === 'off' || (arg === '' && S.keepWarm)) {
992 text = stopKeepWarm('turned off') || 'ℹ️ Keep warm is already off.'
993 } else if (arg === '') {
994 await measureStart($)
995 text = startKeepWarm(defaultKeepWarmMs(), true, true)
996 } else {
997 const ms = parseDuration(arg)
998 if (!ms) return { text: `Usage: \`/${names.keepwarm} 30m\`, \`/${names.keepwarm} 4h\`, or \`/${names.keepwarm} off\`.` }
999 await measureStart($)
1000 text = startKeepWarm(ms, false, true)
1001 }
1002 $.ui.invalidate('ui.render')
1003 return { text }
1004 })
1005
1006 // The band above the prompt: this session's cache at a glance
1007 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
1008 const below = await next(e)
1009 if (e.props && e.props.hasSurvey) return below
1010 if (!S.lastActivity && !S.keepWarm) return below
1011 const mine = drawBand($, e)
1012 const { Box } = $.ui.resolve(e)
1013 return Box({ flexDirection: 'column', children: below ? [mine, below] : [mine] })
1014 })
1015
1016 // The band is terminal-only. The footer also draws in the Desktop app, so it
1017 // carries a short label, only when there's something to act on.
1018 on('ui.render', { component: 'SessionMode' }, async ($, e, next) => {
1019 const label = footerLabel()
1020 if (!label) return next(e)
1021 const modes = Array.isArray(e.props && e.props.modes) ? e.props.modes : []
1022 return next({ ...e, props: { ...e.props, modes: [...modes, label] } })
1023 })
1024}
1025
1026function drawBand($, e) {
1027 const { Box, Text, Button } = $.ui.resolve(e)
1028 const st = cacheState()
1029 const ttl = ttlMin()
1030 const big = isBig()
1031 const rewrite = usd(rewriteCost(S.ctx, S.model, ttl))
1032 const parts = []
1033 if (st.kind === 'kept') parts.push(Text({ color: 'cyan', children: [`◆ kept warm · ${S.pings} ping${S.pings === 1 ? '' : 's'} ${usd(S.pingUsd)} · until ${clock(S.keepWarmUntil)}`] }))
1034 else if (st.kind === 'warm') parts.push(Text({ color: 'green', children: [`● cache warm ${minutes(st.left)}`] }))
1035 else if (st.kind === 'cooling') parts.push(Text({ color: 'yellow', bold: true, children: [`◐ cache cools in ${minutes(st.left)}`] }))
1036 else if (st.kind === 'cold') parts.push(Text(big ? { color: 'red', bold: true, children: [`○ cache cold ${minutes(-st.left)}`] } : { dimColor: true, children: [`○ cache cold ${minutes(-st.left)}`] }))
1037 const pct = S.window ? Math.floor((S.ctx / S.window) * 100) : null
1038 const ctxText = ` │ ctx ${tokens(S.ctx)}${pct === null ? '' : `/${tokens(S.window)} ${pct}%`}`
1039 parts.push(Text(pct >= 80 ? { color: 'red', children: [ctxText] } : pct >= 50 ? { color: 'yellow', children: [ctxText] } : { dimColor: true, children: [ctxText] }))
1040 if (st.kind === 'cold' && big) parts.push(Text({ color: 'red', children: [` │ next send rewrites it ≈ ${rewrite}`] }))
1041 else parts.push(Text({ dimColor: true, children: [` │ rewrite ≈ ${rewrite}`] }))
1042 if (S.rateLimits.length) parts.push(Text(limitTone(S.rateLimits, { children: [' │ ' + limitsText(S.rateLimits)] })))
1043 if (S.coldRestarts.length) parts.push(Text({ dimColor: true, children: [` │ ${S.coldRestarts.length} cold restart${S.coldRestarts.length === 1 ? '' : 's'} ${usd(S.coldRestarts.reduce((a, c) => a + c.usd, 0))}`] }))
1044 if (S.cacheBreaks.length) parts.push(Text({ color: 'yellow', children: [` │ ${S.cacheBreaks.length} cache break${S.cacheBreaks.length === 1 ? '' : 's'} ${usd(S.cacheBreaks.reduce((a, c) => a + c.usd, 0))}`] }))
1045 const row = [Box({ flexDirection: 'row', children: parts })]
1046 if (st.kind === 'cooling' && big) {
1047 row.push(Button({ key: 'keepwarm', label: 'keep warm', hotkey: '1', plain: true, onPress: async () => { now = await $.clock.now(); log($, startKeepWarm(defaultKeepWarmMs(), true, false)); $.ui.invalidate('ui.render') } }))
1048 } else if (st.kind === 'kept') {
1049 row.push(Button({ key: 'keepwarm', label: 'stop warm', hotkey: '1', plain: true, onPress: async () => { now = await $.clock.now(); log($, stopKeepWarm('turned off')); $.ui.invalidate('ui.render') } }))
1050 }
1051 const left = Box({ flexDirection: 'row', columnGap: 2, children: row })
1052 // Right edge: one click runs /session-handoff (h while the band has the focus;
1053 // a letter never fires from the prompt box, unlike a digit). Once a handoff
1054 // is in, c clears this chat and sends it as the fresh chat's first prompt.
1055 const busy = handoffPending || H.armed
1056 const handoff = Button({ key: 'handoff', label: handoffPending ? 'handoff queued' : H.armed ? 'handoff running' : 'handoff', hotkey: 'h', variant: 'primary', dimColor: busy, onPress: () => runHandoff($) })
1057 let right = handoff
1058 if (H.text || H.continuing) {
1059 const go = Button({ key: 'continue', label: H.continuing ? 'clearing' : 'clear and continue', hotkey: 'c', plain: true, onPress: () => clearAndContinue($) })
1060 right = Box({ flexDirection: 'row', columnGap: 2, children: [go, handoff] })
1061 }
1062 const top = Box({ flexDirection: 'row', justifyContent: 'space-between', width: '100%', columnGap: 2, children: [left, right] })
1063 // Second row: the model and what it costs
1064 const info = [Text({ color: 'magenta', children: [modelName(S.model)] })]
1065 if (S.effort) info.push(Text({ dimColor: true, children: [` · ${S.effort}`] }))
1066 info.push(Text({ dimColor: true, children: [` │ session ${usd(S.costUsd)}`] }))
1067 if (S.todayUsd !== null) info.push(Text({ dimColor: true, children: [` · today ${todayText()}`] }))
1068 return Box({ flexDirection: 'column', width: '100%', children: [top, Box({ flexDirection: 'row', children: info })] })
1069}
1070
1071function footerLabel() {
1072 const parts = []
1073 const st = cacheState()
1074 const big = isBig()
1075 if (st.kind === 'kept') parts.push('cache kept warm')
1076 else if (st.kind === 'cooling' && big) parts.push(`cache cools in ${minutes(st.left)} · /${names.keepwarm} or /${names.handoff}`)
1077 else if (st.kind === 'cold' && big) parts.push(`cache cold · rewrite ≈ ${usd(rewriteCost(S.ctx, S.model, ttlMin()))}`)
1078 const high = S.rateLimits.filter((l) => (l.percentUsed || 0) >= 80)
1079 if (high.length) parts.push(limitsText(high))
1080 if (H.text) parts.push(`handoff ready · /${names.handoff} continue`)
1081 return parts.join(' · ')
1082}
1083hooks/pricing.mjs 92 lines1// Shared source. Copied into each mod's hooks/ folder by _dev/sync-shared.mjs.
2// Edit it here, then run: node mods/_dev/sync-shared.mjs
3
4// Dollars per million tokens, from the Claude API pricing table (cached 2026-09-25).
5// Cache writes cost 1.25x input on the 5-minute TTL and 2x input on the 1-hour TTL.
6const TABLE = [
7 ['fable-5-1', { input: 10, output: 50, read: 0.25 }],
8 ['mythos-5-1', { input: 10, output: 50, read: 0.25 }],
9 ['fable-5', { input: 10, output: 50, read: 1.0 }],
10 ['opus-5-5', { input: 4, output: 20, read: 0.2 }],
11 ['opus-5', { input: 5, output: 25, read: 0.5 }],
12 ['opus-4-8', { input: 5, output: 25, read: 0.5 }],
13 ['opus-4-7', { input: 5, output: 25, read: 0.5 }],
14 ['opus-4-6', { input: 5, output: 25, read: 0.5 }],
15 ['sonnet-5-5', { input: 2, output: 10, read: 0.2 }],
16 ['sonnet-5', { input: 2, output: 10, read: 0.2 }],
17 ['sonnet-4-6', { input: 3, output: 15, read: 0.3 }],
18 ['haiku-4-5', { input: 1, output: 5, read: 0.1 }],
19]
20
21const FAMILY = {
22 fable: 'fable-5-1',
23 mythos: 'mythos-5-1',
24 opus: 'opus-5-5',
25 sonnet: 'sonnet-5-5',
26 haiku: 'haiku-4-5',
27}
28
29export function normalizeModel(model) {
30 return String(model || '')
31 .toLowerCase()
32 .replace(/^claude-/, '')
33 .replace(/\[.*?\]/g, '')
34 .replace(/-\d{8}$/, '')
35 .trim()
36}
37
38export function priceFor(model) {
39 const id = normalizeModel(model)
40 for (const [key, price] of TABLE) {
41 if (id === key || id.startsWith(key + '-') || id.startsWith(key)) {
42 // 'opus-5' must not swallow 'opus-5-5': the table lists longer ids first
43 return { id: key, ...price }
44 }
45 }
46 for (const [family, key] of Object.entries(FAMILY)) {
47 if (id.includes(family)) {
48 const found = TABLE.find(([k]) => k === key)
49 return { id: key, ...found[1] }
50 }
51 }
52 return { id: 'opus-5-5', input: 4, output: 20, read: 0.2 }
53}
54
55export function writeRate(model, ttlMinutes) {
56 const p = priceFor(model)
57 return p.input * (ttlMinutes >= 60 ? 2 : 1.25)
58}
59
60// usage: { input_tokens, output_tokens, cache_read_input_tokens, cache_creation_input_tokens }
61export function requestCost(usage, model, ttlMinutes = 60) {
62 if (!usage) return 0
63 const p = priceFor(model || usage.model)
64 const w = writeRate(model || usage.model, ttlMinutes)
65 return (
66 ((usage.input_tokens || 0) * p.input +
67 (usage.cache_read_input_tokens || 0) * p.read +
68 (usage.cache_creation_input_tokens || 0) * w +
69 (usage.output_tokens || 0) * p.output) /
70 1e6
71 )
72}
73
74export function rewriteCost(tokens, model, ttlMinutes = 60) {
75 return ((tokens || 0) * writeRate(model, ttlMinutes)) / 1e6
76}
77
78export function totalInput(usage) {
79 if (!usage) return 0
80 return (
81 (usage.input_tokens || 0) +
82 (usage.cache_read_input_tokens || 0) +
83 (usage.cache_creation_input_tokens || 0)
84 )
85}
86
87// Share of a request's input served from the cache, 0 to 1.
88export function cachedShare(usage) {
89 const total = totalInput(usage)
90 return total > 0 ? (usage.cache_read_input_tokens || 0) / total : 0
91}
92hooks/fmt.mjs 41 lines1// Shared source. Copied into each mod's hooks/ folder by _dev/sync-shared.mjs.
2
3export function tokens(n) {
4 const v = Number(n) || 0
5 if (v >= 1e6) return (v / 1e6).toFixed(v >= 1e7 ? 0 : 1) + 'M'
6 if (v >= 1e3) return (v / 1e3).toFixed(v >= 1e5 ? 0 : 1) + 'k'
7 return String(Math.round(v))
8}
9
10export function usd(n) {
11 const v = Number(n) || 0
12 if (v === 0) return '$0'
13 if (v < 0.01) return '<$0.01'
14 if (v < 10) return '$' + v.toFixed(2)
15 return '$' + v.toFixed(0)
16}
17
18export function minutes(ms) {
19 const m = Math.round((Number(ms) || 0) / 60000)
20 if (m <= 60) return m + 'm'
21 return Math.floor(m / 60) + 'h' + String(m % 60).padStart(2, '0') + 'm'
22}
23
24export function clock(ts) {
25 const d = new Date(ts)
26 let h = d.getHours()
27 const ampm = h >= 12 ? 'pm' : 'am'
28 h = h % 12 || 12
29 return h + ':' + String(d.getMinutes()).padStart(2, '0') + ampm
30}
31
32export function clip(text, max) {
33 const s = String(text ?? '').replace(/\s+/g, ' ').trim()
34 return s.length > max ? s.slice(0, Math.max(0, max - 1)) + '…' : s
35}
36
37export function basename(path) {
38 const parts = String(path || '').split(/[\\/]+/).filter(Boolean)
39 return parts[parts.length - 1] || String(path || '')
40}
41