SLOPSHOPPER

token-keeper

Cache Keeper by Nate Herk, cut down to token economics: cache stats, a dollar-based cold-send guard, keep warm with a break-even, and handoff

newbandspinnerguardcommandprompt
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · token-keeper
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /cache ⎿ token-keeper: **Token Keeper** · cache window **60 min** (default) ⎿ token-keeper: ⎿ token-keeper: - **Context:** **97.4k** / 200k (48%) tokens on **Opus 5.5** ⎿ token-keeper: - – **Cache:** no request yet this session, so its state is unknown ⎿ token-keeper: - ◆ **Keep warm:** 4h00m by default on the 1h TTL, pays off up to **~34h40m** on Opus 5.5 ⎿ token-keeper: - **Plan limits:** 5h 31% │ ctx 97.4k/200k 48% │ rewrite ≈ $0.78 │ 5h 31% [ handoff ] Opus 5.5 │ session $0.42 · today $0 ⟨Claude Code's own drawing⟩ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Band
│ ctx 97.4k/200k 48% │ rewrite ≈ $0.78 │ 5h 31% [ handoff ] Opus 5.5 │ session $0.42 · today $0 ⟨Claude Code's own drawing⟩
README

Token Keeper

A Claude Code mod for token economics, for subscription and API users alike. It is a fork of Cache Keeper by Nate Herk, from nateherkai/claude-code-mods (MIT), cut down to four jobs: show what the cache costs, set it up, keep it warm, and hand a big chat off to a fresh one.

Credits

All of Cache Keeper's original work is Nate Herk's: the warm/cold cache band, /keepwarm, the cold-send guard, TTL detection, /handoff and the session-handoff skill, plus the shared pricing and fmt helpers. Go check out his repo and videos. This fork changes and adds the things listed below. It leaves out Cache Keeper's /board and recording mode for now.

What's different from Cache Keeper

  • Cache-break detector. If a request rewrites most of the context even though the cache was still warm (less than 4.5 minutes idle), Token Keeper reports it as a cache break and names the cause it suspects: a model switch, or a changed prompt prefix (CLAUDE.md, MCP tools, settings, effort). Cache breaks show on the band and in /cache.
  • The threshold is in dollars. A context counts as big when rewriting it costs at least /cache big ($1 by default). Token Keeper works out what that means in tokens for the current model and TTL, for example $1 ≈ 125k tokens on Opus 5.5 with the 1h TTL. Above it, the cold-send guard asks before a cold send, a line in the chat warns before the cache goes cold, and the band shows the keep-warm button.
  • Keep warm knows its break-even. /keepwarm runs for 30 minutes on the 5-minute TTL and for 4 hours on the 1-hour TTL by default. A rewrite costs as much as (cache-write price / cache-read price) pings, whatever the context size, so pinging pays off only up to a point, for example about 1h 28m on Opus 5.5 with the 5-minute TTL. Ask for longer and you get a warning that suggests /handoff or /clear instead.
  • /cache ttl sets the real TTL, and checks it. /cache ttl 5 or 60 sets CLAUDE_CODE_PROMPT_CACHE_TTL for the running Claude Code, so the next request already writes 5-minute or 1-hour cache entries; auto gives the variable back. Every cache write reports which TTL it used, so the cache window is measured rather than guessed, and a line warns once if Claude Code writes another TTL than the one you set (FORCE_PROMPT_CACHING_5M wins over it). Switching while the cache is warm makes the next message rewrite the part of the context cached under the other TTL, so /cache ttl asks first and names the price; once the cache is cold it switches without asking, because that rewrite happens anyway.
  • Handoff or keep warm? A handoff reads the context once, writes the handoff, and starts a fresh chat; after that nothing needs pinging. /keepwarm compares the two for the time you asked for and says when the handoff is cheaper, for example at 300k tokens on Opus 5.5 with the 5-minute TTL after about 14 minutes. It also says what each later turn saves, which is usually worth more: at 300k tokens each turn re-reads $0.06 of context, after a handoff about $0.005. The fresh chat's start is measured from /context, and the handoff's size is the average of your last five. Past /cache handoff tokens (140k by default, a common rule of thumb for where a model starts to lose track of a long context) a notice suggests a handoff or a /compact, which costs about the same, and asks whether to hand off now, be reminded after another 100k tokens, or get no more reminders. Not counted: files the fresh chat reads again because the handoff left them out, so a handoff pays off most when much of the old context is done with.
  • A handoff while the cache is warm. On a cold cache, writing a handoff rewrites the whole context too. So the cooling warning offers /handoff before the cache goes cold.
  • Today's cost across sessions. A cache file per day keeps what each transcript cost, so only transcripts that changed are read again, and this session counts live. A transcript too big to read shows as ≥ instead of being left out without a word.

Where it shows

Token Keeper runs wherever Claude Code runs, but each app draws a different part of it. Claude Code decides which parts each app gets, not the mod.

Terminal (CLI): everything

Token Keeper's band above the prompt in the Claude Code terminal

The band above the prompt shows the cache state (warm, cooling, cold, kept warm), the context size, what a rewrite costs, rate limits, cold restarts and cache breaks. While a big cache is cooling, it shows a keep warm button you press with 1.

Claude Code Desktop: band and footer

Token Keeper in the Claude Code Desktop app

The Desktop app draws the band too, with its buttons as native buttons. The footer also gets a short label, but only when there is something to act on.

VS Code: commands only, no band

Token Keeper in the Claude Code extension for VS Code

The VS Code extension does not draw the band, the footer label or the hint line under the prompt; Claude Code gives mods no place there. What still works:

  • /cache, /keepwarm and /handoff, with their full text output.
  • Questions in Claude Code's own dialog: the cold-send guard, /cache ttl while the cache is warm, and the handoff reminder.
  • Keep-warm pings and cost tracking, which run in the background on every app.

What you miss in VS Code:

  • No live cache state. Type /cache: it shows everything the band does, including the context against the window, the effort and what a rewrite costs.
  • No keep warm button and no hotkey 1. Type /keepwarm instead.
  • After /handoff, type /handoff continue to clear the chat and continue with the handoff.

If you want the band, run claude in VS Code's integrated terminal instead of the extension's panel.

Install

Needs Claude Code 2.1.287 or later, with mods turned on for your account.

claude plugin marketplace add LamAnhNguyenHTW/token-keeper
claude plugin install token-keeper@token-keeper

Don't install it alongside Cache Keeper: both register /cache, /keepwarm and /handoff. If a name is already taken, Token Keeper falls back to /tk-cache and so on.

Commands

CommandWhat it does
/cacheStatus and settings
`/cache ttl 5\60\auto`Sets Claude Code's cache TTL for the main conversation (CLAUDE_CODE_PROMPT_CACHE_TTL, needs Claude Code v2.1.242+). auto leaves it to Claude Code. Asks first while the cache is warm, since the switch rewrites the context once
`/cache guard on\off`Ask before a cold send of a big context
`/cache big $1\default`When a context counts as big: what a rewrite costs. $1 by default
`/cache handoff 140k\off\default`Suggest a handoff once the context passes this size, then every 100k more. 140k by default
`/cache alerts on\off`Warnings in the chat
`/keepwarm [30m\4h\off]`Keep the cache warm. 30m by default on the 5-minute TTL, 4h on the 1-hour TTL
/handoff [continue]Hand the chat off to a fresh one

Data

Everything stays local, under ~/.claude/mods-data/token-keeper/. The mod makes no network calls of its own. The only extra requests are the keep-warm pings, and those run only while /keepwarm is on.

License

MIT. See LICENSE. The original copyright belongs to Nate Herk.

Source 3 files
hooks/register.mjs 1083 lines
1// Token Keeper (based on Cache Keeper by Nate Herk): watches this session's
2// prompt cache, keeps a big cache warm on request, asks before a cold send,
3// and hands a big chat off to a fresh one.
4//
5// Why: each request re-reads the whole context. From the cache that costs about
6// a tenth of normal input. Once the cache expires (1 hour idle on a subscription,
7// 5 minutes on the default API TTL), the next message writes the whole context
8// again at 1.25x to 2x input. A cache read also restarts the timer, so a tiny
9// ping before expiry costs a fraction of a rewrite.
10
11import { rewriteCost, requestCost, totalInput, cachedShare, priceFor, writeRate } from './pricing.mjs'
12import { tokens, usd, minutes, clock, clip, basename } from './fmt.mjs'
13
14const MIN = 60000
15const TICK_EVERY = 30000
16const TODAY_EVERY = 10 * MIN
17const MAX_READ = 4 * 1024 * 1024 // $.fs.read refuses larger files
18// The default cold-send threshold: below a rewrite of about $1 a dialog
19// bothers more than the rewrite costs (about 125k tokens on Opus 5.5, 1h TTL)
20const BIG_USD = 1
21// What a handoff costs besides one read of the context: its output, and a fresh
22// chat's start (system prompt, tools, memory files). Both are measured; these
23// stand in until they are
24const HANDOFF_OUT = 3000
25const FRESH_START = 20000
26// The default context size from which a handoff is suggested (a rule of thumb:
27// past about 140k a model loses track more often), and how much more the
28// context grows before the next reminder
29const HANDOFF_AT = 140000
30const HANDOFF_STEP = 100000
31
32// This session
33const S = {
34  id: '',
35  cwd: '',
36  model: '',
37  lastActivity: 0, // last main-loop request or keep-warm ping that touched the cache
38  ctx: 0,
39  lastTotal: 0, // the last main-loop request's prompt size, what its cache entry holds
40  window: 0,
41  costUsd: 0,
42  ttlMin: 60,
43  ttlSource: 'default',
44  ttlWritten: 0, // the TTL of the last main-loop cache write, in minutes (0: none since /cache ttl)
45  ttlWarned: false, // said once that Claude Code does not write the TTL /cache ttl set
46  coldRestarts: [],
47  working: false,
48  turnId: '', // the main loop's running (or last) turn
49  keepWarm: false,
50  keepWarmUntil: 0,
51  pings: 0,
52  pingUsd: 0,
53  pingsSinceTurn: 0, // keep-warm pings since the last main-loop request
54  rateLimits: [],
55  cacheBreaks: [], // rewrites while the cache was still warm
56  effort: '',
57  todayUsd: null, // every session's spend since local midnight
58  todayPartial: false, // a transcript was too big to read, so todayUsd is a lower bound
59  startTokens: 0, // what a fresh chat starts with, as /context counts it (0: not measured)
60  handoffOuts: [], // output tokens of the last few handoffs, kept across sessions
61  handoffNextAt: 0, // the size of the next handoff reminder (0: /cache handoff)
62}
63
64const settings = { bigUsd: BIG_USD, guard: true, ttlMin: 0, alerts: true, handoffAt: HANDOFF_AT }
65const names = { cache: 'cache', keepwarm: 'keepwarm', handoff: 'handoff' }
66
67let now = 0
68let home = ''
69let warnedFor = 0 // the lastActivity the cooling warning was shown for
70let justCompacted = false
71let ttlSwitched = false // /cache ttl changed while the cache was warm: the next request rewrites
72let handoffPending = false
73let envPrior // CLAUDE_CODE_PROMPT_CACHE_TTL before /cache ttl set it
74
75// The handoff flow. When /session-handoff runs (the band's button, /handoff, a
76// typed command, or Claude calling the skill), the answer of the turn it starts
77// is captured and saved to a file. "clear and continue" then runs /clear and
78// sends the handoff as the fresh chat's first prompt.
79const HANDOFF_SKILL = /(^|:)session-handoff$/
80const H = {
81  armed: false,
82  outTokens: 0, // output of the handoff turn so far
83  notTurn: '', // a turn already running when the handoff was queued: not the handoff's
84  askAfter: false, // /handoff asks "clear and continue?" once the handoff is in
85  text: '',
86  path: '',
87  continuing: false,
88}
89
90function armHandoff(insideTurn) {
91  H.armed = true
92  H.outTokens = 0
93  H.notTurn = insideTurn ? '' : S.working ? S.turnId : ''
94}
95
96// The handoff without any preamble Claude put before its heading
97function handoffBody(answer) {
98  const text = String(answer || '').trim()
99  const i = text.indexOf('# Session Handoff')
100  return i > 0 ? text.slice(i) : text
101}
102
103function fileStamp(ms) {
104  const d = new Date(ms)
105  const two = (n) => String(n).padStart(2, '0')
106  return `${d.getFullYear()}-${two(d.getMonth() + 1)}-${two(d.getDate())}-${two(d.getHours())}${two(d.getMinutes())}`
107}
108
109function dataDir() {
110  return `${home || '.'}/.claude/mods-data/token-keeper`.replace(/\\/g, '/')
111}
112
113async function saveHandoff($, text) {
114  const project = basename(S.cwd).replace(/[^A-Za-z0-9-]+/g, '-') || 'session'
115  const path = `${dataDir()}/handoffs/${fileStamp(now)}-${project}-${S.id.slice(0, 8)}.md`
116  await $.fs.write(path, text + '\n')
117  return path
118}
119
120async function captureHandoff($, e) {
121  H.armed = false
122  H.notTurn = ''
123  const askAfter = H.askAfter
124  H.askAfter = false
125  const body = e.reason === 'answer' ? handoffBody(e.answer) : ''
126  if (body.length < 200) {
127    warn($, 'The handoff turn ended without a handoff, so nothing was saved.')
128    return
129  }
130  H.text = body
131  if (H.outTokens > 0) {
132    S.handoffOuts = [...S.handoffOuts, H.outTokens].slice(-5)
133    await $.store.set('handoffOuts', S.handoffOuts).catch(() => {})
134  }
135  try {
136    H.path = await saveHandoff($, body)
137  } catch {
138    H.path = ''
139  }
140  const saved = H.path ? `Handoff saved to ${H.path}.` : 'Handoff captured (the backup file could not be written).'
141  note($, `${saved} Press c on the band or type /${names.handoff} continue to clear and continue.`)
142  // Off the hook: the turn is ending, and a dialog would hold it open
143  if (askAfter) $.clock.after(300, () => offerContinue($).catch(() => {}))
144}
145
146async function offerContinue($) {
147  if (!H.text) return
148  let answer = ''
149  try {
150    answer = await $.ui.ask('Handoff saved. Clear this chat and continue with it in a fresh context?', ['Clear and continue', 'Keep this chat'])
151  } catch {
152    return // dismissed: the band button and /handoff continue still work
153  }
154  if (answer === 'Clear and continue') await clearAndContinue($)
155}
156
157function continuationPrompt() {
158  const where = H.path ? ` (saved at ${H.path})` : ''
159  return `Handoff from my previous session${where}:\n\n${H.text}`
160}
161
162async function clearAndContinue($) {
163  if (H.continuing) return
164  if (!H.text) {
165    note($, `No handoff ready. Press h on the band or type /${names.handoff} first.`)
166    return
167  }
168  if (S.working) {
169    note($, 'Claude is still working. Clear and continue once this turn ends.')
170    return
171  }
172  const text = continuationPrompt()
173  const path = H.path
174  H.continuing = true
175  $.ui.invalidate('ui.render')
176  try {
177    await $.command.run({ command: 'clear', args: '' })
178  } catch (err) {
179    H.continuing = false
180    $.ui.invalidate('ui.render')
181    warn($, 'Could not run /clear: ' + clip(String((err && err.message) || err), 100) + (path ? ` The handoff is saved at ${path}.` : ''))
182    return
183  }
184  H.text = ''
185  H.path = ''
186  try {
187    await $.prompt.submit({ text, asUser: true })
188  } catch {
189    // Not sent: leave it in the prompt box for one Enter
190    let filled = false
191    try {
192      filled = !!(await $.prompt.fill({ text })).isFilled
193    } catch {
194      filled = false
195    }
196    if (filled) note($, 'Cleared. The handoff is in the prompt box: press Enter to send it.')
197    else warn($, `Cleared, but the handoff could not be sent.${path ? ` It's saved at ${path}.` : ''}`)
198  } finally {
199    H.continuing = false
200    $.ui.invalidate('ui.render')
201  }
202}
203
204// Runs the /session-handoff skill as if you typed it. The engine queues it
205// until Claude finishes the current turn.
206async function runHandoff($) {
207  if (handoffPending) {
208    note($, 'Session handoff is already queued.')
209    return
210  }
211  try {
212    const commands = await $.command.list()
213    const cmd = commands.find((c) => c.name === 'session-handoff') || commands.find((c) => HANDOFF_SKILL.test(c.name))
214    if (!cmd) {
215      H.askAfter = false
216      warn($, 'No /session-handoff skill in this session.')
217      return
218    }
219    handoffPending = true
220    armHandoff(false)
221    $.ui.invalidate('ui.render')
222    note($, S.working ? 'Session handoff queued: it runs when Claude finishes this turn.' : 'Running /session-handoff.')
223    await $.command.run({ command: cmd.name, args: '' })
224  } catch (err) {
225    H.armed = false
226    H.askAfter = false
227    warn($, 'Could not start /session-handoff: ' + clip(String((err && err.message) || err), 100))
228  } finally {
229    handoffPending = false
230    $.ui.invalidate('ui.render')
231  }
232}
233
234// What the cache really uses: the last write's TTL, else /cache ttl, else the estimate
235function ttlMin() {
236  return S.ttlWritten || settings.ttlMin || S.ttlMin
237}
238
239// The session's spend as Claude Code counts it, or null when it does not say
240async function sessionUsd($) {
241  try {
242    const cost = (await $.session.usage()).cost
243    return cost ? cost.usd : null
244  } catch {
245    return null
246  }
247}
248
249// The TTL a cache write used: the usage's 1h/5m split when it has one; else the
250// price that matches what Claude Code charged for the request, since its cost
251// counts the TTL really written (0: too small to tell, or another request in between)
252function writeTtl(u, written, charged) {
253  if (u.cache_creation && written > 0) return (u.cache_creation.ephemeral_1h_input_tokens || 0) * 2 >= written ? 60 : 5
254  if (charged === null || written < 1000) return 0
255  const p5 = requestCost(u, u.model || S.model, 5)
256  const p60 = requestCost(u, u.model || S.model, 60)
257  const near = Math.abs(charged - p60) < Math.abs(charged - p5) ? 60 : 5
258  return Math.abs(charged - (near === 60 ? p60 : p5)) < Math.abs(p60 - p5) / 4 ? near : 0
259}
260
261function ttlEnv(min) {
262  return min === 5 ? '5m' : min === 60 ? '1h' : undefined
263}
264
265// Points Claude Code's own cache TTL at the /cache ttl choice; auto gives the
266// variable back as it was before Token Keeper set it
267async function applyTtl($) {
268  // a refusal must not break /cache: the next write's TTL check reports it
269  await $.env.set('CLAUDE_CODE_PROMPT_CACHE_TTL', settings.ttlMin ? ttlEnv(settings.ttlMin) : envPrior).catch(() => {})
270}
271
272function ttlSourceText() {
273  if (!settings.ttlMin) return S.ttlSource
274  if (!S.ttlWritten) return 'set by you'
275  if (S.ttlWritten === settings.ttlMin) return 'set by you · confirmed'
276  return `⚠️ you set ${settings.ttlMin} min, but Claude Code writes ${S.ttlWritten} min`
277}
278
279function ttlName(ttl) {
280  return ttl >= 60 ? '1h TTL' : '5-min TTL'
281}
282
283// Plan limit windows as the API reports them: five_hour, seven_day, spend_limit
284const LIMIT_LABEL = { five_hour: '5h', seven_day: 'week', spend_limit: 'spend' }
285
286function limitsText(limits) {
287  return limits.map((l) => `${LIMIT_LABEL[l.kind] || l.kind} ${Math.round(l.percentUsed)}%`).join(' · ')
288}
289
290function limitTone(limits, extra) {
291  const top = Math.max(...limits.map((l) => l.percentUsed || 0))
292  if (top >= 95) return { ...extra, color: 'red', bold: true }
293  if (top >= 80) return { ...extra, color: 'yellow' }
294  return { ...extra, dimColor: true }
295}
296
297// 'claude-opus-5-5' -> 'Opus 5.5'
298function modelName(model) {
299  const [family, ...ver] = priceFor(model).id.split('-')
300  return family[0].toUpperCase() + family.slice(1) + ' ' + ver.join('.')
301}
302
303function msLeft(lastActivity, ttl) {
304  if (!lastActivity) return null
305  return ttl * MIN - (now - lastActivity)
306}
307
308function cacheState() {
309  const left = msLeft(S.lastActivity, ttlMin())
310  if (left === null) return { kind: 'unknown', left: 0 }
311  if (S.keepWarm) return { kind: 'kept', left }
312  if (left <= 0) return { kind: 'cold', left }
313  // Cooling: from when a keep-warm ping would go out (5 minutes before on the 1h
314  // TTL, 90s on the 5-minute one), so a 5-minute cache is not cooling right away
315  const ttl = ttlMin()
316  if (left <= Math.min(5 * MIN, ttl * MIN - pingEvery(ttl))) return { kind: 'cooling', left }
317  return { kind: 'warm', left }
318}
319
320// Big: rewriting this context costs at least the threshold, on this model and TTL
321function isBig() {
322  return rewriteCost(S.ctx, S.model, ttlMin()) >= settings.bigUsd
323}
324
325// "$1.00 (default) ≈ 125k tokens on Opus 5.5, 1h TTL"
326function bigText() {
327  const ttl = ttlMin()
328  const atTokens = settings.bigUsd / (writeRate(S.model, ttl) / 1e6)
329  return `${usd(settings.bigUsd)}${settings.bigUsd === BIG_USD ? ' (default)' : ''} ≈ ${tokens(atTokens)} tokens on ${modelName(S.model)}, ${ttlName(ttl)}`
330}
331
332// Keep warm pings this long after the last read, and every ping after it
333function pingEvery(ttl) {
334  return ttl * MIN - (ttl >= 60 ? 8 * MIN : 90000)
335}
336
337// How long pinging stays cheaper than one rewrite: a rewrite costs as much as
338// (write price / read price) pings, whatever the context size
339function breakEvenMs() {
340  const ttl = ttlMin()
341  return (writeRate(S.model, ttl) / priceFor(S.model).read) * pingEvery(ttl)
342}
343
344function defaultKeepWarmMs() {
345  return ttlMin() >= 60 ? 4 * 60 * MIN : 30 * MIN
346}
347
348// "30m", "30min", "2h", "1.5" (hours); capped at 24 hours
349function parseDuration(text) {
350  const m = String(text || '').match(/^(\d+(?:\.\d+)?)\s*(m|min|h)?$/)
351  if (!m || !(Number(m[1]) > 0)) return null
352  const ms = Number(m[1]) * (m[2] === 'm' || m[2] === 'min' ? MIN : 60 * MIN)
353  return Math.min(24 * 60 * MIN, ms)
354}
355
356// "300k", "1.5M", "300000"
357function parseTokens(text) {
358  const m = String(text || '').match(/^(\d+(?:\.\d+)?)\s*(k|m)?$/)
359  if (!m || !(Number(m[1]) > 0)) return null
360  return Math.round(Number(m[1]) * (m[2] === 'k' ? 1e3 : m[2] === 'm' ? 1e6 : 1))
361}
362
363// What a fresh chat starts with: every /context row but the messages. A local
364// estimate, no request
365async function measureStart($) {
366  try {
367    const u = await $.session.usage({ breakdown: 'summary' })
368    const rows = (u.context && u.context.breakdown && u.context.breakdown.categories) || []
369    const sum = rows.filter((r) => r.kind === 'used' && r.name !== 'Messages').reduce((a, r) => a + (r.tokens || 0), 0)
370    if (sum > 0) S.startTokens = sum
371  } catch {
372    // no breakdown: keep the last figure or the default
373  }
374}
375
376// When the context passes the handoff size: a notice, then a choice to hand
377// off now, be reminded HANDOFF_STEP later, or turn the reminders off
378function handoffStep($) {
379  if (!settings.alerts || !settings.handoffAt) return
380  const at = S.handoffNextAt || settings.handoffAt
381  if (S.ctx < at) return
382  let next = at + HANDOFF_STEP
383  while (next <= S.ctx) next += HANDOFF_STEP
384  S.handoffNextAt = next
385  const c = keepWarmVsHandoff(0)
386  // Quality is the reason; the saving is named only when the fresh chat is clearly smaller
387  const saves = c.perTurn >= 0.01 ? `, and every turn after it re-reads about ${tokens(c.fresh)} instead of ${tokens(S.ctx)}: ${usd(c.perTurn)} less per turn` : ''
388  note($, `This chat is at ${tokens(S.ctx)} tokens, and a long context gets less reliable. A /${names.handoff} or /compact (about ${usd(c.handoff)}, either) carries on in a fresh chat${saves}.`)
389  // Off the hook: the turn is ending, and a dialog would hold it open
390  $.clock.after(300, () => offerHandoff($, next).catch(() => {}))
391}
392
393async function offerHandoff($, next) {
394  const remind = `Remind me after another ${tokens(HANDOFF_STEP)} (${tokens(next)})`
395  let answer = ''
396  try {
397    answer = await $.ui.ask(`This chat is at ${tokens(S.ctx)} tokens. Hand it off to a fresh chat?`, ['Handoff now', remind, 'No more reminders'])
398  } catch {
399    return // dismissed: remind at the next step
400  }
401  if (answer === 'Handoff now') {
402    H.askAfter = true
403    await runHandoff($)
404  } else if (answer === 'No more reminders') {
405    settings.handoffAt = 0
406    await $.store.set('settings', settings)
407    note($, `Handoff reminders off. /${names.cache} handoff default turns them back on.`)
408  }
409}
410
411function parseUsd(text) {
412  const m = String(text || '').match(/^\$?(\d+(?:\.\d+)?)$/)
413  return m && Number(m[1]) > 0 ? Number(m[1]) : null
414}
415
416// Lines in the chat (plain text, drawn dim): ⚠️ a warning, ℹ️ a notice
417function warn($, text) {
418  $.ui.log('⚠️ ' + text)
419}
420
421function note($, text) {
422  $.ui.log('ℹ️ ' + text)
423}
424
425// Bold in a command's output, which draws markdown; plain in a chat line
426function bolder(md) {
427  return md ? (s) => `**${s}**` : (s) => s
428}
429
430// Keep warm for ms vs a handoff now, in dollars. A handoff reads the context
431// once, writes the handoff, and the fresh chat writes its start; then no pings
432// The handoff's output: the average of the last few, or the estimate
433function handoffOut() {
434  return S.handoffOuts.length ? S.handoffOuts.reduce((a, n) => a + n, 0) / S.handoffOuts.length : HANDOFF_OUT
435}
436
437// Keep warm for ms vs a handoff now, in dollars, and what each turn after the
438// handoff saves on re-reading the context (a /compact costs about the same)
439function keepWarmVsHandoff(ms) {
440  const ttl = ttlMin()
441  const p = priceFor(S.model)
442  const ping = S.ctx * p.read / 1e6
443  const start = S.startTokens || FRESH_START
444  const fresh = start + handoffOut()
445  return {
446    ping,
447    fresh,
448    keepWarm: Math.floor(ms / pingEvery(ttl)) * ping, // a ping every pingEvery, none at the start
449    handoff: ping + handoffOut() * p.output / 1e6 + start * writeRate(S.model, ttl) / 1e6,
450    perTurn: Math.max(0, S.ctx - fresh) * p.read / 1e6,
451  }
452}
453
454// Starts keeping the cache warm and says so, with ⚠️ past the break-even
455function startKeepWarm(ms, isDefault, md) {
456  // A new run counts its own pings; extending a running one keeps counting
457  if (!S.keepWarm) {
458    S.pings = 0
459    S.pingUsd = 0
460  }
461  S.keepWarm = true
462  S.keepWarmUntil = now + ms
463  const b = bolder(md)
464  const ttl = ttlMin()
465  const be = breakEvenMs()
466  const how = `${b(minutes(ms))}${isDefault ? `, the default on the ${ttlName(ttl)}` : ''}, until ${clock(S.keepWarmUntil)}`
467  const on = `On ${modelName(S.model)} with the ${ttlName(ttl)}`
468  if (ms > be) return `⚠️ Keeping this cache warm for ${how}. ${on} that only pays off up to ${b('~' + minutes(be))}: after that, one rewrite (${usd(rewriteCost(S.ctx, S.model, ttl))}) is cheaper. Better: ${b(`/${names.keepwarm} ${ttl >= 60 ? '4h' : '30m'}`)} (the default), ${b('/' + names.handoff)} now while the cache is warm, or ${b('/clear')}.`
469  const c = keepWarmVsHandoff(ms)
470  const vs = c.handoff < c.keepWarm ? ` A ${b('/' + names.handoff)} now is cheaper: about ${usd(c.handoff)} vs ${usd(c.keepWarm)} for ${minutes(ms)} of pings. After it, each turn re-reads about ${tokens(c.fresh)} instead of ${tokens(S.ctx)}, about ${usd(c.perTurn)} less.` : ''
471  return `ℹ️ Keeping this cache warm for ${how}. ${on} it pays off up to ~${minutes(be)}.${vs}`
472}
473
474// Stops keeping the cache warm and says so ('' when it was off), ⚠️ for a problem
475function stopKeepWarm(why, problem) {
476  if (!S.keepWarm) return ''
477  S.keepWarm = false
478  return `${problem ? '⚠️' : 'ℹ️'} Keep warm off (${why}). ${S.pings} ping(s), ${usd(S.pingUsd)}.`
479}
480
481function log($, text) {
482  if (text) $.ui.log(text)
483}
484
485async function keepWarmStep($) {
486  if (!S.keepWarm) return
487  if (now >= S.keepWarmUntil) return log($, stopKeepWarm('time limit reached'))
488  if (S.working || !S.lastActivity) return
489  const ttl = ttlMin()
490  if (now - S.lastActivity < pingEvery(ttl)) return
491  // Past the TTL the cache is gone (keep warm started late, or the machine slept):
492  // a ping would only write it again
493  if (now - S.lastActivity >= ttl * MIN) return log($, stopKeepWarm(`the cache already went cold ${minutes(now - S.lastActivity - ttl * MIN)} ago, so a ping would only write it again`, true))
494  let reply = null
495  try {
496    reply = await $.model.fork({ prompt: 'token-keeper keep-alive ping. Reply with only: ok' })
497  } catch {
498    reply = null
499  }
500  if (!reply || !reply.usage) {
501    return log($, stopKeepWarm('the ping got no answer, so the cache may already be cold', true))
502  }
503  const u = reply.usage
504  const cost = requestCost(u, S.model, ttl)
505  S.pings += 1
506  S.pingUsd += cost
507  addToday(cost)
508  // The first ping after a turn may write the turn's tail (its prefix ends after the
509  // reply, the main thread's entry before it); a later ping that writes much missed
510  // the cache: stop paying for it
511  const first = S.pingsSinceTurn === 0
512  S.pingsSinceTurn += 1
513  if (!first && (u.cache_creation_input_tokens || 0) > 0.1 * Math.max(1, u.cache_read_input_tokens || 0)) {
514    return log($, stopKeepWarm(`the ping wrote ${tokens(u.cache_creation_input_tokens)} tokens instead of reading the cache`, true))
515  }
516  S.lastActivity = now
517}
518
519// Shortly before a big cache goes cold: keep it warm, or hand off while it's warm
520function warnStep($) {
521  if (!settings.alerts || S.keepWarm || !isBig()) return
522  const st = cacheState()
523  if (st.kind !== 'cooling' || warnedFor === S.lastActivity) return
524  warnedFor = S.lastActivity
525  const cost = usd(rewriteCost(S.ctx, S.model, ttlMin()))
526  warn($, `This ${tokens(S.ctx)}-token cache goes cold in ${minutes(st.left)}. Rewriting it then costs about ${cost}. Back soon: /${names.keepwarm} (or press 1). Done for now: /${names.handoff} while it's still warm.`)
527}
528
529async function tick($) {
530  now = await $.clock.now()
531  await keepWarmStep($)
532  warnStep($)
533  if (now - todayAt > TODAY_EVERY) await refreshToday($).catch(() => {})
534  $.ui.invalidate('ui.render')
535}
536
537// Today's spend at API list prices, every session's (subagents too). A cache
538// file per day keeps each transcript's costs, so only transcripts that changed
539// are read again. This session's own transcripts are read once, then its
540// requests are counted live. Each response counts once across files.
541let todayAt = 0
542let todayDate = ''
543let todayFiles = {} // path -> { mtimeMs, costs: { responseKey: usd } } or { mtimeMs, tooBig: true }
544let othersUsd = 0
545let othersPartial = false
546const own = { scanned: false, base: 0, live: 0, partial: false }
547
548function localDate(ms) {
549  const d = new Date(ms)
550  return `${d.getFullYear()}-${d.getMonth() + 1}-${d.getDate()}`
551}
552
553function setToday() {
554  S.todayUsd = othersUsd + own.base + own.live
555  S.todayPartial = othersPartial || own.partial
556}
557
558// A request of this session: counted live once its transcripts were read
559function addToday(cost) {
560  if (!own.scanned) return // the first read finds it in the transcript
561  own.live += cost
562  setToday()
563}
564
565async function readCosts($, path, mtimeMs, size, date) {
566  if (size > MAX_READ) return { mtimeMs, tooBig: true }
567  const costs = {}
568  const text = await $.fs.read(path).catch(() => null)
569  if (text === null) return { mtimeMs, tooBig: true }
570  for (const line of text.split('\n')) {
571    if (!line.includes('"usage"')) continue
572    let entry
573    try { entry = JSON.parse(line) } catch { continue }
574    const msg = entry.message
575    if (!msg || !msg.usage || !entry.timestamp || msg.model === '<synthetic>') continue
576    if (localDate(Date.parse(entry.timestamp)) !== date) continue
577    const cc = msg.usage.cache_creation || {}
578    const w1h = cc.ephemeral_1h_input_tokens || 0
579    const w5m = cc.ephemeral_5m_input_tokens ?? Math.max(0, (msg.usage.cache_creation_input_tokens || 0) - w1h)
580    costs[`${msg.id}:${entry.requestId}`] = requestCost({ ...msg.usage, cache_creation_input_tokens: w1h }, msg.model, 60) + requestCost({ cache_creation_input_tokens: w5m }, msg.model, 5)
581  }
582  return { mtimeMs, costs }
583}
584
585async function refreshToday($) {
586  todayAt = now
587  const date = localDate(now)
588  const cachePath = `${dataDir()}/today/${date}.json`
589  if (date !== todayDate) {
590    todayDate = date
591    todayFiles = {}
592    try {
593      todayFiles = JSON.parse(await $.fs.read(cachePath)).files || {}
594    } catch {
595      // no cache file yet today
596    }
597    own.scanned = false
598  }
599  const midnight = new Date(now).setHours(0, 0, 0, 0)
600  const root = ((await $.env.get('CLAUDE_CONFIG_DIR')) || home + '/.claude') + '/projects'
601  const files = []
602  const walk = async (dir) => {
603    for (const f of await $.fs.list(dir).catch(() => [])) {
604      const p = dir + '/' + f.name
605      if (f.kind === 'dir') await walk(p)
606      else if (f.name.endsWith('.jsonl') && f.mtimeMs >= midnight) files.push({ p, mtimeMs: f.mtimeMs, size: f.size })
607    }
608  }
609  await walk(root)
610  const scanOwn = !own.scanned
611  if (scanOwn) Object.assign(own, { base: 0, live: 0, partial: false })
612  const seen = new Map()
613  let partial = false
614  let changed = false
615  for (const { p, mtimeMs, size } of files) {
616    const mine = !!S.id && p.includes(S.id) // the transcript and its subagents' folder
617    if (mine && !scanOwn) continue
618    let entry = todayFiles[p]
619    if (!entry || entry.mtimeMs !== mtimeMs) {
620      entry = await readCosts($, p, mtimeMs, size, date)
621      todayFiles[p] = entry
622      changed = true
623    }
624    if (entry.tooBig) {
625      if (mine) own.partial = true
626      else partial = true
627    } else if (mine) {
628      own.base += Object.values(entry.costs).reduce((a, v) => a + v, 0)
629    } else {
630      for (const [k, v] of Object.entries(entry.costs)) seen.set(k, v)
631    }
632  }
633  own.scanned = true
634  othersUsd = [...seen.values()].reduce((a, v) => a + v, 0)
635  othersPartial = partial
636  setToday()
637  if (changed) await $.fs.write(cachePath, JSON.stringify({ files: todayFiles })).catch(() => {})
638}
639
640function todayText() {
641  return (S.todayPartial ? '≥ ' : '') + usd(S.todayUsd)
642}
643
644async function loadSettings($) {
645  const saved = await $.store.get('settings')
646  if (!saved || typeof saved !== 'object') return
647  for (const k of Object.keys(settings)) if (saved[k] !== undefined) settings[k] = saved[k]
648}
649
650function statusText() {
651  const st = cacheState()
652  const ttl = ttlMin()
653  const rewrite = rewriteCost(S.ctx, S.model, ttl)
654  const c = names.cache
655  // Markdown, as a command's output draws: a bold label per line, the band's icons.
656  // Everything the band shows is here too: VS Code draws no band, only this.
657  const lines = []
658  lines.push(`**Token Keeper** · cache window **${ttl} min** (${ttlSourceText()})`)
659  lines.push('')
660  const pct = S.window ? Math.floor((S.ctx / S.window) * 100) : null
661  const full = pct >= 80 ? '⚠️ ' : ''
662  lines.push(`- **Context:** ${full}**${tokens(S.ctx)}**${pct === null ? '' : ` / ${tokens(S.window)} (${pct}%)`} tokens on **${modelName(S.model)}**${S.effort ? ` · ${S.effort}` : ''}`)
663  if (st.kind === 'unknown') lines.push('- – **Cache:** no request yet this session, so its state is unknown')
664  if (st.kind === 'warm') lines.push(`- ● **Cache:** warm for about **${minutes(st.left)}** more · a rewrite would cost about ${usd(rewrite)}`)
665  if (st.kind === 'kept') lines.push(`- ● **Cache:** kept warm (see below) · a rewrite would cost about ${usd(rewrite)}`)
666  if (st.kind === 'cooling') lines.push(`- ◐ **Cache:** ⚠️ cools in **${minutes(st.left)}**, then the next message rewrites it for about **${usd(rewrite)}**`)
667  if (st.kind === 'cold') lines.push(`- ○ **Cache:** ⚠️ cold for ${minutes(-st.left)}, the next message rewrites it for about **${usd(rewrite)}**`)
668  if (H.text) lines.push(`- ✅ **Handoff ready:** \`/${names.handoff} continue\` clears this chat and carries on with it`)
669  if (S.keepWarm) lines.push(`- ◆ **Keep warm:** on until **${clock(S.keepWarmUntil)}**, ${S.pings} ping(s), ${usd(S.pingUsd)} so far`)
670  else lines.push(`- ◆ **Keep warm:** ${minutes(defaultKeepWarmMs())} by default on the ${ttlName(ttl)}, pays off up to **~${minutes(breakEvenMs())}** on ${modelName(S.model)}`)
671  if (S.rateLimits.length) lines.push(`- **Plan limits:** ${S.rateLimits.some((l) => (l.percentUsed || 0) >= 80) ? '⚠️ ' : ''}${limitsText(S.rateLimits)}`)
672  lines.push(`- **Cost:** session **${usd(S.costUsd)}**${S.todayUsd === null ? '' : ` · today **${todayText()}**`} (API list prices)`)
673  let restarts = `- **Cold restarts:** ${S.coldRestarts.length} (${usd(S.coldRestarts.reduce((a, r) => a + r.usd, 0))})`
674  if (S.cacheBreaks.length) restarts += ` · ⚠️ **Cache breaks:** ${S.cacheBreaks.length} (${usd(S.cacheBreaks.reduce((a, r) => a + r.usd, 0))}), last: ${S.cacheBreaks[S.cacheBreaks.length - 1].cause}`
675  lines.push(restarts)
676  lines.push(`- **Cold-send guard:** ${settings.guard ? 'on' : 'off'}, asks above **${bigText()}** · rewriting now ≈ ${usd(rewrite)}, ${isBig() ? 'above' : 'below'} that`)
677  lines.push(`- **Handoff hint:** ${settings.handoffAt ? `from **${tokens(settings.handoffAt)}** tokens${settings.handoffAt === HANDOFF_AT ? ' (default)' : ''}, then every ${tokens(HANDOFF_STEP)} more` : 'off'} · a handoff costs about **${usd(keepWarmVsHandoff(0).handoff)}** now (fresh chat ${tokens(S.startTokens || FRESH_START)}${S.startTokens ? '' : ' estimated'}, handoff ${tokens(handoffOut())}${S.handoffOuts.length ? `, average of the last ${S.handoffOuts.length}` : ' estimated'})`)
678  lines.push(`- **Alerts:** ${settings.alerts ? 'on' : 'off'}`)
679  lines.push(`- **Settings:** \`/${c} ttl 5|60|auto\` · \`/${c} guard on|off\` · \`/${c} big $1|default\` · \`/${c} handoff 140k|off|default\` · \`/${c} alerts on|off\``)
680  return lines.join('\n')
681}
682
683async function registerCommand($, name, description, argumentHint, immediate) {
684  const spec = immediate ? { name, description, argumentHint, immediate: true } : { name, description, argumentHint }
685  try {
686    await $.command.register(spec)
687    return name
688  } catch {
689    try {
690      await $.command.register({ ...spec, name: 'tk-' + name })
691      return 'tk-' + name
692    } catch {
693      return null
694    }
695  }
696}
697
698// The guard's "Compact first, then send", after its prompt.submit hook dropped
699// the message. Whatever fails, the message goes back in the prompt box unsent.
700async function compactThenSend($, held) {
701  let why = ''
702  try {
703    const r = await $.session.compact()
704    if (r && r.skip) why = r.skip
705  } catch (err) {
706    why = String((err && err.message) || err || 'it was refused')
707  }
708  if (why) {
709    log($, `guard: compact failed: ${why}`)
710    await $.prompt.fill({ text: held.text })
711    return warn($, `Not sent: compacting didn't work (${why}). Your message is back in the prompt box.`)
712  }
713  // Pasted images and @file mentions can't be sent again by a plugin
714  if (held.attached) {
715    await $.prompt.fill({ text: held.text })
716    return note($, 'Compacted. Your message is back in the prompt box: attach any images again, then send it.')
717  }
718  await $.prompt.submit({ text: held.text, asUser: true })
719}
720
721export function register(on) {
722  on('session.start', async ($, e, next) => {
723    now = await $.clock.now()
724    home = (await $.env.get('USERPROFILE')) || (await $.env.get('HOME')) || ''
725    S.id = await $.session.id()
726    S.cwd = await $.session.cwd()
727    S.model = await $.session.model()
728    await loadSettings($)
729    // Remember the variable as found, unless it is the value /cache ttl set (a reload)
730    const prior = await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL')
731    envPrior = settings.ttlMin && prior === ttlEnv(settings.ttlMin) ? undefined : prior
732    if (settings.ttlMin) await applyTtl($)
733    const outs = await $.store.get('handoffOuts')
734    S.handoffOuts = Array.isArray(outs) ? outs.filter((n) => n > 0).slice(-5) : []
735    measureStart($).catch(() => {})
736    names.cache = (await registerCommand($, 'cache', 'Token Keeper status and settings', '[ttl 5|60|auto] [guard on|off] [big $1|default] [handoff 140k|off|default] [alerts on|off]')) || names.cache
737    names.keepwarm = (await registerCommand($, 'keepwarm', 'Keep this session\'s prompt cache warm (default 30m on the 5-min TTL, 4h on the 1h TTL), or /keepwarm off', '[30m|4h|off]', true)) || names.keepwarm
738    names.handoff = (await registerCommand($, 'handoff', 'Session handoff, then clear this chat and continue with it (/handoff continue)', '[continue]')) || names.handoff
739    $.clock.every(TICK_EVERY, () => tick($).catch(() => {}))
740    refreshToday($).catch(() => {})
741    return next(e)
742  })
743
744  // /clear, /resume and /branch start from an unknown cache. The process goes
745  // on under a new session id, and no session.start fires for it.
746  on('classic.SessionStart', { source: ['clear', 'resume', 'fork'] }, async ($, e, next) => {
747    S.lastActivity = 0
748    S.keepWarm = false
749    H.armed = false
750    try {
751      S.id = (await $.session.id()) || S.id
752    } catch {
753      // keep the old id
754    }
755    // The old transcript is another session's now: read today's costs afresh
756    own.scanned = false
757    todayAt = 0
758    return next(e)
759  })
760
761  on('prompt.submit', async ($, e, next) => {
762    now = await $.clock.now()
763    // A new message in this chat makes a waiting handoff out of date (its file stays)
764    if (H.text && !H.continuing && e.origin && e.origin.kind === 'composer' && !String(e.text || '').trim().startsWith('/')) {
765      H.text = ''
766      H.path = ''
767      $.ui.invalidate('ui.render')
768    }
769    const ttl = ttlMin()
770    const left = msLeft(S.lastActivity, ttl)
771    const isCold = left !== null && left <= 0
772    const fromUser = e.origin && (e.origin.kind === 'composer' || e.origin.kind === 'bridge')
773    if (!settings.guard || !fromUser || !isCold || !isBig() || e.turnId) return next(e)
774    const cost = usd(rewriteCost(S.ctx, S.model, ttl))
775    let answer = 'Send anyway'
776    try {
777      answer = await $.ui.ask(
778        `Cache went cold ${minutes(-left)} ago. Sending now rewrites ${tokens(S.ctx)} tokens of context (about ${cost}). What should happen?`,
779        ['Send anyway', 'Compact first, then send', 'Cancel'],
780      )
781    } catch {
782      // nobody to ask (a -p run, or the dialog was dismissed): send as typed
783      return next(e)
784    }
785    if (answer === 'Compact first, then send') {
786      // Claude Code refuses compact() from a prompt.submit hook (it would
787      // compact under the turn this hook holds): hold the message, compact
788      // once this hook has returned, then send it
789      // A plugin's prompt expands no pasted images and no @file mentions
790      const attached = !!(e.attachments && e.attachments.length) || /(^|\s)@\S/.test(e.text)
791      const held = { text: e.text, attached }
792      $.clock.after(0, () => { compactThenSend($, held).catch(() => {}) })
793      return { drop: 'Held by token-keeper: compacting first, then sending it' }
794    }
795    if (answer === 'Send anyway') return next(e)
796    // A handoff now would rewrite the cache too: it only saves while warm
797    warn($, `Not sent. Next time, /${names.handoff} before the cache goes cold skips the rewrite.`)
798    return { drop: 'Cancelled by token-keeper before a cold cache rewrite' }
799  })
800
801  on('turn.start', async ($, e, next) => {
802    S.working = true
803    if (e.turnId) S.turnId = e.turnId
804    return next(e)
805  })
806
807  // Each request: what it cost today, and for the main loop, did it read the
808  // cache or rewrite it?
809  on('turn.step', async function* ($, e, next) {
810    const startedAt = await $.clock.now()
811    const usdBefore = e.agentId ? null : await sessionUsd($)
812    const result = yield* next(e)
813    if (!result || !result.usage) return result
814    const u = result.usage
815    addToday(requestCost(u, u.model || S.model, ttlMin()))
816    if (e.agentId) return result
817    now = await $.clock.now()
818    const total = totalInput(u)
819    const written = u.cache_creation_input_tokens || 0
820    // Each write's TTL is the cache window, measured
821    const usdAfter = usdBefore === null ? null : await sessionUsd($)
822    const seen = writeTtl(u, written, usdAfter === null ? null : usdAfter - usdBefore)
823    if (seen) {
824      S.ttlWritten = seen
825      if (!settings.ttlMin) {
826        S.ttlMin = S.ttlWritten
827        S.ttlSource = 'measured'
828      } else if (S.ttlWritten !== settings.ttlMin && !S.ttlWarned) {
829        S.ttlWarned = true
830        if (settings.alerts) warn($, `You set the ${settings.ttlMin}-min cache TTL, but Claude Code still writes ${S.ttlWritten}-min entries. FORCE_PROMPT_CACHING_5M overrides it, and setting it needs Claude Code v2.1.242 or later.`)
831      }
832    }
833    const gap = S.lastActivity ? startedAt - S.lastActivity : 0
834    const prevModel = S.model
835    S.model = u.model || S.model
836    const afterCompact = justCompacted
837    justCompacted = false
838    const afterTtlSwitch = ttlSwitched
839    ttlSwitched = false
840    // What the cache lost: the part of the previous context this request did not read.
841    // New content (a file read, a tool result) is written too, but that is no loss
842    // (against the last request's own size: S.ctx may already count the new content)
843    const lost = Math.max(0, S.lastTotal - (u.cache_read_input_tokens || 0))
844    const lostMuch = S.lastActivity && S.lastTotal > 30000 && lost > 0.2 * S.lastTotal
845    if (afterCompact) {
846      // the first request after a compaction writes the new, shorter context: expected
847    } else if (afterTtlSwitch) {
848      // a /cache ttl switch while warm: the rewrite is expected, so say what it really cost
849      if (written > 0) note($, `Switching to the ${ttlName(ttlMin())} rewrote ${tokens(written)} tokens for about ${usd(rewriteCost(written, S.model, ttlMin()))}.`)
850    } else if (lostMuch && gap < 4.5 * MIN) {
851      // Lost although the cache was still warm: something changed the prompt prefix
852      const cause = prevModel && u.model && prevModel !== u.model ? `the model changed (${priceFor(prevModel).id} → ${priceFor(u.model).id})` : 'the prompt prefix changed (CLAUDE.md, MCP tools, settings, effort, or system prompt)'
853      const cost = rewriteCost(lost, S.model, ttlMin())
854      S.cacheBreaks.push({ at: now, tokens: lost, usd: cost, cause })
855      if (settings.alerts) warn($, `Cache broken while warm: ${cause}. Rewrote ${tokens(lost)} tokens for about ${usd(cost)}.`)
856    } else if (lostMuch) {
857      // Cold: a fifth or more of the previous context was not read after a pause (the
858      // system prompt part often stays warm through other sessions).
859      // A rewrite after 5-60 idle minutes means this session runs on the 5-minute TTL
860      if (!settings.ttlMin && gap > 5.5 * MIN && gap < S.ttlMin * MIN) {
861        S.ttlMin = 5
862        S.ttlSource = 'measured'
863      }
864      S.coldRestarts.push({ at: now, tokens: lost, usd: rewriteCost(lost, S.model, ttlMin()), gapMin: Math.round(gap / MIN) })
865    } else if (S.lastActivity && total > 30000 && gap > 5.5 * MIN && cachedShare(u) > 0.8) {
866      // A hit after more than 5 idle minutes proves the 1-hour TTL
867      if (!settings.ttlMin) {
868        S.ttlMin = 60
869        S.ttlSource = 'measured'
870      }
871    }
872    if (H.armed) H.outTokens += u.output_tokens || 0
873    S.ctx = total
874    S.lastTotal = total
875    S.lastActivity = now
876    S.pingsSinceTurn = 0
877    if (e.effort) S.effort = String(e.effort)
878    return result
879  })
880
881  on('turn.complete', async ($, e, next) => {
882    const r = await next(e)
883    if (e.agentId) return r
884    now = await $.clock.now()
885    S.working = false
886    try {
887      const usage = await $.session.usage()
888      if (usage.context && usage.context.tokens) S.ctx = usage.context.tokens
889      if (usage.context && usage.context.window) S.window = usage.context.window
890      if (usage.cost) S.costUsd = usage.cost.usd
891      if (Array.isArray(usage.rateLimits) && usage.rateLimits.length) S.rateLimits = usage.rateLimits
892    } catch {
893      // usage unavailable: keep the per-request figures
894    }
895    if (H.armed && e.turnId !== H.notTurn) await captureHandoff($, e)
896    else handoffStep($)
897    $.ui.invalidate('ui.render')
898    return r
899  })
900
901  on('session.compact', async ($, e, next) => {
902    const r = await next(e)
903    justCompacted = true
904    return r
905  })
906
907  on('session.measure', async ($, e, next) => {
908    if (e.context && e.context.tokens) S.ctx = e.context.tokens
909    if (e.context && e.context.window) S.window = e.context.window
910    if (e.cost) S.costUsd = e.cost.usd
911    if (Array.isArray(e.rateLimits) && e.rateLimits.length) S.rateLimits = e.rateLimits
912    return next(e)
913  })
914
915  // Claude calling the handoff skill itself: this turn's answer is the handoff
916  on('tool.call', { tool: 'Skill' }, async ($, e, next) => {
917    if (!e.agentId && HANDOFF_SKILL.test(String(e.skill || ''))) armHandoff(true)
918    return next(e)
919  })
920
921  // /session-handoff from anywhere (typed, the band, /handoff): watch for its answer
922  on('command.run', async ($, e, next) => {
923    if (HANDOFF_SKILL.test(String(e.command || ''))) armHandoff(false)
924    return next(e)
925  })
926
927  // /handoff runs the handoff and then asks to clear and continue (the band's
928  // buttons draw only in the terminal); /handoff continue does the second half.
929  // Both run off the hook: a command started inside a hook the session is
930  // waiting on is refused.
931  on('command.run', { command: ['handoff', 'tk-handoff'] }, async ($, e) => {
932    now = await $.clock.now()
933    const arg = String(e.args || '').trim().toLowerCase()
934    if (arg === 'continue' || arg === 'go') {
935      if (!H.text) return { text: `No handoff ready yet. Type /${names.handoff} to make one.` }
936      $.clock.after(50, () => clearAndContinue($).catch(() => {}))
937      return { text: 'Clearing this chat and continuing with the handoff.' }
938    }
939    H.askAfter = true
940    $.clock.after(50, () => runHandoff($).catch(() => {}))
941    return { text: 'Running the session handoff. When it finishes, you can clear this chat and continue with it.' }
942  })
943
944  on('command.run', { command: ['cache', 'tk-cache'] }, async ($, e) => {
945    now = await $.clock.now()
946    const [key, value] = String(e.args || '').trim().toLowerCase().split(/\s+/)
947    if (key === 'ttl') {
948      const want = value === '5' ? 5 : value === '60' ? 60 : 0
949      // While the cache is warm, the next request rewrites what was cached under
950      // the old TTL; once it is cold that rewrite happens anyway, so no question
951      const st = cacheState()
952      if (want && want !== ttlMin() && ['warm', 'cooling', 'kept'].includes(st.kind) && S.ctx > 30000) {
953        const cost = usd(rewriteCost(S.ctx, S.model, want))
954        const keep = `Keep the ${ttlName(ttlMin())}`
955        let answer = keep
956        try {
957          answer = await $.ui.ask(`The cache is warm. Switching to the ${ttlName(want)} makes your next message rewrite up to ${tokens(S.ctx)} tokens (about ${cost}). Once the cache is cold, switching costs nothing extra. Switch now?`, ['Switch now', keep])
958        } catch {
959          // dismissed, or nobody to ask: keep the TTL
960        }
961        if (answer !== 'Switch now') return { text: `Cache TTL unchanged: ${ttlName(ttlMin())}. Switching now would rewrite up to ${tokens(S.ctx)} tokens (about ${cost}); after the cache goes cold it is free.` }
962      }
963      if (want !== settings.ttlMin && ['warm', 'cooling', 'kept'].includes(st.kind)) ttlSwitched = true
964      settings.ttlMin = want
965      S.ttlWritten = 0
966      S.ttlWarned = false
967      await applyTtl($)
968    } else if (key === 'guard' || key === 'alerts') {
969      settings[key] = value !== 'off'
970    } else if (key === 'big') {
971      const n = value === 'default' ? BIG_USD : parseUsd(value)
972      if (!n) return { text: `Usage: \`/${names.cache} big $1\` (a dollar amount), or \`/${names.cache} big default\`.` }
973      settings.bigUsd = n
974    } else if (key === 'handoff') {
975      const n = value === 'default' ? HANDOFF_AT : value === 'off' ? 0 : parseTokens(value)
976      if (n === null) return { text: `Usage: \`/${names.cache} handoff 140k\` (a context size in tokens), \`off\`, or \`default\`.` }
977      settings.handoffAt = n
978      S.handoffNextAt = 0
979    }
980    if (key) {
981      await $.store.set('settings', settings)
982      $.ui.invalidate('ui.render')
983    }
984    return { text: statusText() }
985  })
986
987  on('command.run', { command: ['keepwarm', 'tk-keepwarm'] }, async ($, e) => {
988    now = await $.clock.now()
989    const arg = String(e.args || '').trim().toLowerCase()
990    let text
991    if (arg === 'off' || (arg === '' && S.keepWarm)) {
992      text = stopKeepWarm('turned off') || 'ℹ️ Keep warm is already off.'
993    } else if (arg === '') {
994      await measureStart($)
995      text = startKeepWarm(defaultKeepWarmMs(), true, true)
996    } else {
997      const ms = parseDuration(arg)
998      if (!ms) return { text: `Usage: \`/${names.keepwarm} 30m\`, \`/${names.keepwarm} 4h\`, or \`/${names.keepwarm} off\`.` }
999      await measureStart($)
1000      text = startKeepWarm(ms, false, true)
1001    }
1002    $.ui.invalidate('ui.render')
1003    return { text }
1004  })
1005
1006  // The band above the prompt: this session's cache at a glance
1007  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
1008    const below = await next(e)
1009    if (e.props && e.props.hasSurvey) return below
1010    if (!S.lastActivity && !S.keepWarm) return below
1011    const mine = drawBand($, e)
1012    const { Box } = $.ui.resolve(e)
1013    return Box({ flexDirection: 'column', children: below ? [mine, below] : [mine] })
1014  })
1015
1016  // The band is terminal-only. The footer also draws in the Desktop app, so it
1017  // carries a short label, only when there's something to act on.
1018  on('ui.render', { component: 'SessionMode' }, async ($, e, next) => {
1019    const label = footerLabel()
1020    if (!label) return next(e)
1021    const modes = Array.isArray(e.props && e.props.modes) ? e.props.modes : []
1022    return next({ ...e, props: { ...e.props, modes: [...modes, label] } })
1023  })
1024}
1025
1026function drawBand($, e) {
1027  const { Box, Text, Button } = $.ui.resolve(e)
1028  const st = cacheState()
1029  const ttl = ttlMin()
1030  const big = isBig()
1031  const rewrite = usd(rewriteCost(S.ctx, S.model, ttl))
1032  const parts = []
1033  if (st.kind === 'kept') parts.push(Text({ color: 'cyan', children: [`◆ kept warm · ${S.pings} ping${S.pings === 1 ? '' : 's'} ${usd(S.pingUsd)} · until ${clock(S.keepWarmUntil)}`] }))
1034  else if (st.kind === 'warm') parts.push(Text({ color: 'green', children: [`● cache warm ${minutes(st.left)}`] }))
1035  else if (st.kind === 'cooling') parts.push(Text({ color: 'yellow', bold: true, children: [`◐ cache cools in ${minutes(st.left)}`] }))
1036  else if (st.kind === 'cold') parts.push(Text(big ? { color: 'red', bold: true, children: [`○ cache cold ${minutes(-st.left)}`] } : { dimColor: true, children: [`○ cache cold ${minutes(-st.left)}`] }))
1037  const pct = S.window ? Math.floor((S.ctx / S.window) * 100) : null
1038  const ctxText = ` │ ctx ${tokens(S.ctx)}${pct === null ? '' : `/${tokens(S.window)} ${pct}%`}`
1039  parts.push(Text(pct >= 80 ? { color: 'red', children: [ctxText] } : pct >= 50 ? { color: 'yellow', children: [ctxText] } : { dimColor: true, children: [ctxText] }))
1040  if (st.kind === 'cold' && big) parts.push(Text({ color: 'red', children: [` │ next send rewrites it ≈ ${rewrite}`] }))
1041  else parts.push(Text({ dimColor: true, children: [` │ rewrite ≈ ${rewrite}`] }))
1042  if (S.rateLimits.length) parts.push(Text(limitTone(S.rateLimits, { children: [' │ ' + limitsText(S.rateLimits)] })))
1043  if (S.coldRestarts.length) parts.push(Text({ dimColor: true, children: [` │ ${S.coldRestarts.length} cold restart${S.coldRestarts.length === 1 ? '' : 's'} ${usd(S.coldRestarts.reduce((a, c) => a + c.usd, 0))}`] }))
1044  if (S.cacheBreaks.length) parts.push(Text({ color: 'yellow', children: [` │ ${S.cacheBreaks.length} cache break${S.cacheBreaks.length === 1 ? '' : 's'} ${usd(S.cacheBreaks.reduce((a, c) => a + c.usd, 0))}`] }))
1045  const row = [Box({ flexDirection: 'row', children: parts })]
1046  if (st.kind === 'cooling' && big) {
1047    row.push(Button({ key: 'keepwarm', label: 'keep warm', hotkey: '1', plain: true, onPress: async () => { now = await $.clock.now(); log($, startKeepWarm(defaultKeepWarmMs(), true, false)); $.ui.invalidate('ui.render') } }))
1048  } else if (st.kind === 'kept') {
1049    row.push(Button({ key: 'keepwarm', label: 'stop warm', hotkey: '1', plain: true, onPress: async () => { now = await $.clock.now(); log($, stopKeepWarm('turned off')); $.ui.invalidate('ui.render') } }))
1050  }
1051  const left = Box({ flexDirection: 'row', columnGap: 2, children: row })
1052  // Right edge: one click runs /session-handoff (h while the band has the focus;
1053  // a letter never fires from the prompt box, unlike a digit). Once a handoff
1054  // is in, c clears this chat and sends it as the fresh chat's first prompt.
1055  const busy = handoffPending || H.armed
1056  const handoff = Button({ key: 'handoff', label: handoffPending ? 'handoff queued' : H.armed ? 'handoff running' : 'handoff', hotkey: 'h', variant: 'primary', dimColor: busy, onPress: () => runHandoff($) })
1057  let right = handoff
1058  if (H.text || H.continuing) {
1059    const go = Button({ key: 'continue', label: H.continuing ? 'clearing' : 'clear and continue', hotkey: 'c', plain: true, onPress: () => clearAndContinue($) })
1060    right = Box({ flexDirection: 'row', columnGap: 2, children: [go, handoff] })
1061  }
1062  const top = Box({ flexDirection: 'row', justifyContent: 'space-between', width: '100%', columnGap: 2, children: [left, right] })
1063  // Second row: the model and what it costs
1064  const info = [Text({ color: 'magenta', children: [modelName(S.model)] })]
1065  if (S.effort) info.push(Text({ dimColor: true, children: [` · ${S.effort}`] }))
1066  info.push(Text({ dimColor: true, children: [` │ session ${usd(S.costUsd)}`] }))
1067  if (S.todayUsd !== null) info.push(Text({ dimColor: true, children: [` · today ${todayText()}`] }))
1068  return Box({ flexDirection: 'column', width: '100%', children: [top, Box({ flexDirection: 'row', children: info })] })
1069}
1070
1071function footerLabel() {
1072  const parts = []
1073  const st = cacheState()
1074  const big = isBig()
1075  if (st.kind === 'kept') parts.push('cache kept warm')
1076  else if (st.kind === 'cooling' && big) parts.push(`cache cools in ${minutes(st.left)} · /${names.keepwarm} or /${names.handoff}`)
1077  else if (st.kind === 'cold' && big) parts.push(`cache cold · rewrite ≈ ${usd(rewriteCost(S.ctx, S.model, ttlMin()))}`)
1078  const high = S.rateLimits.filter((l) => (l.percentUsed || 0) >= 80)
1079  if (high.length) parts.push(limitsText(high))
1080  if (H.text) parts.push(`handoff ready · /${names.handoff} continue`)
1081  return parts.join(' · ')
1082}
1083
hooks/pricing.mjs 92 lines
1// Shared source. Copied into each mod's hooks/ folder by _dev/sync-shared.mjs.
2// Edit it here, then run: node mods/_dev/sync-shared.mjs
3
4// Dollars per million tokens, from the Claude API pricing table (cached 2026-09-25).
5// Cache writes cost 1.25x input on the 5-minute TTL and 2x input on the 1-hour TTL.
6const TABLE = [
7  ['fable-5-1', { input: 10, output: 50, read: 0.25 }],
8  ['mythos-5-1', { input: 10, output: 50, read: 0.25 }],
9  ['fable-5', { input: 10, output: 50, read: 1.0 }],
10  ['opus-5-5', { input: 4, output: 20, read: 0.2 }],
11  ['opus-5', { input: 5, output: 25, read: 0.5 }],
12  ['opus-4-8', { input: 5, output: 25, read: 0.5 }],
13  ['opus-4-7', { input: 5, output: 25, read: 0.5 }],
14  ['opus-4-6', { input: 5, output: 25, read: 0.5 }],
15  ['sonnet-5-5', { input: 2, output: 10, read: 0.2 }],
16  ['sonnet-5', { input: 2, output: 10, read: 0.2 }],
17  ['sonnet-4-6', { input: 3, output: 15, read: 0.3 }],
18  ['haiku-4-5', { input: 1, output: 5, read: 0.1 }],
19]
20
21const FAMILY = {
22  fable: 'fable-5-1',
23  mythos: 'mythos-5-1',
24  opus: 'opus-5-5',
25  sonnet: 'sonnet-5-5',
26  haiku: 'haiku-4-5',
27}
28
29export function normalizeModel(model) {
30  return String(model || '')
31    .toLowerCase()
32    .replace(/^claude-/, '')
33    .replace(/\[.*?\]/g, '')
34    .replace(/-\d{8}$/, '')
35    .trim()
36}
37
38export function priceFor(model) {
39  const id = normalizeModel(model)
40  for (const [key, price] of TABLE) {
41    if (id === key || id.startsWith(key + '-') || id.startsWith(key)) {
42      // 'opus-5' must not swallow 'opus-5-5': the table lists longer ids first
43      return { id: key, ...price }
44    }
45  }
46  for (const [family, key] of Object.entries(FAMILY)) {
47    if (id.includes(family)) {
48      const found = TABLE.find(([k]) => k === key)
49      return { id: key, ...found[1] }
50    }
51  }
52  return { id: 'opus-5-5', input: 4, output: 20, read: 0.2 }
53}
54
55export function writeRate(model, ttlMinutes) {
56  const p = priceFor(model)
57  return p.input * (ttlMinutes >= 60 ? 2 : 1.25)
58}
59
60// usage: { input_tokens, output_tokens, cache_read_input_tokens, cache_creation_input_tokens }
61export function requestCost(usage, model, ttlMinutes = 60) {
62  if (!usage) return 0
63  const p = priceFor(model || usage.model)
64  const w = writeRate(model || usage.model, ttlMinutes)
65  return (
66    ((usage.input_tokens || 0) * p.input +
67      (usage.cache_read_input_tokens || 0) * p.read +
68      (usage.cache_creation_input_tokens || 0) * w +
69      (usage.output_tokens || 0) * p.output) /
70    1e6
71  )
72}
73
74export function rewriteCost(tokens, model, ttlMinutes = 60) {
75  return ((tokens || 0) * writeRate(model, ttlMinutes)) / 1e6
76}
77
78export function totalInput(usage) {
79  if (!usage) return 0
80  return (
81    (usage.input_tokens || 0) +
82    (usage.cache_read_input_tokens || 0) +
83    (usage.cache_creation_input_tokens || 0)
84  )
85}
86
87// Share of a request's input served from the cache, 0 to 1.
88export function cachedShare(usage) {
89  const total = totalInput(usage)
90  return total > 0 ? (usage.cache_read_input_tokens || 0) / total : 0
91}
92
hooks/fmt.mjs 41 lines
1// Shared source. Copied into each mod's hooks/ folder by _dev/sync-shared.mjs.
2
3export function tokens(n) {
4  const v = Number(n) || 0
5  if (v >= 1e6) return (v / 1e6).toFixed(v >= 1e7 ? 0 : 1) + 'M'
6  if (v >= 1e3) return (v / 1e3).toFixed(v >= 1e5 ? 0 : 1) + 'k'
7  return String(Math.round(v))
8}
9
10export function usd(n) {
11  const v = Number(n) || 0
12  if (v === 0) return '$0'
13  if (v < 0.01) return '<$0.01'
14  if (v < 10) return '$' + v.toFixed(2)
15  return '$' + v.toFixed(0)
16}
17
18export function minutes(ms) {
19  const m = Math.round((Number(ms) || 0) / 60000)
20  if (m <= 60) return m + 'm'
21  return Math.floor(m / 60) + 'h' + String(m % 60).padStart(2, '0') + 'm'
22}
23
24export function clock(ts) {
25  const d = new Date(ts)
26  let h = d.getHours()
27  const ampm = h >= 12 ? 'pm' : 'am'
28  h = h % 12 || 12
29  return h + ':' + String(d.getMinutes()).padStart(2, '0') + ampm
30}
31
32export function clip(text, max) {
33  const s = String(text ?? '').replace(/\s+/g, ' ').trim()
34  return s.length > max ? s.slice(0, Math.max(0, max - 1)) + '…' : s
35}
36
37export function basename(path) {
38  const parts = String(path || '').split(/[\\/]+/).filter(Boolean)
39  return parts[parts.length - 1] || String(path || '')
40}
41