SLOPSHOPPER

cache-tax

A Cache panel: a countdown gauge for the prompt cache, a hit-rate sparkline, cold-cache warnings and /keepwarm

newpanecommandtoaststatusprompt
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · cache-tax
│ ┃ Cache ✕ › fix the failing auth test and add an audit log call │ ┃ ❄ PROMPT CACHE · EMPTY │ ┃ ──────────────────────────────────────────── ⏺ Read(src/auth.ts) │ ┃ ────────── ⎿ Read 6 lines │ ┃ Nothing cached yet. The gauge starts with ⏺ Update(src/auth.ts) │ ┃ the first response. ⎿ Added 2 lines, removed 1 line │ ┃ ⏺ Bash(bun test) │ ┃ hits · 0% read from cache ⎿ 3 pass, 1 fail │ ┃ saved 0 full-price input tokens │ ┃ COST: calibrating against /cost, a few more ● Done. refresh now rejects expired claims and logs an audit event. │ ┃ requests needed. │ ┃ ✻ Worked for 42s · done 4:20 PM │ ┃ keep warm ON pings 0 │ ┃ [ Stop keeping warm ] [ Ping now ] [ Window: › /keepwarm │ ⎿ cache-tax: Keeping the cache warm: a one-line ping 60s before ea │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · Cache
❄ PROMPT CACHE · EMPTY ────────────────────────────────────────────────────── Nothing cached yet. The gauge starts with the first response. hits · 0% read from cache saved 0 full-price input tokens COST: calibrating against /cost, a few more requests needed. keep warm ON pings 0 [ Stop keeping warm ] [ Ping now ] [ Window: 1h ]
README

Session dashboard for Claude Code

Five side-panel mods for Claude Code (function hooks, early access API):

ModPanelCommand
agent-progressCrew: each sub-agent as a little Claude crab dressed for its job (builder with a hammer, inspector with a monocle, scout with a compass, architect with a set square, librarian with a book), model, live activity, steps, elapsed time, cache hit bar, tokens, and a warning when a worker goes quiet/crew
cache-taxCache: prompt cache countdown gauge, hit-rate sparkline, cold-start warnings, keep-warm pings, and what caching cost: spent, saved by reads, the cold-start tax, what pings cost and what going cold now would add (estimated from /cost with the published cache multipliers)/keepwarm, /keepwarm now, `/keepwarm ttl 5\60, /keepwarm panel`
spend-ledgerSpend: what each prompt and each sub-agent cost, a STATS dashboard modelled on /usage stats: views Overview, Tokens, Cost, Cache and In/Out; ranges This session, 7 days, 30 days and All time; a model filter; an activity heatmap, charts in four styles (line, area, bars, dots) with axes, stat tiles (favorite model, total tokens, sessions, active days, longest session, streaks) and per-model cards. Past days come from Claude Code's own history (~/.claude/stats-cache.json, the file /usage reads), with their in/out/cache split and cost estimated from each model's totals; today is exact, from the mod's own tracking. An active-hours chart shows when you work, a by-model breakdown (cost, share, requests, tokens, cache hit rate), burn rate, budget, running insights. On an API key it budgets in dollars (/spend budget 50, warnings at 80% and 100%). On a Claude subscription it shows the plan's 5-hour and weekly windows with reset times and how long the 5-hour window lasts at your pace, and budgets as a share of it (/spend budget 80%); the dollar figures are then the API-price value of the work, not a charge/spend, /spend budget 50, /spend budget 80%, add hard to refuse new prompts once reached (/spend budget 50 hard), `/spend hard on\off, /spend budget off`
session-linkSession Link: the live Claude sessions on this machine, each with its own label (the session sets it itself early in its first turn; until then its first prompt stands in; /link name overrides) and its own Claude crab badge (a pose and colour no two live sessions share), its directory, branch and the files it edited; the messages the sessions send each other; files two sessions both edited. It also stops one session overwriting another: an edit to a file another live session edited in the last six hours, or a broad git command (add -A, commit -a, reset --hard, checkout, switch, stash, rebase, merge, pull) in a repository another session works in, is refused once with the other session's name so the model messages it first; the same call again goes through/link, /link name A: finish UA, /link release; the model's set_session_label tool
mission-controlMission Control: all four stacked in one scrollable panel, with buttons for every command/mission

Every panel has buttons with hotkeys for its commands.

Install

At the prompt of a terminal session:

/plugin install mission-control --marketplace kimathinjoki/claude-session-dashboard

Answer y to add the marketplace, choose the user scope, then install the others the same way (agent-progress, cache-tax, spend-ledger, session-link). Mission Control reads their state, so it needs all four.

Best viewed

Side panels dock beside the transcript in fullscreen mode ("tui": "fullscreen" in ~/.claude/settings.json, or /config) from 110 columns. They open on their own at session start from 144 columns; at any width the commands above open them. With several open they share the dock as tabs.

The Cache panel follows your promptCacheTtl setting ("1h" keeps the cache for an hour; 1-hour cache writes cost more than 5-minute ones).

Source 2 files
hooks/register.tsx 346 lines
1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { CacheReading } from '../types'
5
6const initial: CacheReading = {
7  lastRequestAt: null, cachedTokens: 0, ttlMinutes: 60, keepWarm: false, pings: 0, lastPingAt: null,
8  history: [], readTokens: 0, writtenTokens: 0, coldStarts: 0,
9  rewrittenTokens: 0, pingReadTokens: 0, weightedAll: 0, sessionUsd: null, usdBaseline: null, mainModel: '',
10}
11const reading = atom({ plugin: 'cache-tax', key: 'reading' } as const, initial)
12const tick = atom({ plugin: 'cache-tax', key: 'tick' } as const, 0)
13
14const PANE = 'cache'
15
16// State saved by an earlier version of this mod may lack fields added since; fill them in.
17async function readReading($: EngineInterface): Promise<CacheReading> {
18  return { ...initial, ...(await read($, reading)) }
19}
20// Anthropic's published multipliers on the base input price: a cache write costs 1.25x
21// (2x for the one-hour cache), a cache read 0.1x. A cold prefix is re-written, not read.
22const READ = 0.1
23const writeMultiplier = (ttlMinutes: number) => (ttlMinutes >= 60 ? 2 : 1.25)
24const MARGIN_MS = 60_000
25const TICK_MS = 15_000
26const HISTORY = 24
27const SPARK = ['▁', '▂', '▃', '▄', '▅', '▆', '▇', '█']
28
29// Relative list prices by model family, as the Spend panel uses them.
30const MODEL_FACTOR: Array<[RegExp, number]> = [[/opus/i, 5], [/sonnet/i, 3], [/haiku/i, 1], [/fable/i, 5]]
31export const modelFactor = (model: string) => MODEL_FACTOR.find(([test]) => test.test(model))?.[1] ?? 3
32
33// A request in units of "base input tokens", so /cost divided by the sum prices one.
34export const weightOf = (
35  u: { input_tokens: number; output_tokens: number; cache_read_input_tokens: number; cache_creation_input_tokens: number },
36  model: string,
37  ttlMinutes: number,
38) =>
39  (u.input_tokens + writeMultiplier(ttlMinutes) * u.cache_creation_input_tokens + READ * u.cache_read_input_tokens + 5 * u.output_tokens) *
40  modelFactor(model)
41
42// The main thread's base input price per token, estimated from what the session actually cost.
43// Below this many weighted units the estimate is noise, and the panel says it is calibrating.
44export const MIN_WEIGHT = 200_000
45
46export const basePrice = (r: Pick<CacheReading, 'sessionUsd' | 'usdBaseline' | 'weightedAll' | 'mainModel'>) => {
47  if (r.sessionUsd == null || r.usdBaseline == null || r.weightedAll < MIN_WEIGHT) return null
48  const spent = r.sessionUsd - r.usdBaseline
49  return spent > 0 ? (spent / r.weightedAll) * modelFactor(r.mainModel || 'opus') : null
50}
51
52export type CacheCosts = { spent: number; saved: number; coldTax: number; pings: number; ifColdNow: number }
53
54export const cacheCosts = (r: CacheReading): CacheCosts | null => {
55  const base = basePrice(r)
56  if (base === null) return null
57  const write = writeMultiplier(r.ttlMinutes)
58  return {
59    spent: (r.writtenTokens * write + r.readTokens * READ) * base,
60    saved: r.readTokens * (1 - READ) * base,
61    coldTax: r.rewrittenTokens * (write - READ) * base,
62    pings: r.pingReadTokens * READ * base,
63    ifColdNow: r.cachedTokens * (write - READ) * base,
64  }
65}
66
67export const money = (n: number) => (n >= 100 ? `$${n.toFixed(0)}` : n >= 1 ? `$${n.toFixed(2)}` : `${Math.max(0, Math.round(n * 100))}¢`)
68
69export const tokens = (n: number) =>
70  n >= 1_000_000 ? `${(n / 1_000_000).toFixed(1)}M` : n >= 1000 ? `${Math.round(n / 1000)}k` : `${n}`
71
72export const remainingMs = (r: CacheReading, now: number) =>
73  r.lastRequestAt === null ? null : r.lastRequestAt + r.ttlMinutes * 60_000 - now
74
75export const coldFactor = (ttlMinutes: number) => Math.round((writeMultiplier(ttlMinutes) / READ) * 10) / 10
76
77export const clock = (ms: number) => {
78  const s = Math.max(0, Math.floor(ms / 1000))
79  return `${Math.floor(s / 60)}:${String(s % 60).padStart(2, '0')}`
80}
81
82export const sparkline = (values: number[]) =>
83  values.map(v => SPARK[Math.max(0, Math.min(SPARK.length - 1, Math.round(v * (SPARK.length - 1))))]).join('')
84
85// Warm, fading (under a quarter of the window), or cold.
86export const mood = (r: CacheReading, now: number) => {
87  const left = remainingMs(r, now)
88  if (left === null || r.cachedTokens === 0) return 'empty' as const
89  if (left <= 0) return 'cold' as const
90  return left < r.ttlMinutes * 60_000 * 0.25 ? ('fading' as const) : ('warm' as const)
91}
92
93export const statusText = (r: CacheReading, now: number) => {
94  const left = remainingMs(r, now)
95  if (left === null || r.cachedTokens === 0) return undefined
96  const warmth = r.keepWarm ? ', kept warm' : ''
97  if (left > 0) return `cache warm ${clock(left)} left (${tokens(r.cachedTokens)})${warmth}`
98  return `cache cold: next prompt re-writes ${tokens(r.cachedTokens)} tokens`
99}
100
101let pinging = false
102
103// A fork re-sends the main thread's own prefix, so a warm entry is read (and its clock
104// restarted) at the read price. Nothing is added to the conversation.
105async function ping($: EngineInterface): Promise<string> {
106  pinging = true
107  try {
108    const reply = await $.model.fork({ prompt: 'Reply with the single word: ok' })
109    if (!reply.isAnswered) {
110      return reply.reason === 'nothing-to-fork' ? 'Nothing cached yet: no response in this session.' : `The ping did not go through (${reply.reason}).`
111    }
112    const now = await $.clock.now()
113    const readBack = reply.usage.cache_read_input_tokens
114    await update($, reading, r => ({ ...initial, ...r, lastRequestAt: now, lastPingAt: now, pings: r.pings + 1, pingReadTokens: (r.pingReadTokens ?? 0) + readBack }))
115    return readBack > 0 ? `Cache pinged: ${tokens(readBack)} tokens read, window restarted.` : 'The cache had already lapsed; the ping re-wrote it.'
116  } finally {
117    pinging = false
118  }
119}
120
121async function toggleKeepWarm($: EngineInterface): Promise<CacheReading> {
122  return update($, reading, current => ({ ...initial, ...current, keepWarm: !current.keepWarm }))
123}
124
125export const register: Register = on => {
126  on('session.start', async ($, e, next) => {
127    // The window Claude Code actually asks for: an explicit /keepwarm ttl wins, then the
128    // promptCacheTtl setting. Never a value remembered from an earlier version of this mod.
129    const saved = await $.store.get('ttlMinutes')
130    const settings = (await $.settings.read()) as { promptCacheTtl?: string }
131    const ttlMinutes = typeof saved === 'number' ? saved : settings.promptCacheTtl === '5m' ? 5 : 60
132    await update($, reading, r => ({ ...initial, ...r, ttlMinutes }))
133
134    await $.command.register({
135      name: 'keepwarm',
136      description: 'Prompt cache: /keepwarm toggles keep-warm, /keepwarm now pings, /keepwarm ttl 5|60, /keepwarm panel opens the Cache panel',
137    })
138    void $.ui.open({ id: PANE, title: 'Cache' })
139
140    $.clock.every(1000, () => void update($, tick, n => n + 1))
141    $.clock.every(TICK_MS, async () => {
142      const now = await $.clock.now()
143      const usage = await $.session.usage()
144      const usd = usage.cost?.usd ?? null
145      await update($, reading, saved => {
146        const r = { ...initial, ...saved }
147        // Start (or restart, after an older version) the measured stretch here: weights so far
148        // were counted against an unknown share of the cost, so they are dropped.
149        if (usd !== null && r.usdBaseline === null) return { ...r, sessionUsd: usd, usdBaseline: usd, weightedAll: 0 }
150        return { ...r, sessionUsd: usd }
151      })
152      const r = await readReading($)
153      $.ui.status(statusText(r, now))
154
155      const left = remainingMs(r, now)
156      if (r.keepWarm && !pinging && left !== null && left > 0 && left < MARGIN_MS + TICK_MS) {
157        await ping($)
158      }
159    })
160
161    return next(e)
162  })
163
164  // The main thread's requests decide what is cached; a sub-agent's have their own prefix.
165  on('turn.step', async function* ($, e, next) {
166    const result = yield* next(e)
167    if (!result.usage) return result
168
169    const u = result.usage
170    const model = u.model || e.model
171    if (e.agentId) {
172      await update($, reading, saved => {
173        const r = { ...initial, ...saved }
174        return { ...r, weightedAll: r.weightedAll + weightOf(u, model, 5) }
175      })
176      return result
177    }
178
179    const cached = u.cache_read_input_tokens + u.cache_creation_input_tokens
180    const total = cached + u.input_tokens
181    const share = total > 0 ? u.cache_read_input_tokens / total : 0
182    const now = await $.clock.now()
183    await update($, reading, saved => {
184      const r = { ...initial, ...saved }
185      const isCold = r.cachedTokens > 20_000 && u.cache_creation_input_tokens > r.cachedTokens * 0.5
186      return {
187      ...r,
188      lastRequestAt: now,
189      cachedTokens: cached || r.cachedTokens,
190      history: [...r.history, share].slice(-HISTORY),
191      readTokens: r.readTokens + u.cache_read_input_tokens,
192      writtenTokens: r.writtenTokens + u.cache_creation_input_tokens,
193      weightedAll: r.weightedAll + weightOf(u, model, r.ttlMinutes),
194      mainModel: model || r.mainModel,
195      // A request that re-wrote most of a large prefix it should have read is a cold start.
196      coldStarts: r.coldStarts + (isCold ? 1 : 0),
197      rewrittenTokens: r.rewrittenTokens + (isCold ? u.cache_creation_input_tokens : 0),
198      }
199    })
200    return result
201  })
202
203  on('prompt.submit', async ($, e, next) => {
204    const r = await readReading($)
205    const left = remainingMs(r, await $.clock.now())
206    if (left !== null && left <= 0 && r.cachedTokens > 0) {
207      $.ui.toast(
208        `Cache is cold: this prompt re-writes about ${tokens(r.cachedTokens)} tokens, ` +
209          `${coldFactor(r.ttlMinutes)}x what reading them warm would cost. /keepwarm stops this.`,
210        { timeoutMs: 8000 },
211      )
212    }
213    return next(e)
214  })
215
216  on('command.run', { command: 'keepwarm' }, async ($, e) => {
217    const [word, value] = e.args.trim().split(/\s+/)
218
219    if (word === 'panel') {
220      await $.ui.open({ id: PANE, title: 'Cache', focus: true })
221      return { text: 'Cache panel open.' }
222    }
223    if (word === 'ttl') {
224      const minutes = Number(value)
225      if (minutes !== 5 && minutes !== 60) return { text: 'The cache window is 5 or 60 minutes: /keepwarm ttl 5 or /keepwarm ttl 60.' }
226      await $.store.set('ttlMinutes', minutes)
227      await update($, reading, r => ({ ...initial, ...r, ttlMinutes: minutes }))
228      return { text: `Cache window set to ${minutes} minutes.` }
229    }
230    if (word === 'now') return { text: await ping($) }
231
232    const r = await toggleKeepWarm($)
233    return {
234      text: r.keepWarm
235        ? `Keeping the cache warm: a one-line ping ${MARGIN_MS / 1000}s before each ${r.ttlMinutes}-minute window ends. Each ping reads the cache at ${READ}x.`
236        : 'Stopped keeping the cache warm.',
237    }
238  })
239
240  on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
241    const { Box, Text, Button } = $.ui.resolve(e)
242    const r = await readReading($)
243    await read($, tick)
244    const now = await $.clock.now()
245    const width = Math.max(24, (e.props.bodyColumns ?? 40) - 2)
246    const gaugeWidth = Math.max(10, width - 16)
247    const state = mood(r, now)
248    const left = remainingMs(r, now) ?? 0
249    const windowMs = r.ttlMinutes * 60_000
250    const share = state === 'cold' || state === 'empty' ? 0 : left / windowMs
251    const hue = { warm: '#22c55e', fading: '#f97316', cold: '#ef4444', empty: '#64748b' }[state]
252    const label = { warm: '● WARM', fading: '◐ FADING', cold: '○ COLD', empty: '· EMPTY' }[state]
253    const filled = Math.round(share * gaugeWidth)
254    const hitRate = r.readTokens + r.writtenTokens > 0 ? r.readTokens / (r.readTokens + r.writtenTokens) : 0
255    // Reading a token warm costs a tenth of full input: everything read is 90% saved.
256    const savedFull = Math.round(r.readTokens * (1 - READ))
257
258    return (
259      <Box flexDirection="column">
260        <Box justifyContent="space-between">
261          <Text bold color="#22d3ee">❄ PROMPT CACHE</Text>
262          <Text bold color={hue} inverse={state === 'cold'}> {label} </Text>
263        </Box>
264        <Text dimColor>{'─'.repeat(width)}</Text>
265
266        {state === 'empty' ? (
267          <Text dimColor>Nothing cached yet. The gauge starts with the first response.</Text>
268        ) : (
269          <Box flexDirection="column">
270            <Text>
271              <Text color={hue}>{'▰'.repeat(filled)}</Text>
272              <Text dimColor>{'▱'.repeat(gaugeWidth - filled)}</Text>
273              <Text bold color={hue}> {state === 'cold' ? 'lapsed' : clock(left)}</Text>
274            </Text>
275            <Text dimColor>
276              {tokens(r.cachedTokens)} tokens held  ·  {r.ttlMinutes === 60 ? '1 hour' : '5 minute'} window
277            </Text>
278            {state === 'fading' && (
279              <Text color="#f97316">⚠ under a quarter of the window left. /keepwarm now keeps it.</Text>
280            )}
281            {state === 'cold' && (
282              <Text color="#ef4444">✗ your next prompt re-writes {tokens(r.cachedTokens)} tokens, {coldFactor(r.ttlMinutes)}x a warm read.</Text>
283            )}
284          </Box>
285        )}
286
287        <Box flexDirection="column" marginTop={1}>
288          <Text>
289            <Text dimColor>hits     </Text>
290            <Text color={hitRate >= 0.8 ? '#22c55e' : hitRate >= 0.5 ? '#f97316' : '#ef4444'}>{sparkline(r.history) || '·'}</Text>
291            <Text bold> {Math.round(hitRate * 100)}%</Text>
292            <Text dimColor> read from cache</Text>
293          </Text>
294          <Text>
295            <Text dimColor>saved    </Text>
296            <Text color="#4ade80" bold>{tokens(savedFull)}</Text>
297            <Text dimColor> full-price input tokens</Text>
298          </Text>
299          {r.coldStarts > 0 && (
300            <Text color="#f97316">⚠ {r.coldStarts} cold {r.coldStarts === 1 ? 'start' : 'starts'} this session</Text>
301          )}
302        </Box>
303
304        {(() => {
305          const c = cacheCosts(r)
306          if (!c) return <Text dimColor>{'\n'}COST: calibrating against /cost, a few more requests needed.</Text>
307          const row = (label: string, value: string, hue: string, note?: string) => (
308            <Text key={label}>
309              <Text dimColor>{label.padEnd(18)}</Text>
310              <Text bold color={hue}>{value.padStart(7)}</Text>
311              {note ? <Text dimColor>  {note}</Text> : null}
312            </Text>
313          )
314          return (
315            <Box flexDirection="column" marginTop={1}>
316              <Text bold dimColor>COST (est.)</Text>
317              {row('caching this run', money(c.spent), '#22d3ee', `${tokens(r.writtenTokens)} written, ${tokens(r.readTokens)} read`)}
318              {row('saved by reads', money(c.saved), '#4ade80', 'vs paying full input')}
319              {row('cold-start tax', money(c.coldTax), c.coldTax > 0 ? '#f97316' : '#64748b', `${r.coldStarts} cold ${r.coldStarts === 1 ? 'start' : 'starts'}, ${tokens(r.rewrittenTokens)} re-written`)}
320              {row('keep-warm pings', money(c.pings), '#93c5fd', `${r.pings} ${r.pings === 1 ? 'ping' : 'pings'}`)}
321              {state !== 'empty' && row(state === 'cold' ? 'next prompt costs' : 'if it goes cold', `+${money(c.ifColdNow)}`, state === 'cold' ? '#ef4444' : '#f97316', `extra to re-write ${tokens(r.cachedTokens)}`)}
322              {r.pings > 0 && c.ifColdNow > 0 && (
323                <Text color="#4ade80">› each ping costs about {money((c.pings || 0) / Math.max(1, r.pings))}; a cold start about {money(c.ifColdNow)}</Text>
324              )}
325              <Text dimColor>priced from /cost with the published cache multipliers</Text>
326            </Box>
327          )
328        })()}
329
330        <Box flexDirection="column" marginTop={1}>
331          <Text>
332            <Text dimColor>keep warm </Text>
333            <Text bold color={r.keepWarm ? '#22c55e' : '#64748b'}>{r.keepWarm ? 'ON' : 'off'}</Text>
334            <Text dimColor>   pings {r.pings}{r.lastPingAt ? `, last ${clock(now - r.lastPingAt)} ago` : ''}</Text>
335          </Text>
336          <Box gap={1}>
337            <Button key="toggle" hotkey="w" label={r.keepWarm ? 'Stop keeping warm' : 'Keep warm'} onPress={() => void toggleKeepWarm($)} />
338            <Button key="ping" hotkey="p" label="Ping now" onPress={() => void ping($).then(text => $.ui.toast(text))} />
339            <Button key="ttl" hotkey="t" label={r.ttlMinutes === 60 ? 'Window: 1h' : 'Window: 5m'} onPress={() => void $.command.run({ command: 'keepwarm', args: `ttl ${r.ttlMinutes === 60 ? 5 : 60}` })} />
340          </Box>
341        </Box>
342      </Box>
343    )
344  })
345}
346
types/index.d.ts 31 lines
1export type CacheReading = {
2  lastRequestAt: number | null
3  cachedTokens: number
4  ttlMinutes: number
5  keepWarm: boolean
6  pings: number
7  lastPingAt: number | null
8  // Share of each main-thread request's input served from cache, newest last.
9  history: number[]
10  readTokens: number
11  writtenTokens: number
12  coldStarts: number
13  // Tokens re-written by cold starts: paid at the write price instead of the read price.
14  rewrittenTokens: number
15  // Tokens the keep-warm pings read back.
16  pingReadTokens: number
17  // Every request's tokens, weighted by price ratio and model, for pricing a token from /cost.
18  weightedAll: number
19  sessionUsd: number | null
20  // The session cost when weighting began: a token is priced from the cost added since, over
21  // the requests weighed since, so both cover the same stretch of the session.
22  usdBaseline: number | null
23  mainModel: string
24}
25
26declare module 'claude-code' {
27  interface PluginState {
28    'cache-tax': { reading: CacheReading; tick: number }
29  }
30}
31