A Cache panel: a countdown gauge for the prompt cache, a hit-rate sparkline, cold-cache warnings and /keepwarm

Five side-panel mods for Claude Code (function hooks, early access API):
| Mod | Panel | Command | |
|---|---|---|---|
| agent-progress | Crew: each sub-agent as a little Claude crab dressed for its job (builder with a hammer, inspector with a monocle, scout with a compass, architect with a set square, librarian with a book), model, live activity, steps, elapsed time, cache hit bar, tokens, and a warning when a worker goes quiet | /crew | |
| cache-tax | Cache: prompt cache countdown gauge, hit-rate sparkline, cold-start warnings, keep-warm pings, and what caching cost: spent, saved by reads, the cold-start tax, what pings cost and what going cold now would add (estimated from /cost with the published cache multipliers) | /keepwarm, /keepwarm now, `/keepwarm ttl 5\ | 60, /keepwarm panel` |
| spend-ledger | Spend: what each prompt and each sub-agent cost, a STATS dashboard modelled on /usage stats: views Overview, Tokens, Cost, Cache and In/Out; ranges This session, 7 days, 30 days and All time; a model filter; an activity heatmap, charts in four styles (line, area, bars, dots) with axes, stat tiles (favorite model, total tokens, sessions, active days, longest session, streaks) and per-model cards. Past days come from Claude Code's own history (~/.claude/stats-cache.json, the file /usage reads), with their in/out/cache split and cost estimated from each model's totals; today is exact, from the mod's own tracking. An active-hours chart shows when you work, a by-model breakdown (cost, share, requests, tokens, cache hit rate), burn rate, budget, running insights. On an API key it budgets in dollars (/spend budget 50, warnings at 80% and 100%). On a Claude subscription it shows the plan's 5-hour and weekly windows with reset times and how long the 5-hour window lasts at your pace, and budgets as a share of it (/spend budget 80%); the dollar figures are then the API-price value of the work, not a charge | /spend, /spend budget 50, /spend budget 80%, add hard to refuse new prompts once reached (/spend budget 50 hard), `/spend hard on\ | off, /spend budget off` |
| session-link | Session Link: the live Claude sessions on this machine, each with its own label (the session sets it itself early in its first turn; until then its first prompt stands in; /link name overrides) and its own Claude crab badge (a pose and colour no two live sessions share), its directory, branch and the files it edited; the messages the sessions send each other; files two sessions both edited. It also stops one session overwriting another: an edit to a file another live session edited in the last six hours, or a broad git command (add -A, commit -a, reset --hard, checkout, switch, stash, rebase, merge, pull) in a repository another session works in, is refused once with the other session's name so the model messages it first; the same call again goes through | /link, /link name A: finish UA, /link release; the model's set_session_label tool | |
| mission-control | Mission Control: all four stacked in one scrollable panel, with buttons for every command | /mission |
Every panel has buttons with hotkeys for its commands.
At the prompt of a terminal session:
/plugin install mission-control --marketplace kimathinjoki/claude-session-dashboard
Answer y to add the marketplace, choose the user scope, then install the others the same way (agent-progress, cache-tax, spend-ledger, session-link). Mission Control reads their state, so it needs all four.
Side panels dock beside the transcript in fullscreen mode ("tui": "fullscreen" in ~/.claude/settings.json, or /config) from 110 columns. They open on their own at session start from 144 columns; at any width the commands above open them. With several open they share the dock as tabs.
The Cache panel follows your promptCacheTtl setting ("1h" keeps the cache for an hour; 1-hour cache writes cost more than 5-minute ones).
hooks/register.tsx 346 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { CacheReading } from '../types'
5
6const initial: CacheReading = {
7 lastRequestAt: null, cachedTokens: 0, ttlMinutes: 60, keepWarm: false, pings: 0, lastPingAt: null,
8 history: [], readTokens: 0, writtenTokens: 0, coldStarts: 0,
9 rewrittenTokens: 0, pingReadTokens: 0, weightedAll: 0, sessionUsd: null, usdBaseline: null, mainModel: '',
10}
11const reading = atom({ plugin: 'cache-tax', key: 'reading' } as const, initial)
12const tick = atom({ plugin: 'cache-tax', key: 'tick' } as const, 0)
13
14const PANE = 'cache'
15
16// State saved by an earlier version of this mod may lack fields added since; fill them in.
17async function readReading($: EngineInterface): Promise<CacheReading> {
18 return { ...initial, ...(await read($, reading)) }
19}
20// Anthropic's published multipliers on the base input price: a cache write costs 1.25x
21// (2x for the one-hour cache), a cache read 0.1x. A cold prefix is re-written, not read.
22const READ = 0.1
23const writeMultiplier = (ttlMinutes: number) => (ttlMinutes >= 60 ? 2 : 1.25)
24const MARGIN_MS = 60_000
25const TICK_MS = 15_000
26const HISTORY = 24
27const SPARK = ['▁', '▂', '▃', '▄', '▅', '▆', '▇', '█']
28
29// Relative list prices by model family, as the Spend panel uses them.
30const MODEL_FACTOR: Array<[RegExp, number]> = [[/opus/i, 5], [/sonnet/i, 3], [/haiku/i, 1], [/fable/i, 5]]
31export const modelFactor = (model: string) => MODEL_FACTOR.find(([test]) => test.test(model))?.[1] ?? 3
32
33// A request in units of "base input tokens", so /cost divided by the sum prices one.
34export const weightOf = (
35 u: { input_tokens: number; output_tokens: number; cache_read_input_tokens: number; cache_creation_input_tokens: number },
36 model: string,
37 ttlMinutes: number,
38) =>
39 (u.input_tokens + writeMultiplier(ttlMinutes) * u.cache_creation_input_tokens + READ * u.cache_read_input_tokens + 5 * u.output_tokens) *
40 modelFactor(model)
41
42// The main thread's base input price per token, estimated from what the session actually cost.
43// Below this many weighted units the estimate is noise, and the panel says it is calibrating.
44export const MIN_WEIGHT = 200_000
45
46export const basePrice = (r: Pick<CacheReading, 'sessionUsd' | 'usdBaseline' | 'weightedAll' | 'mainModel'>) => {
47 if (r.sessionUsd == null || r.usdBaseline == null || r.weightedAll < MIN_WEIGHT) return null
48 const spent = r.sessionUsd - r.usdBaseline
49 return spent > 0 ? (spent / r.weightedAll) * modelFactor(r.mainModel || 'opus') : null
50}
51
52export type CacheCosts = { spent: number; saved: number; coldTax: number; pings: number; ifColdNow: number }
53
54export const cacheCosts = (r: CacheReading): CacheCosts | null => {
55 const base = basePrice(r)
56 if (base === null) return null
57 const write = writeMultiplier(r.ttlMinutes)
58 return {
59 spent: (r.writtenTokens * write + r.readTokens * READ) * base,
60 saved: r.readTokens * (1 - READ) * base,
61 coldTax: r.rewrittenTokens * (write - READ) * base,
62 pings: r.pingReadTokens * READ * base,
63 ifColdNow: r.cachedTokens * (write - READ) * base,
64 }
65}
66
67export const money = (n: number) => (n >= 100 ? `$${n.toFixed(0)}` : n >= 1 ? `$${n.toFixed(2)}` : `${Math.max(0, Math.round(n * 100))}¢`)
68
69export const tokens = (n: number) =>
70 n >= 1_000_000 ? `${(n / 1_000_000).toFixed(1)}M` : n >= 1000 ? `${Math.round(n / 1000)}k` : `${n}`
71
72export const remainingMs = (r: CacheReading, now: number) =>
73 r.lastRequestAt === null ? null : r.lastRequestAt + r.ttlMinutes * 60_000 - now
74
75export const coldFactor = (ttlMinutes: number) => Math.round((writeMultiplier(ttlMinutes) / READ) * 10) / 10
76
77export const clock = (ms: number) => {
78 const s = Math.max(0, Math.floor(ms / 1000))
79 return `${Math.floor(s / 60)}:${String(s % 60).padStart(2, '0')}`
80}
81
82export const sparkline = (values: number[]) =>
83 values.map(v => SPARK[Math.max(0, Math.min(SPARK.length - 1, Math.round(v * (SPARK.length - 1))))]).join('')
84
85// Warm, fading (under a quarter of the window), or cold.
86export const mood = (r: CacheReading, now: number) => {
87 const left = remainingMs(r, now)
88 if (left === null || r.cachedTokens === 0) return 'empty' as const
89 if (left <= 0) return 'cold' as const
90 return left < r.ttlMinutes * 60_000 * 0.25 ? ('fading' as const) : ('warm' as const)
91}
92
93export const statusText = (r: CacheReading, now: number) => {
94 const left = remainingMs(r, now)
95 if (left === null || r.cachedTokens === 0) return undefined
96 const warmth = r.keepWarm ? ', kept warm' : ''
97 if (left > 0) return `cache warm ${clock(left)} left (${tokens(r.cachedTokens)})${warmth}`
98 return `cache cold: next prompt re-writes ${tokens(r.cachedTokens)} tokens`
99}
100
101let pinging = false
102
103// A fork re-sends the main thread's own prefix, so a warm entry is read (and its clock
104// restarted) at the read price. Nothing is added to the conversation.
105async function ping($: EngineInterface): Promise<string> {
106 pinging = true
107 try {
108 const reply = await $.model.fork({ prompt: 'Reply with the single word: ok' })
109 if (!reply.isAnswered) {
110 return reply.reason === 'nothing-to-fork' ? 'Nothing cached yet: no response in this session.' : `The ping did not go through (${reply.reason}).`
111 }
112 const now = await $.clock.now()
113 const readBack = reply.usage.cache_read_input_tokens
114 await update($, reading, r => ({ ...initial, ...r, lastRequestAt: now, lastPingAt: now, pings: r.pings + 1, pingReadTokens: (r.pingReadTokens ?? 0) + readBack }))
115 return readBack > 0 ? `Cache pinged: ${tokens(readBack)} tokens read, window restarted.` : 'The cache had already lapsed; the ping re-wrote it.'
116 } finally {
117 pinging = false
118 }
119}
120
121async function toggleKeepWarm($: EngineInterface): Promise<CacheReading> {
122 return update($, reading, current => ({ ...initial, ...current, keepWarm: !current.keepWarm }))
123}
124
125export const register: Register = on => {
126 on('session.start', async ($, e, next) => {
127 // The window Claude Code actually asks for: an explicit /keepwarm ttl wins, then the
128 // promptCacheTtl setting. Never a value remembered from an earlier version of this mod.
129 const saved = await $.store.get('ttlMinutes')
130 const settings = (await $.settings.read()) as { promptCacheTtl?: string }
131 const ttlMinutes = typeof saved === 'number' ? saved : settings.promptCacheTtl === '5m' ? 5 : 60
132 await update($, reading, r => ({ ...initial, ...r, ttlMinutes }))
133
134 await $.command.register({
135 name: 'keepwarm',
136 description: 'Prompt cache: /keepwarm toggles keep-warm, /keepwarm now pings, /keepwarm ttl 5|60, /keepwarm panel opens the Cache panel',
137 })
138 void $.ui.open({ id: PANE, title: 'Cache' })
139
140 $.clock.every(1000, () => void update($, tick, n => n + 1))
141 $.clock.every(TICK_MS, async () => {
142 const now = await $.clock.now()
143 const usage = await $.session.usage()
144 const usd = usage.cost?.usd ?? null
145 await update($, reading, saved => {
146 const r = { ...initial, ...saved }
147 // Start (or restart, after an older version) the measured stretch here: weights so far
148 // were counted against an unknown share of the cost, so they are dropped.
149 if (usd !== null && r.usdBaseline === null) return { ...r, sessionUsd: usd, usdBaseline: usd, weightedAll: 0 }
150 return { ...r, sessionUsd: usd }
151 })
152 const r = await readReading($)
153 $.ui.status(statusText(r, now))
154
155 const left = remainingMs(r, now)
156 if (r.keepWarm && !pinging && left !== null && left > 0 && left < MARGIN_MS + TICK_MS) {
157 await ping($)
158 }
159 })
160
161 return next(e)
162 })
163
164 // The main thread's requests decide what is cached; a sub-agent's have their own prefix.
165 on('turn.step', async function* ($, e, next) {
166 const result = yield* next(e)
167 if (!result.usage) return result
168
169 const u = result.usage
170 const model = u.model || e.model
171 if (e.agentId) {
172 await update($, reading, saved => {
173 const r = { ...initial, ...saved }
174 return { ...r, weightedAll: r.weightedAll + weightOf(u, model, 5) }
175 })
176 return result
177 }
178
179 const cached = u.cache_read_input_tokens + u.cache_creation_input_tokens
180 const total = cached + u.input_tokens
181 const share = total > 0 ? u.cache_read_input_tokens / total : 0
182 const now = await $.clock.now()
183 await update($, reading, saved => {
184 const r = { ...initial, ...saved }
185 const isCold = r.cachedTokens > 20_000 && u.cache_creation_input_tokens > r.cachedTokens * 0.5
186 return {
187 ...r,
188 lastRequestAt: now,
189 cachedTokens: cached || r.cachedTokens,
190 history: [...r.history, share].slice(-HISTORY),
191 readTokens: r.readTokens + u.cache_read_input_tokens,
192 writtenTokens: r.writtenTokens + u.cache_creation_input_tokens,
193 weightedAll: r.weightedAll + weightOf(u, model, r.ttlMinutes),
194 mainModel: model || r.mainModel,
195 // A request that re-wrote most of a large prefix it should have read is a cold start.
196 coldStarts: r.coldStarts + (isCold ? 1 : 0),
197 rewrittenTokens: r.rewrittenTokens + (isCold ? u.cache_creation_input_tokens : 0),
198 }
199 })
200 return result
201 })
202
203 on('prompt.submit', async ($, e, next) => {
204 const r = await readReading($)
205 const left = remainingMs(r, await $.clock.now())
206 if (left !== null && left <= 0 && r.cachedTokens > 0) {
207 $.ui.toast(
208 `Cache is cold: this prompt re-writes about ${tokens(r.cachedTokens)} tokens, ` +
209 `${coldFactor(r.ttlMinutes)}x what reading them warm would cost. /keepwarm stops this.`,
210 { timeoutMs: 8000 },
211 )
212 }
213 return next(e)
214 })
215
216 on('command.run', { command: 'keepwarm' }, async ($, e) => {
217 const [word, value] = e.args.trim().split(/\s+/)
218
219 if (word === 'panel') {
220 await $.ui.open({ id: PANE, title: 'Cache', focus: true })
221 return { text: 'Cache panel open.' }
222 }
223 if (word === 'ttl') {
224 const minutes = Number(value)
225 if (minutes !== 5 && minutes !== 60) return { text: 'The cache window is 5 or 60 minutes: /keepwarm ttl 5 or /keepwarm ttl 60.' }
226 await $.store.set('ttlMinutes', minutes)
227 await update($, reading, r => ({ ...initial, ...r, ttlMinutes: minutes }))
228 return { text: `Cache window set to ${minutes} minutes.` }
229 }
230 if (word === 'now') return { text: await ping($) }
231
232 const r = await toggleKeepWarm($)
233 return {
234 text: r.keepWarm
235 ? `Keeping the cache warm: a one-line ping ${MARGIN_MS / 1000}s before each ${r.ttlMinutes}-minute window ends. Each ping reads the cache at ${READ}x.`
236 : 'Stopped keeping the cache warm.',
237 }
238 })
239
240 on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
241 const { Box, Text, Button } = $.ui.resolve(e)
242 const r = await readReading($)
243 await read($, tick)
244 const now = await $.clock.now()
245 const width = Math.max(24, (e.props.bodyColumns ?? 40) - 2)
246 const gaugeWidth = Math.max(10, width - 16)
247 const state = mood(r, now)
248 const left = remainingMs(r, now) ?? 0
249 const windowMs = r.ttlMinutes * 60_000
250 const share = state === 'cold' || state === 'empty' ? 0 : left / windowMs
251 const hue = { warm: '#22c55e', fading: '#f97316', cold: '#ef4444', empty: '#64748b' }[state]
252 const label = { warm: '● WARM', fading: '◐ FADING', cold: '○ COLD', empty: '· EMPTY' }[state]
253 const filled = Math.round(share * gaugeWidth)
254 const hitRate = r.readTokens + r.writtenTokens > 0 ? r.readTokens / (r.readTokens + r.writtenTokens) : 0
255 // Reading a token warm costs a tenth of full input: everything read is 90% saved.
256 const savedFull = Math.round(r.readTokens * (1 - READ))
257
258 return (
259 <Box flexDirection="column">
260 <Box justifyContent="space-between">
261 <Text bold color="#22d3ee">❄ PROMPT CACHE</Text>
262 <Text bold color={hue} inverse={state === 'cold'}> {label} </Text>
263 </Box>
264 <Text dimColor>{'─'.repeat(width)}</Text>
265
266 {state === 'empty' ? (
267 <Text dimColor>Nothing cached yet. The gauge starts with the first response.</Text>
268 ) : (
269 <Box flexDirection="column">
270 <Text>
271 <Text color={hue}>{'▰'.repeat(filled)}</Text>
272 <Text dimColor>{'▱'.repeat(gaugeWidth - filled)}</Text>
273 <Text bold color={hue}> {state === 'cold' ? 'lapsed' : clock(left)}</Text>
274 </Text>
275 <Text dimColor>
276 {tokens(r.cachedTokens)} tokens held · {r.ttlMinutes === 60 ? '1 hour' : '5 minute'} window
277 </Text>
278 {state === 'fading' && (
279 <Text color="#f97316">⚠ under a quarter of the window left. /keepwarm now keeps it.</Text>
280 )}
281 {state === 'cold' && (
282 <Text color="#ef4444">✗ your next prompt re-writes {tokens(r.cachedTokens)} tokens, {coldFactor(r.ttlMinutes)}x a warm read.</Text>
283 )}
284 </Box>
285 )}
286
287 <Box flexDirection="column" marginTop={1}>
288 <Text>
289 <Text dimColor>hits </Text>
290 <Text color={hitRate >= 0.8 ? '#22c55e' : hitRate >= 0.5 ? '#f97316' : '#ef4444'}>{sparkline(r.history) || '·'}</Text>
291 <Text bold> {Math.round(hitRate * 100)}%</Text>
292 <Text dimColor> read from cache</Text>
293 </Text>
294 <Text>
295 <Text dimColor>saved </Text>
296 <Text color="#4ade80" bold>{tokens(savedFull)}</Text>
297 <Text dimColor> full-price input tokens</Text>
298 </Text>
299 {r.coldStarts > 0 && (
300 <Text color="#f97316">⚠ {r.coldStarts} cold {r.coldStarts === 1 ? 'start' : 'starts'} this session</Text>
301 )}
302 </Box>
303
304 {(() => {
305 const c = cacheCosts(r)
306 if (!c) return <Text dimColor>{'\n'}COST: calibrating against /cost, a few more requests needed.</Text>
307 const row = (label: string, value: string, hue: string, note?: string) => (
308 <Text key={label}>
309 <Text dimColor>{label.padEnd(18)}</Text>
310 <Text bold color={hue}>{value.padStart(7)}</Text>
311 {note ? <Text dimColor> {note}</Text> : null}
312 </Text>
313 )
314 return (
315 <Box flexDirection="column" marginTop={1}>
316 <Text bold dimColor>COST (est.)</Text>
317 {row('caching this run', money(c.spent), '#22d3ee', `${tokens(r.writtenTokens)} written, ${tokens(r.readTokens)} read`)}
318 {row('saved by reads', money(c.saved), '#4ade80', 'vs paying full input')}
319 {row('cold-start tax', money(c.coldTax), c.coldTax > 0 ? '#f97316' : '#64748b', `${r.coldStarts} cold ${r.coldStarts === 1 ? 'start' : 'starts'}, ${tokens(r.rewrittenTokens)} re-written`)}
320 {row('keep-warm pings', money(c.pings), '#93c5fd', `${r.pings} ${r.pings === 1 ? 'ping' : 'pings'}`)}
321 {state !== 'empty' && row(state === 'cold' ? 'next prompt costs' : 'if it goes cold', `+${money(c.ifColdNow)}`, state === 'cold' ? '#ef4444' : '#f97316', `extra to re-write ${tokens(r.cachedTokens)}`)}
322 {r.pings > 0 && c.ifColdNow > 0 && (
323 <Text color="#4ade80">› each ping costs about {money((c.pings || 0) / Math.max(1, r.pings))}; a cold start about {money(c.ifColdNow)}</Text>
324 )}
325 <Text dimColor>priced from /cost with the published cache multipliers</Text>
326 </Box>
327 )
328 })()}
329
330 <Box flexDirection="column" marginTop={1}>
331 <Text>
332 <Text dimColor>keep warm </Text>
333 <Text bold color={r.keepWarm ? '#22c55e' : '#64748b'}>{r.keepWarm ? 'ON' : 'off'}</Text>
334 <Text dimColor> pings {r.pings}{r.lastPingAt ? `, last ${clock(now - r.lastPingAt)} ago` : ''}</Text>
335 </Text>
336 <Box gap={1}>
337 <Button key="toggle" hotkey="w" label={r.keepWarm ? 'Stop keeping warm' : 'Keep warm'} onPress={() => void toggleKeepWarm($)} />
338 <Button key="ping" hotkey="p" label="Ping now" onPress={() => void ping($).then(text => $.ui.toast(text))} />
339 <Button key="ttl" hotkey="t" label={r.ttlMinutes === 60 ? 'Window: 1h' : 'Window: 5m'} onPress={() => void $.command.run({ command: 'keepwarm', args: `ttl ${r.ttlMinutes === 60 ? 5 : 60}` })} />
340 </Box>
341 </Box>
342 </Box>
343 )
344 })
345}
346types/index.d.ts 31 lines1export type CacheReading = {
2 lastRequestAt: number | null
3 cachedTokens: number
4 ttlMinutes: number
5 keepWarm: boolean
6 pings: number
7 lastPingAt: number | null
8 // Share of each main-thread request's input served from cache, newest last.
9 history: number[]
10 readTokens: number
11 writtenTokens: number
12 coldStarts: number
13 // Tokens re-written by cold starts: paid at the write price instead of the read price.
14 rewrittenTokens: number
15 // Tokens the keep-warm pings read back.
16 pingReadTokens: number
17 // Every request's tokens, weighted by price ratio and model, for pricing a token from /cost.
18 weightedAll: number
19 sessionUsd: number | null
20 // The session cost when weighting began: a token is priced from the cost added since, over
21 // the requests weighed since, so both cover the same stretch of the session.
22 usdBaseline: number | null
23 mainModel: string
24}
25
26declare module 'claude-code' {
27 interface PluginState {
28 'cache-tax': { reading: CacheReading; tick: number }
29 }
30}
31