Live quota and prompt-cache monitor for Claude Code: how fast your plan windows are filling and when they run out, plus your prompt-cache misses and an…

Live quota and prompt-cache monitor for Claude Code: how fast your plan windows are filling and when they run out, plus your prompt-cache misses and an estimate of the tokens each one rewrote.

Status: early. The meter works and is tested, but it has not been tuned against many real sessions yet.
● burnrate 5h 23% +4%/h 7d 41% 98% cached
● burnrate 5h 62% +30%/h full in 1h 16m 7d 41% 98% cached
✖ burnrate 5h 24% +4%/h 7d 41% 0% cached miss: model changed, ~81k rewritten
/burn: a pane with the session's totals, each miss with what is known about it, the plan windows with their reset times, pace and where that pace leads, and the last requests one by one. /burn close shuts it.
Both animations are generated from the mod's own rendering code with a scripted session (demo/make-demo.ts, demo/render-gif.py).
The pace is how fast a plan window is filling: points per hour for the five-hour window, per day for the seven-day one. It is measured from the window's own readings, over the last half hour for the five-hour window and up to six hours for the others, and it falls back to "steady" by itself when you stop sending.
full in 1h 16m appears only when the window would fill before it resets at the current pace. It turns red inside half an hour. A rise of a single point is not enough to raise it.~74% at reset is where the window ends up if the reset comes first.pace: measuring shows until burnrate has enough history: five minutes for the five-hour window, an hour for the others.It is a projection of the recent past, not a forecast: one heavy turn moves it, and it knows nothing about what you will do next. The windows are your account's, so other sessions and devices move them too.
A miss is a request that left a fifth or more of the cached prompt unread and wrote it again. burnrate says what it can tell about it:
| Shown as | What it means |
|---|---|
| model changed | the request went to a different model than the one before |
| after 12m idle | more than five minutes passed between the previous response and this request |
| likely prefix change | neither of the above. The usual reason is that something ahead of the conversation changed (effort, tools, system prompt, CLAUDE.md), but burnrate cannot see that |
The limits:
~. The API reports how many tokens were read and written, not which ones, so new content can hide inside a rewrite./compact or a rewind the rewrite is expected, and it cannot be told from a lost cache./plugin marketplace add nemke82/claude-code-burnrate
/plugin install burnrate@burnrate
Mods are early access: start Claude Code with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 (2.1.259 or later). The mod API may change between releases.
claude plugin validate burnrate
claude plugin test burnrate
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir burnrate # hot reloads on save
MIT
hooks/register.tsx 136 lines1/**
2 * burnrate — Claude Code mod (EARLY ACCESS)
3 *
4 * A quota and prompt-cache monitor: the account's rate-limit windows
5 * (`session.measure`), every main-loop request's cache usage (`turn.step`)
6 * and the misses among them, kept in `$.state`. No dollars: the engine's
7 * figure is list price, which is not what a subscription or a custom
8 * contract pays.
9 *
10 * - a row above the prompt: the plan window that matters most with its pace
11 * (see pace.ts), the last request and the misses (estimates, see detectMiss)
12 * - `/burn`: a pane with the session's totals, each miss and the last requests
13 */
14import { atom, read, update } from 'claude-code'
15import type { EngineInterface, Register } from 'claude-code'
16
17import { EMPTY_METER, record } from './meter'
18import { addReadings } from './pace'
19import { bandSegs, paneRows } from './view'
20import type { Meter, PlanWindow, Reading } from '../types'
21
22const PANE = 'burnrate'
23const COMMAND = 'burn'
24
25const meter = atom({ plugin: 'burnrate', key: 'meter' } as const, EMPTY_METER as Meter)
26const windows = atom({ plugin: 'burnrate', key: 'windows' } as const, [] as PlanWindow[])
27const readings = atom({ plugin: 'burnrate', key: 'readings' } as const, [] as Reading[])
28
29// the windows as they stand, and one more point of each one's history
30async function measure($: EngineInterface, limits: readonly PlanWindow[]) {
31 const now = limits.map(w => ({ ...w }))
32 const at = await $.clock.now()
33 await update($, windows, () => now)
34 await update($, readings, history => addReadings(history, now, at))
35}
36
37export const register: Register = on => {
38 on('session.start', async ($, e, next) => {
39 // the engine pushes the windows only as they move: start from where they stand
40 const usage = await $.session.usage().catch(() => undefined)
41 if (usage && usage.rateLimits.length > 0) await measure($, usage.rateLimits)
42
43 await $.command.register({
44 name: COMMAND,
45 description: 'Cache misses, plan windows and the last requests of this session (close shuts the pane)',
46 argumentHint: '[close]',
47 immediate: true,
48 })
49
50 return next(e)
51 })
52
53 on('command.run', { command: COMMAND }, async ($, e) => {
54 if (e.args.trim().toLowerCase() === 'close') {
55 await $.ui.close({ id: PANE })
56 return { text: 'burnrate pane closed.' }
57 }
58 await $.ui.open({ id: PANE, title: 'burnrate', rows: 24 })
59
60 return { text: `burnrate pane opened. /${COMMAND} close shuts it.` }
61 })
62
63 // each main-loop request: what the cache did with it (a subagent has a prefix of its own)
64 on('turn.step', async function* ($, e, next) {
65 if (e.agentId) return yield* next(e)
66 const startedAt = await $.clock.now()
67 const r = yield* next(e)
68 const usage = r.usage
69 if (usage) {
70 const endedAt = await $.clock.now()
71 await update($, meter, m =>
72 record(m, {
73 turnId: e.turnId,
74 index: e.index,
75 model: usage.model || e.model,
76 startedAt,
77 endedAt,
78 read: usage.cache_read_input_tokens,
79 write: usage.cache_creation_input_tokens,
80 fresh: usage.input_tokens,
81 output: usage.output_tokens,
82 }),
83 )
84 }
85 return r
86 })
87
88 on('session.measure', async ($, e, next) => {
89 await measure($, e.rateLimits)
90 return next(e)
91 })
92
93 // /clear starts a new conversation in the same session: its cache is a new one
94 on('session.end', async ($, e, next) => {
95 if (e.reason === 'clear') await update($, meter, () => EMPTY_METER)
96 return next(e)
97 })
98
99 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
100 if (e.props.hasSurvey) return next(e)
101 const segs = bandSegs(await read($, meter), await read($, windows), await read($, readings), await $.clock.now(), e.props.bodyColumns)
102 const { Box, Text } = $.ui.resolve(e)
103
104 return (
105 <Box flexDirection="row" columnGap={1}>
106 {segs.map(s => (
107 <Text color={s.color} bold={s.bold} dimColor={s.dim} wrap="truncate-end">{s.text}</Text>
108 ))}
109 </Box>
110 )
111 })
112
113 on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
114 const { Box, Text } = $.ui.resolve(e)
115 // HTML collapses runs of spaces; a no-break space keeps the columns
116 const sp = (t: string) => (e.surface === 'terminal' ? t : t.replace(/ /g, ' '))
117 const rows = paneRows(await read($, meter), await read($, windows), await read($, readings), await $.clock.now(), e.props.scroll.bodyRows)
118
119 return (
120 <Box flexDirection="column">
121 {rows.map(row =>
122 row.length === 0 ? (
123 <Text> </Text>
124 ) : (
125 <Box flexDirection="row" columnGap={1}>
126 {row.map(s => (
127 <Text color={s.color} bold={s.bold} dimColor={s.dim}>{sp(s.text)}</Text>
128 ))}
129 </Box>
130 ),
131 )}
132 </Box>
133 )
134 })
135}
136hooks/meter.ts 82 lines1/**
2 * meter.ts — the pure half of burnrate: no `$`, no engine.
3 *
4 * A request's prompt is `fresh` (uncached) + `read` (served by the cache) +
5 * `write` (written to it). A healthy request reads nearly all of the previous
6 * prompt and writes only the new tail; a miss writes the prefix again.
7 */
8import type { Meter, Miss, PlanWindow, Sample, Totals } from '../types'
9
10/** Samples kept for the per-request view; the totals count every request. */
11export const KEEP = 200
12
13// The shortest cache lifetime. After a longer pause the entry may have lapsed;
14// the mod API does not say which lifetime the request asked for, so a miss
15// past it is reported with the pause, not as an expiry.
16export const IDLE_MS = 300_000
17
18export const EMPTY_TOTALS: Totals = { requests: 0, read: 0, write: 0, fresh: 0, output: 0, misses: 0, wasted: 0 }
19
20export const EMPTY_METER: Meter = { samples: [], totals: EMPTY_TOTALS }
21
22export const promptTokens = (s: Pick<Sample, 'read' | 'write' | 'fresh'>) => s.read + s.write + s.fresh
23
24/** Share of the prompt the cache served, 0 to 1; 0 for an empty prompt. */
25export function hitRatio(s: Pick<Sample, 'read' | 'write' | 'fresh'>): number {
26 const total = s.read + s.write + s.fresh
27 return total === 0 ? 0 : s.read / total
28}
29
30// A request that leaves this share of the cached prefix unread, and writes it
31// again, is a miss; less is the ordinary churn at the prompt's tail.
32export const MISS_SHARE = 0.2
33
34/**
35 * The miss `cur` was, given the request before it; undefined when it was not
36 * one. Judged on counts alone: what `prev` left cached that `cur` did not read
37 * and wrote instead. A prompt that shrank (a /compact, a rewind) is left out:
38 * its rewrite cannot be told from a lost cache. A request after an uncached
39 * one had nothing to read.
40 */
41export function detectMiss(prev: Sample | undefined, cur: Omit<Sample, 'n' | 'miss'>): Miss | undefined {
42 if (!prev) return undefined
43 const cached = prev.read + prev.write
44 if (cached === 0 || promptTokens(cur) < promptTokens(prev) * 0.7) return undefined
45 const wasted = Math.min(cur.write, cached - cur.read)
46 if (wasted < cached * MISS_SHARE) return undefined
47 const gapMs = Math.max(0, cur.startedAt - prev.endedAt)
48 if (cur.model !== prev.model) return { cause: 'model', wasted, gapMs, from: prev.model }
49 return { cause: gapMs > IDLE_MS ? 'idle' : 'prefix', wasted, gapMs }
50}
51
52/** The meter with one more request: its miss worked out, the totals moved on. */
53export function record(meter: Meter, raw: Omit<Sample, 'n' | 'miss'>): Meter {
54 const miss = detectMiss(meter.samples[meter.samples.length - 1], raw)
55 const t = meter.totals
56 const n = t.requests + 1
57 const sample: Sample = miss ? { ...raw, n, miss } : { ...raw, n }
58 return {
59 samples: [...meter.samples, sample].slice(-KEEP),
60 totals: {
61 requests: t.requests + 1,
62 read: t.read + sample.read,
63 write: t.write + sample.write,
64 fresh: t.fresh + sample.fresh,
65 output: t.output + sample.output,
66 misses: t.misses + (miss ? 1 : 0),
67 wasted: t.wasted + (miss?.wasted ?? 0),
68 },
69 }
70}
71
72export function fmtTokens(n: number): string {
73 if (n < 1000) return String(n)
74 if (n < 100_000) return `${(n / 1000).toFixed(1).replace(/\.0$/, '')}k`
75 if (n < 1_000_000) return `${Math.round(n / 1000)}k`
76 return `${(n / 1_000_000).toFixed(1).replace(/\.0$/, '')}M`
77}
78
79const WINDOW_LABEL: Record<string, string> = { five_hour: '5h', seven_day: '7d', spend_limit: 'spend' }
80
81export const windowLabel = (w: PlanWindow) => WINDOW_LABEL[w.kind] ?? w.kind
82hooks/pace.ts 86 lines1/**
2 * pace.ts — how fast a plan window is filling, from its own readings over
3 * time. Pure: no `$`, no engine.
4 *
5 * The engine reports a window's `percentUsed` after each turn and when it
6 * moves a whole point. A pace is the rise over the recent past, measured up
7 * to now, so it falls off by itself while nothing is being sent.
8 */
9import type { PlanWindow, Reading } from '../types'
10
11const MINUTE = 60_000
12const HOUR = 60 * MINUTE
13
14/** Readings older than this are dropped. */
15export const KEEP_MS = 6 * HOUR
16
17/** How far back a pace looks: half an hour for the five-hour window, the kept history for the rest. */
18export const lookbackMs = (kind: string) => (kind === 'five_hour' ? 30 * MINUTE : KEEP_MS)
19
20/** Less history than this and the pace is not stated: five minutes for the five-hour window, an hour for the slower ones. */
21export const minSpanMs = (kind: string) => (kind === 'five_hour' ? 5 * MINUTE : HOUR)
22
23// The engine reports a window as it moves a whole point, so a rise of one
24// point may be rounding. A warning needs more than that behind it.
25export const MIN_RISE = 2
26
27export type Pace = {
28 /** points of the window per hour; 0 when it did not move */
29 perHour: number
30 /** ms until the window is full at this pace; absent when it is not rising */
31 fullInMs?: number
32 /** where the window stands when it resets, at this pace, capped at 100; absent without a reset time */
33 atReset?: number
34 /** true when it fills before it resets: the reading that matters */
35 hitsLimit: boolean
36}
37
38/**
39 * The history with the windows' readings at `now`: one per window whose share
40 * moved. A share that fell means the window reset, and its earlier readings
41 * go; so do those past KEEP_MS.
42 */
43export function addReadings(history: readonly Reading[], windows: readonly PlanWindow[], now: number): Reading[] {
44 let next = history.filter(r => now - r.at <= KEEP_MS)
45 for (const w of windows) {
46 const mine = next.filter(r => r.kind === w.kind)
47 const last = mine[mine.length - 1]
48 if (last && w.percentUsed < last.percentUsed) next = next.filter(r => r.kind !== w.kind)
49 else if (last && w.percentUsed === last.percentUsed) continue
50 next = [...next, { kind: w.kind, at: now, percentUsed: w.percentUsed }]
51 }
52 return next
53}
54
55/** The window's pace at `now`; undefined until `minSpanMs` of its history exists. */
56export function paceOf(history: readonly Reading[], w: PlanWindow, now: number): Pace | undefined {
57 const mine = history.filter(r => r.kind === w.kind && r.at <= now)
58 const first = mine[0]
59 if (!first) return undefined
60 const from = now - lookbackMs(w.kind)
61 // where the window stood when the lookback began: the last reading before it, carried forward
62 const before = mine.filter(r => r.at <= from)
63 const base = before.length > 0 ? { at: from, percentUsed: (before[before.length - 1] as Reading).percentUsed } : first
64 const span = now - base.at
65 if (span < minSpanMs(w.kind)) return undefined
66
67 const rise = w.percentUsed - base.percentUsed
68 const perHour = Math.max(0, (rise / span) * HOUR)
69 const resetAt = w.resetsAt ? Date.parse(w.resetsAt) : NaN
70 const untilReset = Number.isFinite(resetAt) && resetAt > now ? resetAt - now : undefined
71 const pace: Pace = { perHour, hitsLimit: false }
72 if (untilReset !== undefined) pace.atReset = Math.min(100, w.percentUsed + (perHour * untilReset) / HOUR)
73 if (perHour > 0 && w.percentUsed < 100) {
74 pace.fullInMs = ((100 - w.percentUsed) / perHour) * HOUR
75 pace.hitsLimit = rise >= MIN_RISE && untilReset !== undefined && pace.fullInMs < untilReset
76 }
77 return pace
78}
79
80/** `+12%/h`, or per day for the seven-day window; one decimal under ten. */
81export function fmtPace(perHour: number, kind: string): string {
82 const [rate, unit] = kind === 'seven_day' ? [perHour * 24, 'd'] : [perHour, 'h']
83 const text = rate >= 10 ? String(Math.round(rate)) : rate.toFixed(1).replace(/\.0$/, '')
84 return `+${text}%/${unit}`
85}
86hooks/view.ts 205 lines1/**
2 * view.ts — what the band and the /burn pane say, as plain data: rows of
3 * coloured text segments. Pure, so a test reads it without a surface.
4 */
5import type { Meter, Miss, PlanWindow, Reading, Sample } from '../types'
6import { fmtTokens, hitRatio, windowLabel } from './meter'
7import { fmtPace, paceOf } from './pace'
8import type { Pace } from './pace'
9
10export type Seg = { text: string; color?: string; bold?: boolean; dim?: boolean }
11
12export type Health = 'cold' | 'uncached' | 'miss' | 'warm' | 'partial'
13
14/** How the cache treated a request: the band's colour and mark. */
15export function health(s: Sample): Health {
16 if (s.miss) return 'miss'
17 if (s.read + s.write === 0) return 'uncached'
18 if (s.n === 1) return 'cold'
19 return hitRatio(s) >= 0.8 ? 'warm' : 'partial'
20}
21
22const MARK: Record<Health, string> = { cold: '○', uncached: '○', miss: '✖', warm: '●', partial: '▲' }
23const COLOR: Record<Health, string | undefined> = { cold: undefined, uncached: undefined, miss: 'red', warm: 'green', partial: 'yellow' }
24
25export function bar(ratio: number, width: number): string {
26 const filled = Math.round(Math.min(1, Math.max(0, ratio)) * width)
27 return '█'.repeat(filled) + '░'.repeat(width - filled)
28}
29
30/** 45s, 12m, 1h 5m, 2d 3h. */
31export function fmtGap(ms: number): string {
32 const s = Math.max(0, Math.round(ms / 1000))
33 if (s < 60) return `${s}s`
34 const m = Math.round(s / 60)
35 if (m < 60) return `${m}m`
36 const h = Math.floor(m / 60)
37 if (h < 24) return m % 60 === 0 ? `${h}h` : `${h}h ${m % 60}m`
38 return h % 24 === 0 ? `${Math.floor(h / 24)}d` : `${Math.floor(h / 24)}d ${h % 24}h`
39}
40
41/** What is known about a miss, not why the cache lost it: a pause is stated, never called an expiry. */
42export function missText(miss: Miss): string {
43 if (miss.cause === 'model') return 'model changed'
44 if (miss.cause === 'idle') return `after ${fmtGap(miss.gapMs)} idle`
45 // neither a model change nor a pause: a changed prefix is the usual reason, not an observed one
46 return 'likely prefix change'
47}
48
49/** `~81.3k`: a miss's rewrite is an estimate. */
50export const fmtWaste = (tokens: number) => `~${fmtTokens(tokens)}`
51
52const windowColor = (percent: number) => (percent >= 90 ? 'red' : percent >= 70 ? 'yellow' : undefined)
53
54function windowSeg(w: PlanWindow): Seg {
55 const color = windowColor(w.percentUsed)
56 const text = `${windowLabel(w)} ${Math.round(w.percentUsed)}%`
57 return color ? { text, color, bold: true } : { text, dim: true }
58}
59
60const plural = (n: number, one: string) => `${n} ${n === 1 ? one : `${one}es`}`
61
62/** Drops the least important segments until the row fits: a higher `rank` goes first, the later of equals before the earlier. */
63export function fit(ranked: readonly (readonly [Seg, number])[], columns: number): Seg[] {
64 const kept = [...ranked]
65 const width = () => kept.reduce((sum, [seg]) => sum + seg.text.length, 0) + Math.max(0, kept.length - 1)
66 while (width() > columns) {
67 const worst = Math.max(...kept.map(([, rank]) => rank))
68 if (worst === 0) break
69 kept.splice(kept.map(([, rank]) => rank).lastIndexOf(worst), 1)
70 }
71 return kept.map(([seg]) => seg)
72}
73
74type Paced = { w: PlanWindow; pace: Pace | undefined }
75
76/** The windows with their paces, the one that fills before it resets first (soonest of several), then by share used. */
77export function byUrgency(windows: readonly PlanWindow[], readings: readonly Reading[], now: number): Paced[] {
78 const paced = windows.map(w => ({ w, pace: paceOf(readings, w, now) }))
79 const due = (p: Paced) => (p.pace?.hitsLimit ? (p.pace.fullInMs ?? 0) : Infinity)
80 return paced.sort((a, b) => due(a) - due(b) || b.w.percentUsed - a.w.percentUsed)
81}
82
83/** `full in 1h 16m`, red inside half an hour; nothing for a window that resets first. */
84function limitSeg(p: Paced): Seg | undefined {
85 if (p.w.percentUsed >= 100) return { text: 'limit reached', color: 'red', bold: true }
86 if (!p.pace?.hitsLimit || p.pace.fullInMs === undefined) return undefined
87 return { text: `full in ${fmtGap(p.pace.fullInMs)}`, color: p.pace.fullInMs < 30 * 60_000 ? 'red' : 'yellow', bold: true }
88}
89
90const isRising = (p: Paced) => p.pace !== undefined && p.pace.perHour >= 0.05
91
92/**
93 * The band: one row, the most urgent plan window first with its pace. In the
94 * order a narrow terminal drops them: the session's miss count, the other
95 * windows, the last miss's detail (its red mark stays) and the pace, the
96 * cached share, then the warning that a window fills before it resets.
97 */
98export function bandSegs(meter: Meter, windows: readonly PlanWindow[], readings: readonly Reading[], now: number, columns: number): Seg[] {
99 const name: Seg = { text: 'burnrate', color: 'cyan', bold: true }
100 const last = meter.samples[meter.samples.length - 1]
101 const [lead, ...others] = byUrgency(windows, readings, now)
102 const plan: [Seg, number][] = []
103 if (lead) {
104 plan.push([windowSeg(lead.w), 1])
105 if (isRising(lead)) plan.push([{ text: fmtPace(lead.pace?.perHour ?? 0, lead.w.kind), dim: true }, 3])
106 const limit = limitSeg(lead)
107 if (limit) plan.push([limit, 1])
108 }
109 for (const p of others) plan.push([windowSeg(p.w), 4])
110 if (!last) return fit([[name, 0], ...plan, [{ text: 'waiting for the first request', dim: true }, 2]], columns)
111
112 const state = health(last)
113 const ranked: [Seg, number][] = [[{ text: MARK[state], color: COLOR[state], bold: true }, 0], [name, 0], ...plan]
114 ranked.push([{ text: `${Math.round(hitRatio(last) * 100)}% cached`, color: COLOR[state] }, 2])
115
116 const { misses, wasted } = meter.totals
117 if (last.miss) ranked.push([{ text: `miss: ${missText(last.miss)}, ${fmtWaste(last.miss.wasted)} rewritten`, color: 'red' }, 3])
118 if (misses > (last.miss ? 1 : 0)) ranked.push([{ text: `${plural(misses, 'miss')} ${fmtWaste(wasted)}`, dim: true }, 5])
119 return fit(ranked, columns)
120}
121
122/** What the pane says of a window's pace: the rate, then where it leads. */
123function paceSegs(p: Paced): Seg[] {
124 if (!p.pace) return [{ text: 'pace: measuring', dim: true }]
125 if (!isRising(p)) return [{ text: 'steady', dim: true }]
126 const segs: Seg[] = [{ text: fmtPace(p.pace.perHour, p.w.kind) }]
127 const limit = limitSeg(p)
128 if (limit) segs.push(limit)
129 else if (p.pace.atReset !== undefined) segs.push({ text: `~${Math.round(p.pace.atReset)}% at reset`, dim: true })
130 return segs
131}
132
133const pad = (text: string, width: number) => text.padStart(width)
134
135/** `resets in 2h 10m`, or nothing when the window names no reset or it has passed. */
136export function resetText(w: PlanWindow, now: number): string {
137 const at = w.resetsAt ? Date.parse(w.resetsAt) : NaN
138 return Number.isFinite(at) && at > now ? `resets in ${fmtGap(at - now)}` : ''
139}
140
141/** The /burn pane, row by row; an empty row is a blank line. The request table takes what `maxRows` leaves, three rows at least. */
142export function paneRows(meter: Meter, windows: readonly PlanWindow[], readings: readonly Reading[], now: number, maxRows: number): Seg[][] {
143 const t = meter.totals
144 const prompt = t.read + t.write + t.fresh
145 const rows: Seg[][] = []
146
147 const summary = [`${t.requests} ${t.requests === 1 ? 'request' : 'requests'}`]
148 if (prompt > 0) summary.push(`${Math.round((t.read / prompt) * 100)}% cached`)
149 summary.push(`${fmtTokens(t.output)} output`)
150 rows.push([{ text: 'Session', bold: true, color: 'cyan' }, { text: summary.join(' · ') }])
151 rows.push([])
152
153 rows.push([
154 { text: 'Cache misses', bold: true, color: 'cyan' },
155 t.misses === 0 ? { text: 'none', dim: true } : { text: `${t.misses} · ${fmtWaste(t.wasted)} tokens rewritten (estimate)`, color: 'yellow' },
156 ])
157 for (const s of meter.samples.filter(s => s.miss).slice(-5)) {
158 const miss = s.miss
159 if (!miss) continue
160 const row: Seg[] = [
161 { text: ` req ${pad(String(s.n), 3)}`, dim: true },
162 { text: pad(fmtWaste(miss.wasted), 7), color: 'red', bold: true },
163 { text: missText(miss) },
164 ]
165 if (miss.from) row.push({ text: `${miss.from} → ${s.model}`, dim: true })
166 rows.push(row)
167 }
168 rows.push([])
169
170 rows.push([{ text: 'Plan windows', bold: true, color: 'cyan' }, ...(windows.length === 0 ? [{ text: 'none reported (API key, or no response yet)', dim: true }] : [])])
171 for (const p of byUrgency(windows, readings, now)) {
172 const w = p.w
173 const color = windowColor(w.percentUsed)
174 const reset = resetText(w, now)
175 rows.push([
176 { text: ` ${windowLabel(w).padEnd(5)}`, dim: true },
177 { text: bar(w.percentUsed / 100, 20), color: color ?? 'green' },
178 { text: pad(`${Math.round(w.percentUsed)}%`, 4), bold: true, color },
179 ...(reset ? [{ text: reset, dim: true }] : []),
180 ...paceSegs(p),
181 ])
182 }
183 rows.push([])
184
185 rows.push([{ text: 'Recent requests', bold: true, color: 'cyan' }, ...(meter.samples.length === 0 ? [{ text: 'none yet', dim: true }] : [])])
186 if (meter.samples.length > 0) {
187 rows.push([{ text: ` ${pad('req', 5)} ${pad('read', 6)} ${pad('wrote', 6)} ${pad('new', 6)} ${pad('hit', 4)}`, dim: true }])
188 // a blank and the legend come out of what is left
189 for (const s of meter.samples.slice(-Math.max(3, maxRows - rows.length - 2))) {
190 const state = health(s)
191 rows.push([
192 { text: ` ${pad(String(s.n), 5)}`, dim: true },
193 { text: pad(fmtTokens(s.read), 6), color: 'green' },
194 { text: pad(fmtTokens(s.write), 6), color: 'yellow' },
195 { text: pad(fmtTokens(s.fresh), 6) },
196 { text: pad(`${Math.round(hitRatio(s) * 100)}%`, 4), bold: true, color: COLOR[state] },
197 ...(s.miss ? [{ text: missText(s.miss), color: 'red' }] : []),
198 ])
199 }
200 rows.push([])
201 rows.push([{ text: 'read: served by the cache · wrote: new cache entry · new: sent uncached', dim: true }])
202 }
203 return rows
204}
205types/index.d.ts 76 lines1/** Why a request wrote the cache where it should have read it. */
2export type MissCause = 'model' | 'idle' | 'prefix'
3
4export type Miss = {
5 cause: MissCause
6 /**
7 * An estimate of the tokens written again: what the previous request left
8 * cached and this one did not read, capped at what this one wrote. The API
9 * reports counts, not which tokens, so new content can hide inside it.
10 */
11 wasted: number
12 /** ms between the end of the previous request and the start of this one */
13 gapMs: number
14 /** the previous request's model, when the cause is a model change */
15 from?: string
16}
17
18/** One main-loop request, as the API reported it. */
19export type Sample = {
20 /** the request's number in the session, from 1 */
21 n: number
22 turnId: string
23 index: number
24 model: string
25 /** `$.clock.now()` when the request started */
26 startedAt: number
27 /** `$.clock.now()` when its response was whole */
28 endedAt: number
29 /** served by the cache */
30 read: number
31 /** written to the cache */
32 write: number
33 /** sent uncached */
34 fresh: number
35 output: number
36 miss?: Miss
37}
38
39/** Sums over every request since the session began or was cleared. */
40export type Totals = {
41 requests: number
42 read: number
43 write: number
44 fresh: number
45 output: number
46 misses: number
47 wasted: number
48}
49
50export type Meter = {
51 /** the most recent requests, oldest first */
52 samples: Sample[]
53 totals: Totals
54}
55
56/** A rate-limit window of the account: `five_hour`, `seven_day`, a gateway's `spend_limit`. */
57export type PlanWindow = {
58 kind: string
59 percentUsed: number
60 resetsAt?: string
61}
62
63/** What a plan window stood at, when: the points a pace is worked out from. */
64export type Reading = {
65 kind: string
66 /** `$.clock.now()` when it was read */
67 at: number
68 percentUsed: number
69}
70
71declare module 'claude-code' {
72 interface PluginState {
73 burnrate: { meter: Meter; windows: PlanWindow[]; readings: Reading[] }
74 }
75}
76