Your Claude plan's 5h and weekly usage, the context window and the prompt cache's time left, as one line of bars above the prompt

A Claude Code mod that keeps your plan's usage in view: one line above the prompt with the 5-hour and weekly windows, the context window, the prompt cache's time left and what the context adds to each turn's cost.
5h ▬▬▬▬▬▬▬▬▬ 93% → 100% 23:10 7d ▬▬▬▬▬▬▬ 72% 11:00 ctx ▬▬▬ 37% 370k cache ▬▬▬ 42m cost ×4.5
→ N% is where the current pace lands by the reset, with a dashed outline on the bar.ctx is how full the context window is; 370k is what the next request re-sends, which a cold cache has to write again.cache drains over the time the prompt cache stays warm after the last response: the length the last response actually wrote, read from the transcript at the end of each turn, so it drops to five minutes in overage. Until the first turn ends (and on Windows, which has no tail) it assumes an hour on a subscription's main thread, five minutes otherwise, or what promptCacheTtl sets. It turns orange in the last five minutes and red once cold. When 50%, 25%, 10% and 5% of the lifetime is left, every thread pings with a toast and, on macOS, a Notification Centre banner titled with the thread's name (osascript, so the banner comes from Script Editor, which macOS may ask to allow once). Set iMessageTo (below) to a phone number or Apple ID to get the same pings on the phone, sent through this Mac's Messages app; macOS asks once to let Claude control Messages.cost ×4.5 is what the last turn cost over the same turn in a fresh thread. Every request in a turn (one per tool call) resends the context, so a long thread costs more per message even with a warm cache. The fresh thread starts from what the session's first request sent (system prompt, tools, rules). Token kinds are weighed at API price ratios: cache read 0.1, cache write 1.25 (5 minutes) or 2 (1 hour), output 5, against uncached input 1. How the plan's quotas weigh them is not published, so treat it as an estimate.cost uses those and says measured.Each session logs its requests' token counts and the 5h and 7d readings to the plugin's store, kept for eight days. Readings from all sessions are cut into intervals in which a window rose at least one point, and a least-squares fit gives how many cache-write tokens make 1% and what a cache read weighs against a cache write. Output stays at the API ratio of 2.5 cache writes per token: it is about a tenth of the quota use, too little to separate from the noise. /quota-weights prints the fit:
5h: 52 intervals, 3 dropped as usage elsewhere; 1% = 2.0M cache-write tokens; cache read = 0.052 (0.048 to 0.058) of a cache write, API 0.05 to 0.08
7d: 6 intervals; 1% = 14.1M cache-write tokens; cache read = 0.061 (0.030 to 0.090) of a cache write, API 0.05 to 0.08
output stays at 2.5 cache writes per token
cost × uses measured 5h weights
The readings cover the whole account, so usage the log cannot see (claude.ai, another machine, subagents) lands in them too. Intervals that used 30% more than the fit predicts are dropped as such, intervals across a 30-minute silence are skipped, and cost switches to measured weights only after 20 intervals with the cache read's 10th to 90th percentile range within ±25%.
Two optional buttons send a prompt of your choice, for example a wrap-up skill. Set them in /config under the plugin's options, or in settings.json:
"pluginConfigs": {
"quota-band@claude-quota-band": {
"options": { "closeCommand": "/close-thread", "renamePrompt": "Rename this thread to cover every topic in it", "iMessageTo": "" }
}
}
Empty hides a button. A press during a turn waits until the turn ends.
It reads the figures Claude Code already receives with each response (session.measure), and the ends of the session's transcript at the end of each turn, so it sends no requests of its own. It writes to its own plugin store and, where the Claude Profiles Mac app is installed, the latest figures per account to ~/Library/Application Support/Claude Profiles/mod-readings/, which that app shows instead of asking Anthropic. The figures appear after the first response of a session, on Pro, Max and Team plans.
Needs Claude Code 2.1.287 or later (Claude Mods). In Claude Code:
/plugin marketplace add andras-gyarmati/claude-quota-band
/plugin install quota-band@claude-quota-band
To pick up new versions on each start, set "autoUpdate": true for this marketplace under extraKnownMarketplaces in settings.json.
Bars draw in the desktop app's Code tab; the terminal shows the same line as text.
hooks/register.tsx 604 lines1import { atom, read, update } from 'claude-code'
2import type { Register } from 'claude-code'
3
4import type { Limit, Multiplier, Turn } from '../types'
5
6const limits = atom({ plugin: 'quota-band', key: 'limits' } as const, [] as Limit[])
7const turn = atom({ plugin: 'quota-band', key: 'turn' } as const, null as Turn | null)
8const multiplier = atom({ plugin: 'quota-band', key: 'multiplier' } as const, null as Multiplier | null)
9const tick = atom({ plugin: 'quota-band', key: 'tick' } as const, 0)
10
11const HOUR = 3600e3
12const FIVE_MINUTES = 5 * 60e3
13const COLD_SOON = FIVE_MINUTES
14/** Shares of the cache lifetime left at which a thread pings, warmest first. */
15export const COLD_STEPS = [0.5, 0.25, 0.1, 0.05]
16const COLD_CHECK_MS = 10e3
17const TRANSCRIPT_TAIL_BYTES = '524288'
18
19/** The fallback before a transcript has been read: subscribers' main thread
20 * caches for an hour, everyone else for five minutes, unless `promptCacheTtl`
21 * says otherwise. */
22async function cacheTtl($: any, subscriber: boolean): Promise<number> {
23 const setting = (await $.settings.read())?.promptCacheTtl
24 if (setting === '1h') return HOUR
25 if (setting === '5m') return FIVE_MINUTES
26 return subscriber ? HOUR : FIVE_MINUTES
27}
28
29/** `$.fs.read` stops at 4 MiB, so transcript ends come from `head` and `tail`;
30 * null where that cannot be read. */
31async function transcriptEnd($: any, tool: 'head' | 'tail', transcriptPath: string): Promise<string[] | null> {
32 const run = await $.process.run([`/usr/bin/${tool}`, '-c', TRANSCRIPT_TAIL_BYTES, transcriptPath]).catch(() => null)
33 return run && run.exitCode === 0 ? run.stdout.split('\n') : null
34}
35
36/** The cache lifetime the last response actually wrote, which drops to five
37 * minutes in overage. */
38function writtenTtl(lines: string[]): number | null {
39 for (let i = lines.length - 1; i >= 0; i--) {
40 if (!lines[i].includes('"cache_creation"')) continue
41 try {
42 const creation = JSON.parse(lines[i])?.message?.usage?.cache_creation
43 if (creation?.ephemeral_1h_input_tokens) return HOUR
44 if (creation?.ephemeral_5m_input_tokens) return FIVE_MINUTES
45 } catch {}
46 }
47 return null
48}
49
50/** The thread's title as the app last set it, from a transcript's end. */
51function threadTitle(lines: string[]): string | null {
52 for (let i = lines.length - 1; i >= 0; i--) {
53 if (!lines[i].includes('"custom-title"')) continue
54 try {
55 const title = JSON.parse(lines[i])?.customTitle
56 if (typeof title === 'string' && title) return title
57 } catch {}
58 }
59 return null
60}
61
62/** The coldest step `left` of `ttlMs` has passed, as an index into
63 * `COLD_STEPS`; -1 above the first, null once cold. */
64export function coldStep(left: number, ttlMs: number): number | null {
65 if (left <= 0) return null
66 let step = -1
67 COLD_STEPS.forEach((share, i) => { if (left <= share * ttlMs) step = i })
68 return step
69}
70
71/** A Notification Centre banner, so a ping reaches him outside the thread. The
72 * text goes in as arguments, never into the script. macOS only. */
73async function notify($: any, title: string, text: string) {
74 await $.process.run(['/usr/bin/osascript',
75 '-e', 'on run argv', '-e', 'display notification (item 2 of argv) with title (item 1 of argv)', '-e', 'end run',
76 title, text]).catch(() => null)
77}
78
79/** An iMessage from this Mac's Messages account, which reaches the phone. */
80async function sendIMessage($: any, to: string, text: string) {
81 await $.process.run(['/usr/bin/osascript',
82 '-e', 'on run argv', '-e', 'tell application "Messages" to send (item 2 of argv) to participant (item 1 of argv) of (1st account whose service type = iMessage)', '-e', 'end run',
83 to, text]).catch(() => null)
84}
85
86type Usage = {
87 input_tokens?: number
88 cache_creation_input_tokens?: number
89 cache_read_input_tokens?: number
90 output_tokens?: number
91 cache_creation?: { ephemeral_1h_input_tokens?: number; ephemeral_5m_input_tokens?: number }
92}
93
94/** API prices relative to uncached input. How the plan's quotas weigh token
95 * kinds is not published, so the multiplier assumes the same ratios. */
96const READ_RATE = 0.1
97const WRITE_5M_RATE = 1.25
98const WRITE_1H_RATE = 2
99const OUTPUT_RATE = 5
100
101const sentTokens = (u: Usage) => (u.input_tokens ?? 0) + (u.cache_creation_input_tokens ?? 0) + (u.cache_read_input_tokens ?? 0)
102
103function writeRate(u: Usage): number {
104 const hour = u.cache_creation?.ephemeral_1h_input_tokens ?? 0
105 const five = u.cache_creation?.ephemeral_5m_input_tokens ?? 0
106 return hour + five > 0 ? (hour * WRITE_1H_RATE + five * WRITE_5M_RATE) / (hour + five) : WRITE_5M_RATE
107}
108
109/** Quota used per token kind, relative to a cache write: measured by
110 * `fitWeights`; null means the API price ratios. */
111type Weights = { write: number; read: number; output: number }
112
113function cost(u: Usage, w: Weights | null): number {
114 const written = (u.input_tokens ?? 0) + (u.cache_creation_input_tokens ?? 0)
115 if (w) return written * w.write + (u.cache_read_input_tokens ?? 0) * w.read + (u.output_tokens ?? 0) * w.output
116 return (u.input_tokens ?? 0) + (u.cache_creation_input_tokens ?? 0) * writeRate(u)
117 + (u.cache_read_input_tokens ?? 0) * READ_RATE + (u.output_tokens ?? 0) * OUTPUT_RATE
118}
119
120/** What resending `old` tokens of earlier context cost this request: read from
121 * the cache first, then written to it, then sent uncached. */
122function carried(u: Usage, old: number, w: Weights | null): number {
123 const read = Math.min(old, u.cache_read_input_tokens ?? 0)
124 const written = Math.min(old - read, u.cache_creation_input_tokens ?? 0)
125 return read * (w?.read ?? READ_RATE) + written * (w?.write ?? writeRate(u)) + (old - read - written) * (w?.write ?? 1)
126}
127
128type Entry = { type?: string; timestamp?: string; isMeta?: boolean; isSidechain?: boolean; message?: { id?: string; content?: unknown; usage?: Usage } }
129
130type Request = { at: number; usage: Usage }
131
132const parse = (line: string): Entry | null => { try { return JSON.parse(line) } catch { return null } }
133
134const isPrompt = (e: Entry) => e.type === 'user' && !e.isMeta && !e.isSidechain
135 && !(Array.isArray(e.message?.content) && e.message!.content.some((b: any) => b?.type === 'tool_result'))
136
137/** Context the session's first request sent: system prompt, tools, rules and
138 * the first prompt, the floor a fresh thread starts from. */
139function baselineTokens(head: string[]): number | null {
140 for (const line of head) {
141 if (!line.includes('"usage"')) continue
142 const e = parse(line)
143 if (e?.type === 'assistant' && !e.isSidechain && e.message?.usage) return sentTokens(e.message.usage)
144 }
145 return null
146}
147
148/** The last turn's requests, newest first. A response spans several transcript
149 * lines that repeat its usage, so each message id counts once. */
150function lastTurn(tail: string[]): Request[] {
151 const seen = new Set<string>()
152 const requests: Request[] = []
153 for (let i = tail.length - 1; i >= 0; i--) {
154 const e = parse(tail[i])
155 if (!e || e.isSidechain) continue
156 if (isPrompt(e)) break
157 const id = e.message?.id
158 if (e.type !== 'assistant' || !e.message?.usage || !id || seen.has(id)) continue
159 seen.add(id)
160 requests.push({ at: Date.parse(e.timestamp ?? '') || 0, usage: e.message.usage })
161 }
162 return requests
163}
164
165/** The turn's cost over what the same turn costs on a fresh thread: every
166 * request in it resends the context the turn started with beyond `baseline`. */
167function turnMultiplier(turn: Request[], baseline: number, w: Weights | null): number | null {
168 if (turn.length === 0) return null
169 const old = Math.max(0, sentTokens(turn[turn.length - 1].usage) - baseline)
170 const actual = turn.reduce((sum, r) => sum + cost(r.usage, w), 0)
171 const fresh = turn.reduce((sum, r) => sum + cost(r.usage, w) - carried(r.usage, old, w), 0)
172 return fresh > 0 ? actual / fresh : null
173}
174
175/** One session's calibration record in `$.store`: requests as [unix s, tokens
176 * written or sent uncached, cache read, output], quota readings as [unix s,
177 * 5h %, 7d %]. */
178type Log = { updated: number; requests: number[][]; readings: (number | null)[][] }
179
180const LOG_PREFIX = 'calibration/'
181const KEEP_MS = 8 * 86400e3
182const READING_EVERY_S = 300
183const REFIT_MS = 10 * 60e3
184/** Guesses: a quota step large enough that the one-decimal reading and request
185 * timing are noise, a gap after which usage elsewhere (claude.ai, another
186 * machine) is likely, and enough intervals to fit two weights. */
187const MIN_STEP = 1
188const MAX_GAP_S = 30 * 60
189const MIN_INTERVALS = 20
190
191const requestRow = (r: Request): number[] => [
192 Math.round(r.at / 1e3),
193 (r.usage.input_tokens ?? 0) + (r.usage.cache_creation_input_tokens ?? 0),
194 r.usage.cache_read_input_tokens ?? 0,
195 r.usage.output_tokens ?? 0,
196]
197
198/** Output per cache-write token at API prices with hour-long writes. Output is
199 * about a tenth of a turn's quota use, below what the readings can resolve, so
200 * it stays at this ratio and only writes and reads are measured. */
201const OUTPUT_PER_WRITE = OUTPUT_RATE / WRITE_1H_RATE
202/** Guesses: an interval using this much more than the fit predicts had usage
203 * this log cannot see; the read weight is trusted once its resampled 10th to
204 * 90th percentile range stays within this share of it. */
205const OUTLIER = 1.3
206const MAX_SPREAD = 0.25
207const RESAMPLES = 100
208
209type Fit = { kind: string; intervals: number; dropped: number; writePerPoint: number | null; weights: Weights | null; readRange: [number, number] | null }
210
211const usable = (fit: Fit | null): fit is Fit & { weights: Weights; readRange: [number, number] } =>
212 !!fit?.weights && !!fit.readRange && fit.intervals >= MIN_INTERVALS
213 && fit.readRange[1] - fit.readRange[0] <= 2 * MAX_SPREAD * fit.weights.read
214
215/** Intervals in which a window rose at least `MIN_STEP`, as rows of [written +
216 * weighted output, cache read] in millions of tokens and the points used. */
217function quotaIntervals(logs: Log[], column: 1 | 2): { rows: number[][]; ys: number[] } {
218 const readings = logs.flatMap(l => l.readings)
219 .filter(r => typeof r[column] === 'number')
220 .map(r => [r[0] as number, r[column] as number])
221 .sort((a, b) => a[0] - b[0])
222 const requests = logs.flatMap(l => l.requests)
223 const rows: number[][] = []
224 const ys: number[] = []
225 let start = readings[0]
226 for (let i = 1; i < readings.length; i++) {
227 const [t, p] = readings[i]
228 const [prevT, prevP] = readings[i - 1]
229 if (p < prevP || t - prevT > MAX_GAP_S) { start = readings[i]; continue }
230 if (p - start[1] < MIN_STEP) continue
231 const row = [0, 0]
232 for (const q of requests) {
233 if (q[0] <= start[0] || q[0] > t) continue
234 row[0] += (q[1] + q[3] * OUTPUT_PER_WRITE) / 1e6
235 row[1] += q[2] / 1e6
236 }
237 rows.push(row)
238 ys.push(p - start[1])
239 start = readings[i]
240 }
241 return { rows, ys }
242}
243
244/** Least squares for points = rows·[W, R] with W > 0 and R ≥ 0. */
245function leastSquares(rows: number[][], ys: number[]): [number, number] | null {
246 let aa = 0, ab = 0, bb = 0, ay = 0, by = 0
247 rows.forEach(([a, b], i) => { aa += a * a; ab += a * b; bb += b * b; ay += a * ys[i]; by += b * ys[i] })
248 const det = aa * bb - ab * ab
249 if (det > 1e-12) {
250 const w = (ay * bb - by * ab) / det
251 const r = (by * aa - ay * ab) / det
252 if (w > 0 && r >= 0) return [w, r]
253 }
254 return aa > 0 && ay > 0 ? [ay / aa, 0] : null
255}
256
257/** Fits the points each interval used, drops intervals well above the fit as
258 * usage elsewhere (claude.ai, another machine) and fits again; the read weight's
259 * range comes from refitting resampled intervals. */
260function fitWeights(logs: Log[], column: 1 | 2, kind: string): Fit {
261 const { rows, ys } = quotaIntervals(logs, column)
262 const none: Fit = { kind, intervals: rows.length, dropped: 0, writePerPoint: null, weights: null, readRange: null }
263 const first = leastSquares(rows, ys)
264 if (!first) return none
265 const kept = rows.map((r, i) => i).filter(i => ys[i] <= OUTLIER * (rows[i][0] * first[0] + rows[i][1] * first[1]))
266 const keptRows = kept.map(i => rows[i])
267 const keptYs = kept.map(i => ys[i])
268 const fit = leastSquares(keptRows, keptYs)
269 if (!fit) return none
270 let seed = 1
271 const random = () => (seed = (seed * 1103515245 + 12345) % 2147483648) / 2147483648
272 const reads: number[] = []
273 for (let n = 0; n < RESAMPLES; n++) {
274 const pick = keptRows.map(() => Math.floor(random() * keptRows.length))
275 const again = leastSquares(pick.map(i => keptRows[i]), pick.map(i => keptYs[i]))
276 if (again) reads.push(again[1] / again[0])
277 }
278 reads.sort((a, b) => a - b)
279 return {
280 kind,
281 intervals: kept.length,
282 dropped: rows.length - kept.length,
283 writePerPoint: 1e6 / fit[0],
284 weights: { write: 1, read: fit[1] / fit[0], output: OUTPUT_PER_WRITE },
285 readRange: reads.length >= RESAMPLES / 2 ? [reads[Math.floor(reads.length * 0.1)], reads[Math.floor(reads.length * 0.9)]] : null,
286 }
287}
288
289function fitText(fit: Fit): string {
290 const head = `${fit.kind}: ${fit.intervals} intervals` + (fit.dropped ? `, ${fit.dropped} dropped as usage elsewhere` : '')
291 if (!fit.weights) return `${head}, no fit yet`
292 const range = fit.readRange ? ` (${fit.readRange[0].toFixed(3)} to ${fit.readRange[1].toFixed(3)})` : ''
293 return `${head}; 1% = ${tokensText(fit.writePerPoint)} cache-write tokens; cache read = ${fit.weights.read.toFixed(3)}${range} of a cache write, API 0.05 to 0.08`
294}
295
296function tokensText(tokens: number | null): string {
297 if (tokens === null) return ''
298 return tokens >= 1e6 ? `${(tokens / 1e6).toFixed(1)}M` : `${Math.round(tokens / 1e3)}k`
299}
300
301const windowMs: Record<string, number> = { five_hour: 5 * 3600e3, seven_day: 7 * 86400e3 }
302const labels: Record<string, string> = { five_hour: '5h', seven_day: '7d', spend_limit: 'spend' }
303const days = ['Sun', 'Mon', 'Tue', 'Wed', 'Thu', 'Fri', 'Sat']
304
305const pad = (n: number) => String(n).padStart(2, '0')
306
307function resetText(limit: Limit, now: number): string {
308 if (!limit.resetsAt) return ''
309 const at = new Date(limit.resetsAt)
310 const time = `${pad(at.getHours())}:${pad(at.getMinutes())}`
311 return at.getTime() - now < 20 * 3600e3 ? time : `${days[at.getDay()]} ${time}`
312}
313
314/** Share of the window gone, 0 to 1, or null without a running window. */
315function elapsed(limit: Limit, now: number): number | null {
316 const length = windowMs[limit.kind]
317 if (!length || !limit.resetsAt) return null
318 const left = new Date(limit.resetsAt).getTime() - now
319 return left > 0 && left <= length ? 1 - left / length : null
320}
321
322/** Red from 95%, orange from 80%, yellow while ahead of an even pace, else white. */
323function tint(percent: number, gone: number | null): string {
324 if (percent >= 95) return '#ff453a'
325 if (percent >= 80) return '#ff8c00'
326 return gone !== null && percent / 100 > gone + 0.05 ? '#ffd60a' : '#f5f5f7'
327}
328
329/** Where the current rate lands by the reset; none in the first tenth of the window. */
330function projected(limit: Limit, now: number): number | null {
331 const gone = elapsed(limit, now)
332 if (gone === null || gone < 0.1 || limit.percentUsed <= 0) return null
333 return Math.min(Math.round(limit.percentUsed / gone), 100)
334}
335
336/** Hour notches for the 5h window, midnights for the week. */
337function notches(limit: Limit): number[] {
338 const length = windowMs[limit.kind]
339 if (!length || !limit.resetsAt) return []
340 const end = new Date(limit.resetsAt).getTime()
341 const start = end - length
342 const step = length > 86400e3 ? 'day' : 'hour'
343 const line = new Date(start)
344 if (step === 'day') line.setHours(0, 0, 0, 0)
345 else line.setMinutes(0, 0, 0)
346 const marks: number[] = []
347 for (;;) {
348 if (step === 'day') line.setDate(line.getDate() + 1)
349 else line.setHours(line.getHours() + 1)
350 if (line.getTime() >= end) return marks
351 marks.push((line.getTime() - start) / length)
352 }
353}
354
355function bar(limit: Limit, now: number, width: number, height: number): string {
356 const x = (share: number) => (Math.min(Math.max(share, 0), 1) * width).toFixed(1)
357 const used = limit.percentUsed / 100
358 const gone = elapsed(limit, now)
359 const color = tint(limit.percentUsed, gone)
360 const ahead = projected(limit, now)
361 let svg = `<svg xmlns="http://www.w3.org/2000/svg" width="${width}" height="${height}" viewBox="0 0 ${width} ${height}">`
362 + `<rect width="${width}" height="${height}" rx="2" fill="#000" fill-opacity="0.4"/>`
363 if (ahead !== null && ahead > limit.percentUsed) {
364 svg += `<rect x="${x(used)}" y="0.5" width="${(Number(x(ahead / 100)) - Number(x(used))).toFixed(1)}" height="${height - 1}" fill="none" stroke="${color}" stroke-opacity="0.5" stroke-dasharray="2 2"/>`
365 }
366 svg += `<rect width="${x(used)}" height="${height}" rx="2" fill="${color}"/>`
367 for (const mark of notches(limit)) svg += `<rect x="${x(mark)}" y="0" width="1" height="${height}" fill="#000" fill-opacity="0.35"/>`
368 if (gone !== null) svg += `<rect x="${(Number(x(gone)) - 1).toFixed(1)}" y="0" width="2" height="${height}" fill="#8b8d98"/>`
369 return svg + `</svg>`
370}
371
372type Options = { closeCommand?: string; renamePrompt?: string; iMessageTo?: string }
373
374/** A plain filled bar for a share of 0 to 1. */
375function meter(share: number, color: string, width: number, height: number): string {
376 const fill = (Math.min(Math.max(share, 0), 1) * width).toFixed(1)
377 return `<svg xmlns="http://www.w3.org/2000/svg" width="${width}" height="${height}" viewBox="0 0 ${width} ${height}">`
378 + `<rect width="${width}" height="${height}" rx="2" fill="#000" fill-opacity="0.4"/>`
379 + `<rect width="${fill}" height="${height}" rx="2" fill="${color}"/></svg>`
380}
381
382/** The Claude Profiles Mac app shows each account's quota; the figures every response carries
383 * spare it a request to Anthropic's rate-limited usage endpoint. Written only where that app is
384 * installed, one file per account (Desktop sessions name it) or per CLI config folder. */
385async function handOff($: any, rateLimits: unknown[]): Promise<void> {
386 const home = await $.env.get('HOME')
387 if (!home) return
388 const root = home + '/Library/Application Support/Claude Profiles'
389 if (!(await $.fs.exists(root))) return
390 const account = await $.env.get('CLAUDE_CODE_ACCOUNT_UUID')
391 const configDir = await $.env.get('CLAUDE_CONFIG_DIR')
392 const name = account ? 'account-' + account : 'config-' + (configDir ?? 'default').replace(/[^A-Za-z0-9]+/g, '-')
393 const email = await $.env.get('CLAUDE_CODE_USER_EMAIL')
394 await $.fs.write(`${root}/mod-readings/${name}.json`, JSON.stringify({ at: new Date().toISOString(), account, email, configDir, rateLimits }))
395}
396
397type Calibration = { own: { key: string; log: Log } | null; fits: Fit[]; fittedAt: number }
398
399async function ownLog($: any, c: Calibration): Promise<{ key: string; log: Log }> {
400 const key = LOG_PREFIX + await $.session.id()
401 if (c.own?.key !== key) c.own = { key, log: ((await $.store.get(key)) as Log | undefined) ?? { updated: 0, requests: [], readings: [] } }
402 return c.own
403}
404
405/** Expired session logs are dropped; when the store refuses a write as
406 * over its 4 MiB, the oldest other log goes and the write is tried again. */
407async function calibrationLogs($: any): Promise<{ key: string; log: Log }[]> {
408 const now = await $.clock.now()
409 const found: { key: string; log: Log }[] = []
410 for (const key of (await $.store.keys()) as string[]) {
411 if (!key.startsWith(LOG_PREFIX)) continue
412 const log = (await $.store.get(key)) as Log | undefined
413 if (!log || now - log.updated > KEEP_MS) await $.store.delete(key)
414 else found.push({ key, log })
415 }
416 return found.sort((a, b) => a.log.updated - b.log.updated)
417}
418
419async function saveLog($: any, mine: { key: string; log: Log }) {
420 for (let attempt = 0; attempt < 3; attempt++) {
421 try {
422 await $.store.set(mine.key, mine.log)
423 return
424 } catch {
425 const oldest = (await calibrationLogs($)).find(l => l.key !== mine.key)
426 if (!oldest) break
427 await $.store.delete(oldest.key)
428 }
429 }
430 $.ui.toast('quota-band: calibration log not saved')
431}
432
433async function refit($: any, c: Calibration) {
434 const all = (await calibrationLogs($)).map(l => (l.key === c.own?.key ? c.own.log : l.log))
435 if (c.own && !all.includes(c.own.log)) all.push(c.own.log)
436 c.fits = [fitWeights(all, 1, '5h'), fitWeights(all, 2, '7d')]
437 c.fittedAt = await $.clock.now()
438}
439
440export const register: Register = (on, options: Options = {}) => {
441 let pinged = { at: 0, step: -1 }
442 let title: string | null = null
443 let observedTtl: number | null = null
444 let baseline: { path: string; tokens: number } | null = null
445 const calibration: Calibration = { own: null, fits: [], fittedAt: 0 }
446
447 on('command.run', { command: 'quota-weights' }, async $ => {
448 await refit($, calibration)
449 const used = usable(calibration.fits[0]) ? 'measured 5h weights' : `API price ratios until 5h has ${MIN_INTERVALS} intervals and a cache read range within ±${MAX_SPREAD * 100}%`
450 return { text: [...calibration.fits.map(fitText), `output stays at ${OUTPUT_PER_WRITE} cache writes per token`, `cost × uses ${used}`].join('\n') }
451 })
452
453 on('session.start', async ($, e, next) => {
454 const result = await next(e)
455 await $.command.register({ name: 'quota-weights', description: 'How much of the 5h and 7d quota cache reads and output use, measured against cache writes' })
456 await ownLog($, calibration)
457 await refit($, calibration)
458 const usage = await $.session.usage()
459 if (usage.rateLimits.length > 0) await update($, limits, () => usage.rateLimits)
460 $.clock.every(60e3, () => { update($, tick, n => n + 1) })
461 $.clock.every(COLD_CHECK_MS, () => {
462 read($, turn).then(async last => {
463 if (!last) return
464 const left = last.at + last.ttlMs - Date.now()
465 const step = coldStep(left, last.ttlMs)
466 if (step === null || step < 0) return
467 if (pinged.at === last.at && pinged.step >= step) return
468 pinged = { at: last.at, step }
469 const minutes = left >= 60e3 ? `${Math.round(left / 60e3)} min` : `${Math.round(left / 1e3)} s`
470 const text = `Prompt cache ${COLD_STEPS[step] * 100}% warm · cold in ${minutes} · ${tokensText(last.tokens)} to rebuild`
471 $.ui.toast(text, { timeoutMs: 10e3 })
472 const folder = (await $.session.cwd().catch(() => '')).split('/').pop()
473 await notify($, title ?? folder ?? 'Claude', text)
474 if (options.iMessageTo) await sendIMessage($, String(options.iMessageTo), `${title ?? folder ?? 'Claude'}: ${text}`)
475 })
476 })
477 return result
478 })
479
480 on('session.measure', async ($, e, next) => {
481 if (e.rateLimits.length > 0) {
482 await update($, limits, () => e.rateLimits)
483 handOff($, e.rateLimits).catch(() => {})
484 }
485 const ttlMs = observedTtl ?? await cacheTtl($, e.rateLimits.length > 0)
486 const at = await $.clock.now()
487 await update($, turn, () => ({ at, tokens: e.context.tokens ?? null, window: e.context.window ?? null, ttlMs }))
488 const five = e.rateLimits.find(l => l.kind === 'five_hour')?.percentUsed ?? null
489 const seven = e.rateLimits.find(l => l.kind === 'seven_day')?.percentUsed ?? null
490 if (five !== null || seven !== null) {
491 const { log } = await ownLog($, calibration)
492 const last = log.readings[log.readings.length - 1]
493 const t = Math.round(at / 1e3)
494 if (!last || last[1] !== five || last[2] !== seven || t - (last[0] as number) >= READING_EVERY_S) log.readings.push([t, five, seven])
495 }
496 return next(e)
497 })
498
499 on('classic.Stop', async ($, e, next) => {
500 const result = await next(e)
501 const path: string | undefined = (e as any).transcript_path
502 if (!path) return result
503 const tail = await transcriptEnd($, 'tail', path)
504 if (!tail) return result
505 title = threadTitle(tail) ?? title
506 const ttlMs = writtenTtl(tail)
507 if (ttlMs !== null) {
508 observedTtl = ttlMs
509 await update($, turn, last => (last ? { ...last, ttlMs } : last))
510 }
511 if (baseline?.path !== path) {
512 const head = await transcriptEnd($, 'head', path)
513 const tokens = head ? baselineTokens(head) : null
514 baseline = tokens === null ? null : { path, tokens }
515 }
516 const requests = lastTurn(tail)
517 const mine = await ownLog($, calibration)
518 const loggedUntil = mine.log.requests[mine.log.requests.length - 1]?.[0] ?? 0
519 mine.log.requests.push(...requests.map(requestRow).filter(r => r[0] > loggedUntil).reverse())
520 mine.log.updated = await $.clock.now()
521 await saveLog($, mine)
522 if (mine.log.updated - calibration.fittedAt >= REFIT_MS) await refit($, calibration)
523 if (baseline) {
524 const fit = calibration.fits[0] ?? null
525 const weights = usable(fit) ? fit.weights : null
526 const value = turnMultiplier(requests, baseline.tokens, weights)
527 await update($, multiplier, () => (value === null ? null : { value, measured: weights !== null }))
528 }
529 return result
530 })
531
532 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
533 const shown = await read($, limits)
534 const last = await read($, turn)
535 const times = await read($, multiplier)
536 await read($, tick)
537 const below = await next(e)
538 if (e.props.hasSurvey || (shown.length === 0 && !last)) return below
539 const now = await $.clock.now()
540 const ui = $.ui.resolve(e)
541 const { Box, Text, Button } = ui
542 const left = last ? last.at + last.ttlMs - now : null
543 const cache = left === null || !last ? null
544 : left > 0 ? { text: `${Math.ceil(left / 60e3)}m`, share: left / last.ttlMs, color: left <= COLD_SOON ? '#ff8c00' : '#f5f5f7' }
545 : { text: 'cold', share: 0, color: '#ff453a' }
546 const context = last?.tokens && last.window ? { share: last.tokens / last.window, percent: Math.round((last.tokens / last.window) * 100) } : null
547 const buttons = [
548 options.closeCommand ? { key: 'close', label: 'close', text: String(options.closeCommand) } : null,
549 options.renamePrompt ? { key: 'rename', label: 'rename', text: String(options.renamePrompt) } : null,
550 ].filter(Boolean) as { key: string; label: string; text: string }[]
551 const line = shown.map(l => `${labels[l.kind] ?? l.kind} ${Math.round(l.percentUsed)}%` + (l.resetsAt ? ` · ${resetText(l, now)}` : '')).join(' ')
552 + (context ? ` ctx ${context.percent}% · ${tokensText(last!.tokens)}` : '')
553 + (cache ? ` cache ${cache.text}` : '')
554 + (times ? ` cost ×${times.value.toFixed(1)}${times.measured ? ' measured' : ''}` : '')
555 // Wider than any terminal; the divider row clips it to the band's width.
556 const divider = below ? <Box height={1} overflow="hidden"><Text dimColor>{'─'.repeat(400)}</Text></Box> : null
557 if (e.surface !== 'desktop' || !('Svg' in ui)) return <Box flexDirection="column"><Text dimColor>{line}</Text>{divider}{below}</Box>
558 const { Svg } = ui as any
559
560 return (
561 <Box flexDirection="column">
562 <Box flexDirection="row" alignItems="center" columnGap={2}>
563 {shown.map(l => (
564 <Box key={l.kind} flexDirection="row" alignItems="center" columnGap={1}>
565 <Text dimColor>{labels[l.kind] ?? l.kind}</Text>
566 <Svg source={bar(l, now, 120, 8)} alt={`${l.kind} ${l.percentUsed}%`} width={120} height={8} />
567 <Text>{Math.round(l.percentUsed)}%</Text>
568 {projected(l, now) !== null && projected(l, now)! > Math.round(l.percentUsed) ? <Text dimColor>→ {projected(l, now)}%</Text> : null}
569 {l.resetsAt ? <Text dimColor>{resetText(l, now)}</Text> : null}
570 </Box>
571 ))}
572 {context ? (
573 <Box flexDirection="row" alignItems="center" columnGap={1}>
574 <Text dimColor>ctx</Text>
575 <Svg source={meter(context.share, tint(context.percent, null), 60, 8)} alt={`context ${context.percent}%`} width={60} height={8} />
576 <Text>{context.percent}%</Text>
577 <Text dimColor>{tokensText(last!.tokens)}</Text>
578 </Box>
579 ) : null}
580 {cache ? (
581 <Box flexDirection="row" alignItems="center" columnGap={1}>
582 <Text dimColor>cache</Text>
583 <Svg source={meter(cache.share, cache.color, 60, 8)} alt={`cache ${cache.text}`} width={60} height={8} />
584 <Text color={cache.color === '#f5f5f7' ? undefined : cache.color}>{cache.text}</Text>
585 </Box>
586 ) : null}
587 {times ? (
588 <Box flexDirection="row" alignItems="center" columnGap={1}>
589 <Text dimColor>cost</Text>
590 <Text>×{times.value.toFixed(1)}</Text>
591 {times.measured ? <Text dimColor>measured</Text> : null}
592 </Box>
593 ) : null}
594 {buttons.map(b => (
595 <Button key={b.key} label={b.label} onPress={() => { $.prompt.submit({ text: b.text, asUser: true }) }} />
596 ))}
597 </Box>
598 {divider}
599 {below}
600 </Box>
601 )
602 })
603}
604types/index.d.ts 12 lines1export type Limit = { kind: string; percentUsed: number; resetsAt?: string }
2
3export type Multiplier = { value: number; measured: boolean }
4
5export type Turn = { at: number; tokens: number | null; window: number | null; ttlMs: number }
6
7declare module 'claude-code' {
8 interface PluginState {
9 'quota-band': { limits: Limit[]; turn: Turn | null; multiplier: Multiplier | null; tick: number }
10 }
11}
12