SLOPSHOPPER

quota-band

Your Claude plan's 5h and weekly usage, the context window and the prompt cache's time left, as one line of bars above the prompt

newbandcommandtoastprocesstimer
v0.7.0MITupdated 2026-10-07andras-gyarmati/claude-quota-band
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · quota-band
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /quota-weights ⎿ quota-band: 5h: 0 intervals, no fit yet ⎿ quota-band: 7d: 0 intervals, no fit yet ⎿ quota-band: output stays at 2.5 cache writes per token ⎿ quota-band: cost × uses API price ratios until 5h has 20 intervals and a cache read range within ±25% 5h 31% · 09:53 ctx 49% · 97k cache 60m ──────────────────────────────────────────────────────────────────────────────────────────────────────────────────── ──────────────────────────────────────────────────────────────────────────────────────────────────────────────────── ──────────────────────────────────────────────────────────────────────────────────────────────────────────────────── ──────────────────────────────────────────────────── ⟨Claude Code's own drawing⟩ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Band
5h 31% · 09:53 ctx 49% · 97k cache 60m ──────────────────────────────────────────────────────────────────────────────────────────────────── ──────────────────────────────────────────────────────────────────────────────────────────────────── ──────────────────────────────────────────────────────────────────────────────────────────────────── ──────────────────────────────────────────────────────────────────────────────────────────────────── ⟨Claude Code's own drawing⟩
README

quota-band

A Claude Code mod that keeps your plan's usage in view: one line above the prompt with the 5-hour and weekly windows, the context window, the prompt cache's time left and what the context adds to each turn's cost.

5h ▬▬▬▬▬▬▬▬▬ 93% → 100% 23:10    7d ▬▬▬▬▬▬▬ 72% 11:00    ctx ▬▬▬ 37% 370k    cache ▬▬▬ 42m    cost ×4.5
  • Each bar fills with the percent used. A grey tick marks how much of the window has passed, and the notches are hours (5h) or days (7d).
  • Colours: red from 95%, orange from 80%, yellow while you are ahead of an even pace, white otherwise.
  • → N% is where the current pace lands by the reset, with a dashed outline on the bar.
  • The time is when the window resets.
  • ctx is how full the context window is; 370k is what the next request re-sends, which a cold cache has to write again.
  • cache drains over the time the prompt cache stays warm after the last response: the length the last response actually wrote, read from the transcript at the end of each turn, so it drops to five minutes in overage. Until the first turn ends (and on Windows, which has no tail) it assumes an hour on a subscription's main thread, five minutes otherwise, or what promptCacheTtl sets. It turns orange in the last five minutes and red once cold. When 50%, 25%, 10% and 5% of the lifetime is left, every thread pings with a toast and, on macOS, a Notification Centre banner titled with the thread's name (osascript, so the banner comes from Script Editor, which macOS may ask to allow once). Set iMessageTo (below) to a phone number or Apple ID to get the same pings on the phone, sent through this Mac's Messages app; macOS asks once to let Claude control Messages.
  • cost ×4.5 is what the last turn cost over the same turn in a fresh thread. Every request in a turn (one per tool call) resends the context, so a long thread costs more per message even with a warm cache. The fresh thread starts from what the session's first request sent (system prompt, tools, rules). Token kinds are weighed at API price ratios: cache read 0.1, cache write 1.25 (5 minutes) or 2 (1 hour), output 5, against uncached input 1. How the plan's quotas weigh them is not published, so treat it as an estimate.
  • Once the plan's own readings have measured the weights (below), cost uses those and says measured.

Measured quota weights

Each session logs its requests' token counts and the 5h and 7d readings to the plugin's store, kept for eight days. Readings from all sessions are cut into intervals in which a window rose at least one point, and a least-squares fit gives how many cache-write tokens make 1% and what a cache read weighs against a cache write. Output stays at the API ratio of 2.5 cache writes per token: it is about a tenth of the quota use, too little to separate from the noise. /quota-weights prints the fit:

5h: 52 intervals, 3 dropped as usage elsewhere; 1% = 2.0M cache-write tokens; cache read = 0.052 (0.048 to 0.058) of a cache write, API 0.05 to 0.08
7d: 6 intervals; 1% = 14.1M cache-write tokens; cache read = 0.061 (0.030 to 0.090) of a cache write, API 0.05 to 0.08
output stays at 2.5 cache writes per token
cost × uses measured 5h weights

The readings cover the whole account, so usage the log cannot see (claude.ai, another machine, subagents) lands in them too. Intervals that used 30% more than the fit predicts are dropped as such, intervals across a 30-minute silence are skipped, and cost switches to measured weights only after 20 intervals with the cache read's 10th to 90th percentile range within ±25%.

Buttons

Two optional buttons send a prompt of your choice, for example a wrap-up skill. Set them in /config under the plugin's options, or in settings.json:

"pluginConfigs": {
  "quota-band@claude-quota-band": {
    "options": { "closeCommand": "/close-thread", "renamePrompt": "Rename this thread to cover every topic in it", "iMessageTo": "" }
  }
}

Empty hides a button. A press during a turn waits until the turn ends.

It reads the figures Claude Code already receives with each response (session.measure), and the ends of the session's transcript at the end of each turn, so it sends no requests of its own. It writes to its own plugin store and, where the Claude Profiles Mac app is installed, the latest figures per account to ~/Library/Application Support/Claude Profiles/mod-readings/, which that app shows instead of asking Anthropic. The figures appear after the first response of a session, on Pro, Max and Team plans.

Install

Needs Claude Code 2.1.287 or later (Claude Mods). In Claude Code:

/plugin marketplace add andras-gyarmati/claude-quota-band
/plugin install quota-band@claude-quota-band

To pick up new versions on each start, set "autoUpdate": true for this marketplace under extraKnownMarketplaces in settings.json.

Bars draw in the desktop app's Code tab; the terminal shows the same line as text.

Source 2 files
hooks/register.tsx 604 lines
1import { atom, read, update } from 'claude-code'
2import type { Register } from 'claude-code'
3
4import type { Limit, Multiplier, Turn } from '../types'
5
6const limits = atom({ plugin: 'quota-band', key: 'limits' } as const, [] as Limit[])
7const turn = atom({ plugin: 'quota-band', key: 'turn' } as const, null as Turn | null)
8const multiplier = atom({ plugin: 'quota-band', key: 'multiplier' } as const, null as Multiplier | null)
9const tick = atom({ plugin: 'quota-band', key: 'tick' } as const, 0)
10
11const HOUR = 3600e3
12const FIVE_MINUTES = 5 * 60e3
13const COLD_SOON = FIVE_MINUTES
14/** Shares of the cache lifetime left at which a thread pings, warmest first. */
15export const COLD_STEPS = [0.5, 0.25, 0.1, 0.05]
16const COLD_CHECK_MS = 10e3
17const TRANSCRIPT_TAIL_BYTES = '524288'
18
19/** The fallback before a transcript has been read: subscribers' main thread
20 * caches for an hour, everyone else for five minutes, unless `promptCacheTtl`
21 * says otherwise. */
22async function cacheTtl($: any, subscriber: boolean): Promise<number> {
23  const setting = (await $.settings.read())?.promptCacheTtl
24  if (setting === '1h') return HOUR
25  if (setting === '5m') return FIVE_MINUTES
26  return subscriber ? HOUR : FIVE_MINUTES
27}
28
29/** `$.fs.read` stops at 4 MiB, so transcript ends come from `head` and `tail`;
30 * null where that cannot be read. */
31async function transcriptEnd($: any, tool: 'head' | 'tail', transcriptPath: string): Promise<string[] | null> {
32  const run = await $.process.run([`/usr/bin/${tool}`, '-c', TRANSCRIPT_TAIL_BYTES, transcriptPath]).catch(() => null)
33  return run && run.exitCode === 0 ? run.stdout.split('\n') : null
34}
35
36/** The cache lifetime the last response actually wrote, which drops to five
37 * minutes in overage. */
38function writtenTtl(lines: string[]): number | null {
39  for (let i = lines.length - 1; i >= 0; i--) {
40    if (!lines[i].includes('"cache_creation"')) continue
41    try {
42      const creation = JSON.parse(lines[i])?.message?.usage?.cache_creation
43      if (creation?.ephemeral_1h_input_tokens) return HOUR
44      if (creation?.ephemeral_5m_input_tokens) return FIVE_MINUTES
45    } catch {}
46  }
47  return null
48}
49
50/** The thread's title as the app last set it, from a transcript's end. */
51function threadTitle(lines: string[]): string | null {
52  for (let i = lines.length - 1; i >= 0; i--) {
53    if (!lines[i].includes('"custom-title"')) continue
54    try {
55      const title = JSON.parse(lines[i])?.customTitle
56      if (typeof title === 'string' && title) return title
57    } catch {}
58  }
59  return null
60}
61
62/** The coldest step `left` of `ttlMs` has passed, as an index into
63 * `COLD_STEPS`; -1 above the first, null once cold. */
64export function coldStep(left: number, ttlMs: number): number | null {
65  if (left <= 0) return null
66  let step = -1
67  COLD_STEPS.forEach((share, i) => { if (left <= share * ttlMs) step = i })
68  return step
69}
70
71/** A Notification Centre banner, so a ping reaches him outside the thread. The
72 * text goes in as arguments, never into the script. macOS only. */
73async function notify($: any, title: string, text: string) {
74  await $.process.run(['/usr/bin/osascript',
75    '-e', 'on run argv', '-e', 'display notification (item 2 of argv) with title (item 1 of argv)', '-e', 'end run',
76    title, text]).catch(() => null)
77}
78
79/** An iMessage from this Mac's Messages account, which reaches the phone. */
80async function sendIMessage($: any, to: string, text: string) {
81  await $.process.run(['/usr/bin/osascript',
82    '-e', 'on run argv', '-e', 'tell application "Messages" to send (item 2 of argv) to participant (item 1 of argv) of (1st account whose service type = iMessage)', '-e', 'end run',
83    to, text]).catch(() => null)
84}
85
86type Usage = {
87  input_tokens?: number
88  cache_creation_input_tokens?: number
89  cache_read_input_tokens?: number
90  output_tokens?: number
91  cache_creation?: { ephemeral_1h_input_tokens?: number; ephemeral_5m_input_tokens?: number }
92}
93
94/** API prices relative to uncached input. How the plan's quotas weigh token
95 * kinds is not published, so the multiplier assumes the same ratios. */
96const READ_RATE = 0.1
97const WRITE_5M_RATE = 1.25
98const WRITE_1H_RATE = 2
99const OUTPUT_RATE = 5
100
101const sentTokens = (u: Usage) => (u.input_tokens ?? 0) + (u.cache_creation_input_tokens ?? 0) + (u.cache_read_input_tokens ?? 0)
102
103function writeRate(u: Usage): number {
104  const hour = u.cache_creation?.ephemeral_1h_input_tokens ?? 0
105  const five = u.cache_creation?.ephemeral_5m_input_tokens ?? 0
106  return hour + five > 0 ? (hour * WRITE_1H_RATE + five * WRITE_5M_RATE) / (hour + five) : WRITE_5M_RATE
107}
108
109/** Quota used per token kind, relative to a cache write: measured by
110 * `fitWeights`; null means the API price ratios. */
111type Weights = { write: number; read: number; output: number }
112
113function cost(u: Usage, w: Weights | null): number {
114  const written = (u.input_tokens ?? 0) + (u.cache_creation_input_tokens ?? 0)
115  if (w) return written * w.write + (u.cache_read_input_tokens ?? 0) * w.read + (u.output_tokens ?? 0) * w.output
116  return (u.input_tokens ?? 0) + (u.cache_creation_input_tokens ?? 0) * writeRate(u)
117    + (u.cache_read_input_tokens ?? 0) * READ_RATE + (u.output_tokens ?? 0) * OUTPUT_RATE
118}
119
120/** What resending `old` tokens of earlier context cost this request: read from
121 * the cache first, then written to it, then sent uncached. */
122function carried(u: Usage, old: number, w: Weights | null): number {
123  const read = Math.min(old, u.cache_read_input_tokens ?? 0)
124  const written = Math.min(old - read, u.cache_creation_input_tokens ?? 0)
125  return read * (w?.read ?? READ_RATE) + written * (w?.write ?? writeRate(u)) + (old - read - written) * (w?.write ?? 1)
126}
127
128type Entry = { type?: string; timestamp?: string; isMeta?: boolean; isSidechain?: boolean; message?: { id?: string; content?: unknown; usage?: Usage } }
129
130type Request = { at: number; usage: Usage }
131
132const parse = (line: string): Entry | null => { try { return JSON.parse(line) } catch { return null } }
133
134const isPrompt = (e: Entry) => e.type === 'user' && !e.isMeta && !e.isSidechain
135  && !(Array.isArray(e.message?.content) && e.message!.content.some((b: any) => b?.type === 'tool_result'))
136
137/** Context the session's first request sent: system prompt, tools, rules and
138 * the first prompt, the floor a fresh thread starts from. */
139function baselineTokens(head: string[]): number | null {
140  for (const line of head) {
141    if (!line.includes('"usage"')) continue
142    const e = parse(line)
143    if (e?.type === 'assistant' && !e.isSidechain && e.message?.usage) return sentTokens(e.message.usage)
144  }
145  return null
146}
147
148/** The last turn's requests, newest first. A response spans several transcript
149 * lines that repeat its usage, so each message id counts once. */
150function lastTurn(tail: string[]): Request[] {
151  const seen = new Set<string>()
152  const requests: Request[] = []
153  for (let i = tail.length - 1; i >= 0; i--) {
154    const e = parse(tail[i])
155    if (!e || e.isSidechain) continue
156    if (isPrompt(e)) break
157    const id = e.message?.id
158    if (e.type !== 'assistant' || !e.message?.usage || !id || seen.has(id)) continue
159    seen.add(id)
160    requests.push({ at: Date.parse(e.timestamp ?? '') || 0, usage: e.message.usage })
161  }
162  return requests
163}
164
165/** The turn's cost over what the same turn costs on a fresh thread: every
166 * request in it resends the context the turn started with beyond `baseline`. */
167function turnMultiplier(turn: Request[], baseline: number, w: Weights | null): number | null {
168  if (turn.length === 0) return null
169  const old = Math.max(0, sentTokens(turn[turn.length - 1].usage) - baseline)
170  const actual = turn.reduce((sum, r) => sum + cost(r.usage, w), 0)
171  const fresh = turn.reduce((sum, r) => sum + cost(r.usage, w) - carried(r.usage, old, w), 0)
172  return fresh > 0 ? actual / fresh : null
173}
174
175/** One session's calibration record in `$.store`: requests as [unix s, tokens
176 * written or sent uncached, cache read, output], quota readings as [unix s,
177 * 5h %, 7d %]. */
178type Log = { updated: number; requests: number[][]; readings: (number | null)[][] }
179
180const LOG_PREFIX = 'calibration/'
181const KEEP_MS = 8 * 86400e3
182const READING_EVERY_S = 300
183const REFIT_MS = 10 * 60e3
184/** Guesses: a quota step large enough that the one-decimal reading and request
185 * timing are noise, a gap after which usage elsewhere (claude.ai, another
186 * machine) is likely, and enough intervals to fit two weights. */
187const MIN_STEP = 1
188const MAX_GAP_S = 30 * 60
189const MIN_INTERVALS = 20
190
191const requestRow = (r: Request): number[] => [
192  Math.round(r.at / 1e3),
193  (r.usage.input_tokens ?? 0) + (r.usage.cache_creation_input_tokens ?? 0),
194  r.usage.cache_read_input_tokens ?? 0,
195  r.usage.output_tokens ?? 0,
196]
197
198/** Output per cache-write token at API prices with hour-long writes. Output is
199 * about a tenth of a turn's quota use, below what the readings can resolve, so
200 * it stays at this ratio and only writes and reads are measured. */
201const OUTPUT_PER_WRITE = OUTPUT_RATE / WRITE_1H_RATE
202/** Guesses: an interval using this much more than the fit predicts had usage
203 * this log cannot see; the read weight is trusted once its resampled 10th to
204 * 90th percentile range stays within this share of it. */
205const OUTLIER = 1.3
206const MAX_SPREAD = 0.25
207const RESAMPLES = 100
208
209type Fit = { kind: string; intervals: number; dropped: number; writePerPoint: number | null; weights: Weights | null; readRange: [number, number] | null }
210
211const usable = (fit: Fit | null): fit is Fit & { weights: Weights; readRange: [number, number] } =>
212  !!fit?.weights && !!fit.readRange && fit.intervals >= MIN_INTERVALS
213  && fit.readRange[1] - fit.readRange[0] <= 2 * MAX_SPREAD * fit.weights.read
214
215/** Intervals in which a window rose at least `MIN_STEP`, as rows of [written +
216 * weighted output, cache read] in millions of tokens and the points used. */
217function quotaIntervals(logs: Log[], column: 1 | 2): { rows: number[][]; ys: number[] } {
218  const readings = logs.flatMap(l => l.readings)
219    .filter(r => typeof r[column] === 'number')
220    .map(r => [r[0] as number, r[column] as number])
221    .sort((a, b) => a[0] - b[0])
222  const requests = logs.flatMap(l => l.requests)
223  const rows: number[][] = []
224  const ys: number[] = []
225  let start = readings[0]
226  for (let i = 1; i < readings.length; i++) {
227    const [t, p] = readings[i]
228    const [prevT, prevP] = readings[i - 1]
229    if (p < prevP || t - prevT > MAX_GAP_S) { start = readings[i]; continue }
230    if (p - start[1] < MIN_STEP) continue
231    const row = [0, 0]
232    for (const q of requests) {
233      if (q[0] <= start[0] || q[0] > t) continue
234      row[0] += (q[1] + q[3] * OUTPUT_PER_WRITE) / 1e6
235      row[1] += q[2] / 1e6
236    }
237    rows.push(row)
238    ys.push(p - start[1])
239    start = readings[i]
240  }
241  return { rows, ys }
242}
243
244/** Least squares for points = rows·[W, R] with W > 0 and R ≥ 0. */
245function leastSquares(rows: number[][], ys: number[]): [number, number] | null {
246  let aa = 0, ab = 0, bb = 0, ay = 0, by = 0
247  rows.forEach(([a, b], i) => { aa += a * a; ab += a * b; bb += b * b; ay += a * ys[i]; by += b * ys[i] })
248  const det = aa * bb - ab * ab
249  if (det > 1e-12) {
250    const w = (ay * bb - by * ab) / det
251    const r = (by * aa - ay * ab) / det
252    if (w > 0 && r >= 0) return [w, r]
253  }
254  return aa > 0 && ay > 0 ? [ay / aa, 0] : null
255}
256
257/** Fits the points each interval used, drops intervals well above the fit as
258 * usage elsewhere (claude.ai, another machine) and fits again; the read weight's
259 * range comes from refitting resampled intervals. */
260function fitWeights(logs: Log[], column: 1 | 2, kind: string): Fit {
261  const { rows, ys } = quotaIntervals(logs, column)
262  const none: Fit = { kind, intervals: rows.length, dropped: 0, writePerPoint: null, weights: null, readRange: null }
263  const first = leastSquares(rows, ys)
264  if (!first) return none
265  const kept = rows.map((r, i) => i).filter(i => ys[i] <= OUTLIER * (rows[i][0] * first[0] + rows[i][1] * first[1]))
266  const keptRows = kept.map(i => rows[i])
267  const keptYs = kept.map(i => ys[i])
268  const fit = leastSquares(keptRows, keptYs)
269  if (!fit) return none
270  let seed = 1
271  const random = () => (seed = (seed * 1103515245 + 12345) % 2147483648) / 2147483648
272  const reads: number[] = []
273  for (let n = 0; n < RESAMPLES; n++) {
274    const pick = keptRows.map(() => Math.floor(random() * keptRows.length))
275    const again = leastSquares(pick.map(i => keptRows[i]), pick.map(i => keptYs[i]))
276    if (again) reads.push(again[1] / again[0])
277  }
278  reads.sort((a, b) => a - b)
279  return {
280    kind,
281    intervals: kept.length,
282    dropped: rows.length - kept.length,
283    writePerPoint: 1e6 / fit[0],
284    weights: { write: 1, read: fit[1] / fit[0], output: OUTPUT_PER_WRITE },
285    readRange: reads.length >= RESAMPLES / 2 ? [reads[Math.floor(reads.length * 0.1)], reads[Math.floor(reads.length * 0.9)]] : null,
286  }
287}
288
289function fitText(fit: Fit): string {
290  const head = `${fit.kind}: ${fit.intervals} intervals` + (fit.dropped ? `, ${fit.dropped} dropped as usage elsewhere` : '')
291  if (!fit.weights) return `${head}, no fit yet`
292  const range = fit.readRange ? ` (${fit.readRange[0].toFixed(3)} to ${fit.readRange[1].toFixed(3)})` : ''
293  return `${head}; 1% = ${tokensText(fit.writePerPoint)} cache-write tokens; cache read = ${fit.weights.read.toFixed(3)}${range} of a cache write, API 0.05 to 0.08`
294}
295
296function tokensText(tokens: number | null): string {
297  if (tokens === null) return ''
298  return tokens >= 1e6 ? `${(tokens / 1e6).toFixed(1)}M` : `${Math.round(tokens / 1e3)}k`
299}
300
301const windowMs: Record<string, number> = { five_hour: 5 * 3600e3, seven_day: 7 * 86400e3 }
302const labels: Record<string, string> = { five_hour: '5h', seven_day: '7d', spend_limit: 'spend' }
303const days = ['Sun', 'Mon', 'Tue', 'Wed', 'Thu', 'Fri', 'Sat']
304
305const pad = (n: number) => String(n).padStart(2, '0')
306
307function resetText(limit: Limit, now: number): string {
308  if (!limit.resetsAt) return ''
309  const at = new Date(limit.resetsAt)
310  const time = `${pad(at.getHours())}:${pad(at.getMinutes())}`
311  return at.getTime() - now < 20 * 3600e3 ? time : `${days[at.getDay()]} ${time}`
312}
313
314/** Share of the window gone, 0 to 1, or null without a running window. */
315function elapsed(limit: Limit, now: number): number | null {
316  const length = windowMs[limit.kind]
317  if (!length || !limit.resetsAt) return null
318  const left = new Date(limit.resetsAt).getTime() - now
319  return left > 0 && left <= length ? 1 - left / length : null
320}
321
322/** Red from 95%, orange from 80%, yellow while ahead of an even pace, else white. */
323function tint(percent: number, gone: number | null): string {
324  if (percent >= 95) return '#ff453a'
325  if (percent >= 80) return '#ff8c00'
326  return gone !== null && percent / 100 > gone + 0.05 ? '#ffd60a' : '#f5f5f7'
327}
328
329/** Where the current rate lands by the reset; none in the first tenth of the window. */
330function projected(limit: Limit, now: number): number | null {
331  const gone = elapsed(limit, now)
332  if (gone === null || gone < 0.1 || limit.percentUsed <= 0) return null
333  return Math.min(Math.round(limit.percentUsed / gone), 100)
334}
335
336/** Hour notches for the 5h window, midnights for the week. */
337function notches(limit: Limit): number[] {
338  const length = windowMs[limit.kind]
339  if (!length || !limit.resetsAt) return []
340  const end = new Date(limit.resetsAt).getTime()
341  const start = end - length
342  const step = length > 86400e3 ? 'day' : 'hour'
343  const line = new Date(start)
344  if (step === 'day') line.setHours(0, 0, 0, 0)
345  else line.setMinutes(0, 0, 0)
346  const marks: number[] = []
347  for (;;) {
348    if (step === 'day') line.setDate(line.getDate() + 1)
349    else line.setHours(line.getHours() + 1)
350    if (line.getTime() >= end) return marks
351    marks.push((line.getTime() - start) / length)
352  }
353}
354
355function bar(limit: Limit, now: number, width: number, height: number): string {
356  const x = (share: number) => (Math.min(Math.max(share, 0), 1) * width).toFixed(1)
357  const used = limit.percentUsed / 100
358  const gone = elapsed(limit, now)
359  const color = tint(limit.percentUsed, gone)
360  const ahead = projected(limit, now)
361  let svg = `<svg xmlns="http://www.w3.org/2000/svg" width="${width}" height="${height}" viewBox="0 0 ${width} ${height}">`
362    + `<rect width="${width}" height="${height}" rx="2" fill="#000" fill-opacity="0.4"/>`
363  if (ahead !== null && ahead > limit.percentUsed) {
364    svg += `<rect x="${x(used)}" y="0.5" width="${(Number(x(ahead / 100)) - Number(x(used))).toFixed(1)}" height="${height - 1}" fill="none" stroke="${color}" stroke-opacity="0.5" stroke-dasharray="2 2"/>`
365  }
366  svg += `<rect width="${x(used)}" height="${height}" rx="2" fill="${color}"/>`
367  for (const mark of notches(limit)) svg += `<rect x="${x(mark)}" y="0" width="1" height="${height}" fill="#000" fill-opacity="0.35"/>`
368  if (gone !== null) svg += `<rect x="${(Number(x(gone)) - 1).toFixed(1)}" y="0" width="2" height="${height}" fill="#8b8d98"/>`
369  return svg + `</svg>`
370}
371
372type Options = { closeCommand?: string; renamePrompt?: string; iMessageTo?: string }
373
374/** A plain filled bar for a share of 0 to 1. */
375function meter(share: number, color: string, width: number, height: number): string {
376  const fill = (Math.min(Math.max(share, 0), 1) * width).toFixed(1)
377  return `<svg xmlns="http://www.w3.org/2000/svg" width="${width}" height="${height}" viewBox="0 0 ${width} ${height}">`
378    + `<rect width="${width}" height="${height}" rx="2" fill="#000" fill-opacity="0.4"/>`
379    + `<rect width="${fill}" height="${height}" rx="2" fill="${color}"/></svg>`
380}
381
382/** The Claude Profiles Mac app shows each account's quota; the figures every response carries
383 * spare it a request to Anthropic's rate-limited usage endpoint. Written only where that app is
384 * installed, one file per account (Desktop sessions name it) or per CLI config folder. */
385async function handOff($: any, rateLimits: unknown[]): Promise<void> {
386  const home = await $.env.get('HOME')
387  if (!home) return
388  const root = home + '/Library/Application Support/Claude Profiles'
389  if (!(await $.fs.exists(root))) return
390  const account = await $.env.get('CLAUDE_CODE_ACCOUNT_UUID')
391  const configDir = await $.env.get('CLAUDE_CONFIG_DIR')
392  const name = account ? 'account-' + account : 'config-' + (configDir ?? 'default').replace(/[^A-Za-z0-9]+/g, '-')
393  const email = await $.env.get('CLAUDE_CODE_USER_EMAIL')
394  await $.fs.write(`${root}/mod-readings/${name}.json`, JSON.stringify({ at: new Date().toISOString(), account, email, configDir, rateLimits }))
395}
396
397type Calibration = { own: { key: string; log: Log } | null; fits: Fit[]; fittedAt: number }
398
399async function ownLog($: any, c: Calibration): Promise<{ key: string; log: Log }> {
400  const key = LOG_PREFIX + await $.session.id()
401  if (c.own?.key !== key) c.own = { key, log: ((await $.store.get(key)) as Log | undefined) ?? { updated: 0, requests: [], readings: [] } }
402  return c.own
403}
404
405/** Expired session logs are dropped; when the store refuses a write as
406 * over its 4 MiB, the oldest other log goes and the write is tried again. */
407async function calibrationLogs($: any): Promise<{ key: string; log: Log }[]> {
408  const now = await $.clock.now()
409  const found: { key: string; log: Log }[] = []
410  for (const key of (await $.store.keys()) as string[]) {
411    if (!key.startsWith(LOG_PREFIX)) continue
412    const log = (await $.store.get(key)) as Log | undefined
413    if (!log || now - log.updated > KEEP_MS) await $.store.delete(key)
414    else found.push({ key, log })
415  }
416  return found.sort((a, b) => a.log.updated - b.log.updated)
417}
418
419async function saveLog($: any, mine: { key: string; log: Log }) {
420  for (let attempt = 0; attempt < 3; attempt++) {
421    try {
422      await $.store.set(mine.key, mine.log)
423      return
424    } catch {
425      const oldest = (await calibrationLogs($)).find(l => l.key !== mine.key)
426      if (!oldest) break
427      await $.store.delete(oldest.key)
428    }
429  }
430  $.ui.toast('quota-band: calibration log not saved')
431}
432
433async function refit($: any, c: Calibration) {
434  const all = (await calibrationLogs($)).map(l => (l.key === c.own?.key ? c.own.log : l.log))
435  if (c.own && !all.includes(c.own.log)) all.push(c.own.log)
436  c.fits = [fitWeights(all, 1, '5h'), fitWeights(all, 2, '7d')]
437  c.fittedAt = await $.clock.now()
438}
439
440export const register: Register = (on, options: Options = {}) => {
441  let pinged = { at: 0, step: -1 }
442  let title: string | null = null
443  let observedTtl: number | null = null
444  let baseline: { path: string; tokens: number } | null = null
445  const calibration: Calibration = { own: null, fits: [], fittedAt: 0 }
446
447  on('command.run', { command: 'quota-weights' }, async $ => {
448    await refit($, calibration)
449    const used = usable(calibration.fits[0]) ? 'measured 5h weights' : `API price ratios until 5h has ${MIN_INTERVALS} intervals and a cache read range within ±${MAX_SPREAD * 100}%`
450    return { text: [...calibration.fits.map(fitText), `output stays at ${OUTPUT_PER_WRITE} cache writes per token`, `cost × uses ${used}`].join('\n') }
451  })
452
453  on('session.start', async ($, e, next) => {
454    const result = await next(e)
455    await $.command.register({ name: 'quota-weights', description: 'How much of the 5h and 7d quota cache reads and output use, measured against cache writes' })
456    await ownLog($, calibration)
457    await refit($, calibration)
458    const usage = await $.session.usage()
459    if (usage.rateLimits.length > 0) await update($, limits, () => usage.rateLimits)
460    $.clock.every(60e3, () => { update($, tick, n => n + 1) })
461    $.clock.every(COLD_CHECK_MS, () => {
462      read($, turn).then(async last => {
463        if (!last) return
464        const left = last.at + last.ttlMs - Date.now()
465        const step = coldStep(left, last.ttlMs)
466        if (step === null || step < 0) return
467        if (pinged.at === last.at && pinged.step >= step) return
468        pinged = { at: last.at, step }
469        const minutes = left >= 60e3 ? `${Math.round(left / 60e3)} min` : `${Math.round(left / 1e3)} s`
470        const text = `Prompt cache ${COLD_STEPS[step] * 100}% warm · cold in ${minutes} · ${tokensText(last.tokens)} to rebuild`
471        $.ui.toast(text, { timeoutMs: 10e3 })
472        const folder = (await $.session.cwd().catch(() => '')).split('/').pop()
473        await notify($, title ?? folder ?? 'Claude', text)
474        if (options.iMessageTo) await sendIMessage($, String(options.iMessageTo), `${title ?? folder ?? 'Claude'}: ${text}`)
475      })
476    })
477    return result
478  })
479
480  on('session.measure', async ($, e, next) => {
481    if (e.rateLimits.length > 0) {
482      await update($, limits, () => e.rateLimits)
483      handOff($, e.rateLimits).catch(() => {})
484    }
485    const ttlMs = observedTtl ?? await cacheTtl($, e.rateLimits.length > 0)
486    const at = await $.clock.now()
487    await update($, turn, () => ({ at, tokens: e.context.tokens ?? null, window: e.context.window ?? null, ttlMs }))
488    const five = e.rateLimits.find(l => l.kind === 'five_hour')?.percentUsed ?? null
489    const seven = e.rateLimits.find(l => l.kind === 'seven_day')?.percentUsed ?? null
490    if (five !== null || seven !== null) {
491      const { log } = await ownLog($, calibration)
492      const last = log.readings[log.readings.length - 1]
493      const t = Math.round(at / 1e3)
494      if (!last || last[1] !== five || last[2] !== seven || t - (last[0] as number) >= READING_EVERY_S) log.readings.push([t, five, seven])
495    }
496    return next(e)
497  })
498
499  on('classic.Stop', async ($, e, next) => {
500    const result = await next(e)
501    const path: string | undefined = (e as any).transcript_path
502    if (!path) return result
503    const tail = await transcriptEnd($, 'tail', path)
504    if (!tail) return result
505    title = threadTitle(tail) ?? title
506    const ttlMs = writtenTtl(tail)
507    if (ttlMs !== null) {
508      observedTtl = ttlMs
509      await update($, turn, last => (last ? { ...last, ttlMs } : last))
510    }
511    if (baseline?.path !== path) {
512      const head = await transcriptEnd($, 'head', path)
513      const tokens = head ? baselineTokens(head) : null
514      baseline = tokens === null ? null : { path, tokens }
515    }
516    const requests = lastTurn(tail)
517    const mine = await ownLog($, calibration)
518    const loggedUntil = mine.log.requests[mine.log.requests.length - 1]?.[0] ?? 0
519    mine.log.requests.push(...requests.map(requestRow).filter(r => r[0] > loggedUntil).reverse())
520    mine.log.updated = await $.clock.now()
521    await saveLog($, mine)
522    if (mine.log.updated - calibration.fittedAt >= REFIT_MS) await refit($, calibration)
523    if (baseline) {
524      const fit = calibration.fits[0] ?? null
525      const weights = usable(fit) ? fit.weights : null
526      const value = turnMultiplier(requests, baseline.tokens, weights)
527      await update($, multiplier, () => (value === null ? null : { value, measured: weights !== null }))
528    }
529    return result
530  })
531
532  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
533    const shown = await read($, limits)
534    const last = await read($, turn)
535    const times = await read($, multiplier)
536    await read($, tick)
537    const below = await next(e)
538    if (e.props.hasSurvey || (shown.length === 0 && !last)) return below
539    const now = await $.clock.now()
540    const ui = $.ui.resolve(e)
541    const { Box, Text, Button } = ui
542    const left = last ? last.at + last.ttlMs - now : null
543    const cache = left === null || !last ? null
544      : left > 0 ? { text: `${Math.ceil(left / 60e3)}m`, share: left / last.ttlMs, color: left <= COLD_SOON ? '#ff8c00' : '#f5f5f7' }
545      : { text: 'cold', share: 0, color: '#ff453a' }
546    const context = last?.tokens && last.window ? { share: last.tokens / last.window, percent: Math.round((last.tokens / last.window) * 100) } : null
547    const buttons = [
548      options.closeCommand ? { key: 'close', label: 'close', text: String(options.closeCommand) } : null,
549      options.renamePrompt ? { key: 'rename', label: 'rename', text: String(options.renamePrompt) } : null,
550    ].filter(Boolean) as { key: string; label: string; text: string }[]
551    const line = shown.map(l => `${labels[l.kind] ?? l.kind} ${Math.round(l.percentUsed)}%` + (l.resetsAt ? ` · ${resetText(l, now)}` : '')).join('   ')
552      + (context ? `   ctx ${context.percent}% · ${tokensText(last!.tokens)}` : '')
553      + (cache ? `   cache ${cache.text}` : '')
554      + (times ? `   cost ×${times.value.toFixed(1)}${times.measured ? ' measured' : ''}` : '')
555    // Wider than any terminal; the divider row clips it to the band's width.
556    const divider = below ? <Box height={1} overflow="hidden"><Text dimColor>{'─'.repeat(400)}</Text></Box> : null
557    if (e.surface !== 'desktop' || !('Svg' in ui)) return <Box flexDirection="column"><Text dimColor>{line}</Text>{divider}{below}</Box>
558    const { Svg } = ui as any
559
560    return (
561      <Box flexDirection="column">
562      <Box flexDirection="row" alignItems="center" columnGap={2}>
563        {shown.map(l => (
564          <Box key={l.kind} flexDirection="row" alignItems="center" columnGap={1}>
565            <Text dimColor>{labels[l.kind] ?? l.kind}</Text>
566            <Svg source={bar(l, now, 120, 8)} alt={`${l.kind} ${l.percentUsed}%`} width={120} height={8} />
567            <Text>{Math.round(l.percentUsed)}%</Text>
568            {projected(l, now) !== null && projected(l, now)! > Math.round(l.percentUsed) ? <Text dimColor>→ {projected(l, now)}%</Text> : null}
569            {l.resetsAt ? <Text dimColor>{resetText(l, now)}</Text> : null}
570          </Box>
571        ))}
572        {context ? (
573          <Box flexDirection="row" alignItems="center" columnGap={1}>
574            <Text dimColor>ctx</Text>
575            <Svg source={meter(context.share, tint(context.percent, null), 60, 8)} alt={`context ${context.percent}%`} width={60} height={8} />
576            <Text>{context.percent}%</Text>
577            <Text dimColor>{tokensText(last!.tokens)}</Text>
578          </Box>
579        ) : null}
580        {cache ? (
581          <Box flexDirection="row" alignItems="center" columnGap={1}>
582            <Text dimColor>cache</Text>
583            <Svg source={meter(cache.share, cache.color, 60, 8)} alt={`cache ${cache.text}`} width={60} height={8} />
584            <Text color={cache.color === '#f5f5f7' ? undefined : cache.color}>{cache.text}</Text>
585          </Box>
586        ) : null}
587        {times ? (
588          <Box flexDirection="row" alignItems="center" columnGap={1}>
589            <Text dimColor>cost</Text>
590            <Text>×{times.value.toFixed(1)}</Text>
591            {times.measured ? <Text dimColor>measured</Text> : null}
592          </Box>
593        ) : null}
594        {buttons.map(b => (
595          <Button key={b.key} label={b.label} onPress={() => { $.prompt.submit({ text: b.text, asUser: true }) }} />
596        ))}
597      </Box>
598      {divider}
599      {below}
600      </Box>
601    )
602  })
603}
604
types/index.d.ts 12 lines
1export type Limit = { kind: string; percentUsed: number; resetsAt?: string }
2
3export type Multiplier = { value: number; measured: boolean }
4
5export type Turn = { at: number; tokens: number | null; window: number | null; ttlMs: number }
6
7declare module 'claude-code' {
8  interface PluginState {
9    'quota-band': { limits: Limit[]; turn: Turn | null; multiplier: Multiplier | null; tick: number }
10  }
11}
12