Your Claude plan limits, a forecast of when you'll run out, and your context window, live in a band above the prompt. Compacts in one click, automatically, or…

Know how much room you have left, and make more of it. A live band above the Claude Code prompt that shows your plan limits, forecasts whether you'll run out before they reset, and shows how full your context window is. Compact in one click, let it compact automatically, or let Claude pick the moment. Keep the prompt cache warm so the first message after a break stays cheap.
Formerly Claude Code Usage Quota Mod.
The 5 Hour and Weekly limits are your whole Claude account's, including what you use in Claude chat and Cowork. The band itself shows in Claude Code: the desktop app and the terminal.
Type /headroom to turn the band off and on.

One row that still shows the time left until each limit resets:

[!TIP] Never hit a full context again: Auto compact, on from the start. Auto compact is on in every new session and compacts it once the context reaches 80% (set 15 to 99). Set it lower to compact sooner: a long context costs more on every reply, so compacting early keeps replies cheaper and your limits lasting longer. Auto compact never interrupts a reply. Want to compact right now? The Compact now button does it in one click, and ends any hold Claude has.
Claude picks the moment: Agent-timed. On by default too, Agent-timed starts compaction sooner (from 30%) but at a good moment: Claude can hold it through a debugging chain or a refactor, release it at a safe point, and leave itself a note that survives. Every 5 minutes of a hold, Claude is asked to keep, release or update it. Your Auto compact % stays the limit no hold can pass.
Pay less for every message: Keep cache warm. Also on from the start. A small refresh keeps the conversation's prompt cache alive through a break, and while Claude waits on a subagent, so the next message reads the cache at a tenth of the price instead of writing it all again. The Cache dropdown in the top row picks the lifetime, auto, 5m or 1h, and the band shows what the warmer cost and saved, this session and all time.
Each session keeps its own settings, through a restart and a
/clear, so you can run them differently side by side:
- Building something big? Leave Auto compact at 80% or turn it off, so Claude keeps the whole picture. Press Compact now yourself at a good stopping point. (With it off, Claude Code's built-in compaction still steps in when the context is nearly full.)
- Everyday sessions? Set Auto compact lower (say 30%), or leave Agent-timed to it. They stay lean, and your 5 Hour and Weekly limits last longer.
- A session that should send nothing extra? Switch Keep cache warm off there; every refresh counts against your plan like any request.
5 Hour and Weekly limits
On course to run out, it warns you and says when you'll hit the limit:


Context window
/context counts it, split into Messages, Tools and Other.Compacting
Prompt cache
If the session is already past your % when you turn it on, it asks first:


/compact, or Claude Code's own), then takes over.Works in light and dark themes and at every width down to the narrowest. Collapsed or expanded is shared across all sessions. Each session keeps its own Auto compact, Agent-timed and Keep cache warm settings, and a /clear keeps them: the cleared session goes on with the same ones, while Claude's hold and note start over with the context.
Early in a 5 hour window:


Running out on the Weekly limit:


| Light theme | Narrowest window |
|---|---|
![]() | ![]() |
![]() | ![]() |
Auto compact fires at a fixed %, whatever Claude is doing. A compaction that lands mid-debugging throws away the context that mattered. With Agent-timed on, Claude chooses the moment, between two numbers you set:
What Claude can do, through a small compaction tool the mod gives it:
While Claude holds, the band says so, with the reason and for how long, and a Release button ends the hold from your side:

What you see and keep:
Agent-timed is on in a new session, from 30%, and each session keeps its own setting. Switching it on adds the tool to that session; switched off again, the tool stays listed until the session is reopened, and answers that the mode is off.
Every message sends the whole conversation. The API keeps it in a prompt cache for 5 minutes or an hour; a message that reads the cache pays about a tenth of the input price, and one after the cache has expired writes it all again at 1.25× (5 minutes) or 2× (1 hour). After a break, the first message pays.
With Keep cache warm on, the mod sends one small request shortly before the cache would expire: a copy of the conversation's last request with one line asking for the word ok. It re-reads the cache, which keeps it alive, and never enters the conversation; a ☕ row in the transcript records each one with what it cost and saves.
/config rows (headroom.idle5m, headroom.idle1h, headroom.warmUntil).FORCE_PROMPT_CACHING_5M, CLAUDE_CODE_PROMPT_CACHE_TTL, ENABLE_PROMPT_CACHING_1H or the promptCacheTtl setting say). Pick 5m or 1h to choose for this session, warming on or off. The first reply of a session writes the cache, so a choice made after it applies to new sessions. headroom.cacheTtl in /config sets what new chats start with./clear, a model switch or a failed refresh start it afresh with your next message.[!IMPORTANT] Each refresh counts against your plan like any other request: a cache read of the conversation and a few output tokens. It is set per session: switch it off in the band for a chat that should send none. The expired-cache warning shows what the refreshes save, in the same 5 Hour %, warming on or off.
With this on, disable cache-warmer if you have it installed: both would warm the same cache.
[!NOTE] In the desktop app, the band appears once a session has started. The brand-new Welcome back screen has no session yet, so no mod can draw there. Send your first message and it appears.
In the terminal, ▲/▼ expands and collapses it. The [-] next to it is Claude Code's own control and hides the band; ctrl+x ctrl+a brings it back.


Needs a recent Claude Code, signed in with a Pro or Max plan.
Desktop app: in the Code tab, send this as a message, allow the claude plugin command if asked, then quit and reopen the app:
Install the headroom plugin from the GitHub marketplace MWEJ/claude-code-headroom
Terminal: inside claude, run:
/plugin install headroom --marketplace MWEJ/claude-code-headroom
Either way installs it for both the desktop app and the terminal.
Update: ask Claude, then restart:
Update the headroom plugin from its marketplace
Uninstall: in the desktop app, ask Claude:
Remove the claude-code-headroom plugin marketplace
or in the terminal:
/plugin marketplace remove claude-code-headroom
then restart. This removes it from both. Just want it out of sight? /headroom hides it without uninstalling.
claude plugin marketplace add MWEJ/claude-code-headroom; claude plugin install headroom@claude-code-headroom
To update:
claude plugin marketplace update claude-code-headroom; claude plugin update headroom@claude-code-headroom
To uninstall:
claude plugin marketplace remove claude-code-headroom
Everything runs on your machine. It reads what Claude Code already has (your context and the limits each reply carries), and about every 2 minutes asks Anthropic's usage service for your limits through your existing Claude login. It never sees your credentials and sends nothing anywhere else. In the desktop app it reads the app's theme setting to match light or dark.
First release: I'd love to hear how it works for you. Open an issue for bugs or ideas. Pull requests welcome.
To work on it, clone the repo, then load it from the folder, or run its tests:
claude --plugin-dir ./claude-code-headroom
claude plugin test ./claude-code-headroom
If you find it useful, a ⭐ helps others find it.
Headroom started as Claude Code Usage Quota Mod by Anant Raghunath (MIT): the limits, the forecast, the context window and Auto compact are his work.
Inspired by I'm liking the new mods feature on r/ClaudeCode. Thanks to u/itsxzy for sharing the original prompt that started this project.
Agent-timed is inspired by compactor by rhwendt (MIT), which lets the agent hold and release Claude Code's own auto-compaction.
Keep cache warm is ported from cache-warmer by Paul B. Kim (MIT), itself a port of the cache warmer in Pi by Mario Zechner, whose rule decides when a refresh pays.
MIT © 2026 Martin Hygge. Started from Claude Code Usage Quota Mod © 2026 Anant Raghunath, MIT.
hooks/register.tsx 2329 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, ModelForkResult, ModelUsage, Register, SessionRateLimit, TurnUsage } from 'claude-code'
3
4import type { AgentTimed, AutoCompact, Category, Limit, PaceOf, Snapshot, Ttl, TtlChoice, Warm, WarmAnchor, WarmRate, WarmSetting, WarmTotals } from '../types'
5import {
6 AT_DEFAULT, DEFAULT_AUTO, EMPTY, REASON_SHOWN, START_DEFAULT, START_MIN, TOOL, TOOL_DESCRIPTION, TOOL_NAME, TOOL_SCHEMA,
7 afterText, answerTool, breakpointOf, breakpointText, capOf, decide, holdReminder, isTimed, nudgeLevel, nudgeText, startOf, stepNudge, toldText, withNote,
8} from './agent-policy'
9import type { Stuck, ToolInput } from './agent-policy'
10import {
11 DEFAULT_OUTPUT_TOKENS, FORK_PROMPT, IDLE_LIMIT_DEFAULT, TTL_MS, WARM_UNTIL_DEFAULT, ZERO_TOTALS,
12 addRate, addTotals, cacheLineOf, costOf, deadlineOf, decide as decideWarm, delayOf, formatDuration, formatTokens, formatUsd, horizonOf, idleStopNotice,
13 idleStopReason, isEnvOn, isTtlChoice, jumpsOf, limitOf, missCostOf, noticeText, outcomeOf, pastLimitOf, planOf, ttlOf,
14 usageOf, warmUntilOf,
15} from './cache-policy'
16import type { ForkReply, Refresh } from './cache-policy'
17
18const snapshot = atom({ plugin: 'headroom', key: 'snapshot' } as const, null)
19const isOn = atom({ plugin: 'headroom', key: 'isOn' } as const, true)
20const isCollapsed = atom({ plugin: 'headroom', key: 'isCollapsed' } as const, false)
21const autoCompact = atom({ plugin: 'headroom', key: 'autoCompact' } as const, DEFAULT_AUTO)
22// what Agent-timed holds for the session: the agent's hold, note and request, and what it has been told
23const agentTimed = atom({ plugin: 'headroom', key: 'agentTimed' } as const, EMPTY as AgentTimed)
24const fieldTick = atom({ plugin: 'headroom', key: 'fieldTick' } as const, 0)
25// the % field's text while typing is cleaned (digits only, 3 at most); null: the set %
26const fieldText = atom({ plugin: 'headroom', key: 'fieldText' } as const, null as string | null)
27// the start % field's own tick and typed text, as fieldTick and fieldText are the cap's
28const startTick = atom({ plugin: 'headroom', key: 'startTick' } as const, 0)
29const startText = atom({ plugin: 'headroom', key: 'startText' } as const, null as string | null)
30const theme = atom({ plugin: 'headroom', key: 'theme' } as const, 'dark' as 'dark' | 'light')
31const autoAsk = atom({ plugin: 'headroom', key: 'autoAsk' } as const, null as { at: number; percent: number } | null)
32// Keep cache warm: the session's chain, lifetime, totals and rate; and this chat's switch
33const WARM_EMPTY: Warm = {
34 chat: null,
35 anchor: null,
36 status: { state: 'waiting' },
37 isRunning: false,
38 outputTokens: DEFAULT_OUTPUT_TOKENS,
39 isLocked: false,
40 ttl: '5m',
41 assumed: null,
42 totals: ZERO_TOTALS,
43 allTime: { ...ZERO_TOTALS, since: 0 },
44 rate: {},
45 lastLimits: null,
46}
47const warm = atom({ plugin: 'headroom', key: 'warm' } as const, WARM_EMPTY)
48const warmSetting = atom({ plugin: 'headroom', key: 'warmSetting' } as const, { isOn: true, ttl: 'auto' } as WarmSetting)
49
50// two palettes, the desktop's dark and light themes; the band draws in the one the
51// app shows (Theme), set at the start of every draw so all colours below follow it
52type Palette = {
53 muted: string; green: string; greenText: string; amber: string; red: string; free: string; buffer: string; track: string
54 palette: string[]; named: Record<string, string>; switchOff: string; switchOn: string; switchLit: string; compactLit: string
55}
56const DARK: Palette = {
57 muted: '#8b90a0', green: '#4ade80', greenText: '#4ade80', amber: '#fbbf24', red: '#f87171',
58 free: '#2d3140', buffer: '#4b5163', track: '#2d3140',
59 palette: ['#a78bfa', '#60a5fa', '#f472b6', '#34d399', '#fbbf24', '#fb923c', '#22d3ee', '#c084fc'],
60 named: {
61 Tools: '#a78bfa', Other: '#8b90a0', 'System prompt': '#a78bfa', 'System tools': '#60a5fa', 'MCP tools': '#f472b6',
62 'Custom agents': '#34d399', 'Memory files': '#fbbf24', Skills: '#fb923c', Messages: '#22d3ee',
63 },
64 switchOff: '#4b5163', switchOn: '#4ade80', switchLit: '#8b919c', compactLit: '#5a606e',
65}
66// light: the same hues, deep enough to read on the pale band; the tracks pale
67const LIGHT: Palette = {
68 muted: '#5b6070', green: '#22b856', greenText: '#16a34a', amber: '#ea8a0c', red: '#ef4444',
69 free: '#ffffff', buffer: '#b9bec9', track: '#ffffff',
70 palette: ['#8b5cf6', '#3b82f6', '#ec4899', '#10b981', '#f59e0b', '#f97316', '#0ea5e9', '#a855f7'],
71 named: {
72 Tools: '#8b5cf6', Other: '#5b6070', 'System prompt': '#8b5cf6', 'System tools': '#3b82f6', 'MCP tools': '#ec4899',
73 'Custom agents': '#10b981', 'Memory files': '#f59e0b', Skills: '#f97316', Messages: '#0ea5e9',
74 },
75 switchOff: '#b4b9c4', switchOn: '#22c55e', switchLit: '#c3c7cf', compactLit: '#d6d9df',
76}
77let MUTED = DARK.muted
78let GREEN = DARK.green
79let GREEN_TEXT = DARK.greenText
80let AMBER = DARK.amber
81let RED = DARK.red
82let FREE = DARK.free
83let BUFFER = DARK.buffer
84let TRACK = DARK.track
85let PALETTE = DARK.palette
86let NAMED = DARK.named
87let SWITCH_OFF = DARK.switchOff
88let SWITCH_ON = DARK.switchOn
89// the light behind the switch under the pointer: drawn beneath the pill, never over it
90let SWITCH_LIT = DARK.switchLit
91// Compact's hover: clearly lighter than its resting chrome, its label still reads
92let COMPACT_LIT = DARK.compactLit
93function usePalette(p: Palette): void {
94 ;({ muted: MUTED, green: GREEN, greenText: GREEN_TEXT, amber: AMBER, red: RED, free: FREE, buffer: BUFFER, track: TRACK } = p)
95 ;({ palette: PALETTE, named: NAMED, switchOff: SWITCH_OFF, switchOn: SWITCH_ON, switchLit: SWITCH_LIT, compactLit: COMPACT_LIT } = p)
96}
97const CELLS = 16
98const HOUR = 3_600_000
99const WINDOWS: Record<string, { label: string; ms: number }> = {
100 five_hour: { label: '5 Hour', ms: 5 * HOUR },
101 seven_day: { label: 'Weekly', ms: 7 * 24 * HOUR },
102}
103
104// 5 hour on the left, Weekly on the right
105const order = (kind: string) => (kind === 'five_hour' ? 0 : 1)
106
107type Forecast = {
108 kind: string
109 label: string
110 percent: number
111 resetsAt: number
112 /** 'idle': no window running (unused, or reset): 0% until the next message starts one */
113 status: 'hit' | 'ok' | 'out' | 'idle'
114 projected: number
115 runOutAt: number
116 color: string
117}
118
119// The pace a window is forecast at: your usual one (the median of how far your past
120// windows got) blended with this window's own, which takes over as the window runs: it
121// counts for half a quarter of the way in. Early jumps barely move it. `skip`: the %
122// the first reply of a reopened chat took re-reading its old context, a one-off left
123// out of the pace (still in the % used).
124const TYPICAL_DEFAULT = 50
125const TRUST = 0.25
126const FINALS_KEPT = 8
127// one window's reset time as two readings give it: a few minutes apart at most
128const SAME_WINDOW_MS = 10 * 60_000
129
130function forecast(limit: Limit, now: number, pace?: PaceOf): Forecast | null {
131 const win = WINDOWS[limit.kind]
132 const resetsAt = limit.resetsAt ? Date.parse(limit.resetsAt) : NaN
133 if (!win) return null
134 const p = limit.percentUsed
135 // no reset time, or one already passed: the window is not running, so it is at 0
136 if (!Number.isFinite(resetsAt) || resetsAt <= now) {
137 return { kind: limit.kind, label: win.label, percent: 0, resetsAt: NaN, status: 'idle', projected: 0, runOutAt: Infinity, color: GREEN }
138 }
139 const remaining = Math.max(0, resetsAt - now)
140 const elapsed = Math.max(1, win.ms - remaining)
141 const base = { kind: limit.kind, label: win.label, percent: p, resetsAt }
142 if (p >= 100) return { ...base, status: 'hit', projected: p, runOutAt: now, color: RED }
143 const typical = pace?.typical ?? TYPICAL_DEFAULT
144 const own = Math.max(0, p - (pace?.skip ?? 0)) / elapsed
145 const weight = elapsed / (elapsed + win.ms * TRUST)
146 const rate = weight * own + (1 - weight) * (typical / win.ms)
147 const projected = p + rate * remaining
148 if (projected < 100 || rate <= 0) {
149 return { ...base, status: 'ok', projected, runOutAt: Infinity, color: worse(GREEN, usedColor(p)) }
150 }
151 return {
152 ...base,
153 status: 'out',
154 projected,
155 runOutAt: now + (100 - p) / rate,
156 color: worse(projected > 130 ? RED : AMBER, usedColor(p)),
157 }
158}
159
160// A limit's colour by what is used, as the band shows it (rounded): amber from 75%,
161// red from 90%. The bar and its % take the worse of this and the forecast's.
162const LIMIT_AMBER_AT = 75
163const LIMIT_RED_AT = 90
164function usedColor(p: number): string {
165 const shown = Math.round(p)
166 return shown >= LIMIT_RED_AT ? RED : shown >= LIMIT_AMBER_AT ? AMBER : GREEN
167}
168function worse(a: string, b: string): string {
169 const rank = (c: string) => (c === RED ? 2 : c === AMBER ? 1 : 0)
170 return rank(b) > rank(a) ? b : a
171}
172
173function duration(ms: number): string {
174 const m = Math.max(0, Math.round(ms / 60_000))
175 const d = Math.floor(m / 1440)
176 const h = Math.floor((m % 1440) / 60)
177 const min = m % 60
178 if (d > 0) return h > 0 ? `${d}d ${h}h` : `${d}d`
179 if (h > 0) return min > 0 ? `${h}h ${min}m` : `${h}h`
180 return `${min}m`
181}
182
183const time = (t: number) =>
184 new Date(t).toLocaleTimeString('en-US', { hour: 'numeric', minute: '2-digit' })
185const weekday = (t: number) => new Date(t).toLocaleDateString('en-US', { weekday: 'long' })
186const isSameDay = (a: number, b: number) =>
187 new Date(a).toLocaleDateString('en-US') === new Date(b).toLocaleDateString('en-US')
188
189// "the 8:40 PM reset" today, "Wednesday's reset" otherwise
190// the reset by its time when it is under a day away (an after-midnight one included),
191// else by its weekday: a time that far off would only lengthen the headline
192function resetName(t: number, now: number): string {
193 return t - now < 24 * 3_600_000 ? `the ${time(t)} reset` : `${weekday(t)}'s reset`
194}
195
196function headline(list: Forecast[], now: number): { text: string; color: string } | null {
197 const hit = list.find(f => f.status === 'hit')
198 if (hit) {
199 const at = isSameDay(hit.resetsAt, now) ? time(hit.resetsAt) : `${weekday(hit.resetsAt)} ${time(hit.resetsAt)}`
200 return { text: `Limit reached. Usage resumes at ${at}.`, color: RED }
201 }
202 const out = list.filter(f => f.status === 'out').sort((a, b) => a.runOutAt - b.runOutAt)[0]
203 if (out) {
204 return { text: `At this pace you'll run out before ${resetName(out.resetsAt, now)}.`, color: out.color }
205 }
206 const lead = list.find(f => f.kind === 'seven_day' && f.status !== 'idle') ?? list.find(f => f.status !== 'idle')
207 if (!lead) return null
208 return {
209 text: `On track. You should reach ${resetName(lead.resetsAt, now)} with room to spare.`,
210 color: GREEN,
211 }
212}
213
214function tokens(n: number): string {
215 // 1M, not 1.0M; 1.5M keeps its decimal
216 if (n >= 1_000_000) return `${(n / 1_000_000).toFixed(1).replace(/\.0$/, '')}M`
217 if (n >= 1000) return `${Math.round(n / 1000)}k`
218 return `${Math.round(n)}`
219}
220
221// Context by what it holds, as shown (rounded): amber from 50% (a long context costs
222// more on every reply: compacting pays), red from 80%
223const CONTEXT_AMBER_AT = 50
224const CONTEXT_RED_AT = 80
225function fillColor(percent: number): string {
226 const shown = Math.round(percent)
227 if (shown >= CONTEXT_RED_AT) return RED
228 if (shown >= CONTEXT_AMBER_AT) return AMBER
229 return GREEN
230}
231
232// text in green, which on the light band needs a deeper green than the bars to read
233function ink(color: string): string {
234 return color === GREEN ? GREEN_TEXT : color
235}
236
237function colorOf(category: Category, index: number): string {
238 if (category.kind === 'free') return FREE
239 if (category.kind === 'buffer') return BUFFER
240 return NAMED[category.name] ?? PALETTE[index % PALETTE.length]!
241}
242
243type Segment = { color: string; share: number }
244
245const BAR_PX = 7.5
246
247// an iOS-style switch: a pill, the knob right and green when on, left and grey when off
248// three cells wide, so the click area laid over it covers all of it
249const SWITCH_W = 30
250const SWITCH_H = 20
251const SWITCH_PAD = 3
252// the cells the switch's box takes: the press under it is clipped to this, so the
253// app's own button tint can't spill onto the text beside it
254const SWITCH_CELLS = 4
255// the box behind the switch measures a pixel off the drawing (19px tall, 31 wide):
256// the pill is drawn half a pixel down and right so it sits in the light's centre
257const SWITCH_NUDGE = 0.5
258// fully see-through: the unlit border, and the press's own hover tint, which drew a
259// square frame round the light (the word "transparent" draws white or nothing)
260const CLEAR = '#00000000'
261// a space a wrap never breaks at
262const NB = ' '
263
264function switchSvg(isOn: boolean): string {
265 // the pill sits inset by PAD: the light under the pointer shows around it
266 const w = SWITCH_W - 2 * SWITCH_PAD
267 const h = SWITCH_H - 2 * SWITCH_PAD
268 const p = SWITCH_PAD
269 const knob = p + (isOn ? w - h / 2 : h / 2)
270 return (
271 `<svg xmlns="http://www.w3.org/2000/svg" width="${SWITCH_W}" height="${SWITCH_H}" viewBox="${-SWITCH_NUDGE} ${-SWITCH_NUDGE} ${SWITCH_W} ${SWITCH_H}">` +
272 `<rect x="${p}" y="${p}" width="${w}" height="${h}" rx="${h / 2}" fill="${isOn ? SWITCH_ON : SWITCH_OFF}"/>` +
273 `<circle cx="${knob}" cy="${p + h / 2}" r="${h / 2 - 2}" fill="#ffffff"/></svg>`
274 )
275}
276
277// a thin rounded bar for surfaces that draw Svg: segments laid out over 0..1000,
278// with an optional grey tick, the bar's own height, at `marker` percent (the forecast at reset)
279function barSvg(segments: Segment[], width: number, marker?: number): string {
280 const height = BAR_PX
281 const y = 0
282 const total = segments.reduce((sum, s) => sum + s.share, 0) || 1
283 let x = 0
284 const rects = segments
285 .map(s => {
286 const w = (s.share / total) * 1000
287 const rect = `<rect x="${x.toFixed(2)}" y="${y}" width="${w.toFixed(2)}" height="${BAR_PX}" fill="${s.color}"/>`
288 x += w
289 return rect
290 })
291 .join('')
292 // a plain upright block 3px wide whatever the drawn width: no rounded corners
293 // (stretched sideways, those drew it as a "D"), inside the bar's rounded ends so
294 // nothing pokes out past them. Placed on a whole pixel and left smooth-edged: snapped
295 // edges on a scaled screen drew it 3 device pixels wide on one bar and 4 on another
296 const TICK_PX = 3
297 const tickW = (TICK_PX / width) * 1000
298 const tickAt = marker === undefined ? 0 : Math.min(width - TICK_PX, Math.max(0, Math.round((marker / 100) * width - TICK_PX / 2)))
299 const tick =
300 marker === undefined
301 ? ''
302 : `<rect x="${((tickAt / width) * 1000).toFixed(3)}" y="${y}" width="${tickW.toFixed(3)}" height="${BAR_PX}" fill="${MUTED}"/>`
303 return (
304 `<svg xmlns="http://www.w3.org/2000/svg" width="${width}" height="${height}" viewBox="0 0 1000 ${height}" preserveAspectRatio="none">` +
305 `<clipPath id="r"><rect y="${y}" width="1000" height="${BAR_PX}" rx="${BAR_PX / 2}" ry="${BAR_PX / 2}"/></clipPath>` +
306 `<g clip-path="url(#r)">${rects}${tick}</g></svg>`
307 )
308}
309
310// the forecast tick's percent, or none when there is nothing ahead to mark
311function markerOf(f: Forecast): number | undefined {
312 if (f.status === 'hit' || f.status === 'idle') return undefined
313 const at = Math.min(100, f.projected)
314 return at - f.percent >= 1 ? at : undefined
315}
316
317// the terminal's bar: one colour per cell, runs of a colour merged into one Box
318function cellRuns(f: Forecast, cells: number): { color: string; width: number }[] {
319 const filled = Math.max(0, Math.min(cells, Math.round((f.percent / 100) * cells)))
320 const marker = markerOf(f)
321 const tick = marker === undefined ? -1 : Math.max(filled, Math.min(cells - 1, Math.round((marker / 100) * cells) - 1))
322 const runs: { color: string; width: number }[] = []
323 for (let i = 0; i < cells; i++) {
324 const color = i < filled ? f.color : i === tick ? MUTED : TRACK
325 const last = runs[runs.length - 1]
326 if (last && last.color === color) last.width += 1
327 else runs.push({ color, width: 1 })
328 }
329 return runs
330}
331
332const USAGE_URL = 'https://api.anthropic.com/api/oauth/usage'
333// The usage service rate-limits per login, and the app's own Usage panel asks it on
334// the same allowance, so ask it seldom: when a chat opens, then every 2 minutes with the
335// exact context count, one ask for the whole app (when it was last asked, the answer and
336// any back-off sit in the shared store). Between asks the figures every reply from
337// Claude carries keep the band current.
338const PLAN_EVERY_MS = 2 * 60_000
339// timers drift: a 2-minute tick a moment early still counts as due
340const PLAN_SLACK_MS = 10_000
341// after a failed ask wait 2, then 5, then 15 minutes, or what the service says
342const PLAN_BACKOFF_MS = [120_000, 300_000, 900_000]
343const LOCAL_EVERY_MS = 3_000
344
345type Backoff = { until: number; failures: number; error: string }
346/** 'no': the 3s tick, never asks; 'due': asks if 2 minutes have passed app-wide; 'now': asks unless backed off */
347type Ask = 'no' | 'due' | 'now'
348
349class AskError extends Error {
350 constructor(
351 message: string,
352 readonly retryAfterMs?: number,
353 ) {
354 super(message)
355 }
356}
357
358// The plan's limits, as the app's "Plan usage limits" panel reads them: the account's
359// usage endpoint, through the session's own credential (the plugin never sees it).
360async function fetchPlan($: EngineInterface): Promise<Limit[] | null> {
361 const auth = await $.session.authorize()
362 // no Claude login (an API key, a gateway): there is no plan to ask about
363 if (!auth || auth.kind !== 'bearer') return null
364 const res = await $.http.fetch(USAGE_URL, {
365 auth: auth.handle,
366 headers: { 'anthropic-beta': 'oauth-2025-04-20', 'content-type': 'application/json' },
367 })
368 if (!res.ok) {
369 const header = Object.entries(res.headers ?? {}).find(([k]) => k.toLowerCase() === 'retry-after')?.[1]
370 const seconds = Number(Array.isArray(header) ? header[0] : header)
371 const retry = Number.isFinite(seconds) && seconds > 0 ? seconds * 1000 : undefined
372 throw new AskError(res.status === 429 ? 'refused (429)' : `failed (${res.status})`, retry)
373 }
374 const body = JSON.parse(res.text) as Record<string, { utilization?: number | null; resets_at?: string | null } | null>
375 const limits: Limit[] = []
376 for (const kind of ['five_hour', 'seven_day']) {
377 const w = body[kind]
378 if (!w || typeof w.utilization !== 'number') continue
379 limits.push({ kind, percentUsed: Math.round(w.utilization * 10) / 10, resetsAt: w.resets_at ?? undefined })
380 }
381 if (limits.length === 0) throw new AskError('answer had no limits')
382 return limits
383}
384
385type Saved = { at: number; limits: Limit[] }
386
387// this chat's own reading from its replies, and when it last changed
388let live: Saved | undefined
389
390async function storeGet<T>($: EngineInterface, key: string): Promise<T | undefined> {
391 try {
392 return (await $.store.get(key)) as T | undefined
393 } catch {
394 return undefined
395 }
396}
397
398async function storeSet($: EngineInterface, key: string, value: unknown): Promise<void> {
399 try {
400 await $.store.set(key, value)
401 } catch {
402 // this chat still has it
403 }
404}
405
406// Two sources, and the newest wins: the usage service (asked seldom, shared by every
407// chat through the store) and the figures each reply from Claude carries (this chat's,
408// stamped when they last changed). A window that has reset since is dropped.
409async function limitsOf(
410 $: EngineInterface,
411 rateLimits: SessionRateLimit[],
412 now: number,
413 ask: Ask,
414): Promise<{ limits: Limit[]; at?: number; error?: string }> {
415 if (rateLimits.length > 0) {
416 const limits = rateLimits.map(r => ({ kind: r.kind, percentUsed: r.percentUsed, resetsAt: r.resetsAt }))
417 if (!live || JSON.stringify(live.limits) !== JSON.stringify(limits)) live = { at: now, limits }
418 }
419 let saved = await storeGet<Saved>($, 'limits')
420 let backoff = await storeGet<Backoff>($, 'planBackoff')
421 const askedAt = (await storeGet<number>($, 'planAskedAt')) ?? 0
422 const isFree = now >= (backoff?.until ?? 0)
423 const isDue = ask === 'now' || (ask === 'due' && now - askedAt >= PLAN_EVERY_MS - PLAN_SLACK_MS)
424 if (isDue && isFree) {
425 // claim the ask first, so the other chats see it taken and skip theirs
426 await storeSet($, 'planAskedAt', now)
427 try {
428 const plan = await fetchPlan($)
429 if (plan) {
430 saved = { at: now, limits: plan }
431 await storeSet($, 'limits', saved)
432 await storeSet($, 'planLimits', saved)
433 }
434 if (backoff) await storeSet($, 'planBackoff', { until: 0, failures: 0, error: '' })
435 backoff = undefined
436 } catch (error) {
437 const failures = (backoff?.failures ?? 0) + 1
438 const wait =
439 error instanceof AskError && error.retryAfterMs !== undefined
440 ? error.retryAfterMs
441 : PLAN_BACKOFF_MS[Math.min(failures - 1, PLAN_BACKOFF_MS.length - 1)]!
442 backoff = { until: now + wait, failures, error: error instanceof Error ? error.message : String(error) }
443 await storeSet($, 'planBackoff', backoff)
444 }
445 }
446 const pick = [saved, live]
447 .filter((r): r is Saved => r !== undefined)
448 .sort((x, y) => y.at - x.at)[0]
449 // newer reply figures go to the store too, so other chats get them
450 if (pick && pick === live && (!saved || live.at > saved.at)) await storeSet($, 'limits', live)
451 // The reply figures run behind the usage service, which the app's panel reads: whole
452 // percents, and seen at 18 for an hour while the service said 19.0. Use only climbs
453 // within a window, so a newer reply figure below the service's last answer is stale:
454 // the higher of the two is kept. Above it, the reply's is news.
455 const plan = await storeGet<Saved>($, 'planLimits')
456 const finer = (l: Limit): Limit => {
457 const p = plan?.limits.find(x => x.kind === l.kind)
458 return p && sameWindow(p.resetsAt, l.resetsAt) && p.percentUsed > l.percentUsed ? { ...l, percentUsed: p.percentUsed } : l
459 }
460 // a window that has reset since stays, at 0: the forecast reads it as not running
461 const limits = (pick?.limits ?? []).map(l =>
462 l.resetsAt !== undefined && Date.parse(l.resetsAt) <= now ? { kind: l.kind, percentUsed: 0 } : finer(l),
463 )
464 return { limits, at: pick?.at, error: backoff?.error || undefined }
465}
466
467// What the forecast learns, shared by every chat through the store: how far each past
468// window got (its peak, the last few kept), and the running one's peak and skipped %.
469type PaceWindow = { resetsAt: number; peak: number; skip: number }
470type PaceStore = { finals: Record<string, number[]>; windows: Record<string, PaceWindow> }
471
472// A chat reopened with a long history re-reads it all on its first reply: one big jump
473// in usage. Its first turn notes the limits it started from; the first new reading
474// that moved is that jump, and is left out of the pace. Only for a chat that opened
475// with this much already in it (a new chat's first reply is real work).
476const RESUME_MIN_TOKENS = 20_000
477let isOpening = true
478let openMessages = 0
479let spikeFrom: { at: number; limits: Limit[] } | undefined
480
481function median(list: number[] | undefined): number | undefined {
482 if (!list || list.length === 0) return undefined
483 const s = [...list].sort((a, b) => a - b)
484 const mid = Math.floor(s.length / 2)
485 return s.length % 2 ? s[mid]! : (s[mid - 1]! + s[mid]!) / 2
486}
487
488const sameWindow = (a: string | undefined, b: string | undefined) =>
489 a !== undefined && b !== undefined && Math.abs(Date.parse(a) - Date.parse(b)) <= SAME_WINDOW_MS
490
491// the jump the reopened chat's first reply made, once a reading after it has moved
492function takeSpike(): Record<string, number> | undefined {
493 if (!spikeFrom || !live || live.at <= spikeFrom.at) return undefined
494 const jump: Record<string, number> = {}
495 for (const l of live.limits) {
496 const before = spikeFrom.limits.find(b => b.kind === l.kind)
497 if (before && sameWindow(before.resetsAt, l.resetsAt) && l.percentUsed > before.percentUsed) {
498 jump[l.kind] = l.percentUsed - before.percentUsed
499 }
500 }
501 if (Object.keys(jump).length === 0) return undefined
502 spikeFrom = undefined
503 return jump
504}
505
506async function trackPace($: EngineInterface, limits: Limit[], now: number): Promise<Record<string, PaceOf>> {
507 const spike = takeSpike()
508 const saved = (await storeGet<PaceStore>($, 'pace')) ?? { finals: {}, windows: {} }
509 let isChanged = false
510 const pace: Record<string, PaceOf> = {}
511 for (const l of limits) {
512 const resetsAt = l.resetsAt ? Date.parse(l.resetsAt) : NaN
513 if (!WINDOWS[l.kind] || !Number.isFinite(resetsAt) || resetsAt <= now) continue
514 let w = saved.windows[l.kind]
515 if (!w || Math.abs(w.resetsAt - resetsAt) > SAME_WINDOW_MS) {
516 // a new window: the last one's peak is how far it got
517 if (w && w.peak > 0) saved.finals[l.kind] = [...(saved.finals[l.kind] ?? []), w.peak].slice(-FINALS_KEPT)
518 w = { resetsAt, peak: l.percentUsed, skip: 0 }
519 saved.windows[l.kind] = w
520 isChanged = true
521 } else if (l.percentUsed > w.peak) {
522 w.peak = l.percentUsed
523 isChanged = true
524 }
525 const jump = spike?.[l.kind]
526 if (jump) {
527 w.skip = Math.min(w.peak, w.skip + jump)
528 isChanged = true
529 }
530 pace[l.kind] = { typical: median(saved.finals[l.kind]) ?? TYPICAL_DEFAULT, skip: w.skip }
531 }
532 if (isChanged) await storeSet($, 'pace', saved)
533 return pace
534}
535
536// The window is shown as three groups: Messages, Tools (System + MCP tools) and Other
537// (the system prompt, skills, memory files, agents, MCP server instructions).
538const TOOL_NAMES = new Set(['System tools', 'MCP tools'])
539type Group = 'Messages' | 'Tools' | 'Other'
540const groupOf = (name: string): Group => (name === 'Messages' ? 'Messages' : TOOL_NAMES.has(name) ? 'Tools' : 'Other')
541
542// The free local breakdown ('summary') estimates each category from its text, which
543// reads tools ~1.3-1.6x heavy. The exact count ('full', what /context and the app's
544// Context window panel show) costs no tokens but sends one token-count request per tool
545// and memory file (~320 here). So run it every 2 minutes, keep the exact/estimate ratio
546// of Tools and Other (in the store, for every session), and scale the live 3s estimate
547// by it; Messages is then the API's real input total less those two.
548const CALIBRATE_EVERY_MS = 2 * 60_000
549let calib: Partial<Record<Group, number>> = {}
550let isCalibrating = false
551// Right after /compact there is no API reply yet, so the window's real total is unknown
552// and the local estimate is all there is. The exact count fills that gap until the next reply.
553let exactCount: { at: number; sums: Record<Group, number> } | undefined
554let noReplySince: number | undefined
555
556function sumGroups(categories: readonly { name: string; tokens: number; kind: string }[]): Record<Group, number> {
557 const sums: Record<Group, number> = { Messages: 0, Tools: 0, Other: 0 }
558 for (const c of categories) if (c.kind === 'used') sums[groupOf(c.name)] += c.tokens
559 return sums
560}
561
562async function calibrate($: EngineInterface): Promise<void> {
563 if (isCalibrating) return
564 isCalibrating = true
565 try {
566 const exact = await $.session.usage({ breakdown: 'full' })
567 const rough = await $.session.usage({ breakdown: 'summary' })
568 const real = sumGroups(exact.context.breakdown?.categories ?? [])
569 const guess = sumGroups(rough.context.breakdown?.categories ?? [])
570 const next = { ...calib }
571 for (const g of ['Tools', 'Other'] as const) if (guess[g] > 0 && real[g] > 0) next[g] = real[g] / guess[g]
572 calib = next
573 exactCount = { at: await $.clock.now(), sums: real }
574 await $.store.set('calib', next)
575 lastSnapshot = ''
576 } catch (error) {
577 $.ui.log(`headroom: exact count failed: ${error instanceof Error ? error.message : String(error)}`, { to: 'debug' })
578 } finally {
579 isCalibrating = false
580 }
581 // the usage service is asked alongside, on the same 2 minutes
582 await refresh($, 'due')
583}
584
585// Auto compact acts whenever the chat is idle (no turn running): at the end of a turn,
586// on opening a chat, on any refresh between turns. What it does then is `decide`'s
587// (agent-policy.ts): at the cap (the %) it compacts; with Agent-timed on it also
588// compacts from the start %, unless the agent holds or its subagents still run. Paused
589// by `waitsForCompact` (the band is asking, or the person chose "after my next
590// compact", which is also what an unanswered ask means: cleared only by a real
591// compaction, never by the % flickering below) and by `isNewChatsOnly` ("only in new
592// chats"). `stuck`: a compaction already ran and the context is still past a %, so it
593// would only repeat; dropping below clears it, and stuck at the start never blocks the
594// cap. The % is compared as the band shows it, rounded.
595let waitsForCompact = false
596// a reply has been seen since the chat opened or was last compacted
597let hadReply = false
598let stuck: Stuck = 'no'
599let isNewChatsOnly = false
600let isBusy = false
601let isCompacting = false
602let lastPercent = 0
603// one look at a time: a look waits on the engine (the agents, a row told), and the 3s
604// tick must not start a second meanwhile
605let isWatching = false
606
607// Agent-timed Auto compact, the part that touches the engine: what the agent holds for
608// the session, its tool, what it is told, and what a compaction does to all of it. The
609// decisions and the words are agent-policy.ts's; this stays in the hooks module's own
610// file because the engine follows $ only into functions declared here.
611
612const reasonOf = (error: unknown) => (error instanceof Error ? error.message : String(error))
613
614// A row for the agent between turns: a user-role row the person does not see as typed.
615// The debug log has every one, appended or not (a test cannot see a plugin's rows).
616async function tell($: EngineInterface, text: string): Promise<void> {
617 let outcome = 'appended'
618 try {
619 const row = await $.session.append({ message: { type: 'user', content: [{ type: 'text', text }] } })
620 if (row.deny !== undefined) outcome = `not appended: ${row.deny}`
621 } catch (error) {
622 outcome = `not appended: ${reasonOf(error)}`
623 }
624 $.ui.log(`headroom: agent-timed row (${outcome}): ${text}`, { to: 'debug' })
625}
626
627// The main agent is waiting on work whose results it must still take in. An agent list
628// that cannot be read counts as none running: compaction is never held on a guess.
629async function hasRunningAgents($: EngineInterface): Promise<boolean> {
630 try {
631 return (await $.agent.list()).some(a => a.status === 'pending' || a.status === 'running' || a.status === 'waiting')
632 } catch {
633 return false
634 }
635}
636
637// the agent's tool, registered once, and only in a chat where Agent-timed is on (a tool
638// in the list rides on every request, and the API has no way to take one out again)
639let isToolRegistered = false
640async function ensureTool($: EngineInterface): Promise<boolean> {
641 if (isToolRegistered) return true
642 try {
643 await $.tool.register({ name: TOOL_NAME, description: TOOL_DESCRIPTION, inputSchema: TOOL_SCHEMA })
644 isToolRegistered = true
645 } catch (error) {
646 $.ui.log(`headroom: the compaction tool could not be registered: ${reasonOf(error)}`, { to: 'debug' })
647 }
648 return isToolRegistered
649}
650
651// Three places notice a compaction of the main conversation (auto compact's own direct
652// call, the session.compact hook, the reply total going blank): the first one handles
653// it, and the others find it handled until a reply has been seen again.
654let isCompactionHandled = false
655
656// The main conversation was compacted: the cycle starts over, and the agent gets its
657// note back, and word of a hold the cap ended, once.
658async function afterCompaction($: EngineInterface): Promise<void> {
659 isCompactionHandled = true
660 // the cache the chain kept is the old conversation's: the next response starts one afresh
661 await forget($, 'conversation compacted')
662 const state = await read($, agentTimed)
663 await update($, agentTimed, s => ({ ...EMPTY, chat: s.chat }))
664 const text = afterText(state.note, state.overridden, capOf(await read($, autoCompact)))
665 if (text !== null) await tell($, text)
666}
667
668// the request to compact is spent by one attempt, and dropped when the turn it was
669// made in is interrupted
670async function dropAsked($: EngineInterface): Promise<void> {
671 if ((await read($, agentTimed)).isAsked) await update($, agentTimed, s => ({ ...s, isAsked: false }))
672}
673
674// What rides on a main-agent tool result: word that the start % is passed (once a
675// cycle), the reminders of a hold growing old, the five-minute ask for where a hold
676// stands, and a breakpoint after a commit or a passing test run (`command`: a Bash
677// command that succeeded, else null).
678async function linesFor($: EngineInterface, command: string | null): Promise<string[]> {
679 const auto = await read($, autoCompact)
680 if (!isTimed(auto)) return []
681 const now = await $.clock.now()
682 const startAt = startOf(auto)
683 const cap = capOf(auto)
684 const percent = lastPercent
685 const shown = Math.round(percent)
686 const state = await read($, agentTimed)
687 const lines: string[] = []
688 let told = state.told
689 let nudge = state.nudge
690 if (told !== 'yes' && shown >= startAt) {
691 lines.push(toldText(percent, startAt, cap))
692 told = 'yes'
693 }
694 if (state.hold) {
695 const step = stepNudge(nudge, nudgeLevel(percent, startAt, cap, true))
696 nudge = step.nudge
697 if (step.isSaid) lines.push(nudgeText(nudge.level === 3 ? 3 : 2, percent, startAt, cap, state.hold.reason))
698 const kind = command === null || shown < startAt || nudge.isBreakpointSaid ? null : breakpointOf(command)
699 if (kind) {
700 lines.push(breakpointText(kind))
701 nudge = { ...nudge, isBreakpointSaid: true }
702 }
703 }
704 let hold = state.hold
705 const reminder = holdReminder(state, now, percent, startAt, cap)
706 if (reminder) {
707 lines.push(reminder.text)
708 hold = reminder.state.hold
709 }
710 if (told !== state.told || nudge !== state.nudge || hold !== state.hold) await update($, agentTimed, s => ({ ...s, told, nudge, hold }))
711 return lines
712}
713
714function compactNow($: EngineInterface): void {
715 if (isCompacting) return
716 isCompacting = true
717 // compacting takes longer than a press or a hook may run, so a timer starts it
718 $.clock.after(0, async () => {
719 try {
720 await $.command.run({ command: 'compact' })
721 } finally {
722 isCompacting = false
723 }
724 })
725}
726
727// Auto compact compacts through the session itself, not by typing /compact: a
728// command run from the end of a turn is refused while the turn winds down, which
729// is why it toasted and then did nothing. The compaction also rejects while a turn
730// runs, so it is tried again a few times, a couple of seconds apart.
731const AUTO_TRIES = 6
732let canCompactDirectly = true
733const AUTO_RETRY_MS = 2_000
734
735function autoCompactNow($: EngineInterface, attempt = 1): void {
736 if (attempt === 1) {
737 if (isCompacting) return
738 isCompacting = true
739 }
740 $.clock.after(attempt === 1 ? 1_000 : AUTO_RETRY_MS, async () => {
741 try {
742 // the agent's handoff note rides along as the summarizer's instructions (a typed
743 // /compact gets it from the session.compact hook instead)
744 const instructions = withNote(undefined, (await read($, agentTimed)).note)
745 // the desktop app (an SDK session) has no direct compaction: there /compact runs
746 // as a turn of its own, so it is typed, as the Compact button does
747 const result = canCompactDirectly
748 ? await $.session.compact(instructions === undefined ? undefined : { instructions })
749 : (await $.command.run({ command: 'compact' }), undefined)
750 // held as stuck while the figures are read again, so no tick fires a second one
751 stuck = 'cap'
752 isCompacting = false
753 // the attempt spends the agent's request, skipped or not: a skip must not loop
754 await dropAsked($)
755 if (result && 'skip' in result && result.skip) $.ui.toast(`Auto compact was skipped: ${result.skip}`)
756 else if (result) await afterCompaction($)
757 lastSnapshot = ''
758 const usage = await $.session.usage({ breakdown: 'summary' })
759 const shown = Math.round(usage.context.percent ?? 0)
760 // still past a % after compacting: say so once, and do not loop on it
761 const auto = await read($, autoCompact)
762 stuck = shown >= capOf(auto) ? 'cap' : isTimed(auto) && shown >= startOf(auto) ? 'start' : 'no'
763 if (stuck === 'cap') $.ui.toast(`Context is still at ${shown}% after compacting, past your ${capOf(auto)}%: auto compact waits until it drops below`)
764 if (stuck === 'start') $.ui.toast(`Context is still at ${shown}% after compacting, past your ${startOf(auto)}% start: Agent-timed waits until it drops below`)
765 await refresh($)
766 } catch (error) {
767 const reason = error instanceof Error ? error.message : String(error)
768 // any refusal of the direct way: type /compact from then on, as the button does
769 if (canCompactDirectly) {
770 canCompactDirectly = false
771 return autoCompactNow($, attempt + 1)
772 }
773 if (attempt < AUTO_TRIES) return autoCompactNow($, attempt + 1)
774 isCompacting = false
775 $.ui.toast(`Auto compact could not start: ${reason}`)
776 $.ui.log(`headroom: auto compact failed: ${reason}`, { to: 'debug' })
777 }
778 })
779}
780
781// called with the context % whenever it is read
782async function watchAuto($: EngineInterface, percent: number): Promise<void> {
783 lastPercent = percent
784 if (isWatching) return
785 isWatching = true
786 try {
787 const auto = await read($, autoCompact)
788 if (!auto.isOn || auto.at === null) return
789 const shown = Math.round(percent)
790 const cap = auto.at
791 const startAt = startOf(auto)
792 const timed = isTimed(auto)
793 // stuck eases as the context drops below what it was stuck past
794 if (stuck === 'cap' && shown < cap) stuck = timed && shown >= startAt ? 'start' : 'no'
795 if (stuck === 'start' && (!timed || shown < startAt)) stuck = 'no'
796 if (isBusy || isCompacting) return
797 const state = await read($, agentTimed)
798 // the agents are asked after only where they can decide: in the zone, nothing else in the way
799 const isOpenZone = timed && shown >= startAt && shown < cap && !state.hold && !state.isAsked
800 const verdict = decide({
801 percent, cap, startAt, isAgentTimed: timed, isPaused: waitsForCompact || isNewChatsOnly, stuck, state,
802 hasRunningAgents: isOpenZone && (await hasRunningAgents($)),
803 })
804 // a turn may have started while the engine was asked
805 if (isBusy || isCompacting) return
806 if (verdict.action === 'tell') {
807 await update($, agentTimed, s => ({ ...s, told: 'next' as const }))
808 await tell($, toldText(percent, startAt, cap))
809 return
810 }
811 if (verdict.action !== 'compact') return
812 if (verdict.why === 'cap' && state.hold) {
813 // no hold survives the cap: it ends here, and the row after the compaction says so
814 const { reason } = state.hold
815 await update($, agentTimed, s => ({ ...s, hold: null, overridden: { reason, percent: shown } }))
816 $.ui.toast(`Context at ${shown}%: Claude's hold ends at your ${cap}%, auto compacting`)
817 } else if (verdict.why === 'cap') $.ui.toast(`Context at ${shown}%: auto compacting (set at ${cap}%)`)
818 else if (verdict.why === 'start') $.ui.toast(`Context at ${shown}%: auto compacting (Agent-timed from ${startAt}%)`)
819 else $.ui.toast('Compacting as Claude asked')
820 autoCompactNow($)
821 } finally {
822 isWatching = false
823 }
824}
825
826async function saveAuto($: EngineInterface, next: AutoCompact): Promise<void> {
827 await update($, autoCompact, () => next)
828 await storeSet($, `autoCompact:${autoChat ?? (await chatId($))}`, next)
829}
830
831// auto compact is set per chat: a chat that never had it opens with the default (on at
832// 80%, Agent-timed from 30%); a chat reopened, or the app restarted, gets back its own. Read again
833// whenever the chat's id changes (a /clear goes on under a new one, unannounced)
834let autoChat: string | undefined
835// a /clear keeps the settings: the cleared chat's Auto compact and Keep cache warm go
836// to the id that follows it, saved under it as if set there, and the hold, the note and
837// the cache chain start over with the context. Nothing else takes a new id unannounced
838let carried: { auto: AutoCompact; warm: WarmSetting } | null = null
839async function carryOnClear($: EngineInterface): Promise<void> {
840 carried = { auto: await read($, autoCompact), warm: await read($, warmSetting) }
841}
842async function chatId($: EngineInterface): Promise<string> {
843 try {
844 return await $.session.id()
845 } catch {
846 return 'chat'
847 }
848}
849async function loadAuto($: EngineInterface): Promise<void> {
850 const id = await chatId($)
851 if (id === autoChat) return
852 autoChat = id
853 const carry = carried
854 carried = null
855 const saved = (await storeGet<AutoCompact>($, `autoCompact:${id}`)) ?? carry?.auto
856 isNewChatsOnly = false
857 stuck = 'no'
858 waitsForCompact = false
859 const auto: AutoCompact = saved ? { ...saved, at: saved.at ?? AT_DEFAULT } : DEFAULT_AUTO
860 await update($, autoCompact, () => auto)
861 if (carry) {
862 await storeSet($, `autoCompact:${id}`, auto)
863 await storeSet($, `warm:${id}`, carry.warm)
864 }
865 await update($, autoAsk, () => null)
866 // another chat: what the agent held, noted or asked for belonged to the last one. A
867 // reload of the module is not another chat: the state keeps the hold and the note
868 if ((await read($, agentTimed)).chat !== id) await update($, agentTimed, () => ({ ...EMPTY, chat: id }))
869 if (isTimed(auto)) await ensureTool($)
870 await loadWarm($, id)
871}
872
873// Keep cache warm, the part that touches the engine: the chain of refreshes (ported
874// from cache-warmer's register.tsx, MIT, see cache-policy.ts), the lifetime, the totals
875// and the rate the band learns. The prices, the rule and the words are
876// cache-policy.ts's; this stays here because the engine follows $ only into functions
877// declared in this file.
878
879// set from the /config rows when the module loads, and by their config.set hooks
880let idleLimits: Record<Ttl, number> = { '5m': IDLE_LIMIT_DEFAULT, '1h': IDLE_LIMIT_DEFAULT }
881let warmUntil = WARM_UNTIL_DEFAULT
882let defaultChoice: TtlChoice = 'auto'
883let warmTimer: { cancel(): void } | undefined
884// The anchor a refresh has claimed until chain() records it: a schedule() meanwhile,
885// from a turn that ended, cannot fork it again
886let forking: { at: number } | undefined
887// main prompts started this process: a warning from before the latest one is stale
888let prompts = 0
889// the lifetime this mod put in the variable (null: none), and what was there before it
890let envSet: Ttl | null = null
891let envBefore: string | undefined
892// the live window's last response as refresh() last read it, and when: what anchors stand on
893let lastApi: ModelUsage | null = null
894let lastApiAt = -Infinity
895// the response the anchor stands on, and when it was seen: the same counts read again
896// are no new response
897let anchoredApi = ''
898let anchoredAt = -Infinity
899// when the running (or last) main turn started
900let turnStartedAt = -Infinity
901// what refreshes spent since the last main turn ended: the meter's next jump holds it too
902let spentSince = 0
903
904// the variables Claude Code picks the lifetime by, each named outright (the engine lists
905// what a module reads); one that cannot be read counts as unset
906async function lifetimeVars($: EngineInterface): Promise<{ force5m?: string; ttl?: string; enable1h?: string }> {
907 const vars: { force5m?: string; ttl?: string; enable1h?: string } = {}
908 try {
909 vars.force5m = await $.env.get('FORCE_PROMPT_CACHING_5M')
910 vars.ttl = await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL')
911 vars.enable1h = await $.env.get('ENABLE_PROMPT_CACHING_1H')
912 } catch {
913 // the environment cannot be read: Claude Code's own default stands
914 }
915 return vars
916}
917
918// the promptCacheTtl setting; no row listed reads as unset
919async function settingTtlOf($: EngineInterface): Promise<string | undefined> {
920 try {
921 const row = (await $.config.list()).find(r => r.key === 'promptCacheTtl')
922 return row === undefined ? undefined : String(row.value)
923 } catch {
924 return undefined
925 }
926}
927
928// The lifetime in force, as Claude Code decides it (cache-policy's ttlOf), unless a
929// refresh found a 1h cache gone: then 5m for the rest of the session
930async function lifetimeOf($: EngineInterface, limits: readonly Limit[]): Promise<Ttl> {
931 const { assumed } = await read($, warm)
932 if (assumed) return assumed
933 const vars = await lifetimeVars($)
934 return ttlOf(
935 { force5m: isEnvOn(vars.force5m), envTtl: vars.ttl, settingTtl: await settingTtlOf($), enable1h: isEnvOn(vars.enable1h), ...planOf(limits) },
936 envSet ?? 'auto',
937 )
938}
939
940// A chosen 5m or 1h goes in the variable, for the process, from the next request,
941// warming on or off; auto puts back what was there. Once the first response wrote the cache
942// its lifetime holds for the session, as cache-warmer has it: a change waits for a new
943// one, and a press says so (`isPressed`)
944async function applyTtl($: EngineInterface, isPressed: boolean): Promise<void> {
945 const setting = await read($, warmSetting)
946 const want = setting.ttl !== 'auto' ? setting.ttl : null
947 if (want === envSet) return
948 const state = await read($, warm)
949 if (state.isLocked) {
950 if (isPressed) $.ui.toast(`Cache lifetime ${setting.ttl} applies to new sessions: this one's cache is written at ${state.ttl}`)
951 return
952 }
953 try {
954 if (envSet === null) envBefore = (await lifetimeVars($)).ttl
955 await $.env.set('CLAUDE_CODE_PROMPT_CACHE_TTL', want ?? envBefore)
956 envSet = want
957 } catch (error) {
958 $.ui.log(`headroom: the cache lifetime could not be set: ${reasonOf(error)}`, { to: 'debug' })
959 }
960}
961
962// a notice row: the transcript keeps it, no request carries it; the debug log has every
963// one, appended or not (a test cannot see a plugin's rows)
964async function cacheRow($: EngineInterface, text: string): Promise<void> {
965 let outcome = 'appended'
966 try {
967 const row = await $.session.append({ message: { type: 'system', content: [{ type: 'text', text }] } })
968 if (row.deny !== undefined) outcome = `not appended: ${row.deny}`
969 } catch (error) {
970 outcome = `not appended: ${reasonOf(error)}`
971 }
972 $.ui.log(`headroom: cache row (${outcome}): ${text}`, { to: 'debug' })
973}
974
975// this session's totals and the all-time ones in the store take the same delta
976async function addToTotals($: EngineInterface, delta: Partial<WarmTotals>): Promise<void> {
977 await update($, warm, s => ({ ...s, totals: addTotals(s.totals, delta) }))
978 const stored = (await storeGet<Warm['allTime']>($, 'warmAllTime')) ?? { ...ZERO_TOTALS, since: await $.clock.now() }
979 const allTime = addTotals(stored, delta)
980 await storeSet($, 'warmAllTime', allTime)
981 await update($, warm, s => ({ ...s, allTime }))
982}
983
984async function reportStop($: EngineInterface, reason: string, isPaused = false): Promise<void> {
985 await update($, warm, s => ({ ...s, status: isPaused ? { state: 'stopped' as const, reason, isPaused } : { state: 'stopped' as const, reason } }))
986 $.ui.log(`headroom: cache warming stopped: ${reason}`, { to: 'debug' })
987}
988
989// Compaction, /clear, a session's end and a model switch forget the chain: only the
990// next response starts it again, and its fee is wasted now, since no prompt will judge it
991async function forget($: EngineInterface, reason: string): Promise<void> {
992 warmTimer?.cancel()
993 warmTimer = undefined
994 const { anchor } = await read($, warm)
995 if (anchor && anchor.feeUsd > 0) await addToTotals($, { wastedUsd: anchor.feeUsd })
996 await update($, warm, s => ({ ...s, anchor: null }))
997 await reportStop($, reason)
998}
999
1000// Stops the chain anchored at `at`, adding `feeUsd` to it, until the next response; a
1001// response that replaced it meanwhile keeps its own warming. Answers whether it stopped it
1002async function stopChain($: EngineInterface, at: number, reason: string, feeUsd = 0, isPaused = false): Promise<boolean> {
1003 let isStopped = false
1004 await update($, warm, s => {
1005 const { anchor } = s
1006 isStopped = anchor?.at === at && !anchor.isStopped
1007 if (!anchor || !isStopped) return s
1008 return { ...s, anchor: { ...anchor, feeUsd: anchor.feeUsd + feeUsd, isStopped: true } }
1009 })
1010 if (isStopped) await reportStop($, reason, isPaused)
1011 return isStopped
1012}
1013
1014// a limit at or past warmUntil: no refresh spends more of it
1015async function pastLimit($: EngineInterface): Promise<string | null> {
1016 return pastLimitOf((await read($, snapshot))?.limits ?? [], warmUntil)
1017}
1018
1019// The next refresh, at 90% of the lifetime after the last request or refresh, if warming
1020// is on and one pays. A running turn stops at the run horizon, an idle session after its
1021// lifetime's idle limit. The anchor is read last, so the timer is set from it as it stands
1022async function schedule($: EngineInterface): Promise<void> {
1023 const prompt = prompts
1024 if (!(await read($, warmSetting)).isOn) return
1025 const past = await pastLimit($)
1026 const now = await $.clock.now()
1027 const { anchor: current, isRunning, outputTokens } = await read($, warm)
1028 if (!current || current.isStopped || current.at === forking?.at) return
1029 const phase: 'run' | 'idle' = isRunning ? 'run' : 'idle'
1030 const nextAt = current.lastAt + delayOf(current.ttl)
1031 const horizon = horizonOf(current.ttl)
1032 if (phase === 'run' && nextAt > current.at + horizon) {
1033 await stopChain($, current.at, `${formatDuration(horizon)} run limit reached`)
1034 return
1035 }
1036 if (past) {
1037 await stopChain($, current.at, past, 0, true)
1038 return
1039 }
1040 const limit = idleLimits[current.ttl]
1041 if (phase === 'idle' && current.idleRefreshes >= limit) {
1042 const expiresAt = current.lastAt + TTL_MS[current.ttl]
1043 // a limit of 0 turned idle warming off on purpose: it needs no warning
1044 const isStopped = await stopChain($, current.at, limit > 0 ? idleStopReason(limit, expiresAt) : 'no idle refreshes (set to 0)')
1045 if (isStopped && limit > 0 && prompt === prompts) await cacheRow($, idleStopNotice(current.ttl, limit, expiresAt))
1046 return
1047 }
1048 const decision = decideWarm(current.model, current.promptTokens, current.ttl, phase, outputTokens)
1049 if (!decision) {
1050 await stopChain($, current.at, `no price for ${current.model}`)
1051 return
1052 }
1053 if (decision.reason) {
1054 await stopChain($, current.at, decision.reason)
1055 return
1056 }
1057 warmTimer?.cancel()
1058 warmTimer = $.clock.after(Math.max(0, nextAt - now), () => void refreshCache($))
1059 await update($, warm, s => ({ ...s, status: { state: 'scheduled' as const, nextAt, phase, expectedUsd: decision.expectedUsd } }))
1060 $.ui.log(`headroom: cache refresh in ${formatDuration(nextAt - now)} (${current.ttl}, ${phase}, expected saving ${formatUsd(decision.expectedUsd)})`, { to: 'debug' })
1061}
1062
1063// the timer's refresh: it claims its anchor until chain() records it there
1064async function refreshCache($: EngineInterface): Promise<void> {
1065 warmTimer = undefined
1066 if (!(await read($, warmSetting)).isOn) return
1067 const { anchor: current, isRunning, outputTokens } = await read($, warm)
1068 if (!current || current.isStopped || current.at === forking?.at) return
1069 const claim = { at: current.at }
1070 forking = claim
1071 try {
1072 await forkFor($, current, isRunning ? 'run' : 'idle', outputTokens)
1073 } finally {
1074 if (forking === claim) forking = undefined
1075 }
1076}
1077
1078async function forkFor($: EngineInterface, current: WarmAnchor, phase: 'run' | 'idle', outputTokens: number): Promise<void> {
1079 const at = await $.clock.now()
1080 if (at > deadlineOf(current.lastAt, current.ttl)) {
1081 await stopChain($, current.at, 'refresh deadline missed')
1082 return
1083 }
1084 const past = await pastLimit($)
1085 if (past) {
1086 await stopChain($, current.at, past, 0, true)
1087 return
1088 }
1089 const decision = decideWarm(current.model, current.promptTokens, current.ttl, phase, outputTokens)
1090 if (!decision || decision.reason) {
1091 await stopChain($, current.at, decision?.reason ?? `no price for ${current.model}`)
1092 return
1093 }
1094 await update($, warm, s => ({ ...s, status: { state: 'refreshing' as const } }))
1095 let reply: ModelForkResult
1096 try {
1097 reply = await $.model.fork({ prompt: FORK_PROMPT })
1098 } catch (error) {
1099 await stopChain($, current.at, `refresh failed (${reasonOf(error)})`)
1100 return
1101 }
1102 if (!reply.isAnswered && reply.reason === 'nothing-to-fork') {
1103 await stopChain($, current.at, 'nothing to refresh')
1104 return
1105 }
1106 await settle($, current, at, phase, decision.missUsd, reply)
1107}
1108
1109async function settle($: EngineInterface, current: WarmAnchor, at: number, phase: 'run' | 'idle', missUsd: number, reply: ForkReply): Promise<void> {
1110 const usage = usageOf(reply.usage)
1111 // the fork's own write is its short tail: priced at the cache's lifetime it is overstated at most
1112 const costUsd = costOf(current.model, usage, current.ttl)
1113 const outcome = outcomeOf(reply, usage, current.promptTokens)
1114 const savesUsd = outcome.result === 'warmed' && costUsd !== null ? missUsd - costUsd : null
1115 const entry: Refresh = { at, model: current.model, usage, costUsd, savesUsd, ...outcome }
1116 await addToTotals($, { refreshes: 1, costUsd: costUsd ?? 0 })
1117 spentSince += costUsd ?? 0
1118 await cacheRow($, noticeText(current.ttl, entry))
1119 await chain($, current, entry, phase)
1120}
1121
1122// moves a warm refresh's chain on, when the chain is still current; answers whether it was
1123async function extend($: EngineInterface, current: WarmAnchor, entry: Refresh, phase: 'run' | 'idle'): Promise<boolean> {
1124 let isChained = false
1125 await update($, warm, s => {
1126 const { anchor } = s
1127 isChained = anchor?.at === current.at && !anchor.isStopped
1128 if (!anchor || !isChained) return s
1129 return {
1130 ...s,
1131 outputTokens: entry.usage?.output || DEFAULT_OUTPUT_TOKENS,
1132 anchor: {
1133 ...anchor,
1134 lastAt: entry.at,
1135 refreshes: anchor.refreshes + 1,
1136 idleRefreshes: anchor.idleRefreshes + (phase === 'idle' ? 1 : 0),
1137 feeUsd: anchor.feeUsd + (entry.costUsd ?? 0),
1138 },
1139 }
1140 })
1141 return isChained
1142}
1143
1144// The chain carries each refresh's fee. One a response overtook during its fork kept
1145// nothing that response read, so its fee is wasted; so is one whose chain was forgotten.
1146// A 1h cache a refresh found gone means the lifetime was 5m: assumed so from then
1147async function chain($: EngineInterface, current: WarmAnchor, entry: Refresh, phase: 'run' | 'idle'): Promise<void> {
1148 const costUsd = entry.costUsd ?? 0
1149 const isWarmed = entry.result === 'warmed'
1150 const isAssumed = entry.result === 'expired' && current.ttl === '1h'
1151 const reason =
1152 entry.result === 'expired'
1153 ? isAssumed ? 'the cache had expired: 5m assumed for this session' : 'the cache had expired'
1154 : `refresh failed${entry.detail ? ` (${entry.detail})` : ''}`
1155 if (isAssumed) {
1156 await update($, warm, s => ({
1157 ...s,
1158 assumed: '5m' as const,
1159 ttl: '5m' as const,
1160 anchor: s.anchor && s.anchor.at === current.at ? { ...s.anchor, ttl: '5m' as const } : s.anchor,
1161 }))
1162 }
1163 const isChained = isWarmed ? await extend($, current, entry, phase) : await stopChain($, current.at, reason, costUsd)
1164 if (forking?.at === current.at) forking = undefined
1165 if (!isChained) {
1166 if (costUsd > 0) await addToTotals($, { wastedUsd: costUsd })
1167 if ((await read($, warm)).status.state === 'refreshing') await update($, warm, s => ({ ...s, status: { state: 'waiting' as const } }))
1168 await schedule($)
1169 return
1170 }
1171 if (isWarmed) await schedule($)
1172}
1173
1174// A prompt is kept when it read a cache that would have expired without the refreshes
1175// since the last one; it avoided rewriting what it read. Otherwise the chain's fee is wasted
1176async function judgeChain($: EngineInterface, previous: WarmAnchor | null, at: number, cacheRead: number): Promise<void> {
1177 if (!previous) return
1178 const isKept = previous.refreshes > 0 && at - previous.at > TTL_MS[previous.ttl] && cacheRead >= previous.promptTokens / 2
1179 const keptUsd = isKept ? missCostOf(previous.model, cacheRead, previous.ttl) : null
1180 if (keptUsd !== null) await addToTotals($, { kept: 1, keptUsd })
1181 else if (previous.feeUsd > 0) await addToTotals($, { wastedUsd: previous.feeUsd })
1182}
1183
1184// A response of the main conversation was seen (the live window's counts moved): the
1185// chain before it is judged, and a new one starts from it. `model`: the turn's at its
1186// end; mid-turn the anchor's own (a fork's response is not the live window's). A turn
1187// that ended with usage had a response even if its counts read the same as the last
1188async function anchorOn($: EngineInterface, api: ModelUsage, at: number, model: string | undefined, isNew = false): Promise<boolean> {
1189 const key = JSON.stringify(api)
1190 const promptTokens = api.input_tokens + api.cache_read_input_tokens + api.cache_creation_input_tokens
1191 if ((key === anchoredApi && !isNew) || promptTokens <= 0) return false
1192 const previous = (await read($, warm)).anchor
1193 const named = model ?? previous?.model
1194 if (named === undefined) return false
1195 anchoredApi = key
1196 anchoredAt = at
1197 await judgeChain($, previous, at, api.cache_read_input_tokens)
1198 const { ttl } = await read($, warm)
1199 await update($, warm, s => ({
1200 ...s,hooks/agent-policy.ts 268 lines1import type { AgentTimed, AutoCompact } from '../types'
2
3// Agent-timed Auto compact, the part with no engine in it: when to compact, what the
4// agent's tool answers, and every word the agent reads. The idea (the agent holds
5// compaction through fragile work and releases it at a safe point, leaving a note that
6// survives) is compactor's, github.com/rhwendt/compactor (MIT), as is the shape of the
7// breakpoint patterns.
8
9export const TOOL_NAME = 'compaction'
10export const TOOL = 'mcp__headroom__compaction'
11export const TOOL_DESCRIPTION =
12 'Controls when this conversation is compacted. Compaction runs when your turn ends once context passes the start %. ' +
13 '"hold" (with a reason) before fragile multi-step work whose state lives only in this conversation; "release" at a safe point; ' +
14 '"compact" to compact when this turn ends; "note" to save what must survive (current hypothesis, next steps, file:line references). ' +
15 'At the cap % compaction runs whatever is held.'
16export const TOOL_SCHEMA = {
17 type: 'object',
18 properties: {
19 action: { enum: ['hold', 'release', 'compact', 'note', 'status'] },
20 reason: { type: 'string' },
21 note: { type: 'string' },
22 },
23 required: ['action'],
24}
25
26// the cap (the auto compact %) whenever there is none: a blank field, a first switch-on
27export const AT_DEFAULT = 80
28// the start %: 10 at the least, always below the cap; 30 where none was set
29export const START_MIN = 10
30export const START_DEFAULT = 30
31// a chat that never had a setting: Auto compact on at the cap, Agent-timed on from the start %
32export const DEFAULT_AUTO: AutoCompact = { isOn: true, at: AT_DEFAULT, isAgentTimed: true, startAt: START_DEFAULT }
33export const NOTE_MAX = 4_000
34// a reason is kept to this, and drawn in the band to less
35export const REASON_MAX = 500
36export const REASON_SHOWN = 80
37// "compact" under this context % is refused: there is nothing worth compacting
38export const ASK_MIN = 10
39// within this many points of the cap, a holding agent is told the cap is close
40export const NEAR_CAP = 5
41// main tool calls between repeats of that last warning
42export const NUDGE_EVERY = 10
43// while a hold lasts, the agent is asked this often to say where the hold stands
44export const HOLD_REMIND_MS = 5 * 60_000
45
46export const EMPTY: AgentTimed = {
47 chat: null,
48 hold: null,
49 note: null,
50 isAsked: false,
51 told: 'no',
52 nudge: { level: 0, calls: 0, isBreakpointSaid: false },
53 overridden: null,
54}
55
56// a chat's setting as it stands: the cap, whether Agent-timed is on (it needs auto
57// compact on), and the start %: the saved one or the default, never under the least,
58// always below the cap
59export const capOf = (auto: AutoCompact): number => auto.at ?? AT_DEFAULT
60export const isTimed = (auto: AutoCompact): boolean => auto.isOn && auto.isAgentTimed === true
61export const startOf = (auto: AutoCompact): number => Math.min(Math.max(START_MIN, auto.startAt ?? START_DEFAULT), capOf(auto) - 1)
62
63/** a compaction ran and the context is still past a %: `start` never blocks the cap */
64export type Stuck = 'no' | 'start' | 'cap'
65
66export type DecideInput = {
67 percent: number
68 cap: number
69 startAt: number
70 isAgentTimed: boolean
71 /** the band is asking, "After my next compact", or "Only in new chats" */
72 isPaused: boolean
73 stuck: Stuck
74 state: Pick<AgentTimed, 'hold' | 'isAsked' | 'told'>
75 hasRunningAgents: boolean
76}
77
78export type Decision =
79 | { action: 'compact'; why: 'asked' | 'cap' | 'start' }
80 | { action: 'tell' }
81 | { action: 'wait'; why: 'paused' | 'stuck' | 'below' | 'held' | 'agents' | 'told' }
82
83// What auto compact does with an idle chat: the first rule that applies decides. The %
84// is compared as the band shows it, rounded.
85export function decide(i: DecideInput): Decision {
86 const shown = Math.round(i.percent)
87 // the agent's own request is as deliberate as the Compact button: no pause stops it
88 if (i.state.isAsked) return { action: 'compact', why: 'asked' }
89 if (i.isPaused) return { action: 'wait', why: 'paused' }
90 // the cap: no hold, no subagent and no untold agent survives it
91 if (shown >= i.cap) return i.stuck === 'cap' ? { action: 'wait', why: 'stuck' } : { action: 'compact', why: 'cap' }
92 if (!i.isAgentTimed || shown < i.startAt) return { action: 'wait', why: 'below' }
93 if (i.stuck !== 'no') return { action: 'wait', why: 'stuck' }
94 if (i.state.hold) return { action: 'wait', why: 'held' }
95 // the main agent is waiting on work whose results it must still take in
96 if (i.hasRunningAgents) return { action: 'wait', why: 'agents' }
97 // never compacted unawares: told first, and given the coming turn to hold
98 if (i.state.told === 'no') return { action: 'tell' }
99 if (i.state.told === 'next') return { action: 'wait', why: 'told' }
100 return { action: 'compact', why: 'start' }
101}
102
103// How firmly a holding agent is reminded: 2 from halfway between the start % and the
104// cap, 3 from NEAR_CAP points under the cap.
105export function nudgeLevel(percent: number, startAt: number, cap: number, isHolding: boolean): 0 | 2 | 3 {
106 const shown = Math.round(percent)
107 if (!isHolding || shown < startAt) return 0
108 if (shown >= cap - NEAR_CAP) return 3
109 if (shown >= startAt + (cap - startAt) / 2) return 2
110 return 0
111}
112
113// One main tool call later: a level is said when first reached, and level 3 again
114// every NUDGE_EVERY calls while it lasts.
115export function stepNudge(nudge: AgentTimed['nudge'], level: 0 | 2 | 3): { isSaid: boolean; nudge: AgentTimed['nudge'] } {
116 if (level === 0) return { isSaid: false, nudge }
117 if (level > nudge.level) return { isSaid: true, nudge: { ...nudge, level, calls: 0 } }
118 if (level < 3) return { isSaid: false, nudge }
119 const calls = nudge.calls + 1
120 return calls >= NUDGE_EVERY ? { isSaid: true, nudge: { ...nudge, calls: 0 } } : { isSaid: false, nudge: { ...nudge, calls } }
121}
122
123// a command word: at the start, or after a space or a shell separator
124const SEP = String.raw`(?:^|[\s;&|(])`
125const END = String.raw`(?=$|[\s;&|)])`
126const COMMIT = new RegExp(SEP + String.raw`git(?:\s+-[Cc]\s+\S+)*\s+commit` + END)
127const TESTS = new RegExp(
128 SEP +
129 '(?:' +
130 String.raw`pytest|py\.test|python3?\s+-m\s+(?:pytest|unittest)` +
131 String.raw`|(?:npm|pnpm|yarn|bun)\s+(?:run\s+)?test(?::[\w-]+)?` +
132 String.raw`|go\s+test|cargo\s+test|make\s+(?:test|check)|mvn(?:\s+\S+)*?\s+test` +
133 String.raw`|(?:\./)?gradlew?(?:\s+\S+)*?\s+test|rspec|(?:npx\s+)?(?:jest|vitest)|claude\s+plugin\s+test` +
134 ')' +
135 END,
136)
137const SEGMENTS = /&&|\|\||[;|&\n]/
138
139// Which natural breakpoint a Bash command that succeeded was: read one shell segment
140// at a time, so "echo git commit" and a --dry-run are not one.
141export function breakpointOf(command: string): 'commit' | 'tests' | null {
142 const segments = command.split(SEGMENTS).filter(s => s.trim() !== '')
143 if (segments.some(s => COMMIT.test(s) && !s.includes('--dry-run'))) return 'commit'
144 if (segments.some(s => TESTS.test(s))) return 'tests'
145 return null
146}
147
148export function figures(percent: number, startAt: number, cap: number): string {
149 return `Context ${Math.round(percent)}%. Agent-timed compaction starts at ${startAt}%; at ${cap}% it runs whatever is held.`
150}
151
152export function toldText(percent: number, startAt: number, cap: number): string {
153 return (
154 `Agent-timed compaction: context is at ${Math.round(percent)}% (starts at ${startAt}%, cap ${cap}%). ` +
155 'This conversation will be compacted when your turn ends. ' +
156 'If you are mid-task, call the compaction tool with action "hold" and a reason. Otherwise save a "note" of what must survive.'
157 )
158}
159
160export function nudgeText(level: 2 | 3, percent: number, startAt: number, cap: number, reason: string): string {
161 const body =
162 level === 3
163 ? `the cap is close. At ${cap}% compaction runs when your turn ends, whatever is held. Save a note now.`
164 : `you are holding (${reason}) well past the start. Finish the current step, save a note, and release.`
165 return `Agent-timed compaction: ${body}\n${figures(percent, startAt, cap)}`
166}
167
168// Every HOLD_REMIND_MS of a hold, on the next main tool result: the agent is asked to
169// say where the hold stands (kept, released, or a note). `null` while none is due; the
170// state handed back is due again HOLD_REMIND_MS on.
171export function holdReminder(state: AgentTimed, now: number, percent: number, startAt: number, cap: number): { text: string; state: AgentTimed } | null {
172 if (!state.hold || now < state.hold.remindAt) return null
173 const held = Math.max(1, Math.round((now - state.hold.since) / 60_000))
174 const text =
175 `Agent-timed compaction: you have held compaction for ${held}m (${state.hold.reason}). ` +
176 'Update the hold: call the compaction tool with action "hold" and the current reason to keep it, "release" if the fragile step is done, or "note" what must survive.\n' +
177 figures(percent, startAt, cap)
178 return { text, state: { ...state, hold: { ...state.hold, remindAt: now + HOLD_REMIND_MS } } }
179}
180
181export function breakpointText(kind: 'commit' | 'tests'): string {
182 const what = kind === 'commit' ? 'a commit just landed' : 'tests just passed'
183 return `Agent-timed compaction: ${what}, a natural breakpoint. Consider "release" or "compact", with a note.`
184}
185
186const NOTE_HEAD = 'The agent left this handoff note. Keep what it says matters:'
187
188// The note as the summarizer's instructions, after whatever was asked for already.
189// Adding it twice adds it once: a compaction may pass more than one place that adds it.
190export function withNote(instructions: string | undefined, note: string | null): string | undefined {
191 if (note === null) return instructions
192 const block = `${NOTE_HEAD}\n${note}`
193 if (instructions?.includes(block)) return instructions
194 return instructions ? `${instructions}\n\n${block}` : block
195}
196
197// The row the agent reads after a compaction, or null with nothing to say.
198export function afterText(note: string | null, overridden: AgentTimed['overridden'], cap: number): string | null {
199 if (note === null && overridden === null) return null
200 const lines = ['Agent-timed compaction: the conversation was just compacted.']
201 if (overridden) lines.push(`Your hold (${overridden.reason}) ended at the ${cap}% cap. Hold again if the work is still fragile.`)
202 if (note !== null) lines.push(`Handoff note you left:\n${note}`)
203 return lines.join('\n')
204}
205
206export type ToolInput = { action?: unknown; reason?: unknown; note?: unknown }
207export type ToolContext = {
208 /** Agent-timed is on in this chat */
209 isOn: boolean
210 /** the call came from a subagent's or a teammate's loop */
211 isSubagent: boolean
212 percent: number
213 startAt: number
214 cap: number
215 now: number
216}
217
218const ACTIONS = ['hold', 'release', 'compact', 'note', 'status']
219
220function statusText(state: AgentTimed, now: number): string {
221 const lines = [
222 state.hold ? `hold: ${state.hold.reason} (${Math.max(0, Math.round((now - state.hold.since) / 60_000))}m)` : 'hold: none',
223 state.note === null ? 'note: none' : `note: ${state.note}`,
224 ]
225 if (state.isAsked) lines.push('compaction asked for: when this turn ends')
226 return lines.join('\n')
227}
228
229// What a call of the tool does and answers. A refusal says what was wrong and changes
230// nothing: the state handed back is then the very one handed in.
231export function answerTool(state: AgentTimed, input: ToolInput, ctx: ToolContext): { state: AgentTimed; text: string } {
232 if (!ctx.isOn) return { state, text: 'Agent-timed compaction is off in this chat. Nothing changed.' }
233 const say = (text: string, next: AgentTimed = state) => ({ state: next, text: `${text}\n${figures(ctx.percent, ctx.startAt, ctx.cap)}` })
234 const { action } = input
235 if (typeof action !== 'string' || !ACTIONS.includes(action)) {
236 return say('Unknown action. Call it with action "hold" (and a reason), "release", "compact", "note" or "status". Nothing changed.')
237 }
238 if (action === 'status') return say(statusText(state, ctx.now))
239 if (ctx.isSubagent) {
240 return say('Only the main agent can hold, release, compact or write notes: the hold and the note are its own. Finish your task and report back. Nothing changed.')
241 }
242 if (action === 'hold') {
243 const reason = typeof input.reason === 'string' ? input.reason.trim().slice(0, REASON_MAX) : ''
244 if (reason === '') return say('A hold needs a reason: action "hold", reason "<what is fragile>". Nothing changed.')
245 const hold = { reason, since: state.hold?.since ?? ctx.now, remindAt: ctx.now + HOLD_REMIND_MS }
246 return say(`${state.hold ? 'Hold updated' : 'Hold set'}: compaction waits until you release, or until the cap. Reason: ${reason}.`, {
247 ...state,
248 hold,
249 isAsked: false,
250 nudge: { ...state.nudge, isBreakpointSaid: false },
251 })
252 }
253 const note = typeof input.note === 'string' ? input.note : undefined
254 if (note !== undefined && note.length > NOTE_MAX) return say(`The note is ${note.length} characters; the most is ${NOTE_MAX}. Nothing changed.`)
255 const noted = (next: AgentTimed): AgentTimed => (note === undefined ? next : { ...next, note: note === '' ? null : note })
256 const saved = note === undefined ? '' : note === '' ? ' Handoff note cleared.' : ' Handoff note saved; it comes back after the compaction.'
257 if (action === 'release') {
258 const when = Math.round(ctx.percent) >= ctx.startAt ? 'Compaction can run when this turn ends.' : `Compaction waits until context reaches ${ctx.startAt}%.`
259 return say(`${state.hold ? 'Released.' : 'No hold was set.'} ${when}${saved}`, noted({ ...state, hold: null }))
260 }
261 if (action === 'compact') {
262 if (Math.round(ctx.percent) < ASK_MIN) return say(`Context is under ${ASK_MIN}%: there is nothing worth compacting. Nothing changed.`)
263 return say(`Compaction runs when this turn ends.${saved}`, noted({ ...state, hold: null, isAsked: true }))
264 }
265 if (note === undefined) return say('A note needs text: action "note", note "<what must survive>" (an empty note clears it). Nothing changed.')
266 return say(note === '' ? 'Handoff note cleared.' : `Handoff note saved (${note.length} characters). It goes to the summarizer and comes back after the next compaction.`, noted(state))
267}
268hooks/cache-policy.ts 467 lines1import type { ConfigValue, ModelForkResult, ModelUsage } from 'claude-code'
2
3import type { CacheUsage, Limit, Ttl, TtlChoice, WarmAnchor, WarmRate, WarmStatus, WarmTotals } from '../types'
4
5// Keep cache warm, the part with no engine in it: the prices, when a refresh pays,
6// what one costs and saves, the lifetime Claude Code picks, and every word the band
7// and the transcript show.
8//
9// Ported from cache-warmer 0.12.0 (github.com/paulbkim-dev/claude-code-cache-warmer),
10// its hooks/warmer.ts, under the MIT licence:
11//
12// Copyright (c) 2026 Paul B. Kim
13//
14// Permission is hereby granted, free of charge, to any person obtaining a copy
15// of this software and associated documentation files (the "Software"), to deal
16// in the Software without restriction, including without limitation the rights
17// to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
18// copies of the Software, and to permit persons to whom the Software is
19// furnished to do so, subject to the following conditions:
20//
21// The above copyright notice and this permission notice shall be included in all
22// copies or substantial portions of the Software.
23//
24// THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
25// IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
26// FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
27// AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
28// LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
29// OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
30// SOFTWARE.
31//
32// The rule of when a refresh pays (`decide`) is the cache warmer's in Pi
33// (github.com/earendil-works/pi) by Mario Zechner, which cache-warmer ports.
34
35// the prices below were read on this day: a model added since gets no warming, never a wrong one
36export const PRICES_AS_OF = '2026-10-04'
37
38type Price = { input: number; cacheRead: number; output: number }
39
40const price = (input: number, cacheRead: number, output: number): Price => ({ input, cacheRead, output })
41
42// USD per million tokens at standard rates (platform.claude.com/docs/en/about-claude/pricing).
43// A cache write costs 1.25x input for the 5-minute lifetime, 2x for the 1-hour one
44const PRICES: Record<string, Price> = {
45 'claude-fable-5-1': price(10, 0.25, 50),
46 'claude-mythos-5-1': price(10, 0.25, 50),
47 'claude-fable-5': price(10, 1, 50),
48 'claude-mythos-5': price(10, 1, 50),
49 'claude-opus-5-5': price(4, 0.2, 20),
50 'claude-opus-5': price(5, 0.5, 25),
51 'claude-opus-4-8': price(5, 0.5, 25),
52 'claude-opus-4-7': price(5, 0.5, 25),
53 'claude-opus-4-6': price(5, 0.5, 25),
54 'claude-opus-4-5': price(5, 0.5, 25),
55 'claude-sonnet-5-5': price(2, 0.2, 10),
56 'claude-sonnet-5': price(2, 0.2, 10),
57 'claude-sonnet-4-6': price(3, 0.3, 15),
58 'claude-sonnet-4-5': price(3, 0.3, 15),
59 'claude-haiku-4-5': price(1, 0.1, 5),
60}
61
62const WRITE_MULTIPLIER: Record<Ttl, number> = { '5m': 1.25, '1h': 2 }
63
64export const TTL_MS: Record<Ttl, number> = { '5m': 300_000, '1h': 3_600_000 }
65
66// Pi's rule: a refresh is sent only when it is expected to save this much
67export const MIN_SAVINGS_USD = 0.05
68// Pi's measured chance that a prompt comes before the cache expires while the session is idle
69export const IDLE_CONTINUATION = 0.15
70// a fork has no output cap; Opus 5.5 at high effort answered a refresh in 182 output tokens
71export const DEFAULT_OUTPUT_TOKENS = 200
72// the refreshes each lifetime may send while no turn runs, by default and at most
73export const IDLE_LIMIT_DEFAULT = 5
74export const IDLE_LIMIT_MAX = 20
75// no refresh while either limit is at or past this %: a refresh spends the plan like any request
76export const WARM_UNTIL_DEFAULT = 85
77// the band warns this long before the cache expires
78export const EXPIRING_MS = 2 * 60_000
79// one window's reset time as two readings give it: a few minutes apart at most
80const SAME_WINDOW_MS = 10 * 60_000
81
82export const ZERO_TOTALS: WarmTotals = { refreshes: 0, costUsd: 0, wastedUsd: 0, kept: 0, keptUsd: 0 }
83
84// the one user message a refresh appends to the fork: it says what it is, so a
85// transcript reader (or the model) never takes it for the person
86export const FORK_PROMPT =
87 '[headroom] Automated prompt cache refresh by the headroom plugin, not a message from the user. Reply with the single word ok.'
88
89const WINDOW_LABELS: Record<string, string> = { five_hour: '5 Hour', seven_day: 'Weekly' }
90
91const isPriceKey = (id: string): boolean => Object.prototype.hasOwnProperty.call(PRICES, id)
92
93// a gateway's prefix, a [1m] context suffix and a dated id all name the same prices
94function rates(model: string): Price | undefined {
95 const id = model.replace(/^(?:claude-gateway|anthropic)\//, '').replace(/\[\d+(?:k|m)\]$/, '')
96 const key = isPriceKey(id) ? id : id.match(/^(claude-[a-z]+-\d{1,2}(?:-\d{1,2})?)-\d{8}$/)?.[1]
97 return key !== undefined && isPriceKey(key) ? PRICES[key] : undefined
98}
99
100/** a fork that made a request; the other kind found nothing to fork */
101export type ForkReply = Exclude<ModelForkResult, { reason: 'nothing-to-fork' }>
102
103/** one refresh as it settled */
104export type Refresh = {
105 at: number
106 model: string
107 usage: CacheUsage | null
108 costUsd: number | null
109 /** the rewrite the next prompt avoids, less this refresh: earned only if a prompt follows in time */
110 savesUsd: number | null
111 result: 'warmed' | 'expired' | 'failed'
112 detail?: string
113}
114
115export const isTtl = (value: unknown): value is Ttl => value === '5m' || value === '1h'
116export const isTtlChoice = (value: unknown): value is TtlChoice => value === 'auto' || isTtl(value)
117
118// the band's lifetime button goes round: auto, 5m, 1h, auto
119export const nextChoice = (choice: TtlChoice): TtlChoice => (choice === 'auto' ? '5m' : choice === '5m' ? '1h' : 'auto')
120
121export const usageOf = (usage: ModelUsage): CacheUsage => ({
122 input: usage.input_tokens,
123 output: usage.output_tokens,
124 cacheRead: usage.cache_read_input_tokens,
125 cacheWrite: usage.cache_creation_input_tokens,
126})
127
128export const promptTokensOf = (usage: CacheUsage): number => usage.input + usage.cacheRead + usage.cacheWrite
129
130// a refresh at 90% of the lifetime, at least ten seconds before it ends
131export const delayOf = (ttl: Ttl): number => Math.floor(Math.min(TTL_MS[ttl] * 0.9, TTL_MS[ttl] - 10_000))
132
133// Pi stops warming a running turn 60 minutes after its last request; the 1-hour
134// lifetime stretches that to two lifetimes so that it warms at all
135export const horizonOf = (ttl: Ttl): number => Math.max(60 * 60_000, 2 * TTL_MS[ttl])
136
137// a /config value as an idle limit: a whole number from 0 to IDLE_LIMIT_MAX
138export function limitOf(value: ConfigValue | undefined): number {
139 const count = Math.round(Number(value))
140 return value !== undefined && Number.isFinite(count) ? Math.min(IDLE_LIMIT_MAX, Math.max(0, count)) : IDLE_LIMIT_DEFAULT
141}
142
143// a /config value as the warmUntil %: a whole number from 1 to 100
144export function warmUntilOf(value: ConfigValue | undefined): number {
145 const percent = Math.round(Number(value))
146 return value !== undefined && Number.isFinite(percent) ? Math.min(100, Math.max(1, percent)) : WARM_UNTIL_DEFAULT
147}
148
149// how long an idle cache stays warm: the limit's refreshes, then one lifetime
150export const idleSpanOf = (count: number, ttl: Ttl): number => count * delayOf(ttl) + TTL_MS[ttl]
151
152export const addTotals = <T extends WarmTotals>(base: T, delta: Partial<WarmTotals>): T => ({
153 ...base,
154 refreshes: base.refreshes + (delta.refreshes ?? 0),
155 costUsd: base.costUsd + (delta.costUsd ?? 0),
156 wastedUsd: base.wastedUsd + (delta.wastedUsd ?? 0),
157 kept: base.kept + (delta.kept ?? 0),
158 keptUsd: base.keptUsd + (delta.keptUsd ?? 0),
159})
160
161// a timer that fires this late would likely find the cache gone and pay a whole write
162export const deadlineOf = (lastAt: number, ttl: Ttl): number => lastAt + delayOf(ttl) + Math.floor((TTL_MS[ttl] - delayOf(ttl)) / 2)
163
164export function costOf(model: string, usage: CacheUsage, ttl: Ttl): number | null {
165 const p = rates(model)
166 if (!p) return null
167 return (usage.input * p.input + usage.cacheWrite * p.input * WRITE_MULTIPLIER[ttl] + usage.cacheRead * p.cacheRead + usage.output * p.output) / 1_000_000
168}
169
170// what losing the cache adds to the next prompt: writing the prefix again instead of reading it
171export function missCostOf(model: string, promptTokens: number, ttl: Ttl): number | null {
172 const p = rates(model)
173 if (!p) return null
174 return Math.max(0, (promptTokens * (p.input * WRITE_MULTIPLIER[ttl] - p.cacheRead)) / 1_000_000)
175}
176
177// reading the prefix from a warm cache: what the next message pays for it while warm
178export function readCostOf(model: string, promptTokens: number): number | null {
179 const p = rates(model)
180 return p ? (promptTokens * p.cacheRead) / 1_000_000 : null
181}
182
183export type Decision = {
184 warmUsd: number
185 missUsd: number
186 probability: number
187 expectedUsd: number
188 isWarm: boolean
189 /** why a refresh does not pay: the prompt is under the size where the expected saving reaches MIN_SAVINGS_USD */
190 reason?: string
191}
192
193function breakEvenOf(p: Price, ttl: Ttl, probability: number, outputTokens: number): number | null {
194 const savedPerToken = probability * (p.input * WRITE_MULTIPLIER[ttl] - p.cacheRead) - p.cacheRead
195 if (savedPerToken <= 0) return null
196 return Math.ceil((MIN_SAVINGS_USD * 1_000_000 + outputTokens * p.output) / savedPerToken)
197}
198
199// Pi's rule: chance of a prompt before expiry × the rewrite it avoids − the refresh ≥ $0.05.
200// The chance is 1 while a turn runs, IDLE_CONTINUATION while idle. null: no price
201export function decide(model: string, promptTokens: number, ttl: Ttl, phase: 'run' | 'idle', outputTokens: number): Decision | null {
202 const p = rates(model)
203 const missUsd = missCostOf(model, promptTokens, ttl)
204 if (!p || missUsd === null || promptTokens <= 0) return null
205 const warmUsd = (promptTokens * p.cacheRead + outputTokens * p.output) / 1_000_000
206 const probability = phase === 'idle' ? IDLE_CONTINUATION : 1
207 const expectedUsd = probability * missUsd - warmUsd
208 const isWarm = expectedUsd >= MIN_SAVINGS_USD
209 const breakEven = isWarm ? null : breakEvenOf(p, ttl, probability, outputTokens)
210 const where = `${model} at ${ttl} (${phase})`
211 const reason = isWarm
212 ? undefined
213 : breakEven === null
214 ? `a refresh never saves on ${where}`
215 : `${formatTokens(promptTokens)} tokens is below the ${formatTokens(breakEven)} break-even on ${where}`
216 return { warmUsd, missUsd, probability, expectedUsd, isWarm, reason }
217}
218
219// a refresh that read under half the prefix found the cache gone and wrote it again
220export function outcomeOf(reply: ForkReply, usage: CacheUsage, promptTokens: number): Pick<Refresh, 'result' | 'detail'> {
221 if (usage.cacheRead >= promptTokens / 2) return { result: 'warmed' }
222 if (reply.isAnswered) return { result: 'expired' }
223 return {
224 result: 'failed',
225 detail: reply.reason === 'api-error' ? `${reply.error} ${reply.status ?? ''}`.trim() : reply.reason,
226 }
227}
228
229// the transcript's record of a refresh: ☕ 1h · Cache warmed · read 48.2k · $0.01 · saves $0.09 vs rewrite
230export function noticeText(ttl: Ttl, entry: Refresh): string {
231 const cost = `${entry.usage ? `read ${formatTokens(entry.usage.cacheRead)} · ` : ''}${entry.costUsd === null ? 'cost unknown' : formatUsd(entry.costUsd)}`
232 if (entry.result === 'expired') return `☕ ${ttl} · Cache had expired · the refresh rewrote it · ${cost}`
233 if (entry.result === 'failed') return `☕ ${ttl} · Cache refresh failed · ${entry.detail ? `${entry.detail} · ${cost}` : cost}`
234 return `☕ ${ttl} · Cache warmed · ${cost} · saves ${entry.savesUsd === null ? 'unknown' : formatUsd(entry.savesUsd)} vs rewrite`
235}
236
237// the last idle refresh has been sent: when the cache it warmed goes
238export const idleStopReason = (limit: number, expiresAt: number): string =>
239 `all ${limit} idle refreshes used, cache expires at ${clockOf(expiresAt)}`
240
241export const idleStopNotice = (ttl: Ttl, limit: number, expiresAt: number): string =>
242 `☕ ${ttl} · Warming stopped · all ${limit} idle refreshes used · cache expires at ${clockOf(expiresAt)}`
243
244/** Claude Code's own choice of the main conversation's lifetime, and what it reads to make it */
245export type TtlInputs = {
246 /** FORCE_PROMPT_CACHING_5M set */
247 force5m: boolean
248 /** CLAUDE_CODE_PROMPT_CACHE_TTL */
249 envTtl?: string
250 /** the promptCacheTtl setting */
251 settingTtl?: string
252 /** ENABLE_PROMPT_CACHING_1H set */
253 enable1h: boolean
254 /** a Claude subscription: the snapshot has a 5 Hour or Weekly limit */
255 isSubscription: boolean
256 /** a limit at or past 100%: the subscription runs on overage */
257 isOverage: boolean
258}
259
260// The lifetime in force, in Claude Code 2.1.291's order (read from its binary): the
261// force switch, then the variable (which a chosen 5m or 1h sets), the setting, the 1h
262// switch, and otherwise 1h on a subscription within its limits, 5m on an API key or in
263// overage. `choice`: the lifetime this mod set in the variable, `auto` when it set none
264export function ttlOf(i: TtlInputs, choice: TtlChoice = 'auto'): Ttl {
265 if (i.force5m) return '5m'
266 if (choice !== 'auto') return choice
267 if (isTtl(i.envTtl)) return i.envTtl
268 if (isTtl(i.settingTtl)) return i.settingTtl
269 if (i.enable1h) return '1h'
270 return i.isSubscription && !i.isOverage ? '1h' : '5m'
271}
272
273// a switch variable as Claude Code reads one: set, and not a word for off
274export const isEnvOn = (value: string | undefined): boolean =>
275 value !== undefined && !['', '0', 'false', 'no', 'off'].includes(value.trim().toLowerCase())
276
277const WINDOW_KINDS = ['five_hour', 'seven_day']
278
279// a subscription has the 5 Hour or Weekly limit; an API key has none
280export function planOf(limits: readonly Limit[]): { isSubscription: boolean; isOverage: boolean } {
281 const windows = limits.filter(l => WINDOW_KINDS.includes(l.kind))
282 return { isSubscription: windows.length > 0, isOverage: windows.some(l => l.percentUsed >= 100) }
283}
284
285// the first limit at or past warmUntil: no refresh then, as the status line says
286export function pastLimitOf(limits: readonly Limit[], warmUntil: number): string | null {
287 for (const kind of WINDOW_KINDS) {
288 const l = limits.find(x => x.kind === kind)
289 if (l && Math.round(l.percentUsed) >= warmUntil) return `${WINDOW_LABELS[kind]} at ${Math.round(l.percentUsed)}%, past your ${warmUntil}%`
290 }
291 return null
292}
293
294const sameWindow = (a: string | undefined, b: string | undefined): boolean =>
295 a !== undefined && b !== undefined && Math.abs(Date.parse(a) - Date.parse(b)) <= SAME_WINDOW_MS
296
297// How far each window's meter moved between two turn ends, in one window (by its
298// reset): a window that reset meanwhile, or has no reset, says nothing. A window that
299// did not move is a jump of 0, so the dollars it took still count against the rate
300export function jumpsOf(before: readonly Limit[] | null, now: readonly Limit[]): Record<string, number> {
301 const jumps: Record<string, number> = {}
302 if (!before) return jumps
303 for (const kind of WINDOW_KINDS) {
304 const b = before.find(l => l.kind === kind)
305 const n = now.find(l => l.kind === kind)
306 if (b && n && sameWindow(b.resetsAt, n.resetsAt)) jumps[kind] = Math.max(0, n.percentUsed - b.percentUsed)
307 }
308 return jumps
309}
310
311// the turn's dollars go to every window that measured a jump
312export function addRate(rate: WarmRate, jumps: Record<string, number>, usd: number): WarmRate {
313 const next = { ...rate }
314 for (const [kind, jump] of Object.entries(jumps)) {
315 const r = next[kind] ?? { jump: 0, usd: 0 }
316 next[kind] = { jump: r.jump + jump, usd: r.usd + usd }
317 }
318 return next
319}
320
321// How much of a window a dollar moves: Σjump / Σusd. The meter has one decimal at best
322// (whole points from the replies), so a rate is trusted only once a whole point has moved
323export function percentPerUsd(r: { jump: number; usd: number } | undefined): number | null {
324 return r && r.jump >= 1 && r.usd > 0 ? r.jump / r.usd : null
325}
326
327/** the window a cost is shown in, and how much of it a dollar moves */
328export type CostView = { label: string; perUsd: number }
329
330// the 5 Hour window when it has a rate (the session's own limit), else Weekly; none on an API key
331export function viewOf(rate: WarmRate, limits: readonly Limit[]): CostView | null {
332 for (const kind of WINDOW_KINDS) {
333 const perUsd = percentPerUsd(rate[kind])
334 if (perUsd !== null && limits.some(l => l.kind === kind)) return { label: WINDOW_LABELS[kind]!, perUsd }
335 }
336 return null
337}
338
339export const formatPercent = (p: number): string => `${p >= 0.1 ? p.toFixed(1) : p > 0 ? p.toPrecision(1) : '0'}%`
340
341export function formatUsd(usd: number): string {
342 const sign = usd < 0 ? '-' : ''
343 const value = Math.abs(usd)
344 return `${sign}$${value > 0 && value < 0.01 ? value.toPrecision(2) : value.toFixed(2)}`
345}
346
347export const formatTokens = (n: number): string =>
348 n >= 1_000_000 ? `${(n / 1_000_000).toFixed(1)}M` : n >= 1000 ? `${(n / 1000).toFixed(1)}k` : String(n)
349
350export function formatDuration(ms: number): string {
351 const s = Math.max(0, Math.round(ms / 1000))
352 if (s >= 3600) return `${Math.floor(s / 3600)}h${Math.floor((s % 3600) / 60)}m`
353 if (s >= 60) return `${Math.floor(s / 60)}m${s % 60 ? `${s % 60}s` : ''}`
354 return `${s}s`
355}
356
357// a span in the band's own words, to the minute: 38m, 1h 2m; under a minute said so
358export function spanOf(ms: number): string {
359 const m = Math.floor(Math.max(0, ms) / 60_000)
360 if (m < 1) return 'under a minute'
361 const h = Math.floor(m / 60)
362 const min = m % 60
363 if (h > 0) return min > 0 ? `${h}h ${min}m` : `${h}h`
364 return `${min}m`
365}
366
367// a time as the band says one elsewhere: 2:32 PM
368const clockOf = (t: number): string => new Date(t).toLocaleTimeString('en-US', { hour: 'numeric', minute: '2-digit' })
369
370/** a line of the band: what it says, and whether it calls for attention */
371export type CacheLine = { text: string; tone: 'muted' | 'amber' }
372
373const plural = (n: number, one: string, many: string) => `${n} ${n === 1 ? one : many}`
374
375// what the refreshes cost and saved: "3 refreshes this session, $0.04, saved $0.31"
376function totalsText(totals: WarmTotals, span: string): string {
377 return `${plural(totals.refreshes, 'refresh', 'refreshes')} ${span}, ${formatUsd(totals.costUsd)}${totals.keptUsd > 0 ? `, saved ${formatUsd(totals.keptUsd)}` : ''}`
378}
379
380// the warmer's own line while it is on: scheduled, refreshing, or stopped and why; with
381// this session's totals, and all time's where earlier sessions add to them
382export function warmStatusText(status: WarmStatus, totals: WarmTotals, now: number, assumed: Ttl | null = null, allTime: WarmTotals | null = null): CacheLine | null {
383 if (status.state === 'refreshing') return { text: 'Refreshing the cache…', tone: 'amber' }
384 if (status.state === 'stopped') return { text: `Warming ${status.isPaused ? 'paused' : 'stopped'}: ${status.reason}`, tone: 'muted' }
385 if (status.state !== 'scheduled') return null
386 const parts = [`Cache warm`, `refresh in ${spanOf(status.nextAt - now)}`]
387 if (assumed) parts.push(`${assumed} assumed`)
388 if (totals.refreshes > 0) parts.push(totalsText(totals, 'this session'))
389 if (allTime && allTime.refreshes > totals.refreshes) parts.push(totalsText(allTime, 'all time'))
390 return { text: parts.join(' · '), tone: 'muted' }
391}
392
393/** what the warning needs: the prefix, its dollars, and the window to show them in */
394export type CostOfNext = {
395 promptTokens: number
396 /** the rewrite over the warm read (missCostOf); null: no price */
397 missUsd: number | null
398 /** the warm read alone */
399 readUsd: number | null
400 view: CostView | null
401}
402
403// a dollar figure in the window's % when a rate is known, else in dollars
404const amountOf = (usd: number, view: CostView | null): string => (view ? formatPercent(usd * view.perUsd) : formatUsd(usd))
405
406function rewriteOf(c: CostOfNext, isWarmShown: boolean): string {
407 const tokens = `${formatTokens(c.promptTokens)} tokens`
408 if (c.missUsd === null) return tokens
409 const about = `${tokens}, about ${amountOf(c.missUsd, c.view)}${c.view ? ` of ${c.view.label}` : ''}`
410 return isWarmShown && c.readUsd !== null ? `${about} (warm: ${amountOf(c.readUsd, c.view)})` : about
411}
412
413export const expiredText = (agoMs: number, c: CostOfNext): string =>
414 `Cache expired ${spanOf(agoMs)} ago: your next message rewrites ${rewriteOf(c, true)}`
415
416export const expiringText = (inMs: number, c: CostOfNext): string =>
417 `Cache expires in ${spanOf(inMs)}: the next message after that rewrites ${rewriteOf(c, false)}`
418
419// with warming off, what a refresh would have cost instead
420export function nudgeText(refreshUsd: number | null, view: CostView | null, isExpired: boolean): string {
421 if (refreshUsd === null) return ''
422 return ` · Keep cache warm ${isExpired ? 'would have kept' : 'would keep'} it for about ${amountOf(refreshUsd, view)}`
423}
424
425export type CacheLineInput = {
426 isOn: boolean
427 anchor: WarmAnchor | null
428 status: WarmStatus
429 totals: WarmTotals
430 /** every session's totals, this one's included */
431 allTime: WarmTotals | null
432 assumed: Ttl | null
433 rate: WarmRate
434 limits: readonly Limit[]
435 outputTokens: number
436 now: number
437}
438
439// The band's cache line, above the bars: the warmer's status while it is on; else, and
440// once a stopped warmer's cache has gone, what the next message costs (amber once
441// expired, muted in the minutes before). Nothing before the first reply.
442export function cacheLineOf(i: CacheLineInput): CacheLine | null {
443 const a = i.anchor
444 if (!a) return null
445 const expiresAt = a.lastAt + TTL_MS[a.ttl]
446 const isExpired = i.now >= expiresAt
447 if (i.isOn) {
448 const line = warmStatusText(i.status, i.totals, i.now, i.assumed, i.allTime)
449 if (line && !(i.status.state === 'stopped' && isExpired)) return line
450 }
451 const view = viewOf(i.rate, i.limits)
452 const cost: CostOfNext = {
453 promptTokens: a.promptTokens,
454 missUsd: missCostOf(a.model, a.promptTokens, a.ttl),
455 readUsd: readCostOf(a.model, a.promptTokens),
456 view,
457 }
458 const refreshUsd = decide(a.model, a.promptTokens, a.ttl, 'idle', i.outputTokens)?.warmUsd ?? null
459 if (isExpired) {
460 return { text: expiredText(i.now - expiresAt, cost) + (i.isOn ? '' : nudgeText(refreshUsd, view, true)), tone: 'amber' }
461 }
462 if (!i.isOn && expiresAt - i.now <= EXPIRING_MS) {
463 return { text: expiringText(expiresAt - i.now, cost) + nudgeText(refreshUsd, view, false), tone: 'muted' }
464 }
465 return null
466}
467types/index.d.ts 129 lines1export type Category = { name: string; tokens: number; kind: 'used' | 'free' | 'buffer' }
2export type Limit = { kind: string; percentUsed: number; resetsAt?: string }
3export type PaceOf = { typical: number; skip: number }
4export type Snapshot = {
5 window: number
6 tokens: number
7 percent: number
8 total: number
9 categories: Category[]
10 limits: Limit[]
11 /** when the limit figures were read */
12 limitsAt?: number
13 /** why the usage service last refused, while it is backed off */
14 limitsError?: string
15 /** per window kind: the usual final % and the one-off % left out of the pace */
16 pace?: Record<string, PaceOf>
17 /** the minute it was taken in: redraws countdowns at least once a minute */
18 minute?: number
19}
20
21/**
22 * auto compact: on or off, and the context % that sets it off (kept for every chat).
23 * Agent-timed: from `startAt` % the agent chooses the moment, up to `at` %, the cap
24 */
25export type AutoCompact = { isOn: boolean; at: number | null; isAgentTimed?: boolean; startAt?: number | null }
26
27/** what Agent-timed holds for the session; every compaction of the main conversation starts it over */
28export type AgentTimed = {
29 /** the chat the state belongs to: a reload keeps it, another chat starts it over */
30 chat: string | null
31 /** the agent's request to defer compaction below the cap */
32 hold: { reason: string; since: number; remindAt: number } | null
33 /** the handoff note: to the summarizer, back to the agent, then cleared */
34 note: string | null
35 /** the agent asked to compact when this turn ends */
36 isAsked: boolean
37 /** whether the agent knows it is past the start %: `next` until the coming turn begins */
38 told: 'no' | 'next' | 'yes'
39 /** the highest nudge sent this cycle, main tool calls since, and the breakpoint hint */
40 nudge: { level: number; calls: number; isBreakpointSaid: boolean }
41 /** a hold the cap ended, until the row after the compaction says so */
42 overridden: { reason: string; percent: number } | null
43}
44
45/** Keep cache warm: the prompt cache's lifetime, and the chat's choice of it (`auto`: Claude Code's own) */
46export type Ttl = '5m' | '1h'
47export type TtlChoice = 'auto' | Ttl
48/** a request's four token counts, in the policy's spelling */
49export type CacheUsage = { input: number; output: number; cacheRead: number; cacheWrite: number }
50export type WarmTotals = {
51 refreshes: number
52 costUsd: number
53 /** the fees of refresh chains no kept prompt followed */
54 wastedUsd: number
55 kept: number
56 /** the rewrites kept prompts avoided */
57 keptUsd: number
58}
59/** the main conversation's last request: the prefix a fork replays and keeps warm */
60export type WarmAnchor = {
61 at: number
62 lastAt: number
63 model: string
64 promptTokens: number
65 ttl: Ttl
66 refreshes: number
67 /** refreshes sent while no turn ran: the idle limit counts these */
68 idleRefreshes: number
69 /** what this chain's refreshes cost: wasted unless the next prompt is kept */
70 feeUsd: number
71 isStopped: boolean
72}
73export type WarmStatus =
74 | { state: 'waiting' }
75 | { state: 'scheduled'; nextAt: number; phase: 'run' | 'idle'; expectedUsd: number }
76 | { state: 'refreshing' }
77 /** `isPaused`: stopped by the warmUntil threshold, which the band words as a pause */
78 | { state: 'stopped'; reason: string; isPaused?: boolean }
79/** per window kind: the % the meter moved and the dollars spent meanwhile, summed */
80export type WarmRate = Record<string, { jump: number; usd: number }>
81/** Keep cache warm, the session's side: the chain, the lifetime in force, the totals, the learned rate */
82export type Warm = {
83 /** the chat the chain and the totals belong to: another chat starts them afresh */
84 chat: string | null
85 anchor: WarmAnchor | null
86 status: WarmStatus
87 isRunning: boolean
88 outputTokens: number
89 /** the first main response wrote the cache: its lifetime holds for the session */
90 isLocked: boolean
91 /** the lifetime in force */
92 ttl: Ttl
93 /** set when a refresh found a 1h cache gone: 5m from then, for the session */
94 assumed: Ttl | null
95 totals: WarmTotals
96 allTime: WarmTotals & { since: number }
97 rate: WarmRate
98 /** the limits at the last main turn's end: the jump is measured from them */
99 lastLimits: Limit[] | null
100}
101/** Keep cache warm, the chat's setting (kept for every chat): on or off, and the lifetime */
102export type WarmSetting = { isOn: boolean; ttl: TtlChoice }
103
104declare module 'claude-code' {
105 interface PluginState {
106 'headroom': {
107 snapshot: Snapshot | null
108 isOn: boolean
109 isCollapsed: boolean
110 autoCompact: AutoCompact
111 /** the choice the band is asking for: the context was already past a new % */
112 autoAsk: { at: number; percent: number } | null
113 /** bumped on every % set: draws the field afresh even when the value is unchanged */
114 fieldTick: number
115 fieldText: string | null
116 agentTimed: AgentTimed
117 /** the start % field's own tick and typed text, as fieldTick and fieldText are the cap's */
118 startTick: number
119 startText: string | null
120 /** the desktop app's theme, as its settings (or the OS) have it */
121 theme: 'dark' | 'light'
122 /** Keep cache warm: the chain, the lifetime, the totals and the learned rate */
123 warm: Warm
124 /** Keep cache warm, this chat's switch and lifetime */
125 warmSetting: WarmSetting
126 }
127 }
128}
129