SLOPSHOPPER

headroom

Your Claude plan limits, a forecast of when you'll run out, and your context window, live in a band above the prompt. Compacts in one click, automatically, or…

newbandguardcommandtoaststatus
v0.1.0MITupdated 2026-10-07MWEJ/claude-code-headroom
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · headroom
› fix the failing auth test and add an audit log call ╭────────────────────────────────────────────╮ │ headroom │ ● headroom: headroom: agent-timed row (appended): Agent-timed compactio│ Context at 49%: auto compacting │ ● headroom: headroom: the app's theme could not be read: ENOENT: /Users│ (Agent-timed from 30%) │ ⏺ Read(src/auth.ts) ╰────────────────────────────────────────────╯ ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /headroom ⎿ headroom: Headroom off. ● headroom: headroom: cache rate: nothing measured (no earlier reading of the same window) ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts
README

Headroom: a Claude Code mod for your limits, context and cache

Version: v0.1.0 Auto compact: 80% default Compact: one click 5 Hour + Weekly: whole account Forecast: before reset Works in: Desktop + Terminal License: MIT

Know how much room you have left, and make more of it. A live band above the Claude Code prompt that shows your plan limits, forecasts whether you'll run out before they reset, and shows how full your context window is. Compact in one click, let it compact automatically, or let Claude pick the moment. Keep the prompt cache warm so the first message after a break stays cheap.

Formerly Claude Code Usage Quota Mod.

The 5 Hour and Weekly limits are your whole Claude account's, including what you use in Claude chat and Cowork. The band itself shows in Claude Code: the desktop app and the terminal.

Type /headroom to turn the band off and on.

Expanded

The band, expanded, above the prompt in the Claude desktop app

Collapsed

One row that still shows the time left until each limit resets:

The band, collapsed, above the prompt in the Claude desktop app

[!TIP] Never hit a full context again: Auto compact, on from the start. Auto compact is on in every new session and compacts it once the context reaches 80% (set 15 to 99). Set it lower to compact sooner: a long context costs more on every reply, so compacting early keeps replies cheaper and your limits lasting longer. Auto compact never interrupts a reply. Want to compact right now? The Compact now button does it in one click, and ends any hold Claude has.

Claude picks the moment: Agent-timed. On by default too, Agent-timed starts compaction sooner (from 30%) but at a good moment: Claude can hold it through a debugging chain or a refactor, release it at a safe point, and leave itself a note that survives. Every 5 minutes of a hold, Claude is asked to keep, release or update it. Your Auto compact % stays the limit no hold can pass.

Pay less for every message: Keep cache warm. Also on from the start. A small refresh keeps the conversation's prompt cache alive through a break, and while Claude waits on a subagent, so the next message reads the cache at a tenth of the price instead of writing it all again. The Cache dropdown in the top row picks the lifetime, auto, 5m or 1h, and the band shows what the warmer cost and saved, this session and all time.

Each session keeps its own settings, through a restart and a /clear, so you can run them differently side by side:

  • Building something big? Leave Auto compact at 80% or turn it off, so Claude keeps the whole picture. Press Compact now yourself at a good stopping point. (With it off, Claude Code's built-in compaction still steps in when the context is nearly full.)
  • Everyday sessions? Set Auto compact lower (say 30%), or leave Agent-timed to it. They stay lean, and your 5 Hour and Weekly limits last longer.
  • A session that should send nothing extra? Switch Keep cache warm off there; every refresh counts against your plan like any request.

What it shows

5 Hour and Weekly limits

  • The % used, matching the app's own Plan usage limits panel.
  • A forecast of where you'll be at reset: a grey tick on the bar and a note, e.g. On pace for about 87% by reset – resets in 23m.
  • A headline at a glance: On track. You should reach Wednesday's reset with room to spare.
  • The forecast starts from your usual pace (learned from your past windows) and shifts to your actual pace as the window goes on, so an early burst doesn't set it off.
  • Bars turn amber at 75% and red at 90%, or sooner if you're on course to run out.

On course to run out, it warns you and says when you'll hit the limit:

The band, expanded, on course to run out before the 5 hour reset

The band, collapsed, on course to run out

Context window

  • The % used and space free, as /context counts it, split into Messages, Tools and Other.
  • Amber at 50%, red at 80%: the point to compact.

Compacting

  • Compact now compacts in one click. While Claude holds, it ends the hold too.
  • Auto compact runs at your % (80% by default, 15 to 99), never mid-reply. Each session keeps its own setting; a new session starts with it on, Agent-timed included.
  • Agent-timed (optional, beside Auto compact) lets Claude choose the moment below that %. See Agent-timed.

Prompt cache

  • Once the cache has expired, the band says what your next message costs to write it again, in your 5 Hour limit's %: Cache expired 4m ago: your next message rewrites 48.2k tokens, about 0.9% of 5 Hour (warm: 0.05%). In the two minutes before, it says it is about to. The % is learned from how far your limit moves per dollar of replies; until a whole point has moved (and on an API key) it says dollars.
  • Keep cache warm (on by default, set per session) refreshes the cache before it expires, so the first message after a break reads it instead. See Keep cache warm.

If the session is already past your % when you turn it on, it asks first:

Auto compact asking what to do, with the context already at 72%

The same question with the band collapsed

  • Now compacts straight away.
  • After my next compact waits until the session is compacted some other way (Compact, /compact, or Claude Code's own), then takes over.
  • Only in new chats leaves this session alone until it's next opened.

Works in light and dark themes and at every width down to the narrowest. Collapsed or expanded is shared across all sessions. Each session keeps its own Auto compact, Agent-timed and Keep cache warm settings, and a /clear keeps them: the cleared session goes on with the same ones, while Claude's hold and note start over with the context.

Early in a 5 hour window:

The band, expanded, early in a 5 hour window

The band, collapsed, early in a 5 hour window

Running out on the Weekly limit:

The band, expanded, on course to run out before the weekly reset

The band, collapsed, on course to run out before the weekly reset

Light themeNarrowest window
Light, expandedNarrow, expanded
Light, collapsedNarrow, collapsed

Agent-timed

Auto compact fires at a fixed %, whatever Claude is doing. A compaction that lands mid-debugging throws away the context that mattered. With Agent-timed on, Claude chooses the moment, between two numbers you set:

  • From the start % (30% by default, its own box in the band) the session is compacted when a turn ends, unless Claude holds it.
  • At your Auto compact % it is compacted when the turn ends, whatever Claude holds. No hold passes it.

What Claude can do, through a small compaction tool the mod gives it:

  • Hold compaction through fragile work, with a reason.
  • Release it at a safe point, or ask to compact when the turn ends.
  • Leave a handoff note. It goes to the summarizer and comes back to Claude after the compaction, once.

While Claude holds, the band says so, with the reason and for how long, and a Release button ends the hold from your side:

The band with Claude holding compaction: "Held by Claude 3m: implementing task 3 of 5" and a Release button, the context at 33%

What you see and keep:

  • While Claude holds, the band says so: Held by Claude 12m: mid-refactor of auth, with Release to overrule it, or Compact now to overrule it and compact.
  • Every 5 minutes of a hold, Claude is asked where it stands: keep the hold with its current reason, release it, or leave a note.
  • Claude is told when the context passes the start %, so a compaction never comes unannounced.
  • While subagents Claude is waiting on are still running, compaction waits for them, up to your Auto compact %.
  • Only the main agent can hold. Subagents cannot.

Agent-timed is on in a new session, from 30%, and each session keeps its own setting. Switching it on adds the tool to that session; switched off again, the tool stays listed until the session is reopened, and answers that the mode is off.

Keep cache warm

Every message sends the whole conversation. The API keeps it in a prompt cache for 5 minutes or an hour; a message that reads the cache pays about a tenth of the input price, and one after the cache has expired writes it all again at 1.25× (5 minutes) or 2× (1 hour). After a break, the first message pays.

With Keep cache warm on, the mod sends one small request shortly before the cache would expire: a copy of the conversation's last request with one line asking for the word ok. It re-reads the cache, which keeps it alive, and never enters the conversation; a ☕ row in the transcript records each one with what it cost and saves.

  • It sends only when it pays. A refresh goes out only when it is expected to save at least $0.05: always worth it while Claude is working, and at a 15% chance of your next message while you are away. A small conversation is below that, and is not warmed.
  • Two numbers bound it: at most 5 refreshes per lifetime while you are away (set 0 to 20, separately for 5m and 1h), and none while your 5 Hour or Weekly limit is at or past 85%. Both are /config rows (headroom.idle5m, headroom.idle1h, headroom.warmUntil).
  • The lifetime follows Claude Code. The Cache dropdown in the band's top row reads auto by default: 1 hour on a Claude subscription within its limits, 5 minutes on an API key or once in overage, as Claude Code chooses (and as FORCE_PROMPT_CACHING_5M, CLAUDE_CODE_PROMPT_CACHE_TTL, ENABLE_PROMPT_CACHING_1H or the promptCacheTtl setting say). Pick 5m or 1h to choose for this session, warming on or off. The first reply of a session writes the cache, so a choice made after it applies to new sessions. headroom.cacheTtl in /config sets what new chats start with.
  • The band says what it is doing and what it has cost and saved: Cache warm · refresh in 38m · 3 refreshes this session, $0.04, saved $0.31 · 42 refreshes all time, $0.60, saved $4.10, or why it stopped. Compacting, /clear, a model switch or a failed refresh start it afresh with your next message.

[!IMPORTANT] Each refresh counts against your plan like any other request: a cache read of the conversation and a few output tokens. It is set per session: switch it off in the band for a chat that should send none. The expired-cache warning shows what the refreshes save, in the same 5 Hour %, warming on or off.

With this on, disable cache-warmer if you have it installed: both would warm the same cache.

Where it works

  • Claude desktop app, Code tab
  • Claude Code in a terminal

[!NOTE] In the desktop app, the band appears once a session has started. The brand-new Welcome back screen has no session yet, so no mod can draw there. Send your first message and it appears.

In the terminal, ▲/▼ expands and collapses it. The [-] next to it is Claude Code's own control and hides the band; ctrl+x ctrl+a brings it back.

The band in the terminal, expanded

The band in the terminal, collapsed

Install

Needs a recent Claude Code, signed in with a Pro or Max plan.

Desktop app: in the Code tab, send this as a message, allow the claude plugin command if asked, then quit and reopen the app:

Install the headroom plugin from the GitHub marketplace MWEJ/claude-code-headroom

Terminal: inside claude, run:

/plugin install headroom --marketplace MWEJ/claude-code-headroom

Either way installs it for both the desktop app and the terminal.

Update: ask Claude, then restart:

Update the headroom plugin from its marketplace

Uninstall: in the desktop app, ask Claude:

Remove the claude-code-headroom plugin marketplace

or in the terminal:

/plugin marketplace remove claude-code-headroom

then restart. This removes it from both. Just want it out of sight? /headroom hides it without uninstalling.

claude plugin marketplace add MWEJ/claude-code-headroom; claude plugin install headroom@claude-code-headroom

To update:

claude plugin marketplace update claude-code-headroom; claude plugin update headroom@claude-code-headroom

To uninstall:

claude plugin marketplace remove claude-code-headroom

Privacy

Everything runs on your machine. It reads what Claude Code already has (your context and the limits each reply carries), and about every 2 minutes asks Anthropic's usage service for your limits through your existing Claude login. It never sees your credentials and sends nothing anywhere else. In the desktop app it reads the app's theme setting to match light or dark.

Feedback and contributing

First release: I'd love to hear how it works for you. Open an issue for bugs or ideas. Pull requests welcome.

To work on it, clone the repo, then load it from the folder, or run its tests:

claude --plugin-dir ./claude-code-headroom
claude plugin test ./claude-code-headroom

If you find it useful, a ⭐ helps others find it.

Credits

Headroom started as Claude Code Usage Quota Mod by Anant Raghunath (MIT): the limits, the forecast, the context window and Auto compact are his work.

Inspired by I'm liking the new mods feature on r/ClaudeCode. Thanks to u/itsxzy for sharing the original prompt that started this project.

Agent-timed is inspired by compactor by rhwendt (MIT), which lets the agent hold and release Claude Code's own auto-compaction.

Keep cache warm is ported from cache-warmer by Paul B. Kim (MIT), itself a port of the cache warmer in Pi by Mario Zechner, whose rule decides when a refresh pays.

License

MIT © 2026 Martin Hygge. Started from Claude Code Usage Quota Mod © 2026 Anant Raghunath, MIT.

Source 4 files
hooks/register.tsx 2329 lines
1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, ModelForkResult, ModelUsage, Register, SessionRateLimit, TurnUsage } from 'claude-code'
3
4import type { AgentTimed, AutoCompact, Category, Limit, PaceOf, Snapshot, Ttl, TtlChoice, Warm, WarmAnchor, WarmRate, WarmSetting, WarmTotals } from '../types'
5import {
6  AT_DEFAULT, DEFAULT_AUTO, EMPTY, REASON_SHOWN, START_DEFAULT, START_MIN, TOOL, TOOL_DESCRIPTION, TOOL_NAME, TOOL_SCHEMA,
7  afterText, answerTool, breakpointOf, breakpointText, capOf, decide, holdReminder, isTimed, nudgeLevel, nudgeText, startOf, stepNudge, toldText, withNote,
8} from './agent-policy'
9import type { Stuck, ToolInput } from './agent-policy'
10import {
11  DEFAULT_OUTPUT_TOKENS, FORK_PROMPT, IDLE_LIMIT_DEFAULT, TTL_MS, WARM_UNTIL_DEFAULT, ZERO_TOTALS,
12  addRate, addTotals, cacheLineOf, costOf, deadlineOf, decide as decideWarm, delayOf, formatDuration, formatTokens, formatUsd, horizonOf, idleStopNotice,
13  idleStopReason, isEnvOn, isTtlChoice, jumpsOf, limitOf, missCostOf, noticeText, outcomeOf, pastLimitOf, planOf, ttlOf,
14  usageOf, warmUntilOf,
15} from './cache-policy'
16import type { ForkReply, Refresh } from './cache-policy'
17
18const snapshot = atom({ plugin: 'headroom', key: 'snapshot' } as const, null)
19const isOn = atom({ plugin: 'headroom', key: 'isOn' } as const, true)
20const isCollapsed = atom({ plugin: 'headroom', key: 'isCollapsed' } as const, false)
21const autoCompact = atom({ plugin: 'headroom', key: 'autoCompact' } as const, DEFAULT_AUTO)
22// what Agent-timed holds for the session: the agent's hold, note and request, and what it has been told
23const agentTimed = atom({ plugin: 'headroom', key: 'agentTimed' } as const, EMPTY as AgentTimed)
24const fieldTick = atom({ plugin: 'headroom', key: 'fieldTick' } as const, 0)
25// the % field's text while typing is cleaned (digits only, 3 at most); null: the set %
26const fieldText = atom({ plugin: 'headroom', key: 'fieldText' } as const, null as string | null)
27// the start % field's own tick and typed text, as fieldTick and fieldText are the cap's
28const startTick = atom({ plugin: 'headroom', key: 'startTick' } as const, 0)
29const startText = atom({ plugin: 'headroom', key: 'startText' } as const, null as string | null)
30const theme = atom({ plugin: 'headroom', key: 'theme' } as const, 'dark' as 'dark' | 'light')
31const autoAsk = atom({ plugin: 'headroom', key: 'autoAsk' } as const, null as { at: number; percent: number } | null)
32// Keep cache warm: the session's chain, lifetime, totals and rate; and this chat's switch
33const WARM_EMPTY: Warm = {
34  chat: null,
35  anchor: null,
36  status: { state: 'waiting' },
37  isRunning: false,
38  outputTokens: DEFAULT_OUTPUT_TOKENS,
39  isLocked: false,
40  ttl: '5m',
41  assumed: null,
42  totals: ZERO_TOTALS,
43  allTime: { ...ZERO_TOTALS, since: 0 },
44  rate: {},
45  lastLimits: null,
46}
47const warm = atom({ plugin: 'headroom', key: 'warm' } as const, WARM_EMPTY)
48const warmSetting = atom({ plugin: 'headroom', key: 'warmSetting' } as const, { isOn: true, ttl: 'auto' } as WarmSetting)
49
50// two palettes, the desktop's dark and light themes; the band draws in the one the
51// app shows (Theme), set at the start of every draw so all colours below follow it
52type Palette = {
53  muted: string; green: string; greenText: string; amber: string; red: string; free: string; buffer: string; track: string
54  palette: string[]; named: Record<string, string>; switchOff: string; switchOn: string; switchLit: string; compactLit: string
55}
56const DARK: Palette = {
57  muted: '#8b90a0', green: '#4ade80', greenText: '#4ade80', amber: '#fbbf24', red: '#f87171',
58  free: '#2d3140', buffer: '#4b5163', track: '#2d3140',
59  palette: ['#a78bfa', '#60a5fa', '#f472b6', '#34d399', '#fbbf24', '#fb923c', '#22d3ee', '#c084fc'],
60  named: {
61    Tools: '#a78bfa', Other: '#8b90a0', 'System prompt': '#a78bfa', 'System tools': '#60a5fa', 'MCP tools': '#f472b6',
62    'Custom agents': '#34d399', 'Memory files': '#fbbf24', Skills: '#fb923c', Messages: '#22d3ee',
63  },
64  switchOff: '#4b5163', switchOn: '#4ade80', switchLit: '#8b919c', compactLit: '#5a606e',
65}
66// light: the same hues, deep enough to read on the pale band; the tracks pale
67const LIGHT: Palette = {
68  muted: '#5b6070', green: '#22b856', greenText: '#16a34a', amber: '#ea8a0c', red: '#ef4444',
69  free: '#ffffff', buffer: '#b9bec9', track: '#ffffff',
70  palette: ['#8b5cf6', '#3b82f6', '#ec4899', '#10b981', '#f59e0b', '#f97316', '#0ea5e9', '#a855f7'],
71  named: {
72    Tools: '#8b5cf6', Other: '#5b6070', 'System prompt': '#8b5cf6', 'System tools': '#3b82f6', 'MCP tools': '#ec4899',
73    'Custom agents': '#10b981', 'Memory files': '#f59e0b', Skills: '#f97316', Messages: '#0ea5e9',
74  },
75  switchOff: '#b4b9c4', switchOn: '#22c55e', switchLit: '#c3c7cf', compactLit: '#d6d9df',
76}
77let MUTED = DARK.muted
78let GREEN = DARK.green
79let GREEN_TEXT = DARK.greenText
80let AMBER = DARK.amber
81let RED = DARK.red
82let FREE = DARK.free
83let BUFFER = DARK.buffer
84let TRACK = DARK.track
85let PALETTE = DARK.palette
86let NAMED = DARK.named
87let SWITCH_OFF = DARK.switchOff
88let SWITCH_ON = DARK.switchOn
89// the light behind the switch under the pointer: drawn beneath the pill, never over it
90let SWITCH_LIT = DARK.switchLit
91// Compact's hover: clearly lighter than its resting chrome, its label still reads
92let COMPACT_LIT = DARK.compactLit
93function usePalette(p: Palette): void {
94  ;({ muted: MUTED, green: GREEN, greenText: GREEN_TEXT, amber: AMBER, red: RED, free: FREE, buffer: BUFFER, track: TRACK } = p)
95  ;({ palette: PALETTE, named: NAMED, switchOff: SWITCH_OFF, switchOn: SWITCH_ON, switchLit: SWITCH_LIT, compactLit: COMPACT_LIT } = p)
96}
97const CELLS = 16
98const HOUR = 3_600_000
99const WINDOWS: Record<string, { label: string; ms: number }> = {
100  five_hour: { label: '5 Hour', ms: 5 * HOUR },
101  seven_day: { label: 'Weekly', ms: 7 * 24 * HOUR },
102}
103
104// 5 hour on the left, Weekly on the right
105const order = (kind: string) => (kind === 'five_hour' ? 0 : 1)
106
107type Forecast = {
108  kind: string
109  label: string
110  percent: number
111  resetsAt: number
112  /** 'idle': no window running (unused, or reset): 0% until the next message starts one */
113  status: 'hit' | 'ok' | 'out' | 'idle'
114  projected: number
115  runOutAt: number
116  color: string
117}
118
119// The pace a window is forecast at: your usual one (the median of how far your past
120// windows got) blended with this window's own, which takes over as the window runs: it
121// counts for half a quarter of the way in. Early jumps barely move it. `skip`: the %
122// the first reply of a reopened chat took re-reading its old context, a one-off left
123// out of the pace (still in the % used).
124const TYPICAL_DEFAULT = 50
125const TRUST = 0.25
126const FINALS_KEPT = 8
127// one window's reset time as two readings give it: a few minutes apart at most
128const SAME_WINDOW_MS = 10 * 60_000
129
130function forecast(limit: Limit, now: number, pace?: PaceOf): Forecast | null {
131  const win = WINDOWS[limit.kind]
132  const resetsAt = limit.resetsAt ? Date.parse(limit.resetsAt) : NaN
133  if (!win) return null
134  const p = limit.percentUsed
135  // no reset time, or one already passed: the window is not running, so it is at 0
136  if (!Number.isFinite(resetsAt) || resetsAt <= now) {
137    return { kind: limit.kind, label: win.label, percent: 0, resetsAt: NaN, status: 'idle', projected: 0, runOutAt: Infinity, color: GREEN }
138  }
139  const remaining = Math.max(0, resetsAt - now)
140  const elapsed = Math.max(1, win.ms - remaining)
141  const base = { kind: limit.kind, label: win.label, percent: p, resetsAt }
142  if (p >= 100) return { ...base, status: 'hit', projected: p, runOutAt: now, color: RED }
143  const typical = pace?.typical ?? TYPICAL_DEFAULT
144  const own = Math.max(0, p - (pace?.skip ?? 0)) / elapsed
145  const weight = elapsed / (elapsed + win.ms * TRUST)
146  const rate = weight * own + (1 - weight) * (typical / win.ms)
147  const projected = p + rate * remaining
148  if (projected < 100 || rate <= 0) {
149    return { ...base, status: 'ok', projected, runOutAt: Infinity, color: worse(GREEN, usedColor(p)) }
150  }
151  return {
152    ...base,
153    status: 'out',
154    projected,
155    runOutAt: now + (100 - p) / rate,
156    color: worse(projected > 130 ? RED : AMBER, usedColor(p)),
157  }
158}
159
160// A limit's colour by what is used, as the band shows it (rounded): amber from 75%,
161// red from 90%. The bar and its % take the worse of this and the forecast's.
162const LIMIT_AMBER_AT = 75
163const LIMIT_RED_AT = 90
164function usedColor(p: number): string {
165  const shown = Math.round(p)
166  return shown >= LIMIT_RED_AT ? RED : shown >= LIMIT_AMBER_AT ? AMBER : GREEN
167}
168function worse(a: string, b: string): string {
169  const rank = (c: string) => (c === RED ? 2 : c === AMBER ? 1 : 0)
170  return rank(b) > rank(a) ? b : a
171}
172
173function duration(ms: number): string {
174  const m = Math.max(0, Math.round(ms / 60_000))
175  const d = Math.floor(m / 1440)
176  const h = Math.floor((m % 1440) / 60)
177  const min = m % 60
178  if (d > 0) return h > 0 ? `${d}d ${h}h` : `${d}d`
179  if (h > 0) return min > 0 ? `${h}h ${min}m` : `${h}h`
180  return `${min}m`
181}
182
183const time = (t: number) =>
184  new Date(t).toLocaleTimeString('en-US', { hour: 'numeric', minute: '2-digit' })
185const weekday = (t: number) => new Date(t).toLocaleDateString('en-US', { weekday: 'long' })
186const isSameDay = (a: number, b: number) =>
187  new Date(a).toLocaleDateString('en-US') === new Date(b).toLocaleDateString('en-US')
188
189// "the 8:40 PM reset" today, "Wednesday's reset" otherwise
190// the reset by its time when it is under a day away (an after-midnight one included),
191// else by its weekday: a time that far off would only lengthen the headline
192function resetName(t: number, now: number): string {
193  return t - now < 24 * 3_600_000 ? `the ${time(t)} reset` : `${weekday(t)}'s reset`
194}
195
196function headline(list: Forecast[], now: number): { text: string; color: string } | null {
197  const hit = list.find(f => f.status === 'hit')
198  if (hit) {
199    const at = isSameDay(hit.resetsAt, now) ? time(hit.resetsAt) : `${weekday(hit.resetsAt)} ${time(hit.resetsAt)}`
200    return { text: `Limit reached. Usage resumes at ${at}.`, color: RED }
201  }
202  const out = list.filter(f => f.status === 'out').sort((a, b) => a.runOutAt - b.runOutAt)[0]
203  if (out) {
204    return { text: `At this pace you'll run out before ${resetName(out.resetsAt, now)}.`, color: out.color }
205  }
206  const lead = list.find(f => f.kind === 'seven_day' && f.status !== 'idle') ?? list.find(f => f.status !== 'idle')
207  if (!lead) return null
208  return {
209    text: `On track. You should reach ${resetName(lead.resetsAt, now)} with room to spare.`,
210    color: GREEN,
211  }
212}
213
214function tokens(n: number): string {
215  // 1M, not 1.0M; 1.5M keeps its decimal
216  if (n >= 1_000_000) return `${(n / 1_000_000).toFixed(1).replace(/\.0$/, '')}M`
217  if (n >= 1000) return `${Math.round(n / 1000)}k`
218  return `${Math.round(n)}`
219}
220
221// Context by what it holds, as shown (rounded): amber from 50% (a long context costs
222// more on every reply: compacting pays), red from 80%
223const CONTEXT_AMBER_AT = 50
224const CONTEXT_RED_AT = 80
225function fillColor(percent: number): string {
226  const shown = Math.round(percent)
227  if (shown >= CONTEXT_RED_AT) return RED
228  if (shown >= CONTEXT_AMBER_AT) return AMBER
229  return GREEN
230}
231
232// text in green, which on the light band needs a deeper green than the bars to read
233function ink(color: string): string {
234  return color === GREEN ? GREEN_TEXT : color
235}
236
237function colorOf(category: Category, index: number): string {
238  if (category.kind === 'free') return FREE
239  if (category.kind === 'buffer') return BUFFER
240  return NAMED[category.name] ?? PALETTE[index % PALETTE.length]!
241}
242
243type Segment = { color: string; share: number }
244
245const BAR_PX = 7.5
246
247// an iOS-style switch: a pill, the knob right and green when on, left and grey when off
248// three cells wide, so the click area laid over it covers all of it
249const SWITCH_W = 30
250const SWITCH_H = 20
251const SWITCH_PAD = 3
252// the cells the switch's box takes: the press under it is clipped to this, so the
253// app's own button tint can't spill onto the text beside it
254const SWITCH_CELLS = 4
255// the box behind the switch measures a pixel off the drawing (19px tall, 31 wide):
256// the pill is drawn half a pixel down and right so it sits in the light's centre
257const SWITCH_NUDGE = 0.5
258// fully see-through: the unlit border, and the press's own hover tint, which drew a
259// square frame round the light (the word "transparent" draws white or nothing)
260const CLEAR = '#00000000'
261// a space a wrap never breaks at
262const NB = ' '
263
264function switchSvg(isOn: boolean): string {
265  // the pill sits inset by PAD: the light under the pointer shows around it
266  const w = SWITCH_W - 2 * SWITCH_PAD
267  const h = SWITCH_H - 2 * SWITCH_PAD
268  const p = SWITCH_PAD
269  const knob = p + (isOn ? w - h / 2 : h / 2)
270  return (
271    `<svg xmlns="http://www.w3.org/2000/svg" width="${SWITCH_W}" height="${SWITCH_H}" viewBox="${-SWITCH_NUDGE} ${-SWITCH_NUDGE} ${SWITCH_W} ${SWITCH_H}">` +
272    `<rect x="${p}" y="${p}" width="${w}" height="${h}" rx="${h / 2}" fill="${isOn ? SWITCH_ON : SWITCH_OFF}"/>` +
273    `<circle cx="${knob}" cy="${p + h / 2}" r="${h / 2 - 2}" fill="#ffffff"/></svg>`
274  )
275}
276
277// a thin rounded bar for surfaces that draw Svg: segments laid out over 0..1000,
278// with an optional grey tick, the bar's own height, at `marker` percent (the forecast at reset)
279function barSvg(segments: Segment[], width: number, marker?: number): string {
280  const height = BAR_PX
281  const y = 0
282  const total = segments.reduce((sum, s) => sum + s.share, 0) || 1
283  let x = 0
284  const rects = segments
285    .map(s => {
286      const w = (s.share / total) * 1000
287      const rect = `<rect x="${x.toFixed(2)}" y="${y}" width="${w.toFixed(2)}" height="${BAR_PX}" fill="${s.color}"/>`
288      x += w
289      return rect
290    })
291    .join('')
292  // a plain upright block 3px wide whatever the drawn width: no rounded corners
293  // (stretched sideways, those drew it as a "D"), inside the bar's rounded ends so
294  // nothing pokes out past them. Placed on a whole pixel and left smooth-edged: snapped
295  // edges on a scaled screen drew it 3 device pixels wide on one bar and 4 on another
296  const TICK_PX = 3
297  const tickW = (TICK_PX / width) * 1000
298  const tickAt = marker === undefined ? 0 : Math.min(width - TICK_PX, Math.max(0, Math.round((marker / 100) * width - TICK_PX / 2)))
299  const tick =
300    marker === undefined
301      ? ''
302      : `<rect x="${((tickAt / width) * 1000).toFixed(3)}" y="${y}" width="${tickW.toFixed(3)}" height="${BAR_PX}" fill="${MUTED}"/>`
303  return (
304    `<svg xmlns="http://www.w3.org/2000/svg" width="${width}" height="${height}" viewBox="0 0 1000 ${height}" preserveAspectRatio="none">` +
305    `<clipPath id="r"><rect y="${y}" width="1000" height="${BAR_PX}" rx="${BAR_PX / 2}" ry="${BAR_PX / 2}"/></clipPath>` +
306    `<g clip-path="url(#r)">${rects}${tick}</g></svg>`
307  )
308}
309
310// the forecast tick's percent, or none when there is nothing ahead to mark
311function markerOf(f: Forecast): number | undefined {
312  if (f.status === 'hit' || f.status === 'idle') return undefined
313  const at = Math.min(100, f.projected)
314  return at - f.percent >= 1 ? at : undefined
315}
316
317// the terminal's bar: one colour per cell, runs of a colour merged into one Box
318function cellRuns(f: Forecast, cells: number): { color: string; width: number }[] {
319  const filled = Math.max(0, Math.min(cells, Math.round((f.percent / 100) * cells)))
320  const marker = markerOf(f)
321  const tick = marker === undefined ? -1 : Math.max(filled, Math.min(cells - 1, Math.round((marker / 100) * cells) - 1))
322  const runs: { color: string; width: number }[] = []
323  for (let i = 0; i < cells; i++) {
324    const color = i < filled ? f.color : i === tick ? MUTED : TRACK
325    const last = runs[runs.length - 1]
326    if (last && last.color === color) last.width += 1
327    else runs.push({ color, width: 1 })
328  }
329  return runs
330}
331
332const USAGE_URL = 'https://api.anthropic.com/api/oauth/usage'
333// The usage service rate-limits per login, and the app's own Usage panel asks it on
334// the same allowance, so ask it seldom: when a chat opens, then every 2 minutes with the
335// exact context count, one ask for the whole app (when it was last asked, the answer and
336// any back-off sit in the shared store). Between asks the figures every reply from
337// Claude carries keep the band current.
338const PLAN_EVERY_MS = 2 * 60_000
339// timers drift: a 2-minute tick a moment early still counts as due
340const PLAN_SLACK_MS = 10_000
341// after a failed ask wait 2, then 5, then 15 minutes, or what the service says
342const PLAN_BACKOFF_MS = [120_000, 300_000, 900_000]
343const LOCAL_EVERY_MS = 3_000
344
345type Backoff = { until: number; failures: number; error: string }
346/** 'no': the 3s tick, never asks; 'due': asks if 2 minutes have passed app-wide; 'now': asks unless backed off */
347type Ask = 'no' | 'due' | 'now'
348
349class AskError extends Error {
350  constructor(
351    message: string,
352    readonly retryAfterMs?: number,
353  ) {
354    super(message)
355  }
356}
357
358// The plan's limits, as the app's "Plan usage limits" panel reads them: the account's
359// usage endpoint, through the session's own credential (the plugin never sees it).
360async function fetchPlan($: EngineInterface): Promise<Limit[] | null> {
361  const auth = await $.session.authorize()
362  // no Claude login (an API key, a gateway): there is no plan to ask about
363  if (!auth || auth.kind !== 'bearer') return null
364  const res = await $.http.fetch(USAGE_URL, {
365    auth: auth.handle,
366    headers: { 'anthropic-beta': 'oauth-2025-04-20', 'content-type': 'application/json' },
367  })
368  if (!res.ok) {
369    const header = Object.entries(res.headers ?? {}).find(([k]) => k.toLowerCase() === 'retry-after')?.[1]
370    const seconds = Number(Array.isArray(header) ? header[0] : header)
371    const retry = Number.isFinite(seconds) && seconds > 0 ? seconds * 1000 : undefined
372    throw new AskError(res.status === 429 ? 'refused (429)' : `failed (${res.status})`, retry)
373  }
374  const body = JSON.parse(res.text) as Record<string, { utilization?: number | null; resets_at?: string | null } | null>
375  const limits: Limit[] = []
376  for (const kind of ['five_hour', 'seven_day']) {
377    const w = body[kind]
378    if (!w || typeof w.utilization !== 'number') continue
379    limits.push({ kind, percentUsed: Math.round(w.utilization * 10) / 10, resetsAt: w.resets_at ?? undefined })
380  }
381  if (limits.length === 0) throw new AskError('answer had no limits')
382  return limits
383}
384
385type Saved = { at: number; limits: Limit[] }
386
387// this chat's own reading from its replies, and when it last changed
388let live: Saved | undefined
389
390async function storeGet<T>($: EngineInterface, key: string): Promise<T | undefined> {
391  try {
392    return (await $.store.get(key)) as T | undefined
393  } catch {
394    return undefined
395  }
396}
397
398async function storeSet($: EngineInterface, key: string, value: unknown): Promise<void> {
399  try {
400    await $.store.set(key, value)
401  } catch {
402    // this chat still has it
403  }
404}
405
406// Two sources, and the newest wins: the usage service (asked seldom, shared by every
407// chat through the store) and the figures each reply from Claude carries (this chat's,
408// stamped when they last changed). A window that has reset since is dropped.
409async function limitsOf(
410  $: EngineInterface,
411  rateLimits: SessionRateLimit[],
412  now: number,
413  ask: Ask,
414): Promise<{ limits: Limit[]; at?: number; error?: string }> {
415  if (rateLimits.length > 0) {
416    const limits = rateLimits.map(r => ({ kind: r.kind, percentUsed: r.percentUsed, resetsAt: r.resetsAt }))
417    if (!live || JSON.stringify(live.limits) !== JSON.stringify(limits)) live = { at: now, limits }
418  }
419  let saved = await storeGet<Saved>($, 'limits')
420  let backoff = await storeGet<Backoff>($, 'planBackoff')
421  const askedAt = (await storeGet<number>($, 'planAskedAt')) ?? 0
422  const isFree = now >= (backoff?.until ?? 0)
423  const isDue = ask === 'now' || (ask === 'due' && now - askedAt >= PLAN_EVERY_MS - PLAN_SLACK_MS)
424  if (isDue && isFree) {
425    // claim the ask first, so the other chats see it taken and skip theirs
426    await storeSet($, 'planAskedAt', now)
427    try {
428      const plan = await fetchPlan($)
429      if (plan) {
430        saved = { at: now, limits: plan }
431        await storeSet($, 'limits', saved)
432        await storeSet($, 'planLimits', saved)
433      }
434      if (backoff) await storeSet($, 'planBackoff', { until: 0, failures: 0, error: '' })
435      backoff = undefined
436    } catch (error) {
437      const failures = (backoff?.failures ?? 0) + 1
438      const wait =
439        error instanceof AskError && error.retryAfterMs !== undefined
440          ? error.retryAfterMs
441          : PLAN_BACKOFF_MS[Math.min(failures - 1, PLAN_BACKOFF_MS.length - 1)]!
442      backoff = { until: now + wait, failures, error: error instanceof Error ? error.message : String(error) }
443      await storeSet($, 'planBackoff', backoff)
444    }
445  }
446  const pick = [saved, live]
447    .filter((r): r is Saved => r !== undefined)
448    .sort((x, y) => y.at - x.at)[0]
449  // newer reply figures go to the store too, so other chats get them
450  if (pick && pick === live && (!saved || live.at > saved.at)) await storeSet($, 'limits', live)
451  // The reply figures run behind the usage service, which the app's panel reads: whole
452  // percents, and seen at 18 for an hour while the service said 19.0. Use only climbs
453  // within a window, so a newer reply figure below the service's last answer is stale:
454  // the higher of the two is kept. Above it, the reply's is news.
455  const plan = await storeGet<Saved>($, 'planLimits')
456  const finer = (l: Limit): Limit => {
457    const p = plan?.limits.find(x => x.kind === l.kind)
458    return p && sameWindow(p.resetsAt, l.resetsAt) && p.percentUsed > l.percentUsed ? { ...l, percentUsed: p.percentUsed } : l
459  }
460  // a window that has reset since stays, at 0: the forecast reads it as not running
461  const limits = (pick?.limits ?? []).map(l =>
462    l.resetsAt !== undefined && Date.parse(l.resetsAt) <= now ? { kind: l.kind, percentUsed: 0 } : finer(l),
463  )
464  return { limits, at: pick?.at, error: backoff?.error || undefined }
465}
466
467// What the forecast learns, shared by every chat through the store: how far each past
468// window got (its peak, the last few kept), and the running one's peak and skipped %.
469type PaceWindow = { resetsAt: number; peak: number; skip: number }
470type PaceStore = { finals: Record<string, number[]>; windows: Record<string, PaceWindow> }
471
472// A chat reopened with a long history re-reads it all on its first reply: one big jump
473// in usage. Its first turn notes the limits it started from; the first new reading
474// that moved is that jump, and is left out of the pace. Only for a chat that opened
475// with this much already in it (a new chat's first reply is real work).
476const RESUME_MIN_TOKENS = 20_000
477let isOpening = true
478let openMessages = 0
479let spikeFrom: { at: number; limits: Limit[] } | undefined
480
481function median(list: number[] | undefined): number | undefined {
482  if (!list || list.length === 0) return undefined
483  const s = [...list].sort((a, b) => a - b)
484  const mid = Math.floor(s.length / 2)
485  return s.length % 2 ? s[mid]! : (s[mid - 1]! + s[mid]!) / 2
486}
487
488const sameWindow = (a: string | undefined, b: string | undefined) =>
489  a !== undefined && b !== undefined && Math.abs(Date.parse(a) - Date.parse(b)) <= SAME_WINDOW_MS
490
491// the jump the reopened chat's first reply made, once a reading after it has moved
492function takeSpike(): Record<string, number> | undefined {
493  if (!spikeFrom || !live || live.at <= spikeFrom.at) return undefined
494  const jump: Record<string, number> = {}
495  for (const l of live.limits) {
496    const before = spikeFrom.limits.find(b => b.kind === l.kind)
497    if (before && sameWindow(before.resetsAt, l.resetsAt) && l.percentUsed > before.percentUsed) {
498      jump[l.kind] = l.percentUsed - before.percentUsed
499    }
500  }
501  if (Object.keys(jump).length === 0) return undefined
502  spikeFrom = undefined
503  return jump
504}
505
506async function trackPace($: EngineInterface, limits: Limit[], now: number): Promise<Record<string, PaceOf>> {
507  const spike = takeSpike()
508  const saved = (await storeGet<PaceStore>($, 'pace')) ?? { finals: {}, windows: {} }
509  let isChanged = false
510  const pace: Record<string, PaceOf> = {}
511  for (const l of limits) {
512    const resetsAt = l.resetsAt ? Date.parse(l.resetsAt) : NaN
513    if (!WINDOWS[l.kind] || !Number.isFinite(resetsAt) || resetsAt <= now) continue
514    let w = saved.windows[l.kind]
515    if (!w || Math.abs(w.resetsAt - resetsAt) > SAME_WINDOW_MS) {
516      // a new window: the last one's peak is how far it got
517      if (w && w.peak > 0) saved.finals[l.kind] = [...(saved.finals[l.kind] ?? []), w.peak].slice(-FINALS_KEPT)
518      w = { resetsAt, peak: l.percentUsed, skip: 0 }
519      saved.windows[l.kind] = w
520      isChanged = true
521    } else if (l.percentUsed > w.peak) {
522      w.peak = l.percentUsed
523      isChanged = true
524    }
525    const jump = spike?.[l.kind]
526    if (jump) {
527      w.skip = Math.min(w.peak, w.skip + jump)
528      isChanged = true
529    }
530    pace[l.kind] = { typical: median(saved.finals[l.kind]) ?? TYPICAL_DEFAULT, skip: w.skip }
531  }
532  if (isChanged) await storeSet($, 'pace', saved)
533  return pace
534}
535
536// The window is shown as three groups: Messages, Tools (System + MCP tools) and Other
537// (the system prompt, skills, memory files, agents, MCP server instructions).
538const TOOL_NAMES = new Set(['System tools', 'MCP tools'])
539type Group = 'Messages' | 'Tools' | 'Other'
540const groupOf = (name: string): Group => (name === 'Messages' ? 'Messages' : TOOL_NAMES.has(name) ? 'Tools' : 'Other')
541
542// The free local breakdown ('summary') estimates each category from its text, which
543// reads tools ~1.3-1.6x heavy. The exact count ('full', what /context and the app's
544// Context window panel show) costs no tokens but sends one token-count request per tool
545// and memory file (~320 here). So run it every 2 minutes, keep the exact/estimate ratio
546// of Tools and Other (in the store, for every session), and scale the live 3s estimate
547// by it; Messages is then the API's real input total less those two.
548const CALIBRATE_EVERY_MS = 2 * 60_000
549let calib: Partial<Record<Group, number>> = {}
550let isCalibrating = false
551// Right after /compact there is no API reply yet, so the window's real total is unknown
552// and the local estimate is all there is. The exact count fills that gap until the next reply.
553let exactCount: { at: number; sums: Record<Group, number> } | undefined
554let noReplySince: number | undefined
555
556function sumGroups(categories: readonly { name: string; tokens: number; kind: string }[]): Record<Group, number> {
557  const sums: Record<Group, number> = { Messages: 0, Tools: 0, Other: 0 }
558  for (const c of categories) if (c.kind === 'used') sums[groupOf(c.name)] += c.tokens
559  return sums
560}
561
562async function calibrate($: EngineInterface): Promise<void> {
563  if (isCalibrating) return
564  isCalibrating = true
565  try {
566    const exact = await $.session.usage({ breakdown: 'full' })
567    const rough = await $.session.usage({ breakdown: 'summary' })
568    const real = sumGroups(exact.context.breakdown?.categories ?? [])
569    const guess = sumGroups(rough.context.breakdown?.categories ?? [])
570    const next = { ...calib }
571    for (const g of ['Tools', 'Other'] as const) if (guess[g] > 0 && real[g] > 0) next[g] = real[g] / guess[g]
572    calib = next
573    exactCount = { at: await $.clock.now(), sums: real }
574    await $.store.set('calib', next)
575    lastSnapshot = ''
576  } catch (error) {
577    $.ui.log(`headroom: exact count failed: ${error instanceof Error ? error.message : String(error)}`, { to: 'debug' })
578  } finally {
579    isCalibrating = false
580  }
581  // the usage service is asked alongside, on the same 2 minutes
582  await refresh($, 'due')
583}
584
585// Auto compact acts whenever the chat is idle (no turn running): at the end of a turn,
586// on opening a chat, on any refresh between turns. What it does then is `decide`'s
587// (agent-policy.ts): at the cap (the %) it compacts; with Agent-timed on it also
588// compacts from the start %, unless the agent holds or its subagents still run. Paused
589// by `waitsForCompact` (the band is asking, or the person chose "after my next
590// compact", which is also what an unanswered ask means: cleared only by a real
591// compaction, never by the % flickering below) and by `isNewChatsOnly` ("only in new
592// chats"). `stuck`: a compaction already ran and the context is still past a %, so it
593// would only repeat; dropping below clears it, and stuck at the start never blocks the
594// cap. The % is compared as the band shows it, rounded.
595let waitsForCompact = false
596// a reply has been seen since the chat opened or was last compacted
597let hadReply = false
598let stuck: Stuck = 'no'
599let isNewChatsOnly = false
600let isBusy = false
601let isCompacting = false
602let lastPercent = 0
603// one look at a time: a look waits on the engine (the agents, a row told), and the 3s
604// tick must not start a second meanwhile
605let isWatching = false
606
607// Agent-timed Auto compact, the part that touches the engine: what the agent holds for
608// the session, its tool, what it is told, and what a compaction does to all of it. The
609// decisions and the words are agent-policy.ts's; this stays in the hooks module's own
610// file because the engine follows $ only into functions declared here.
611
612const reasonOf = (error: unknown) => (error instanceof Error ? error.message : String(error))
613
614// A row for the agent between turns: a user-role row the person does not see as typed.
615// The debug log has every one, appended or not (a test cannot see a plugin's rows).
616async function tell($: EngineInterface, text: string): Promise<void> {
617  let outcome = 'appended'
618  try {
619    const row = await $.session.append({ message: { type: 'user', content: [{ type: 'text', text }] } })
620    if (row.deny !== undefined) outcome = `not appended: ${row.deny}`
621  } catch (error) {
622    outcome = `not appended: ${reasonOf(error)}`
623  }
624  $.ui.log(`headroom: agent-timed row (${outcome}): ${text}`, { to: 'debug' })
625}
626
627// The main agent is waiting on work whose results it must still take in. An agent list
628// that cannot be read counts as none running: compaction is never held on a guess.
629async function hasRunningAgents($: EngineInterface): Promise<boolean> {
630  try {
631    return (await $.agent.list()).some(a => a.status === 'pending' || a.status === 'running' || a.status === 'waiting')
632  } catch {
633    return false
634  }
635}
636
637// the agent's tool, registered once, and only in a chat where Agent-timed is on (a tool
638// in the list rides on every request, and the API has no way to take one out again)
639let isToolRegistered = false
640async function ensureTool($: EngineInterface): Promise<boolean> {
641  if (isToolRegistered) return true
642  try {
643    await $.tool.register({ name: TOOL_NAME, description: TOOL_DESCRIPTION, inputSchema: TOOL_SCHEMA })
644    isToolRegistered = true
645  } catch (error) {
646    $.ui.log(`headroom: the compaction tool could not be registered: ${reasonOf(error)}`, { to: 'debug' })
647  }
648  return isToolRegistered
649}
650
651// Three places notice a compaction of the main conversation (auto compact's own direct
652// call, the session.compact hook, the reply total going blank): the first one handles
653// it, and the others find it handled until a reply has been seen again.
654let isCompactionHandled = false
655
656// The main conversation was compacted: the cycle starts over, and the agent gets its
657// note back, and word of a hold the cap ended, once.
658async function afterCompaction($: EngineInterface): Promise<void> {
659  isCompactionHandled = true
660  // the cache the chain kept is the old conversation's: the next response starts one afresh
661  await forget($, 'conversation compacted')
662  const state = await read($, agentTimed)
663  await update($, agentTimed, s => ({ ...EMPTY, chat: s.chat }))
664  const text = afterText(state.note, state.overridden, capOf(await read($, autoCompact)))
665  if (text !== null) await tell($, text)
666}
667
668// the request to compact is spent by one attempt, and dropped when the turn it was
669// made in is interrupted
670async function dropAsked($: EngineInterface): Promise<void> {
671  if ((await read($, agentTimed)).isAsked) await update($, agentTimed, s => ({ ...s, isAsked: false }))
672}
673
674// What rides on a main-agent tool result: word that the start % is passed (once a
675// cycle), the reminders of a hold growing old, the five-minute ask for where a hold
676// stands, and a breakpoint after a commit or a passing test run (`command`: a Bash
677// command that succeeded, else null).
678async function linesFor($: EngineInterface, command: string | null): Promise<string[]> {
679  const auto = await read($, autoCompact)
680  if (!isTimed(auto)) return []
681  const now = await $.clock.now()
682  const startAt = startOf(auto)
683  const cap = capOf(auto)
684  const percent = lastPercent
685  const shown = Math.round(percent)
686  const state = await read($, agentTimed)
687  const lines: string[] = []
688  let told = state.told
689  let nudge = state.nudge
690  if (told !== 'yes' && shown >= startAt) {
691    lines.push(toldText(percent, startAt, cap))
692    told = 'yes'
693  }
694  if (state.hold) {
695    const step = stepNudge(nudge, nudgeLevel(percent, startAt, cap, true))
696    nudge = step.nudge
697    if (step.isSaid) lines.push(nudgeText(nudge.level === 3 ? 3 : 2, percent, startAt, cap, state.hold.reason))
698    const kind = command === null || shown < startAt || nudge.isBreakpointSaid ? null : breakpointOf(command)
699    if (kind) {
700      lines.push(breakpointText(kind))
701      nudge = { ...nudge, isBreakpointSaid: true }
702    }
703  }
704  let hold = state.hold
705  const reminder = holdReminder(state, now, percent, startAt, cap)
706  if (reminder) {
707    lines.push(reminder.text)
708    hold = reminder.state.hold
709  }
710  if (told !== state.told || nudge !== state.nudge || hold !== state.hold) await update($, agentTimed, s => ({ ...s, told, nudge, hold }))
711  return lines
712}
713
714function compactNow($: EngineInterface): void {
715  if (isCompacting) return
716  isCompacting = true
717  // compacting takes longer than a press or a hook may run, so a timer starts it
718  $.clock.after(0, async () => {
719    try {
720      await $.command.run({ command: 'compact' })
721    } finally {
722      isCompacting = false
723    }
724  })
725}
726
727// Auto compact compacts through the session itself, not by typing /compact: a
728// command run from the end of a turn is refused while the turn winds down, which
729// is why it toasted and then did nothing. The compaction also rejects while a turn
730// runs, so it is tried again a few times, a couple of seconds apart.
731const AUTO_TRIES = 6
732let canCompactDirectly = true
733const AUTO_RETRY_MS = 2_000
734
735function autoCompactNow($: EngineInterface, attempt = 1): void {
736  if (attempt === 1) {
737    if (isCompacting) return
738    isCompacting = true
739  }
740  $.clock.after(attempt === 1 ? 1_000 : AUTO_RETRY_MS, async () => {
741    try {
742      // the agent's handoff note rides along as the summarizer's instructions (a typed
743      // /compact gets it from the session.compact hook instead)
744      const instructions = withNote(undefined, (await read($, agentTimed)).note)
745      // the desktop app (an SDK session) has no direct compaction: there /compact runs
746      // as a turn of its own, so it is typed, as the Compact button does
747      const result = canCompactDirectly
748        ? await $.session.compact(instructions === undefined ? undefined : { instructions })
749        : (await $.command.run({ command: 'compact' }), undefined)
750      // held as stuck while the figures are read again, so no tick fires a second one
751      stuck = 'cap'
752      isCompacting = false
753      // the attempt spends the agent's request, skipped or not: a skip must not loop
754      await dropAsked($)
755      if (result && 'skip' in result && result.skip) $.ui.toast(`Auto compact was skipped: ${result.skip}`)
756      else if (result) await afterCompaction($)
757      lastSnapshot = ''
758      const usage = await $.session.usage({ breakdown: 'summary' })
759      const shown = Math.round(usage.context.percent ?? 0)
760      // still past a % after compacting: say so once, and do not loop on it
761      const auto = await read($, autoCompact)
762      stuck = shown >= capOf(auto) ? 'cap' : isTimed(auto) && shown >= startOf(auto) ? 'start' : 'no'
763      if (stuck === 'cap') $.ui.toast(`Context is still at ${shown}% after compacting, past your ${capOf(auto)}%: auto compact waits until it drops below`)
764      if (stuck === 'start') $.ui.toast(`Context is still at ${shown}% after compacting, past your ${startOf(auto)}% start: Agent-timed waits until it drops below`)
765      await refresh($)
766    } catch (error) {
767      const reason = error instanceof Error ? error.message : String(error)
768      // any refusal of the direct way: type /compact from then on, as the button does
769      if (canCompactDirectly) {
770        canCompactDirectly = false
771        return autoCompactNow($, attempt + 1)
772      }
773      if (attempt < AUTO_TRIES) return autoCompactNow($, attempt + 1)
774      isCompacting = false
775      $.ui.toast(`Auto compact could not start: ${reason}`)
776      $.ui.log(`headroom: auto compact failed: ${reason}`, { to: 'debug' })
777    }
778  })
779}
780
781// called with the context % whenever it is read
782async function watchAuto($: EngineInterface, percent: number): Promise<void> {
783  lastPercent = percent
784  if (isWatching) return
785  isWatching = true
786  try {
787    const auto = await read($, autoCompact)
788    if (!auto.isOn || auto.at === null) return
789    const shown = Math.round(percent)
790    const cap = auto.at
791    const startAt = startOf(auto)
792    const timed = isTimed(auto)
793    // stuck eases as the context drops below what it was stuck past
794    if (stuck === 'cap' && shown < cap) stuck = timed && shown >= startAt ? 'start' : 'no'
795    if (stuck === 'start' && (!timed || shown < startAt)) stuck = 'no'
796    if (isBusy || isCompacting) return
797    const state = await read($, agentTimed)
798    // the agents are asked after only where they can decide: in the zone, nothing else in the way
799    const isOpenZone = timed && shown >= startAt && shown < cap && !state.hold && !state.isAsked
800    const verdict = decide({
801      percent, cap, startAt, isAgentTimed: timed, isPaused: waitsForCompact || isNewChatsOnly, stuck, state,
802      hasRunningAgents: isOpenZone && (await hasRunningAgents($)),
803    })
804    // a turn may have started while the engine was asked
805    if (isBusy || isCompacting) return
806    if (verdict.action === 'tell') {
807      await update($, agentTimed, s => ({ ...s, told: 'next' as const }))
808      await tell($, toldText(percent, startAt, cap))
809      return
810    }
811    if (verdict.action !== 'compact') return
812    if (verdict.why === 'cap' && state.hold) {
813      // no hold survives the cap: it ends here, and the row after the compaction says so
814      const { reason } = state.hold
815      await update($, agentTimed, s => ({ ...s, hold: null, overridden: { reason, percent: shown } }))
816      $.ui.toast(`Context at ${shown}%: Claude's hold ends at your ${cap}%, auto compacting`)
817    } else if (verdict.why === 'cap') $.ui.toast(`Context at ${shown}%: auto compacting (set at ${cap}%)`)
818    else if (verdict.why === 'start') $.ui.toast(`Context at ${shown}%: auto compacting (Agent-timed from ${startAt}%)`)
819    else $.ui.toast('Compacting as Claude asked')
820    autoCompactNow($)
821  } finally {
822    isWatching = false
823  }
824}
825
826async function saveAuto($: EngineInterface, next: AutoCompact): Promise<void> {
827  await update($, autoCompact, () => next)
828  await storeSet($, `autoCompact:${autoChat ?? (await chatId($))}`, next)
829}
830
831// auto compact is set per chat: a chat that never had it opens with the default (on at
832// 80%, Agent-timed from 30%); a chat reopened, or the app restarted, gets back its own. Read again
833// whenever the chat's id changes (a /clear goes on under a new one, unannounced)
834let autoChat: string | undefined
835// a /clear keeps the settings: the cleared chat's Auto compact and Keep cache warm go
836// to the id that follows it, saved under it as if set there, and the hold, the note and
837// the cache chain start over with the context. Nothing else takes a new id unannounced
838let carried: { auto: AutoCompact; warm: WarmSetting } | null = null
839async function carryOnClear($: EngineInterface): Promise<void> {
840  carried = { auto: await read($, autoCompact), warm: await read($, warmSetting) }
841}
842async function chatId($: EngineInterface): Promise<string> {
843  try {
844    return await $.session.id()
845  } catch {
846    return 'chat'
847  }
848}
849async function loadAuto($: EngineInterface): Promise<void> {
850  const id = await chatId($)
851  if (id === autoChat) return
852  autoChat = id
853  const carry = carried
854  carried = null
855  const saved = (await storeGet<AutoCompact>($, `autoCompact:${id}`)) ?? carry?.auto
856  isNewChatsOnly = false
857  stuck = 'no'
858  waitsForCompact = false
859  const auto: AutoCompact = saved ? { ...saved, at: saved.at ?? AT_DEFAULT } : DEFAULT_AUTO
860  await update($, autoCompact, () => auto)
861  if (carry) {
862    await storeSet($, `autoCompact:${id}`, auto)
863    await storeSet($, `warm:${id}`, carry.warm)
864  }
865  await update($, autoAsk, () => null)
866  // another chat: what the agent held, noted or asked for belonged to the last one. A
867  // reload of the module is not another chat: the state keeps the hold and the note
868  if ((await read($, agentTimed)).chat !== id) await update($, agentTimed, () => ({ ...EMPTY, chat: id }))
869  if (isTimed(auto)) await ensureTool($)
870  await loadWarm($, id)
871}
872
873// Keep cache warm, the part that touches the engine: the chain of refreshes (ported
874// from cache-warmer's register.tsx, MIT, see cache-policy.ts), the lifetime, the totals
875// and the rate the band learns. The prices, the rule and the words are
876// cache-policy.ts's; this stays here because the engine follows $ only into functions
877// declared in this file.
878
879// set from the /config rows when the module loads, and by their config.set hooks
880let idleLimits: Record<Ttl, number> = { '5m': IDLE_LIMIT_DEFAULT, '1h': IDLE_LIMIT_DEFAULT }
881let warmUntil = WARM_UNTIL_DEFAULT
882let defaultChoice: TtlChoice = 'auto'
883let warmTimer: { cancel(): void } | undefined
884// The anchor a refresh has claimed until chain() records it: a schedule() meanwhile,
885// from a turn that ended, cannot fork it again
886let forking: { at: number } | undefined
887// main prompts started this process: a warning from before the latest one is stale
888let prompts = 0
889// the lifetime this mod put in the variable (null: none), and what was there before it
890let envSet: Ttl | null = null
891let envBefore: string | undefined
892// the live window's last response as refresh() last read it, and when: what anchors stand on
893let lastApi: ModelUsage | null = null
894let lastApiAt = -Infinity
895// the response the anchor stands on, and when it was seen: the same counts read again
896// are no new response
897let anchoredApi = ''
898let anchoredAt = -Infinity
899// when the running (or last) main turn started
900let turnStartedAt = -Infinity
901// what refreshes spent since the last main turn ended: the meter's next jump holds it too
902let spentSince = 0
903
904// the variables Claude Code picks the lifetime by, each named outright (the engine lists
905// what a module reads); one that cannot be read counts as unset
906async function lifetimeVars($: EngineInterface): Promise<{ force5m?: string; ttl?: string; enable1h?: string }> {
907  const vars: { force5m?: string; ttl?: string; enable1h?: string } = {}
908  try {
909    vars.force5m = await $.env.get('FORCE_PROMPT_CACHING_5M')
910    vars.ttl = await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL')
911    vars.enable1h = await $.env.get('ENABLE_PROMPT_CACHING_1H')
912  } catch {
913    // the environment cannot be read: Claude Code's own default stands
914  }
915  return vars
916}
917
918// the promptCacheTtl setting; no row listed reads as unset
919async function settingTtlOf($: EngineInterface): Promise<string | undefined> {
920  try {
921    const row = (await $.config.list()).find(r => r.key === 'promptCacheTtl')
922    return row === undefined ? undefined : String(row.value)
923  } catch {
924    return undefined
925  }
926}
927
928// The lifetime in force, as Claude Code decides it (cache-policy's ttlOf), unless a
929// refresh found a 1h cache gone: then 5m for the rest of the session
930async function lifetimeOf($: EngineInterface, limits: readonly Limit[]): Promise<Ttl> {
931  const { assumed } = await read($, warm)
932  if (assumed) return assumed
933  const vars = await lifetimeVars($)
934  return ttlOf(
935    { force5m: isEnvOn(vars.force5m), envTtl: vars.ttl, settingTtl: await settingTtlOf($), enable1h: isEnvOn(vars.enable1h), ...planOf(limits) },
936    envSet ?? 'auto',
937  )
938}
939
940// A chosen 5m or 1h goes in the variable, for the process, from the next request,
941// warming on or off; auto puts back what was there. Once the first response wrote the cache
942// its lifetime holds for the session, as cache-warmer has it: a change waits for a new
943// one, and a press says so (`isPressed`)
944async function applyTtl($: EngineInterface, isPressed: boolean): Promise<void> {
945  const setting = await read($, warmSetting)
946  const want = setting.ttl !== 'auto' ? setting.ttl : null
947  if (want === envSet) return
948  const state = await read($, warm)
949  if (state.isLocked) {
950    if (isPressed) $.ui.toast(`Cache lifetime ${setting.ttl} applies to new sessions: this one's cache is written at ${state.ttl}`)
951    return
952  }
953  try {
954    if (envSet === null) envBefore = (await lifetimeVars($)).ttl
955    await $.env.set('CLAUDE_CODE_PROMPT_CACHE_TTL', want ?? envBefore)
956    envSet = want
957  } catch (error) {
958    $.ui.log(`headroom: the cache lifetime could not be set: ${reasonOf(error)}`, { to: 'debug' })
959  }
960}
961
962// a notice row: the transcript keeps it, no request carries it; the debug log has every
963// one, appended or not (a test cannot see a plugin's rows)
964async function cacheRow($: EngineInterface, text: string): Promise<void> {
965  let outcome = 'appended'
966  try {
967    const row = await $.session.append({ message: { type: 'system', content: [{ type: 'text', text }] } })
968    if (row.deny !== undefined) outcome = `not appended: ${row.deny}`
969  } catch (error) {
970    outcome = `not appended: ${reasonOf(error)}`
971  }
972  $.ui.log(`headroom: cache row (${outcome}): ${text}`, { to: 'debug' })
973}
974
975// this session's totals and the all-time ones in the store take the same delta
976async function addToTotals($: EngineInterface, delta: Partial<WarmTotals>): Promise<void> {
977  await update($, warm, s => ({ ...s, totals: addTotals(s.totals, delta) }))
978  const stored = (await storeGet<Warm['allTime']>($, 'warmAllTime')) ?? { ...ZERO_TOTALS, since: await $.clock.now() }
979  const allTime = addTotals(stored, delta)
980  await storeSet($, 'warmAllTime', allTime)
981  await update($, warm, s => ({ ...s, allTime }))
982}
983
984async function reportStop($: EngineInterface, reason: string, isPaused = false): Promise<void> {
985  await update($, warm, s => ({ ...s, status: isPaused ? { state: 'stopped' as const, reason, isPaused } : { state: 'stopped' as const, reason } }))
986  $.ui.log(`headroom: cache warming stopped: ${reason}`, { to: 'debug' })
987}
988
989// Compaction, /clear, a session's end and a model switch forget the chain: only the
990// next response starts it again, and its fee is wasted now, since no prompt will judge it
991async function forget($: EngineInterface, reason: string): Promise<void> {
992  warmTimer?.cancel()
993  warmTimer = undefined
994  const { anchor } = await read($, warm)
995  if (anchor && anchor.feeUsd > 0) await addToTotals($, { wastedUsd: anchor.feeUsd })
996  await update($, warm, s => ({ ...s, anchor: null }))
997  await reportStop($, reason)
998}
999
1000// Stops the chain anchored at `at`, adding `feeUsd` to it, until the next response; a
1001// response that replaced it meanwhile keeps its own warming. Answers whether it stopped it
1002async function stopChain($: EngineInterface, at: number, reason: string, feeUsd = 0, isPaused = false): Promise<boolean> {
1003  let isStopped = false
1004  await update($, warm, s => {
1005    const { anchor } = s
1006    isStopped = anchor?.at === at && !anchor.isStopped
1007    if (!anchor || !isStopped) return s
1008    return { ...s, anchor: { ...anchor, feeUsd: anchor.feeUsd + feeUsd, isStopped: true } }
1009  })
1010  if (isStopped) await reportStop($, reason, isPaused)
1011  return isStopped
1012}
1013
1014// a limit at or past warmUntil: no refresh spends more of it
1015async function pastLimit($: EngineInterface): Promise<string | null> {
1016  return pastLimitOf((await read($, snapshot))?.limits ?? [], warmUntil)
1017}
1018
1019// The next refresh, at 90% of the lifetime after the last request or refresh, if warming
1020// is on and one pays. A running turn stops at the run horizon, an idle session after its
1021// lifetime's idle limit. The anchor is read last, so the timer is set from it as it stands
1022async function schedule($: EngineInterface): Promise<void> {
1023  const prompt = prompts
1024  if (!(await read($, warmSetting)).isOn) return
1025  const past = await pastLimit($)
1026  const now = await $.clock.now()
1027  const { anchor: current, isRunning, outputTokens } = await read($, warm)
1028  if (!current || current.isStopped || current.at === forking?.at) return
1029  const phase: 'run' | 'idle' = isRunning ? 'run' : 'idle'
1030  const nextAt = current.lastAt + delayOf(current.ttl)
1031  const horizon = horizonOf(current.ttl)
1032  if (phase === 'run' && nextAt > current.at + horizon) {
1033    await stopChain($, current.at, `${formatDuration(horizon)} run limit reached`)
1034    return
1035  }
1036  if (past) {
1037    await stopChain($, current.at, past, 0, true)
1038    return
1039  }
1040  const limit = idleLimits[current.ttl]
1041  if (phase === 'idle' && current.idleRefreshes >= limit) {
1042    const expiresAt = current.lastAt + TTL_MS[current.ttl]
1043    // a limit of 0 turned idle warming off on purpose: it needs no warning
1044    const isStopped = await stopChain($, current.at, limit > 0 ? idleStopReason(limit, expiresAt) : 'no idle refreshes (set to 0)')
1045    if (isStopped && limit > 0 && prompt === prompts) await cacheRow($, idleStopNotice(current.ttl, limit, expiresAt))
1046    return
1047  }
1048  const decision = decideWarm(current.model, current.promptTokens, current.ttl, phase, outputTokens)
1049  if (!decision) {
1050    await stopChain($, current.at, `no price for ${current.model}`)
1051    return
1052  }
1053  if (decision.reason) {
1054    await stopChain($, current.at, decision.reason)
1055    return
1056  }
1057  warmTimer?.cancel()
1058  warmTimer = $.clock.after(Math.max(0, nextAt - now), () => void refreshCache($))
1059  await update($, warm, s => ({ ...s, status: { state: 'scheduled' as const, nextAt, phase, expectedUsd: decision.expectedUsd } }))
1060  $.ui.log(`headroom: cache refresh in ${formatDuration(nextAt - now)} (${current.ttl}, ${phase}, expected saving ${formatUsd(decision.expectedUsd)})`, { to: 'debug' })
1061}
1062
1063// the timer's refresh: it claims its anchor until chain() records it there
1064async function refreshCache($: EngineInterface): Promise<void> {
1065  warmTimer = undefined
1066  if (!(await read($, warmSetting)).isOn) return
1067  const { anchor: current, isRunning, outputTokens } = await read($, warm)
1068  if (!current || current.isStopped || current.at === forking?.at) return
1069  const claim = { at: current.at }
1070  forking = claim
1071  try {
1072    await forkFor($, current, isRunning ? 'run' : 'idle', outputTokens)
1073  } finally {
1074    if (forking === claim) forking = undefined
1075  }
1076}
1077
1078async function forkFor($: EngineInterface, current: WarmAnchor, phase: 'run' | 'idle', outputTokens: number): Promise<void> {
1079  const at = await $.clock.now()
1080  if (at > deadlineOf(current.lastAt, current.ttl)) {
1081    await stopChain($, current.at, 'refresh deadline missed')
1082    return
1083  }
1084  const past = await pastLimit($)
1085  if (past) {
1086    await stopChain($, current.at, past, 0, true)
1087    return
1088  }
1089  const decision = decideWarm(current.model, current.promptTokens, current.ttl, phase, outputTokens)
1090  if (!decision || decision.reason) {
1091    await stopChain($, current.at, decision?.reason ?? `no price for ${current.model}`)
1092    return
1093  }
1094  await update($, warm, s => ({ ...s, status: { state: 'refreshing' as const } }))
1095  let reply: ModelForkResult
1096  try {
1097    reply = await $.model.fork({ prompt: FORK_PROMPT })
1098  } catch (error) {
1099    await stopChain($, current.at, `refresh failed (${reasonOf(error)})`)
1100    return
1101  }
1102  if (!reply.isAnswered && reply.reason === 'nothing-to-fork') {
1103    await stopChain($, current.at, 'nothing to refresh')
1104    return
1105  }
1106  await settle($, current, at, phase, decision.missUsd, reply)
1107}
1108
1109async function settle($: EngineInterface, current: WarmAnchor, at: number, phase: 'run' | 'idle', missUsd: number, reply: ForkReply): Promise<void> {
1110  const usage = usageOf(reply.usage)
1111  // the fork's own write is its short tail: priced at the cache's lifetime it is overstated at most
1112  const costUsd = costOf(current.model, usage, current.ttl)
1113  const outcome = outcomeOf(reply, usage, current.promptTokens)
1114  const savesUsd = outcome.result === 'warmed' && costUsd !== null ? missUsd - costUsd : null
1115  const entry: Refresh = { at, model: current.model, usage, costUsd, savesUsd, ...outcome }
1116  await addToTotals($, { refreshes: 1, costUsd: costUsd ?? 0 })
1117  spentSince += costUsd ?? 0
1118  await cacheRow($, noticeText(current.ttl, entry))
1119  await chain($, current, entry, phase)
1120}
1121
1122// moves a warm refresh's chain on, when the chain is still current; answers whether it was
1123async function extend($: EngineInterface, current: WarmAnchor, entry: Refresh, phase: 'run' | 'idle'): Promise<boolean> {
1124  let isChained = false
1125  await update($, warm, s => {
1126    const { anchor } = s
1127    isChained = anchor?.at === current.at && !anchor.isStopped
1128    if (!anchor || !isChained) return s
1129    return {
1130      ...s,
1131      outputTokens: entry.usage?.output || DEFAULT_OUTPUT_TOKENS,
1132      anchor: {
1133        ...anchor,
1134        lastAt: entry.at,
1135        refreshes: anchor.refreshes + 1,
1136        idleRefreshes: anchor.idleRefreshes + (phase === 'idle' ? 1 : 0),
1137        feeUsd: anchor.feeUsd + (entry.costUsd ?? 0),
1138      },
1139    }
1140  })
1141  return isChained
1142}
1143
1144// The chain carries each refresh's fee. One a response overtook during its fork kept
1145// nothing that response read, so its fee is wasted; so is one whose chain was forgotten.
1146// A 1h cache a refresh found gone means the lifetime was 5m: assumed so from then
1147async function chain($: EngineInterface, current: WarmAnchor, entry: Refresh, phase: 'run' | 'idle'): Promise<void> {
1148  const costUsd = entry.costUsd ?? 0
1149  const isWarmed = entry.result === 'warmed'
1150  const isAssumed = entry.result === 'expired' && current.ttl === '1h'
1151  const reason =
1152    entry.result === 'expired'
1153      ? isAssumed ? 'the cache had expired: 5m assumed for this session' : 'the cache had expired'
1154      : `refresh failed${entry.detail ? ` (${entry.detail})` : ''}`
1155  if (isAssumed) {
1156    await update($, warm, s => ({
1157      ...s,
1158      assumed: '5m' as const,
1159      ttl: '5m' as const,
1160      anchor: s.anchor && s.anchor.at === current.at ? { ...s.anchor, ttl: '5m' as const } : s.anchor,
1161    }))
1162  }
1163  const isChained = isWarmed ? await extend($, current, entry, phase) : await stopChain($, current.at, reason, costUsd)
1164  if (forking?.at === current.at) forking = undefined
1165  if (!isChained) {
1166    if (costUsd > 0) await addToTotals($, { wastedUsd: costUsd })
1167    if ((await read($, warm)).status.state === 'refreshing') await update($, warm, s => ({ ...s, status: { state: 'waiting' as const } }))
1168    await schedule($)
1169    return
1170  }
1171  if (isWarmed) await schedule($)
1172}
1173
1174// A prompt is kept when it read a cache that would have expired without the refreshes
1175// since the last one; it avoided rewriting what it read. Otherwise the chain's fee is wasted
1176async function judgeChain($: EngineInterface, previous: WarmAnchor | null, at: number, cacheRead: number): Promise<void> {
1177  if (!previous) return
1178  const isKept = previous.refreshes > 0 && at - previous.at > TTL_MS[previous.ttl] && cacheRead >= previous.promptTokens / 2
1179  const keptUsd = isKept ? missCostOf(previous.model, cacheRead, previous.ttl) : null
1180  if (keptUsd !== null) await addToTotals($, { kept: 1, keptUsd })
1181  else if (previous.feeUsd > 0) await addToTotals($, { wastedUsd: previous.feeUsd })
1182}
1183
1184// A response of the main conversation was seen (the live window's counts moved): the
1185// chain before it is judged, and a new one starts from it. `model`: the turn's at its
1186// end; mid-turn the anchor's own (a fork's response is not the live window's). A turn
1187// that ended with usage had a response even if its counts read the same as the last
1188async function anchorOn($: EngineInterface, api: ModelUsage, at: number, model: string | undefined, isNew = false): Promise<boolean> {
1189  const key = JSON.stringify(api)
1190  const promptTokens = api.input_tokens + api.cache_read_input_tokens + api.cache_creation_input_tokens
1191  if ((key === anchoredApi && !isNew) || promptTokens <= 0) return false
1192  const previous = (await read($, warm)).anchor
1193  const named = model ?? previous?.model
1194  if (named === undefined) return false
1195  anchoredApi = key
1196  anchoredAt = at
1197  await judgeChain($, previous, at, api.cache_read_input_tokens)
1198  const { ttl } = await read($, warm)
1199  await update($, warm, s => ({
1200    ...s,
hooks/agent-policy.ts 268 lines
1import type { AgentTimed, AutoCompact } from '../types'
2
3// Agent-timed Auto compact, the part with no engine in it: when to compact, what the
4// agent's tool answers, and every word the agent reads. The idea (the agent holds
5// compaction through fragile work and releases it at a safe point, leaving a note that
6// survives) is compactor's, github.com/rhwendt/compactor (MIT), as is the shape of the
7// breakpoint patterns.
8
9export const TOOL_NAME = 'compaction'
10export const TOOL = 'mcp__headroom__compaction'
11export const TOOL_DESCRIPTION =
12  'Controls when this conversation is compacted. Compaction runs when your turn ends once context passes the start %. ' +
13  '"hold" (with a reason) before fragile multi-step work whose state lives only in this conversation; "release" at a safe point; ' +
14  '"compact" to compact when this turn ends; "note" to save what must survive (current hypothesis, next steps, file:line references). ' +
15  'At the cap % compaction runs whatever is held.'
16export const TOOL_SCHEMA = {
17  type: 'object',
18  properties: {
19    action: { enum: ['hold', 'release', 'compact', 'note', 'status'] },
20    reason: { type: 'string' },
21    note: { type: 'string' },
22  },
23  required: ['action'],
24}
25
26// the cap (the auto compact %) whenever there is none: a blank field, a first switch-on
27export const AT_DEFAULT = 80
28// the start %: 10 at the least, always below the cap; 30 where none was set
29export const START_MIN = 10
30export const START_DEFAULT = 30
31// a chat that never had a setting: Auto compact on at the cap, Agent-timed on from the start %
32export const DEFAULT_AUTO: AutoCompact = { isOn: true, at: AT_DEFAULT, isAgentTimed: true, startAt: START_DEFAULT }
33export const NOTE_MAX = 4_000
34// a reason is kept to this, and drawn in the band to less
35export const REASON_MAX = 500
36export const REASON_SHOWN = 80
37// "compact" under this context % is refused: there is nothing worth compacting
38export const ASK_MIN = 10
39// within this many points of the cap, a holding agent is told the cap is close
40export const NEAR_CAP = 5
41// main tool calls between repeats of that last warning
42export const NUDGE_EVERY = 10
43// while a hold lasts, the agent is asked this often to say where the hold stands
44export const HOLD_REMIND_MS = 5 * 60_000
45
46export const EMPTY: AgentTimed = {
47  chat: null,
48  hold: null,
49  note: null,
50  isAsked: false,
51  told: 'no',
52  nudge: { level: 0, calls: 0, isBreakpointSaid: false },
53  overridden: null,
54}
55
56// a chat's setting as it stands: the cap, whether Agent-timed is on (it needs auto
57// compact on), and the start %: the saved one or the default, never under the least,
58// always below the cap
59export const capOf = (auto: AutoCompact): number => auto.at ?? AT_DEFAULT
60export const isTimed = (auto: AutoCompact): boolean => auto.isOn && auto.isAgentTimed === true
61export const startOf = (auto: AutoCompact): number => Math.min(Math.max(START_MIN, auto.startAt ?? START_DEFAULT), capOf(auto) - 1)
62
63/** a compaction ran and the context is still past a %: `start` never blocks the cap */
64export type Stuck = 'no' | 'start' | 'cap'
65
66export type DecideInput = {
67  percent: number
68  cap: number
69  startAt: number
70  isAgentTimed: boolean
71  /** the band is asking, "After my next compact", or "Only in new chats" */
72  isPaused: boolean
73  stuck: Stuck
74  state: Pick<AgentTimed, 'hold' | 'isAsked' | 'told'>
75  hasRunningAgents: boolean
76}
77
78export type Decision =
79  | { action: 'compact'; why: 'asked' | 'cap' | 'start' }
80  | { action: 'tell' }
81  | { action: 'wait'; why: 'paused' | 'stuck' | 'below' | 'held' | 'agents' | 'told' }
82
83// What auto compact does with an idle chat: the first rule that applies decides. The %
84// is compared as the band shows it, rounded.
85export function decide(i: DecideInput): Decision {
86  const shown = Math.round(i.percent)
87  // the agent's own request is as deliberate as the Compact button: no pause stops it
88  if (i.state.isAsked) return { action: 'compact', why: 'asked' }
89  if (i.isPaused) return { action: 'wait', why: 'paused' }
90  // the cap: no hold, no subagent and no untold agent survives it
91  if (shown >= i.cap) return i.stuck === 'cap' ? { action: 'wait', why: 'stuck' } : { action: 'compact', why: 'cap' }
92  if (!i.isAgentTimed || shown < i.startAt) return { action: 'wait', why: 'below' }
93  if (i.stuck !== 'no') return { action: 'wait', why: 'stuck' }
94  if (i.state.hold) return { action: 'wait', why: 'held' }
95  // the main agent is waiting on work whose results it must still take in
96  if (i.hasRunningAgents) return { action: 'wait', why: 'agents' }
97  // never compacted unawares: told first, and given the coming turn to hold
98  if (i.state.told === 'no') return { action: 'tell' }
99  if (i.state.told === 'next') return { action: 'wait', why: 'told' }
100  return { action: 'compact', why: 'start' }
101}
102
103// How firmly a holding agent is reminded: 2 from halfway between the start % and the
104// cap, 3 from NEAR_CAP points under the cap.
105export function nudgeLevel(percent: number, startAt: number, cap: number, isHolding: boolean): 0 | 2 | 3 {
106  const shown = Math.round(percent)
107  if (!isHolding || shown < startAt) return 0
108  if (shown >= cap - NEAR_CAP) return 3
109  if (shown >= startAt + (cap - startAt) / 2) return 2
110  return 0
111}
112
113// One main tool call later: a level is said when first reached, and level 3 again
114// every NUDGE_EVERY calls while it lasts.
115export function stepNudge(nudge: AgentTimed['nudge'], level: 0 | 2 | 3): { isSaid: boolean; nudge: AgentTimed['nudge'] } {
116  if (level === 0) return { isSaid: false, nudge }
117  if (level > nudge.level) return { isSaid: true, nudge: { ...nudge, level, calls: 0 } }
118  if (level < 3) return { isSaid: false, nudge }
119  const calls = nudge.calls + 1
120  return calls >= NUDGE_EVERY ? { isSaid: true, nudge: { ...nudge, calls: 0 } } : { isSaid: false, nudge: { ...nudge, calls } }
121}
122
123// a command word: at the start, or after a space or a shell separator
124const SEP = String.raw`(?:^|[\s;&|(])`
125const END = String.raw`(?=$|[\s;&|)])`
126const COMMIT = new RegExp(SEP + String.raw`git(?:\s+-[Cc]\s+\S+)*\s+commit` + END)
127const TESTS = new RegExp(
128  SEP +
129    '(?:' +
130    String.raw`pytest|py\.test|python3?\s+-m\s+(?:pytest|unittest)` +
131    String.raw`|(?:npm|pnpm|yarn|bun)\s+(?:run\s+)?test(?::[\w-]+)?` +
132    String.raw`|go\s+test|cargo\s+test|make\s+(?:test|check)|mvn(?:\s+\S+)*?\s+test` +
133    String.raw`|(?:\./)?gradlew?(?:\s+\S+)*?\s+test|rspec|(?:npx\s+)?(?:jest|vitest)|claude\s+plugin\s+test` +
134    ')' +
135    END,
136)
137const SEGMENTS = /&&|\|\||[;|&\n]/
138
139// Which natural breakpoint a Bash command that succeeded was: read one shell segment
140// at a time, so "echo git commit" and a --dry-run are not one.
141export function breakpointOf(command: string): 'commit' | 'tests' | null {
142  const segments = command.split(SEGMENTS).filter(s => s.trim() !== '')
143  if (segments.some(s => COMMIT.test(s) && !s.includes('--dry-run'))) return 'commit'
144  if (segments.some(s => TESTS.test(s))) return 'tests'
145  return null
146}
147
148export function figures(percent: number, startAt: number, cap: number): string {
149  return `Context ${Math.round(percent)}%. Agent-timed compaction starts at ${startAt}%; at ${cap}% it runs whatever is held.`
150}
151
152export function toldText(percent: number, startAt: number, cap: number): string {
153  return (
154    `Agent-timed compaction: context is at ${Math.round(percent)}% (starts at ${startAt}%, cap ${cap}%). ` +
155    'This conversation will be compacted when your turn ends. ' +
156    'If you are mid-task, call the compaction tool with action "hold" and a reason. Otherwise save a "note" of what must survive.'
157  )
158}
159
160export function nudgeText(level: 2 | 3, percent: number, startAt: number, cap: number, reason: string): string {
161  const body =
162    level === 3
163      ? `the cap is close. At ${cap}% compaction runs when your turn ends, whatever is held. Save a note now.`
164      : `you are holding (${reason}) well past the start. Finish the current step, save a note, and release.`
165  return `Agent-timed compaction: ${body}\n${figures(percent, startAt, cap)}`
166}
167
168// Every HOLD_REMIND_MS of a hold, on the next main tool result: the agent is asked to
169// say where the hold stands (kept, released, or a note). `null` while none is due; the
170// state handed back is due again HOLD_REMIND_MS on.
171export function holdReminder(state: AgentTimed, now: number, percent: number, startAt: number, cap: number): { text: string; state: AgentTimed } | null {
172  if (!state.hold || now < state.hold.remindAt) return null
173  const held = Math.max(1, Math.round((now - state.hold.since) / 60_000))
174  const text =
175    `Agent-timed compaction: you have held compaction for ${held}m (${state.hold.reason}). ` +
176    'Update the hold: call the compaction tool with action "hold" and the current reason to keep it, "release" if the fragile step is done, or "note" what must survive.\n' +
177    figures(percent, startAt, cap)
178  return { text, state: { ...state, hold: { ...state.hold, remindAt: now + HOLD_REMIND_MS } } }
179}
180
181export function breakpointText(kind: 'commit' | 'tests'): string {
182  const what = kind === 'commit' ? 'a commit just landed' : 'tests just passed'
183  return `Agent-timed compaction: ${what}, a natural breakpoint. Consider "release" or "compact", with a note.`
184}
185
186const NOTE_HEAD = 'The agent left this handoff note. Keep what it says matters:'
187
188// The note as the summarizer's instructions, after whatever was asked for already.
189// Adding it twice adds it once: a compaction may pass more than one place that adds it.
190export function withNote(instructions: string | undefined, note: string | null): string | undefined {
191  if (note === null) return instructions
192  const block = `${NOTE_HEAD}\n${note}`
193  if (instructions?.includes(block)) return instructions
194  return instructions ? `${instructions}\n\n${block}` : block
195}
196
197// The row the agent reads after a compaction, or null with nothing to say.
198export function afterText(note: string | null, overridden: AgentTimed['overridden'], cap: number): string | null {
199  if (note === null && overridden === null) return null
200  const lines = ['Agent-timed compaction: the conversation was just compacted.']
201  if (overridden) lines.push(`Your hold (${overridden.reason}) ended at the ${cap}% cap. Hold again if the work is still fragile.`)
202  if (note !== null) lines.push(`Handoff note you left:\n${note}`)
203  return lines.join('\n')
204}
205
206export type ToolInput = { action?: unknown; reason?: unknown; note?: unknown }
207export type ToolContext = {
208  /** Agent-timed is on in this chat */
209  isOn: boolean
210  /** the call came from a subagent's or a teammate's loop */
211  isSubagent: boolean
212  percent: number
213  startAt: number
214  cap: number
215  now: number
216}
217
218const ACTIONS = ['hold', 'release', 'compact', 'note', 'status']
219
220function statusText(state: AgentTimed, now: number): string {
221  const lines = [
222    state.hold ? `hold: ${state.hold.reason} (${Math.max(0, Math.round((now - state.hold.since) / 60_000))}m)` : 'hold: none',
223    state.note === null ? 'note: none' : `note: ${state.note}`,
224  ]
225  if (state.isAsked) lines.push('compaction asked for: when this turn ends')
226  return lines.join('\n')
227}
228
229// What a call of the tool does and answers. A refusal says what was wrong and changes
230// nothing: the state handed back is then the very one handed in.
231export function answerTool(state: AgentTimed, input: ToolInput, ctx: ToolContext): { state: AgentTimed; text: string } {
232  if (!ctx.isOn) return { state, text: 'Agent-timed compaction is off in this chat. Nothing changed.' }
233  const say = (text: string, next: AgentTimed = state) => ({ state: next, text: `${text}\n${figures(ctx.percent, ctx.startAt, ctx.cap)}` })
234  const { action } = input
235  if (typeof action !== 'string' || !ACTIONS.includes(action)) {
236    return say('Unknown action. Call it with action "hold" (and a reason), "release", "compact", "note" or "status". Nothing changed.')
237  }
238  if (action === 'status') return say(statusText(state, ctx.now))
239  if (ctx.isSubagent) {
240    return say('Only the main agent can hold, release, compact or write notes: the hold and the note are its own. Finish your task and report back. Nothing changed.')
241  }
242  if (action === 'hold') {
243    const reason = typeof input.reason === 'string' ? input.reason.trim().slice(0, REASON_MAX) : ''
244    if (reason === '') return say('A hold needs a reason: action "hold", reason "<what is fragile>". Nothing changed.')
245    const hold = { reason, since: state.hold?.since ?? ctx.now, remindAt: ctx.now + HOLD_REMIND_MS }
246    return say(`${state.hold ? 'Hold updated' : 'Hold set'}: compaction waits until you release, or until the cap. Reason: ${reason}.`, {
247      ...state,
248      hold,
249      isAsked: false,
250      nudge: { ...state.nudge, isBreakpointSaid: false },
251    })
252  }
253  const note = typeof input.note === 'string' ? input.note : undefined
254  if (note !== undefined && note.length > NOTE_MAX) return say(`The note is ${note.length} characters; the most is ${NOTE_MAX}. Nothing changed.`)
255  const noted = (next: AgentTimed): AgentTimed => (note === undefined ? next : { ...next, note: note === '' ? null : note })
256  const saved = note === undefined ? '' : note === '' ? ' Handoff note cleared.' : ' Handoff note saved; it comes back after the compaction.'
257  if (action === 'release') {
258    const when = Math.round(ctx.percent) >= ctx.startAt ? 'Compaction can run when this turn ends.' : `Compaction waits until context reaches ${ctx.startAt}%.`
259    return say(`${state.hold ? 'Released.' : 'No hold was set.'} ${when}${saved}`, noted({ ...state, hold: null }))
260  }
261  if (action === 'compact') {
262    if (Math.round(ctx.percent) < ASK_MIN) return say(`Context is under ${ASK_MIN}%: there is nothing worth compacting. Nothing changed.`)
263    return say(`Compaction runs when this turn ends.${saved}`, noted({ ...state, hold: null, isAsked: true }))
264  }
265  if (note === undefined) return say('A note needs text: action "note", note "<what must survive>" (an empty note clears it). Nothing changed.')
266  return say(note === '' ? 'Handoff note cleared.' : `Handoff note saved (${note.length} characters). It goes to the summarizer and comes back after the next compaction.`, noted(state))
267}
268
hooks/cache-policy.ts 467 lines
1import type { ConfigValue, ModelForkResult, ModelUsage } from 'claude-code'
2
3import type { CacheUsage, Limit, Ttl, TtlChoice, WarmAnchor, WarmRate, WarmStatus, WarmTotals } from '../types'
4
5// Keep cache warm, the part with no engine in it: the prices, when a refresh pays,
6// what one costs and saves, the lifetime Claude Code picks, and every word the band
7// and the transcript show.
8//
9// Ported from cache-warmer 0.12.0 (github.com/paulbkim-dev/claude-code-cache-warmer),
10// its hooks/warmer.ts, under the MIT licence:
11//
12//   Copyright (c) 2026 Paul B. Kim
13//
14//   Permission is hereby granted, free of charge, to any person obtaining a copy
15//   of this software and associated documentation files (the "Software"), to deal
16//   in the Software without restriction, including without limitation the rights
17//   to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
18//   copies of the Software, and to permit persons to whom the Software is
19//   furnished to do so, subject to the following conditions:
20//
21//   The above copyright notice and this permission notice shall be included in all
22//   copies or substantial portions of the Software.
23//
24//   THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
25//   IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
26//   FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
27//   AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
28//   LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
29//   OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
30//   SOFTWARE.
31//
32// The rule of when a refresh pays (`decide`) is the cache warmer's in Pi
33// (github.com/earendil-works/pi) by Mario Zechner, which cache-warmer ports.
34
35// the prices below were read on this day: a model added since gets no warming, never a wrong one
36export const PRICES_AS_OF = '2026-10-04'
37
38type Price = { input: number; cacheRead: number; output: number }
39
40const price = (input: number, cacheRead: number, output: number): Price => ({ input, cacheRead, output })
41
42// USD per million tokens at standard rates (platform.claude.com/docs/en/about-claude/pricing).
43// A cache write costs 1.25x input for the 5-minute lifetime, 2x for the 1-hour one
44const PRICES: Record<string, Price> = {
45  'claude-fable-5-1': price(10, 0.25, 50),
46  'claude-mythos-5-1': price(10, 0.25, 50),
47  'claude-fable-5': price(10, 1, 50),
48  'claude-mythos-5': price(10, 1, 50),
49  'claude-opus-5-5': price(4, 0.2, 20),
50  'claude-opus-5': price(5, 0.5, 25),
51  'claude-opus-4-8': price(5, 0.5, 25),
52  'claude-opus-4-7': price(5, 0.5, 25),
53  'claude-opus-4-6': price(5, 0.5, 25),
54  'claude-opus-4-5': price(5, 0.5, 25),
55  'claude-sonnet-5-5': price(2, 0.2, 10),
56  'claude-sonnet-5': price(2, 0.2, 10),
57  'claude-sonnet-4-6': price(3, 0.3, 15),
58  'claude-sonnet-4-5': price(3, 0.3, 15),
59  'claude-haiku-4-5': price(1, 0.1, 5),
60}
61
62const WRITE_MULTIPLIER: Record<Ttl, number> = { '5m': 1.25, '1h': 2 }
63
64export const TTL_MS: Record<Ttl, number> = { '5m': 300_000, '1h': 3_600_000 }
65
66// Pi's rule: a refresh is sent only when it is expected to save this much
67export const MIN_SAVINGS_USD = 0.05
68// Pi's measured chance that a prompt comes before the cache expires while the session is idle
69export const IDLE_CONTINUATION = 0.15
70// a fork has no output cap; Opus 5.5 at high effort answered a refresh in 182 output tokens
71export const DEFAULT_OUTPUT_TOKENS = 200
72// the refreshes each lifetime may send while no turn runs, by default and at most
73export const IDLE_LIMIT_DEFAULT = 5
74export const IDLE_LIMIT_MAX = 20
75// no refresh while either limit is at or past this %: a refresh spends the plan like any request
76export const WARM_UNTIL_DEFAULT = 85
77// the band warns this long before the cache expires
78export const EXPIRING_MS = 2 * 60_000
79// one window's reset time as two readings give it: a few minutes apart at most
80const SAME_WINDOW_MS = 10 * 60_000
81
82export const ZERO_TOTALS: WarmTotals = { refreshes: 0, costUsd: 0, wastedUsd: 0, kept: 0, keptUsd: 0 }
83
84// the one user message a refresh appends to the fork: it says what it is, so a
85// transcript reader (or the model) never takes it for the person
86export const FORK_PROMPT =
87  '[headroom] Automated prompt cache refresh by the headroom plugin, not a message from the user. Reply with the single word ok.'
88
89const WINDOW_LABELS: Record<string, string> = { five_hour: '5 Hour', seven_day: 'Weekly' }
90
91const isPriceKey = (id: string): boolean => Object.prototype.hasOwnProperty.call(PRICES, id)
92
93// a gateway's prefix, a [1m] context suffix and a dated id all name the same prices
94function rates(model: string): Price | undefined {
95  const id = model.replace(/^(?:claude-gateway|anthropic)\//, '').replace(/\[\d+(?:k|m)\]$/, '')
96  const key = isPriceKey(id) ? id : id.match(/^(claude-[a-z]+-\d{1,2}(?:-\d{1,2})?)-\d{8}$/)?.[1]
97  return key !== undefined && isPriceKey(key) ? PRICES[key] : undefined
98}
99
100/** a fork that made a request; the other kind found nothing to fork */
101export type ForkReply = Exclude<ModelForkResult, { reason: 'nothing-to-fork' }>
102
103/** one refresh as it settled */
104export type Refresh = {
105  at: number
106  model: string
107  usage: CacheUsage | null
108  costUsd: number | null
109  /** the rewrite the next prompt avoids, less this refresh: earned only if a prompt follows in time */
110  savesUsd: number | null
111  result: 'warmed' | 'expired' | 'failed'
112  detail?: string
113}
114
115export const isTtl = (value: unknown): value is Ttl => value === '5m' || value === '1h'
116export const isTtlChoice = (value: unknown): value is TtlChoice => value === 'auto' || isTtl(value)
117
118// the band's lifetime button goes round: auto, 5m, 1h, auto
119export const nextChoice = (choice: TtlChoice): TtlChoice => (choice === 'auto' ? '5m' : choice === '5m' ? '1h' : 'auto')
120
121export const usageOf = (usage: ModelUsage): CacheUsage => ({
122  input: usage.input_tokens,
123  output: usage.output_tokens,
124  cacheRead: usage.cache_read_input_tokens,
125  cacheWrite: usage.cache_creation_input_tokens,
126})
127
128export const promptTokensOf = (usage: CacheUsage): number => usage.input + usage.cacheRead + usage.cacheWrite
129
130// a refresh at 90% of the lifetime, at least ten seconds before it ends
131export const delayOf = (ttl: Ttl): number => Math.floor(Math.min(TTL_MS[ttl] * 0.9, TTL_MS[ttl] - 10_000))
132
133// Pi stops warming a running turn 60 minutes after its last request; the 1-hour
134// lifetime stretches that to two lifetimes so that it warms at all
135export const horizonOf = (ttl: Ttl): number => Math.max(60 * 60_000, 2 * TTL_MS[ttl])
136
137// a /config value as an idle limit: a whole number from 0 to IDLE_LIMIT_MAX
138export function limitOf(value: ConfigValue | undefined): number {
139  const count = Math.round(Number(value))
140  return value !== undefined && Number.isFinite(count) ? Math.min(IDLE_LIMIT_MAX, Math.max(0, count)) : IDLE_LIMIT_DEFAULT
141}
142
143// a /config value as the warmUntil %: a whole number from 1 to 100
144export function warmUntilOf(value: ConfigValue | undefined): number {
145  const percent = Math.round(Number(value))
146  return value !== undefined && Number.isFinite(percent) ? Math.min(100, Math.max(1, percent)) : WARM_UNTIL_DEFAULT
147}
148
149// how long an idle cache stays warm: the limit's refreshes, then one lifetime
150export const idleSpanOf = (count: number, ttl: Ttl): number => count * delayOf(ttl) + TTL_MS[ttl]
151
152export const addTotals = <T extends WarmTotals>(base: T, delta: Partial<WarmTotals>): T => ({
153  ...base,
154  refreshes: base.refreshes + (delta.refreshes ?? 0),
155  costUsd: base.costUsd + (delta.costUsd ?? 0),
156  wastedUsd: base.wastedUsd + (delta.wastedUsd ?? 0),
157  kept: base.kept + (delta.kept ?? 0),
158  keptUsd: base.keptUsd + (delta.keptUsd ?? 0),
159})
160
161// a timer that fires this late would likely find the cache gone and pay a whole write
162export const deadlineOf = (lastAt: number, ttl: Ttl): number => lastAt + delayOf(ttl) + Math.floor((TTL_MS[ttl] - delayOf(ttl)) / 2)
163
164export function costOf(model: string, usage: CacheUsage, ttl: Ttl): number | null {
165  const p = rates(model)
166  if (!p) return null
167  return (usage.input * p.input + usage.cacheWrite * p.input * WRITE_MULTIPLIER[ttl] + usage.cacheRead * p.cacheRead + usage.output * p.output) / 1_000_000
168}
169
170// what losing the cache adds to the next prompt: writing the prefix again instead of reading it
171export function missCostOf(model: string, promptTokens: number, ttl: Ttl): number | null {
172  const p = rates(model)
173  if (!p) return null
174  return Math.max(0, (promptTokens * (p.input * WRITE_MULTIPLIER[ttl] - p.cacheRead)) / 1_000_000)
175}
176
177// reading the prefix from a warm cache: what the next message pays for it while warm
178export function readCostOf(model: string, promptTokens: number): number | null {
179  const p = rates(model)
180  return p ? (promptTokens * p.cacheRead) / 1_000_000 : null
181}
182
183export type Decision = {
184  warmUsd: number
185  missUsd: number
186  probability: number
187  expectedUsd: number
188  isWarm: boolean
189  /** why a refresh does not pay: the prompt is under the size where the expected saving reaches MIN_SAVINGS_USD */
190  reason?: string
191}
192
193function breakEvenOf(p: Price, ttl: Ttl, probability: number, outputTokens: number): number | null {
194  const savedPerToken = probability * (p.input * WRITE_MULTIPLIER[ttl] - p.cacheRead) - p.cacheRead
195  if (savedPerToken <= 0) return null
196  return Math.ceil((MIN_SAVINGS_USD * 1_000_000 + outputTokens * p.output) / savedPerToken)
197}
198
199// Pi's rule: chance of a prompt before expiry × the rewrite it avoids − the refresh ≥ $0.05.
200// The chance is 1 while a turn runs, IDLE_CONTINUATION while idle. null: no price
201export function decide(model: string, promptTokens: number, ttl: Ttl, phase: 'run' | 'idle', outputTokens: number): Decision | null {
202  const p = rates(model)
203  const missUsd = missCostOf(model, promptTokens, ttl)
204  if (!p || missUsd === null || promptTokens <= 0) return null
205  const warmUsd = (promptTokens * p.cacheRead + outputTokens * p.output) / 1_000_000
206  const probability = phase === 'idle' ? IDLE_CONTINUATION : 1
207  const expectedUsd = probability * missUsd - warmUsd
208  const isWarm = expectedUsd >= MIN_SAVINGS_USD
209  const breakEven = isWarm ? null : breakEvenOf(p, ttl, probability, outputTokens)
210  const where = `${model} at ${ttl} (${phase})`
211  const reason = isWarm
212    ? undefined
213    : breakEven === null
214      ? `a refresh never saves on ${where}`
215      : `${formatTokens(promptTokens)} tokens is below the ${formatTokens(breakEven)} break-even on ${where}`
216  return { warmUsd, missUsd, probability, expectedUsd, isWarm, reason }
217}
218
219// a refresh that read under half the prefix found the cache gone and wrote it again
220export function outcomeOf(reply: ForkReply, usage: CacheUsage, promptTokens: number): Pick<Refresh, 'result' | 'detail'> {
221  if (usage.cacheRead >= promptTokens / 2) return { result: 'warmed' }
222  if (reply.isAnswered) return { result: 'expired' }
223  return {
224    result: 'failed',
225    detail: reply.reason === 'api-error' ? `${reply.error} ${reply.status ?? ''}`.trim() : reply.reason,
226  }
227}
228
229// the transcript's record of a refresh: ☕ 1h · Cache warmed · read 48.2k · $0.01 · saves $0.09 vs rewrite
230export function noticeText(ttl: Ttl, entry: Refresh): string {
231  const cost = `${entry.usage ? `read ${formatTokens(entry.usage.cacheRead)} · ` : ''}${entry.costUsd === null ? 'cost unknown' : formatUsd(entry.costUsd)}`
232  if (entry.result === 'expired') return `☕ ${ttl} · Cache had expired · the refresh rewrote it · ${cost}`
233  if (entry.result === 'failed') return `☕ ${ttl} · Cache refresh failed · ${entry.detail ? `${entry.detail} · ${cost}` : cost}`
234  return `☕ ${ttl} · Cache warmed · ${cost} · saves ${entry.savesUsd === null ? 'unknown' : formatUsd(entry.savesUsd)} vs rewrite`
235}
236
237// the last idle refresh has been sent: when the cache it warmed goes
238export const idleStopReason = (limit: number, expiresAt: number): string =>
239  `all ${limit} idle refreshes used, cache expires at ${clockOf(expiresAt)}`
240
241export const idleStopNotice = (ttl: Ttl, limit: number, expiresAt: number): string =>
242  `☕ ${ttl} · Warming stopped · all ${limit} idle refreshes used · cache expires at ${clockOf(expiresAt)}`
243
244/** Claude Code's own choice of the main conversation's lifetime, and what it reads to make it */
245export type TtlInputs = {
246  /** FORCE_PROMPT_CACHING_5M set */
247  force5m: boolean
248  /** CLAUDE_CODE_PROMPT_CACHE_TTL */
249  envTtl?: string
250  /** the promptCacheTtl setting */
251  settingTtl?: string
252  /** ENABLE_PROMPT_CACHING_1H set */
253  enable1h: boolean
254  /** a Claude subscription: the snapshot has a 5 Hour or Weekly limit */
255  isSubscription: boolean
256  /** a limit at or past 100%: the subscription runs on overage */
257  isOverage: boolean
258}
259
260// The lifetime in force, in Claude Code 2.1.291's order (read from its binary): the
261// force switch, then the variable (which a chosen 5m or 1h sets), the setting, the 1h
262// switch, and otherwise 1h on a subscription within its limits, 5m on an API key or in
263// overage. `choice`: the lifetime this mod set in the variable, `auto` when it set none
264export function ttlOf(i: TtlInputs, choice: TtlChoice = 'auto'): Ttl {
265  if (i.force5m) return '5m'
266  if (choice !== 'auto') return choice
267  if (isTtl(i.envTtl)) return i.envTtl
268  if (isTtl(i.settingTtl)) return i.settingTtl
269  if (i.enable1h) return '1h'
270  return i.isSubscription && !i.isOverage ? '1h' : '5m'
271}
272
273// a switch variable as Claude Code reads one: set, and not a word for off
274export const isEnvOn = (value: string | undefined): boolean =>
275  value !== undefined && !['', '0', 'false', 'no', 'off'].includes(value.trim().toLowerCase())
276
277const WINDOW_KINDS = ['five_hour', 'seven_day']
278
279// a subscription has the 5 Hour or Weekly limit; an API key has none
280export function planOf(limits: readonly Limit[]): { isSubscription: boolean; isOverage: boolean } {
281  const windows = limits.filter(l => WINDOW_KINDS.includes(l.kind))
282  return { isSubscription: windows.length > 0, isOverage: windows.some(l => l.percentUsed >= 100) }
283}
284
285// the first limit at or past warmUntil: no refresh then, as the status line says
286export function pastLimitOf(limits: readonly Limit[], warmUntil: number): string | null {
287  for (const kind of WINDOW_KINDS) {
288    const l = limits.find(x => x.kind === kind)
289    if (l && Math.round(l.percentUsed) >= warmUntil) return `${WINDOW_LABELS[kind]} at ${Math.round(l.percentUsed)}%, past your ${warmUntil}%`
290  }
291  return null
292}
293
294const sameWindow = (a: string | undefined, b: string | undefined): boolean =>
295  a !== undefined && b !== undefined && Math.abs(Date.parse(a) - Date.parse(b)) <= SAME_WINDOW_MS
296
297// How far each window's meter moved between two turn ends, in one window (by its
298// reset): a window that reset meanwhile, or has no reset, says nothing. A window that
299// did not move is a jump of 0, so the dollars it took still count against the rate
300export function jumpsOf(before: readonly Limit[] | null, now: readonly Limit[]): Record<string, number> {
301  const jumps: Record<string, number> = {}
302  if (!before) return jumps
303  for (const kind of WINDOW_KINDS) {
304    const b = before.find(l => l.kind === kind)
305    const n = now.find(l => l.kind === kind)
306    if (b && n && sameWindow(b.resetsAt, n.resetsAt)) jumps[kind] = Math.max(0, n.percentUsed - b.percentUsed)
307  }
308  return jumps
309}
310
311// the turn's dollars go to every window that measured a jump
312export function addRate(rate: WarmRate, jumps: Record<string, number>, usd: number): WarmRate {
313  const next = { ...rate }
314  for (const [kind, jump] of Object.entries(jumps)) {
315    const r = next[kind] ?? { jump: 0, usd: 0 }
316    next[kind] = { jump: r.jump + jump, usd: r.usd + usd }
317  }
318  return next
319}
320
321// How much of a window a dollar moves: Σjump / Σusd. The meter has one decimal at best
322// (whole points from the replies), so a rate is trusted only once a whole point has moved
323export function percentPerUsd(r: { jump: number; usd: number } | undefined): number | null {
324  return r && r.jump >= 1 && r.usd > 0 ? r.jump / r.usd : null
325}
326
327/** the window a cost is shown in, and how much of it a dollar moves */
328export type CostView = { label: string; perUsd: number }
329
330// the 5 Hour window when it has a rate (the session's own limit), else Weekly; none on an API key
331export function viewOf(rate: WarmRate, limits: readonly Limit[]): CostView | null {
332  for (const kind of WINDOW_KINDS) {
333    const perUsd = percentPerUsd(rate[kind])
334    if (perUsd !== null && limits.some(l => l.kind === kind)) return { label: WINDOW_LABELS[kind]!, perUsd }
335  }
336  return null
337}
338
339export const formatPercent = (p: number): string => `${p >= 0.1 ? p.toFixed(1) : p > 0 ? p.toPrecision(1) : '0'}%`
340
341export function formatUsd(usd: number): string {
342  const sign = usd < 0 ? '-' : ''
343  const value = Math.abs(usd)
344  return `${sign}$${value > 0 && value < 0.01 ? value.toPrecision(2) : value.toFixed(2)}`
345}
346
347export const formatTokens = (n: number): string =>
348  n >= 1_000_000 ? `${(n / 1_000_000).toFixed(1)}M` : n >= 1000 ? `${(n / 1000).toFixed(1)}k` : String(n)
349
350export function formatDuration(ms: number): string {
351  const s = Math.max(0, Math.round(ms / 1000))
352  if (s >= 3600) return `${Math.floor(s / 3600)}h${Math.floor((s % 3600) / 60)}m`
353  if (s >= 60) return `${Math.floor(s / 60)}m${s % 60 ? `${s % 60}s` : ''}`
354  return `${s}s`
355}
356
357// a span in the band's own words, to the minute: 38m, 1h 2m; under a minute said so
358export function spanOf(ms: number): string {
359  const m = Math.floor(Math.max(0, ms) / 60_000)
360  if (m < 1) return 'under a minute'
361  const h = Math.floor(m / 60)
362  const min = m % 60
363  if (h > 0) return min > 0 ? `${h}h ${min}m` : `${h}h`
364  return `${min}m`
365}
366
367// a time as the band says one elsewhere: 2:32 PM
368const clockOf = (t: number): string => new Date(t).toLocaleTimeString('en-US', { hour: 'numeric', minute: '2-digit' })
369
370/** a line of the band: what it says, and whether it calls for attention */
371export type CacheLine = { text: string; tone: 'muted' | 'amber' }
372
373const plural = (n: number, one: string, many: string) => `${n} ${n === 1 ? one : many}`
374
375// what the refreshes cost and saved: "3 refreshes this session, $0.04, saved $0.31"
376function totalsText(totals: WarmTotals, span: string): string {
377  return `${plural(totals.refreshes, 'refresh', 'refreshes')} ${span}, ${formatUsd(totals.costUsd)}${totals.keptUsd > 0 ? `, saved ${formatUsd(totals.keptUsd)}` : ''}`
378}
379
380// the warmer's own line while it is on: scheduled, refreshing, or stopped and why; with
381// this session's totals, and all time's where earlier sessions add to them
382export function warmStatusText(status: WarmStatus, totals: WarmTotals, now: number, assumed: Ttl | null = null, allTime: WarmTotals | null = null): CacheLine | null {
383  if (status.state === 'refreshing') return { text: 'Refreshing the cache…', tone: 'amber' }
384  if (status.state === 'stopped') return { text: `Warming ${status.isPaused ? 'paused' : 'stopped'}: ${status.reason}`, tone: 'muted' }
385  if (status.state !== 'scheduled') return null
386  const parts = [`Cache warm`, `refresh in ${spanOf(status.nextAt - now)}`]
387  if (assumed) parts.push(`${assumed} assumed`)
388  if (totals.refreshes > 0) parts.push(totalsText(totals, 'this session'))
389  if (allTime && allTime.refreshes > totals.refreshes) parts.push(totalsText(allTime, 'all time'))
390  return { text: parts.join(' · '), tone: 'muted' }
391}
392
393/** what the warning needs: the prefix, its dollars, and the window to show them in */
394export type CostOfNext = {
395  promptTokens: number
396  /** the rewrite over the warm read (missCostOf); null: no price */
397  missUsd: number | null
398  /** the warm read alone */
399  readUsd: number | null
400  view: CostView | null
401}
402
403// a dollar figure in the window's % when a rate is known, else in dollars
404const amountOf = (usd: number, view: CostView | null): string => (view ? formatPercent(usd * view.perUsd) : formatUsd(usd))
405
406function rewriteOf(c: CostOfNext, isWarmShown: boolean): string {
407  const tokens = `${formatTokens(c.promptTokens)} tokens`
408  if (c.missUsd === null) return tokens
409  const about = `${tokens}, about ${amountOf(c.missUsd, c.view)}${c.view ? ` of ${c.view.label}` : ''}`
410  return isWarmShown && c.readUsd !== null ? `${about} (warm: ${amountOf(c.readUsd, c.view)})` : about
411}
412
413export const expiredText = (agoMs: number, c: CostOfNext): string =>
414  `Cache expired ${spanOf(agoMs)} ago: your next message rewrites ${rewriteOf(c, true)}`
415
416export const expiringText = (inMs: number, c: CostOfNext): string =>
417  `Cache expires in ${spanOf(inMs)}: the next message after that rewrites ${rewriteOf(c, false)}`
418
419// with warming off, what a refresh would have cost instead
420export function nudgeText(refreshUsd: number | null, view: CostView | null, isExpired: boolean): string {
421  if (refreshUsd === null) return ''
422  return ` · Keep cache warm ${isExpired ? 'would have kept' : 'would keep'} it for about ${amountOf(refreshUsd, view)}`
423}
424
425export type CacheLineInput = {
426  isOn: boolean
427  anchor: WarmAnchor | null
428  status: WarmStatus
429  totals: WarmTotals
430  /** every session's totals, this one's included */
431  allTime: WarmTotals | null
432  assumed: Ttl | null
433  rate: WarmRate
434  limits: readonly Limit[]
435  outputTokens: number
436  now: number
437}
438
439// The band's cache line, above the bars: the warmer's status while it is on; else, and
440// once a stopped warmer's cache has gone, what the next message costs (amber once
441// expired, muted in the minutes before). Nothing before the first reply.
442export function cacheLineOf(i: CacheLineInput): CacheLine | null {
443  const a = i.anchor
444  if (!a) return null
445  const expiresAt = a.lastAt + TTL_MS[a.ttl]
446  const isExpired = i.now >= expiresAt
447  if (i.isOn) {
448    const line = warmStatusText(i.status, i.totals, i.now, i.assumed, i.allTime)
449    if (line && !(i.status.state === 'stopped' && isExpired)) return line
450  }
451  const view = viewOf(i.rate, i.limits)
452  const cost: CostOfNext = {
453    promptTokens: a.promptTokens,
454    missUsd: missCostOf(a.model, a.promptTokens, a.ttl),
455    readUsd: readCostOf(a.model, a.promptTokens),
456    view,
457  }
458  const refreshUsd = decide(a.model, a.promptTokens, a.ttl, 'idle', i.outputTokens)?.warmUsd ?? null
459  if (isExpired) {
460    return { text: expiredText(i.now - expiresAt, cost) + (i.isOn ? '' : nudgeText(refreshUsd, view, true)), tone: 'amber' }
461  }
462  if (!i.isOn && expiresAt - i.now <= EXPIRING_MS) {
463    return { text: expiringText(expiresAt - i.now, cost) + nudgeText(refreshUsd, view, false), tone: 'muted' }
464  }
465  return null
466}
467
types/index.d.ts 129 lines
1export type Category = { name: string; tokens: number; kind: 'used' | 'free' | 'buffer' }
2export type Limit = { kind: string; percentUsed: number; resetsAt?: string }
3export type PaceOf = { typical: number; skip: number }
4export type Snapshot = {
5  window: number
6  tokens: number
7  percent: number
8  total: number
9  categories: Category[]
10  limits: Limit[]
11  /** when the limit figures were read */
12  limitsAt?: number
13  /** why the usage service last refused, while it is backed off */
14  limitsError?: string
15  /** per window kind: the usual final % and the one-off % left out of the pace */
16  pace?: Record<string, PaceOf>
17  /** the minute it was taken in: redraws countdowns at least once a minute */
18  minute?: number
19}
20
21/**
22 * auto compact: on or off, and the context % that sets it off (kept for every chat).
23 * Agent-timed: from `startAt` % the agent chooses the moment, up to `at` %, the cap
24 */
25export type AutoCompact = { isOn: boolean; at: number | null; isAgentTimed?: boolean; startAt?: number | null }
26
27/** what Agent-timed holds for the session; every compaction of the main conversation starts it over */
28export type AgentTimed = {
29  /** the chat the state belongs to: a reload keeps it, another chat starts it over */
30  chat: string | null
31  /** the agent's request to defer compaction below the cap */
32  hold: { reason: string; since: number; remindAt: number } | null
33  /** the handoff note: to the summarizer, back to the agent, then cleared */
34  note: string | null
35  /** the agent asked to compact when this turn ends */
36  isAsked: boolean
37  /** whether the agent knows it is past the start %: `next` until the coming turn begins */
38  told: 'no' | 'next' | 'yes'
39  /** the highest nudge sent this cycle, main tool calls since, and the breakpoint hint */
40  nudge: { level: number; calls: number; isBreakpointSaid: boolean }
41  /** a hold the cap ended, until the row after the compaction says so */
42  overridden: { reason: string; percent: number } | null
43}
44
45/** Keep cache warm: the prompt cache's lifetime, and the chat's choice of it (`auto`: Claude Code's own) */
46export type Ttl = '5m' | '1h'
47export type TtlChoice = 'auto' | Ttl
48/** a request's four token counts, in the policy's spelling */
49export type CacheUsage = { input: number; output: number; cacheRead: number; cacheWrite: number }
50export type WarmTotals = {
51  refreshes: number
52  costUsd: number
53  /** the fees of refresh chains no kept prompt followed */
54  wastedUsd: number
55  kept: number
56  /** the rewrites kept prompts avoided */
57  keptUsd: number
58}
59/** the main conversation's last request: the prefix a fork replays and keeps warm */
60export type WarmAnchor = {
61  at: number
62  lastAt: number
63  model: string
64  promptTokens: number
65  ttl: Ttl
66  refreshes: number
67  /** refreshes sent while no turn ran: the idle limit counts these */
68  idleRefreshes: number
69  /** what this chain's refreshes cost: wasted unless the next prompt is kept */
70  feeUsd: number
71  isStopped: boolean
72}
73export type WarmStatus =
74  | { state: 'waiting' }
75  | { state: 'scheduled'; nextAt: number; phase: 'run' | 'idle'; expectedUsd: number }
76  | { state: 'refreshing' }
77  /** `isPaused`: stopped by the warmUntil threshold, which the band words as a pause */
78  | { state: 'stopped'; reason: string; isPaused?: boolean }
79/** per window kind: the % the meter moved and the dollars spent meanwhile, summed */
80export type WarmRate = Record<string, { jump: number; usd: number }>
81/** Keep cache warm, the session's side: the chain, the lifetime in force, the totals, the learned rate */
82export type Warm = {
83  /** the chat the chain and the totals belong to: another chat starts them afresh */
84  chat: string | null
85  anchor: WarmAnchor | null
86  status: WarmStatus
87  isRunning: boolean
88  outputTokens: number
89  /** the first main response wrote the cache: its lifetime holds for the session */
90  isLocked: boolean
91  /** the lifetime in force */
92  ttl: Ttl
93  /** set when a refresh found a 1h cache gone: 5m from then, for the session */
94  assumed: Ttl | null
95  totals: WarmTotals
96  allTime: WarmTotals & { since: number }
97  rate: WarmRate
98  /** the limits at the last main turn's end: the jump is measured from them */
99  lastLimits: Limit[] | null
100}
101/** Keep cache warm, the chat's setting (kept for every chat): on or off, and the lifetime */
102export type WarmSetting = { isOn: boolean; ttl: TtlChoice }
103
104declare module 'claude-code' {
105  interface PluginState {
106    'headroom': {
107      snapshot: Snapshot | null
108      isOn: boolean
109      isCollapsed: boolean
110      autoCompact: AutoCompact
111      /** the choice the band is asking for: the context was already past a new % */
112      autoAsk: { at: number; percent: number } | null
113      /** bumped on every % set: draws the field afresh even when the value is unchanged */
114      fieldTick: number
115      fieldText: string | null
116      agentTimed: AgentTimed
117      /** the start % field's own tick and typed text, as fieldTick and fieldText are the cap's */
118      startTick: number
119      startText: string | null
120      /** the desktop app's theme, as its settings (or the OS) have it */
121      theme: 'dark' | 'light'
122      /** Keep cache warm: the chain, the lifetime, the totals and the learned rate */
123      warm: Warm
124      /** Keep cache warm, this chat's switch and lifetime */
125      warmSetting: WarmSetting
126    }
127  }
128}
129