Context weather and cache tax above the prompt: a coloured forecast of how full the window /autocompact measures against is, beside what the cache does to the…

Two questions about a Claude Code session, answered in one coloured band above the prompt:
☀ Clear, ☁ Cloudy, ☂ Showers, ☇ Storm, ↯ Compact soon — measured against the window /autocompact actually fires on, not the model's raw limit.☂ Showers 55% of 362k · ▃▄▅▅▆▇ · ▲ +12k last turn · cached 141.9k · 6h00m left · ping in 50m · -$0.12
To the left is the weather; to the right, the tax.
/plugin marketplace add <owner>/<repo>
/plugin install token-tax-weather@token-tax-weather
Or, with the claude CLI:
claude plugin marketplace add <owner>/<repo>
claude plugin install token-tax-weather@token-tax-weather
The band is raised on the desktop and terminal surfaces. On the surfaces that do not raise it the same facts go to the status line instead, without colour, so the two never both draw the readout.
The band carries the continuously-updating facts: the forecast, the cache read, the countdown, the price. The one-off notices are different. A cold resume, a cold send the guard warns about, and a cold write it records are pinned on the status line under the prompt — not in the transcript — so they are read once where you are working, rather than scrolled back to. A notice stays up while the turn it belongs to runs and clears when that turn finishes; a second notice simply replaces the first. Unlike the readout, a notice is shown on every surface, the band's own included.
| Command | What it does |
|---|---|
/keepwarm | Arms a six-hour window that pings the cache so it does not expire |
/keepwarm 90m, 2h30m, 6h | Arms a window of that length |
/keepwarm every 50m | Arms it, pinging every 50 minutes |
/keepwarm always | Arms it with no deadline, this session |
/keepwarm off | Disarms it |
/keepwarm status | Reports what is armed |
/cache-tax | The full card: state, window, prices, the cache's effect, break-even |
/cache-tax-diagnostic | Read-only. What the host's pricing interface exposes, and nothing else |
/keepwarm is the only thing that spends anything, and it is always yours to run.n/a rather than a number from another model.Unit prices in USD per million tokens, in hooks/register.ts:
export const RATES: Record<string, Rates> = {
'claude-fable-5': { input: 0.43, cacheRead: 0.03, cacheWrite: 0, output: 0.87 },
'claude-opus-5-5': { input: 0.0007, cacheRead: 0.005, cacheWrite: 0, output: 0.6 },
}
Matched on the engine's exact model ID, not the name the picker shows. Add a row for each model you use; nothing else needs changing.
The final number on the band is context × (cacheRead − input) — what the cache does to the bill for that context, not an indicator of whether the cache is currently warm. It is decided entirely by the model's rates and the context size, so it is the same number whether the cache is hot or has been cold for a day.
Whether the cache is warm right now is the <time> left field, not this one.
Everything it reads is local session state: token counts the engine already reports, and the context breakdown it computes itself with no network request. It sends nothing anywhere, and /cache-tax-diagnostic prints no settings values — only whether the host's pricing interface exposes any.
claude plugin test ./token-tax-weather
claude plugin validate ./token-tax-weather
The tests mount the real surface and read the drawn tree, so a tree the surface would refuse fails the suite rather than passing a stub. The engine writes this build's API declarations into .claude-plugin/types/ when it loads the plugin; that folder is generated and gitignored, and tsc -p token-tax-weather needs a load first.
This supersedes cache-tax 2.1.3. It registers /keepwarm and /cache-tax under the same names, so run only one of the two — whichever loads second skips the names already taken and says so.
hooks/register.ts 1023 lines1import type {
2 BoxProps,
3 CommandSpec,
4 ElementConstructor,
5 EngineInterface,
6 Register,
7 RenderElement,
8 SettingsReadArgs,
9 SettingsSource,
10 TextProps,
11} from 'claude-code'
12
13// token-tax-weather — supersedes the Original Mod (cache-tax 2.1.3).
14//
15// Ported from the Original Mod with six deliberate changes:
16// 1. the hard-coded first-party PRICES table is replaced by user-supplied rates
17// 2. uncached input is its own rate, not half the cache-write rate
18// 3. the guard warns only; no prompt is ever dropped
19// 4. warming is never armed automatically, only by an explicit command
20// 5. the status line is repainted as state changes — session start, every
21// turn, every command — instead of being written once and left to go stale
22// 6. output never repeats the plugin name the host already prints
23
24const TTL_MS = 60 * 60 * 1000
25const PING_AFTER_MS = 50 * 60 * 1000
26const MIN_PING_MS = 60 * 1000
27const DEFAULT_WINDOW_MS = 6 * 60 * 60 * 1000
28const BIG_TOKENS = 50000
29const PING_PROMPT = 'Reply with the single word: warm'
30
31// Namespaced so the Custom Mod never inherits the Original Mod's window, ping
32// period, or `always` flag — inheriting `always` would arm paid warming on its
33// own at session start.
34const KEY_DEADLINE = 'token-tax-weather:deadline'
35const KEY_EVERY = 'token-tax-weather:every'
36const KEY_GUARD = 'token-tax-weather:guard'
37
38const rateKeys = ['input', 'cacheRead', 'cacheWrite', 'output'] as const
39type Rates = { input: number; cacheRead: number; cacheWrite: number; output: number }
40
41// USER-SUPPLIED rates in USD per million tokens, keyed by the exact model ID the
42// engine reports. Desktop does not hand its managed four rates to a plugin, so
43// every cost below is only as right as this table, and is always labeled
44// user-supplied. Zero is a valid rate — cacheWrite is zero here, which means a
45// cold re-write is billed at nothing and break-even reports not-applicable. A
46// model with no row shows an explicit unknown cost rather than borrowing
47// another model's rates.
48//
49// These are the operator's own figures, moved between models on request. On
50// 2026-10-10 the Fable row took over the numbers first given for Opus, and Opus
51// took the second set. Both are stated per million tokens; if an operator quotes
52// a provider's per-1K numbers instead, the whole table is 1000x out.
53//
54// Note on the Opus row: its cacheRead (0.005) is about 7x its input (0.0007), so
55// reading the prompt cache costs more than sending the same context uncached.
56// The card reports what the arithmetic says; it does not pretend the figures are
57// typical.
58//
59// Add rows in this shape, then update the plugin:
60//
61// 'claude-opus-5-5': { input: 0, cacheRead: 0, cacheWrite: 0, output: 0 },
62//
63export const RATES: Record<string, Rates> = {
64 'claude-fable-5': { input: 0.43, cacheRead: 0.03, cacheWrite: 0, output: 0.87 },
65 'claude-opus-5-5': { input: 0.0007, cacheRead: 0.005, cacheWrite: 0, output: 0.6 },
66}
67
68const RATE_SOURCE = 'user-supplied rates'
69
70type PingRecord = { at: number; read: number; write: number; usd: number | null; warm: boolean }
71type Miss = { at: number; tokens: number; usd: number | null }
72type GuardMode = 'warn'
73
74export type State = {
75 sid: string
76 deadline: number
77 every: number
78 lastRequestAt: number
79 lastModel: string | null
80 ctx: number
81 compacted: boolean
82 guard: GuardMode
83 coldWritePending: boolean
84 misses: Miss[]
85 pending: { cancel: () => void } | null
86 last: PingRecord | null
87 stopped: string | null
88 /** Cache-read tokens of the last completed main turn: what the cache actually served. */
89 lastRead: number
90 /** True once this mod's command names are the session's, so the footer slot is ours to write. */
91 owns: boolean
92 /** The window the fill is measured against: the compaction window /autocompact sets, else the model's. */
93 window: number
94 /** The fill as a whole percentage of `window`. */
95 percent: number
96 /** Where that window came from (the engine's `ContextWindowSource`), or null when only the model's is known. */
97 windowSource: string | null
98 /** Recent readings of the context size, oldest first: the sparkline and the per-turn trend. */
99 history: number[]
100 /**
101 * The clock as of the last event that could change what the band says. A
102 * `ui.render` hook draws synchronously and `$.clock.now()` only answers a
103 * promise, so the countdown reads this rather than asking again; every event
104 * that repaints stamps it first.
105 */
106 now: number
107 /** The surface this session runs on: the band is raised on `desktop` and `terminal` only. */
108 surface: string
109 /**
110 * The one-off line pinned on the status line ahead of the readout, or null for
111 * none. It carries a cache-state notice — a cold resume, a cold send, a recorded
112 * cold write — rather than the continuously-updating facts, so it is read once
113 * and not scrolled back to. A second notice simply replaces the first: last
114 * writer wins, which is what the status line itself does.
115 */
116 notice: string | null
117}
118
119/** The configured rates for a model, or null when none are configured. Exact ID only. */
120export function ratesFor(model: string | null): Rates | null {
121 if (!model) return null
122 const row = RATES[model] ?? RATES[model.toLowerCase()]
123 if (!row) return null
124 const ok = rateKeys.every(k => typeof row[k] === 'number' && Number.isFinite(row[k]) && row[k] >= 0)
125 return ok ? row : null
126}
127
128export function parseDuration(text: string): number | null {
129 const m = /^(?:(\d+)h)?(?:(\d+)m)?$/.exec(text.trim())
130 if (!m || (m[1] === undefined && m[2] === undefined)) return null
131 return (Number(m[1] ?? 0) * 60 + Number(m[2] ?? 0)) * 60 * 1000
132}
133
134export function fmtDuration(ms: number): string {
135 const total = Math.max(0, Math.round(ms / 60000))
136 const h = Math.floor(total / 60)
137 const m = total % 60
138 if (h >= 48) return `${Math.floor(h / 24)}d ${h % 24}h`
139 return h > 0 ? `${h}h${String(m).padStart(2, '0')}m` : `${m}m`
140}
141
142function fmtUsd(usd: number | null): string {
143 if (usd == null) return 'n/a'
144 // Sub-cent figures are normal at these rates; two decimals would read as $0.00.
145 if (usd > 0 && usd < 0.01) return '$' + usd.toFixed(4)
146 return '$' + (usd >= 100 ? usd.toFixed(0) : usd.toFixed(2))
147}
148
149/** A signed difference: `+$0.12` costs more, `-$0.00` saves. */
150function fmtSignedUsd(usd: number | null): string {
151 if (usd == null) return 'n/a'
152 return (usd >= 0 ? '+' : '-') + fmtUsd(Math.abs(usd))
153}
154
155/**
156 * Rounding to whole thousands hid the difference the line exists to show: a 300k
157 * context served from a 299k cache and a 300k one served from 141k read the same.
158 * One decimal for the thousands and millions keeps the two figures apart, and the
159 * trailing `.0` is dropped so an exact thousand still reads as a round number.
160 */
161function fmtTok(n: number): string {
162 if (n >= 1e6) return trimZero((n / 1e6).toFixed(1)) + 'M'
163 if (n >= 1000) return trimZero((n / 1000).toFixed(1)) + 'k'
164 return String(n)
165}
166
167function trimZero(s: string): string {
168 return s.endsWith('.0') ? s.slice(0, -2) : s
169}
170
171/** ctx × the cache-write rate: what re-writing the context after expiry is billed. */
172function coldUsd(s: State): number | null {
173 const r = ratesFor(s.lastModel)
174 return r ? s.ctx * r.cacheWrite / 1e6 : null
175}
176
177/** ctx × the cache-read rate: what a turn whose context is served from cache costs. */
178function warmUsd(s: State): number | null {
179 const r = ratesFor(s.lastModel)
180 return r ? s.ctx * r.cacheRead / 1e6 : null
181}
182
183/** ctx × the uncached input rate: what the same context costs with no cache reuse at all. */
184function uncachedUsd(s: State): number | null {
185 const r = ratesFor(s.lastModel)
186 return r ? s.ctx * r.input / 1e6 : null
187}
188
189/**
190 * The adjustment the cache makes to the bill for this one context: from cache
191 * minus no cache, i.e. the cache-read rate against the input rate. **Negative
192 * means the cache lowers the bill** — it is a saving of that much, and the usual
193 * case where the read rate is far below the input rate (Fable: 0.03 against
194 * 0.43). **Positive means the cache raises it** — the read rate is above the
195 * input rate, so serving from cache costs more than not caching and the cache is
196 * a surcharge, which some operators' rates genuinely do (Opus: 0.005 against
197 * 0.0007).
198 *
199 * This is the reverse of "what a cold send costs": that is the same magnitude
200 * with the opposite sign. The sign here reads as an effect on the bill, so with
201 * a colour beside it negative is green and positive is red.
202 */
203function cacheEffectUsd(s: State): number | null {
204 const uncached = uncachedUsd(s)
205 const warm = warmUsd(s)
206 return uncached == null || warm == null ? null : warm - uncached
207}
208
209/** Everything a fork bills: the cache read, any cache write, uncached input at its own rate, and the output. */
210function pingUsd(u: { input_tokens: number; output_tokens: number; cache_read_input_tokens: number; cache_creation_input_tokens: number }, r: Rates): number {
211 return (u.cache_read_input_tokens * r.cacheRead + u.cache_creation_input_tokens * r.cacheWrite + u.input_tokens * r.input + u.output_tokens * r.output) / 1e6
212}
213
214/** The read-only upper bound: pings at the cache-read rate that cost as much as one cold write of the same context. */
215function breakEvenPings(s: State): number | null {
216 const r = ratesFor(s.lastModel)
217 if (!r || s.ctx <= 0 || r.cacheRead <= 0) return null
218 return Math.floor(r.cacheWrite / r.cacheRead)
219}
220
221function isCold(s: State, now: number): boolean {
222 return s.lastRequestAt > 0 && !s.compacted && now - s.lastRequestAt >= TTL_MS
223}
224
225function guardText(s: State, now: number): string {
226 const r = ratesFor(s.lastModel)
227 const rate = r ? `$${r.cacheWrite}/MTok` : 'the cache-write rate'
228 const warm = warmUsd(s)
229 const basis = r ? ` (${RATE_SOURCE})` : ' (no rates configured for this model)'
230 return `the prompt cache went cold ${fmtDuration(now - s.lastRequestAt - TTL_MS)} ago. Sending this re-writes ` +
231 `up to ${s.ctx.toLocaleString('en-US')} tokens at ${rate} = ${fmtUsd(coldUsd(s))}` +
232 (warm == null ? '' : ` (a warm turn would have cost ${fmtUsd(warm)})`) + basis + '.'
233}
234
235export type ResumeFields = {
236 source: string
237 model?: string
238 context_tokens?: number
239 seconds_since_last_response?: number
240 prompt_cache_likely_expired?: boolean
241 estimated_cache_write_usd?: number
242}
243
244/** /clear starts a new conversation in the same process; nothing priced before it still exists. */
245export function resetForClear(s: State) {
246 s.ctx = 0
247 s.lastRequestAt = 0
248 s.compacted = false
249 s.coldWritePending = false
250 s.misses = []
251 s.lastRead = 0
252 s.percent = 0
253 s.history = []
254 s.notice = null
255 disarm(s)
256}
257
258/** Applies a resumed session's fields to the state; returns the line to log, if any. */
259export function seedFromResume(s: State, e: ResumeFields, now: number): string | null {
260 if (e.source !== 'resume' && e.source !== 'fork') return null
261 if (typeof e.context_tokens === 'number' && e.context_tokens > 0) s.ctx = e.context_tokens
262 if (typeof e.seconds_since_last_response === 'number') s.lastRequestAt = now - e.seconds_since_last_response * 1000
263 if (typeof e.model === 'string') s.lastModel = e.model
264 s.compacted = false
265 if (e.prompt_cache_likely_expired !== true || s.ctx < BIG_TOKENS) return null
266 const usd = typeof e.estimated_cache_write_usd === 'number' ? fmtUsd(e.estimated_cache_write_usd) : fmtUsd(coldUsd(s))
267 return `resuming cold. The first message re-writes ${s.ctx.toLocaleString('en-US')} tokens, about ${usd}. /clear and paste a summary if you only need the conclusions.`
268}
269
270/** The warming slot on its own: off, a countdown, or why it stopped. */
271function warmingText(s: State, now: number): string {
272 if (s.stopped) return `stopped: ${s.stopped}`
273 if (!s.deadline) return 'off'
274 return `${fmtDuration(s.deadline - now)} left`
275}
276
277/** When the next ping can come, or null while nothing is armed. */
278function nextPingText(s: State, now: number): string | null {
279 if (!s.deadline) return null
280 if (!s.lastRequestAt) return 'ping after first turn'
281 if (isCold(s, now)) return 'ping after next turn'
282 return `ping in ${fmtDuration(s.lastRequestAt + s.every - now)}`
283}
284
285// The under-composer line, in one fixed field order so the same fact always
286// appears in the same place — durations first, then token counts, then money:
287//
288// <warming> · [ping] · <context> · [cached] · [±]
289//
290// warming off | <time> left | stopped: <why>
291// ping present only while a window is armed
292// context no context yet | ctx <tokens> — what the cache could be holding
293// cached absent until a turn reports a cache read: what it actually served
294// final absent until a context exists: what the cache does to the bill for
295// this context, bare and signed. `-` is a saving — the read rate is
296// below the input rate, so being warm lowers the bill by that much (the
297// usual case); `+` is a surcharge — the read rate is above the input
298// rate, so the cache costs more than it saves. `n/a` while this mod has
299// no rates for the model in use, so a figure it cannot compute is never
300// shown as zero. The sign is left to carry the meaning on its own; the
301// wording that would explain it is what the card is for.
302//
303// The model itself is left out: the picker beside this line already names it, and
304// `/cache-tax` prints the exact ID the rates are keyed on.
305function statusText(s: State, now: number): string {
306 const parts = [warmingText(s, now)]
307 const ping = nextPingText(s, now)
308 if (ping) parts.push(ping)
309 parts.push(s.ctx > 0 ? `ctx ${fmtTok(s.ctx)}` : 'no context yet')
310 if (s.lastRead > 0) parts.push(`cached ${fmtTok(s.lastRead)}`)
311 if (s.ctx > 0) {
312 const effect = cacheEffectUsd(s)
313 parts.push(effect == null ? 'n/a' : fmtSignedUsd(effect))
314 }
315 return parts.join(' · ')
316}
317
318function disarm(s: State) {
319 if (s.pending) s.pending.cancel()
320 s.pending = null
321}
322
323// --- The AbovePrompt band ----------------------------------------------------
324// The status line takes a plain string and cannot colour, so once this mod owns
325// the session the same facts move to the band above the prompt, where a tree's
326// `Text` does take a colour. The band is one instance and one plugin draws it:
327// a hook that returns a tree answers the site and the chain stops beneath it, so
328// a second plugin hooked on AbovePrompt never draws.
329
330/** Readings the sparkline shows, oldest first. */
331const HISTORY = 12
332const SPARK = '▁▂▃▄▅▆▇█'
333
334/**
335 * The forecast, by share of the window used. The colours are literal rather than
336 * theme keys on purpose: this is a ramp the person reads as a ramp, and it stays
337 * the same ramp whatever the theme does to `warning` or `error`.
338 */
339const FORECAST = [
340 { upTo: 25, icon: '☀', word: 'Clear', color: 'yellow' },
341 { upTo: 50, icon: '☁', word: 'Cloudy', color: 'cyan' },
342 { upTo: 75, icon: '☂', word: 'Showers', color: 'blue' },
343 { upTo: 90, icon: '☇', word: 'Storm', color: 'magenta' },
344 { upTo: Infinity, icon: '↯', word: 'Compact soon', color: 'red' },
345]
346
347function forecastFor(percent: number) {
348 return FORECAST.find(f => percent < f.upTo) ?? FORECAST[FORECAST.length - 1]!
349}
350
351/**
352 * One bar per reading, its height the share of the window it filled — not the
353 * share of the busiest reading shown, which is what token-weather scales by and
354 * which makes the same bar mean a different number of tokens in every window.
355 * Scaling to the window keeps a bar comparable with the percentage beside it.
356 */
357function sparkline(history: number[], window: number): string {
358 if (!window || !history.length) return ''
359 return history
360 .map(t => {
361 const level = Math.floor((t / window) * (SPARK.length - 1))
362 return SPARK[Math.min(SPARK.length - 1, Math.max(0, level))]!
363 })
364 .join('')
365}
366
367/** How much the context moved on the last turn, or null while there is no turn to compare. */
368function trendText(history: number[]): string | null {
369 if (history.length < 2) return null
370 const delta = history[history.length - 1]! - history[history.length - 2]!
371 if (delta === 0) return 'steady'
372 return `${delta > 0 ? '▲ +' : '▼ '}${fmtTok(Math.abs(delta))} last turn`
373}
374
375/**
376 * The window the fill is measured against, and the fill itself. `rawMaxTokens` is
377 * what /autocompact actually fires against — the model's limit, or the smaller
378 * compaction window when a setting names one — so it is the honest denominator;
379 * `context.window` is only the model's and would understate the fill whenever
380 * autocompact has set something smaller, which is the case that matters.
381 *
382 * `breakdown: 'summary'` estimates locally and sends no request. `full` would
383 * send one count request per tool and per memory file, which the standing
384 * local-only constraint rules out.
385 */
386async function refreshWindow($: EngineInterface, s: State) {
387 try {
388 const usage = await $.session.usage({ breakdown: 'summary' })
389 const context = usage.context
390 const breakdown = context.breakdown
391 if (breakdown && breakdown.rawMaxTokens > 0) {
392 s.window = breakdown.rawMaxTokens
393 s.percent = Math.round(breakdown.percentage)
394 s.windowSource = breakdown.autocompactSource
395 return
396 }
397 if (context.window > 0) {
398 s.window = context.window
399 s.percent = Math.round(context.percent ?? (s.ctx / context.window) * 100)
400 s.windowSource = null
401 }
402 } catch {
403 // No session bound, or a host that answers no breakdown: keep the last reading.
404 }
405}
406
407/**
408 * The band's tree. Sized to `columns` (the band's own box, which is narrower than
409 * the viewport while a pane is docked beside the transcript): the sparkline and
410 * the trend are the first things to go, then the cache read.
411 *
412 * The field order groups like kinds — the forecast and its measurement, then the
413 * token counts, then the durations, then the one price.
414 */
415function band(
416 Box: ElementConstructor<BoxProps>,
417 Text: ElementConstructor<TextProps>,
418 props: { bodyColumns: number },
419 s: State,
420) {
421 const columns = props.bodyColumns || 80
422 const f = forecastFor(s.percent)
423 const now = s.now || 0
424 // The table's constructors take children among the props and hand back an
425 // element with them beside `props`, which is the shape the engine draws.
426 const t = (textProps: TextProps, text: string) => Text({ ...textProps, children: text })
427 const dim = (text: string) => t({ dimColor: true }, text)
428 const parts: RenderElement[] = [
429 t({ color: f.color, bold: true }, `${f.icon} ${f.word}`),
430 t({ color: f.color }, ` ${s.percent}% of ${fmtTok(s.window)}`),
431 ]
432 const spark = columns >= 56 ? sparkline(s.history, s.window) : ''
433 if (spark) parts.push(dim(' · '), t({ color: f.color }, spark))
434 const trend = columns >= 72 ? trendText(s.history) : null
435 if (trend) parts.push(dim(` · ${trend}`))
436 if (s.lastRead > 0) parts.push(dim(` · cached ${fmtTok(s.lastRead)}`))
437 parts.push(dim(` · ${warmingText(s, now)}`))
438 const ping = nextPingText(s, now)
439 if (ping) parts.push(dim(` · ${ping}`))
440 if (s.ctx > 0) {
441 const effect = cacheEffectUsd(s)
442 if (effect == null || effect === 0) {
443 // `n/a` is a price this mod cannot compute, and a configured zero is a real
444 // answer; neither is a saving, so neither is coloured as one.
445 parts.push(dim(` · ${fmtSignedUsd(effect)}`))
446 } else {
447 // Negative lowers the bill (green), positive raises it (red). The theme's
448 // own keys, so the colours follow the person's theme rather than a guess.
449 parts.push(dim(' · '), t({ color: effect < 0 ? 'success' : 'error' }, fmtSignedUsd(effect)))
450 }
451 }
452 return Box({ flexDirection: 'row', paddingX: 1, children: parts })
453}
454
455// The line this mod keeps in its own slot under the prompt. The continuously
456// updating readout — what is armed when a window is, and otherwise what the
457// costing is based on — so the slot says something true about this mod rather
458// than sitting empty. It names the model the engine reports, because that exact
459// ID decides whether rates are found.
460//
461// The host already prints the plugin's own name ahead of this text, so the line
462// never repeats it.
463//
464// The model comes from the last turn's own usage and never from a session lookup:
465// at session start the host may not yet have told the engine which model to use,
466// and a footer must not be left waiting on that. Until the first turn the line
467// says so instead of guessing.
468//
469// Written only while this mod owns the session's command names: the slot holds one
470// entry per plugin, so painting into it while the Original Mod owns the session
471// would clobber that mod's footer. Repainted on session start, before each prompt,
472// and after each turn or command, because a single paint can be lost while the
473// surface is still attaching.
474//
475// The band above the prompt is raised on the desktop and terminal surfaces only,
476// and where it is raised it carries the readout in colour. So the status line is
477// the fallback for the surfaces the band never reaches, and the two never both
478// draw the readout: the same thing is never said twice. A live notice is not the
479// readout — it is pinned wherever it is set, the band's own surfaces included, and
480// it takes the line until the turn that raised it is over.
481const BAND_SURFACES = new Set(['desktop', 'terminal'])
482
483// The one place that writes the status line, so a notice and the readout never
484// race each other for it. A live notice outranks everything and shows on every
485// surface, the band's included: it is a one-off thing to read now, not the facts
486// the band keeps current, so it is worth saying even where the band is up.
487async function publish($: EngineInterface, s: State) {
488 // Stamped whatever happens: the band draws synchronously and reads this rather
489 // than asking the clock again.
490 s.now = await $.clock.now()
491 if (!s.owns) return
492 if (s.notice) {
493 $.ui.status(s.notice)
494 return
495 }
496 // With no notice the band says all of it in colour, so the status line is
497 // cleared rather than left stale — which is also what takes a finished notice
498 // off the screen.
499 if (BAND_SURFACES.has(s.surface)) {
500 $.ui.status(undefined)
501 return
502 }
503 $.ui.status(statusText(s, s.now))
504}
505
506// The window and its ping period belong to the session that armed them, so a
507// second session, or one resumed from another transcript, never inherits them
508// and cannot turn them off. The guard mode stays global.
509function deadlineKey(s: State): string {
510 return `${KEY_DEADLINE}:${s.sid}`
511}
512
513function everyKey(s: State): string {
514 return `${KEY_EVERY}:${s.sid}`
515}
516
517/** Clears this session's own dead window. Other sessions' keys are never touched. */
518async function prune($: EngineInterface, s: State, now: number) {
519 for (const key of [deadlineKey(s)]) {
520 const deadline = await $.store.get(key)
521 if (deadline === undefined) continue
522 if (typeof deadline === 'number' && deadline > now) continue
523 await $.store.delete(key)
524 await $.store.delete(everyKey(s))
525 }
526}
527
528async function stop($: EngineInterface, s: State, why: string | null) {
529 s.deadline = 0
530 s.every = PING_AFTER_MS
531 s.stopped = why
532 disarm(s)
533 await $.store.delete(deadlineKey(s))
534 await $.store.delete(everyKey(s))
535 await publish($, s)
536}
537
538async function arm($: EngineInterface, s: State) {
539 disarm(s)
540 if (!s.deadline) return
541 const now = await $.clock.now()
542 if (now >= s.deadline) return stop($, s, null)
543 // A cold window still needs expiry cleanup, but must not send a model request.
544 if (s.lastRequestAt && !s.compacted && !isCold(s, now)) {
545 const delay = Math.min(s.deadline - now, Math.max(1000, s.lastRequestAt + s.every - now))
546 s.pending = $.clock.after(delay, () => { void ping($, s) })
547 } else {
548 s.pending = $.clock.after(s.deadline - now, () => { void arm($, s) })
549 }
550 await publish($, s)
551}
552
553async function ping($: EngineInterface, s: State) {
554 s.pending = null
555 if (!s.deadline) return
556 const now = await $.clock.now()
557 if (now >= s.deadline) return arm($, s)
558 // A turn in the meantime re-armed the timer; this callback is stale.
559 if (now - s.lastRequestAt < s.every - 1000) return
560 if (isCold(s, now)) return arm($, s)
561 let reply
562 try {
563 reply = await $.model.fork({ prompt: PING_PROMPT })
564 } catch (err) {
565 return stop($, s, `ping failed: ${err instanceof Error ? err.message : String(err)}`)
566 }
567 // This build's fork never answers null: the union is an answer, or a reason.
568 if (!reply) return stop($, s, 'ping not sent')
569 if (!reply.isAnswered) return stop($, s, `ping not sent (${reply.reason})`)
570 const u = reply.usage
571 const r = ratesFor(s.lastModel)
572 // A warm ping reads the prefix and writes only its own message; a write of a tenth of the read or more means the prefix broke.
573 const warm = u.cache_read_input_tokens > 0 && u.cache_creation_input_tokens < 0.1 * u.cache_read_input_tokens
574 const usd = r ? pingUsd(u, r) : null
575 s.last = { at: now, read: u.cache_read_input_tokens, write: u.cache_creation_input_tokens, usd, warm }
576 if (!warm) return stop($, s, `ping read ${fmtTok(u.cache_read_input_tokens)}, wrote ${fmtTok(u.cache_creation_input_tokens)} (${fmtUsd(usd)}): the cache was gone`)
577 s.lastRequestAt = now
578 await arm($, s)
579}
580
581/** The reply to an arming command; on a cold cache it says when the first ping can come. */
582function armedText(s: State, now: number, windowMs: number): string {
583 if (isCold(s, now)) return `keepwarm on for ${fmtDuration(windowMs)}. The cache is cold now, so the first ping comes ${fmtDuration(s.every)} after the next turn`
584 return `keepwarm on for ${fmtDuration(windowMs)}, a ping ${fmtDuration(s.every)} after each idle stretch keeps the cache read, not re-written`
585}
586
587async function startWindow($: EngineInterface, s: State, windowMs: number, every: number) {
588 const now = await $.clock.now()
589 s.every = every
590 if (every === PING_AFTER_MS) await $.store.delete(everyKey(s))
591 else await $.store.set(everyKey(s), every)
592 s.deadline = now + windowMs
593 s.stopped = null
594 await $.store.set(deadlineKey(s), s.deadline)
595 await arm($, s)
596}
597
598/**
599 * Where the window the band measures against came from, said plainly. `settings`
600 * and `env` are a person's own `/autocompact` value; `model-default` is the
601 * model's own limit; `unknown-model` means the engine could not settle one and
602 * the percentage should be read as a guess, which is worth knowing when a
603 * gateway serves a model ID the engine does not recognise.
604 */
605function windowNote(s: State): string {
606 switch (s.windowSource) {
607 case 'settings': return ', set by /autocompact'
608 case 'env': return ', set by the environment'
609 case 'clientdata': return ', set by the client'
610 case 'experiment': return ', set by an experiment'
611 case 'model-default': return ", the model's own window"
612 case 'auto': return ', settled automatically'
613 case 'unknown-model': return ', the model is unknown to the engine, so this is a guess'
614 default: return ', no compaction window reported, so the model window is used'
615 }
616}
617
618function card(s: State, now: number): string {
619 const lines: string[] = []
620 const r = ratesFor(s.lastModel)
621 lines.push(`${s.lastModel ?? 'model not seen yet'}`)
622 if (s.compacted) lines.push('state reset by compaction, waiting for the first turn')
623 else if (!s.lastRequestAt) lines.push('state no request yet this session')
624 else if (isCold(s, now)) lines.push(`state COLD, last request ${fmtDuration(now - s.lastRequestAt)} ago`)
625 else lines.push(`state warm, ${fmtDuration(s.lastRequestAt + TTL_MS - now)} left`)
626 lines.push(`context ${s.ctx.toLocaleString('en-US')} tokens, the most the cache can be holding`)
627 lines.push(`window ${s.window ? `${fmtTok(s.window)} (${s.percent}% used)` : 'not reported yet'}${windowNote(s)}`)
628 // What the cache actually served last turn, rather than what it might hold.
629 if (s.lastRead > 0 && s.ctx > 0) {
630 const share = Math.min(100, Math.round((100 * s.lastRead) / s.ctx))
631 lines.push(`cache read ${s.lastRead.toLocaleString('en-US')} tokens served from cache last turn (${share}% of context)`)
632 } else if (s.lastRequestAt) {
633 lines.push('cache read none reported by the last turn, so nothing was served from cache')
634 }
635 if (r) {
636 lines.push(`unit prices input ${r.input} · read ${r.cacheRead} · write ${r.cacheWrite} · output ${r.output} USD/MTok (${RATE_SOURCE}; prices, not costs)`)
637 lines.push(`for ${fmtTok(s.ctx)} tokens re-write ${fmtUsd(coldUsd(s))} (write rate) · from cache ${fmtUsd(warmUsd(s))} (read rate) · no cache ${fmtUsd(uncachedUsd(s))} (input rate)`)
638 const effect = cacheEffectUsd(s)
639 lines.push(`cache effect ${fmtSignedUsd(effect)} for this context: "from cache" minus "no cache". Negative is a saving — the read rate is below the input rate, so being warm lowers the bill by that much. Positive is a surcharge — the read rate is above the input rate, so the cache costs more than it saves`)
640 } else {
641 lines.push(`unit prices none configured for this model, so no cost can be shown (n/a)`)
642 }
643 const armed = s.deadline ? `on, ${warmingText(s, now)}` : 'off (/keepwarm to arm it for 6h00m)'
644 lines.push(`keepwarm ${s.stopped ? `stopped, ${s.stopped}` : armed}`)
645 if (s.last) lines.push(`last ping read ${fmtTok(s.last.read)} from cache, wrote ${fmtTok(s.last.write)}, ${fmtUsd(s.last.usd)}${s.last.warm ? ' (warm)' : ' (the cache was gone)'}`)
646 const pings = breakEvenPings(s)
647 if (pings != null && pings > 0) lines.push(`break-even up to ${pings} pings at the read rate cost one cold write, about ${fmtDuration(pings * s.every)} of idle at one ping per ${fmtDuration(s.every)}`)
648 else if (r && r.cacheWrite <= 0) lines.push('break-even warming cannot pay for itself at these rates: a cold re-write is billed at nothing, so every ping is a net cost')
649 else if (r) lines.push('break-even not applicable: a zero cache-read rate, or no context yet')
650 lines.push(`guard warn only, always. This mod never drops a prompt.`)
651 const paid = s.misses.reduce((a, m) => a + (m.usd ?? 0), 0)
652 lines.push(`session ${s.misses.length} cold write${s.misses.length === 1 ? '' : 's'} paid, ${fmtUsd(paid)}`)
653 return lines.join('\n')
654}
655
656export function freshState(): State {
657 return {
658 sid: '', deadline: 0, every: PING_AFTER_MS, lastRequestAt: 0, lastModel: null, ctx: 0, compacted: false,
659 guard: 'warn', coldWritePending: false, misses: [], pending: null, last: null, stopped: null, lastRead: 0, owns: false,
660 window: 0, percent: 0, windowSource: null, history: [], now: 0, surface: '', notice: null,
661 }
662}
663
664// --- Host pricing provenance -------------------------------------------------
665// Desktop does not expose its managed four rates to a plugin. This section
666// reports what the settings interface actually shows, so the absence of
667// accepted pricing is visible rather than guessed at. It never becomes the
668// source of a cost: costs come from the user-supplied RATES table above.
669
670type Pricing = { rates: Rates; multiplier: number }
671type Snapshot = { available: boolean; present: boolean; exact: boolean; pricing?: Pricing }
672type Diagnostic = {
673 model: string
674 state: 'accepted pricing available' | 'candidate/unconfirmed' | 'pricing unavailable'
675 reason: string
676 pricing?: Pricing
677 sources: string[]
678}
679
680function record(value: unknown): Record<string, unknown> | undefined {
681 return value !== null && typeof value === 'object' && !Array.isArray(value)
682 ? value as Record<string, unknown>
683 : undefined
684}
685
686function validRate(value: unknown): value is number {
687 return typeof value === 'number' && Number.isFinite(value) && value >= 0 && value <= 10000
688}
689
690function validMultiplier(value: unknown): value is number {
691 return typeof value === 'number' && Number.isFinite(value) && value > 0 && value <= 10
692}
693
694// Select pricing fields immediately. Never retain, print, or persist settings.
695async function snapshot($: EngineInterface, model: string, args?: SettingsReadArgs): Promise<Snapshot> {
696 try {
697 const { modelPricing } = await $.settings.read(args)
698 const config = record(modelPricing)
699 const overrides = record(config?.overrides)
700 const exact = overrides !== undefined && Object.prototype.hasOwnProperty.call(overrides, model)
701 const row = exact ? record(overrides?.[model]) : undefined
702 const multiplier = config?.multiplier === undefined ? 1 : config.multiplier
703 let pricing: Pricing | undefined
704 if (row && rateKeys.every(key => validRate(row[key])) && validMultiplier(multiplier)) {
705 pricing = {
706 rates: { input: row.input as number, output: row.output as number,
707 cacheRead: row.cacheRead as number, cacheWrite: row.cacheWrite as number },
708 multiplier,
709 }
710 }
711 return { available: true, present: modelPricing !== undefined, exact, pricing }
712 } catch {
713 // A host exception can contain settings or credentials; do not echo it.
714 return { available: false, present: false, exact: false }
715 }
716}
717
718function samePricing(a: Pricing | undefined, b: Pricing): boolean {
719 return a !== undefined && a.multiplier === b.multiplier &&
720 rateKeys.every(key => a.rates[key] === b.rates[key])
721}
722
723async function diagnose($: EngineInterface): Promise<Diagnostic> {
724 let model: string
725 try {
726 model = await $.session.model()
727 if (!model) throw new Error('no model')
728 } catch {
729 return { model: '(unavailable)', state: 'pricing unavailable',
730 reason: 'Active model is not exposed by the host.', sources: [] }
731 }
732
733 const policy = await snapshot($, model, { source: 'policy' })
734 const ordinary: Array<{ source: SettingsSource; snapshot: Snapshot }> = []
735 for (const source of ['user', 'project', 'local', 'flag'] as const) {
736 ordinary.push({ source, snapshot: await snapshot($, model, { source }) })
737 }
738 const merged = await snapshot($, model)
739 const sources = [
740 `policy: ${!policy.available ? 'unavailable' : !policy.present ? 'absent' :
741 policy.exact ? 'exact override present' : 'no exact override'}`,
742 ...ordinary.map(({ source, snapshot: s }) =>
743 `${source}: ${!s.available ? 'unavailable' : s.present ? 'modelPricing present (ignored by host pricing)' : 'absent'}`),
744 `merged: ${!merged.available ? 'unavailable' : merged.present ? 'modelPricing present (not proof of acceptance)' : 'absent'}`,
745 ]
746 const base = { model, sources }
747 if (policy.pricing) {
748 return { ...base, state: 'accepted pricing available', pricing: policy.pricing,
749 reason: 'Managed policy exact model override; documented accepted pricing source. Multiplier defaults to 1 only when absent.' }
750 }
751 if (policy.exact) {
752 return { ...base, state: 'pricing unavailable',
753 reason: 'Managed exact override is invalid: four finite rates in 0..10000 and, when supplied, a finite multiplier >0 and <=10 are required. No fallback guessed.' }
754 }
755 if (merged.pricing) {
756 const ignored = ordinary.filter(({ snapshot: s }) => samePricing(s.pricing, merged.pricing!))
757 if (policy.available && ignored.length > 0) {
758 return { ...base, state: 'pricing unavailable',
759 reason: `Merged exact row matches ignored ${ignored.map(s => s.source).join('/')} settings; it is not accepted pricing. SDK fallback is not exposed by this settings interface.` }
760 }
761 return { ...base, state: 'candidate/unconfirmed', pricing: merged.pricing,
762 reason: 'Merged exact row has no verified accepted provenance. Do not use for costs; managed source or host SDK fallback acceptance cannot be established.' }
763 }
764 return { ...base, state: 'pricing unavailable',
765 reason: !policy.available ? 'Managed pricing snapshot unavailable; host SDK fallback is not exposed.' :
766 'No valid managed exact override. Ordinary sources are ignored; host SDK fallback may not appear in settings. Canonical matching is unimplemented.' }
767}
768
769function display(d: Diagnostic, model: string): string {
770 const r = ratesFor(model)
771 const lines = ['Pricing diagnostic (read-only)', `Active model: ${d.model}`,
772 r ? `Cost rates this mod uses: ${RATE_SOURCE} (USD / million tokens)`
773 : `Cost rates this mod uses: none configured for this model — costs show as n/a. Add a row to RATES in hooks/register.ts.`]
774 if (r) {
775 lines.push(` input=${r.input}; output=${r.output}; cacheRead=${r.cacheRead}; cacheWrite=${r.cacheWrite}`)
776 }
777 lines.push(`Desktop-managed pricing: ${d.state}`, `Provenance / reason: ${d.reason}`)
778 if (d.pricing) {
779 const { rates, multiplier } = d.pricing
780 lines.push(`Host-selected rates (USD / million tokens, before multiplier${d.state === 'candidate/unconfirmed' ? '; UNCONFIRMED' : ''}):`,
781 ` input=${rates.input}; output=${rates.output}; cacheRead=${rates.cacheRead}; cacheWrite=${rates.cacheWrite}`,
782 `Multiplier: ${multiplier}`)
783 }
784 lines.push('Source snapshots (pricing fields only):', ...d.sources.map(source => ` ${source}`),
785 'Matching: exact model ID only; canonical matching is unimplemented.',
786 'Behavior: nothing here warms or sends on its own — a model call happens only while a /keepwarm window is armed. No prompt is ever dropped, and no settings are written.',
787 'Prompt Cache / Response Cache observations: not exposed to a plugin by this engine build; the gateway-side response-cache controls are a separate ticket.')
788 return lines.join('\n')
789}
790
791async function refresh($: EngineInterface, s: State): Promise<string> {
792 const d = await diagnose($)
793 await publish($, s)
794 return display(d, d.model)
795}
796
797export const register: Register = on => {
798 const s = freshState()
799
800 on('session.start', async ($, e, next) => {
801 const r = await next(e)
802 s.sid = await $.session.id()
803 // The band is raised on the desktop and terminal surfaces only, so which one
804 // this is decides whether the status line is the fallback or is left empty.
805 s.surface = e.surface ?? ''
806 const now = await $.clock.now()
807 await prune($, s, now)
808 const saved = await $.store.get(deadlineKey(s))
809 const savedEvery = await $.store.get(everyKey(s))
810 const savedGuard = await $.store.get(KEY_GUARD)
811 s.deadline = typeof saved === 'number' && saved > now ? saved : 0
812 s.every = typeof savedEvery === 'number' && savedEvery >= MIN_PING_MS ? savedEvery : PING_AFTER_MS
813 s.guard = savedGuard === 'warn' ? 'warn' : 'warn'
814 const usage = await $.session.usage()
815 if (usage.context.tokens) s.ctx = usage.context.tokens
816 // Registering a name another plugin already owns would double the hook.
817 // While the Original Mod is enabled its names are taken, so say so once.
818 const taken = new Set((await $.command.list()).map(c => c.name))
819 let claimed = false
820 const specs: CommandSpec[] = [
821 { name: 'keepwarm',
822 description: 'Keep the prompt cache warm: bare for 6h, a window such as 90m, always, off, or status (token-tax-weather)',
823 argumentHint: '[6h | always | off | status]', immediate: true },
824 { name: 'cache-tax',
825 description: 'Prompt cache state, cold price and this session\'s cold writes (token-tax-weather)',
826 argumentHint: '[status]', immediate: true },
827 ]
828 for (const spec of specs) {
829 if (taken.has(spec.name)) {
830 $.ui.log(`/token-tax-weather: /${spec.name} is already registered by another plugin, so this mod did not claim it. Disable the Original Mod (cache-tax) to use this mod's /${spec.name}.`)
831 continue
832 }
833 await $.command.register(spec)
834 claimed = true
835 }
836 const diagnostic: CommandSpec = { name: 'cache-tax-diagnostic',
837 description: 'Read-only pricing provenance diagnostic; no inference or warming.', immediate: true }
838 await $.command.register(diagnostic)
839 // The under-composer line has one slot per plugin and the Original Mod also
840 // draws into it, so paint only when this mod owns the session — otherwise
841 // the two footers clobber each other and truncate.
842 s.owns = claimed
843 // The window /autocompact measures against, so the band's percentage is the
844 // one compaction will actually fire on rather than the model's raw limit.
845 await refreshWindow($, s)
846 await publish($, s)
847 // The Desktop surface attaches a moment after this event, so the paint above
848 // can be dropped; paint once more shortly after, and after that the prompts,
849 // turns and commands keep the line current.
850 $.clock.after(3000, () => { void publish($, s) })
851 return r
852 })
853
854 // The band above the prompt. One instance and one plugin: a hook that returns a
855 // tree answers the site and the chain stops beneath it, so a second plugin
856 // hooked here never draws. A survey holds the band while it is up, and nothing
857 // is drawn at all until this mod owns the session's command names — otherwise
858 // this band and the Original Mod's status line would both speak for one session.
859 on('ui.render', { component: 'AbovePrompt' }, ($, e, next) => {
860 if (!s.owns || e.props.hasSurvey || !s.window) return next(e)
861 const { Box, Text } = $.ui.resolve(e)
862 return band(Box, Text, e.props, s)
863 })
864
865 // The resume fields Claude Code computes for settings hooks seed the guard
866 // before any turn of the resumed session has run.
867 on('classic.SessionStart', async ($, e, next) => {
868 const r = await next(e)
869 if (e.source === 'clear') {
870 await stop($, s, null)
871 s.stopped = null
872 resetForClear(s)
873 s.sid = await $.session.id()
874 return r
875 }
876 const line = seedFromResume(s, e, await $.clock.now())
877 // The resume payload may omit the model; without it the guard cannot price the cold write.
878 if (!s.lastModel) s.lastModel = await $.session.model()
879 // A cold resume is a notice on the status line, not a transcript line: it is
880 // read now, and it stays up until the first turn finishes.
881 if (line) s.notice = line
882 // The seeded clock decides whether a restored window pings before the first turn: never when it is cold.
883 await arm($, s)
884 // arm() only paints when a window is armed; a bare resume still has to show.
885 if (line) await publish($, s)
886 return r
887 })
888
889 on('command.run', { command: 'keepwarm' }, async ($, e) => {
890 const words = String(e.args ?? '').trim().split(/\s+/).filter(Boolean)
891 const now = await $.clock.now()
892 if (words[0] === 'off') {
893 await stop($, s, null)
894 return { text: 'keepwarm is off' }
895 }
896 if (words[0] === 'always') {
897 // Kept for parity with the Original Mod, but nothing arms itself here:
898 // `always` only pre-arms this session, it is not remembered across them.
899 await startWindow($, s, DEFAULT_WINDOW_MS, PING_AFTER_MS)
900 return { text: `keepwarm on for ${fmtDuration(DEFAULT_WINDOW_MS)} this session. This mod never re-arms itself at session start, so run /keepwarm again in a new session.` }
901 }
902 if (!words.length) {
903 await startWindow($, s, DEFAULT_WINDOW_MS, PING_AFTER_MS)
904 return { text: armedText(s, now, DEFAULT_WINDOW_MS) }
905 }
906 if (words[0] !== 'status') {
907 const window = parseDuration(words[0] ?? '')
908 if (window == null) return { text: 'keepwarm takes a window such as 6h or 90m, or always, off, or status' }
909 let every = PING_AFTER_MS
910 if (words[1] === 'every') {
911 const period = parseDuration(words[2] ?? '')
912 if (period == null || period < MIN_PING_MS) return { text: 'every takes a period of at least 1m' }
913 every = period
914 }
915 await startWindow($, s, window, every)
916 return { text: armedText(s, now, window) }
917 }
918 return { text: s.deadline || s.stopped ? statusText(s, now) : 'keepwarm is off' }
919 })
920
921 on('command.run', { command: 'cache-tax' }, async ($, e) => {
922 const words = String(e.args ?? '').trim().split(/\s+/).filter(Boolean)
923 const now = await $.clock.now()
924 if (words[0] === 'guard') {
925 return { text: 'guard is warn only. This mod shows the cold price and sends your prompt anyway; it never drops one.' }
926 }
927 return { text: card(s, now) }
928 })
929
930 on('command.run', { command: 'cache-tax-diagnostic' }, async $ => {
931 return { text: await refresh($, s) }
932 })
933
934 // The message that pays. Only its first character is read.
935 on('prompt.submit', async ($, e, next) => {
936 if (e.origin.kind === 'plugin') return next(e)
937 if (typeof e.text !== 'string' || e.text.trimStart().startsWith('/')) return next(e)
938 const now = await $.clock.now()
939 if (isCold(s, now) && s.ctx >= BIG_TOKENS) {
940 // Warn only: the prompt always goes through, whatever the guard says. The
941 // warning is a notice on the status line, not a transcript line, and it
942 // stays up while the turn it belongs to runs.
943 s.notice = `${guardText(s, now)} Sending anyway.`
944 s.coldWritePending = true
945 }
946 // The surface may still have been attaching when session start painted, so
947 // repaint here: by the first prompt the slot is certainly live.
948 await publish($, s)
949 return next(e)
950 })
951
952 on('turn.step', async function* ($, e, next) {
953 if (!e.agentId) s.lastRequestAt = await $.clock.now()
954 yield* next(e)
955 })
956
957 on('turn.complete', async ($, e, next) => {
958 const r = await next(e)
959 if (e.agentId) return r
960 const now = await $.clock.now()
961 // A sleeping host may deliver this turn before the expired window's timer.
962 if (s.deadline && now >= s.deadline) await stop($, s, null)
963 // turn.step stamps the exact request time; when no step of this turn did, the turn's end is the floor.
964 if (now - s.lastRequestAt > e.durationMs) s.lastRequestAt = now
965 s.compacted = false
966 const u = e.usage
967 // The notice this turn raises, if any. It is set after the clear below, so a
968 // cold write — a call to action — outlives its own turn by one, unlike the
969 // notice from the prompt or the resume it supersedes.
970 let notice: string | null = null
971 if (u) {
972 if (u.model) s.lastModel = u.model
973 const prev = s.ctx
974 // A turn's usage is its responses summed, so a ten-step turn reports ten
975 // reads of the context. The live window is the engine's figure; the sum
976 // is only the fallback for a host that reports no tokens.
977 const write = u.cache_creation_input_tokens
978 const live = (await $.session.usage()).context.tokens
979 s.ctx = live && live > 0 ? live : u.input_tokens + u.cache_read_input_tokens + write
980 // What the cache actually served this turn — the honest "warmed" figure,
981 // as against the context size, which is only what it could be holding.
982 s.lastRead = u.cache_read_input_tokens
983 const full = prev > 20000 && write >= 0.5 * prev
984 if (full || s.coldWritePending) {
985 const rates = ratesFor(s.lastModel)
986 const usd = rates ? write * rates.cacheWrite / 1e6 : null
987 s.misses.push({ at: now, tokens: write, usd })
988 // No warming is armed here: a cold write is recorded and reported, never
989 // answered with spending. /keepwarm is the only thing that arms a window.
990 notice = `cold write of ${fmtTok(write)} tokens${usd == null ? '' : ` (${fmtUsd(usd)}, ${RATE_SOURCE})`} recorded. Run /keepwarm to hold the cache; this mod will not arm it for you.`
991 }
992 }
993 s.coldWritePending = false
994 // The band's history: one reading per completed turn, oldest first.
995 s.history.push(s.ctx)
996 if (s.history.length > HISTORY) s.history = s.history.slice(-HISTORY)
997 // A stop notice is a one-shot: it stays until the turn after it. This turn is
998 // also what reveals which model the engine is really using, so the idle footer
999 // is repainted here; arm() repaints an armed one itself.
1000 s.stopped = null
1001 // A notice is a one-shot too: it stays up while the turn it belongs to runs,
1002 // and the ordinary readout returns once that turn is over. Cleared here, at the
1003 // end of the turn, before the paints below; the cold write this turn raised is
1004 // the one thing set after the clear, so it stays until the next turn.
1005 s.notice = null
1006 if (notice) s.notice = notice
1007 await refreshWindow($, s)
1008 await arm($, s)
1009 if (!s.deadline) await publish($, s)
1010 return r
1011 })
1012
1013 on('session.compact', async ($, e, next) => {
1014 const r = await next(e)
1015 if (!e.agentId) {
1016 s.compacted = true
1017 s.ctx = 0
1018 disarm(s)
1019 }
1020 return r
1021 })
1022}
1023