SLOPSHOPPER

token-tax-weather

Context weather and cache tax above the prompt: a coloured forecast of how full the window /autocompact measures against is, beside what the cache does to the…

newbandcommandstatuspromptmodel
v0.4.0no licenseupdated 2026-10-10gelwuyl/token-tax-weather/token-tax-weather
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · token-tax-weather
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /keepwarm ⎿ token-tax-weather: keepwarm on for 6h00m, a ping 50m after each idle stretch keeps the cache read, not re-written ☁ Cloudy 49% of 200k · ▄ · cached 91k · 6h00m left · ping in 50m · +$0.0004 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Band
☁ Cloudy 49% of 200k · ▄ · cached 91k · 6h00m left · ping in 50m · +$0.0004
README

token-tax-weather

Two questions about a Claude Code session, answered in one coloured band above the prompt:

  • How full is the context window? A forecast — ☀ Clear, ☁ Cloudy, ☂ Showers, ☇ Storm, ↯ Compact soon — measured against the window /autocompact actually fires on, not the model's raw limit.
  • What is the prompt cache doing to the bill? A signed figure, green when being warm is saving you money and red when the cache-read rate is above the plain input rate and it is costing you instead.
☂  Showers  55% of 362k  ·  ▃▄▅▅▆▇  ·  ▲ +12k last turn  ·  cached 141.9k  ·  6h00m left · ping in 50m  ·  -$0.12

To the left is the weather; to the right, the tax.

Install

/plugin marketplace add <owner>/<repo>
/plugin install token-tax-weather@token-tax-weather

Or, with the claude CLI:

claude plugin marketplace add <owner>/<repo>
claude plugin install token-tax-weather@token-tax-weather

The band is raised on the desktop and terminal surfaces. On the surfaces that do not raise it the same facts go to the status line instead, without colour, so the two never both draw the readout.

Notices

The band carries the continuously-updating facts: the forecast, the cache read, the countdown, the price. The one-off notices are different. A cold resume, a cold send the guard warns about, and a cold write it records are pinned on the status line under the prompt — not in the transcript — so they are read once where you are working, rather than scrolled back to. A notice stays up while the turn it belongs to runs and clears when that turn finishes; a second notice simply replaces the first. Unlike the readout, a notice is shown on every surface, the band's own included.

Commands

CommandWhat it does
/keepwarmArms a six-hour window that pings the cache so it does not expire
/keepwarm 90m, 2h30m, 6hArms a window of that length
/keepwarm every 50mArms it, pinging every 50 minutes
/keepwarm alwaysArms it with no deadline, this session
/keepwarm offDisarms it
/keepwarm statusReports what is armed
/cache-taxThe full card: state, window, prices, the cache's effect, break-even
/cache-tax-diagnosticRead-only. What the host's pricing interface exposes, and nothing else

What it never does

  • It never drops a prompt. The guard warns that the cache went cold and sends your message anyway.
  • It never warms on its own. No timer arms itself; /keepwarm is the only thing that spends anything, and it is always yours to run.
  • It never invents a price. Costs come from the rate table below, which you edit. A model with no row shows n/a rather than a number from another model.

The rates

Unit prices in USD per million tokens, in hooks/register.ts:

export const RATES: Record<string, Rates> = {
  'claude-fable-5': { input: 0.43, cacheRead: 0.03, cacheWrite: 0, output: 0.87 },
  'claude-opus-5-5': { input: 0.0007, cacheRead: 0.005, cacheWrite: 0, output: 0.6 },
}

Matched on the engine's exact model ID, not the name the picker shows. Add a row for each model you use; nothing else needs changing.

Reading the figure

The final number on the band is context × (cacheRead − input) — what the cache does to the bill for that context, not an indicator of whether the cache is currently warm. It is decided entirely by the model's rates and the context size, so it is the same number whether the cache is hot or has been cold for a day.

  • Negative (green): the read rate is below the input rate, so being warm lowers the bill by that much. The usual case.
  • Positive (red): the read rate is above the input rate, so serving from cache costs more than not caching and losing the cache would be the saving.

Whether the cache is warm right now is the <time> left field, not this one.

Privacy

Everything it reads is local session state: token counts the engine already reports, and the context breakdown it computes itself with no network request. It sends nothing anywhere, and /cache-tax-diagnostic prints no settings values — only whether the host's pricing interface exposes any.

Development

claude plugin test ./token-tax-weather
claude plugin validate ./token-tax-weather

The tests mount the real surface and read the drawn tree, so a tree the surface would refuse fails the suite rather than passing a stub. The engine writes this build's API declarations into .claude-plugin/types/ when it loads the plugin; that folder is generated and gitignored, and tsc -p token-tax-weather needs a load first.

Relation to cache-tax

This supersedes cache-tax 2.1.3. It registers /keepwarm and /cache-tax under the same names, so run only one of the two — whichever loads second skips the names already taken and says so.

Source 1 files
hooks/register.ts 1023 lines
1import type {
2  BoxProps,
3  CommandSpec,
4  ElementConstructor,
5  EngineInterface,
6  Register,
7  RenderElement,
8  SettingsReadArgs,
9  SettingsSource,
10  TextProps,
11} from 'claude-code'
12
13// token-tax-weather — supersedes the Original Mod (cache-tax 2.1.3).
14//
15// Ported from the Original Mod with six deliberate changes:
16//   1. the hard-coded first-party PRICES table is replaced by user-supplied rates
17//   2. uncached input is its own rate, not half the cache-write rate
18//   3. the guard warns only; no prompt is ever dropped
19//   4. warming is never armed automatically, only by an explicit command
20//   5. the status line is repainted as state changes — session start, every
21//      turn, every command — instead of being written once and left to go stale
22//   6. output never repeats the plugin name the host already prints
23
24const TTL_MS = 60 * 60 * 1000
25const PING_AFTER_MS = 50 * 60 * 1000
26const MIN_PING_MS = 60 * 1000
27const DEFAULT_WINDOW_MS = 6 * 60 * 60 * 1000
28const BIG_TOKENS = 50000
29const PING_PROMPT = 'Reply with the single word: warm'
30
31// Namespaced so the Custom Mod never inherits the Original Mod's window, ping
32// period, or `always` flag — inheriting `always` would arm paid warming on its
33// own at session start.
34const KEY_DEADLINE = 'token-tax-weather:deadline'
35const KEY_EVERY = 'token-tax-weather:every'
36const KEY_GUARD = 'token-tax-weather:guard'
37
38const rateKeys = ['input', 'cacheRead', 'cacheWrite', 'output'] as const
39type Rates = { input: number; cacheRead: number; cacheWrite: number; output: number }
40
41// USER-SUPPLIED rates in USD per million tokens, keyed by the exact model ID the
42// engine reports. Desktop does not hand its managed four rates to a plugin, so
43// every cost below is only as right as this table, and is always labeled
44// user-supplied. Zero is a valid rate — cacheWrite is zero here, which means a
45// cold re-write is billed at nothing and break-even reports not-applicable. A
46// model with no row shows an explicit unknown cost rather than borrowing
47// another model's rates.
48//
49// These are the operator's own figures, moved between models on request. On
50// 2026-10-10 the Fable row took over the numbers first given for Opus, and Opus
51// took the second set. Both are stated per million tokens; if an operator quotes
52// a provider's per-1K numbers instead, the whole table is 1000x out.
53//
54// Note on the Opus row: its cacheRead (0.005) is about 7x its input (0.0007), so
55// reading the prompt cache costs more than sending the same context uncached.
56// The card reports what the arithmetic says; it does not pretend the figures are
57// typical.
58//
59// Add rows in this shape, then update the plugin:
60//
61//   'claude-opus-5-5': { input: 0, cacheRead: 0, cacheWrite: 0, output: 0 },
62//
63export const RATES: Record<string, Rates> = {
64  'claude-fable-5': { input: 0.43, cacheRead: 0.03, cacheWrite: 0, output: 0.87 },
65  'claude-opus-5-5': { input: 0.0007, cacheRead: 0.005, cacheWrite: 0, output: 0.6 },
66}
67
68const RATE_SOURCE = 'user-supplied rates'
69
70type PingRecord = { at: number; read: number; write: number; usd: number | null; warm: boolean }
71type Miss = { at: number; tokens: number; usd: number | null }
72type GuardMode = 'warn'
73
74export type State = {
75  sid: string
76  deadline: number
77  every: number
78  lastRequestAt: number
79  lastModel: string | null
80  ctx: number
81  compacted: boolean
82  guard: GuardMode
83  coldWritePending: boolean
84  misses: Miss[]
85  pending: { cancel: () => void } | null
86  last: PingRecord | null
87  stopped: string | null
88  /** Cache-read tokens of the last completed main turn: what the cache actually served. */
89  lastRead: number
90  /** True once this mod's command names are the session's, so the footer slot is ours to write. */
91  owns: boolean
92  /** The window the fill is measured against: the compaction window /autocompact sets, else the model's. */
93  window: number
94  /** The fill as a whole percentage of `window`. */
95  percent: number
96  /** Where that window came from (the engine's `ContextWindowSource`), or null when only the model's is known. */
97  windowSource: string | null
98  /** Recent readings of the context size, oldest first: the sparkline and the per-turn trend. */
99  history: number[]
100  /**
101   * The clock as of the last event that could change what the band says. A
102   * `ui.render` hook draws synchronously and `$.clock.now()` only answers a
103   * promise, so the countdown reads this rather than asking again; every event
104   * that repaints stamps it first.
105   */
106  now: number
107  /** The surface this session runs on: the band is raised on `desktop` and `terminal` only. */
108  surface: string
109  /**
110   * The one-off line pinned on the status line ahead of the readout, or null for
111   * none. It carries a cache-state notice — a cold resume, a cold send, a recorded
112   * cold write — rather than the continuously-updating facts, so it is read once
113   * and not scrolled back to. A second notice simply replaces the first: last
114   * writer wins, which is what the status line itself does.
115   */
116  notice: string | null
117}
118
119/** The configured rates for a model, or null when none are configured. Exact ID only. */
120export function ratesFor(model: string | null): Rates | null {
121  if (!model) return null
122  const row = RATES[model] ?? RATES[model.toLowerCase()]
123  if (!row) return null
124  const ok = rateKeys.every(k => typeof row[k] === 'number' && Number.isFinite(row[k]) && row[k] >= 0)
125  return ok ? row : null
126}
127
128export function parseDuration(text: string): number | null {
129  const m = /^(?:(\d+)h)?(?:(\d+)m)?$/.exec(text.trim())
130  if (!m || (m[1] === undefined && m[2] === undefined)) return null
131  return (Number(m[1] ?? 0) * 60 + Number(m[2] ?? 0)) * 60 * 1000
132}
133
134export function fmtDuration(ms: number): string {
135  const total = Math.max(0, Math.round(ms / 60000))
136  const h = Math.floor(total / 60)
137  const m = total % 60
138  if (h >= 48) return `${Math.floor(h / 24)}d ${h % 24}h`
139  return h > 0 ? `${h}h${String(m).padStart(2, '0')}m` : `${m}m`
140}
141
142function fmtUsd(usd: number | null): string {
143  if (usd == null) return 'n/a'
144  // Sub-cent figures are normal at these rates; two decimals would read as $0.00.
145  if (usd > 0 && usd < 0.01) return '$' + usd.toFixed(4)
146  return '$' + (usd >= 100 ? usd.toFixed(0) : usd.toFixed(2))
147}
148
149/** A signed difference: `+$0.12` costs more, `-$0.00` saves. */
150function fmtSignedUsd(usd: number | null): string {
151  if (usd == null) return 'n/a'
152  return (usd >= 0 ? '+' : '-') + fmtUsd(Math.abs(usd))
153}
154
155/**
156 * Rounding to whole thousands hid the difference the line exists to show: a 300k
157 * context served from a 299k cache and a 300k one served from 141k read the same.
158 * One decimal for the thousands and millions keeps the two figures apart, and the
159 * trailing `.0` is dropped so an exact thousand still reads as a round number.
160 */
161function fmtTok(n: number): string {
162  if (n >= 1e6) return trimZero((n / 1e6).toFixed(1)) + 'M'
163  if (n >= 1000) return trimZero((n / 1000).toFixed(1)) + 'k'
164  return String(n)
165}
166
167function trimZero(s: string): string {
168  return s.endsWith('.0') ? s.slice(0, -2) : s
169}
170
171/** ctx × the cache-write rate: what re-writing the context after expiry is billed. */
172function coldUsd(s: State): number | null {
173  const r = ratesFor(s.lastModel)
174  return r ? s.ctx * r.cacheWrite / 1e6 : null
175}
176
177/** ctx × the cache-read rate: what a turn whose context is served from cache costs. */
178function warmUsd(s: State): number | null {
179  const r = ratesFor(s.lastModel)
180  return r ? s.ctx * r.cacheRead / 1e6 : null
181}
182
183/** ctx × the uncached input rate: what the same context costs with no cache reuse at all. */
184function uncachedUsd(s: State): number | null {
185  const r = ratesFor(s.lastModel)
186  return r ? s.ctx * r.input / 1e6 : null
187}
188
189/**
190 * The adjustment the cache makes to the bill for this one context: from cache
191 * minus no cache, i.e. the cache-read rate against the input rate. **Negative
192 * means the cache lowers the bill** — it is a saving of that much, and the usual
193 * case where the read rate is far below the input rate (Fable: 0.03 against
194 * 0.43). **Positive means the cache raises it** — the read rate is above the
195 * input rate, so serving from cache costs more than not caching and the cache is
196 * a surcharge, which some operators' rates genuinely do (Opus: 0.005 against
197 * 0.0007).
198 *
199 * This is the reverse of "what a cold send costs": that is the same magnitude
200 * with the opposite sign. The sign here reads as an effect on the bill, so with
201 * a colour beside it negative is green and positive is red.
202 */
203function cacheEffectUsd(s: State): number | null {
204  const uncached = uncachedUsd(s)
205  const warm = warmUsd(s)
206  return uncached == null || warm == null ? null : warm - uncached
207}
208
209/** Everything a fork bills: the cache read, any cache write, uncached input at its own rate, and the output. */
210function pingUsd(u: { input_tokens: number; output_tokens: number; cache_read_input_tokens: number; cache_creation_input_tokens: number }, r: Rates): number {
211  return (u.cache_read_input_tokens * r.cacheRead + u.cache_creation_input_tokens * r.cacheWrite + u.input_tokens * r.input + u.output_tokens * r.output) / 1e6
212}
213
214/** The read-only upper bound: pings at the cache-read rate that cost as much as one cold write of the same context. */
215function breakEvenPings(s: State): number | null {
216  const r = ratesFor(s.lastModel)
217  if (!r || s.ctx <= 0 || r.cacheRead <= 0) return null
218  return Math.floor(r.cacheWrite / r.cacheRead)
219}
220
221function isCold(s: State, now: number): boolean {
222  return s.lastRequestAt > 0 && !s.compacted && now - s.lastRequestAt >= TTL_MS
223}
224
225function guardText(s: State, now: number): string {
226  const r = ratesFor(s.lastModel)
227  const rate = r ? `$${r.cacheWrite}/MTok` : 'the cache-write rate'
228  const warm = warmUsd(s)
229  const basis = r ? ` (${RATE_SOURCE})` : ' (no rates configured for this model)'
230  return `the prompt cache went cold ${fmtDuration(now - s.lastRequestAt - TTL_MS)} ago. Sending this re-writes ` +
231    `up to ${s.ctx.toLocaleString('en-US')} tokens at ${rate} = ${fmtUsd(coldUsd(s))}` +
232    (warm == null ? '' : ` (a warm turn would have cost ${fmtUsd(warm)})`) + basis + '.'
233}
234
235export type ResumeFields = {
236  source: string
237  model?: string
238  context_tokens?: number
239  seconds_since_last_response?: number
240  prompt_cache_likely_expired?: boolean
241  estimated_cache_write_usd?: number
242}
243
244/** /clear starts a new conversation in the same process; nothing priced before it still exists. */
245export function resetForClear(s: State) {
246  s.ctx = 0
247  s.lastRequestAt = 0
248  s.compacted = false
249  s.coldWritePending = false
250  s.misses = []
251  s.lastRead = 0
252  s.percent = 0
253  s.history = []
254  s.notice = null
255  disarm(s)
256}
257
258/** Applies a resumed session's fields to the state; returns the line to log, if any. */
259export function seedFromResume(s: State, e: ResumeFields, now: number): string | null {
260  if (e.source !== 'resume' && e.source !== 'fork') return null
261  if (typeof e.context_tokens === 'number' && e.context_tokens > 0) s.ctx = e.context_tokens
262  if (typeof e.seconds_since_last_response === 'number') s.lastRequestAt = now - e.seconds_since_last_response * 1000
263  if (typeof e.model === 'string') s.lastModel = e.model
264  s.compacted = false
265  if (e.prompt_cache_likely_expired !== true || s.ctx < BIG_TOKENS) return null
266  const usd = typeof e.estimated_cache_write_usd === 'number' ? fmtUsd(e.estimated_cache_write_usd) : fmtUsd(coldUsd(s))
267  return `resuming cold. The first message re-writes ${s.ctx.toLocaleString('en-US')} tokens, about ${usd}. /clear and paste a summary if you only need the conclusions.`
268}
269
270/** The warming slot on its own: off, a countdown, or why it stopped. */
271function warmingText(s: State, now: number): string {
272  if (s.stopped) return `stopped: ${s.stopped}`
273  if (!s.deadline) return 'off'
274  return `${fmtDuration(s.deadline - now)} left`
275}
276
277/** When the next ping can come, or null while nothing is armed. */
278function nextPingText(s: State, now: number): string | null {
279  if (!s.deadline) return null
280  if (!s.lastRequestAt) return 'ping after first turn'
281  if (isCold(s, now)) return 'ping after next turn'
282  return `ping in ${fmtDuration(s.lastRequestAt + s.every - now)}`
283}
284
285// The under-composer line, in one fixed field order so the same fact always
286// appears in the same place — durations first, then token counts, then money:
287//
288//   <warming> · [ping] · <context> · [cached] · [±]
289//
290// warming  off | <time> left | stopped: <why>
291// ping     present only while a window is armed
292// context  no context yet | ctx <tokens> — what the cache could be holding
293// cached   absent until a turn reports a cache read: what it actually served
294// final    absent until a context exists: what the cache does to the bill for
295//          this context, bare and signed. `-` is a saving — the read rate is
296//          below the input rate, so being warm lowers the bill by that much (the
297//          usual case); `+` is a surcharge — the read rate is above the input
298//          rate, so the cache costs more than it saves. `n/a` while this mod has
299//          no rates for the model in use, so a figure it cannot compute is never
300//          shown as zero. The sign is left to carry the meaning on its own; the
301//          wording that would explain it is what the card is for.
302//
303// The model itself is left out: the picker beside this line already names it, and
304// `/cache-tax` prints the exact ID the rates are keyed on.
305function statusText(s: State, now: number): string {
306  const parts = [warmingText(s, now)]
307  const ping = nextPingText(s, now)
308  if (ping) parts.push(ping)
309  parts.push(s.ctx > 0 ? `ctx ${fmtTok(s.ctx)}` : 'no context yet')
310  if (s.lastRead > 0) parts.push(`cached ${fmtTok(s.lastRead)}`)
311  if (s.ctx > 0) {
312    const effect = cacheEffectUsd(s)
313    parts.push(effect == null ? 'n/a' : fmtSignedUsd(effect))
314  }
315  return parts.join(' · ')
316}
317
318function disarm(s: State) {
319  if (s.pending) s.pending.cancel()
320  s.pending = null
321}
322
323// --- The AbovePrompt band ----------------------------------------------------
324// The status line takes a plain string and cannot colour, so once this mod owns
325// the session the same facts move to the band above the prompt, where a tree's
326// `Text` does take a colour. The band is one instance and one plugin draws it:
327// a hook that returns a tree answers the site and the chain stops beneath it, so
328// a second plugin hooked on AbovePrompt never draws.
329
330/** Readings the sparkline shows, oldest first. */
331const HISTORY = 12
332const SPARK = '▁▂▃▄▅▆▇█'
333
334/**
335 * The forecast, by share of the window used. The colours are literal rather than
336 * theme keys on purpose: this is a ramp the person reads as a ramp, and it stays
337 * the same ramp whatever the theme does to `warning` or `error`.
338 */
339const FORECAST = [
340  { upTo: 25, icon: '☀', word: 'Clear', color: 'yellow' },
341  { upTo: 50, icon: '☁', word: 'Cloudy', color: 'cyan' },
342  { upTo: 75, icon: '☂', word: 'Showers', color: 'blue' },
343  { upTo: 90, icon: '☇', word: 'Storm', color: 'magenta' },
344  { upTo: Infinity, icon: '↯', word: 'Compact soon', color: 'red' },
345]
346
347function forecastFor(percent: number) {
348  return FORECAST.find(f => percent < f.upTo) ?? FORECAST[FORECAST.length - 1]!
349}
350
351/**
352 * One bar per reading, its height the share of the window it filled — not the
353 * share of the busiest reading shown, which is what token-weather scales by and
354 * which makes the same bar mean a different number of tokens in every window.
355 * Scaling to the window keeps a bar comparable with the percentage beside it.
356 */
357function sparkline(history: number[], window: number): string {
358  if (!window || !history.length) return ''
359  return history
360    .map(t => {
361      const level = Math.floor((t / window) * (SPARK.length - 1))
362      return SPARK[Math.min(SPARK.length - 1, Math.max(0, level))]!
363    })
364    .join('')
365}
366
367/** How much the context moved on the last turn, or null while there is no turn to compare. */
368function trendText(history: number[]): string | null {
369  if (history.length < 2) return null
370  const delta = history[history.length - 1]! - history[history.length - 2]!
371  if (delta === 0) return 'steady'
372  return `${delta > 0 ? '▲ +' : '▼ '}${fmtTok(Math.abs(delta))} last turn`
373}
374
375/**
376 * The window the fill is measured against, and the fill itself. `rawMaxTokens` is
377 * what /autocompact actually fires against — the model's limit, or the smaller
378 * compaction window when a setting names one — so it is the honest denominator;
379 * `context.window` is only the model's and would understate the fill whenever
380 * autocompact has set something smaller, which is the case that matters.
381 *
382 * `breakdown: 'summary'` estimates locally and sends no request. `full` would
383 * send one count request per tool and per memory file, which the standing
384 * local-only constraint rules out.
385 */
386async function refreshWindow($: EngineInterface, s: State) {
387  try {
388    const usage = await $.session.usage({ breakdown: 'summary' })
389    const context = usage.context
390    const breakdown = context.breakdown
391    if (breakdown && breakdown.rawMaxTokens > 0) {
392      s.window = breakdown.rawMaxTokens
393      s.percent = Math.round(breakdown.percentage)
394      s.windowSource = breakdown.autocompactSource
395      return
396    }
397    if (context.window > 0) {
398      s.window = context.window
399      s.percent = Math.round(context.percent ?? (s.ctx / context.window) * 100)
400      s.windowSource = null
401    }
402  } catch {
403    // No session bound, or a host that answers no breakdown: keep the last reading.
404  }
405}
406
407/**
408 * The band's tree. Sized to `columns` (the band's own box, which is narrower than
409 * the viewport while a pane is docked beside the transcript): the sparkline and
410 * the trend are the first things to go, then the cache read.
411 *
412 * The field order groups like kinds — the forecast and its measurement, then the
413 * token counts, then the durations, then the one price.
414 */
415function band(
416  Box: ElementConstructor<BoxProps>,
417  Text: ElementConstructor<TextProps>,
418  props: { bodyColumns: number },
419  s: State,
420) {
421  const columns = props.bodyColumns || 80
422  const f = forecastFor(s.percent)
423  const now = s.now || 0
424  // The table's constructors take children among the props and hand back an
425  // element with them beside `props`, which is the shape the engine draws.
426  const t = (textProps: TextProps, text: string) => Text({ ...textProps, children: text })
427  const dim = (text: string) => t({ dimColor: true }, text)
428  const parts: RenderElement[] = [
429    t({ color: f.color, bold: true }, `${f.icon}  ${f.word}`),
430    t({ color: f.color }, `  ${s.percent}% of ${fmtTok(s.window)}`),
431  ]
432  const spark = columns >= 56 ? sparkline(s.history, s.window) : ''
433  if (spark) parts.push(dim('  ·  '), t({ color: f.color }, spark))
434  const trend = columns >= 72 ? trendText(s.history) : null
435  if (trend) parts.push(dim(`  ·  ${trend}`))
436  if (s.lastRead > 0) parts.push(dim(`  ·  cached ${fmtTok(s.lastRead)}`))
437  parts.push(dim(`  ·  ${warmingText(s, now)}`))
438  const ping = nextPingText(s, now)
439  if (ping) parts.push(dim(` · ${ping}`))
440  if (s.ctx > 0) {
441    const effect = cacheEffectUsd(s)
442    if (effect == null || effect === 0) {
443      // `n/a` is a price this mod cannot compute, and a configured zero is a real
444      // answer; neither is a saving, so neither is coloured as one.
445      parts.push(dim(`  ·  ${fmtSignedUsd(effect)}`))
446    } else {
447      // Negative lowers the bill (green), positive raises it (red). The theme's
448      // own keys, so the colours follow the person's theme rather than a guess.
449      parts.push(dim('  ·  '), t({ color: effect < 0 ? 'success' : 'error' }, fmtSignedUsd(effect)))
450    }
451  }
452  return Box({ flexDirection: 'row', paddingX: 1, children: parts })
453}
454
455// The line this mod keeps in its own slot under the prompt. The continuously
456// updating readout — what is armed when a window is, and otherwise what the
457// costing is based on — so the slot says something true about this mod rather
458// than sitting empty. It names the model the engine reports, because that exact
459// ID decides whether rates are found.
460//
461// The host already prints the plugin's own name ahead of this text, so the line
462// never repeats it.
463//
464// The model comes from the last turn's own usage and never from a session lookup:
465// at session start the host may not yet have told the engine which model to use,
466// and a footer must not be left waiting on that. Until the first turn the line
467// says so instead of guessing.
468//
469// Written only while this mod owns the session's command names: the slot holds one
470// entry per plugin, so painting into it while the Original Mod owns the session
471// would clobber that mod's footer. Repainted on session start, before each prompt,
472// and after each turn or command, because a single paint can be lost while the
473// surface is still attaching.
474//
475// The band above the prompt is raised on the desktop and terminal surfaces only,
476// and where it is raised it carries the readout in colour. So the status line is
477// the fallback for the surfaces the band never reaches, and the two never both
478// draw the readout: the same thing is never said twice. A live notice is not the
479// readout — it is pinned wherever it is set, the band's own surfaces included, and
480// it takes the line until the turn that raised it is over.
481const BAND_SURFACES = new Set(['desktop', 'terminal'])
482
483// The one place that writes the status line, so a notice and the readout never
484// race each other for it. A live notice outranks everything and shows on every
485// surface, the band's included: it is a one-off thing to read now, not the facts
486// the band keeps current, so it is worth saying even where the band is up.
487async function publish($: EngineInterface, s: State) {
488  // Stamped whatever happens: the band draws synchronously and reads this rather
489  // than asking the clock again.
490  s.now = await $.clock.now()
491  if (!s.owns) return
492  if (s.notice) {
493    $.ui.status(s.notice)
494    return
495  }
496  // With no notice the band says all of it in colour, so the status line is
497  // cleared rather than left stale — which is also what takes a finished notice
498  // off the screen.
499  if (BAND_SURFACES.has(s.surface)) {
500    $.ui.status(undefined)
501    return
502  }
503  $.ui.status(statusText(s, s.now))
504}
505
506// The window and its ping period belong to the session that armed them, so a
507// second session, or one resumed from another transcript, never inherits them
508// and cannot turn them off. The guard mode stays global.
509function deadlineKey(s: State): string {
510  return `${KEY_DEADLINE}:${s.sid}`
511}
512
513function everyKey(s: State): string {
514  return `${KEY_EVERY}:${s.sid}`
515}
516
517/** Clears this session's own dead window. Other sessions' keys are never touched. */
518async function prune($: EngineInterface, s: State, now: number) {
519  for (const key of [deadlineKey(s)]) {
520    const deadline = await $.store.get(key)
521    if (deadline === undefined) continue
522    if (typeof deadline === 'number' && deadline > now) continue
523    await $.store.delete(key)
524    await $.store.delete(everyKey(s))
525  }
526}
527
528async function stop($: EngineInterface, s: State, why: string | null) {
529  s.deadline = 0
530  s.every = PING_AFTER_MS
531  s.stopped = why
532  disarm(s)
533  await $.store.delete(deadlineKey(s))
534  await $.store.delete(everyKey(s))
535  await publish($, s)
536}
537
538async function arm($: EngineInterface, s: State) {
539  disarm(s)
540  if (!s.deadline) return
541  const now = await $.clock.now()
542  if (now >= s.deadline) return stop($, s, null)
543  // A cold window still needs expiry cleanup, but must not send a model request.
544  if (s.lastRequestAt && !s.compacted && !isCold(s, now)) {
545    const delay = Math.min(s.deadline - now, Math.max(1000, s.lastRequestAt + s.every - now))
546    s.pending = $.clock.after(delay, () => { void ping($, s) })
547  } else {
548    s.pending = $.clock.after(s.deadline - now, () => { void arm($, s) })
549  }
550  await publish($, s)
551}
552
553async function ping($: EngineInterface, s: State) {
554  s.pending = null
555  if (!s.deadline) return
556  const now = await $.clock.now()
557  if (now >= s.deadline) return arm($, s)
558  // A turn in the meantime re-armed the timer; this callback is stale.
559  if (now - s.lastRequestAt < s.every - 1000) return
560  if (isCold(s, now)) return arm($, s)
561  let reply
562  try {
563    reply = await $.model.fork({ prompt: PING_PROMPT })
564  } catch (err) {
565    return stop($, s, `ping failed: ${err instanceof Error ? err.message : String(err)}`)
566  }
567  // This build's fork never answers null: the union is an answer, or a reason.
568  if (!reply) return stop($, s, 'ping not sent')
569  if (!reply.isAnswered) return stop($, s, `ping not sent (${reply.reason})`)
570  const u = reply.usage
571  const r = ratesFor(s.lastModel)
572  // A warm ping reads the prefix and writes only its own message; a write of a tenth of the read or more means the prefix broke.
573  const warm = u.cache_read_input_tokens > 0 && u.cache_creation_input_tokens < 0.1 * u.cache_read_input_tokens
574  const usd = r ? pingUsd(u, r) : null
575  s.last = { at: now, read: u.cache_read_input_tokens, write: u.cache_creation_input_tokens, usd, warm }
576  if (!warm) return stop($, s, `ping read ${fmtTok(u.cache_read_input_tokens)}, wrote ${fmtTok(u.cache_creation_input_tokens)} (${fmtUsd(usd)}): the cache was gone`)
577  s.lastRequestAt = now
578  await arm($, s)
579}
580
581/** The reply to an arming command; on a cold cache it says when the first ping can come. */
582function armedText(s: State, now: number, windowMs: number): string {
583  if (isCold(s, now)) return `keepwarm on for ${fmtDuration(windowMs)}. The cache is cold now, so the first ping comes ${fmtDuration(s.every)} after the next turn`
584  return `keepwarm on for ${fmtDuration(windowMs)}, a ping ${fmtDuration(s.every)} after each idle stretch keeps the cache read, not re-written`
585}
586
587async function startWindow($: EngineInterface, s: State, windowMs: number, every: number) {
588  const now = await $.clock.now()
589  s.every = every
590  if (every === PING_AFTER_MS) await $.store.delete(everyKey(s))
591  else await $.store.set(everyKey(s), every)
592  s.deadline = now + windowMs
593  s.stopped = null
594  await $.store.set(deadlineKey(s), s.deadline)
595  await arm($, s)
596}
597
598/**
599 * Where the window the band measures against came from, said plainly. `settings`
600 * and `env` are a person's own `/autocompact` value; `model-default` is the
601 * model's own limit; `unknown-model` means the engine could not settle one and
602 * the percentage should be read as a guess, which is worth knowing when a
603 * gateway serves a model ID the engine does not recognise.
604 */
605function windowNote(s: State): string {
606  switch (s.windowSource) {
607    case 'settings': return ', set by /autocompact'
608    case 'env': return ', set by the environment'
609    case 'clientdata': return ', set by the client'
610    case 'experiment': return ', set by an experiment'
611    case 'model-default': return ", the model's own window"
612    case 'auto': return ', settled automatically'
613    case 'unknown-model': return ', the model is unknown to the engine, so this is a guess'
614    default: return ', no compaction window reported, so the model window is used'
615  }
616}
617
618function card(s: State, now: number): string {
619  const lines: string[] = []
620  const r = ratesFor(s.lastModel)
621  lines.push(`${s.lastModel ?? 'model not seen yet'}`)
622  if (s.compacted) lines.push('state       reset by compaction, waiting for the first turn')
623  else if (!s.lastRequestAt) lines.push('state       no request yet this session')
624  else if (isCold(s, now)) lines.push(`state       COLD, last request ${fmtDuration(now - s.lastRequestAt)} ago`)
625  else lines.push(`state       warm, ${fmtDuration(s.lastRequestAt + TTL_MS - now)} left`)
626  lines.push(`context     ${s.ctx.toLocaleString('en-US')} tokens, the most the cache can be holding`)
627  lines.push(`window      ${s.window ? `${fmtTok(s.window)} (${s.percent}% used)` : 'not reported yet'}${windowNote(s)}`)
628  // What the cache actually served last turn, rather than what it might hold.
629  if (s.lastRead > 0 && s.ctx > 0) {
630    const share = Math.min(100, Math.round((100 * s.lastRead) / s.ctx))
631    lines.push(`cache read  ${s.lastRead.toLocaleString('en-US')} tokens served from cache last turn (${share}% of context)`)
632  } else if (s.lastRequestAt) {
633    lines.push('cache read  none reported by the last turn, so nothing was served from cache')
634  }
635  if (r) {
636    lines.push(`unit prices input ${r.input} · read ${r.cacheRead} · write ${r.cacheWrite} · output ${r.output} USD/MTok (${RATE_SOURCE}; prices, not costs)`)
637    lines.push(`for ${fmtTok(s.ctx)} tokens   re-write ${fmtUsd(coldUsd(s))} (write rate) · from cache ${fmtUsd(warmUsd(s))} (read rate) · no cache ${fmtUsd(uncachedUsd(s))} (input rate)`)
638    const effect = cacheEffectUsd(s)
639    lines.push(`cache effect  ${fmtSignedUsd(effect)} for this context: "from cache" minus "no cache". Negative is a saving — the read rate is below the input rate, so being warm lowers the bill by that much. Positive is a surcharge — the read rate is above the input rate, so the cache costs more than it saves`)
640  } else {
641    lines.push(`unit prices none configured for this model, so no cost can be shown (n/a)`)
642  }
643  const armed = s.deadline ? `on, ${warmingText(s, now)}` : 'off (/keepwarm to arm it for 6h00m)'
644  lines.push(`keepwarm    ${s.stopped ? `stopped, ${s.stopped}` : armed}`)
645  if (s.last) lines.push(`last ping   read ${fmtTok(s.last.read)} from cache, wrote ${fmtTok(s.last.write)}, ${fmtUsd(s.last.usd)}${s.last.warm ? ' (warm)' : ' (the cache was gone)'}`)
646  const pings = breakEvenPings(s)
647  if (pings != null && pings > 0) lines.push(`break-even  up to ${pings} pings at the read rate cost one cold write, about ${fmtDuration(pings * s.every)} of idle at one ping per ${fmtDuration(s.every)}`)
648  else if (r && r.cacheWrite <= 0) lines.push('break-even  warming cannot pay for itself at these rates: a cold re-write is billed at nothing, so every ping is a net cost')
649  else if (r) lines.push('break-even  not applicable: a zero cache-read rate, or no context yet')
650  lines.push(`guard       warn only, always. This mod never drops a prompt.`)
651  const paid = s.misses.reduce((a, m) => a + (m.usd ?? 0), 0)
652  lines.push(`session     ${s.misses.length} cold write${s.misses.length === 1 ? '' : 's'} paid, ${fmtUsd(paid)}`)
653  return lines.join('\n')
654}
655
656export function freshState(): State {
657  return {
658    sid: '', deadline: 0, every: PING_AFTER_MS, lastRequestAt: 0, lastModel: null, ctx: 0, compacted: false,
659    guard: 'warn', coldWritePending: false, misses: [], pending: null, last: null, stopped: null, lastRead: 0, owns: false,
660    window: 0, percent: 0, windowSource: null, history: [], now: 0, surface: '', notice: null,
661  }
662}
663
664// --- Host pricing provenance -------------------------------------------------
665// Desktop does not expose its managed four rates to a plugin. This section
666// reports what the settings interface actually shows, so the absence of
667// accepted pricing is visible rather than guessed at. It never becomes the
668// source of a cost: costs come from the user-supplied RATES table above.
669
670type Pricing = { rates: Rates; multiplier: number }
671type Snapshot = { available: boolean; present: boolean; exact: boolean; pricing?: Pricing }
672type Diagnostic = {
673  model: string
674  state: 'accepted pricing available' | 'candidate/unconfirmed' | 'pricing unavailable'
675  reason: string
676  pricing?: Pricing
677  sources: string[]
678}
679
680function record(value: unknown): Record<string, unknown> | undefined {
681  return value !== null && typeof value === 'object' && !Array.isArray(value)
682    ? value as Record<string, unknown>
683    : undefined
684}
685
686function validRate(value: unknown): value is number {
687  return typeof value === 'number' && Number.isFinite(value) && value >= 0 && value <= 10000
688}
689
690function validMultiplier(value: unknown): value is number {
691  return typeof value === 'number' && Number.isFinite(value) && value > 0 && value <= 10
692}
693
694// Select pricing fields immediately. Never retain, print, or persist settings.
695async function snapshot($: EngineInterface, model: string, args?: SettingsReadArgs): Promise<Snapshot> {
696  try {
697    const { modelPricing } = await $.settings.read(args)
698    const config = record(modelPricing)
699    const overrides = record(config?.overrides)
700    const exact = overrides !== undefined && Object.prototype.hasOwnProperty.call(overrides, model)
701    const row = exact ? record(overrides?.[model]) : undefined
702    const multiplier = config?.multiplier === undefined ? 1 : config.multiplier
703    let pricing: Pricing | undefined
704    if (row && rateKeys.every(key => validRate(row[key])) && validMultiplier(multiplier)) {
705      pricing = {
706        rates: { input: row.input as number, output: row.output as number,
707          cacheRead: row.cacheRead as number, cacheWrite: row.cacheWrite as number },
708        multiplier,
709      }
710    }
711    return { available: true, present: modelPricing !== undefined, exact, pricing }
712  } catch {
713    // A host exception can contain settings or credentials; do not echo it.
714    return { available: false, present: false, exact: false }
715  }
716}
717
718function samePricing(a: Pricing | undefined, b: Pricing): boolean {
719  return a !== undefined && a.multiplier === b.multiplier &&
720    rateKeys.every(key => a.rates[key] === b.rates[key])
721}
722
723async function diagnose($: EngineInterface): Promise<Diagnostic> {
724  let model: string
725  try {
726    model = await $.session.model()
727    if (!model) throw new Error('no model')
728  } catch {
729    return { model: '(unavailable)', state: 'pricing unavailable',
730      reason: 'Active model is not exposed by the host.', sources: [] }
731  }
732
733  const policy = await snapshot($, model, { source: 'policy' })
734  const ordinary: Array<{ source: SettingsSource; snapshot: Snapshot }> = []
735  for (const source of ['user', 'project', 'local', 'flag'] as const) {
736    ordinary.push({ source, snapshot: await snapshot($, model, { source }) })
737  }
738  const merged = await snapshot($, model)
739  const sources = [
740    `policy: ${!policy.available ? 'unavailable' : !policy.present ? 'absent' :
741      policy.exact ? 'exact override present' : 'no exact override'}`,
742    ...ordinary.map(({ source, snapshot: s }) =>
743      `${source}: ${!s.available ? 'unavailable' : s.present ? 'modelPricing present (ignored by host pricing)' : 'absent'}`),
744    `merged: ${!merged.available ? 'unavailable' : merged.present ? 'modelPricing present (not proof of acceptance)' : 'absent'}`,
745  ]
746  const base = { model, sources }
747  if (policy.pricing) {
748    return { ...base, state: 'accepted pricing available', pricing: policy.pricing,
749      reason: 'Managed policy exact model override; documented accepted pricing source. Multiplier defaults to 1 only when absent.' }
750  }
751  if (policy.exact) {
752    return { ...base, state: 'pricing unavailable',
753      reason: 'Managed exact override is invalid: four finite rates in 0..10000 and, when supplied, a finite multiplier >0 and <=10 are required. No fallback guessed.' }
754  }
755  if (merged.pricing) {
756    const ignored = ordinary.filter(({ snapshot: s }) => samePricing(s.pricing, merged.pricing!))
757    if (policy.available && ignored.length > 0) {
758      return { ...base, state: 'pricing unavailable',
759        reason: `Merged exact row matches ignored ${ignored.map(s => s.source).join('/')} settings; it is not accepted pricing. SDK fallback is not exposed by this settings interface.` }
760    }
761    return { ...base, state: 'candidate/unconfirmed', pricing: merged.pricing,
762      reason: 'Merged exact row has no verified accepted provenance. Do not use for costs; managed source or host SDK fallback acceptance cannot be established.' }
763  }
764  return { ...base, state: 'pricing unavailable',
765    reason: !policy.available ? 'Managed pricing snapshot unavailable; host SDK fallback is not exposed.' :
766      'No valid managed exact override. Ordinary sources are ignored; host SDK fallback may not appear in settings. Canonical matching is unimplemented.' }
767}
768
769function display(d: Diagnostic, model: string): string {
770  const r = ratesFor(model)
771  const lines = ['Pricing diagnostic (read-only)', `Active model: ${d.model}`,
772    r ? `Cost rates this mod uses: ${RATE_SOURCE} (USD / million tokens)`
773      : `Cost rates this mod uses: none configured for this model — costs show as n/a. Add a row to RATES in hooks/register.ts.`]
774  if (r) {
775    lines.push(`  input=${r.input}; output=${r.output}; cacheRead=${r.cacheRead}; cacheWrite=${r.cacheWrite}`)
776  }
777  lines.push(`Desktop-managed pricing: ${d.state}`, `Provenance / reason: ${d.reason}`)
778  if (d.pricing) {
779    const { rates, multiplier } = d.pricing
780    lines.push(`Host-selected rates (USD / million tokens, before multiplier${d.state === 'candidate/unconfirmed' ? '; UNCONFIRMED' : ''}):`,
781      `  input=${rates.input}; output=${rates.output}; cacheRead=${rates.cacheRead}; cacheWrite=${rates.cacheWrite}`,
782      `Multiplier: ${multiplier}`)
783  }
784  lines.push('Source snapshots (pricing fields only):', ...d.sources.map(source => `  ${source}`),
785    'Matching: exact model ID only; canonical matching is unimplemented.',
786    'Behavior: nothing here warms or sends on its own — a model call happens only while a /keepwarm window is armed. No prompt is ever dropped, and no settings are written.',
787    'Prompt Cache / Response Cache observations: not exposed to a plugin by this engine build; the gateway-side response-cache controls are a separate ticket.')
788  return lines.join('\n')
789}
790
791async function refresh($: EngineInterface, s: State): Promise<string> {
792  const d = await diagnose($)
793  await publish($, s)
794  return display(d, d.model)
795}
796
797export const register: Register = on => {
798  const s = freshState()
799
800  on('session.start', async ($, e, next) => {
801    const r = await next(e)
802    s.sid = await $.session.id()
803    // The band is raised on the desktop and terminal surfaces only, so which one
804    // this is decides whether the status line is the fallback or is left empty.
805    s.surface = e.surface ?? ''
806    const now = await $.clock.now()
807    await prune($, s, now)
808    const saved = await $.store.get(deadlineKey(s))
809    const savedEvery = await $.store.get(everyKey(s))
810    const savedGuard = await $.store.get(KEY_GUARD)
811    s.deadline = typeof saved === 'number' && saved > now ? saved : 0
812    s.every = typeof savedEvery === 'number' && savedEvery >= MIN_PING_MS ? savedEvery : PING_AFTER_MS
813    s.guard = savedGuard === 'warn' ? 'warn' : 'warn'
814    const usage = await $.session.usage()
815    if (usage.context.tokens) s.ctx = usage.context.tokens
816    // Registering a name another plugin already owns would double the hook.
817    // While the Original Mod is enabled its names are taken, so say so once.
818    const taken = new Set((await $.command.list()).map(c => c.name))
819    let claimed = false
820    const specs: CommandSpec[] = [
821      { name: 'keepwarm',
822        description: 'Keep the prompt cache warm: bare for 6h, a window such as 90m, always, off, or status (token-tax-weather)',
823        argumentHint: '[6h | always | off | status]', immediate: true },
824      { name: 'cache-tax',
825        description: 'Prompt cache state, cold price and this session\'s cold writes (token-tax-weather)',
826        argumentHint: '[status]', immediate: true },
827    ]
828    for (const spec of specs) {
829      if (taken.has(spec.name)) {
830        $.ui.log(`/token-tax-weather: /${spec.name} is already registered by another plugin, so this mod did not claim it. Disable the Original Mod (cache-tax) to use this mod's /${spec.name}.`)
831        continue
832      }
833      await $.command.register(spec)
834      claimed = true
835    }
836    const diagnostic: CommandSpec = { name: 'cache-tax-diagnostic',
837      description: 'Read-only pricing provenance diagnostic; no inference or warming.', immediate: true }
838    await $.command.register(diagnostic)
839    // The under-composer line has one slot per plugin and the Original Mod also
840    // draws into it, so paint only when this mod owns the session — otherwise
841    // the two footers clobber each other and truncate.
842    s.owns = claimed
843    // The window /autocompact measures against, so the band's percentage is the
844    // one compaction will actually fire on rather than the model's raw limit.
845    await refreshWindow($, s)
846    await publish($, s)
847    // The Desktop surface attaches a moment after this event, so the paint above
848    // can be dropped; paint once more shortly after, and after that the prompts,
849    // turns and commands keep the line current.
850    $.clock.after(3000, () => { void publish($, s) })
851    return r
852  })
853
854  // The band above the prompt. One instance and one plugin: a hook that returns a
855  // tree answers the site and the chain stops beneath it, so a second plugin
856  // hooked here never draws. A survey holds the band while it is up, and nothing
857  // is drawn at all until this mod owns the session's command names — otherwise
858  // this band and the Original Mod's status line would both speak for one session.
859  on('ui.render', { component: 'AbovePrompt' }, ($, e, next) => {
860    if (!s.owns || e.props.hasSurvey || !s.window) return next(e)
861    const { Box, Text } = $.ui.resolve(e)
862    return band(Box, Text, e.props, s)
863  })
864
865  // The resume fields Claude Code computes for settings hooks seed the guard
866  // before any turn of the resumed session has run.
867  on('classic.SessionStart', async ($, e, next) => {
868    const r = await next(e)
869    if (e.source === 'clear') {
870      await stop($, s, null)
871      s.stopped = null
872      resetForClear(s)
873      s.sid = await $.session.id()
874      return r
875    }
876    const line = seedFromResume(s, e, await $.clock.now())
877    // The resume payload may omit the model; without it the guard cannot price the cold write.
878    if (!s.lastModel) s.lastModel = await $.session.model()
879    // A cold resume is a notice on the status line, not a transcript line: it is
880    // read now, and it stays up until the first turn finishes.
881    if (line) s.notice = line
882    // The seeded clock decides whether a restored window pings before the first turn: never when it is cold.
883    await arm($, s)
884    // arm() only paints when a window is armed; a bare resume still has to show.
885    if (line) await publish($, s)
886    return r
887  })
888
889  on('command.run', { command: 'keepwarm' }, async ($, e) => {
890    const words = String(e.args ?? '').trim().split(/\s+/).filter(Boolean)
891    const now = await $.clock.now()
892    if (words[0] === 'off') {
893      await stop($, s, null)
894      return { text: 'keepwarm is off' }
895    }
896    if (words[0] === 'always') {
897      // Kept for parity with the Original Mod, but nothing arms itself here:
898      // `always` only pre-arms this session, it is not remembered across them.
899      await startWindow($, s, DEFAULT_WINDOW_MS, PING_AFTER_MS)
900      return { text: `keepwarm on for ${fmtDuration(DEFAULT_WINDOW_MS)} this session. This mod never re-arms itself at session start, so run /keepwarm again in a new session.` }
901    }
902    if (!words.length) {
903      await startWindow($, s, DEFAULT_WINDOW_MS, PING_AFTER_MS)
904      return { text: armedText(s, now, DEFAULT_WINDOW_MS) }
905    }
906    if (words[0] !== 'status') {
907      const window = parseDuration(words[0] ?? '')
908      if (window == null) return { text: 'keepwarm takes a window such as 6h or 90m, or always, off, or status' }
909      let every = PING_AFTER_MS
910      if (words[1] === 'every') {
911        const period = parseDuration(words[2] ?? '')
912        if (period == null || period < MIN_PING_MS) return { text: 'every takes a period of at least 1m' }
913        every = period
914      }
915      await startWindow($, s, window, every)
916      return { text: armedText(s, now, window) }
917    }
918    return { text: s.deadline || s.stopped ? statusText(s, now) : 'keepwarm is off' }
919  })
920
921  on('command.run', { command: 'cache-tax' }, async ($, e) => {
922    const words = String(e.args ?? '').trim().split(/\s+/).filter(Boolean)
923    const now = await $.clock.now()
924    if (words[0] === 'guard') {
925      return { text: 'guard is warn only. This mod shows the cold price and sends your prompt anyway; it never drops one.' }
926    }
927    return { text: card(s, now) }
928  })
929
930  on('command.run', { command: 'cache-tax-diagnostic' }, async $ => {
931    return { text: await refresh($, s) }
932  })
933
934  // The message that pays. Only its first character is read.
935  on('prompt.submit', async ($, e, next) => {
936    if (e.origin.kind === 'plugin') return next(e)
937    if (typeof e.text !== 'string' || e.text.trimStart().startsWith('/')) return next(e)
938    const now = await $.clock.now()
939    if (isCold(s, now) && s.ctx >= BIG_TOKENS) {
940      // Warn only: the prompt always goes through, whatever the guard says. The
941      // warning is a notice on the status line, not a transcript line, and it
942      // stays up while the turn it belongs to runs.
943      s.notice = `${guardText(s, now)} Sending anyway.`
944      s.coldWritePending = true
945    }
946    // The surface may still have been attaching when session start painted, so
947    // repaint here: by the first prompt the slot is certainly live.
948    await publish($, s)
949    return next(e)
950  })
951
952  on('turn.step', async function* ($, e, next) {
953    if (!e.agentId) s.lastRequestAt = await $.clock.now()
954    yield* next(e)
955  })
956
957  on('turn.complete', async ($, e, next) => {
958    const r = await next(e)
959    if (e.agentId) return r
960    const now = await $.clock.now()
961    // A sleeping host may deliver this turn before the expired window's timer.
962    if (s.deadline && now >= s.deadline) await stop($, s, null)
963    // turn.step stamps the exact request time; when no step of this turn did, the turn's end is the floor.
964    if (now - s.lastRequestAt > e.durationMs) s.lastRequestAt = now
965    s.compacted = false
966    const u = e.usage
967    // The notice this turn raises, if any. It is set after the clear below, so a
968    // cold write — a call to action — outlives its own turn by one, unlike the
969    // notice from the prompt or the resume it supersedes.
970    let notice: string | null = null
971    if (u) {
972      if (u.model) s.lastModel = u.model
973      const prev = s.ctx
974      // A turn's usage is its responses summed, so a ten-step turn reports ten
975      // reads of the context. The live window is the engine's figure; the sum
976      // is only the fallback for a host that reports no tokens.
977      const write = u.cache_creation_input_tokens
978      const live = (await $.session.usage()).context.tokens
979      s.ctx = live && live > 0 ? live : u.input_tokens + u.cache_read_input_tokens + write
980      // What the cache actually served this turn — the honest "warmed" figure,
981      // as against the context size, which is only what it could be holding.
982      s.lastRead = u.cache_read_input_tokens
983      const full = prev > 20000 && write >= 0.5 * prev
984      if (full || s.coldWritePending) {
985        const rates = ratesFor(s.lastModel)
986        const usd = rates ? write * rates.cacheWrite / 1e6 : null
987        s.misses.push({ at: now, tokens: write, usd })
988        // No warming is armed here: a cold write is recorded and reported, never
989        // answered with spending. /keepwarm is the only thing that arms a window.
990        notice = `cold write of ${fmtTok(write)} tokens${usd == null ? '' : ` (${fmtUsd(usd)}, ${RATE_SOURCE})`} recorded. Run /keepwarm to hold the cache; this mod will not arm it for you.`
991      }
992    }
993    s.coldWritePending = false
994    // The band's history: one reading per completed turn, oldest first.
995    s.history.push(s.ctx)
996    if (s.history.length > HISTORY) s.history = s.history.slice(-HISTORY)
997    // A stop notice is a one-shot: it stays until the turn after it. This turn is
998    // also what reveals which model the engine is really using, so the idle footer
999    // is repainted here; arm() repaints an armed one itself.
1000    s.stopped = null
1001    // A notice is a one-shot too: it stays up while the turn it belongs to runs,
1002    // and the ordinary readout returns once that turn is over. Cleared here, at the
1003    // end of the turn, before the paints below; the cold write this turn raised is
1004    // the one thing set after the clear, so it stays until the next turn.
1005    s.notice = null
1006    if (notice) s.notice = notice
1007    await refreshWindow($, s)
1008    await arm($, s)
1009    if (!s.deadline) await publish($, s)
1010    return r
1011  })
1012
1013  on('session.compact', async ($, e, next) => {
1014    const r = await next(e)
1015    if (!e.agentId) {
1016      s.compacted = true
1017      s.ctx = 0
1018      disarm(s)
1019    }
1020    return r
1021  })
1022}
1023