SLOPSHOPPER

warmline

Claude Code's prompt-cache state, drawn above the prompt where no statusline renders: the Desktop app's Code tab.

newbandtimer
★ 3v2.8.0MITupdated 2026-10-06Miguel-Barroso/claude-warmline/plugin
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · warmline
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM warmline cache ? no response yet [ Hide ] ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Band
warmline cache ? no response yet [ Hide ]
README

claude-warmline

English | 日本語 | 繁體中文 | 简体中文

tests version license

Warmline makes Claude Code's prompt-cache state visible.

Claude Code can silently go cold

A working session carries tens or hundreds of thousands of tokens — system prompt, tools, every file it read, every reply — and Claude Code re-sends all of it on every turn. While the prompt cache holds, that context is read back at roughly 0.1× the normal input price. When the cache goes cold, the same context is processed again at about 2×. It goes cold after about an hour of inactivity, and instantly whenever the start of the conversation is rewritten — /compact and auto-compaction both do that. Come back to a big session after lunch and your next message pays the rebuild on all of it, at exactly the moment you wanted results, and nothing in the session tells you so.

Warmline puts the cache state directly in your statusline, and gives you the tools to see what it has been doing historically.

Opus 5 | claude-warmline | ctx 64% (127k) | cache HOT (127k, cold ~11:58) | 5h 78%

No guessing, and no timing-based inference. Since v2.1.251 Claude Code hands the statusline its own prompt_cache object — whether the prefix is warm, which TTL it is on, the second it expires — and warmline prints what that object says.

Then, for everything before this turn:

warmline audit

grades every recorded API turn of a session from the usage Claude Code wrote for it, and --all ranks every session on this machine.

The warmline statusline in five states: cache HOT in green with its expiry time, cache HOT in yellow as the expiry nears, cache COLD in red, and cache off and cache ? dimmed when Claude Code reports no cache data

Install

curl -fsSL https://raw.githubusercontent.com/Miguel-Barroso/claude-warmline/main/install.sh | bash

or with Homebrew, on macOS:

brew install Miguel-Barroso/warmline/warmline

Then:

CommandWhat it does
warmline statuswhat's installed and on
warmline auditthis session, turn by turn
warmline audit --allevery session on this machine, ranked
warmline watchevery session's warmth, live, until ctrl-c
warmline afk enableoptional, at your own risk: type afk and the cache stays warm until you're back

Needs python3 and bash, nothing else. On Windows, run it from Git Bash (with Python 3.7+ installed) or inside WSL. No curl? wget -qO- <same URL> | bash works the same way, and the installer downloads with whichever one it finds. Whichever route you take it is one command: brew install wires the statusline and the Desktop band too, brew upgrade re-wires them, and brew uninstall unwires them again.

Install details, flags, pinning a release, package managers, Windows →

Why warmline

Any statusline can now show you whether the cache is warm. Since v2.1.251 Claude Code hands every statusline script the same prompt_cache object warmline reads — warmline's own gauge is, deliberately, nothing more than what that object says. The live light matters, and it is no longer the point.

Warmline exists for the questions no live light can answer:

When did the cache go cold, how often, why — and did it cost enough to care about?

That is warmline audit: every recorded turn of your history graded from the usage Claude Code wrote for it, each cold event attributed to a cause only where the transcript proves one, and the avoidable premium priced at rates derived from your own sessions. Around the audit sits the full loop — observe, explain, measure, mitigate (keep-warm, wait-for --until-cold, awake) — so a cold event goes from noticed to explained to priced to, when it earns it, prevented.

A statusline tells you what is happening in your session. Warmline tells you whether your context is still being reused — and what it has cost you when it wasn't.

Everything here exists for the 100k+ case. Going cold on 20k tokens is cheap.

From visibility to control

Observe — the cache state in your statusline, live, from Claude Code's own data.

Explain — what HOT, COLD, off and the expiry clock each mean, and which of them you can do anything about.

Measure — warmline audit: the cold events across your history, what caused them where the transcript proves a cause, and an estimate of what they were worth.

Mitigate — warmline keep-warm, if it turns out you need it.

That is a progression, not a feature list. It takes you from "did my cache go cold?" to "how often does this happen?" to "is it costing me enough to care?" to "do I want to do something about it?" — and the honest answer to the last one is often no.

The first three are read-only observation, and they are the point of the project. The fourth is opt-in, off by default, and deliberately bounded — see optional: keep warm.

Observe: the statusline

Opus 5 | claude-warmline | ctx 64% (127k) | cache HOT (127k, cold ~11:58) | 5h 78%
FieldMeaning
ctx 64% (127k)context-window utilization — yellow within 10k tokens of the threshold auto-compaction actually fires at (window - 33000)
cache HOT (127k, cold ~11:58)the cached prefix is warm, a rebuild would re-cache 127k tokens, and it leaves its TTL at the time shown; yellow within 15 minutes of it
cache HOT 5mas above, on the 5-minute TTL — usage credits, an API key, a cloud provider. The 1-hour case is the norm and goes unlabelled
5h 78% / 7d 91%the plan window nearest its cap, hidden below 50%, with its reset time once it turns yellow. Absent on API keys and cloud providers, which have no plan window
cache COLDthe prefix is outside its TTL; the next turn re-caches it
cache offprompt caching is off, or this provider or gateway never reports cache tokens. Nothing here will warm up
cache ?no cache data — Claude Code before v2.1.251, or before the session's first API response

Warmline does not try to work out whether your cache is warm by timing how long Claude Code takes to answer. Every verdict above is a field Claude Code handed the statusline: prompt_cache.warm for the state, caching_observed for whether caching is happening at all, ttl for the bucket, expires_at for the clock, recache_tokens_if_cold for the stake, rate_limits for the plan window. Warmline formats and colors them. It used to infer all of this from the gap between turns, and that machinery is gone.

The two numbers that decide anything are the stake and the quota. (127k) is what a rebuild would cost you — the size of the next cache write if this prefix goes cold, which is what makes "worth keeping warm" a number rather than a feeling. 5h 78% is the currency a subscription actually runs out of: plan windows, not dollars, are what stop work mid-task.

COLD, off and ? are kept distinct on purpose: "the cache expired", "caching isn't happening here" and "warmline can't see" are three different facts, and collapsing them is how a cache gauge starts lying.

The expiry is absolute wall-clock, never a countdown — a frozen countdown is wrong, while a frozen clock is still true.

Every field, colors, troubleshooting →

When you come back cold

The cache is gone; that context gets processed once more at the uncached rate whatever you do next. The only question is what that one pass buys:

  • You still need the conversation history → /compact. The expensive pass was coming anyway; this way it produces a summary you carry cheaply from then on.
  • Your state is written down outside the conversation (memory files, a plan document, the code) → /clear. It skips even the summarization pass.
  • Small context → do nothing. Rebuilding 20k tokens is cheap.
  • Never while HOT, unless you are out of context window: compacting destroys a cache you already paid ~2× to build.

The full reasoning →

Explain and measure: the audit

A single HOT on your statusline is useful. Seeing that your sessions went cold 105 times in fourteen weeks is more useful.

warmline audit grades every recorded API request of a session from the usage fields Claude Code wrote for it — no one-turn lag, unlike the statusline. --all does the same across every session on this machine and ranks them. (It runs the installed warmline-audit; both spellings work, and scripts pinned to the hyphenated one keep working.) Real output from 14 weeks of history, priced at a flat --price 3 for a stable example — a bare --price solves your real rate instead, per project:

$ warmline-audit --all --price 3
74 sessions under /Users/mb/.claude/projects  (4 more without API turns; ttl per session from its cache buckets, 60m fallback)

cache health  █████████████████████████░  97% hot  (18,577 of 19,182 turns)
cold events   105  (36 rebuilt, 69 ttl) -- <1% of all turns

start        project                 turns    hot  part  rebuilt   ttl  avoidable cold  share    premium
09-23 17:57  MimirBlue                 287    282     0        1     4       2,635,789    18%     $15.02
   ⋮
TOTAL                                19182  18577   500       36    69      14,473,962   100%     $82.50

where the cold came from  (by estimated premium over warm reads; events beside)
  inactivity          ██████████████████████████  ~$82.50 (63%)  64 events
  auto-compact        ███████████  ~$35.07 (27%)  164 events  not avoidable
  model change        ██  ~$7.47 (5.7%)  6 events  chosen
  session start       █  ~$3.92 (3.0%)  30 events  not avoidable
  inactivity+compact  █  ~$1.22 (<1%)  5 events  not avoidable
  /compact            █  ~$1.21 (<1%)  5 events  not avoidable

estimated avoidable premium ~$82.50  (top 5 sessions: $43.86, other 69: $38.64)

That is the difference between observability and a decorative statusline: a pattern you can act on, or decide not to.

Each cold turn carries a cause where the transcript proves one — /compact, auto-compact, model change, inactivity, and claude upgrade: every transcript entry records the Claude Code build, so a cold rebuild whose version differs from the previous turn's is the upgrade rewriting the system prompt and tools. Everything else lands in unknown, which is a residual bucket, not a finding: the transcript recorded no proof, so warmline declines to name a cause. Anthropic documents several prefix-invalidating actions a transcript never witnesses — changing the effort level, turning on fast mode, denying a whole tool, enabling or disabling a plugin, connecting an MCP server whose tools load into the prefix — and any of them lands here. Editing CLAUDE.md mid-session does not: Anthropic lists that under the actions that keep the cache. Where unknown is large, read it as "the largest share is unexplained", which is a reason to look, not a diagnosis. (This history happens to have none.) The chart ranks causes by what they cost, not by how often they happened: here auto-compact fired more than twice as often as anything else, but walking away past the TTL re-cached the most, and it is the only cause the premium counts. (Verdicts come from recorded usage, but the split between COLD(rebuilt) and COLD(ttl) rests on the TTL, auto-detected per session from its own cache-bucket records.)

The dollar figures estimate exposure from token counts in your own transcripts. They are not billing data — warmline never sees, and cannot see, what Anthropic actually billed you. The closing line is labeled estimated avoidable premium. "Avoidable" leaves out each session's first cache write and every compaction's write — the compacted context, cached for the first time, which no timing could have spared once the compaction happened — and a model change, which you chose and which is reported on its own line. Some of what it still counts was never preventable, such as a TTL expiry while the laptop was asleep.

Warmline ships no price sheet. A baked-in table would go stale, and it would be wrong the moment you switched from Sonnet to Opus. A bare --price solves for the base input rate you are actually paying, from the cost Claude Code recorded for your last session in that project, using the multiples every Claude pricing tier shares (output 5× base input, a 1-hour cache write 2×, a 5-minute write 1.25×, a warm read 0.1×). --all prices each project at its own solved rate, every report prints where the number came from, and --price N still overrides. Output tokens are never cached, so the audit prices them on their own line rather than pretending warmth could save them.

Where the audit grades the past, warmline watch shows the present: every session's warmth, live, until ctrl-c. Desktop-app sessions appear too — they write the same transcripts even though they can't render a statusline.

Full walkthrough, verdicts, causes, --all, --live, --json →

Optional: keep warm

Everything above observes. Keep Warm is the optional fourth capability: a short instruction block in your ~/.claude/CLAUDE.md telling the agent to ping a long, quiet wait about every 50 minutes, so the cache is still warm when results land.

warmline keep-warm on        # off by default; global; reversible
warmline keep-warm status    # ON / OFF / INCONSISTENT (exit 0 / 1 / 2)

It is not a daemon — no cron, no process, nothing outside a running Claude Code session. It skips when background tasks are already keeping the cache warm for free, when the stake is small, and on the 5-minute cache, where ~12 pings an hour would cost more than the 1.15× rebuild they prevent. It stops the moment work resumes, and gives up after ~10 hours. Each ping is an ordinary billed request against your own plan: it trades one expensive rebuild for a few cheap reads, and bypasses nothing.

Better still, when the wait is something this machine can watch, don't schedule a ping at all:

warmline wait-for --pidfile /tmp/job.pid --until-cold

returns when the job ends or just before the cache expires, whichever comes first — read from this session's own transcript, so a job that finishes at minute 12 costs no ping at all.

What it is and isn't, wait-for, no-sleep mode, limits, terms →

Optional, at your own risk: AFK mode

Keep Warm pings only while a job is running, never just because you walked away. AFK mode is for walking away. Type afk in a Claude Code session, in the terminal or the desktop app, and the session keeps its own cache warm until you type again:

> afk
  AFK: keeping the cache warm until you're back (at most until 23:10).
  ...
> ok, where were we?
  warmline: welcome back -- AFK ended (away 2h40m, 3 keep-warm pings)

[!WARNING] Those pings are automated requests from a session nobody is attending. That may breach Anthropic's terms and can get your account rate-limited, suspended or banned. On a Pro or Max subscription the risk is highest. AFK mode is off until you turn it on yourself:

warmline afk enable     # explains the risk; you type "I accept the risk"

afk, brb, afk 3h or /afk 2h starts it, and any other message ends it. Under the hood it's a background waiter that wakes the session a few minutes before the cache would expire. The agent re-arms it with a one-line turn, and that turn is the ping. It stays bounded: one waiter per session, a limit per stretch (10h by default, 24h at most), no 5-minute caches, and it stops rather than pay for a rebuild if the cache goes cold anyway. The agent can't enable it for you.

How it works, the bounds, the risk in full →

Local by design

Warmline runs on your machine. It does not phone home, collect telemetry, or require an account, and it has no dependencies beyond python3 and bash.

Everything it shows comes from data Claude Code already produced locally: the JSON payload Claude Code pipes to the statusline, and the session transcripts under ~/.claude/projects. Warmline observes what Claude Code exposes locally — it has no special access to anything of Anthropic's, and there is nothing to log into.

Where it works

Front endstatuslinewarmline audit / watchkeep-warm / AFK
Terminal CLI✅✅✅
Desktop app (local Code tab)✅ as a band above the prompt✅✅
VS Code / JetBrains panel❌✅✅
Cloud / Cowork sessions❌❌❌

The graphical front ends don't render custom statuslines (open request). In the Desktop app's Code tab warmline draws the gauge itself, as a band above the prompt: a small Claude Code plugin the installer puts in place, grading each response the way the audit does. The IDE panels run the same engine, share the same ~/.claude and write the same transcripts, so the audit, warmline watch and keep-warm work there unchanged. Cloud and Cowork sessions are the exception: no part of warmline reaches them.

The full surface matrix, and how it was verified →

Evidence

The TTL is measured, not assumed. In a clean-room two-arm probe, a session idle 50 minutes read its full 71,312-token prefix back from cache; an identical one idle 70 minutes found the cache gone and re-wrote all 45,033 tokens. Warm at 50, cold at 70 — and reads refresh the clock. Everything else on this page comes from the same corpus the audit above grades.

Numbers, method, how to reproduce →

Docs

Statuslineevery field, colors, gap mechanics, troubleshooting
Auditverdicts, cause attribution, --all, the live watch view, what "avoidable" means
Keep Warmthe policy, no-sleep mode (warmline awake), limits, the terms question
AFK modeopt-in, at your own risk: type afk, the cache stays warm until you're back
Where it worksterminal, desktop, IDE, SSH, cloud
Installinstall, update, uninstall, configure, Windows, tests
Measurementsthe evidence behind every claim
Changelogtagged releases

License

MIT

Buy me a coffee ☕

BTC: bc1qjsvtd3dd44llyu4rwz2ucl4kp9wd9kvpsj6tk5

Source 2 files
hooks/register.tsx 161 lines
1// warmline's cache gauge for the surfaces that render no statusLine.
2//
3// Claude Code Desktop's Code tab draws no `statusLine`, so the gauge that
4// `warmline-statusline.py` prints never reaches it. This mod draws the band
5// above the prompt instead, from what a plugin can see: the four token counts
6// of every main-thread response. Each one is graded as `warmline audit` grades
7// it (verdict_of in warmline-audit), and the session's tally is kept.
8//
9// What it deliberately does not say is whether the prefix is warm *now*.
10// Claude Code's `prompt_cache` object (warm, TTL, expiry) reaches the
11// statusline alone; the plugin API exposes none of it. So the band reports
12// the last response's grade and how long ago it landed, both facts, and
13// leaves the TTL arithmetic to the reader.
14//
15// On the terminal, where warmline's statusLine is wired, the band yields: the
16// statusline is the better instrument there, and one gauge is enough.
17import { atom, read, update } from 'claude-code'
18import type { EngineInterface, Register, TurnUsage } from 'claude-code'
19import type { Counts, Verdict } from '../types'
20
21const gauge = atom({ plugin: 'warmline', key: 'gauge' } as const, null)
22const isHidden = atom({ plugin: 'warmline', key: 'isHidden' } as const, false)
23const tick = atom({ plugin: 'warmline', key: 'tick' } as const, 0)
24
25const COLOR: Record<Verdict, string | undefined> = {
26  HOT: 'green',
27  PARTIAL: 'yellow',
28  COLD: 'red',
29  '?': undefined,
30}
31
32/**
33 * verdict_of() from warmline-audit, less its COLD(ttl)/COLD(rebuilt) split:
34 * that needs the gap against the TTL, and a mod sees no TTL.
35 */
36function verdictOf(cacheRead: number, cacheWrite: number): Verdict {
37  if (cacheRead > 0) return cacheWrite > cacheRead ? 'PARTIAL' : 'HOT'
38  if (cacheWrite > 0) return 'COLD'
39  return '?'
40}
41
42/** fmt_tokens() from statusline.py: 127000 -> 127k, 1240000 -> 1.2M. */
43function fmtTokens(n: number): string {
44  if (n >= 999500) return `${(n / 1e6).toFixed(1).replace(/\.0$/, '')}M`
45  return `${Math.round(n / 1000)}k`
46}
47
48function fmtAgo(ms: number): string {
49  const min = Math.floor(ms / 60_000)
50  if (min < 1) return 'just now'
51  if (min < 60) return `${min} min ago`
52  const h = Math.floor(min / 60)
53  const m = min % 60
54  return m === 0 ? `${h} h ago` : `${h} h ${m} min ago`
55}
56
57function stake(verdict: Verdict, cacheRead: number, cacheWrite: number): string {
58  switch (verdict) {
59    case 'HOT':
60      return `read ${fmtTokens(cacheRead)}`
61    case 'PARTIAL':
62      return `wrote ${fmtTokens(cacheWrite)}, read ${fmtTokens(cacheRead)}`
63    case 'COLD':
64      return `wrote ${fmtTokens(cacheWrite)}`
65    default:
66      return 'no cache tokens'
67  }
68}
69
70/** The terminal renders `statusLine`; where warmline's is wired, that is the gauge to read. */
71async function hasWarmlineStatusLine($: EngineInterface): Promise<boolean> {
72  try {
73    const settings = await $.settings.read()
74    const statusLine = settings.statusLine
75    const command =
76      statusLine && typeof statusLine === 'object' ? (statusLine as { command?: unknown }).command : undefined
77    return typeof command === 'string' && /warmline/i.test(command)
78  } catch {
79    return false
80  }
81}
82
83async function record($: EngineInterface, usage: TurnUsage | null): Promise<void> {
84  if (!usage) return
85  const cacheRead = usage.cache_read_input_tokens
86  const cacheWrite = usage.cache_creation_input_tokens
87  // Nothing on the input side is a synthetic entry with no request behind it; the auditor skips those too.
88  if (!cacheRead && !cacheWrite && !usage.input_tokens) return
89  const verdict = verdictOf(cacheRead, cacheWrite)
90  const at = await $.clock.now()
91  await update($, gauge, (g) => {
92    const counts: Counts = { ...(g?.counts ?? { HOT: 0, PARTIAL: 0, COLD: 0 }) }
93    if (verdict !== '?') counts[verdict] += 1
94    return { last: { verdict, cacheRead, cacheWrite, at }, counts }
95  })
96}
97
98export const register: Register = (on) => {
99  on('session.start', ($, e, next) => {
100    // Once a minute, so "n min ago" stays current while the session idles.
101    $.clock.every(60_000, () => update($, tick, (n) => n + 1))
102    return next(e)
103  })
104
105  on('turn.step', async function* ($, e, next) {
106    const result = yield* next(e)
107    if (e.agentId === undefined) {
108      try {
109        await record($, result.usage)
110      } catch {
111        // A failed record must never touch the response.
112      }
113    }
114    return result
115  })
116
117  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
118    if (e.props.hasSurvey || (await read($, isHidden))) return next(e)
119    if (e.surface === 'terminal' && (await hasWarmlineStatusLine($))) return next(e)
120
121    const g = await read($, gauge)
122    await read($, tick) // subscribes the band to the minute tick
123    const { Box, Button, Text } = $.ui.resolve(e)
124
125    if (g === null) {
126      return (
127        <Box gap={1}>
128          <Text dimColor>warmline</Text>
129          <Text dimColor>
130            cache ?
131          </Text>
132          <Text dimColor>
133            no response yet
134          </Text>
135          <Button key="hide" label="Hide" onPress={() => update($, isHidden, () => true)} />
136        </Box>
137      )
138    }
139
140    const { last, counts } = g
141    const ago = fmtAgo((await $.clock.now()) - last.at)
142    return (
143      <Box gap={1}>
144        <Text dimColor>warmline</Text>
145        <Text>last turn</Text>
146        <Text color={COLOR[last.verdict]} bold>
147          {last.verdict}
148        </Text>
149        <Text>{`(${stake(last.verdict, last.cacheRead, last.cacheWrite)})`}</Text>
150        <Text dimColor>
151          {ago}
152        </Text>
153        <Text dimColor>
154          {`· session ${counts.HOT} hot ${counts.PARTIAL} partial ${counts.COLD} cold`}
155        </Text>
156        <Button key="hide" label="Hide" onPress={() => update($, isHidden, () => true)} />
157      </Box>
158    )
159  })
160}
161
types/index.d.ts 20 lines
1// The values the warmline mod keeps in `$.state`, declared so that
2// `claude plugin validate` can hold the module to them.
3
4/** The grade of one API response, as `warmline audit` spells it. */
5export type Verdict = 'HOT' | 'PARTIAL' | 'COLD' | '?'
6
7/** How many of this session's main-thread responses took each grade. */
8export type Counts = { HOT: number; PARTIAL: number; COLD: number }
9
10/** The last main-thread response: its grade, its cache traffic, when it landed. */
11export type LastTurn = { verdict: Verdict; cacheRead: number; cacheWrite: number; at: number }
12
13export type Gauge = { last: LastTurn; counts: Counts }
14
15declare module 'claude-code' {
16  interface PluginState {
17    warmline: { gauge: Gauge | null; isHidden: boolean; tick: number }
18  }
19}
20