Claude Code's prompt-cache state, drawn above the prompt where no statusline renders: the Desktop app's Code tab.

Warmline makes Claude Code's prompt-cache state visible.
A working session carries tens or hundreds of thousands of tokens — system prompt, tools, every file it read, every reply — and Claude Code re-sends all of it on every turn. While the prompt cache holds, that context is read back at roughly 0.1× the normal input price. When the cache goes cold, the same context is processed again at about 2×. It goes cold after about an hour of inactivity, and instantly whenever the start of the conversation is rewritten — /compact and auto-compaction both do that. Come back to a big session after lunch and your next message pays the rebuild on all of it, at exactly the moment you wanted results, and nothing in the session tells you so.
Warmline puts the cache state directly in your statusline, and gives you the tools to see what it has been doing historically.
Opus 5 | claude-warmline | ctx 64% (127k) | cache HOT (127k, cold ~11:58) | 5h 78%
No guessing, and no timing-based inference. Since v2.1.251 Claude Code hands the statusline its own prompt_cache object — whether the prefix is warm, which TTL it is on, the second it expires — and warmline prints what that object says.
Then, for everything before this turn:
warmline audit
grades every recorded API turn of a session from the usage Claude Code wrote for it, and --all ranks every session on this machine.
curl -fsSL https://raw.githubusercontent.com/Miguel-Barroso/claude-warmline/main/install.sh | bash
or with Homebrew, on macOS:
brew install Miguel-Barroso/warmline/warmline
Then:
| Command | What it does |
|---|---|
warmline status | what's installed and on |
warmline audit | this session, turn by turn |
warmline audit --all | every session on this machine, ranked |
warmline watch | every session's warmth, live, until ctrl-c |
warmline afk enable | optional, at your own risk: type afk and the cache stays warm until you're back |
Needs python3 and bash, nothing else. On Windows, run it from Git Bash (with Python 3.7+ installed) or inside WSL. No curl? wget -qO- <same URL> | bash works the same way, and the installer downloads with whichever one it finds. Whichever route you take it is one command: brew install wires the statusline and the Desktop band too, brew upgrade re-wires them, and brew uninstall unwires them again.
Install details, flags, pinning a release, package managers, Windows →
Any statusline can now show you whether the cache is warm. Since v2.1.251 Claude Code hands every statusline script the same prompt_cache object warmline reads — warmline's own gauge is, deliberately, nothing more than what that object says. The live light matters, and it is no longer the point.
Warmline exists for the questions no live light can answer:
When did the cache go cold, how often, why — and did it cost enough to care about?
That is warmline audit: every recorded turn of your history graded from the usage Claude Code wrote for it, each cold event attributed to a cause only where the transcript proves one, and the avoidable premium priced at rates derived from your own sessions. Around the audit sits the full loop — observe, explain, measure, mitigate (keep-warm, wait-for --until-cold, awake) — so a cold event goes from noticed to explained to priced to, when it earns it, prevented.
A statusline tells you what is happening in your session. Warmline tells you whether your context is still being reused — and what it has cost you when it wasn't.
Everything here exists for the 100k+ case. Going cold on 20k tokens is cheap.
Observe — the cache state in your statusline, live, from Claude Code's own data.
Explain — what HOT, COLD, off and the expiry clock each mean, and which of them you can do anything about.
Measure — warmline audit: the cold events across your history, what caused them where the transcript proves a cause, and an estimate of what they were worth.
Mitigate — warmline keep-warm, if it turns out you need it.
That is a progression, not a feature list. It takes you from "did my cache go cold?" to "how often does this happen?" to "is it costing me enough to care?" to "do I want to do something about it?" — and the honest answer to the last one is often no.
The first three are read-only observation, and they are the point of the project. The fourth is opt-in, off by default, and deliberately bounded — see optional: keep warm.
Opus 5 | claude-warmline | ctx 64% (127k) | cache HOT (127k, cold ~11:58) | 5h 78%
| Field | Meaning |
|---|---|
ctx 64% (127k) | context-window utilization — yellow within 10k tokens of the threshold auto-compaction actually fires at (window - 33000) |
cache HOT (127k, cold ~11:58) | the cached prefix is warm, a rebuild would re-cache 127k tokens, and it leaves its TTL at the time shown; yellow within 15 minutes of it |
cache HOT 5m | as above, on the 5-minute TTL — usage credits, an API key, a cloud provider. The 1-hour case is the norm and goes unlabelled |
5h 78% / 7d 91% | the plan window nearest its cap, hidden below 50%, with its reset time once it turns yellow. Absent on API keys and cloud providers, which have no plan window |
cache COLD | the prefix is outside its TTL; the next turn re-caches it |
cache off | prompt caching is off, or this provider or gateway never reports cache tokens. Nothing here will warm up |
cache ? | no cache data — Claude Code before v2.1.251, or before the session's first API response |
Warmline does not try to work out whether your cache is warm by timing how long Claude Code takes to answer. Every verdict above is a field Claude Code handed the statusline: prompt_cache.warm for the state, caching_observed for whether caching is happening at all, ttl for the bucket, expires_at for the clock, recache_tokens_if_cold for the stake, rate_limits for the plan window. Warmline formats and colors them. It used to infer all of this from the gap between turns, and that machinery is gone.
The two numbers that decide anything are the stake and the quota. (127k) is what a rebuild would cost you — the size of the next cache write if this prefix goes cold, which is what makes "worth keeping warm" a number rather than a feeling. 5h 78% is the currency a subscription actually runs out of: plan windows, not dollars, are what stop work mid-task.
COLD, off and ? are kept distinct on purpose: "the cache expired", "caching isn't happening here" and "warmline can't see" are three different facts, and collapsing them is how a cache gauge starts lying.
The expiry is absolute wall-clock, never a countdown — a frozen countdown is wrong, while a frozen clock is still true.
Every field, colors, troubleshooting →
The cache is gone; that context gets processed once more at the uncached rate whatever you do next. The only question is what that one pass buys:
/compact. The expensive pass was coming anyway; this way it produces a summary you carry cheaply from then on./clear. It skips even the summarization pass.HOT, unless you are out of context window: compacting destroys a cache you already paid ~2× to build.A single HOT on your statusline is useful. Seeing that your sessions went cold 105 times in fourteen weeks is more useful.
warmline audit grades every recorded API request of a session from the usage fields Claude Code wrote for it — no one-turn lag, unlike the statusline. --all does the same across every session on this machine and ranks them. (It runs the installed warmline-audit; both spellings work, and scripts pinned to the hyphenated one keep working.) Real output from 14 weeks of history, priced at a flat --price 3 for a stable example — a bare --price solves your real rate instead, per project:
$ warmline-audit --all --price 3
74 sessions under /Users/mb/.claude/projects (4 more without API turns; ttl per session from its cache buckets, 60m fallback)
cache health █████████████████████████░ 97% hot (18,577 of 19,182 turns)
cold events 105 (36 rebuilt, 69 ttl) -- <1% of all turns
start project turns hot part rebuilt ttl avoidable cold share premium
09-23 17:57 MimirBlue 287 282 0 1 4 2,635,789 18% $15.02
⋮
TOTAL 19182 18577 500 36 69 14,473,962 100% $82.50
where the cold came from (by estimated premium over warm reads; events beside)
inactivity ██████████████████████████ ~$82.50 (63%) 64 events
auto-compact ███████████ ~$35.07 (27%) 164 events not avoidable
model change ██ ~$7.47 (5.7%) 6 events chosen
session start █ ~$3.92 (3.0%) 30 events not avoidable
inactivity+compact █ ~$1.22 (<1%) 5 events not avoidable
/compact █ ~$1.21 (<1%) 5 events not avoidable
estimated avoidable premium ~$82.50 (top 5 sessions: $43.86, other 69: $38.64)
That is the difference between observability and a decorative statusline: a pattern you can act on, or decide not to.
Each cold turn carries a cause where the transcript proves one — /compact, auto-compact, model change, inactivity, and claude upgrade: every transcript entry records the Claude Code build, so a cold rebuild whose version differs from the previous turn's is the upgrade rewriting the system prompt and tools. Everything else lands in unknown, which is a residual bucket, not a finding: the transcript recorded no proof, so warmline declines to name a cause. Anthropic documents several prefix-invalidating actions a transcript never witnesses — changing the effort level, turning on fast mode, denying a whole tool, enabling or disabling a plugin, connecting an MCP server whose tools load into the prefix — and any of them lands here. Editing CLAUDE.md mid-session does not: Anthropic lists that under the actions that keep the cache. Where unknown is large, read it as "the largest share is unexplained", which is a reason to look, not a diagnosis. (This history happens to have none.) The chart ranks causes by what they cost, not by how often they happened: here auto-compact fired more than twice as often as anything else, but walking away past the TTL re-cached the most, and it is the only cause the premium counts. (Verdicts come from recorded usage, but the split between COLD(rebuilt) and COLD(ttl) rests on the TTL, auto-detected per session from its own cache-bucket records.)
The dollar figures estimate exposure from token counts in your own transcripts. They are not billing data — warmline never sees, and cannot see, what Anthropic actually billed you. The closing line is labeled estimated avoidable premium. "Avoidable" leaves out each session's first cache write and every compaction's write — the compacted context, cached for the first time, which no timing could have spared once the compaction happened — and a model change, which you chose and which is reported on its own line. Some of what it still counts was never preventable, such as a TTL expiry while the laptop was asleep.
Warmline ships no price sheet. A baked-in table would go stale, and it would be wrong the moment you switched from Sonnet to Opus. A bare --price solves for the base input rate you are actually paying, from the cost Claude Code recorded for your last session in that project, using the multiples every Claude pricing tier shares (output 5× base input, a 1-hour cache write 2×, a 5-minute write 1.25×, a warm read 0.1×). --all prices each project at its own solved rate, every report prints where the number came from, and --price N still overrides. Output tokens are never cached, so the audit prices them on their own line rather than pretending warmth could save them.
Where the audit grades the past, warmline watch shows the present: every session's warmth, live, until ctrl-c. Desktop-app sessions appear too — they write the same transcripts even though they can't render a statusline.
Full walkthrough, verdicts, causes, --all, --live, --json →
Everything above observes. Keep Warm is the optional fourth capability: a short instruction block in your ~/.claude/CLAUDE.md telling the agent to ping a long, quiet wait about every 50 minutes, so the cache is still warm when results land.
warmline keep-warm on # off by default; global; reversible
warmline keep-warm status # ON / OFF / INCONSISTENT (exit 0 / 1 / 2)
It is not a daemon — no cron, no process, nothing outside a running Claude Code session. It skips when background tasks are already keeping the cache warm for free, when the stake is small, and on the 5-minute cache, where ~12 pings an hour would cost more than the 1.15× rebuild they prevent. It stops the moment work resumes, and gives up after ~10 hours. Each ping is an ordinary billed request against your own plan: it trades one expensive rebuild for a few cheap reads, and bypasses nothing.
Better still, when the wait is something this machine can watch, don't schedule a ping at all:
warmline wait-for --pidfile /tmp/job.pid --until-cold
returns when the job ends or just before the cache expires, whichever comes first — read from this session's own transcript, so a job that finishes at minute 12 costs no ping at all.
What it is and isn't, wait-for, no-sleep mode, limits, terms →
Keep Warm pings only while a job is running, never just because you walked away. AFK mode is for walking away. Type afk in a Claude Code session, in the terminal or the desktop app, and the session keeps its own cache warm until you type again:
> afk
AFK: keeping the cache warm until you're back (at most until 23:10).
...
> ok, where were we?
warmline: welcome back -- AFK ended (away 2h40m, 3 keep-warm pings)
[!WARNING] Those pings are automated requests from a session nobody is attending. That may breach Anthropic's terms and can get your account rate-limited, suspended or banned. On a Pro or Max subscription the risk is highest. AFK mode is off until you turn it on yourself:
warmline afk enable # explains the risk; you type "I accept the risk"
afk, brb, afk 3h or /afk 2h starts it, and any other message ends it. Under the hood it's a background waiter that wakes the session a few minutes before the cache would expire. The agent re-arms it with a one-line turn, and that turn is the ping. It stays bounded: one waiter per session, a limit per stretch (10h by default, 24h at most), no 5-minute caches, and it stops rather than pay for a rebuild if the cache goes cold anyway. The agent can't enable it for you.
How it works, the bounds, the risk in full →
Warmline runs on your machine. It does not phone home, collect telemetry, or require an account, and it has no dependencies beyond python3 and bash.
Everything it shows comes from data Claude Code already produced locally: the JSON payload Claude Code pipes to the statusline, and the session transcripts under ~/.claude/projects. Warmline observes what Claude Code exposes locally — it has no special access to anything of Anthropic's, and there is nothing to log into.
| Front end | statusline | warmline audit / watch | keep-warm / AFK |
|---|---|---|---|
| Terminal CLI | ✅ | ✅ | ✅ |
| Desktop app (local Code tab) | ✅ as a band above the prompt | ✅ | ✅ |
| VS Code / JetBrains panel | ❌ | ✅ | ✅ |
| Cloud / Cowork sessions | ❌ | ❌ | ❌ |
The graphical front ends don't render custom statuslines (open request). In the Desktop app's Code tab warmline draws the gauge itself, as a band above the prompt: a small Claude Code plugin the installer puts in place, grading each response the way the audit does. The IDE panels run the same engine, share the same ~/.claude and write the same transcripts, so the audit, warmline watch and keep-warm work there unchanged. Cloud and Cowork sessions are the exception: no part of warmline reaches them.
The full surface matrix, and how it was verified →
The TTL is measured, not assumed. In a clean-room two-arm probe, a session idle 50 minutes read its full 71,312-token prefix back from cache; an identical one idle 70 minutes found the cache gone and re-wrote all 45,033 tokens. Warm at 50, cold at 70 — and reads refresh the clock. Everything else on this page comes from the same corpus the audit above grades.
Numbers, method, how to reproduce →
| Statusline | every field, colors, gap mechanics, troubleshooting |
| Audit | verdicts, cause attribution, --all, the live watch view, what "avoidable" means |
| Keep Warm | the policy, no-sleep mode (warmline awake), limits, the terms question |
| AFK mode | opt-in, at your own risk: type afk, the cache stays warm until you're back |
| Where it works | terminal, desktop, IDE, SSH, cloud |
| Install | install, update, uninstall, configure, Windows, tests |
| Measurements | the evidence behind every claim |
| Changelog | tagged releases |
BTC: bc1qjsvtd3dd44llyu4rwz2ucl4kp9wd9kvpsj6tk5
hooks/register.tsx 161 lines1// warmline's cache gauge for the surfaces that render no statusLine.
2//
3// Claude Code Desktop's Code tab draws no `statusLine`, so the gauge that
4// `warmline-statusline.py` prints never reaches it. This mod draws the band
5// above the prompt instead, from what a plugin can see: the four token counts
6// of every main-thread response. Each one is graded as `warmline audit` grades
7// it (verdict_of in warmline-audit), and the session's tally is kept.
8//
9// What it deliberately does not say is whether the prefix is warm *now*.
10// Claude Code's `prompt_cache` object (warm, TTL, expiry) reaches the
11// statusline alone; the plugin API exposes none of it. So the band reports
12// the last response's grade and how long ago it landed, both facts, and
13// leaves the TTL arithmetic to the reader.
14//
15// On the terminal, where warmline's statusLine is wired, the band yields: the
16// statusline is the better instrument there, and one gauge is enough.
17import { atom, read, update } from 'claude-code'
18import type { EngineInterface, Register, TurnUsage } from 'claude-code'
19import type { Counts, Verdict } from '../types'
20
21const gauge = atom({ plugin: 'warmline', key: 'gauge' } as const, null)
22const isHidden = atom({ plugin: 'warmline', key: 'isHidden' } as const, false)
23const tick = atom({ plugin: 'warmline', key: 'tick' } as const, 0)
24
25const COLOR: Record<Verdict, string | undefined> = {
26 HOT: 'green',
27 PARTIAL: 'yellow',
28 COLD: 'red',
29 '?': undefined,
30}
31
32/**
33 * verdict_of() from warmline-audit, less its COLD(ttl)/COLD(rebuilt) split:
34 * that needs the gap against the TTL, and a mod sees no TTL.
35 */
36function verdictOf(cacheRead: number, cacheWrite: number): Verdict {
37 if (cacheRead > 0) return cacheWrite > cacheRead ? 'PARTIAL' : 'HOT'
38 if (cacheWrite > 0) return 'COLD'
39 return '?'
40}
41
42/** fmt_tokens() from statusline.py: 127000 -> 127k, 1240000 -> 1.2M. */
43function fmtTokens(n: number): string {
44 if (n >= 999500) return `${(n / 1e6).toFixed(1).replace(/\.0$/, '')}M`
45 return `${Math.round(n / 1000)}k`
46}
47
48function fmtAgo(ms: number): string {
49 const min = Math.floor(ms / 60_000)
50 if (min < 1) return 'just now'
51 if (min < 60) return `${min} min ago`
52 const h = Math.floor(min / 60)
53 const m = min % 60
54 return m === 0 ? `${h} h ago` : `${h} h ${m} min ago`
55}
56
57function stake(verdict: Verdict, cacheRead: number, cacheWrite: number): string {
58 switch (verdict) {
59 case 'HOT':
60 return `read ${fmtTokens(cacheRead)}`
61 case 'PARTIAL':
62 return `wrote ${fmtTokens(cacheWrite)}, read ${fmtTokens(cacheRead)}`
63 case 'COLD':
64 return `wrote ${fmtTokens(cacheWrite)}`
65 default:
66 return 'no cache tokens'
67 }
68}
69
70/** The terminal renders `statusLine`; where warmline's is wired, that is the gauge to read. */
71async function hasWarmlineStatusLine($: EngineInterface): Promise<boolean> {
72 try {
73 const settings = await $.settings.read()
74 const statusLine = settings.statusLine
75 const command =
76 statusLine && typeof statusLine === 'object' ? (statusLine as { command?: unknown }).command : undefined
77 return typeof command === 'string' && /warmline/i.test(command)
78 } catch {
79 return false
80 }
81}
82
83async function record($: EngineInterface, usage: TurnUsage | null): Promise<void> {
84 if (!usage) return
85 const cacheRead = usage.cache_read_input_tokens
86 const cacheWrite = usage.cache_creation_input_tokens
87 // Nothing on the input side is a synthetic entry with no request behind it; the auditor skips those too.
88 if (!cacheRead && !cacheWrite && !usage.input_tokens) return
89 const verdict = verdictOf(cacheRead, cacheWrite)
90 const at = await $.clock.now()
91 await update($, gauge, (g) => {
92 const counts: Counts = { ...(g?.counts ?? { HOT: 0, PARTIAL: 0, COLD: 0 }) }
93 if (verdict !== '?') counts[verdict] += 1
94 return { last: { verdict, cacheRead, cacheWrite, at }, counts }
95 })
96}
97
98export const register: Register = (on) => {
99 on('session.start', ($, e, next) => {
100 // Once a minute, so "n min ago" stays current while the session idles.
101 $.clock.every(60_000, () => update($, tick, (n) => n + 1))
102 return next(e)
103 })
104
105 on('turn.step', async function* ($, e, next) {
106 const result = yield* next(e)
107 if (e.agentId === undefined) {
108 try {
109 await record($, result.usage)
110 } catch {
111 // A failed record must never touch the response.
112 }
113 }
114 return result
115 })
116
117 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
118 if (e.props.hasSurvey || (await read($, isHidden))) return next(e)
119 if (e.surface === 'terminal' && (await hasWarmlineStatusLine($))) return next(e)
120
121 const g = await read($, gauge)
122 await read($, tick) // subscribes the band to the minute tick
123 const { Box, Button, Text } = $.ui.resolve(e)
124
125 if (g === null) {
126 return (
127 <Box gap={1}>
128 <Text dimColor>warmline</Text>
129 <Text dimColor>
130 cache ?
131 </Text>
132 <Text dimColor>
133 no response yet
134 </Text>
135 <Button key="hide" label="Hide" onPress={() => update($, isHidden, () => true)} />
136 </Box>
137 )
138 }
139
140 const { last, counts } = g
141 const ago = fmtAgo((await $.clock.now()) - last.at)
142 return (
143 <Box gap={1}>
144 <Text dimColor>warmline</Text>
145 <Text>last turn</Text>
146 <Text color={COLOR[last.verdict]} bold>
147 {last.verdict}
148 </Text>
149 <Text>{`(${stake(last.verdict, last.cacheRead, last.cacheWrite)})`}</Text>
150 <Text dimColor>
151 {ago}
152 </Text>
153 <Text dimColor>
154 {`· session ${counts.HOT} hot ${counts.PARTIAL} partial ${counts.COLD} cold`}
155 </Text>
156 <Button key="hide" label="Hide" onPress={() => update($, isHidden, () => true)} />
157 </Box>
158 )
159 })
160}
161types/index.d.ts 20 lines1// The values the warmline mod keeps in `$.state`, declared so that
2// `claude plugin validate` can hold the module to them.
3
4/** The grade of one API response, as `warmline audit` spells it. */
5export type Verdict = 'HOT' | 'PARTIAL' | 'COLD' | '?'
6
7/** How many of this session's main-thread responses took each grade. */
8export type Counts = { HOT: number; PARTIAL: number; COLD: number }
9
10/** The last main-thread response: its grade, its cache traffic, when it landed. */
11export type LastTurn = { verdict: Verdict; cacheRead: number; cacheWrite: number; at: number }
12
13export type Gauge = { last: LastTurn; counts: Counts }
14
15declare module 'claude-code' {
16 interface PluginState {
17 warmline: { gauge: Gauge | null; isHidden: boolean; tick: number }
18 }
19}
20