Gauges how long the conversation's prompt cache stays warm, its hit ratio and misses, and what the next message re-caches once it has gone cold

会話のプロンプトキャッシュがあと何分もつかを、プロンプト入力欄の上の帯にゲージで常時表示する Claude Code の mod です。 キャッシュが切れたあとは、次のメッセージで何トークンを書き直すことになるかを表示します。
cache ● ~1h ████████░░ 48m left · hit 91% · misses 0
cache ● ~1h █░░░░░░░░░ 53s left · hit 91% · misses 0
cache ○ cold · next message re-caches 128k tokens
Claude Code はメッセージのたびに会話全体を送り直していて、キャッシュはその読み直しを安くしています。 ところが、キャッシュの残り時間は、/usage を開くかステータスラインのスクリプトを書かない限り目に入りません。 席を外して戻ってきた直後の 1 ターンが遅く、使用量も多く消費するのは、その間に TTL が切れていたからです。 残りが少ないときに聞きたいことがあるなら、切れる前に聞いておくほうが安く済みます。
cold へ変わるexpired、モデルの切替なら model switch)は括弧で添える。/compact のあとの作り直しは数えないキャッシュが温まっていない状態は、次のように表示します。
| 表示 | 状態 |
|---|---|
cache ○ cold · next message re-caches N tokens | TTL を過ぎた。N は直前の応答の入力(キャッシュの読み書きを含む)と出力の合計 |
cache ○ cold · model switch · … | モデルを切り替えた。キャッシュはモデルごとなので、次の送信は会話全体を書き直す |
cache ○ rebuilding · … | /compact のあと。会話部分のキャッシュは、次の送信で要約から作り直される |
cache ○ off · no cache tokens reported | 応答にキャッシュの読み書きがなかった。キャッシュを無効にしているか、ゲートウェイが報告していない |
cache ○ off · DISABLE_PROMPT_CACHING | 環境変数でキャッシュを切っている |
最初の応答が返るまでと、/clear の直後は何も表示しません。
残り時間は、メイン会話の最後のリクエストを送った時刻から数えます。 キャッシュを読んだリクエストは、そのたびにタイマーを戻すからです。 サブエージェントやワークフローは別のキャッシュを持ち、TTL も別なので、そのリクエストでは数え直しません。 ただし fork は会話をそのまま引き継ぐので、最初のリクエストがメイン会話のキャッシュを読み、残り時間を戻します。
TTL は、Claude Code が採用するのと同じ優先順で決めます。
PostModelSwitch の cache_ttl)FORCE_PROMPT_CACHING_5M=1 なら 5 分CLAUDE_CODE_PROMPT_CACHE_TTL(5m か 1h)promptCacheTtl(5m か 1h)ENABLE_PROMPT_CACHING_1H=1 なら 1 時間6 の既定値だけは推定なので、TTL の前に ~ を付けます(~1h)。 mods API はプランの種類を直接は渡さないため、OAuth ログインならサブスクリプション、5 時間か 7 日の制限が 100% に達していたら usage credits、と見なしています。 1 は実際の値ですが、モデルを切り替えたときにしか届きません。
TTL はリクエストごとに、応答が返った時点で決めて記録します。 Claude Code もリクエストごとに TTL を選ぶので、usage credits で書いた 5 分のキャッシュは、そのあと利用枠が戻っても 5 分で切れます。
Claude Code v2.1.292 で、読み込みと、実際の応答からの表示を確認しています。 /plugin でのインストールと更新の手順は、リポジトリの README にまとめてあります。 インストールせずに試すときは、clone したリポジトリのルートで次のように起動します。
claude --plugin-dir ./plugins/cache-meter
次の操作は mod から見えないので、表示に反映されません。
/rewind:巻き戻した先の接頭辞はそれ以後のリクエストが読み続けているので、キャッシュは残る。ただし再キャッシュのトークン数は巻き戻す前の大きさのまま表示されるopusplan でのプランモードの出入り:モデルの切替にあたるが、切替のイベントが来ない場合は、次の応答が返るまで前のモデルの残り時間を表示するDISABLE_PROMPT_CACHING_SONNET などのモデル別の無効化:そのファミリーのエイリアスが指すモデルにだけ効くが、mods API からはエイリアスの解決先が分からない。そのモデルに切り替えた直後は cold と表示し、最初の応答で off に変わる公式のステータスラインスクリプトなら、prompt_cache オブジェクト(v2.1.251 以降)から expires_at や recache_tokens_if_cold を Claude Code 自身の計算で受け取れます。 この mod は、その値に手が届かない mods API の中で、同じものを応答のトークン数から組み立てています。
すべて How Claude Code uses prompt caching によります。
opusplan ではプランモードの出入りがモデルの切替になる/compact は会話部分のキャッシュを作り直し、/rewind は巻き戻した先のキャッシュをそのまま読むcd plugins/cache-meter
claude plugin testhooks/register.js 417 lines1// Draws a gauge in the band above the prompt for the main conversation's prompt cache: how long
2// it stays warm, its hit ratio and misses, and once it has gone cold, how many tokens the next
3// message writes to the cache again.
4
5// The last main-thread request that came back with usage: { seq, at, ttl, model, tokens, prefix, cached, expired }
6// - seq: which main request it was, counted across the module's life, so a fork can tell whether
7// the entry it read is still the newest
8// - at: when it was sent, in $.clock.now() milliseconds. Each request that reads the cache resets
9// its TTL, so the TTL runs from here; the response may end minutes later.
10// - ttl: the TTL resolved as it came back, { ttl, estimated }, since Claude Code picks one per
11// request; null when unknown (a resumed conversation), which takes the TTL as it stands now
12// - model: the model that answered it, null when unknown (a resumed conversation); each model has
13// a cache of its own
14// - tokens: what the next request re-sends, input + cache read + cache write + output, as the
15// engine counts it for a model switch
16// - prefix: what the cache held after it, cache read + cache write; the next request reads about
17// this much when the cache is still there
18// - cached: whether the response reported any cache tokens at all
19// - expired: the engine said, on resume, that the cache had likely gone cold
20let last = null
21// The model the main loop runs now, as the last step or model switch named it
22let model = null
23// This conversation's main requests: input tokens in all and from the cache, and the requests
24// that wrote back what the cache had held, with the likely cause of the last one
25let stats = emptyStats()
26// The TTL the engine stated at the last model switch: { ttl, overLimit }, kept only while the
27// plan usage stays on the side of the limit it was on then
28let engineTtl = null
29// The TTL as last resolved, { ttl, estimated }, for a request whose own is unknown, and whether
30// caching is switched off
31let ttl = defaultTtl()
32let disabled = false
33// The rate-limit windows, for the TTL estimate: past a window's limit, usage credits pay
34let rateLimits = []
35// After /compact, the next request rebuilds the conversation layer: not a miss
36let compacted = false
37// The subagents whose first request has been seen: only a fork's first reads the main
38// conversation's entry, hit or miss
39let agentsSeen = new Set()
40// Numbers the main requests (`last.seq`)
41let mainRequests = 0
42// The timer that redraws the gauge when its text next changes
43let timer = null
44// Counts refreshes and resets, so only the latest refresh redraws: one begun before a later one,
45// or before /clear and the rest, leaves the gauge alone
46let refreshes = 0
47
48const TTL_MS = { '5m': 5 * 60_000, '1h': 60 * 60_000 }
49// The session.end reasons after which this module stops; /clear, /resume and logout leave it
50// running
51const FINAL_REASONS = ['prompt_input_exit', 'other']
52// The windows a subscription's plan usage is measured in; past either, Claude Code draws on usage
53// credits and drops the main conversation to five minutes
54const PLAN_WINDOWS = ['five_hour', 'seven_day']
55// A request that read less than this share of what the cache held before it is a miss
56const MISS_BELOW_SHARE = 0.5
57// Under this share of the TTL left, the gauge turns yellow
58const WARN_BELOW_SHARE = 0.2
59
60const BAR_CELLS = 10
61// The band's last column, which the terminal may draw over
62const BAND_RESERVED_COLUMNS = 2
63const SVG_BAR = { width: 96, height: 10 }
64const SVG_COLORS = { success: '#4caf50', warning: '#e0a526', error: '#e5534b', track: 'rgba(128,128,128,0.3)' }
65
66export function register(on) {
67 // Fires again on an enable or a worker respawn, which may keep this module's variables
68 on('session.start', async ($, e, next) => {
69 reset()
70 engineTtl = null
71 ttl = defaultTtl()
72 disabled = false
73 rateLimits = (await $.session.usage()).rateLimits
74 void refresh($)
75 return next(e)
76 })
77
78 on('session.end', async ($, e, next) => {
79 if (!FINAL_REASONS.includes(e.reason)) return next(e)
80 refreshes += 1
81 timer?.cancel()
82 timer = null
83 return next(e)
84 })
85
86 on('session.measure', async ($, e, next) => {
87 if (e.changed.includes('rateLimits')) {
88 rateLimits = e.rateLimits
89 void refresh($)
90 }
91 return next(e)
92 })
93
94 // The main loop's requests (no agentId) move the gauge. Subagents and workflows keep caches of
95 // their own, with their own TTL; a fork inherits the main conversation whole, so its first
96 // request reads the main entry and resets its timer.
97 on('turn.step', async function* ($, e, next) {
98 const sentAt = await $.clock.now()
99 // The main entry as this request went out, which a fork's request read if it read any
100 const parent = last
101 const response = yield* next(e)
102 const usage = response?.usage
103 if (!usage) return response
104 // A later mod may have sent the request to another model than the step named
105 const answeredBy = usage.model || e.model
106 if (e.agentId) {
107 // Forks that overlap may answer out of order: the newest read stands, and a main request
108 // recorded since holds an entry the fork did not read
109 const read = await readsMainEntry($, e.agentId, parent, answeredBy, usage).catch(() => false)
110 if (read && last?.seq === parent.seq && sentAt > last.at) {
111 last = { ...last, at: sentAt }
112 void refresh($)
113 }
114 return response
115 }
116 // Unresolved, the request takes the TTL as it stands when drawn
117 const resolved = await resolveTtl($).catch(() => null)
118 const read = usage.cache_read_input_tokens
119 const written = usage.cache_creation_input_tokens
120 count(answeredBy, sentAt, usage)
121 last = {
122 seq: ++mainRequests,
123 at: sentAt,
124 ttl: resolved,
125 model: answeredBy,
126 tokens: usage.input_tokens + read + written + usage.output_tokens,
127 prefix: read + written,
128 cached: read + written > 0,
129 expired: false,
130 }
131 model = answeredBy
132 compacted = false
133 void refresh($)
134 return response
135 })
136
137 // The engine states the TTL it uses here, which no other event carries
138 on('classic.PostModelSwitch', async ($, e, next) => {
139 model = e.to_model
140 engineTtl = { ttl: e.cache_ttl, overLimit: overLimit() }
141 if (last && e.context_tokens > 0) last = { ...last, tokens: e.context_tokens }
142 void refresh($)
143 return next(e)
144 })
145
146 // /clear starts a conversation with nothing cached for it yet. /compact replaces the history,
147 // so the next request writes the summary's cache. A resumed or forked conversation carries how
148 // long ago it last had a response, which dates its cache.
149 on('classic.SessionStart', { source: ['clear', 'compact', 'resume', 'fork'] }, async ($, e, next) => {
150 if (e.source === 'compact') {
151 compacted = true
152 } else {
153 reset()
154 if (typeof e.seconds_since_last_response === 'number' && typeof e.context_tokens === 'number') {
155 last = {
156 seq: ++mainRequests,
157 at: (await $.clock.now()) - e.seconds_since_last_response * 1000,
158 ttl: null,
159 model: null,
160 tokens: e.context_tokens,
161 prefix: e.context_tokens,
162 cached: true,
163 expired: e.prompt_cache_likely_expired === true,
164 }
165 }
166 }
167 void refresh($)
168 return next(e)
169 })
170
171 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
172 const rest = await next(e)
173 const view = viewAt(await $.clock.now())
174 if (!view) return rest
175 const elements = $.ui.resolve(e)
176 // Only a warm cache has time left to draw as a bar
177 let gauge = 'none'
178 if (view.kind === 'warm') gauge = e.surface === 'desktop' ? 'svg' : fits(view, e.props.bodyColumns ?? 0) ? 'text' : 'none'
179 const line = gaugeLine(elements, gauge, view)
180 if (!rest) return line
181 return elements.Box({ flexDirection: 'column', children: [line, rest] })
182 })
183}
184
185function defaultTtl() {
186 return { ttl: '1h', estimated: true }
187}
188
189function emptyStats() {
190 return { input: 0, read: 0, misses: 0, lastMissCause: null }
191}
192
193function reset() {
194 refreshes += 1
195 timer?.cancel()
196 timer = null
197 last = null
198 model = null
199 stats = emptyStats()
200 compacted = false
201 agentsSeen = new Set()
202}
203
204// Whether a subagent's request was a fork's first and read `parent`, the main conversation's
205// cached prefix as it went out: the same model, and at least as much read as a request that hit it
206// would. Only an agent's first request counts, hit or miss: its later ones read its own entry.
207async function readsMainEntry($, agentId, parent, answeredBy, usage) {
208 if (agentsSeen.has(agentId)) return false
209 agentsSeen.add(agentId)
210 if (!parent?.cached) return false
211 if (parent.model && parent.model !== answeredBy) return false
212 if (usage.cache_read_input_tokens < parent.prefix * MISS_BELOW_SHARE) return false
213 const agent = (await $.agent.list()).find((a) => a.id === agentId)
214 return agent?.type === 'fork'
215}
216
217// Adds a main request to the conversation's figures. A request that reads well under what the
218// cache held after the one before re-processed what had been cached: a miss, unless /compact
219// rewrote the conversation, which rebuilds on purpose.
220function count(stepModel, sentAt, usage) {
221 const read = usage.cache_read_input_tokens
222 stats.input += usage.input_tokens + read + usage.cache_creation_input_tokens
223 stats.read += read
224 if (!last || !last.cached || compacted || read >= last.prefix * MISS_BELOW_SHARE) return
225 stats.misses += 1
226 if (last.model && last.model !== stepModel) stats.lastMissCause = 'model switch'
227 else if (last.expired || sentAt - last.at > TTL_MS[(last.ttl ?? ttl).ttl]) stats.lastMissCause = 'expired'
228 else stats.lastMissCause = null
229}
230
231// Takes up the TTL and the switches as they stand, redraws, and sets a timer for the moment the
232// gauge's text next changes. Never awaited by a hook, so it keeps its failures to itself: the next
233// event or timer tries again.
234async function refresh($) {
235 const started = ++refreshes
236 try {
237 const resolved = await resolveTtl($)
238 const off = (await $.env.get('DISABLE_PROMPT_CACHING')) === '1'
239 const now = await $.clock.now()
240 if (started !== refreshes) return
241 ttl = resolved
242 disabled = off
243 redraw($, now)
244 } catch {
245 // A module being unloaded is refused its calls
246 }
247}
248
249function redraw($, now) {
250 timer?.cancel()
251 timer = null
252 $.ui.invalidate('ui.render')
253 const view = viewAt(now)
254 if (view?.kind !== 'warm') return
255 // Minutes count down by the minute, the last one by the second, and the last tick lands on the
256 // expiry itself
257 const step = view.remaining > 60_000 ? 60_000 : 1_000
258 timer = $.clock.after(view.remaining % step || step, () => {
259 void $.clock.now().then(
260 (at) => redraw($, at),
261 () => {},
262 )
263 })
264}
265
266// What the gauge shows at `now`, or null before there is anything to show
267function viewAt(now) {
268 if (disabled) return { kind: 'off', reason: 'DISABLE_PROMPT_CACHING' }
269 if (compacted) return { kind: 'rebuilding' }
270 if (!last) return null
271 if (!last.cached) return { kind: 'off', reason: 'no cache tokens reported' }
272 if (last.model && model && last.model !== model) return { kind: 'cold', reason: 'model switch', tokens: last.tokens }
273 const own = last.ttl ?? ttl
274 const ttlMs = TTL_MS[own.ttl]
275 const remaining = last.at + ttlMs - now
276 if (last.expired || remaining <= 0) return { kind: 'cold', reason: null, tokens: last.tokens }
277 return {
278 kind: 'warm',
279 remaining,
280 share: remaining / ttlMs,
281 ttl: (own.estimated ? '~' : '') + own.ttl,
282 hit: stats.input > 0 ? Math.round((stats.read / stats.input) * 100) : null,
283 misses: stats.misses,
284 lastMissCause: stats.lastMissCause,
285 }
286}
287
288// The gauge's runs of text with their styles: the bar goes after `head`, `value` in the status
289// color after it, and `tail` in the default color last
290function parts(view) {
291 if (view.kind === 'warm') {
292 const status = view.share < WARN_BELOW_SHARE ? 'warning' : 'success'
293 const after = []
294 if (view.hit != null) after.push('hit ' + view.hit + '%')
295 after.push('misses ' + view.misses + (view.misses > 0 && view.lastMissCause ? ' (' + view.lastMissCause + ')' : ''))
296 return {
297 status,
298 head: { text: 'cache ● ' + view.ttl, style: { color: status } },
299 value: { text: timeLeft(view.remaining) + ' left', style: { color: status } },
300 tail: ' · ' + after.join(' · '),
301 }
302 }
303 if (view.kind === 'cold') {
304 return {
305 status: 'error',
306 head: { text: 'cache ○ cold', style: { color: 'error' } },
307 tail: ' · ' + (view.reason ? view.reason + ' · ' : '') + 'next message re-caches ' + tokenCount(view.tokens) + ' tokens',
308 }
309 }
310 if (view.kind === 'rebuilding') {
311 return {
312 status: 'warning',
313 head: { text: 'cache ○ rebuilding', style: { color: 'warning' } },
314 tail: ' · next message caches the /compact summary',
315 }
316 }
317 return { status: null, head: { text: 'cache ○ off', style: { dimColor: true } }, tail: ' · ' + view.reason }
318}
319
320// Whether the line fits with its text bar; every character drawn is one cell wide
321function fits(view, columns) {
322 const { head, value, tail } = parts(view)
323 const width = [...head.text].length + 1 + BAR_CELLS + 1 + (value ? [...value.text].length : 0) + [...tail].length
324 return width <= columns - BAND_RESERVED_COLUMNS
325}
326
327function gaugeLine({ Box, Text, Svg }, gauge, view) {
328 const { status, head, value, tail } = parts(view)
329 const share = view.share ?? 0
330 const headText = Text({ ...head.style, children: [head.text] })
331 // Without a value there is no bar either: the head and what follows stay one run
332 if (!value) return Box({ key: 'cache-meter', flexDirection: 'row', children: [Text({ children: [headText, tail] })] })
333 const children = [headText]
334 if (gauge === 'svg') {
335 children.push(
336 Svg({
337 source: svgBar(share, status),
338 alt: head.text + ' ' + value.text + tail,
339 width: SVG_BAR.width,
340 height: SVG_BAR.height,
341 }),
342 )
343 } else if (gauge === 'text') {
344 children.push(textBar(Text, share, status))
345 }
346 // The value and what follows it stay one run, so no gap opens before the separator
347 children.push(Text({ children: [Text({ ...value.style, children: [value.text] }), tail] }))
348 return Box({
349 key: 'cache-meter',
350 flexDirection: 'row',
351 columnGap: 1,
352 alignItems: 'center',
353 ...(gauge === 'svg' && { flexWrap: 'wrap' }),
354 children,
355 })
356}
357
358// The bar as two runs: the share of the TTL left in the status color, the rest dim
359function textBar(Text, share, status) {
360 const filled = Math.round(Math.min(Math.max(share, 0), 1) * BAR_CELLS)
361 const runs = []
362 if (filled > 0) runs.push(Text({ color: status ?? 'success', children: ['█'.repeat(filled)] }))
363 if (filled < BAR_CELLS) runs.push(Text({ dimColor: true, children: ['░'.repeat(BAR_CELLS - filled)] }))
364 return Text({ children: runs })
365}
366
367function svgBar(share, status) {
368 const { width, height } = SVG_BAR
369 const r = height / 2
370 const fill = Math.round(Math.min(Math.max(share, 0), 1) * width)
371 const parts = [
372 `<svg xmlns="http://www.w3.org/2000/svg" width="${width}" height="${height}" viewBox="0 0 ${width} ${height}">`,
373 `<clipPath id="c"><rect width="${width}" height="${height}" rx="${r}"/></clipPath>`,
374 `<g clip-path="url(#c)">`,
375 `<rect width="${width}" height="${height}" fill="${SVG_COLORS.track}"/>`,
376 ]
377 if (fill > 0) parts.push(`<rect width="${fill}" height="${height}" fill="${SVG_COLORS[status ?? 'success']}"/>`)
378 parts.push('</g>', '</svg>')
379 return parts.join('')
380}
381
382// The main conversation's TTL, in the order Claude Code takes it: the engine's own word at a model
383// switch, then FORCE_PROMPT_CACHING_5M, CLAUDE_CODE_PROMPT_CACHE_TTL, the promptCacheTtl setting
384// and ENABLE_PROMPT_CACHING_1H. Without any of them, the default is estimated: one hour on a
385// subscription within its plan usage, five minutes past it and on an API key or cloud provider.
386async function resolveTtl($) {
387 if (engineTtl && engineTtl.overLimit === overLimit()) return { ttl: engineTtl.ttl, estimated: false }
388 if ((await $.env.get('FORCE_PROMPT_CACHING_5M')) === '1') return { ttl: '5m', estimated: false }
389 const fromEnv = await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL')
390 if (isTtl(fromEnv)) return { ttl: fromEnv, estimated: false }
391 const fromSettings = (await $.settings.read()).promptCacheTtl
392 if (isTtl(fromSettings)) return { ttl: fromSettings, estimated: false }
393 if ((await $.env.get('ENABLE_PROMPT_CACHING_1H')) === '1') return { ttl: '1h', estimated: false }
394 const auth = await $.session.authorize()
395 return { ttl: auth?.kind === 'bearer' && !overLimit() ? '1h' : '5m', estimated: true }
396}
397
398// Only the two values Claude Code takes; it ignores any other
399function isTtl(value) {
400 return value === '5m' || value === '1h'
401}
402
403function overLimit() {
404 return rateLimits.some((limit) => PLAN_WINDOWS.includes(limit.kind) && limit.percentUsed >= 100)
405}
406
407function timeLeft(ms) {
408 if (ms > 60_000) return Math.ceil(ms / 60_000) + 'm'
409 return Math.ceil(ms / 1_000) + 's'
410}
411
412function tokenCount(tokens) {
413 if (tokens >= 1_000_000) return (tokens / 1_000_000).toFixed(1) + 'M'
414 if (tokens >= 1_000) return Math.round(tokens / 1_000) + 'k'
415 return String(tokens)
416}
417