SLOPSHOPPER

cache-meter

Gauges how long the conversation's prompt cache stays warm, its hit ratio and misses, and what the next message re-caches once it has gone cold

newbandtimer
★ 17v0.1.0no licenseupdated 2026-10-07HolyGrail/claude-mods/plugins/cache-meter
A shopper browsing a rack in a slop shop
README

cache-meter

会話のプロンプトキャッシュがあと何分もつかを、プロンプト入力欄の上の帯にゲージで常時表示する Claude Code の mod です。 キャッシュが切れたあとは、次のメッセージで何トークンを書き直すことになるかを表示します。

cache ● ~1h ████████░░ 48m left · hit 91% · misses 0
cache ● ~1h █░░░░░░░░░ 53s left · hit 91% · misses 0
cache ○ cold · next message re-caches 128k tokens

Claude Code はメッセージのたびに会話全体を送り直していて、キャッシュはその読み直しを安くしています。 ところが、キャッシュの残り時間は、/usage を開くかステータスラインのスクリプトを書かない限り目に入りません。 席を外して戻ってきた直後の 1 ターンが遅く、使用量も多く消費するのは、その間に TTL が切れていたからです。 残りが少ないときに聞きたいことがあるなら、切れる前に聞いておくほうが安く済みます。

表示

  • ゲージ:TTL のうち残っている割合を塗る。Desktop アプリでは SVG、ターミナルでは文字のバーで描き、ターミナルで 1 行に収まらないときはバーを省く。usage-meter など、ほかの mod の行とは別の行に描く
  • 色:残りが TTL の 2 割以上なら緑、2 割を切ると黄、切れたら赤
  • 残り時間:1 分ごとに減り、最後の 1 分は秒で数える。切れた瞬間に cold へ変わる
  • hit:この会話のメインのリクエストで、入力トークンのうちキャッシュから読んだ割合
  • misses:直前のリクエストのあとにキャッシュが持っていた量の半分も読めなかったリクエストの数。原因が分かるとき(TTL 切れなら expired、モデルの切替なら model switch)は括弧で添える。/compact のあとの作り直しは数えない

キャッシュが温まっていない状態は、次のように表示します。

表示状態
cache ○ cold · next message re-caches N tokensTTL を過ぎた。N は直前の応答の入力(キャッシュの読み書きを含む)と出力の合計
cache ○ cold · model switch · …モデルを切り替えた。キャッシュはモデルごとなので、次の送信は会話全体を書き直す
cache ○ rebuilding · …/compact のあと。会話部分のキャッシュは、次の送信で要約から作り直される
cache ○ off · no cache tokens reported応答にキャッシュの読み書きがなかった。キャッシュを無効にしているか、ゲートウェイが報告していない
cache ○ off · DISABLE_PROMPT_CACHING環境変数でキャッシュを切っている

最初の応答が返るまでと、/clear の直後は何も表示しません。

TTL の決め方

残り時間は、メイン会話の最後のリクエストを送った時刻から数えます。 キャッシュを読んだリクエストは、そのたびにタイマーを戻すからです。 サブエージェントやワークフローは別のキャッシュを持ち、TTL も別なので、そのリクエストでは数え直しません。 ただし fork は会話をそのまま引き継ぐので、最初のリクエストがメイン会話のキャッシュを読み、残り時間を戻します。

TTL は、Claude Code が採用するのと同じ優先順で決めます。

  1. モデル切替のときに Claude Code 自身が知らせる TTL(PostModelSwitch の cache_ttl)
  2. FORCE_PROMPT_CACHING_5M=1 なら 5 分
  3. 環境変数 CLAUDE_CODE_PROMPT_CACHE_TTL(5m か 1h)
  4. 設定 promptCacheTtl(5m か 1h)
  5. ENABLE_PROMPT_CACHING_1H=1 なら 1 時間
  6. どれもなければ既定値。サブスクリプションでプランの利用枠内なら 1 時間、利用枠を超えて usage credits で払っている場合と、API キーやクラウドプロバイダーなら 5 分

6 の既定値だけは推定なので、TTL の前に ~ を付けます(~1h)。 mods API はプランの種類を直接は渡さないため、OAuth ログインならサブスクリプション、5 時間か 7 日の制限が 100% に達していたら usage credits、と見なしています。 1 は実際の値ですが、モデルを切り替えたときにしか届きません。

TTL はリクエストごとに、応答が返った時点で決めて記録します。 Claude Code もリクエストごとに TTL を選ぶので、usage credits で書いた 5 分のキャッシュは、そのあと利用枠が戻っても 5 分で切れます。

使い方

Claude Code v2.1.292 で、読み込みと、実際の応答からの表示を確認しています。 /plugin でのインストールと更新の手順は、リポジトリの README にまとめてあります。 インストールせずに試すときは、clone したリポジトリのルートで次のように起動します。

claude --plugin-dir ./plugins/cache-meter

制約

次の操作は mod から見えないので、表示に反映されません。

  • /rewind:巻き戻した先の接頭辞はそれ以後のリクエストが読み続けているので、キャッシュは残る。ただし再キャッシュのトークン数は巻き戻す前の大きさのまま表示される
  • opusplan でのプランモードの出入り:モデルの切替にあたるが、切替のイベントが来ない場合は、次の応答が返るまで前のモデルの残り時間を表示する
  • effort の変更、fast mode を有効にした直後、MCP サーバーの接続:多くの場合キャッシュを失うが、表示は次の応答まで変わらない
  • DISABLE_PROMPT_CACHING_SONNET などのモデル別の無効化:そのファミリーのエイリアスが指すモデルにだけ効くが、mods API からはエイリアスの解決先が分からない。そのモデルに切り替えた直後は cold と表示し、最初の応答で off に変わる

公式のステータスラインスクリプトなら、prompt_cache オブジェクト(v2.1.251 以降)から expires_at や recache_tokens_if_cold を Claude Code 自身の計算で受け取れます。 この mod は、その値に手が届かない mods API の中で、同じものを応答のトークン数から組み立てています。

前提にした仕様

すべて How Claude Code uses prompt caching によります。

  • キャッシュを読んだリクエストは、そのたびに TTL のタイマーを戻す
  • サブスクリプションでプランの利用枠内なら、メイン会話は 1 時間、サブエージェントやワークフローは 5 分。usage credits、API キー、クラウドプロバイダーではメイン会話も 5 分
  • キャッシュはモデルごとで、opusplan ではプランモードの出入りがモデルの切替になる
  • /compact は会話部分のキャッシュを作り直し、/rewind は巻き戻した先のキャッシュをそのまま読む

テスト

cd plugins/cache-meter
claude plugin test
Source 1 files
hooks/register.js 417 lines
1// Draws a gauge in the band above the prompt for the main conversation's prompt cache: how long
2// it stays warm, its hit ratio and misses, and once it has gone cold, how many tokens the next
3// message writes to the cache again.
4
5// The last main-thread request that came back with usage: { seq, at, ttl, model, tokens, prefix, cached, expired }
6// - seq: which main request it was, counted across the module's life, so a fork can tell whether
7//   the entry it read is still the newest
8// - at: when it was sent, in $.clock.now() milliseconds. Each request that reads the cache resets
9//   its TTL, so the TTL runs from here; the response may end minutes later.
10// - ttl: the TTL resolved as it came back, { ttl, estimated }, since Claude Code picks one per
11//   request; null when unknown (a resumed conversation), which takes the TTL as it stands now
12// - model: the model that answered it, null when unknown (a resumed conversation); each model has
13//   a cache of its own
14// - tokens: what the next request re-sends, input + cache read + cache write + output, as the
15//   engine counts it for a model switch
16// - prefix: what the cache held after it, cache read + cache write; the next request reads about
17//   this much when the cache is still there
18// - cached: whether the response reported any cache tokens at all
19// - expired: the engine said, on resume, that the cache had likely gone cold
20let last = null
21// The model the main loop runs now, as the last step or model switch named it
22let model = null
23// This conversation's main requests: input tokens in all and from the cache, and the requests
24// that wrote back what the cache had held, with the likely cause of the last one
25let stats = emptyStats()
26// The TTL the engine stated at the last model switch: { ttl, overLimit }, kept only while the
27// plan usage stays on the side of the limit it was on then
28let engineTtl = null
29// The TTL as last resolved, { ttl, estimated }, for a request whose own is unknown, and whether
30// caching is switched off
31let ttl = defaultTtl()
32let disabled = false
33// The rate-limit windows, for the TTL estimate: past a window's limit, usage credits pay
34let rateLimits = []
35// After /compact, the next request rebuilds the conversation layer: not a miss
36let compacted = false
37// The subagents whose first request has been seen: only a fork's first reads the main
38// conversation's entry, hit or miss
39let agentsSeen = new Set()
40// Numbers the main requests (`last.seq`)
41let mainRequests = 0
42// The timer that redraws the gauge when its text next changes
43let timer = null
44// Counts refreshes and resets, so only the latest refresh redraws: one begun before a later one,
45// or before /clear and the rest, leaves the gauge alone
46let refreshes = 0
47
48const TTL_MS = { '5m': 5 * 60_000, '1h': 60 * 60_000 }
49// The session.end reasons after which this module stops; /clear, /resume and logout leave it
50// running
51const FINAL_REASONS = ['prompt_input_exit', 'other']
52// The windows a subscription's plan usage is measured in; past either, Claude Code draws on usage
53// credits and drops the main conversation to five minutes
54const PLAN_WINDOWS = ['five_hour', 'seven_day']
55// A request that read less than this share of what the cache held before it is a miss
56const MISS_BELOW_SHARE = 0.5
57// Under this share of the TTL left, the gauge turns yellow
58const WARN_BELOW_SHARE = 0.2
59
60const BAR_CELLS = 10
61// The band's last column, which the terminal may draw over
62const BAND_RESERVED_COLUMNS = 2
63const SVG_BAR = { width: 96, height: 10 }
64const SVG_COLORS = { success: '#4caf50', warning: '#e0a526', error: '#e5534b', track: 'rgba(128,128,128,0.3)' }
65
66export function register(on) {
67  // Fires again on an enable or a worker respawn, which may keep this module's variables
68  on('session.start', async ($, e, next) => {
69    reset()
70    engineTtl = null
71    ttl = defaultTtl()
72    disabled = false
73    rateLimits = (await $.session.usage()).rateLimits
74    void refresh($)
75    return next(e)
76  })
77
78  on('session.end', async ($, e, next) => {
79    if (!FINAL_REASONS.includes(e.reason)) return next(e)
80    refreshes += 1
81    timer?.cancel()
82    timer = null
83    return next(e)
84  })
85
86  on('session.measure', async ($, e, next) => {
87    if (e.changed.includes('rateLimits')) {
88      rateLimits = e.rateLimits
89      void refresh($)
90    }
91    return next(e)
92  })
93
94  // The main loop's requests (no agentId) move the gauge. Subagents and workflows keep caches of
95  // their own, with their own TTL; a fork inherits the main conversation whole, so its first
96  // request reads the main entry and resets its timer.
97  on('turn.step', async function* ($, e, next) {
98    const sentAt = await $.clock.now()
99    // The main entry as this request went out, which a fork's request read if it read any
100    const parent = last
101    const response = yield* next(e)
102    const usage = response?.usage
103    if (!usage) return response
104    // A later mod may have sent the request to another model than the step named
105    const answeredBy = usage.model || e.model
106    if (e.agentId) {
107      // Forks that overlap may answer out of order: the newest read stands, and a main request
108      // recorded since holds an entry the fork did not read
109      const read = await readsMainEntry($, e.agentId, parent, answeredBy, usage).catch(() => false)
110      if (read && last?.seq === parent.seq && sentAt > last.at) {
111        last = { ...last, at: sentAt }
112        void refresh($)
113      }
114      return response
115    }
116    // Unresolved, the request takes the TTL as it stands when drawn
117    const resolved = await resolveTtl($).catch(() => null)
118    const read = usage.cache_read_input_tokens
119    const written = usage.cache_creation_input_tokens
120    count(answeredBy, sentAt, usage)
121    last = {
122      seq: ++mainRequests,
123      at: sentAt,
124      ttl: resolved,
125      model: answeredBy,
126      tokens: usage.input_tokens + read + written + usage.output_tokens,
127      prefix: read + written,
128      cached: read + written > 0,
129      expired: false,
130    }
131    model = answeredBy
132    compacted = false
133    void refresh($)
134    return response
135  })
136
137  // The engine states the TTL it uses here, which no other event carries
138  on('classic.PostModelSwitch', async ($, e, next) => {
139    model = e.to_model
140    engineTtl = { ttl: e.cache_ttl, overLimit: overLimit() }
141    if (last && e.context_tokens > 0) last = { ...last, tokens: e.context_tokens }
142    void refresh($)
143    return next(e)
144  })
145
146  // /clear starts a conversation with nothing cached for it yet. /compact replaces the history,
147  // so the next request writes the summary's cache. A resumed or forked conversation carries how
148  // long ago it last had a response, which dates its cache.
149  on('classic.SessionStart', { source: ['clear', 'compact', 'resume', 'fork'] }, async ($, e, next) => {
150    if (e.source === 'compact') {
151      compacted = true
152    } else {
153      reset()
154      if (typeof e.seconds_since_last_response === 'number' && typeof e.context_tokens === 'number') {
155        last = {
156          seq: ++mainRequests,
157          at: (await $.clock.now()) - e.seconds_since_last_response * 1000,
158          ttl: null,
159          model: null,
160          tokens: e.context_tokens,
161          prefix: e.context_tokens,
162          cached: true,
163          expired: e.prompt_cache_likely_expired === true,
164        }
165      }
166    }
167    void refresh($)
168    return next(e)
169  })
170
171  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
172    const rest = await next(e)
173    const view = viewAt(await $.clock.now())
174    if (!view) return rest
175    const elements = $.ui.resolve(e)
176    // Only a warm cache has time left to draw as a bar
177    let gauge = 'none'
178    if (view.kind === 'warm') gauge = e.surface === 'desktop' ? 'svg' : fits(view, e.props.bodyColumns ?? 0) ? 'text' : 'none'
179    const line = gaugeLine(elements, gauge, view)
180    if (!rest) return line
181    return elements.Box({ flexDirection: 'column', children: [line, rest] })
182  })
183}
184
185function defaultTtl() {
186  return { ttl: '1h', estimated: true }
187}
188
189function emptyStats() {
190  return { input: 0, read: 0, misses: 0, lastMissCause: null }
191}
192
193function reset() {
194  refreshes += 1
195  timer?.cancel()
196  timer = null
197  last = null
198  model = null
199  stats = emptyStats()
200  compacted = false
201  agentsSeen = new Set()
202}
203
204// Whether a subagent's request was a fork's first and read `parent`, the main conversation's
205// cached prefix as it went out: the same model, and at least as much read as a request that hit it
206// would. Only an agent's first request counts, hit or miss: its later ones read its own entry.
207async function readsMainEntry($, agentId, parent, answeredBy, usage) {
208  if (agentsSeen.has(agentId)) return false
209  agentsSeen.add(agentId)
210  if (!parent?.cached) return false
211  if (parent.model && parent.model !== answeredBy) return false
212  if (usage.cache_read_input_tokens < parent.prefix * MISS_BELOW_SHARE) return false
213  const agent = (await $.agent.list()).find((a) => a.id === agentId)
214  return agent?.type === 'fork'
215}
216
217// Adds a main request to the conversation's figures. A request that reads well under what the
218// cache held after the one before re-processed what had been cached: a miss, unless /compact
219// rewrote the conversation, which rebuilds on purpose.
220function count(stepModel, sentAt, usage) {
221  const read = usage.cache_read_input_tokens
222  stats.input += usage.input_tokens + read + usage.cache_creation_input_tokens
223  stats.read += read
224  if (!last || !last.cached || compacted || read >= last.prefix * MISS_BELOW_SHARE) return
225  stats.misses += 1
226  if (last.model && last.model !== stepModel) stats.lastMissCause = 'model switch'
227  else if (last.expired || sentAt - last.at > TTL_MS[(last.ttl ?? ttl).ttl]) stats.lastMissCause = 'expired'
228  else stats.lastMissCause = null
229}
230
231// Takes up the TTL and the switches as they stand, redraws, and sets a timer for the moment the
232// gauge's text next changes. Never awaited by a hook, so it keeps its failures to itself: the next
233// event or timer tries again.
234async function refresh($) {
235  const started = ++refreshes
236  try {
237    const resolved = await resolveTtl($)
238    const off = (await $.env.get('DISABLE_PROMPT_CACHING')) === '1'
239    const now = await $.clock.now()
240    if (started !== refreshes) return
241    ttl = resolved
242    disabled = off
243    redraw($, now)
244  } catch {
245    // A module being unloaded is refused its calls
246  }
247}
248
249function redraw($, now) {
250  timer?.cancel()
251  timer = null
252  $.ui.invalidate('ui.render')
253  const view = viewAt(now)
254  if (view?.kind !== 'warm') return
255  // Minutes count down by the minute, the last one by the second, and the last tick lands on the
256  // expiry itself
257  const step = view.remaining > 60_000 ? 60_000 : 1_000
258  timer = $.clock.after(view.remaining % step || step, () => {
259    void $.clock.now().then(
260      (at) => redraw($, at),
261      () => {},
262    )
263  })
264}
265
266// What the gauge shows at `now`, or null before there is anything to show
267function viewAt(now) {
268  if (disabled) return { kind: 'off', reason: 'DISABLE_PROMPT_CACHING' }
269  if (compacted) return { kind: 'rebuilding' }
270  if (!last) return null
271  if (!last.cached) return { kind: 'off', reason: 'no cache tokens reported' }
272  if (last.model && model && last.model !== model) return { kind: 'cold', reason: 'model switch', tokens: last.tokens }
273  const own = last.ttl ?? ttl
274  const ttlMs = TTL_MS[own.ttl]
275  const remaining = last.at + ttlMs - now
276  if (last.expired || remaining <= 0) return { kind: 'cold', reason: null, tokens: last.tokens }
277  return {
278    kind: 'warm',
279    remaining,
280    share: remaining / ttlMs,
281    ttl: (own.estimated ? '~' : '') + own.ttl,
282    hit: stats.input > 0 ? Math.round((stats.read / stats.input) * 100) : null,
283    misses: stats.misses,
284    lastMissCause: stats.lastMissCause,
285  }
286}
287
288// The gauge's runs of text with their styles: the bar goes after `head`, `value` in the status
289// color after it, and `tail` in the default color last
290function parts(view) {
291  if (view.kind === 'warm') {
292    const status = view.share < WARN_BELOW_SHARE ? 'warning' : 'success'
293    const after = []
294    if (view.hit != null) after.push('hit ' + view.hit + '%')
295    after.push('misses ' + view.misses + (view.misses > 0 && view.lastMissCause ? ' (' + view.lastMissCause + ')' : ''))
296    return {
297      status,
298      head: { text: 'cache ● ' + view.ttl, style: { color: status } },
299      value: { text: timeLeft(view.remaining) + ' left', style: { color: status } },
300      tail: ' · ' + after.join(' · '),
301    }
302  }
303  if (view.kind === 'cold') {
304    return {
305      status: 'error',
306      head: { text: 'cache ○ cold', style: { color: 'error' } },
307      tail: ' · ' + (view.reason ? view.reason + ' · ' : '') + 'next message re-caches ' + tokenCount(view.tokens) + ' tokens',
308    }
309  }
310  if (view.kind === 'rebuilding') {
311    return {
312      status: 'warning',
313      head: { text: 'cache ○ rebuilding', style: { color: 'warning' } },
314      tail: ' · next message caches the /compact summary',
315    }
316  }
317  return { status: null, head: { text: 'cache ○ off', style: { dimColor: true } }, tail: ' · ' + view.reason }
318}
319
320// Whether the line fits with its text bar; every character drawn is one cell wide
321function fits(view, columns) {
322  const { head, value, tail } = parts(view)
323  const width = [...head.text].length + 1 + BAR_CELLS + 1 + (value ? [...value.text].length : 0) + [...tail].length
324  return width <= columns - BAND_RESERVED_COLUMNS
325}
326
327function gaugeLine({ Box, Text, Svg }, gauge, view) {
328  const { status, head, value, tail } = parts(view)
329  const share = view.share ?? 0
330  const headText = Text({ ...head.style, children: [head.text] })
331  // Without a value there is no bar either: the head and what follows stay one run
332  if (!value) return Box({ key: 'cache-meter', flexDirection: 'row', children: [Text({ children: [headText, tail] })] })
333  const children = [headText]
334  if (gauge === 'svg') {
335    children.push(
336      Svg({
337        source: svgBar(share, status),
338        alt: head.text + ' ' + value.text + tail,
339        width: SVG_BAR.width,
340        height: SVG_BAR.height,
341      }),
342    )
343  } else if (gauge === 'text') {
344    children.push(textBar(Text, share, status))
345  }
346  // The value and what follows it stay one run, so no gap opens before the separator
347  children.push(Text({ children: [Text({ ...value.style, children: [value.text] }), tail] }))
348  return Box({
349    key: 'cache-meter',
350    flexDirection: 'row',
351    columnGap: 1,
352    alignItems: 'center',
353    ...(gauge === 'svg' && { flexWrap: 'wrap' }),
354    children,
355  })
356}
357
358// The bar as two runs: the share of the TTL left in the status color, the rest dim
359function textBar(Text, share, status) {
360  const filled = Math.round(Math.min(Math.max(share, 0), 1) * BAR_CELLS)
361  const runs = []
362  if (filled > 0) runs.push(Text({ color: status ?? 'success', children: ['█'.repeat(filled)] }))
363  if (filled < BAR_CELLS) runs.push(Text({ dimColor: true, children: ['░'.repeat(BAR_CELLS - filled)] }))
364  return Text({ children: runs })
365}
366
367function svgBar(share, status) {
368  const { width, height } = SVG_BAR
369  const r = height / 2
370  const fill = Math.round(Math.min(Math.max(share, 0), 1) * width)
371  const parts = [
372    `<svg xmlns="http://www.w3.org/2000/svg" width="${width}" height="${height}" viewBox="0 0 ${width} ${height}">`,
373    `<clipPath id="c"><rect width="${width}" height="${height}" rx="${r}"/></clipPath>`,
374    `<g clip-path="url(#c)">`,
375    `<rect width="${width}" height="${height}" fill="${SVG_COLORS.track}"/>`,
376  ]
377  if (fill > 0) parts.push(`<rect width="${fill}" height="${height}" fill="${SVG_COLORS[status ?? 'success']}"/>`)
378  parts.push('</g>', '</svg>')
379  return parts.join('')
380}
381
382// The main conversation's TTL, in the order Claude Code takes it: the engine's own word at a model
383// switch, then FORCE_PROMPT_CACHING_5M, CLAUDE_CODE_PROMPT_CACHE_TTL, the promptCacheTtl setting
384// and ENABLE_PROMPT_CACHING_1H. Without any of them, the default is estimated: one hour on a
385// subscription within its plan usage, five minutes past it and on an API key or cloud provider.
386async function resolveTtl($) {
387  if (engineTtl && engineTtl.overLimit === overLimit()) return { ttl: engineTtl.ttl, estimated: false }
388  if ((await $.env.get('FORCE_PROMPT_CACHING_5M')) === '1') return { ttl: '5m', estimated: false }
389  const fromEnv = await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL')
390  if (isTtl(fromEnv)) return { ttl: fromEnv, estimated: false }
391  const fromSettings = (await $.settings.read()).promptCacheTtl
392  if (isTtl(fromSettings)) return { ttl: fromSettings, estimated: false }
393  if ((await $.env.get('ENABLE_PROMPT_CACHING_1H')) === '1') return { ttl: '1h', estimated: false }
394  const auth = await $.session.authorize()
395  return { ttl: auth?.kind === 'bearer' && !overLimit() ? '1h' : '5m', estimated: true }
396}
397
398// Only the two values Claude Code takes; it ignores any other
399function isTtl(value) {
400  return value === '5m' || value === '1h'
401}
402
403function overLimit() {
404  return rateLimits.some((limit) => PLAN_WINDOWS.includes(limit.kind) && limit.percentUsed >= 100)
405}
406
407function timeLeft(ms) {
408  if (ms > 60_000) return Math.ceil(ms / 60_000) + 'm'
409  return Math.ceil(ms / 1_000) + 's'
410}
411
412function tokenCount(tokens) {
413  if (tokens >= 1_000_000) return (tokens / 1_000_000).toFixed(1) + 'M'
414  if (tokens >= 1_000) return Math.round(tokens / 1_000) + 'k'
415  return String(tokens)
416}
417