SLOPSHOPPER

cache-timer

A countdown to when the conversation's prompt cache expires, in the prompt footer before the model name

newspinnerprocesstimer
★ 1v0.1.0no licenseupdated 2026-10-03EricJamie/claude-code-mods/plugins/cache-timer
A shopper browsing a rack in a slop shop
README

claude-code-mods

Mods for Claude Code, distributed as a plugin marketplace:

  • context-band: a band above the prompt with rate limits, tokens, speed, cache, cost and context
  • cache-timer: a countdown to when the conversation's prompt cache expires, before the model name

And the band for other agents: Grok Build (a status line script) and Codex (built-in status line settings).

中文说明

context-band

A band above the prompt, in the desktop app and the CLI, that updates after every turn:

  • 5h / 7d rate-limit usage and the time until each resets
  • in / out tokens, t/s output speed, cache tokens and hit rate
  • $ the session's cost at API prices, ctx how full the context window is

<img alt="The band: 5h and 7d limits, tokens in and out, speed, cache, cost and context" src="docs/screenshots/band-light.png" width="800">

The band stays on one line. When some pills don't fit, a +N button at the end expands it to show every pill, and Less folds it back.

Hover the 5h or 7d pill for a one-line estimate of what that window is worth at API prices. The 📈 button opens two views of both windows:

  • Chart: where the window ends at your pace (or when you hit the limit), your average spend rate beside the rate that lands on 100% at the reset, and a strip chart against the limit and last week
  • By model: for each model, the tokens left if you use only that model (with bars to compare), and the tokens it has used

<img alt="Chart view: where each window ends at your pace, and the spend rate that lands on 100% at the reset" src="docs/screenshots/chart-light.png" width="800">

<img alt="By model view: tokens left in each window if you use only one model, and tokens each model used" src="docs/screenshots/models-light.png" width="800">

<sub>Screenshots use demo figures.</sub>

In the terminal, 📈 shows the same two views as small tables, one row per window with the columns lined up, and drops the least important columns when the window is narrow.

Installed or reloaded partway through a session (/reload-plugins), the band starts from the tokens the session has already used, read from its transcript; t/s appears after the next reply.

The band follows the app's light or dark theme; ◐ (in the 📈 panel) cycles auto, light and dark.

Install

In Claude Code (CLI or desktop):

/plugin marketplace add EricJamie/claude-code-mods
/plugin install context-band@claude-code-mods

It loads in the next session.

Commands

/context-band auto|light|dark theme · hide / show the band · reset the token counters

How the estimates work

  • A window runs from its reset time back 5 hours or 7 days.
  • $ at 100% = what Claude Code spent in the window at API list prices ÷ the % of the window used. Under 5% used it is marked "(rough)".
  • Tokens left if one model does it all = the dollars left in the window ÷ what a token costs on that model at your own mix of input, output and cache (your last 7 days, priced as if they had all gone to that model).
  • Spend comes from the Claude Code transcripts on this machine; usage elsewhere (claude.ai, other machines) counts toward the % but not the $, so the estimates read low if you use those a lot. The estimates assume the limits weigh models by API price.

Requirements and privacy

  • python3 (the estimator, plugins/context-band/bin/api_estimate.py)
  • macOS for the automatic theme and the desktop usage history; elsewhere those fall back quietly
  • Everything stays on your machine: the mod reads local files and makes no network calls. Its cache is ~/.cache/context-band/.

Prices and new models

List prices live in PRICES in bin/api_estimate.py, from Anthropic's pricing page (https://platform.claude.com/docs/en/about-claude/pricing).

A model the table doesn't know still works: it gets its own row, priced like its family's current model (a new family is priced like Opus), and its figures are marked ≈ until its price is learned. The band learns it from Claude Code's own cost figure for each session, which always uses current prices, set against the tokens each model used in that session. The same check corrects a listed price that has changed. Updating the table when a model ships is still the quickest fix.

Development

claude plugin validate plugins/context-band
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude plugin test plugins/context-band

Mods are an early-access Claude Code feature; the test runner needs that variable.

cache-timer

A countdown in the prompt footer, before the model name: cache 41:41 until the conversation's prompt cache expires. It turns orange in the last sixth of the entry's life (10 minutes of an hour), red in the last 2 minutes, and says expired after.

/plugin install cache-timer@claude-code-mods

Claude Code caches the conversation so far, so each message reads it back instead of sending it again. A cache entry lives a fixed time from the start of the last request that read or wrote it, and every message restarts that time. The main conversation uses 1-hour entries on some plans and 5-minute ones on others (subagents use 5-minute ones); the timer reads which from the session's transcript, so it needs python3.

Reading the cache is cheap but not free: on Opus 5.5 a cache read costs $0.20 per million tokens against $4 for fresh input, and writing costs 1.25× the input price for a 5-minute entry or 2× for a 1-hour one. Once an entry expires, the next message writes the whole context again: for a 180k-token Opus 5.5 conversation that is ≈ $1.44 instead of a ≈ $0.04 read.

The countdown starts from the main conversation's own requests. Requests it does not see (some background ones) can refresh the cache too, so the real expiry can be a little later than shown.


中文说明

Claude Code 的插件(mod)集合,以插件市场的形式发布:

  • context-band:输入框上方的状态栏,显示用量限额、token、速度、缓存、花费和上下文
  • cache-timer:在模型名称前面显示对话缓存到期的倒计时

其他工具的状态栏:Grok Build(状态行脚本)和 Codex(内置状态行配置)。

context-band 状态栏

显示在输入框上方(桌面端和命令行都支持),每轮对话后自动更新:

  • 5h / 7d:5 小时和 7 天用量限额的使用比例,以及距离重置的时间
  • in / out:输入、输出 token;t/s:输出速度;cache:缓存 token 和命中率
  • $:本次会话按 API 价格计算的花费;ctx:上下文窗口的使用比例

<img alt="The band: 5h and 7d limits, tokens in and out, speed, cache, cost and context" src="docs/screenshots/band-light.png" width="800">

状态栏默认只占一行。放不下的指标会收进末尾的 +N 按钮,点击即可展开显示全部,点击 Less 收起。

鼠标悬停在 5h 或 7d 上,会显示该窗口按 API 价格折算的估值。点击 📈 打开两个视图:

  • Chart(图表):按当前速度到重置时会用到多少(或何时触顶)、目前的平均花费速度和刚好在重置时用满的速度,以及与额度线和上周对比的小图
  • By model(按模型):如果只用某个模型还能用多少 token(带对比条),以及该模型已用的 token

<img alt="Chart view: where each window ends at your pace, and the spend rate that lands on 100% at the reset" src="docs/screenshots/chart-light.png" width="800">

<img alt="By model view: tokens left in each window if you use only one model, and tokens each model used" src="docs/screenshots/models-light.png" width="800">

<sub>截图中的数字为演示数据。</sub>

在命令行里,📈 用对齐的小表格显示同样的两个视图(每个窗口一行),终端较窄时会先省略次要的列。

如果在会话中途安装或重新加载插件(/reload-plugins),状态栏会从对话记录里读出本次会话已用的 token 作为起点;t/s 会在下一次回复后出现。

状态栏会自动跟随应用的浅色/深色主题;📈 面板里的 ◐ 按钮可在自动、浅色、深色之间切换。

安装

在 Claude Code(命令行或桌面端)中运行:

/plugin marketplace add EricJamie/claude-code-mods
/plugin install context-band@claude-code-mods

下一个会话开始生效。

命令

/context-band auto|light|dark 切换主题 · hide / show 隐藏或显示 · reset 重置 token 计数

估算方法

  • 窗口从重置时间往前推 5 小时或 7 天。
  • 100% 估值 = 窗口内 Claude Code 按 API 价格的花费 ÷ 已用比例。已用不足 5% 时标注"(rough)",表示还不准确。
  • 只用某个模型还能用多少 token = 窗口剩余金额 ÷ 该模型在你的使用习惯下每个 token 的价格(把你最近 7 天的请求全部按该模型重新计价)。
  • 花费来自本机的 Claude Code 对话记录。在其他地方的使用(claude.ai、其他电脑)会占用额度,但不会计入金额,所以如果你经常用这些,估值会偏低。估算假设额度按 API 价格对不同模型加权。

依赖与隐私

  • 需要 python3(估算脚本 plugins/context-band/bin/api_estimate.py)
  • 自动主题和桌面端用量历史仅支持 macOS,其他系统会自动跳过
  • 所有数据都留在本机:插件只读取本地文件,不联网。缓存位于 ~/.cache/context-band/。

价格与新模型

价格表写在 bin/api_estimate.py 的 PRICES 里,来源是 Anthropic 官方价格页面(https://platform.claude.com/docs/en/about-claude/pricing)。

价格表里没有的新模型也能正常显示:它会有自己的一行,先按同系列当前模型的价格估算(全新系列按 Opus 估算),在价格学到之前数字前会标 ≈。插件会用 Claude Code 自己统计的每个会话花费(始终按当前价格计算)对照该会话里各模型用掉的 token,自动学出新模型的真实价格;已有模型如果调价,也会被同样校正。新模型发布时更新价格表仍然是最快的办法。

开发

claude plugin validate plugins/context-band
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude plugin test plugins/context-band

插件(mod)是 Claude Code 的早期功能,运行测试需要设置这个环境变量。

cache-timer 缓存倒计时

在输入框下方、模型名称前面显示 cache 41:41,即离对话缓存过期还有多久。剩最后六分之一时(1 小时缓存即最后 10 分钟)变橙色,最后 2 分钟变红色,过期后显示 expired。

/plugin install cache-timer@claude-code-mods

Claude Code 会把已有的对话内容缓存起来,之后每条消息直接读缓存,不用重新发送全部内容。每条缓存有固定寿命,从最近一次读写它的请求开始时算起,每发一条消息都会重新计时。主对话的缓存,有的套餐是 1 小时,有的是 5 分钟(子代理用 5 分钟);插件会从本次会话的对话记录里读出是哪一种,所以需要 python3。

读缓存很便宜,但不免费:Opus 5.5 读缓存每百万 token $0.20,正常输入是 $4;写缓存是输入价的 1.25 倍(5 分钟缓存)或 2 倍(1 小时缓存)。缓存过期后,下一条消息要把整段上下文重新写入缓存:一段 18 万 token 的 Opus 5.5 对话,这部分费用会从约 $0.04 变成约 $1.44。

倒计时以主对话自己的请求为准。插件看不到的一些后台请求也可能刷新缓存,所以实际过期时间可能比显示的稍晚一些。

Source 2 files
hooks/register.tsx 120 lines
1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { CacheTimerCache } from '../types'
5
6// Claude Code caches the conversation so far, so each message reads it back instead of sending it
7// again. A cache entry lives a fixed time (1 hour or 5 minutes) from the start of the last request
8// that read or wrote it, and every request restarts that time; once it lapses, the next message
9// writes the whole context to the cache again, at the write price. This counts down to that lapse
10// in the prompt footer, before the model name.
11
12const EMPTY: CacheTimerCache = { startedAt: 0, ttlMs: null }
13const cache = atom({ plugin: 'cache-timer', key: 'cache' } as const, EMPTY)
14// The second the countdown last moved: written each second while it runs, so the footer redraws.
15const tick = atom({ plugin: 'cache-timer', key: 'tick' } as const, 0)
16
17// Mid tones, so they read on light and dark footers alike.
18const COLORS = { label: '#8A8A8A', warm: '#3E9E6E', warn: '#D08A1E', danger: '#D64545', expired: '#8A8A8A' }
19
20// How often the lifetime is read again once known: it follows the plan, so it rarely changes.
21const RECHECK_MS = 10 * 60_000
22const lookup = { at: 0, isRunning: false }
23
24// Reads the lifetime from the session's transcript (bin/cache_ttl.py): the newest cache write says.
25async function readTtl($: EngineInterface) {
26  const t = await $.clock.now()
27  const known = (await read($, cache)).ttlMs
28  if (lookup.isRunning || (known !== null && t - lookup.at < RECHECK_MS)) return
29  lookup.isRunning = true
30  lookup.at = t
31  try {
32    const id = await $.session.id()
33    const { exitCode, stdout } = await $.process.run(['python3', `${$.plugin.root}/bin/cache_ttl.py`, id], { timeoutMs: 20_000 })
34    const ttlMs = exitCode === 0 ? Number(stdout.trim()) : NaN
35    if (ttlMs > 0) await update($, cache, c => (c.ttlMs === ttlMs ? c : { ...c, ttlMs }))
36  } catch {
37  } finally {
38    lookup.isRunning = false
39  }
40}
41
42// Seconds while the countdown runs; once the entry has expired (and the footer has said so) the
43// ticking stops until the next request starts another.
44async function tickTimer($: EngineInterface) {
45  const c = await read($, cache)
46  if (!c.startedAt || !c.ttlMs) return
47  const t = await $.clock.now()
48  if (t > c.startedAt + c.ttlMs + 2000) return
49  const second = t - (t % 1000)
50  if ((await read($, tick)) !== second) await update($, tick, () => second)
51}
52
53// "59:41", "07:05" for 1-hour entries (two-digit minutes, so the footer keeps its width); "4:59" for 5 minutes.
54const fmtClock = (ms: number, ttlMs: number) => {
55  const seconds = Math.max(0, Math.ceil(ms / 1000))
56  const minutes = String(Math.floor(seconds / 60))
57  return `${ttlMs >= 600_000 ? minutes.padStart(2, '0') : minutes}:${String(seconds % 60).padStart(2, '0')}`
58}
59
60// Green while warm, orange in the last sixth of the entry's life (10 minutes of an hour), red in
61// the last thirtieth (2 minutes), grey once expired.
62const colorOf = (leftMs: number, ttlMs: number) =>
63  leftMs <= 0 ? COLORS.expired : leftMs <= ttlMs / 30 ? COLORS.danger : leftMs <= ttlMs / 6 ? COLORS.warn : COLORS.warm
64
65export const register: Register = on => {
66  on('session.start', async ($, e, next) => {
67    const started = await next(e)
68    $.clock.every(1000, () => void tickTimer($))
69    // A resumed session already has cache writes to read the lifetime from.
70    void readTtl($)
71    return started
72  })
73
74  // The main conversation's requests: when each starts, and whether it read or wrote the cache.
75  on('turn.step', async function* ($, e, next) {
76    const startedAt = e.agentId ? 0 : await $.clock.now()
77    const stream = next(e)
78    for await (const chunk of stream) yield chunk
79    const result = await stream.result
80    try {
81      const u = result.usage
82      if (startedAt > 0 && u && (u.cache_read_input_tokens > 0 || u.cache_creation_input_tokens > 0)) {
83        await update($, cache, c => (startedAt > c.startedAt ? { ...c, startedAt } : c))
84      }
85    } catch {}
86    return result
87  })
88
89  // By the end of a turn its cache writes are in the transcript.
90  on('turn.complete', async ($, e, next) => {
91    const done = await next(e)
92    if (!e.agentId) void readTtl($)
93    return done
94  })
95
96  // After /clear the new conversation has no cache entry of its own yet.
97  on('session.end', async ($, e, next) => {
98    if (e.reason === 'clear') await update($, cache, c => ({ ...c, startedAt: 0 }))
99    return next(e)
100  })
101
102  // The footer's mode labels, then the countdown.
103  on('ui.render', { component: 'SessionMode' }, async ($, e, next) => {
104    await read($, tick)
105    const c = await read($, cache)
106    if (!c.startedAt || !c.ttlMs) return next(e)
107    const leftMs = c.startedAt + c.ttlMs - (await $.clock.now())
108    const { Box, Text } = $.ui.resolve(e)
109    return (
110      <Box flexDirection="row">
111        {e.props.modes.length > 0 ? <Text dimColor>{`${e.props.modes.join(' & ')} · `}</Text> : null}
112        <Text color={COLORS.label}>cache </Text>
113        <Text color={colorOf(leftMs, c.ttlMs)} bold={leftMs > 0}>
114          {leftMs > 0 ? fmtClock(leftMs, c.ttlMs) : 'expired'}
115        </Text>
116      </Box>
117    )
118  })
119}
120
types/index.d.ts 17 lines
1// The main conversation's prompt cache: when the last request that read or wrote it started (an
2// entry lives from the start of the request that last touched it; 0 before one) and how long an
3// entry lives (1 hour or 5 minutes, null until the transcript has shown a cache write).
4export type CacheTimerCache = {
5  startedAt: number
6  ttlMs: number | null
7}
8
9declare module 'claude-code' {
10  interface PluginState {
11    'cache-timer': {
12      cache: CacheTimerCache
13      tick: number
14    }
15  }
16}
17