Counts down to when the prompt cache expires, so you can send the next message before the whole conversation has to be cached again

Counts down to when the prompt cache expires, so you can send your next message while it is still warm.

Claude Code caches the start of the conversation. The next message reads that part from the cache, which costs a fraction of normal input. If you wait too long, the cache expires and the next message writes the whole conversation to the cache again, at a higher rate than normal input. In a long session that is a large one-time charge.
How long the cache lives depends on how you use Claude Code:
| You use | Cache lives | What an expired cache costs you |
|---|---|---|
| An API key, Bedrock, Vertex or Foundry | 5 minutes | Money: the whole context billed at the cache-write rate |
| A Claude subscription, within its limits | 1 hour | No money; the re-write likely counts toward your 5-hour and 7-day windows |
| A Claude subscription past its limits (extra usage) | 5 minutes | Money, as with an API key |
So the timer helps most on 5 minutes, where a coffee break is enough to lose the cache. On a subscription it matters after long breaks, and for big contexts that eat into your limits.
A line just above the prompt, counting from the last request of the main conversation. It sits apart from the status lines under the prompt, so its tick every second does not shuffle them:
cache 54:12: the cache is warm.cache 0:48 · send soon: a fifth of its life or less is left.cache cold · next message re-writes 182k tokens: it has expired; the number is how much the next message caches again.Subagent requests do not count; they do not keep the main conversation's cache alive. The line clears after /clear and /compact, since the next message writes a new cache either way.
FORCE_PROMPT_CACHING_5M, then CLAUDE_CODE_PROMPT_CACHE_TTL, then the promptCacheTtl setting, then ENABLE_PROMPT_CACHING_1H, then 1 hour on a subscription within its limits and 5 minutes otherwise. It remembers your limits from the last reading, so a new session starts with the right lifetime; on the very first run it assumes 5 minutes until the first reply reports them. After you switch between a subscription and an API key, the first reply can show the old lifetime for a moment./model drops the cache; the timer does not know until the next reply.claude --resume the line stays empty until the first reply.claude plugin marketplace add arasovic/claude-code-mods
claude plugin install cache-timer@claude-code-mods
Restart Claude Code. The timer appears above the prompt after the first reply.
claude plugin validate .
claude plugin test .
../typecheck.sh cache-timerhooks/register.tsx 110 lines1import type { EngineInterface, Register, SessionRateLimit } from 'claude-code'
2
3type Ttl = '5m' | '1h'
4const TTL_MS: Record<Ttl, number> = { '5m': 300_000, '1h': 3_600_000 }
5const isTtl = (v: unknown): v is Ttl => v === '5m' || v === '1h'
6const isOn = (v: string | undefined) => ['1', 'true', 'yes', 'on'].includes((v ?? '').toLowerCase())
7
8type TtlInputs = { force5m: boolean; envTtl?: string; settingTtl?: unknown; enable1h: boolean; subscribed: boolean; overLimit: boolean }
9
10// The engine's own order as of 2.1.288; the mod API does not expose the TTL it picked.
11export function cacheTtl(o: TtlInputs): Ttl {
12 if (o.force5m) return '5m'
13 if (isTtl(o.envTtl)) return o.envTtl
14 if (isTtl(o.settingTtl)) return o.settingTtl
15 if (o.enable1h) return '1h'
16 return o.subscribed && !o.overLimit ? '1h' : '5m'
17}
18
19// ponytail: "over a limit" stands in for the engine's overage flag, which no mod can read.
20export function limitState(limits: readonly SessionRateLimit[]) {
21 const windows = limits.filter(l => l.kind === 'five_hour' || l.kind === 'seven_day')
22 return { subscribed: windows.length > 0, overLimit: windows.some(l => l.percentUsed >= 100) }
23}
24
25const tokens = (n: number) => (n >= 1000 ? `${Math.round(n / 1000)}k` : `${n}`)
26
27export function statusText(now: number, requestAt: number | undefined, ttl: Ttl, context: number): string | undefined {
28 if (requestAt === undefined) return undefined
29 const left = requestAt + TTL_MS[ttl] - now
30 if (left <= 0) return `🔴 cache cold · next message re-writes ${tokens(context)} tokens`
31 const secs = Math.ceil(left / 1000)
32 const clock = `${Math.floor(secs / 60)}:${String(secs % 60).padStart(2, '0')}`
33 return left <= TTL_MS[ttl] / 5 ? `🟡 cache ${clock} · send soon` : `🟢 cache ${clock}`
34}
35
36const inputs: TtlInputs = { force5m: false, enable1h: false, subscribed: false, overLimit: false }
37let requestAt: number | undefined
38let context = 0
39let shown: string | undefined
40const LIMITS = 'limits'
41
42async function tick($: EngineInterface) {
43 const text = statusText(await $.clock.now(), requestAt, cacheTtl(inputs), context)
44 if (text !== shown) {
45 shown = text
46 $.ui.invalidate('ui.render')
47 }
48}
49
50async function start($: EngineInterface) {
51 inputs.force5m = isOn(await $.env.get('FORCE_PROMPT_CACHING_5M'))
52 inputs.envTtl = await $.env.get('CLAUDE_CODE_PROMPT_CACHE_TTL')
53 inputs.enable1h = isOn(await $.env.get('ENABLE_PROMPT_CACHING_1H'))
54 inputs.settingTtl = (await $.settings.read()).promptCacheTtl
55 // The last session's limits, so the first reply after a start or reload shows the right lifetime.
56 Object.assign(inputs, await $.store.get(LIMITS))
57 // A timer started inside a turn.step dispatch dies with it; session.start's runs until the module reloads.
58 $.clock.every(1000, () => void tick($))
59}
60
61async function forget($: EngineInterface) {
62 requestAt = undefined
63 await tick($)
64}
65
66export const register: Register = on => {
67 on('turn.step', async function* ($, e, next) {
68 // A subagent's requests carry their own prefix; only the main thread's keep its cache alive.
69 if (e.agentId !== undefined) return yield* next(e)
70 // The cache entry is refreshed when the request is served, so the clock starts before the reply streams.
71 const at = await $.clock.now()
72 const result = yield* next(e)
73 const usage = result.usage
74 if (usage) {
75 requestAt = at
76 context = usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens + usage.output_tokens
77 }
78 return result
79 })
80
81 on('session.measure', async ($, e, next) => {
82 // Empty limits mean "no subscription" only once a response has reported its tokens; before that, "no reading yet".
83 if (e.rateLimits.length > 0 || e.context.tokens !== undefined) {
84 const limits = limitState(e.rateLimits)
85 Object.assign(inputs, limits)
86 await $.store.set(LIMITS, limits)
87 }
88 return next(e)
89 })
90
91 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
92 // The band is shared: other plugins (next-steps) draw beneath this hook, so keep what they drew.
93 const below = await next(e)
94 if (e.props.hasSurvey || shown === undefined) return below
95 const { Box, Text } = $.ui.resolve(e)
96 return (
97 <Box flexDirection="column">
98 {below}
99 <Text>{shown}</Text>
100 </Box>
101 )
102 })
103
104 on('session.start', async ($, e, next) => (await start($), next(e)))
105
106 // After /clear, /resume or a compaction the next request writes a new prefix whatever the clock says.
107 on('session.end', async ($, e, next) => (await forget($), next(e)))
108 on('session.compact', async ($, e, next) => (await forget($), next(e)))
109}
110