Squish your context with Haiku when the prompt cache goes cold, so your main model re-caches a short summary instead of the whole transcript. One line above…

Squish cold Claude Code sessions with Haiku before your main model re-caches them.
🧊 cache cold · squish saves ~$1.70 of $2.00 [ squish ]
🌀 haiku squishing 204k…
✨ saved ~$1.68 · 204k → 9k
Claude Code keeps your conversation in the prompt cache for a while (5 minutes or 1 hour). Step away longer and the cache goes cold: your next message makes the main model re-read and re-cache the whole transcript at full price. On a 200k-token Opus session that is a real cost, paid again every time you come back.
haiku-compact summarizes the conversation with Haiku instead. Haiku reads the transcript at a fraction of the price, and your main model then re-caches a short summary. Your session model never changes.
/plugin install haiku-compact --marketplace kjhq/haiku-compact
Answer y to add the marketplace, then pick a scope.
While the cache is warm you see nothing. Once it goes cold, one line shows above the prompt:
| 🧊 cold | what re-caching costs and what a squish would save, with a squish button |
| 🌀 squishing | Haiku is reading the transcript |
| ✨ saved | what the squish saved you, net of the summary; gone after 30 seconds |
Prices come from Claude Code itself: its estimate on resume, or the rate implied by what your last turn cost. Until it knows a price, the line shows tokens (204k to re-cache). Small contexts (under 40k tokens) get no line and are never touched.
/compact haiku is put in the prompt box. Press Enter to squish; your held prompt is sent right after. Or send your prompt again to skip it and use the full context./compact haiku
/compact haiku keep the deployment plan and the open bugs
Anything after haiku tells the summarizer what to keep or stress. Plain /compact and Claude Code's own auto-compaction are left alone.
If Haiku fails, the compaction is skipped and your transcript is left as it was. It never falls back to compacting on the expensive model.
Set in /config or under pluginConfigs in settings.json:
| Option | Default | |
|---|---|---|
summaryModel | haiku | Model that summarizes, and the word after /compact that selects it. |
keepMessages | 4 | Recent messages kept verbatim. |
cacheTtlMinutes | 60 | Your cache TTL (60 or 5). The line shows once it has passed. |
minTokens | 40000 | Smaller contexts get no line and no compaction. |
autoCompactOnResume | true | Squish at once when resuming a cold session. |
holdPromptWhenCold | true | Hold the first prompt after the cache went cold and offer /compact haiku. |
/compact haiku, or turn the automatic parts off and squish only when you choose.~).claude plugin validate .
claude plugin test .
The Claude Code function-hooks API is early access and may change between releases.
MIT © kjhq
hooks/register.tsx 372 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, ModelUsage, Register, SessionCompactResult, SessionMessage, Timer } from 'claude-code'
3
4import type { CacheState } from '../types'
5
6
7// One Haiku request reads at most this much transcript (~100k tokens); a longer
8// one is summarized in order, each request carrying the summary so far.
9const CHUNK_CHARS = 400_000
10const TOOL_INPUT_CHARS = 2_000
11const TOOL_TEXT_CHARS = 4_000
12const SUMMARY_MAX_TOKENS = 8_000
13const CALL_TIMEOUT_MS = 5 * 60_000
14
15const SYSTEM = `You compress the transcript of a coding-agent session so the agent can continue the work from your summary alone, without the original transcript.
16
17Write a dense, factual summary with these sections:
181. User's requests and intent: every explicit request, in order, and what the user wants overall.
192. User's corrections and preferences: anything the user told the agent to do differently, verbatim where short.
203. Key technical context: languages, frameworks, services, hosts, commands that matter.
214. Files and code: every file path read or changed, what changed and why; include short code snippets that later work depends on.
225. Errors and fixes: what broke, the cause, and how it was fixed.
236. Decisions made and open questions.
247. Current state: what was being worked on in the latest messages, exactly where it stopped.
258. Next step: the immediate next action, only if the transcript makes it clear.
26
27Keep exact file paths, identifiers, commands, URLs, versions and numbers. Do not invent anything that is not in the transcript. No preamble.`
28
29const HEADER =
30 'The earlier conversation was compacted into this summary (written by a smaller model). The most recent messages follow it word for word.\n\n'
31
32const ZERO: ModelUsage = {
33 input_tokens: 0,
34 output_tokens: 0,
35 cache_read_input_tokens: 0,
36 cache_creation_input_tokens: 0,
37}
38
39const clip = (text: string, max: number) =>
40 text.length <= max ? text : `${text.slice(0, max)}… [${text.length - max} chars cut]`
41
42const isPrompt = (m: SessionMessage) => m.role === 'user' && !m.toolResults?.length && m.text.trim() !== ''
43
44/** The transcript as plain text; tool results are read from the assistant's tool calls. */
45export function renderTranscript(messages: readonly SessionMessage[]): string {
46 const parts: string[] = []
47 for (const m of messages) {
48 if (m.role === 'user') {
49 if (m.text.trim()) parts.push(`## User\n${m.text}`)
50 continue
51 }
52 const lines = m.text.trim() ? [m.text] : []
53 for (const use of m.toolUses) {
54 lines.push(`→ ${use.tool} ${clip(JSON.stringify(use.input), TOOL_INPUT_CHARS)}`)
55 if (use.text !== undefined) lines.push(` ⇒ ${use.isError ? 'ERROR ' : ''}${clip(use.text, TOOL_TEXT_CHARS)}`)
56 }
57 if (lines.length) parts.push(`## Assistant\n${lines.join('\n')}`)
58 }
59 return parts.join('\n\n')
60}
61
62/** Pieces of at most `max` chars, cut at a message boundary where one is near. */
63export function chunk(text: string, max: number): string[] {
64 const out: string[] = []
65 let rest = text
66 while (rest.length > max) {
67 const cut = rest.lastIndexOf('\n\n## ', max)
68 const at = cut > max / 2 ? cut : max
69 out.push(rest.slice(0, at))
70 rest = rest.slice(at)
71 }
72 if (rest) out.push(rest)
73 return out
74}
75
76/**
77 * Where the kept tail starts: the earliest typed prompt among the last `keep`
78 * messages, so no tool call is separated from its result. `messages.length`
79 * when there is none: everything is summarized.
80 */
81export function keepFrom(messages: readonly SessionMessage[], keep: number): number {
82 for (let i = Math.max(0, messages.length - keep); i < messages.length; i++) {
83 if (isPrompt(messages[i]!)) return i
84 }
85 return messages.length
86}
87
88const add = (a: ModelUsage, b: ModelUsage): ModelUsage => ({
89 input_tokens: a.input_tokens + b.input_tokens,
90 output_tokens: a.output_tokens + b.output_tokens,
91 cache_read_input_tokens: a.cache_read_input_tokens + b.cache_read_input_tokens,
92 cache_creation_input_tokens: a.cache_creation_input_tokens + b.cache_creation_input_tokens,
93})
94
95
96// `/compact haiku [notes]` (the first word being the model option) is
97// summarized by this mod; any other compaction is core's.
98const WORD = /^\s*(\S+)\s*/
99const MIN = 60_000
100// How long the "saved" line stays.
101const DONE_MS = 30_000
102
103const cache = atom({ plugin: 'haiku-compact', key: 'cache' } as const, {
104 phase: 'none',
105 tokens: 0,
106 warmUntil: 0,
107} as CacheState)
108
109// Set while this mod's own compaction runs, so its session.compact hook answers it.
110let isOurs = false
111// The person was told the cache is cold: their next prompt goes through as typed.
112let isWarned = false
113// The prompt held back for a `/compact <model>` the mod offered, sent after it.
114let held: string | undefined
115let coldTimer: Timer | undefined
116let doneTimer: Timer | undefined
117
118/** The compaction this mod installs: the older messages summarized by `model`, the tail kept. */
119export async function summarize(
120 $: EngineInterface,
121 model: string,
122 messages: readonly SessionMessage[],
123 keep: number,
124 notes?: string,
125): Promise<SessionCompactResult> {
126 const from = keepFrom(messages, keep)
127 const older = messages.slice(0, from)
128 if (older.length === 0) return { skip: 'nothing older than the kept messages to summarize' }
129
130 const system = notes ? `${SYSTEM}\n\nThe user asks the summary to keep or stress: ${notes}` : SYSTEM
131 let summary = ''
132 let usage = ZERO
133 for (const piece of chunk(renderTranscript(older), CHUNK_CHARS)) {
134 const prompt = summary
135 ? `Summary of the transcript so far:\n\n${summary}\n\nThe transcript continues below. Rewrite the summary so it covers everything.\n\n${piece}`
136 : `Transcript:\n\n${piece}`
137 const r = await $.model.complete({ model, system, prompt, maxTokens: SUMMARY_MAX_TOKENS, timeoutMs: CALL_TIMEOUT_MS })
138 usage = add(usage, r.usage)
139 // Never fall through to core here: that would summarize on the expensive model.
140 if (!r.isAnswered) return { skip: `${model} could not summarize (${r.reason})` }
141 summary = r.text
142 }
143
144 return {
145 messages: [{ role: 'user', text: HEADER + summary, toolUses: [] }, ...messages.slice(from)],
146 usage,
147 }
148}
149
150/** A rough token count of a transcript: about four characters a token. */
151export const estimateTokens = (messages: readonly SessionMessage[]) =>
152 Math.round(messages.reduce((n, m) => n + m.text.length + JSON.stringify(m.toolUses).length, 0) / 4)
153
154export const k = (tokens: number) => `${Math.max(1, Math.round(tokens / 1000))}k`
155
156export const usd = (dollars: number) => `$${dollars < 0.1 ? dollars.toFixed(3) : dollars.toFixed(2)}`
157
158/**
159 * The input-token price a turn implies: what it cost over its tokens, each
160 * weighed by the API's ratio to plain input (cache write `writeWeight`, cache
161 * read 0.1, output 5). Undefined when the turn says too little to tell.
162 */
163export function inputRate(costUsd: number, u: ModelUsage, writeWeight: number): number | undefined {
164 const units =
165 u.input_tokens + writeWeight * u.cache_creation_input_tokens + 0.1 * u.cache_read_input_tokens + 5 * u.output_tokens
166 return costUsd > 0 && units >= 1_000 ? costUsd / units : undefined
167}
168
169/** What re-caching `tokens` costs, and what a squish would save of it, at `rate` dollars an input token. */
170export function savings(tokens: number, rate: number, writeWeight: number, share: number) {
171 const full = tokens * writeWeight * rate
172 return { full, saved: full * (1 - share) }
173}
174
175/** The session's cost so far, or undefined where the engine does not say. */
176const costNow = ($: EngineInterface) => $.session.usage().then(u => u.cost?.usd, () => undefined)
177
178/** Compacts now through this mod; throws where the engine refuses a compaction (mid-turn). */
179async function compactCold($: EngineInterface): Promise<void> {
180 isOurs = true
181 try {
182 await $.session.compact()
183 } finally {
184 isOurs = false
185 }
186}
187
188/** Holds `text` back and offers `/compact <model>` in the prompt box; it is resent once that ran. */
189async function offerCompact($: EngineInterface, model: string, text: string | undefined): Promise<string> {
190 held = text
191 isWarned = true
192 const command = `/compact ${model}`
193 const filled = await $.prompt.fill({ text: command }).then(
194 r => r.isFilled,
195 () => false,
196 )
197 const run = filled ? 'Press Enter' : `Run ${command}`
198 const after = text === undefined ? '' : ' (your prompt is sent right after)'
199 return `Cache cold. ${run} to squish it with ${model} first${after}, or send the prompt again to use the full context.`
200}
201
202/** The band's button: compact now, or where the engine refuses that here, offer the command. */
203async function pressCompact($: EngineInterface, model: string): Promise<void> {
204 try {
205 await compactCold($)
206 } catch {
207 $.ui.toast(await offerCompact($, model, undefined))
208 }
209}
210
211export const register: Register = (on, options) => {
212 const model = String(options.summaryModel ?? 'haiku')
213 const keep = Math.max(0, Number(options.keepMessages ?? 4))
214 const ttlMs = Math.max(1, Number(options.cacheTtlMinutes ?? 60)) * MIN
215 const minTokens = Number(options.minTokens ?? 40_000)
216 const isAutoResume = options.autoCompactOnResume !== false
217 const isHolding = options.holdPromptWhenCold !== false
218
219 // Dollars a plain input token costs on the session model, learned from the engine.
220 let rate: number | undefined
221 // The share of the context a squish keeps; the last one sets it.
222 let share = 0.15
223 // The session cost when the current turn began.
224 let costAtStart: number | undefined
225
226 // A cache write costs this many plain input tokens: 2x for the 1h cache, 1.25x for 5m.
227 const writeWeight = ttlMs >= 60 * MIN ? 2 : 1.25
228
229 /** Cold, with the price of re-caching `tokens` when it is known. */
230 const goCold = (c: CacheState, tokens: number, fullUsd?: number): CacheState => {
231 const full = fullUsd ?? (rate === undefined ? undefined : savings(tokens, rate, writeWeight, share).full)
232 return { ...c, phase: 'cold', tokens, warmUntil: 0, costUsd: full, note: undefined }
233 }
234
235 on('classic.SessionStart', async ($, e, next) => {
236 const done = await next(e)
237 const isResume = e.source === 'resume' || e.source === 'fork'
238 const tokens = e.context_tokens ?? 0
239 if (!isResume || !e.prompt_cache_likely_expired || tokens < minTokens) return done
240
241 const est = e.estimated_cache_write_usd ?? 0
242 if (est > 0) rate = est / (tokens * writeWeight)
243 await update($, cache, c => goCold(c, tokens, est || undefined))
244 if (!isAutoResume) return done
245 try {
246 await compactCold($)
247 } catch {
248 // Refused at this start: the band stays cold, and the first prompt is held for it.
249 }
250 return done
251 })
252
253 on('turn.start', async ($, e, next) => {
254 coldTimer?.cancel()
255 coldTimer = undefined
256 costAtStart = await costNow($)
257 return next(e)
258 })
259
260 on('turn.complete', async ($, e, next) => {
261 const done = await next(e)
262 if (e.agentId !== undefined) return done
263 isWarned = false
264 const usage = await $.session.usage()
265 if (e.usage && costAtStart !== undefined && usage.cost) {
266 rate = inputRate(usage.cost.usd - costAtStart, e.usage, writeWeight) ?? rate
267 }
268 const tokens = usage.context.tokens ?? 0
269 if (tokens < minTokens) {
270 await update($, cache, () => ({ phase: 'none' as const, tokens, warmUntil: 0 }))
271 return done
272 }
273 // Warm: nothing to show. The band appears once the cache goes cold.
274 const at = await $.clock.now()
275 // A "saved" line still showing keeps its few seconds.
276 await update($, cache, c => ({ ...c, phase: c.phase === 'done' ? c.phase : ('warm' as const), tokens, warmUntil: at + ttlMs, note: undefined }))
277 coldTimer?.cancel()
278 coldTimer = $.clock.after(ttlMs, () => {
279 void update($, cache, c => (c.phase === 'warm' || c.phase === 'done' ? goCold(c, c.tokens) : c))
280 })
281 return done
282 })
283
284 on('prompt.submit', async ($, e, next) => {
285 // Leave alone: prompts typed mid-turn, slash commands, prompts with images
286 // (a resend would lose them), other plugins' prompts (this one's resend
287 // included), and the prompt after a warning, which goes through as typed.
288 const isPlain =
289 e.turnId === undefined && !e.text.startsWith('/') && !e.attachments?.length && e.origin.kind !== 'plugin'
290 if (!e.text.startsWith('/')) held = undefined
291 if (!isPlain || isWarned || !isHolding) return next(e)
292
293 const state = await read($, cache)
294 if (state.phase !== 'cold') return next(e)
295 return { drop: await offerCompact($, model, e.text) }
296 })
297
298 on('session.compact', async ($, e, next) => {
299 const word = e.instructions?.match(WORD)?.[1]
300 const isAsked = e.trigger === 'manual' && word?.toLowerCase() === model.toLowerCase()
301 if ((!isOurs && !isAsked) || e.agentId !== undefined) return next(e)
302 if (!Array.isArray(e.messages)) return { skip: 'no transcript to summarize' }
303
304 const state = await read($, cache)
305 const before = state.tokens || estimateTokens(e.messages)
306 const fullUsd = state.costUsd ?? (rate === undefined ? undefined : savings(before, rate, writeWeight, share).full)
307 await update($, cache, c => ({ ...c, phase: 'compacting' as const, tokens: before, note: undefined }))
308 const notes = isAsked ? e.instructions!.replace(WORD, '').trim() || undefined : undefined
309 const costBefore = await costNow($)
310 const done = await summarize($, model, e.messages, keep, notes)
311
312 if (done.skip === undefined) {
313 const after = Math.min(before, estimateTokens(done.messages))
314 share = after / before
315 // The summary's own price, where the session ledger counted it.
316 const costAfter = await costNow($)
317 const spent = costBefore === undefined || costAfter === undefined ? 0 : Math.max(0, costAfter - costBefore)
318 const savedUsd = fullUsd === undefined ? undefined : fullUsd * (1 - share) - spent
319 await update($, cache, () => ({ phase: 'done' as const, tokens: after, warmUntil: 0, before, after, savedUsd }))
320 doneTimer?.cancel()
321 doneTimer = $.clock.after(DONE_MS, () => {
322 void update($, cache, c => (c.phase === 'done' ? { ...c, phase: c.warmUntil ? ('warm' as const) : ('none' as const) } : c))
323 })
324 } else {
325 await update($, cache, c => ({ ...c, phase: 'cold' as const, note: done.skip }))
326 }
327 // The held prompt goes either way: on a skip it meets the full context.
328 if (held !== undefined) {
329 const text = held
330 held = undefined
331 // Queued: it runs once the compaction is in and the session idle.
332 await $.prompt.submit({ text, asUser: true }).catch(() => $.prompt.fill({ text }))
333 }
334 return done
335 }).catch(($, e, next) =>
336 // A failed summary must not fall through to core's expensive-model compaction.
337 next.called ? next(e) : { skip: `${model} summary failed` },
338 )
339
340 // One line, only when there is something to act on or to celebrate.
341 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
342 if (e.props.hasSurvey) return next(e)
343 const state = await read($, cache)
344 const { Box, Button, Text } = $.ui.resolve(e)
345
346 if (state.phase === 'cold') {
347 const price =
348 state.costUsd === undefined
349 ? `${k(state.tokens)} to re-cache`
350 : `squish saves ~${usd(state.costUsd * (1 - share))} of ${usd(state.costUsd)}`
351 return (
352 <Box gap={1}>
353 <Text color="suggestion">{`🧊 cache cold · ${price}${state.note ? ` · ${state.note}` : ''}`}</Text>
354 <Button key="squish" label="squish" variant="primary" onPress={() => pressCompact($, model)} />
355 </Box>
356 )
357 }
358
359 if (state.phase === 'compacting') {
360 return <Text color="claude">{`🌀 ${model} squishing ${k(state.tokens)}…`}</Text>
361 }
362
363 if (state.phase === 'done') {
364 const what = `${k(state.before ?? 0)} → ${k(state.after ?? 0)}`
365 const saved = state.savedUsd === undefined ? what : `saved ~${usd(Math.max(0, state.savedUsd))} · ${what}`
366 return <Text color="success">{`✨ ${saved}`}</Text>
367 }
368
369 return next(e)
370 })
371}
372types/index.d.ts 26 lines1/** Where the prompt cache stands, as the band above the prompt draws it. */
2export type CachePhase = 'none' | 'warm' | 'cold' | 'compacting' | 'done'
3
4export type CacheState = {
5 phase: CachePhase
6 /** Context tokens the next request re-sends. */
7 tokens: number
8 /** When the cache is expected to go cold, in `$.clock.now()` milliseconds. */
9 warmUntil: number
10 /** Estimated cost of re-caching `tokens`, when the engine said. */
11 costUsd?: number
12 /** Context before and after the last compaction, in tokens. */
13 before?: number
14 after?: number
15 /** Dollars the last squish saved, net of the summary, when prices were known. */
16 savedUsd?: number
17 /** One line about the last thing that went wrong, if anything. */
18 note?: string
19}
20
21declare module 'claude-code' {
22 interface PluginState {
23 'haiku-compact': { cache: CacheState }
24 }
25}
26