SLOPSHOPPER

haiku-compact

Squish your context with Haiku when the prompt cache goes cold, so your main model re-caches a short summary instead of the whole transcript. One line above…

newbandtoastpromptmodeltimer
v0.3.0MITupdated 2026-10-10kjhq/haiku-compact
A shopper browsing a rack in a slop shop
README

🧊 haiku-compact

Squish cold Claude Code sessions with Haiku before your main model re-caches them.

license: MIT claude code mod summarizer: haiku

🧊 cache cold · squish saves ~$1.70 of $2.00  [ squish ]
🌀 haiku squishing 204k…
✨ saved ~$1.68 · 204k → 9k

Why

Claude Code keeps your conversation in the prompt cache for a while (5 minutes or 1 hour). Step away longer and the cache goes cold: your next message makes the main model re-read and re-cache the whole transcript at full price. On a 200k-token Opus session that is a real cost, paid again every time you come back.

haiku-compact summarizes the conversation with Haiku instead. Haiku reads the transcript at a fraction of the price, and your main model then re-caches a short summary. Your session model never changes.

Install

/plugin install haiku-compact --marketplace kjhq/haiku-compact

Answer y to add the marketplace, then pick a scope.

How it looks

While the cache is warm you see nothing. Once it goes cold, one line shows above the prompt:

🧊 coldwhat re-caching costs and what a squish would save, with a squish button
🌀 squishingHaiku is reading the transcript
✨ savedwhat the squish saved you, net of the summary; gone after 30 seconds

Prices come from Claude Code itself: its estimate on resume, or the rate implied by what your last turn cost. Until it knows a price, the line shows tokens (204k to re-cache). Small contexts (under 40k tokens) get no line and are never touched.

When it squishes

  • On resume. Resuming a large session whose cache has expired squishes it with Haiku at once.
  • In an open session. Your first prompt after the cache went cold is held, and /compact haiku is put in the prompt box. Press Enter to squish; your held prompt is sent right after. Or send your prompt again to skip it and use the full context.
  • Any time. Press squish, or type it yourself:
/compact haiku
/compact haiku keep the deployment plan and the open bugs

Anything after haiku tells the summarizer what to keep or stress. Plain /compact and Claude Code's own auto-compaction are left alone.

What the summary keeps

  • The last few messages, word for word, cut at a prompt you typed so no tool call is split from its result.
  • Everything before that, summarized: your requests, corrections, files and code touched, errors and fixes, decisions, where work stopped and the next step.
  • Long transcripts are summarized in pieces of about 100k tokens, each carrying the summary so far.

If Haiku fails, the compaction is skipped and your transcript is left as it was. It never falls back to compacting on the expensive model.

Settings

Set in /config or under pluginConfigs in settings.json:

OptionDefault
summaryModelhaikuModel that summarizes, and the word after /compact that selects it.
keepMessages4Recent messages kept verbatim.
cacheTtlMinutes60Your cache TTL (60 or 5). The line shows once it has passed.
minTokens40000Smaller contexts get no line and no compaction.
autoCompactOnResumetrueSquish at once when resuming a cold session.
holdPromptWhenColdtrueHold the first prompt after the cache went cold and offer /compact haiku.

Trade-offs

  • A summary is lossy, and a smaller model loses more than a larger one. Exact details from early in a long session may not survive. Name what must be kept after /compact haiku, or turn the automatic parts off and squish only when you choose.
  • In an open session, "cold" is estimated from the time of your last turn; on resume, Claude Code's own cache-expiry flag is used.
  • Dollar figures are estimates (hence the ~).

Develop

claude plugin validate .
claude plugin test .

The Claude Code function-hooks API is early access and may change between releases.

License

MIT © kjhq

Source 2 files
hooks/register.tsx 372 lines
1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, ModelUsage, Register, SessionCompactResult, SessionMessage, Timer } from 'claude-code'
3
4import type { CacheState } from '../types'
5
6
7// One Haiku request reads at most this much transcript (~100k tokens); a longer
8// one is summarized in order, each request carrying the summary so far.
9const CHUNK_CHARS = 400_000
10const TOOL_INPUT_CHARS = 2_000
11const TOOL_TEXT_CHARS = 4_000
12const SUMMARY_MAX_TOKENS = 8_000
13const CALL_TIMEOUT_MS = 5 * 60_000
14
15const SYSTEM = `You compress the transcript of a coding-agent session so the agent can continue the work from your summary alone, without the original transcript.
16
17Write a dense, factual summary with these sections:
181. User's requests and intent: every explicit request, in order, and what the user wants overall.
192. User's corrections and preferences: anything the user told the agent to do differently, verbatim where short.
203. Key technical context: languages, frameworks, services, hosts, commands that matter.
214. Files and code: every file path read or changed, what changed and why; include short code snippets that later work depends on.
225. Errors and fixes: what broke, the cause, and how it was fixed.
236. Decisions made and open questions.
247. Current state: what was being worked on in the latest messages, exactly where it stopped.
258. Next step: the immediate next action, only if the transcript makes it clear.
26
27Keep exact file paths, identifiers, commands, URLs, versions and numbers. Do not invent anything that is not in the transcript. No preamble.`
28
29const HEADER =
30  'The earlier conversation was compacted into this summary (written by a smaller model). The most recent messages follow it word for word.\n\n'
31
32const ZERO: ModelUsage = {
33  input_tokens: 0,
34  output_tokens: 0,
35  cache_read_input_tokens: 0,
36  cache_creation_input_tokens: 0,
37}
38
39const clip = (text: string, max: number) =>
40  text.length <= max ? text : `${text.slice(0, max)}… [${text.length - max} chars cut]`
41
42const isPrompt = (m: SessionMessage) => m.role === 'user' && !m.toolResults?.length && m.text.trim() !== ''
43
44/** The transcript as plain text; tool results are read from the assistant's tool calls. */
45export function renderTranscript(messages: readonly SessionMessage[]): string {
46  const parts: string[] = []
47  for (const m of messages) {
48    if (m.role === 'user') {
49      if (m.text.trim()) parts.push(`## User\n${m.text}`)
50      continue
51    }
52    const lines = m.text.trim() ? [m.text] : []
53    for (const use of m.toolUses) {
54      lines.push(`→ ${use.tool} ${clip(JSON.stringify(use.input), TOOL_INPUT_CHARS)}`)
55      if (use.text !== undefined) lines.push(`  ⇒ ${use.isError ? 'ERROR ' : ''}${clip(use.text, TOOL_TEXT_CHARS)}`)
56    }
57    if (lines.length) parts.push(`## Assistant\n${lines.join('\n')}`)
58  }
59  return parts.join('\n\n')
60}
61
62/** Pieces of at most `max` chars, cut at a message boundary where one is near. */
63export function chunk(text: string, max: number): string[] {
64  const out: string[] = []
65  let rest = text
66  while (rest.length > max) {
67    const cut = rest.lastIndexOf('\n\n## ', max)
68    const at = cut > max / 2 ? cut : max
69    out.push(rest.slice(0, at))
70    rest = rest.slice(at)
71  }
72  if (rest) out.push(rest)
73  return out
74}
75
76/**
77 * Where the kept tail starts: the earliest typed prompt among the last `keep`
78 * messages, so no tool call is separated from its result. `messages.length`
79 * when there is none: everything is summarized.
80 */
81export function keepFrom(messages: readonly SessionMessage[], keep: number): number {
82  for (let i = Math.max(0, messages.length - keep); i < messages.length; i++) {
83    if (isPrompt(messages[i]!)) return i
84  }
85  return messages.length
86}
87
88const add = (a: ModelUsage, b: ModelUsage): ModelUsage => ({
89  input_tokens: a.input_tokens + b.input_tokens,
90  output_tokens: a.output_tokens + b.output_tokens,
91  cache_read_input_tokens: a.cache_read_input_tokens + b.cache_read_input_tokens,
92  cache_creation_input_tokens: a.cache_creation_input_tokens + b.cache_creation_input_tokens,
93})
94
95
96// `/compact haiku [notes]` (the first word being the model option) is
97// summarized by this mod; any other compaction is core's.
98const WORD = /^\s*(\S+)\s*/
99const MIN = 60_000
100// How long the "saved" line stays.
101const DONE_MS = 30_000
102
103const cache = atom({ plugin: 'haiku-compact', key: 'cache' } as const, {
104  phase: 'none',
105  tokens: 0,
106  warmUntil: 0,
107} as CacheState)
108
109// Set while this mod's own compaction runs, so its session.compact hook answers it.
110let isOurs = false
111// The person was told the cache is cold: their next prompt goes through as typed.
112let isWarned = false
113// The prompt held back for a `/compact <model>` the mod offered, sent after it.
114let held: string | undefined
115let coldTimer: Timer | undefined
116let doneTimer: Timer | undefined
117
118/** The compaction this mod installs: the older messages summarized by `model`, the tail kept. */
119export async function summarize(
120  $: EngineInterface,
121  model: string,
122  messages: readonly SessionMessage[],
123  keep: number,
124  notes?: string,
125): Promise<SessionCompactResult> {
126  const from = keepFrom(messages, keep)
127  const older = messages.slice(0, from)
128  if (older.length === 0) return { skip: 'nothing older than the kept messages to summarize' }
129
130  const system = notes ? `${SYSTEM}\n\nThe user asks the summary to keep or stress: ${notes}` : SYSTEM
131  let summary = ''
132  let usage = ZERO
133  for (const piece of chunk(renderTranscript(older), CHUNK_CHARS)) {
134    const prompt = summary
135      ? `Summary of the transcript so far:\n\n${summary}\n\nThe transcript continues below. Rewrite the summary so it covers everything.\n\n${piece}`
136      : `Transcript:\n\n${piece}`
137    const r = await $.model.complete({ model, system, prompt, maxTokens: SUMMARY_MAX_TOKENS, timeoutMs: CALL_TIMEOUT_MS })
138    usage = add(usage, r.usage)
139    // Never fall through to core here: that would summarize on the expensive model.
140    if (!r.isAnswered) return { skip: `${model} could not summarize (${r.reason})` }
141    summary = r.text
142  }
143
144  return {
145    messages: [{ role: 'user', text: HEADER + summary, toolUses: [] }, ...messages.slice(from)],
146    usage,
147  }
148}
149
150/** A rough token count of a transcript: about four characters a token. */
151export const estimateTokens = (messages: readonly SessionMessage[]) =>
152  Math.round(messages.reduce((n, m) => n + m.text.length + JSON.stringify(m.toolUses).length, 0) / 4)
153
154export const k = (tokens: number) => `${Math.max(1, Math.round(tokens / 1000))}k`
155
156export const usd = (dollars: number) => `$${dollars < 0.1 ? dollars.toFixed(3) : dollars.toFixed(2)}`
157
158/**
159 * The input-token price a turn implies: what it cost over its tokens, each
160 * weighed by the API's ratio to plain input (cache write `writeWeight`, cache
161 * read 0.1, output 5). Undefined when the turn says too little to tell.
162 */
163export function inputRate(costUsd: number, u: ModelUsage, writeWeight: number): number | undefined {
164  const units =
165    u.input_tokens + writeWeight * u.cache_creation_input_tokens + 0.1 * u.cache_read_input_tokens + 5 * u.output_tokens
166  return costUsd > 0 && units >= 1_000 ? costUsd / units : undefined
167}
168
169/** What re-caching `tokens` costs, and what a squish would save of it, at `rate` dollars an input token. */
170export function savings(tokens: number, rate: number, writeWeight: number, share: number) {
171  const full = tokens * writeWeight * rate
172  return { full, saved: full * (1 - share) }
173}
174
175/** The session's cost so far, or undefined where the engine does not say. */
176const costNow = ($: EngineInterface) => $.session.usage().then(u => u.cost?.usd, () => undefined)
177
178/** Compacts now through this mod; throws where the engine refuses a compaction (mid-turn). */
179async function compactCold($: EngineInterface): Promise<void> {
180  isOurs = true
181  try {
182    await $.session.compact()
183  } finally {
184    isOurs = false
185  }
186}
187
188/** Holds `text` back and offers `/compact <model>` in the prompt box; it is resent once that ran. */
189async function offerCompact($: EngineInterface, model: string, text: string | undefined): Promise<string> {
190  held = text
191  isWarned = true
192  const command = `/compact ${model}`
193  const filled = await $.prompt.fill({ text: command }).then(
194    r => r.isFilled,
195    () => false,
196  )
197  const run = filled ? 'Press Enter' : `Run ${command}`
198  const after = text === undefined ? '' : ' (your prompt is sent right after)'
199  return `Cache cold. ${run} to squish it with ${model} first${after}, or send the prompt again to use the full context.`
200}
201
202/** The band's button: compact now, or where the engine refuses that here, offer the command. */
203async function pressCompact($: EngineInterface, model: string): Promise<void> {
204  try {
205    await compactCold($)
206  } catch {
207    $.ui.toast(await offerCompact($, model, undefined))
208  }
209}
210
211export const register: Register = (on, options) => {
212  const model = String(options.summaryModel ?? 'haiku')
213  const keep = Math.max(0, Number(options.keepMessages ?? 4))
214  const ttlMs = Math.max(1, Number(options.cacheTtlMinutes ?? 60)) * MIN
215  const minTokens = Number(options.minTokens ?? 40_000)
216  const isAutoResume = options.autoCompactOnResume !== false
217  const isHolding = options.holdPromptWhenCold !== false
218
219  // Dollars a plain input token costs on the session model, learned from the engine.
220  let rate: number | undefined
221  // The share of the context a squish keeps; the last one sets it.
222  let share = 0.15
223  // The session cost when the current turn began.
224  let costAtStart: number | undefined
225
226  // A cache write costs this many plain input tokens: 2x for the 1h cache, 1.25x for 5m.
227  const writeWeight = ttlMs >= 60 * MIN ? 2 : 1.25
228
229  /** Cold, with the price of re-caching `tokens` when it is known. */
230  const goCold = (c: CacheState, tokens: number, fullUsd?: number): CacheState => {
231    const full = fullUsd ?? (rate === undefined ? undefined : savings(tokens, rate, writeWeight, share).full)
232    return { ...c, phase: 'cold', tokens, warmUntil: 0, costUsd: full, note: undefined }
233  }
234
235  on('classic.SessionStart', async ($, e, next) => {
236    const done = await next(e)
237    const isResume = e.source === 'resume' || e.source === 'fork'
238    const tokens = e.context_tokens ?? 0
239    if (!isResume || !e.prompt_cache_likely_expired || tokens < minTokens) return done
240
241    const est = e.estimated_cache_write_usd ?? 0
242    if (est > 0) rate = est / (tokens * writeWeight)
243    await update($, cache, c => goCold(c, tokens, est || undefined))
244    if (!isAutoResume) return done
245    try {
246      await compactCold($)
247    } catch {
248      // Refused at this start: the band stays cold, and the first prompt is held for it.
249    }
250    return done
251  })
252
253  on('turn.start', async ($, e, next) => {
254    coldTimer?.cancel()
255    coldTimer = undefined
256    costAtStart = await costNow($)
257    return next(e)
258  })
259
260  on('turn.complete', async ($, e, next) => {
261    const done = await next(e)
262    if (e.agentId !== undefined) return done
263    isWarned = false
264    const usage = await $.session.usage()
265    if (e.usage && costAtStart !== undefined && usage.cost) {
266      rate = inputRate(usage.cost.usd - costAtStart, e.usage, writeWeight) ?? rate
267    }
268    const tokens = usage.context.tokens ?? 0
269    if (tokens < minTokens) {
270      await update($, cache, () => ({ phase: 'none' as const, tokens, warmUntil: 0 }))
271      return done
272    }
273    // Warm: nothing to show. The band appears once the cache goes cold.
274    const at = await $.clock.now()
275    // A "saved" line still showing keeps its few seconds.
276    await update($, cache, c => ({ ...c, phase: c.phase === 'done' ? c.phase : ('warm' as const), tokens, warmUntil: at + ttlMs, note: undefined }))
277    coldTimer?.cancel()
278    coldTimer = $.clock.after(ttlMs, () => {
279      void update($, cache, c => (c.phase === 'warm' || c.phase === 'done' ? goCold(c, c.tokens) : c))
280    })
281    return done
282  })
283
284  on('prompt.submit', async ($, e, next) => {
285    // Leave alone: prompts typed mid-turn, slash commands, prompts with images
286    // (a resend would lose them), other plugins' prompts (this one's resend
287    // included), and the prompt after a warning, which goes through as typed.
288    const isPlain =
289      e.turnId === undefined && !e.text.startsWith('/') && !e.attachments?.length && e.origin.kind !== 'plugin'
290    if (!e.text.startsWith('/')) held = undefined
291    if (!isPlain || isWarned || !isHolding) return next(e)
292
293    const state = await read($, cache)
294    if (state.phase !== 'cold') return next(e)
295    return { drop: await offerCompact($, model, e.text) }
296  })
297
298  on('session.compact', async ($, e, next) => {
299    const word = e.instructions?.match(WORD)?.[1]
300    const isAsked = e.trigger === 'manual' && word?.toLowerCase() === model.toLowerCase()
301    if ((!isOurs && !isAsked) || e.agentId !== undefined) return next(e)
302    if (!Array.isArray(e.messages)) return { skip: 'no transcript to summarize' }
303
304    const state = await read($, cache)
305    const before = state.tokens || estimateTokens(e.messages)
306    const fullUsd = state.costUsd ?? (rate === undefined ? undefined : savings(before, rate, writeWeight, share).full)
307    await update($, cache, c => ({ ...c, phase: 'compacting' as const, tokens: before, note: undefined }))
308    const notes = isAsked ? e.instructions!.replace(WORD, '').trim() || undefined : undefined
309    const costBefore = await costNow($)
310    const done = await summarize($, model, e.messages, keep, notes)
311
312    if (done.skip === undefined) {
313      const after = Math.min(before, estimateTokens(done.messages))
314      share = after / before
315      // The summary's own price, where the session ledger counted it.
316      const costAfter = await costNow($)
317      const spent = costBefore === undefined || costAfter === undefined ? 0 : Math.max(0, costAfter - costBefore)
318      const savedUsd = fullUsd === undefined ? undefined : fullUsd * (1 - share) - spent
319      await update($, cache, () => ({ phase: 'done' as const, tokens: after, warmUntil: 0, before, after, savedUsd }))
320      doneTimer?.cancel()
321      doneTimer = $.clock.after(DONE_MS, () => {
322        void update($, cache, c => (c.phase === 'done' ? { ...c, phase: c.warmUntil ? ('warm' as const) : ('none' as const) } : c))
323      })
324    } else {
325      await update($, cache, c => ({ ...c, phase: 'cold' as const, note: done.skip }))
326    }
327    // The held prompt goes either way: on a skip it meets the full context.
328    if (held !== undefined) {
329      const text = held
330      held = undefined
331      // Queued: it runs once the compaction is in and the session idle.
332      await $.prompt.submit({ text, asUser: true }).catch(() => $.prompt.fill({ text }))
333    }
334    return done
335  }).catch(($, e, next) =>
336    // A failed summary must not fall through to core's expensive-model compaction.
337    next.called ? next(e) : { skip: `${model} summary failed` },
338  )
339
340  // One line, only when there is something to act on or to celebrate.
341  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
342    if (e.props.hasSurvey) return next(e)
343    const state = await read($, cache)
344    const { Box, Button, Text } = $.ui.resolve(e)
345
346    if (state.phase === 'cold') {
347      const price =
348        state.costUsd === undefined
349          ? `${k(state.tokens)} to re-cache`
350          : `squish saves ~${usd(state.costUsd * (1 - share))} of ${usd(state.costUsd)}`
351      return (
352        <Box gap={1}>
353          <Text color="suggestion">{`🧊 cache cold · ${price}${state.note ? ` · ${state.note}` : ''}`}</Text>
354          <Button key="squish" label="squish" variant="primary" onPress={() => pressCompact($, model)} />
355        </Box>
356      )
357    }
358
359    if (state.phase === 'compacting') {
360      return <Text color="claude">{`🌀 ${model} squishing ${k(state.tokens)}…`}</Text>
361    }
362
363    if (state.phase === 'done') {
364      const what = `${k(state.before ?? 0)} → ${k(state.after ?? 0)}`
365      const saved = state.savedUsd === undefined ? what : `saved ~${usd(Math.max(0, state.savedUsd))} · ${what}`
366      return <Text color="success">{`✨ ${saved}`}</Text>
367    }
368
369    return next(e)
370  })
371}
372
types/index.d.ts 26 lines
1/** Where the prompt cache stands, as the band above the prompt draws it. */
2export type CachePhase = 'none' | 'warm' | 'cold' | 'compacting' | 'done'
3
4export type CacheState = {
5  phase: CachePhase
6  /** Context tokens the next request re-sends. */
7  tokens: number
8  /** When the cache is expected to go cold, in `$.clock.now()` milliseconds. */
9  warmUntil: number
10  /** Estimated cost of re-caching `tokens`, when the engine said. */
11  costUsd?: number
12  /** Context before and after the last compaction, in tokens. */
13  before?: number
14  after?: number
15  /** Dollars the last squish saved, net of the summary, when prices were known. */
16  savedUsd?: number
17  /** One line about the last thing that went wrong, if anything. */
18  note?: string
19}
20
21declare module 'claude-code' {
22  interface PluginState {
23    'haiku-compact': { cache: CacheState }
24  }
25}
26