SLOPSHOPPER

gemini-compact

Moves compaction from Claude to Gemini: in summary mode Gemini summarizes the conversation and the newest messages stay verbatim, in prune mode every message…

newcommandtoastnetwork
A shopper browsing a rack in a slop shop
README

gemini-compact

When the context fills up, Claude compacts it with one more Claude request: it reads the whole context and writes a summary, and that request counts against your Claude usage. This mod hands that job to Gemini. It has two modes:

  • summary (the default): Gemini summarizes the conversation before the newest 6 messages, and those messages stay verbatim after the summary. Claude writes no summary; the engine's built-in summary runs only when Gemini fails.
  • prune: Gemini answers, for every older tool call, whether the call and its output stay, stay with a cut output, or go. Every user and assistant message stays verbatim. Nothing is summarized.

The prune idea follows fast-jev-compaction by Tamara Tran (tamaratran/fast-jev-compaction), which asks TypeSafe Jev. The code here is new and asks Gemini.

Which mode

Both modes replace the Claude request of the built-in compaction with a Gemini request.

After the compaction, every Claude request reads what is left. In summary mode that is the summary and the newest messages, close to what the built-in summary leaves. In prune mode it is every message of the conversation less the dropped tool output, which is larger, so each following request reads more. The sizes were not measured on a long session.

Pick summary mode to save the most Claude usage. Pick prune mode when the exact wording of every message matters more than the size.

Summary mode

  1. At /compact, at the engine's own compaction, and after a main-loop turn that ends with an answer and the context over the threshold, the session.compact hook takes the conversation.
  2. The newest 6 messages stay. The cut moves back to an assistant message, so a tool result is never kept without its call and the kept part opens with an assistant message after the summary. With no assistant message to cut at, everything is summarized.
  3. One generateContent request sends everything before the cut to Gemini: every message, and each call with its input and its full output. When the text is over summaryMaxInputChars (2,000,000 characters by default), the longest outputs are cut to one common length, each keeping its head and tail; a conversation over the limit even without any output fails.
  4. Gemini writes a plain-text summary in nine sections: the request and intent, technical concepts, files and code, errors and fixes, problem solving, every user message verbatim, pending tasks, the current work, and the next step. The text you write after /compact goes to Gemini with it.
  5. The conversation becomes one user message (a note, then the summary) followed by the kept messages, which go back as the engine's own messages.
  6. The built-in summary runs, and one line says why, when there is no key, nothing lies before the newest messages, Gemini fails, the summary is under 200 characters or cut at the output limit (32,768 tokens), or the result is not smaller than the conversation.

In both modes gemini-core builds the request with the key, the model and the thinking level it holds for gemini-compact, and reads the answer. After an HTTP 503 ("high demand") the mod asks again after 1 s, 2 s and 3 s, at most four attempts in all, and no attempt starts whose wait would end past 60 s. After a 429 or a key error, gemini-core hands over the request with its next key, when it holds one.

In a live check on 2.1.277, /compact took 2.6 seconds with gemini-3.5-flash-lite, Claude sent no compaction request, and after it the model named a word and a file that appeared only in the summarized part.

Prune mode

  1. The same three triggers reach the session.compact hook.
  2. A tool call in the first message or in the newest 6 messages, or whose result lies in one of them, is kept whole. Every other call gets an id (c1, c2, ...). With no such call, the built-in summary runs.
  3. One generateContent request sends the conversation to Gemini: every message, and each call with its input and its full output. Over maxInputChars (400,000 characters by default) the longest outputs are cut the same way as in summary mode. A response schema allows exactly one answer per id: keep, truncate or drop.
  4. The mod checks the answer (every id exactly once, no other id) and rebuilds the conversation:
  5. keep: the call and its output stay.
  6. truncate: the call stays, and the output keeps its first 300 characters and one line that says it was cut. An output at most 120 characters longer than that stays whole.
  7. drop: the call and its output go. A note on the nearest assistant message names the removed calls, for example [gemini-compact removed 1 earlier tool call(s) and their output after a compaction; they ran: Bash(ls -la /usr/bin)]. Without the note, the model read a reply whose work was gone and said it had never done that work (measured on 2.1.277).
  8. A message the answer does not touch goes back as the engine's own message.
  9. When the result is less than 25% smaller, or anything fails (no key, an HTTP error, an answer that breaks the schema), the engine's built-in summary runs and one line says why.

In a live check on 2.1.277, /compact took 1.1 seconds with gemini-3.5-flash-lite. Gemini dropped an ls listing and kept the cat output of a file the user was about to edit, and the conversation became 93% smaller.

What it shows

One line in the transcript, not sent to the model, and a toast that stays for 15 seconds:

gemini-compact: summary: 58 → 7 messages · 91% smaller · 312k in, 5k out gemini-compact: kept 9/11 messages · 93% smaller · 1 dropped, 0 truncated · 5k in, 59 out gemini-compact: built-in summary: under 25% smaller (kept 14/16 messages · 3% smaller · ...) gemini-compact: built-in summary: Gemini HTTP 429: Resource has been exhausted

On the free tier the toast adds · sent to Gemini free tier. An automatic compaction the engine skips or that fails writes one line too (automatic compaction skipped: ..., automatic compaction failed: ...).

Command

/gemini-compact on or off, mode, the model and thinking level gemini-core holds, threshold, tier, whether a key is set, the last result /gemini-compact on | off off leaves every compaction to the built-in summary; on is refused while gemini-core has no key /gemini-compact mode summary | mode prune /gemini-compact at <1-99> compact after a turn that ends with the context over this percentage /gemini-compact at off no automatic compaction; /compact and the engine's own compaction still ask Gemini /gemini-compact reset back to the plugin options, and off

The mod is off after an install: every compaction is the built-in one, none starts on the mod's threshold, and nothing goes to Gemini until /gemini-compact on. The command settings are kept across sessions and take effect at once. After a compaction it started, the mod starts no other one until a turn ends with the context under the threshold, so a context that stays over it does not compact after every turn.

The key, the tier, the model (default gemini-3.5-flash-lite) and the thinking level belong to gemini-core:

/gemini-core model compact gemini-3.5-flash /gemini-core thinking compact low /gemini-core paid

Free tier or paid tier

The conversation holds your prompts, the commands the model ran and the contents of the files it read. On the free tier Google may use them and human reviewers may read them; the gemini-core README quotes the Gemini API Additional Terms. On a project you would not show to Google, use a key with billing enabled and set /gemini-core paid. No mod can tell which tier a key is on; the tier setting only chooses the warning.

Google AI Studio shows the free tier limits per model; the documentation does not. They were not measured. A summary of a long conversation is one large request, so a per-minute token limit can refuse it with HTTP 429; the built-in summary then runs.

Install

claude plugin marketplace add KilimcininKorOglu/claude-code-mods claude plugin install gemini-compact@kilimcininkoroglu-mods

It depends on gemini-core, which claude plugin install adds. Function hooks are early access. Claude Code 2.1.288 and later load them by default, so there is nothing to switch on.

To load it from a local checkout for one session, put gemini-core beside it:

claude --plugin-dir plugins/gemini-core --plugin-dir plugins/gemini-compact

After installing

  1. Set the Gemini key and the tier in gemini-core, as its After installing section says, then restart Claude Code.
  2. Run /gemini-compact on. Without a key it answers still off: gemini-core has no Gemini key and stays off.
  3. Run /gemini-compact. The first line reads on · summary · <model> · thinking ... · automatic at 60% · <tier> tier · key set.
  4. Run /compact once. The transcript line should start with gemini-compact: summary:. A line that starts with built-in summary: names why Gemini was not used.

After an update from 0.2.x: claude plugin update does not add gemini-core (measured on 2.1.278), so run claude plugin install gemini-core@kilimcininkoroglu-mods once. Version 0.3.0 moved the key, tier and model to gemini-core; the apiKey, tier and model options and the settings /gemini-compact free|paid|model stored before are no longer read, so set them again in gemini-core. The mode and at settings stay. Version 0.4.0 made the mod off by default: after an update from an earlier version it is off unless you ran /gemini-compact on before, so run /gemini-compact on once.

Options

OptionDefaultWhat it sets
modesummarysummary or prune; /gemini-compact mode overrides it
compactAtPercent60The automatic threshold, 0 to 99; 0 turns it off; /gemini-compact at overrides it
keepRecent6Newest messages kept verbatim (summary) or whose calls are never sent for a decision (prune); 0 to 1000
minReduction0.25Prune mode: below this fraction (0 to 1) the built-in summary runs
headChars300Prune mode: characters kept of a truncated output; 0 to 100,000
maxInputChars400000Prune mode: characters sent to Gemini at most; 10,000 to 4,000,000
summaryMaxInputChars2000000Summary mode: characters sent to Gemini at most; 10,000 to 4,000,000

A value outside its range falls back to the default.

What it can reach

Validated with claude plugin validate on Claude Code 2.1.283:

❯ ./register.ts hooks: session.start, command.run{command=gemini-compact}, session.compact, turn.complete ❯ ./register.ts calls: $.clock.now (via askGemini), $.clock.sleep (via askGemini), $.command.register, $.gemini.enroll, $.gemini.read (via askGemini), $.gemini.request (via askGemini), $.gemini.settings (via compactWithGemini, runCommand, storePatch), $.http.fetch (via askGemini), $.session.compact (via maybeCompact), $.session.usage (via maybeCompact), $.store.delete (via runCommand), $.store.get (via loadConfig), $.store.set (via storePatch), $.ui.log, $.ui.toast (via report)

Reach L3, reaches the network.

  1. Reads: the conversation at each compaction (messages, tool inputs and outputs); the context fill after each main-loop turn; its own $.store; from gemini-core, the request with the key
  2. Runs: no process; one $.session.compact after a turn that ends over the threshold, at most once until the context was under it again
  3. Sends: the conversation (summary: all but the newest messages; prune: all of it), one request per compaction (up to four after a 503, and once more per extra key after a 429 or a key error), to the URL gemini-core builds (generativelanguage.googleapis.com) with the key in the x-goog-api-key header, never in the URL
  4. Persists: in $.store, the three command settings (enabled, mode, atPercent); the last result lives in memory
  5. Hostile input: the Gemini answer is untrusted: a prune answer is applied only in the schema shape with every candidate id once; a summary becomes the text of one user message the model reads, so a hostile summary can steer the model, as text in a file it reads can; anything malformed falls back to the built-in summary

Limits

  • A summary and a drop are a model's judgment. A summary loses detail the newest messages do not repeat. The prune note tells the model which calls ran, so it can run a tool again.
  • After the compaction the context is written to the cache again. In prune mode it stays larger than a built-in summary, so the next message pays a larger cache write.
  • A summary of a long conversation takes Gemini longer, and the compaction waits for it. Only short conversations were timed.
  • A subagent's own compaction is left to the engine.
  • A compaction the engine precomputes (precompute) also asks Gemini. Whether the engine reuses that result for the compaction that follows was not measured.
  • The test engine of claude plugin test passes no trigger to a $.session.compact() call. The plugin trigger and the hook that answers it were measured in a live session.

Development

make install # eslint, typescript-eslint, typescript make lint # complexity limit 10, the build fails above it make typecheck # needs .claude/types/ from /plugin-types make validate make test # claude plugin test

Source 7 files
hooks/register.ts 202 lines
1import type { EngineInterface, Register, SessionCompactInput, SessionMessage } from 'claude-code'
2import { actionsByUse, applyActions, sizeOf } from './apply.ts'
3import { changeText, NO_KEY_ON, parseCommand, RESET_TEXT, STORE_KEYS, statusText, storedValue, type Patch, type Settings } from './command.ts'
4import { CONSUMER, configFrom, DEADLINE_MS, DEFAULT_MODEL, outcomeText, summaryOutcomeText, type Config } from './config.ts'
5import { buildPruneBody, buildSummaryBody, type Answer } from './gemini.ts'
6import { collectCalls, parseDecisions, renderTranscript } from './prune.ts'
7import { parseSummary, summaryMessage, tailStart } from './summary.ts'
8
9/**
10 * What the hooks share: whether a compaction this mod started runs, whether
11 * the context was under the threshold since the last one it started (so a
12 * context that stays over it does not compact after every turn), and the last
13 * outcome.
14 */
15type State = { compacting: boolean; armed: boolean; last?: string }
16
17type Tier = 'free' | 'paid'
18
19type Outcome = { messages: SessionMessage[]; ratio: number; text: string; tier: Tier }
20
21function message(err: unknown): string {
22  return err instanceof Error ? err.message : String(err)
23}
24
25/** The plugin options with the settings /gemini-compact stored on top. */
26async function loadConfig($: EngineInterface, base: Config): Promise<Config> {
27  const config: Config = { ...base }
28  for (const key of Object.keys(STORE_KEYS) as (keyof Settings)[]) {
29    const value = storedValue(key, await $.store.get(STORE_KEYS[key]))
30    if (value !== undefined) Object.assign(config, { [key]: value })
31  }
32  return config
33}
34
35/**
36 * gemini-core builds the request and reads each answer; the request is sent
37 * here, again after a 503 while it allows, and with the next key after a 429
38 * or a key error. The answer carries the tier it went to.
39 */
40async function askGemini($: EngineInterface, body: Record<string, unknown>): Promise<Answer & { tier: Tier }> {
41  const prepared = await $.gemini.request({ consumer: CONSUMER, body })
42  if ('error' in prepared) throw new Error(prepared.error)
43  const started = await $.clock.now()
44  let http = prepared.http
45  for (let attempt = 1; ; attempt++) {
46    const r = await $.http.fetch(http.url, http.init)
47    const read = await $.gemini.read({ http, status: r.status, ok: r.ok, text: r.text, attempt, elapsedMs: (await $.clock.now()) - started, deadlineMs: DEADLINE_MS })
48    if ('answer' in read) return { ...read.answer, tier: prepared.tier }
49    if ('error' in read) throw new Error(read.error)
50    if ('next' in read) http = read.next
51    else await $.clock.sleep(read.retryInMs)
52  }
53}
54
55/** The fraction of characters the compaction removed. */
56function reduction(before: readonly SessionMessage[], after: readonly SessionMessage[]): number {
57  const size = sizeOf(before)
58  return size === 0 ? 0 : (size - sizeOf(after)) / size
59}
60
61/** Asks Gemini about every call outside the pinned messages and applies its answer. */
62async function prune($: EngineInterface, config: Config, e: SessionCompactInput): Promise<Outcome> {
63  const calls = collectCalls(e.messages, config.keepRecent)
64  const ids = calls.filter(c => !c.pinned).map(c => c.id)
65  if (ids.length === 0) throw new Error('no tool call outside the newest messages')
66  const transcript = renderTranscript(e.messages, calls, config.maxInputChars)
67  const answer = await askGemini($, buildPruneBody(transcript, ids, e.instructions))
68  const actions = actionsByUse(calls, parseDecisions(answer.text, ids))
69  const messages = applyActions(e.messages, actions, config.headChars)
70  const ratio = reduction(e.messages, messages)
71  const tally = { kept: messages.length, total: e.messages.length, ratio, actions: actions.values(), ...answer }
72  return { messages, ratio, text: outcomeText(tally), tier: answer.tier }
73}
74
75/** Has Gemini summarize everything before the newest messages, which stay as they are. */
76async function summarize($: EngineInterface, config: Config, e: SessionCompactInput): Promise<Outcome> {
77  const start = tailStart(e.messages, config.keepRecent)
78  const head = e.messages.slice(0, start)
79  if (head.length === 0) throw new Error('nothing to summarize before the newest messages')
80  const transcript = renderTranscript(head, collectCalls(head, 0), config.summaryMaxInputChars, true)
81  const answer = await askGemini($, buildSummaryBody(transcript, config.summaryMaxOutputTokens, e.instructions))
82  const messages = [summaryMessage(parseSummary(answer.text, answer.finishReason)), ...e.messages.slice(start)]
83  const ratio = reduction(e.messages, messages)
84  return { messages, ratio, text: summaryOutcomeText({ kept: messages.length, total: e.messages.length, ratio, ...answer }), tier: answer.tier }
85}
86
87/** One line in the transcript (not sent to the model) and a toast; free tier adds its warning to the toast. */
88function report($: EngineInterface, state: State, tier: Tier, text: string): void {
89  state.last = text
90  $.ui.log(text)
91  $.ui.toast(tier === 'free' ? `${text} · sent to Gemini free tier` : text, { timeoutMs: 15_000 })
92}
93
94/**
95 * Why a result goes to the built-in summary, or undefined when it is taken.
96 * A summary is taken whenever it is smaller, so the built-in summary runs only
97 * when Gemini fails; a prune must reach `minReduction`.
98 */
99function refusal(config: Config, ratio: number): string | undefined {
100  if (config.mode === 'summary') return ratio > 0 ? undefined : 'the summary is not smaller'
101  return ratio >= config.minReduction ? undefined : `under ${Math.round(config.minReduction * 100)}% smaller`
102}
103
104async function compactWithGemini($: EngineInterface, state: State, config: Config, e: SessionCompactInput): Promise<SessionMessage[] | undefined> {
105  const settings = await $.gemini.settings({ consumer: CONSUMER })
106  if (!settings.hasKey) {
107    report($, state, settings.tier, 'built-in summary: no Gemini key (set GEMINI_API_KEY or the gemini-core apiKey option)')
108    return undefined
109  }
110  try {
111    const outcome = config.mode === 'summary' ? await summarize($, config, e) : await prune($, config, e)
112    const refused = refusal(config, outcome.ratio)
113    if (refused === undefined) {
114      report($, state, outcome.tier, outcome.text)
115      return outcome.messages
116    }
117    report($, state, outcome.tier, `built-in summary: ${refused} (${outcome.text})`)
118  } catch (err) {
119    report($, state, settings.tier, `built-in summary: ${message(err)}`)
120  }
121  return undefined
122}
123
124/** Stores a change; `on` is refused while gemini-core has no key. */
125async function storePatch($: EngineInterface, patch: Patch): Promise<string> {
126  if (patch.enabled === true && !(await $.gemini.settings({ consumer: CONSUMER })).hasKey) return NO_KEY_ON
127  for (const [key, value] of Object.entries(patch)) await $.store.set(STORE_KEYS[key as keyof Settings], value)
128  return changeText(patch)
129}
130
131async function runCommand($: EngineInterface, state: State, base: Config, args: string): Promise<string> {
132  const command = parseCommand(args)
133  if (command.kind === 'error') return command.text
134  if (command.kind === 'reset') {
135    for (const key of Object.values(STORE_KEYS)) await $.store.delete(key)
136    return RESET_TEXT
137  }
138  if (command.kind === 'set') return storePatch($, command.patch)
139  return statusText(await loadConfig($, base), await $.gemini.settings({ consumer: CONSUMER }), state.last)
140}
141
142/** Starts a compaction when the context passed the threshold; not awaited, so the next prompt is not held. */
143async function maybeCompact($: EngineInterface, state: State, base: Config): Promise<void> {
144  const config = await loadConfig($, base)
145  if (!config.enabled || config.atPercent === 0) return
146  const { context } = await $.session.usage()
147  if ((context.percent ?? 0) < config.atPercent) {
148    state.armed = true
149    return
150  }
151  if (!state.armed) return
152  state.armed = false
153  state.compacting = true
154  // A compaction started from a timer skips this plugin's own hook (measured
155  // on 2.1.277), so it starts here, inside the dispatch, without an await.
156  void $.session
157    .compact()
158    .then(r => {
159      if (r.skip !== undefined) $.ui.log(`automatic compaction skipped: ${r.skip}`)
160    })
161    .catch((err: unknown) => $.ui.log(`automatic compaction failed: ${message(err)}`))
162    .finally(() => {
163      state.compacting = false
164    })
165}
166
167export const register: Register = (on, options) => {
168  const base = configFrom(options)
169  const state: State = { compacting: false, armed: true }
170
171  on('session.start', async ($, e, next) => {
172    const r = await next(e)
173    await $.gemini.enroll({ consumer: CONSUMER, defaultModel: DEFAULT_MODEL })
174    await $.command.register({
175      name: 'gemini-compact',
176      description: 'Gemini compaction: status, on, off, mode summary|prune, at <1-99>, at off, reset; /gemini-core sets the model, thinking and tier (gemini-compact)',
177      argumentHint: '[on | off | mode summary|prune | at <N> | at off | reset]',
178    })
179    return r
180  })
181
182  on('command.run', { command: 'gemini-compact' }, async ($, e) => ({ text: await runCommand($, state, base, String(e.args ?? '')) }))
183
184  on('session.compact', async ($, e, next) => {
185    const config = await loadConfig($, base)
186    if (!config.enabled || e.agentId !== undefined) return next(e)
187    const messages = await compactWithGemini($, state, config, e)
188    return messages === undefined ? next(e) : { messages }
189  })
190
191  on('turn.complete', async ($, e, next) => {
192    const r = await next(e)
193    if (e.agentId !== undefined || e.reason !== 'answer' || state.compacting) return r
194    try {
195      await maybeCompact($, state, base)
196    } catch (err) {
197      $.ui.log(`automatic compaction not started: ${message(err)}`)
198    }
199    return r
200  })
201}
202
hooks/apply.ts 125 lines
1/**
2 * Rebuilds the conversation from Gemini's decisions. A message the decisions
3 * do not touch is returned as the same object, so it keeps the engine's handle.
4 */
5import type { SessionMessage, ToolResultSummary, ToolUseSummary } from 'claude-code'
6import type { Action, Call } from './prune.ts'
7
8/** The action for each tool_use_id; a call without one is kept. */
9export function actionsByUse(calls: readonly Call[], decisions: ReadonlyMap<string, Action>): Map<string, Action> {
10  const actions = new Map<string, Action>()
11  for (const call of calls) {
12    const action = decisions.get(call.id)
13    if (!call.pinned && action !== undefined && action !== 'keep') actions.set(call.toolUseId, action)
14  }
15  return actions
16}
17
18/** The head of a result Gemini let go, and one line that says so. */
19export function truncated(text: string, headChars: number): string {
20  if (text.length <= headChars + 120) return text
21  return `${text.slice(0, headChars)}\n[gemini-compact cut ${text.length - headChars} chars of this output after a compaction; run the tool again if you need it]`
22}
23
24/** `Bash(ls -la)`: the tool and the first string of its input, for the note that names a removed call. */
25export function callLabel(use: ToolUseSummary): string {
26  const first = Object.values(use.input).find((v): v is string => typeof v === 'string') ?? ''
27  const arg = first.replace(/\s+/g, ' ').trim()
28  return `${use.tool}(${arg.length > 60 ? `${arg.slice(0, 57)}...` : arg})`
29}
30
31/**
32 * The line that replaces removed calls. Without it the model reads a reply
33 * whose work is gone and takes the work as never done (measured on 2.1.277).
34 */
35export function removedNote(labels: readonly string[]): string {
36  return `[gemini-compact removed ${labels.length} earlier tool call(s) and their output after a compaction; they ran: ${labels.join(', ')}]`
37}
38
39function withText(m: SessionMessage, text: string): SessionMessage {
40  const built: SessionMessage = { role: m.role, text, toolUses: m.toolUses }
41  if (m.toolResults !== undefined && m.toolResults.length > 0) built.toolResults = m.toolResults
42  return built
43}
44
45function joinText(a: string, b: string): string {
46  return a.trim() === '' ? b : `${a}\n\n${b}`
47}
48
49function editUses(m: SessionMessage, actions: ReadonlyMap<string, Action>, headChars: number): { uses: ToolUseSummary[]; removed: string[] } {
50  const uses: ToolUseSummary[] = []
51  const removed: string[] = []
52  for (const use of m.toolUses) {
53    const action = actions.get(use.tool_use_id)
54    if (action === 'drop') removed.push(callLabel(use))
55    else if (action === 'truncate' && use.text !== undefined) uses.push({ ...use, text: truncated(use.text, headChars) })
56    else uses.push(use)
57  }
58  return { uses, removed }
59}
60
61function editResults(m: SessionMessage, actions: ReadonlyMap<string, Action>, headChars: number): ToolResultSummary[] {
62  const results: ToolResultSummary[] = []
63  for (const r of m.toolResults ?? []) {
64    const action = actions.get(r.tool_use_id)
65    if (action === 'truncate') results.push({ ...r, text: truncated(r.text, headChars) })
66    else if (action !== 'drop') results.push(r)
67  }
68  return results
69}
70
71function touches(m: SessionMessage, actions: ReadonlyMap<string, Action>): boolean {
72  return m.toolUses.some(u => actions.has(u.tool_use_id)) || (m.toolResults ?? []).some(r => actions.has(r.tool_use_id))
73}
74
75type Edited = { message: SessionMessage | undefined; removed: string[] }
76
77/** One message after the decisions: undefined when nothing of it is left. */
78function editMessage(m: SessionMessage, actions: ReadonlyMap<string, Action>, headChars: number): Edited {
79  if (!touches(m, actions)) return { message: m, removed: [] }
80  const { uses, removed } = editUses(m, actions, headChars)
81  const results = editResults(m, actions, headChars)
82  if (m.text.trim() === '' && uses.length === 0 && results.length === 0) return { message: undefined, removed }
83  const built: SessionMessage = { role: m.role, text: m.text, toolUses: uses }
84  if (results.length > 0) built.toolResults = results
85  return { message: built, removed }
86}
87
88/** Puts the note on the last assistant message at or before `at`, or on the first after it. */
89function placeNote(out: SessionMessage[], at: number, note: string): void {
90  let target = -1
91  for (let i = Math.min(at, out.length - 1); i >= 0 && target === -1; i--) if (out[i]?.role === 'assistant') target = i
92  if (target === -1) target = out.findIndex(m => m.role === 'assistant')
93  const m = out[target]
94  if (m === undefined) return
95  out[target] = withText(m, joinText(m.text, note))
96}
97
98/**
99 * Applies the actions: `drop` removes a call with its result, `truncate` keeps
100 * the call and the head of its result. A message left empty is removed, and the
101 * calls removed from it are named in a note on the nearest assistant message.
102 */
103export function applyActions(messages: readonly SessionMessage[], actions: ReadonlyMap<string, Action>, headChars: number): SessionMessage[] {
104  const out: SessionMessage[] = []
105  const notes: { at: number; labels: string[] }[] = []
106  for (const m of messages) {
107    const { message, removed } = editMessage(m, actions, headChars)
108    if (message !== undefined) out.push(message)
109    if (removed.length > 0) notes.push({ at: out.length - 1, labels: removed })
110  }
111  for (const { at, labels } of notes) placeNote(out, at, removedNote(labels))
112  return out
113}
114
115/** Characters of text, tool input and tool result the conversation holds. */
116export function sizeOf(messages: readonly SessionMessage[]): number {
117  let total = 0
118  for (const m of messages) {
119    total += m.text.length
120    for (const u of m.toolUses) total += (JSON.stringify(u.input) ?? '').length
121    for (const r of m.toolResults ?? []) total += r.text.length
122  }
123  return total
124}
125
hooks/command.ts 98 lines
1/** The settings /gemini-compact changes, and the reading of its argument. */
2import type { EngineInterface } from 'claude-code'
3
4/** What gemini-core says this mod runs with. */
5export type GeminiSettings = Awaited<ReturnType<EngineInterface['gemini']['settings']>>
6
7/** summary: Gemini summarizes the conversation; prune: Gemini decides on each old tool call. */
8export type Mode = 'summary' | 'prune'
9
10/** What the command can change; 0 in `atPercent` turns the automatic trigger off. */
11export type Settings = { enabled: boolean; mode: Mode; atPercent: number }
12
13export type Patch = Partial<Settings>
14
15export type Command =
16  | { kind: 'status' }
17  | { kind: 'reset' }
18  | { kind: 'set'; patch: Patch }
19  | { kind: 'error'; text: string }
20
21export const USAGE = 'expects on, off, mode summary, mode prune, at <1-99>, at off, or reset; /gemini-core sets the model, the thinking level and the tier'
22
23/** The store key of each setting. */
24export const STORE_KEYS: Record<keyof Settings, string> = {
25  enabled: 'enabled',
26  mode: 'mode',
27  atPercent: 'atPercent',
28}
29
30function parseAt(value: string | undefined): Command {
31  if (value === 'off') return { kind: 'set', patch: { atPercent: 0 } }
32  const n = Number(value)
33  if (value === undefined || !/^\d+$/.test(value) || n < 1 || n > 99) return { kind: 'error', text: 'at takes a whole percentage from 1 to 99, or off' }
34  return { kind: 'set', patch: { atPercent: n } }
35}
36
37function parseMode(value: string | undefined): Command {
38  if (value !== 'summary' && value !== 'prune') return { kind: 'error', text: 'mode takes summary or prune' }
39  return { kind: 'set', patch: { mode: value } }
40}
41
42const WORDS: Record<string, Command> = {
43  '': { kind: 'status' },
44  status: { kind: 'status' },
45  reset: { kind: 'reset' },
46  on: { kind: 'set', patch: { enabled: true } },
47  off: { kind: 'set', patch: { enabled: false } },
48}
49
50/** Reads the argument of /gemini-compact. */
51export function parseCommand(args: string): Command {
52  const words = args.trim().split(/\s+/).filter(Boolean)
53  const [first = '', second, ...rest] = words
54  if (rest.length > 0) return { kind: 'error', text: USAGE }
55  if (first === 'at') return parseAt(second)
56  if (first === 'mode') return parseMode(second)
57  const word = WORDS[first]
58  return word !== undefined && second === undefined ? word : { kind: 'error', text: USAGE }
59}
60
61/** A stored value of the right type, or undefined so the plugin option applies. */
62export function storedValue<K extends keyof Settings>(key: K, value: unknown): Settings[K] | undefined {
63  const valid: Record<keyof Settings, (v: unknown) => boolean> = {
64    enabled: v => typeof v === 'boolean',
65    mode: v => v === 'summary' || v === 'prune',
66    atPercent: v => typeof v === 'number' && Number.isInteger(v) && v >= 0 && v <= 99,
67  }
68  return valid[key](value) ? (value as Settings[K]) : undefined
69}
70
71const MODE_TEXT: Record<Mode, string> = {
72  summary: 'mode summary: Gemini summarizes the conversation and the newest messages stay verbatim',
73  prune: 'mode prune: every message stays and Gemini keeps, truncates or drops each old tool call',
74}
75
76/** What `on` answers while gemini-core has no key; nothing is stored. */
77export const NO_KEY_ON = 'still off: gemini-core has no Gemini key. Set GEMINI_API_KEY or the gemini-core apiKey option, restart Claude Code, then run /gemini-compact on'
78
79/** What `reset` answers: the plugin options, and off until it is turned on. */
80export const RESET_TEXT = 'settings reset to the plugin options; off until /gemini-compact on'
81
82/** The line /gemini-compact prints for a change. */
83export function changeText(patch: Patch): string {
84  if (patch.enabled !== undefined) return patch.enabled ? 'on' : 'off: compaction uses the built-in summary'
85  if (patch.mode !== undefined) return MODE_TEXT[patch.mode]
86  return patch.atPercent === 0 ? 'automatic compaction off; /compact and the engine\'s own compaction still use Gemini' : `compacts when the context passes ${patch.atPercent}%`
87}
88
89/** The status line of /gemini-compact, with what gemini-core says it runs with. */
90export function statusText(s: Settings, g: GeminiSettings, last: string | undefined): string {
91  const at = s.atPercent === 0 ? 'automatic off' : `automatic at ${s.atPercent}%`
92  const key = g.hasKey ? 'key set' : 'no key: set GEMINI_API_KEY or the gemini-core apiKey option'
93  const lines = [`${s.enabled ? 'on' : 'off'} · ${s.mode} · ${g.model} · thinking ${g.thinking ?? 'model default'} · ${at} · ${g.tier} tier · ${key}`]
94  if (!s.enabled) lines.push('off until /gemini-compact on; every compaction uses the built-in summary; the model, the thinking level and the tier are /gemini-core settings')
95  if (last !== undefined) lines.push(`last: ${last}`)
96  return lines.join('\n')
97}
98
hooks/config.ts 75 lines
1/** The plugin options with their defaults, and the line a compaction reports. */
2import type { PluginOptions } from 'claude-code'
3import { storedValue, type Settings } from './command.ts'
4import type { Action } from './prune.ts'
5
6export type Config = Settings & {
7  keepRecent: number
8  minReduction: number
9  headChars: number
10  maxInputChars: number
11  summaryMaxInputChars: number
12  summaryMaxOutputTokens: number
13}
14
15/** The plugin name gemini-core knows this mod by, and the model it uses until /gemini-core sets another. */
16export const CONSUMER = 'gemini-compact'
17export const DEFAULT_MODEL = 'gemini-3.5-flash-lite'
18
19/** No new Gemini attempt starts once this much has passed. */
20export const DEADLINE_MS = 60_000
21
22/** Off until the user turns it on, so a fresh install sends nothing to Gemini. */
23export const DEFAULTS: Config = {
24  enabled: false,
25  mode: 'summary',
26  atPercent: 60,
27  keepRecent: 6,
28  minReduction: 0.25,
29  headChars: 300,
30  maxInputChars: 400_000,
31  summaryMaxInputChars: 2_000_000,
32  summaryMaxOutputTokens: 32_768,
33}
34
35function numberOption(options: PluginOptions, key: string, fallback: number, min: number, max: number): number {
36  const value = options[key]
37  return typeof value === 'number' && Number.isFinite(value) && value >= min && value <= max ? value : fallback
38}
39
40/** The plugin options, each out-of-range or missing one replaced by its default. */
41export function configFrom(options: PluginOptions): Config {
42  return {
43    enabled: DEFAULTS.enabled,
44    mode: storedValue('mode', options.mode) ?? DEFAULTS.mode,
45    atPercent: numberOption(options, 'compactAtPercent', DEFAULTS.atPercent, 0, 99),
46    keepRecent: numberOption(options, 'keepRecent', DEFAULTS.keepRecent, 0, 1000),
47    minReduction: numberOption(options, 'minReduction', DEFAULTS.minReduction, 0, 1),
48    headChars: numberOption(options, 'headChars', DEFAULTS.headChars, 0, 100_000),
49    maxInputChars: numberOption(options, 'maxInputChars', DEFAULTS.maxInputChars, 10_000, 4_000_000),
50    summaryMaxInputChars: numberOption(options, 'summaryMaxInputChars', DEFAULTS.summaryMaxInputChars, 10_000, 4_000_000),
51    summaryMaxOutputTokens: DEFAULTS.summaryMaxOutputTokens,
52  }
53}
54
55function tokens(n: number): string {
56  return n >= 1000 ? `${Math.round(n / 1000)}k` : String(n)
57}
58
59export type Tally = { kept: number; total: number; ratio: number; actions: Iterable<Action>; inputTokens: number; outputTokens: number }
60
61/** `kept 41/58 messages · 52% smaller · 18 dropped, 9 truncated · 31k in, 1k out` */
62export function outcomeText(t: Tally): string {
63  const all = [...t.actions]
64  const dropped = all.filter(a => a === 'drop').length
65  const cut = all.filter(a => a === 'truncate').length
66  return `kept ${t.kept}/${t.total} messages · ${Math.round(t.ratio * 100)}% smaller · ${dropped} dropped, ${cut} truncated · ${tokens(t.inputTokens)} in, ${tokens(t.outputTokens)} out`
67}
68
69export type SummaryTally = Omit<Tally, 'actions'>
70
71/** `summary: 58 → 7 messages · 91% smaller · 312k in, 5k out` */
72export function summaryOutcomeText(t: SummaryTally): string {
73  return `summary: ${t.total} → ${t.kept} messages · ${Math.round(t.ratio * 100)}% smaller · ${tokens(t.inputTokens)} in, ${tokens(t.outputTokens)} out`
74}
75
hooks/gemini.ts 57 lines
1/** The generateContent bodies this mod sends; gemini-core builds the request around them and reads the answer. */
2import type { EngineInterface } from 'claude-code'
3import { ACTIONS } from './prune.ts'
4import { SUMMARY_TASK } from './summary.ts'
5
6/** Gemini's answer as gemini-core reads it; `finishReason` is Gemini's, such as `STOP` or `MAX_TOKENS`. */
7export type Answer = Extract<Awaited<ReturnType<EngineInterface['gemini']['read']>>, { answer: unknown }>['answer']
8
9const TASK = `You decide which tool calls of a Claude Code conversation can leave its context, because the context is being compacted. Every message stays; only tool calls and their outputs can go.
10
11For every call with an id such as [c7], answer one action:
12- keep: the assistant still needs the full output (a file it is editing, an error it is fixing, data it will quote, an output that cannot be produced again).
13- truncate: knowing the call ran and its input still matters, but the full output does not; only its first lines stay.
14- drop: neither the call nor its output matters for what comes next (a superseded read of a file changed since, a listing or search already acted on, a passing test run, a finished side task).
15
16Calls marked [fixed] stay whatever you answer; read them as context only. The latest user messages state the current task. When unsure, answer keep.`
17
18function prompt(task: string, transcript: string, instructions: string | undefined): string {
19  const focus = instructions === undefined || instructions.trim() === '' ? '' : `\n\nThe user asked the compaction to keep in mind: ${instructions.trim()}`
20  return `${task}${focus}\n\nThe conversation:\n\n${transcript}`
21}
22
23/** The body that asks for a plain-text summary of the conversation. */
24export function buildSummaryBody(transcript: string, maxOutputTokens: number, instructions?: string): Record<string, unknown> {
25  return {
26    contents: [{ role: 'user', parts: [{ text: prompt(SUMMARY_TASK, transcript, instructions) }] }],
27    generationConfig: { maxOutputTokens },
28  }
29}
30
31/** The body that asks for one action per candidate id, in a schema Gemini must follow. */
32export function buildPruneBody(transcript: string, ids: readonly string[], instructions?: string): Record<string, unknown> {
33  return {
34    contents: [{ role: 'user', parts: [{ text: prompt(TASK, transcript, instructions) }] }],
35    generationConfig: {
36      responseMimeType: 'application/json',
37      responseSchema: {
38        type: 'OBJECT',
39        properties: {
40          decisions: {
41            type: 'ARRAY',
42            items: {
43              type: 'OBJECT',
44              properties: {
45                id: { type: 'STRING', enum: [...ids] },
46                action: { type: 'STRING', enum: [...ACTIONS] },
47              },
48              required: ['id', 'action'],
49            },
50          },
51        },
52        required: ['decisions'],
53      },
54    },
55  }
56}
57
hooks/prune.ts 153 lines
1/**
2 * The pure part of gemini-compact: which tool calls Gemini decides on, the
3 * transcript it reads, its answer, and the conversation rebuilt from it.
4 */
5import type { SessionMessage, ToolUseSummary } from 'claude-code'
6
7export type Action = 'keep' | 'truncate' | 'drop'
8
9export const ACTIONS: readonly Action[] = ['keep', 'truncate', 'drop']
10
11/** One tool call with its result. Only a call that is not pinned gets a decision. */
12export type Call = {
13  /** The short id Gemini answers with (`c1`), or '' for a pinned call. */
14  id: string
15  toolUseId: string
16  tool: string
17  input: Record<string, unknown>
18  output: string
19  isError: boolean
20  pinned: boolean
21}
22
23function outputOf(messages: readonly SessionMessage[], use: ToolUseSummary): string {
24  for (const m of messages) {
25    const result = m.toolResults?.find(r => r.tool_use_id === use.tool_use_id)
26    if (result !== undefined) return result.text
27  }
28  return use.text ?? ''
29}
30
31/** The index of the message that holds each tool result, by tool_use_id. */
32function resultIndex(messages: readonly SessionMessage[]): Map<string, number> {
33  const at = new Map<string, number>()
34  messages.forEach((m, i) => {
35    for (const r of m.toolResults ?? []) at.set(r.tool_use_id, i)
36  })
37  return at
38}
39
40/**
41 * Pairs every tool call with its result. A call in the first message or in the
42 * newest `keepRecent` messages, or whose result is in one of them, is pinned.
43 */
44export function collectCalls(messages: readonly SessionMessage[], keepRecent: number): Call[] {
45  const firstRecent = messages.length - Math.max(0, keepRecent)
46  const isPinned = (i: number): boolean => i === 0 || i >= firstRecent
47  const results = resultIndex(messages)
48  const calls: Call[] = []
49  let next = 1
50  messages.forEach((m, i) => {
51    for (const use of m.toolUses) {
52      const pinned = isPinned(i) || isPinned(results.get(use.tool_use_id) ?? i)
53      const id = pinned ? '' : `c${next++}`
54      const output = outputOf(messages, use)
55      calls.push({ id, toolUseId: use.tool_use_id, tool: use.tool, input: use.input, output, isError: use.isError === true, pinned })
56    }
57  })
58  return calls
59}
60
61function inputText(input: Record<string, unknown>): string {
62  return JSON.stringify(input) ?? '{}'
63}
64
65/** Cuts a text to about `max` characters: its head and its tail around one note. */
66export function clip(text: string, max: number): string {
67  if (text.length <= max) return text
68  const half = Math.max(0, Math.floor(max / 2))
69  return `${text.slice(0, half)}\n[… ${text.length - 2 * half} chars omitted …]\n${text.slice(text.length - half)}`
70}
71
72/** The longest output length that makes the transcript fit, found by bisection. */
73function outputCap(fixed: number, outputs: readonly number[], max: number): number | undefined {
74  const size = (cap: number): number => fixed + outputs.reduce((sum, n) => sum + Math.min(n, cap + 40), 0)
75  if (size(0) > max) return undefined
76  let low = 0
77  let high = Math.max(0, ...outputs)
78  while (low < high) {
79    const mid = Math.ceil((low + high) / 2)
80    if (size(mid) <= max) low = mid
81    else high = mid - 1
82  }
83  return low
84}
85
86function callLines(call: Call, cap: number, plain: boolean): string[] {
87  const label = plain ? '[call]' : call.pinned ? '[fixed]' : `[${call.id}]`
88  const status = call.isError ? 'error' : 'output'
89  return [`  ${label} ${call.tool} ${inputText(call.input)}`, `  ${status}: ${clip(call.output, cap)}`]
90}
91
92function render(messages: readonly SessionMessage[], byUse: ReadonlyMap<string, Call>, cap: number, plain: boolean): string {
93  const lines: string[] = []
94  messages.forEach((m, i) => {
95    lines.push(`#${i + 1} ${m.role}: ${m.text}`)
96    for (const use of m.toolUses) {
97      const call = byUse.get(use.tool_use_id)
98      if (call !== undefined) lines.push(...callLines(call, cap, plain))
99    }
100  })
101  return lines.join('\n')
102}
103
104/**
105 * The conversation as Gemini reads it: every message in order, each call with
106 * its id (or `[call]` when `plain`), input and output. When it is over
107 * `maxChars`, the longest outputs are cut to one common length; a
108 * conversation over it without any output throws.
109 */
110export function renderTranscript(messages: readonly SessionMessage[], calls: readonly Call[], maxChars: number, plain = false): string {
111  const byUse = new Map(calls.map(c => [c.toolUseId, c]))
112  const full = render(messages, byUse, Infinity, plain)
113  if (full.length <= maxChars) return full
114  const outputs = calls.map(c => c.output.length)
115  const fixed = full.length - outputs.reduce((a, b) => a + b, 0)
116  const cap = outputCap(fixed, outputs, maxChars)
117  if (cap === undefined) throw new Error(`the conversation is over ${maxChars} characters even without tool output`)
118  return render(messages, byUse, cap, plain)
119}
120
121function isRecord(value: unknown): value is Record<string, unknown> {
122  return typeof value === 'object' && value !== null && !Array.isArray(value)
123}
124
125function decisionOf(value: unknown): [string, Action] | string {
126  if (!isRecord(value) || typeof value.id !== 'string') return 'a decision has no id'
127  const action = ACTIONS.find(a => a === value.action)
128  return action === undefined ? `${value.id}: unknown action ${JSON.stringify(value.action)}` : [value.id, action]
129}
130
131/** Reads Gemini's answer: exactly one known action for every candidate id, nothing else. */
132export function parseDecisions(text: string, ids: readonly string[]): Map<string, Action> {
133  let value: unknown
134  try {
135    value = JSON.parse(text)
136  } catch {
137    throw new Error('the answer is not JSON')
138  }
139  if (!isRecord(value) || !Array.isArray(value.decisions)) throw new Error('the answer has no decisions list')
140  const known = new Set(ids)
141  const decisions = new Map<string, Action>()
142  for (const raw of value.decisions) {
143    const d = decisionOf(raw)
144    if (typeof d === 'string') throw new Error(d)
145    if (!known.has(d[0])) throw new Error(`unknown call id ${d[0]}`)
146    if (decisions.has(d[0])) throw new Error(`call ${d[0]} decided twice`)
147    decisions.set(d[0], d[1])
148  }
149  const missing = ids.filter(id => !decisions.has(id))
150  if (missing.length > 0) throw new Error(`no decision for ${missing.slice(0, 5).join(', ')}${missing.length > 5 ? '…' : ''}`)
151  return decisions
152}
153
hooks/summary.ts 54 lines
1/**
2 * The pure part of the summary mode: where the verbatim tail starts, what
3 * Gemini is asked, and the message its summary becomes.
4 */
5import type { SessionMessage } from 'claude-code'
6
7/**
8 * The index where the newest `keepRecent` messages start, moved back to an
9 * assistant message: the tail then holds every tool result whose call it
10 * holds, and it follows the summary (a user message) with an assistant one.
11 * With no assistant message after the first, everything is summarized.
12 */
13export function tailStart(messages: readonly SessionMessage[], keepRecent: number): number {
14  if (keepRecent <= 0) return messages.length
15  for (let i = Math.max(1, messages.length - keepRecent); i >= 1; i--) {
16    if (messages[i]?.role === 'assistant') return i
17  }
18  return messages.length
19}
20
21export const SUMMARY_TASK = `You write the summary that replaces the older part of a Claude Code conversation, because its context is being compacted. The assistant continues the work from your summary and the newest messages, which stay verbatim after it. It cannot see anything you leave out.
22
23Write plain text with these sections:
241. Primary request and intent: everything the user asked for, in detail.
252. Key technical concepts: technologies, frameworks and facts the work relies on.
263. Files and code: every file read, created or changed, with its full path, why it matters, and the code snippets the work still needs.
274. Errors and fixes: each error met, how it was fixed, and what the user said about it.
285. Problem solving: problems solved and open investigations.
296. All user messages: every user message that is not a tool result, verbatim.
307. Pending tasks: what the user asked for that is not done.
318. Current work: what was being done right before this summary, with file names and code.
329. Next step: the step that follows from the latest request, quoting the user's words.
33
34Keep exact names, paths, commands, numbers and error texts. Do not invent anything the conversation does not say.`
35
36/** The note in front of the summary, so the assistant knows what it reads. */
37export const SUMMARY_NOTE =
38  'This session continues an earlier conversation that was compacted. Gemini wrote the summary below of the part before the newest messages, which follow it verbatim.'
39
40/** The first message of the compacted conversation; without a handle, the engine builds it from its role and text. */
41export function summaryMessage(summary: string): SessionMessage {
42  return { role: 'user', text: `${SUMMARY_NOTE}\n\n${summary}`, toolUses: [] }
43}
44
45const MIN_SUMMARY_CHARS = 200
46
47/** The summary text; one Gemini cut at its output limit, an empty one or a too short one throws. */
48export function parseSummary(text: string, finishReason?: string): string {
49  if (finishReason === 'MAX_TOKENS') throw new Error('the summary hit the output token limit')
50  const summary = text.trim()
51  if (summary.length < MIN_SUMMARY_CHARS) throw new Error(`the summary is too short (${summary.length} chars)`)
52  return summary
53}
54