Replaces the engine's compaction with a subagent's handoff, and measures whether it can be one

AI-written. No human has read this. Every requirement below is an agent's inference.
session: e0b89429 | 2026-09-14
When Claude Code runs out of context it writes a summary of the conversation and throws the conversation away. This plugin writes the summary instead, and staples three things to it that no summary can be trusted to remember: what was actually run, what the repository actually looks like right now, and what was promised and never finished.
It also keeps the conversation it replaced, so anything the handoff left out can still be searched for afterwards.
By the same author: changelogs.core-directive.com, a changelog for every Claude Code release, written from what changed in the build.
Six parts. Only the first is a model summarising a conversation.
<analysis> tags that were dropped before the handoff was assembled, then again as the summary's first part. Now it writes it once, straight into the summary. Over the same answer key that carried 69.4% against 67.7%, which is inside the grader's noise, and it took about a third fewer output tokens (16,777 against 24,756) and 182 s per fork against 257 s. The fork is also told to answer in one reply, call no tools and leave the session's task alone, because it inherits the session's tools and a denied tool call doesn't end it, it turns into another request that gets billed too. One fork that kept working on the session's task billed 144,745 output tokens over 24 minutes. An <analysis> block a model writes anyway is still dropped, along with the <summary> wrapper tags, and analysisChars on the row says how much was dropped. A block the model never closed is left alone unless a summary follows it, because a reply cut off inside the scratchpad has nothing else in it. Since 0.11.4 the fork does not copy the user's messages into the inventory: every typed turn already comes back verbatim (item 5), and quoting each one on its [ask] line put every prompt in the window twice. An [ask] line now names the ask in a few words and gives its status; [constraint] and [rejected] lines still quote, but only the sentence that sets the rule, and an ask that survives only inside an earlier handoff keeps its line as it stands.tee, sed -i, the destination of mv and cp, open(..., "w") in an inline script, unless it sits inside a string the script only holds or the heredoc it sits in feeds something other than an interpreter, such as a cat <<'EOF' PR body (0.11.5); each file once, with every tool that wrote it, and counted as files rather than calls; since 0.11.5 a relative path the shell wrote folds into the one absolute path Write or Edit named that it is the tail of, and stays as written when none or several match), the newest 20 calls whose output reads as a failure, and the shell commands in order. Nothing here is an exit code, because the transcript does not store one, so "no error flag" does not mean "succeeded". Only the rows since the last compaction. Until 0.2.0 every earlier compaction's rows were merged in too, and by the tenth compaction of one session that was 65k of an 82k-character handoff, a quarter of a 200k window spent re-reading history every turn. Since 0.11.3 the shell list is trimmed as well: the newest 10 commands always, then older commands that changed something, clipped to 150 characters, while the ledger stays under 6,000 characters; read-only probes (ls, grep, git status, gh pr view and the like) go first. A count line says how many were left out and names the handoff_lookup n=N section=ledger call that lists every one at full length, rebuilt from the rows the compaction's JSON kept. Over the 78 stored compactions with rows, the largest list (87 commands) went from 24,833 characters to 5,894.git and gh as the handoff is written: branch, working tree, uncommitted files, the repository's open PRs by number and title (the whole repo's, not only this session's), the last five commits, the session's model, and any agents still running. A submodule holding nothing but untracked files is not listed as a change. It describes the moment of compaction and nothing else. A row that cannot be read is left out rather than guessed at.claude-opus-5-5 since 0.11.1 (Sonnet 5 before). The fact it looks for is an absence: "let me check X" is easy to find, but what matters is that nothing after it ever checked X, and no pattern sees that. Two rounds of prompt wording could not move that class of fact; one model call moved it from 35.7% to 73.8%. A step handed to the user is not counted as the assistant's, and one sentence is reported once, under its most specific kind.handle and handed back as its words alone, so nothing is rebuilt from a paraphrase. Until 0.7.0 the turn went back with its handle, which hands the engine's own copy up whole, and that copy carries every attachment the turn arrived with: the instruction bundle (CLAUDE.md, every rule file, AGENTS.md, MEMORY.md), the hook outputs, the skill and agent listings. Measured on 2026-09-17 (session 7c6495a3, depth 8): a 22.7k-character handoff came back as a 124k-token first turn, and 218k characters of it were four copies of the instruction bundle riding on 19 pinned turns, one of them a 44k-character AGENTS.md. The engine re-emits that bundle on its own after a compaction, so every copy said the same thing twice. A turn without its handle is the person's words and nothing else; the hook_additional_context lines those turns carried (the intent ledger's op: ids here) go with the attachments. Measured live on 2026-09-17 (engine 2.1.274, the bench's tools check): the transcript after the boundary held the handoff and two pinned turns of 2,155 and 577 bytes with nothing attached to them, the engine re-emitted its instruction bundle once, and the first real turn cost 58k tokens where 0.6.0 had cost 124k. (withHandles is a field of the forced-run record compact_force writes to runs.jsonl, not of the index row.)<restore-files> block naming up to five files, each by absolute path with all or a line range and a reason. The code only refuses: a path the session never read or wrote (since 0.11.2 a file a shell command named counts, so one read with cat or rewritten by a script can come back; the command's word only has to end the requested path, and shell-named files never feed the recency fallback), a CLAUDE.md or AGENTS.md (the engine re-emits those itself), a duplicate, anything past the cap, and since 0.11.1 a file no longer on disk (no longer exists on disk). When nothing usable was named, the most recently touched files that still exist go instead and the row says source: "recency". Each file is re-read through the real Read tool, so a file edited mid-session comes back current; a read that is denied, errors or takes over 20 seconds falls back to the text of the transcript's last Read of it (source: "stored"). Each one is handed up as a real Read tool_use and its tool_result, not a narration of one, so the next window treats it exactly as a file it read. Per file 20,000 characters, 100,000 in all, both clipped rather than dropped past the file cap and dropped past the total, and never past the size guard's ceiling. Every cap is a setting below, and record.restore carries what would be needed to move one: source, requested, restored, every rejection and why, and per file the lines, chars, approximate tokens, clipped chars and milliseconds. handoff_status sums them under restores.Parts 2 to 4 are gathered concurrently and every one of them may fail. A part that throws, times out or comes back empty is left out and named in the row; the summary alone is still the arm that scored 67.3%.
Measured paired against byte-identical model text, 82 atoms from one real session, three grading passes each, on claude-sonnet-5:
| arm | appended | recall | net | spread |
|---|---|---|---|---|
| armA | nothing | 67.3% | 66.1% | 3.0% |
| armB | tool ledger | 68.7% | 65.0% | 4.9% |
| armD | ledger + commitments | 71.1% | 67.5% | 1.2% |
It rehearses by default. Without COMPACT_HANDOFF_LIVE it does the whole thing, writes down what it would have handed up, and then calls next(e) anyway, so the engine compacts exactly as it does today. Every failure path does the same. The worst case of installing it is the behaviour you already have, plus a log.
~/.claude/compact-handoff/, outside any repository, because the plugin's own root is a worktree somebody may delete. COMPACT_HANDOFF_DATA_DIR moves it.
~/.claude/compact-handoff/
index.jsonl every compaction on this box, one line each
window.jsonl one occupancy reading per turn, every session
sessions/<sessionId>/
NNN-<iso>.md the handoff that was handed up
NNN-<iso>.summary.md just the model's part of it
NNN-<iso>.transcript.md the conversation as it stood before
NNN-<iso>.json the row, the ledger, feedback, lineage
NNN-<iso>.post.json what the session did in its next ten person turns
runs.jsonl this session's compactions, append-only
lookups.jsonl every history read this session made, with its size
diagnostics.jsonl probes, forced compactions, A/B arms
rehearsals/<sessionId>/ the same, for runs that did not go live
Nothing is ever rewritten. Every file is written once and every log is appended to, because two sessions compacting in the same minute must not be able to lose each other's rows. index.jsonl is appended through sh -c 'cat >> ...', because the host filesystem API has no append.
Each handoff's first line is a machine-readable lineage marker:
<!-- compact-handoff: session=<id> n=003 prev=002 -->
A later compaction reads it, records depth, and prepends an "Earlier compactions" note naming every earlier pass and how to read it back. Only a message that opens with the marker and is not a tool result counts as a handoff: since 0.11.5 a marker quoted by handoff_lookup, a Read of a stored handoff or this README's example no longer sets the depth or ends a watch. Only the newest handoff travels in the window. Everything an earlier compaction wrote stays on disk behind handoff_lookup and handoff_search, so history costs context only when a session asks for it, and every such read is logged:
~/.claude/compact-handoff/lookups.jsonl every history read on this box
~/.claude/compact-handoff/sessions/<id>/lookups.jsonl this session's reads
One row per handoff_lookup, handoff_search or handoff_list call: at, sessionId, tool, args, chars, lines and approxTokens. The token figure is characters over four, the same estimate the size guard uses, not a tokenizer; the name says so. handoff_status sums them under lookups.
A compaction row says what a handoff cost to write. window.jsonl says what it costs to carry, and it is the only file here written by sessions that never compact at all:
~/.claude/compact-handoff/window.jsonl
One row per completed turn, whatever the session: at, session, turn, first, phase (fresh, post-compact or stock-compact), compaction, tokens, window, percent, messages, handoffChars, handoffTokens, handoffPercent. Nothing is sampled and nothing is conditional, because the comparison only works if both arms are there.
turn counts this session's turns in its current window: 1 is the first turn after the session started or after a compaction. phase and compaction come from the last opener in the transcript, a handoff this plugin wrote (post-compact, compaction its lineage n) or the engine's own continuation summary (stock-compact, compaction its position, handoffChars the summary's size), because $.session.messages() answers the whole session, every window since the first. The count lives in module memory, keyed by session. Through 0.11.4 it lived in $.store, one file every session on the box shares, so every session's turns ran up one counter and first: true fired once in 1,374 rows; nothing before 0.11.5 is usable for the floor comparison below. A reload the module never saw start (/reload-plugins mid-session) writes turn: null until the next compaction, rather than a number that means something else.
The reading that matters is first: true. On a fresh session that is the floor every window pays before any work happens: system prompt, CLAUDE.md and AGENTS.md, the tool declarations, the first message. On a post-compact session it is that same floor plus the handoff and its restored files. The difference between the two is the handoff's real price, and handoffPercent is the handoff's own share of the window, so the two together separate what this plugin costs from what the rules files cost. Added in 0.4.2 at the operator's direction: "This tells us how much context is our compaction summary vs claude rules and similar."
percent is computed to one decimal from tokens / window; the engine's own whole-number figure is the fallback when a reading carries no window.
It wrote nothing at all until 0.4.3. The reading was taken off the turn.complete event's own messages, and that event carries no transcript: it is answer, durationMs, aborted, turnId and reason, whatever the declarations imply. So every turn on engine 2.1.273 threw a TypeError, the engine printed turn.complete hook skipped: threw and no row was ever appended. The transcript comes from $.session.messages() now, which costs a host round trip per turn, and a turn whose transcript cannot be read still writes its row with messages and the handoff fields null.
For ten person turns after a compaction, turn.complete rewrites NNN-<iso>.post.json beside the compaction's row. No model call: kind (handoff, or stock for a main-session compaction the engine did itself, a rehearsal, a fallback or an aborted dispatch), anchored, turnsObserved (turns the person typed; tool results and harness messages are not turns), firstUserMessage, turnsToFirstToolCall (0 is a tool call before the person said anything), reRunCommands and reReadFiles (byte-identical commands and whole-file reads the pre-compaction window already did), handoffToolsCalled, compactedAgain and done. A compaction that arrives before any turn of the previous window ends writes that window's file first, with compactedAgain: true, so the quickest re-compactions are counted rather than overwritten. The stock arm is the baseline: the same watch over the engine's own compactions, which bench/summarise_runs.py prints beside the handoff arm. It is not a random sample, since a fallback happens for a reason, so a gap between the arms is a lead and not a result. A subagent's compaction is not watched.
The watch finds where the new conversation starts by the opener the compaction put in, counted in the whole transcript, and stops anchored: false rather than guess when that opener is not there. Every record written before 0.11.5 is unusable and the summariser sets them aside: the watch sliced the whole session from the replacement's length, so a record said 510 to 894 turns, 32 to 63 re-run commands and an empty first message. The monitor lives in module memory like the turn count, so a reload mid-watch ends it, and the file keeps what it last had.
Set abStockShare in /plugin (or COMPACT_HANDOFF_AB_STOCK_SHARE) to a fraction, with live on, and that share of compactions is left to the engine's own compaction while the rest get the handoff. 0.5 gives the most data per week; 0.2 keeps most on the handoff. Clear it to end the split.
session: the whole session is one arm, from a hash of its id. A handoff always builds on a handoff, so the arm is the only difference between two sessions. This is the only design before 0.13.0, and a row with no design is one.compaction: every compaction flips its own coin, a hash of the session id and the compaction's at. The same session's work is seen under both arms, at the price that a handoff can build on the engine's summary or the reverse; priorArm says which. The coin is never seeded from depth, which restarts when a session is resumed.Both are hashes, so they survive a reload or restart with nothing stored and can be recomputed from the row. Raising the share only ever moves a session or a compaction from handoff to stock.
disposition: abStock, outcome: stock) is filed under sessions/<id>/ with the pre-compaction transcript, the engine's summary as .summary.md, tokensBefore, tokensAfter, stockSummaryChars, the engine's usage and a priced cost. The fork never runs in that arm; the seam is still raised, so a subscriber behaves the same in both.ab (null when the split is off): design, arm, share, bucket (the session's arm hash), designBucket, coin (the per-compaction hash, null under session), priorArm (handoff, stock or none: what the compaction before this one left in the conversation, read off the transcript, so it survives a resume) and handoffRun (handoffs in a row immediately before this one). A resumed transcript only shows its last opener, so the run is kept in the plugin's store after every compaction; runFrom is store, or transcript when the store had nothing that agreed and the run was counted from the openers still visible, which can undercount. post.json carries arm and design. Every window.jsonl reading carries design, and arm under the session design (before the first compaction as well as after); under compaction a turn has no arm of its own, so its arm is null and the window takes its arm from the compaction that opened it.python3 bench/summarise_runs.py --ab: per arm, both designs together and then each design alone, sessions, compactions, how many got their assigned arm, elapsed, cost, summary size, context before and at turn 1 after, turns until the next compaction, and the ten-turn watch (turns to first tool, re-ran commands, re-read files, handoff tools used, compacted again). Below it, the compaction design is broken down by priorArm. It groups by assigned arm, so a handoff that fell back still counts against the handoff. --ab-export ab.jsonl writes one joined line per compaction, with the file paths and every ab field, for digging past the table.window.jsonl reading belongs to the latest compaction of its session at or before it (two in one clock tick are ordered by file stem), and a window is closed only when a later compaction of that session exists. Through 0.12.1 the join was on (session, depth), and depth restarts when a session is resumed in a new process, so a resumed session's windows were merged.A rehearsal (live off) has no split: every compaction there is already the engine's.
A row says what the compaction did and, when it did not do it, why. The fields that matter for that second case:
| field | when | what it says |
|---|---|---|
disposition | always | replaced, rehearsed, abStock (the A/B split left this compaction to the engine), fellBack, passedThrough, abortedFallback, skipped (the engine tried to compact this plugin's own fork loop and was declined). Historical rows only: overBudget, written by 0.10.0 and earlier when a session had spent past its USD ceiling. That ceiling is gone, and those rows still read back through every tool |
engine | always | the Claude Code version, from $.session.version() since 0.11.5, with CLAUDE_CODE_VERSION the fallback. Null on nearly every earlier row, because that variable is unset in an ordinary session |
fallbackReason | every disposition but replaced | one line naming why, e.g. noHandoff: no handoff has been written yet (fork nothing-to-fork: the fork found no warm main-thread transcript) |
forkOutcome, forkDetail | every row whose fork was not handed up | why the fork's own answer was not used, and the fallback went to the handoff on disk. On engine 2.1.280 and later an unanswered fork records the engine's own reason: nothing-to-fork (no warm transcript yet), api-error (the detail carries the HTTP status and error kind), empty-reply or aborted, and the spend of any that made a request stays in usage. timeout is a fork that had not answered after five minutes: the compaction stops waiting and falls back, but the engine gives a plugin no way to cancel a fork, so the request runs on and still spends, and its answer is dropped when it arrives (that spend is on no row). mismatch is an answer refused by the forkInput check below, threw anything unexpected. Rows from 0.10.0 and earlier say cold where they now say nothing-to-fork (or threw, on engine 2.1.280 and later) |
cost | always | {forkUsd, commitmentsUsd, commitmentsBasis, totalUsd, cacheReadWaivedUsd, basis, forkUsage, model, priced, pricesTaken}, or null. Since 0.4.1 cache reads are stored in forkUsage.cacheRead and priced into cacheReadWaivedUsd at list, and never added to forkUsd or totalUsd, because on a subscription they cost nothing; basis says so. Rows from 0.4.0 and earlier charged them. commitmentsBasis is measured: the usage the result reported after 0.10.0, estimate: chars/4, no cache from 0.4.0 (and since, on a result that arrives with no usage), none when no commitments pass ran, measured on rows from 0.3.0 and earlier |
| maxHandoff | every replaced and rehearsed row | the handoff ceiling the size guard used and how it was arrived at: {window, fraction, capTokens, chars, decider}, where decider is fraction, cap, default: window unknown or override; an override past the safe cap adds overSafeCapChars and `overSafeC
hooks/module.js 3155 lines1/**
2 * compact-handoff: the engine's compaction, replaced by a subagent's handoff.
3 *
4 * A `session.compact` hook is handed the whole transcript and its answer *is*
5 * the conversation afterwards, so returning `{ messages }` without calling
6 * `next` means the engine's summariser never runs. That much is real and this
7 * module does it. What the module cannot do is write the handoff while the
8 * event is waiting, and finding that out is what shaped everything below.
9 *
10 * **A hook dispatch is cut off after about ten seconds.** Measured at 2.1.269:
11 * a hook sitting on a file nobody writes is aborted at 10052 ms, and a real
12 * `/compact` was aborted at 10525 ms. A subagent takes thirty seconds to two
13 * minutes. So the handoff cannot be written inside the compaction, and any
14 * design that waits for one there falls back to the engine every single time.
15 *
16 * So the handoff is written *between* turns and only read at compaction. Every
17 * turn's end promotes a finished handoff into place and, when the conversation
18 * has moved enough, starts the subagent that writes the next one; `$.agent.spawn`
19 * returns in about 400 ms without waiting for it, which is a defect everywhere
20 * else and exactly what is wanted here. `session.compact` then does nothing but
21 * read a file that is already on disk, which costs milliseconds.
22 *
23 * The cost is staleness: the handoff describes the conversation as of the last
24 * refresh, not as of the compaction. `MAX_STALE_MESSAGES` bounds it, and past
25 * that bound the hook hands the compaction back to the engine rather than hand
26 * up a handoff that is missing the last hour.
27 *
28 * It rehearses by default. Unless `COMPACT_HANDOFF_LIVE` is on it does the whole
29 * thing, writes down the conversation it would have handed up, and then calls
30 * `next(e)` anyway, which is the engine compacting exactly as it does today.
31 * Every failure path does the same. The worst case of running this is the
32 * behaviour you already have, plus a log.
33 */
34
35const SCRATCH = ".runs";
36
37/** The handoff a compaction reads. Only ever written whole, by a promotion. */
38const LATEST = `${SCRATCH}/latest.md`;
39
40/**
41 * A subagent's Write lands in its own time, so it writes somewhere else and a
42 * later turn moves the finished text into `LATEST`. Without that, a compaction
43 * landing mid-write reads half a handoff and cannot tell.
44 */
45const SENTINEL = "<!-- handoff-complete -->";
46
47/** Set by the `force_compact` tool, read and cleared at the turn's end. */
48const ARMED_KEY = "armed";
49
50/** The refresh that is in flight: when it started and what it is reading. */
51const PENDING_KEY = "pending";
52
53/** The handoff on disk: when it landed and how long the transcript was then. */
54const READY_KEY = "ready";
55
56/** When this session last forked for a handoff, wall clock ms, for the fork's context row. */
57const LAST_FORK_KEY = "lastForkAt";
58
59/**
60 * What the last A/B compaction of a session left in its conversation and the
61 * handoffs in a row it ended, under `abChain:<session>`. One small key per
62 * session that ever compacted under the split; a resume reads it back, because
63 * a resumed transcript only holds its last opener.
64 */
65const abChainKey = (session) => `abChain:${session}`;
66
67/** The agent type this plugin defines for the handoff writer, `<plugin>:<name>`. */
68const HANDOFF_AGENT = "compact-handoff:handoff";
69
70/** Whether the type took this session, so a spawn knows which one to name. */
71const HANDOFF_AGENT_KEY = "handoffAgent";
72
73/**
74 * The A/B bench. A variant is a file rather than a tool argument so the
75 * instruction never enters the transcript: every arm is then asked over a
76 * byte-identical context, which is the only way two arms are comparable.
77 */
78const AB_PROMPTS = `${SCRATCH}/ab/prompts`;
79
80/** Where each arm's answer lands, one file per replicate, for scoring later. */
81const AB_OUT = `${SCRATCH}/ab/out`;
82
83/**
84 * Three environment variables steer this, and the static scan will only take
85 * them spelled out at the call site, so they are named here and nowhere else:
86 *
87 * - `COMPACT_HANDOFF_LIVE` off, the hook rehearses and lets the engine
88 * compact; on, it answers the event itself and the summariser never runs.
89 * - `COMPACT_HANDOFF_MODEL` an alias (`haiku`) or a full id for the subagent;
90 * unset lets the agent's own model, then the parent's, decide.
91 * - `COMPACT_HANDOFF_REFRESH_MS` the shortest gap between two refreshes.
92 */
93const DEFAULT_REFRESH_MS = 10 * 60 * 1000;
94
95/** No handoff at all until the conversation is long enough to need one. */
96const MIN_MESSAGES = 20;
97
98/** And no refresh until it has moved enough to be worth paying to re-read. */
99const MIN_NEW_MESSAGES = 20;
100
101/**
102 * How far behind the handoff may be and still be used. Past this the engine
103 * takes the compaction back, because a handoff this stale would drop the work
104 * the session is in the middle of.
105 */
106const MAX_STALE_MESSAGES = 60;
107
108/** A refresh still unfinished after this is treated as dead and restarted. */
109const ABANDON_AFTER_MS = 15 * 60 * 1000;
110
111/** How often the dispatch-budget probe looks for the file it is waiting on. */
112const POLL_MS = 1000;
113
114/**
115 * `$.fs` rejects a write over 4 MiB, and a prompt pointing at an enormous file
116 * is its own problem, so the transcript is clamped well under that.
117 */
118const MAX_TRANSCRIPT_CHARS = 1_500_000;
119
120/** How much of one tool result is worth keeping for a handoff to read. */
121const MAX_TOOL_TEXT_CHARS = 600;
122
123/** What a fork is asked when the probe names nothing: the handoff itself. */
124/**
125 * The manifest's `userConfig`, as `register` was handed it. Every setting here
126 * was an environment variable first and both spellings still work, the
127 * manifest's winning: a config-menu row is the discoverable half, and the
128 * variable is what a cron line already exports around the session.
129 */
130let options = {};
131
132/**
133 * Per-session watch state, keyed by session id: the compaction being watched,
134 * and each session's turn count for `window.jsonl`. Module memory rather than
135 * `$.store`, because the store is one file every session on the box shares:
136 * a counter kept there counted every session's turns as one, and a monitor
137 * kept there could be read and ended by whichever session ticked first. A hot
138 * reload forgets both; the post file keeps what it had and the next rows say
139 * `turn: null`.
140 */
141let watches = new Map();
142let windows = new Map();
143
144/**
145 * One `userConfig` value as a non-empty string, or `undefined` so that `??`
146 * falls through to the variable. Stringified, because every parser below takes
147 * the text `$.env.get` returns and a declared number or boolean has to read
148 * the same way to them. A number row left at `0` counts as unset, which is why
149 * each numeric default is `0` and the real default lives in the constant.
150 * A boolean row declares no default at all: the engine fills a declared default
151 * in as though the person had set it, and `false` is a setting, so a declared
152 * `false` would keep the variable from ever being read.
153 */
154const opt = (key) => {
155 const value = options[key];
156
157 if (value === undefined || value === null || value === "" || value === 0) return undefined;
158
159 return String(value).trim() || undefined;
160};
161
162/**
163 * The compaction instruction, and the bench's winner over five rounds.
164 *
165 * Rounds 1 to 3 varied the wording of the engine's own summariser prompt, which
166 * asks for nine numbered prose sections. Round 4 asked whether that inherited
167 * shape is the right one at all, and it is not. Prose makes the writer choose
168 * what is interesting, and a fact nobody finds interesting is exactly the fact
169 * the next session needed. So this prompt makes it enumerate before it narrates:
170 * a tagged one-line-per-fact inventory first, the prose reading second.
171 *
172 * Measured over the same 82-atom answer key, graded blind, two forks per arm:
173 * 70.0% carried against 63.0% for the nine-section prompt we shipped before, and
174 * it is the only arm whose worst run beat that prompt's best. Two arms that also
175 * restructured (an answer sheet of lettered sections, and the successor's ten
176 * questions as headings) scored 63.4% and 66.5% with spreads of 18.3 and 15.9,
177 * against this one's 7.3. High variance is disqualifying for a prompt that gets
178 * one attempt per compaction.
179 *
180 * In round 4 it cost about 2,600 more output tokens than the prompt it
181 * replaced, roughly 40% more, against a post-compaction floor measured at
182 * 54,600 to 66,058 tokens in a real session.
183 *
184 * Round 5 took out the <analysis> block that round 4's version wrote the
185 * inventory in first. PART 1 was a copy of it, so every compaction wrote the
186 * inventory twice and paid for both. Written once, straight into the summary,
187 * over the same fixture and answer key, two forks per arm graded three times
188 * each: 69.4% recall (spread 5.5) against 67.7% (spread 6.8) for the two-copy
189 * prompt, inside the grader's 3.5-point noise, for 16,777 output tokens against
190 * 24,756 and 182 s against 257 s per fork. `withoutScratchpad` still strips the
191 * block if a model writes one anyway.
192 *
193 * The <restore-files> block below is appended mechanically and was not part of
194 * the benched arm, the same way the Work ledger was appended to the old one.
195 *
196 * The one-reply, no-tools paragraph is there because a fork inherits the
197 * session's tools. A denied tool call does not end the fork, it becomes another
198 * request, and the engine sums the usage across them; a transcript that stops
199 * mid-task reads to the model like a turn in that task. One fork took the
200 * prompt that way, carried on the session's work and billed 144,745 output
201 * tokens over 24 minutes before it wrote a handoff.
202 */
203const FORK_PROMPT = `Your task is to compact this conversation into a handoff for the next session. It is the only thing that survives.
204
205Answer in a single reply and call no tools; everything you need is already in the conversation. Do not continue, finish or act on the task the conversation was working on, even if it stopped mid-step. Your only job now is to write the handoff.
206
207Summaries lose facts because they are written as prose, and prose makes the writer choose what is interesting. You will not choose. You will enumerate first, then narrate.
208
209In <summary> tags, produce the handoff in two parts:
210
211PART 1 — THE INVENTORY. Sweep the conversation from the first message to the last and emit an inventory: one line per discrete fact, no prose, no commentary. A discrete fact is anything a successor could be wrong about — a request, a constraint, a rejection, a decision, a reason, a file, a command run and its result, a number, an identifier, an error, a promise, an unfinished item, a correction. Aim for completeness over elegance; a hundred lines is normal and a short inventory means you skipped. Mark each line with one tag in brackets at the start: [ask] [constraint] [rejected] [decision] [file] [command] [identifier] [error] [promise] [pending] [state]. Group the lines by tag, in that tag order, keeping conversation order within each group. Do not drop a line because it seems minor, and do not merge two lines into one. Every message the user typed comes back word for word right after this handoff, so do not copy one onto an [ask] line: name the ask in a dozen words or fewer, then its status and any op id. An ask that appears only inside an earlier handoff, never as a message the user typed, does not come back; keep its line as it stands. On a [constraint] or [rejected] line quote the user verbatim, the sentence that sets the rule and not the whole message.
212
213PART 2 — THE READING. Now, and only now, write the prose a successor needs to make sense of Part 1: what the work is, what has been done, what is being done right now, and what the next step is, naming the user's most recent request without quoting it. Keep this short. It explains the inventory; it does not replace it.
214
215Two rules that override any instinct toward brevity. Never write "various", "several", "etc.", "and similar", or any phrase that stands in for items you could have named. Never state a fact the conversation did not establish — if you do not know the branch, the file or the number, write "not established", because a plausible invention is worse to a successor than a gap.
216
217There may be additional summarization instructions in the included context; follow them too.
218
219After the closing </summary> tag, add a <restore-files> block naming up to 5 files the next window should have open before it does anything: the file being edited, the spec or test it is being written against, the file the current step depends on. Only files this conversation actually read or wrote, through any tool including shell commands, by absolute path. One file per line, three fields separated by |: the absolute path, then either all or a line range like 120-260, then a one-line reason. Prefer a line range when only part of a large file matters. Do not name CLAUDE.md or AGENTS.md files; they come back on their own. Leave the block empty if nothing qualifies.
220<restore-files>
221/abs/path/to/file.ts | 40-120 | the function being changed
222</restore-files>`;
223
224import {
225 RESTORE_FILE_CHARS,
226 RESTORE_MAX_FILES,
227 RESTORE_TOTAL_CHARS,
228 chooseRestores,
229 clipRestore,
230 fitRestores,
231 parseRestoreRequests,
232 restoreCandidates,
233 shellMentions,
234 restorePair,
235 restoreRow,
236} from "./restore.js";
237import {
238 abArmFor,
239 abChainAfter,
240 abPriorOf,
241 applySizeGuard,
242 assistantTurns,
243 closedByCompaction,
244 commitmentsFrom,
245 commitmentsOutcome,
246 costOf,
247 costOfNothing,
248 estimatedUsage,
249 forkInputOf,
250 handoffCeiling,
251 fallbackReasonFor,
252 isPinnable,
253 ledgerRows,
254 lineageLine,
255 nameOf,
256 monitorSeed,
257 observePost,
258 openersIn,
259 originOf,
260 pinSkipCounts,
261 priorHandoffIn,
262 sectionOf,
263 stockSummaryOf,
264 priceUsage,
265 renderCommitments,
266 lookupRecord,
267 renderLedger,
268 replacementFor,
269 subagentRanThisTurn,
270 summarisePrs,
271 toolIndex,
272 windowReading,
273 withoutScratchpad,
274} from "./lib.js";
275
276export const register = (on, pluginOptions) => {
277 options = pluginOptions ?? {};
278 watches = new Map();
279 windows = new Map();
280
281 on("session.start", async ($, e, next) => {
282 await safely($, () => startWindow($));
283
284 await $.tool.register({
285 name: "force_compact",
286 description:
287 "Force a compaction of this conversation now, and report what it resolved to. " +
288 "For testing the compact-handoff plugin without waiting for the context to fill.",
289 inputSchema: {
290 type: "object",
291 properties: {
292 instructions: {
293 type: "string",
294 description: "What the handoff should keep or stress.",
295 },
296 },
297 },
298 });
299
300 await $.tool.register({
301 name: "handoff_status",
302 description:
303 "What compact-handoff would do if this conversation compacted right now: whether it is live, " +
304 "where its data lives, how many compactions this session has already been through, what the " +
305 "last one did, and what this session has spent on handoffs so far.",
306 inputSchema: {
307 type: "object",
308 properties: {
309 limit: { type: "number", description: "How many recent rows, newest last. Default 10." },
310 },
311 },
312 });
313
314 await $.tool.register({
315 name: "handoff_list",
316 description:
317 "List the compactions this session has been through, oldest first: when each ran, what it cost, " +
318 "how large its handoff was, and which files it left behind.",
319 inputSchema: { type: "object", properties: {} },
320 });
321
322 await $.tool.register({
323 name: "handoff_lookup",
324 description:
325 "Read back a stored compaction. Without arguments, the whole of the most recent handoff. " +
326 "`section` reads one part: summary, ledger, state, commitments, or transcript for the whole " +
327 "conversation as it stood before that compaction. Long sections page with offset and limit.",
328 inputSchema: {
329 type: "object",
330 properties: {
331 n: { type: "number", description: "Which compaction. Default the most recent." },
332 section: {
333 type: "string",
334 description: "full, summary, ledger, state, commitments, or transcript. Default full.",
335 },
336 offset: { type: "number", description: "First line returned, 0-based. Default 0." },
337 limit: { type: "number", description: "How many lines. Default 400." },
338 },
339 },
340 });
341
342 await $.tool.register({
343 name: "handoff_search",
344 description:
345 "Search every stored handoff and pre-compaction transcript of this session for a pattern and " +
346 "return the matching lines with the compaction they came from and their line numbers. This is " +
347 "how to find something a handoff did not carry up.",
348 inputSchema: {
349 type: "object",
350 properties: {
351 pattern: { type: "string", description: "A basic regular expression, as grep reads it." },
352 limit: { type: "number", description: "How many matching lines. Default 50." },
353 },
354 required: ["pattern"],
355 },
356 });
357
358 await $.tool.register({
359 name: "handoff_feedback",
360 description:
361 "Record that a handoff was missing something or wrong about something, against the compaction " +
362 "it came from. Written beside that compaction's own record, where the bench reads it.",
363 inputSchema: {
364 type: "object",
365 properties: {
366 note: { type: "string", description: "What was missing or wrong, in your own words." },
367 n: { type: "number", description: "Which compaction. Default the most recent." },
368 },
369 required: ["note"],
370 },
371 });
372
373 if (await isDev($)) {
374 await registerBenchTools($);
375 }
376
377 await registerHandoffAgent($);
378
379 return next(e);
380 });
381
382 // Nothing should delegate to this agent but this plugin. `agent.offer`
383 // hides it from the model's listing without hiding it from
384 // `$.agent.spawn`, which is what a runner-only type is for.
385 on("agent.offer", { agent: HANDOFF_AGENT }, () => ({ isOffered: false }));
386
387
388 // The tool only arms it. `$.session.compact` refuses to run under a hook
389 // that is holding the turn, and says so: "called from a tool.call hook, it
390 // would compact under the turn this hook is holding; call it from a later
391 // event (turn.complete)". So the tool writes a flag and the turn's end
392 // reads it.
393 on("tool.call", { tool: "mcp__compact-handoff__force_compact" }, async ($, e) => {
394 const instructions = typeof e.instructions === "string" ? e.instructions.trim() : "";
395
396 await $.store.set(ARMED_KEY, { instructions });
397
398 return { result: "Armed. The compaction runs when this turn ends; read it back with handoff_status." };
399 });
400
401 on("tool.call", { tool: "mcp__compact-handoff__refresh_handoff" }, async ($) => {
402 await $.store.set(ARMED_KEY, { refresh: true });
403
404 return { result: "Armed. The refresh starts when this turn ends; watch it with handoff_status." };
405 });
406
407 on("tool.call", { tool: "mcp__compact-handoff__probe_spawn" }, async ($, e, next) => {
408 const seconds = typeof e.seconds === "number" && e.seconds > 0 ? Math.floor(e.seconds) : 30;
409
410 // Two hooks can hold a spawn and they need not behave alike, so the
411 // probe runs in whichever the caller names.
412 if (e.where === "wait") {
413 await probeWait($, seconds, next.signal);
414
415 return { result: "Waited inside the tool.call hook; read it back with handoff_status." };
416 }
417
418 if (e.where === "tool") {
419 await probeSpawn($, seconds, next.signal);
420
421 return { result: "Probed from inside the tool.call hook; read it back with handoff_status." };
422 }
423
424 await $.store.set(ARMED_KEY, { probeSeconds: seconds });
425
426 return { result: `Armed. A subagent sleeping ${seconds}s runs when this turn ends; read it back with handoff_status.` };
427 });
428
429 on("tool.call", { tool: "mcp__compact-handoff__probe_budget" }, async ($, e, next) => {
430 // A fork reads the main thread's transcript, and a tool.call hook is
431 // holding the turn that transcript belongs to, so the probe can also
432 // be run from the turn's end instead.
433 if (e.at === "turn") {
434 await $.store.set(ARMED_KEY, { forkProbe: true });
435
436 return { result: "Armed. The probe runs when this turn ends; read it back with handoff_status." };
437 }
438
439 return { result: JSON.stringify(await probeBudget($, e, next.signal), null, 2) };
440 });
441
442 // Same reason `force_compact` only arms: a fork reads the main thread's
443 // transcript and this hook is holding the turn that transcript belongs to,
444 // so from here it waits out the turn and answers null.
445 on("tool.call", { tool: "mcp__compact-handoff__ab_fork" }, async ($, e) => {
446 const label = typeof e.label === "string" ? e.label.trim() : "";
447
448 if (label === "" || !/^[\w.-]+$/.test(label)) {
449 return { result: "A label must be a plain filename: letters, digits, dot, dash, underscore." };
450 }
451
452 const replicates = typeof e.replicates === "number" && e.replicates > 0 ? Math.min(Math.floor(e.replicates), 10) : 1;
453
454 if (!(await $.fs.exists(atRoot($, `${AB_PROMPTS}/${label}.txt`)))) {
455 return { result: `No variant at ${AB_PROMPTS}/${label}.txt.` };
456 }
457
458 await $.store.set(ARMED_KEY, { ab: { label, replicates } });
459
460 return { result: `Armed ${label} x${replicates}. It runs when this turn ends; read it back with handoff_status.` };
461 });
462
463 on("tool.call", { tool: "mcp__compact-handoff__handoff_status" }, async ($, e) => {
464 return { result: JSON.stringify(await readRuns($, e), null, 2) };
465 });
466
467 on("tool.call", { tool: "mcp__compact-handoff__handoff_list" }, async ($, e) => {
468 return { result: await logged($, "handoff_list", e, JSON.stringify(await listHandoffs($), null, 2)) };
469 });
470
471 on("tool.call", { tool: "mcp__compact-handoff__handoff_lookup" }, async ($, e) => {
472 return { result: await logged($, "handoff_lookup", e, await lookupHandoff($, e)) };
473 });
474
475 on("tool.call", { tool: "mcp__compact-handoff__handoff_search" }, async ($, e) => {
476 return { result: await logged($, "handoff_search", e, await searchHandoffs($, e)) };
477 });
478
479 on("tool.call", { tool: "mcp__compact-handoff__handoff_feedback" }, async ($, e) => {
480 return { result: await recordFeedback($, e) };
481 });
482
483 // Everything expensive happens here, between turns, where a dispatch that
484 // runs long delays nothing the user is waiting on.
485 //
486 // One step that throws used to cost the four after it and `next(e)` with
487 // them, which is how a single `TypeError` in the watch took the whole
488 // dispatch down on every turn at 2.1.273. Each step is wrapped now, so a
489 // failing one costs its own reading and nothing else.
490 on("turn.complete", async ($, e, next) => {
491 const armed = await safely($, () => $.store.get(ARMED_KEY));
492
493 if (armed !== undefined && armed !== null) {
494 await safely($, () => $.store.delete(ARMED_KEY));
495 await safely($, () => runArmed($, armed, next.signal));
496 }
497
498 await safely($, () => promotePending($));
499 await safely($, () => maybeRefresh($, armed?.refresh === true));
500 await safely($, () => watchPost($));
501 await safely($, () => watchWindow($));
502
503 return next(e);
504 });
505
506 // `precompute` is the engine building a compaction before it is needed, so
507 // that the real one is instant. Answering it with `next(e)` was a hole: the
508 // declarations say a precomputed result "is kept for the compaction that
509 // comes, if the conversation it ran over still leads", which is exactly the
510 // automatic path, so an auto-compaction could be served an engine summary
511 // this plugin had waved through and no row would ever say so. Declining the
512 // precompute forces the engine to dispatch a real event when it actually
513 // compacts. This is the declarations' own example for the case.
514 on("session.compact", { trigger: "precompute" }, () => ({
515 skip: "compact-handoff answers compactions live",
516 }));
517
518 on("session.compact", async ($, e, next) => {
519 const startedAt = Date.now();
520 const record = {
521 at: new Date().toISOString(),
522 sessionId: await safely($, () => $.session.id()),
523 cwd: await safely($, () => $.session.cwd()),
524 model: await safely($, () => $.session.model()),
525 trigger: e.trigger,
526 agentId: e.agentId ?? null,
527 messagesIn: e.messages.length,
528 // What the fork was charged to read, against what this session
529 // holds; null on every row that never reached a fork. Issue #690:
530 // a fork answering over a transcript that is not this conversation
531 // is invisible in every other field on the row.
532 forkInput: null,
533 pinnable: e.messages.filter(isPinnable).length,
534 pinnedSkipped: pinSkipCounts(e.messages),
535 plugin: await pluginVersion($),
536 engine: await engineVersion($),
537 // The arm of the live A/B split this compaction is in, the design
538 // that picked it and what it was decided on, or null when the
539 // split is off. Set below, once the compaction is known to be the
540 // main conversation's own.
541 ab: /** @type {AbRow | null} */ (null),
542 };
543
544 // The fork this plugin is waiting on is itself a query loop, and the
545 // engine checks it for auto-compaction like any other. Passing it
546 // through lets the engine summarise it, and the fork then answers over
547 // that summary instead of this conversation: a cold fork. Declining it
548 // lets the fork read the real transcript.
549 if (isOwnForkLoop(e)) {
550 record.outcome = "ownFork";
551 record.detail = "the engine tried to compact this plugin's own handoff fork; declined so the fork reads the conversation";
552 record.elapsedMs = Date.now() - startedAt;
553
554 await finish($, record, "skipped");
555
556 return { skip: "compact-handoff's own handoff fork reads the conversation uncompacted" };
557 }
558
559 // A subagent's compaction is a different conversation with a different
560 // owner, and nothing here has been measured against one. Pass it
561 // through with a row rather than reshape a transcript this plugin has
562 // never been graded on.
563 if (e.agentId !== null && e.agentId !== undefined && !(await handlesSubagents($))) {
564 record.outcome = "subagent";
565 record.elapsedMs = Date.now() - startedAt;
566
567 await finish($, record, "passedThrough");
568
569 return next(e);
570 }
571
572 record.ab = await abRowFor($, record);
573
574 // A subagent's compaction is another conversation; the main session's
575 // watch is not ended by it.
576 if (record.agentId === null) {
577 await safely($, () => closeWatch($, record.sessionId));
578 }
579
580 if (record.ab?.arm === "stock") {
581 return compactStock($, e, next, record, startedAt);
582 }
583
584 // Another plugin's work starts here, before anything is awaited, so it
585 // runs beside this plugin's fork over the same pre-compaction
586 // transcript and both read one warm cache. With nobody subscribed this
587 // is one array copy.
588 const seam = dispatchSeam($, e);
589
590 const context = await lineageContext($, e, record);
591
592 let handoff = { outcome: "threw", detail: "" };
593
594 try {
595 // A fork answers over this session's own transcript, so it needs
596 // nothing written in advance and it can honour the instructions a
597 // `/compact <instructions>` carries. The handoff on disk is only
598 // what stands when the fork has no warm transcript to read.
599 handoff = await forkHandoff($, e, record);
600
601 if (handoff.outcome !== "handoff") {
602 record.forkOutcome = handoff.outcome;
603 record.forkDetail = handoff.detail;
604 handoff = await readHandoff($, e);
605 }
606 } catch (error) {
607 handoff = { outcome: "threw", detail: String(error) };
608 }
609
610 record.seam = await seam;
611 record.outcome = handoff.outcome;
612 // A fork that ran and was refused still spent its tokens, so a usage
613 // already on the record outlives the fallback that replaced it.
614 record.usage = handoff.usage ?? record.usage ?? null;
615 record.staleMessages = handoff.staleMessages ?? null;
616 record.elapsedMs = Date.now() - startedAt;
617 record.aborted = next.signal.aborted;
618
619 if (handoff.detail !== "") {
620 record.detail = handoff.detail.slice(0, 2000);
621 }
622
623 if (handoff.outcome !== "handoff") {
624 priceRun(record, { commitmentsUsd: 0 });
625
626 await finish($, record, "fellBack", { transcript: renderTranscript(e.messages) });
627
628 const ordinal = await openerOrdinal($, record);
629 const result = await next(e);
630
631 armStock(record, ordinal, e.messages, result);
632
633 return result;
634 }
635
636 // The model's summary is only the first of four parts. The rest are read
637 // or gathered rather than recalled, and each one is allowed to fail: a
638 // summary on its own is still the arm that scored 67.3%.
639 const assembled = await assembledHandoff($, e, handoff.text, context, record.depth ?? null);
640 const maxHandoff = await handoffCeilingFor($);
641 const guarded = applySizeGuard(replacementFor(e.messages, assembled.text), {
642 maxChars: maxHandoff.chars,
643 pointerFor: ({ index, chars }) => trimPointer(record, index, chars),
644 });
645
646 record.maxHandoff = maxHandoff;
647
648 const restored = await restoreFiles($, e, handoff.requests ?? [], {
649 depth: record.depth ?? 0,
650 totalChars: Math.max(0, maxHandoff.chars - guarded.charsAfter),
651 });
652 // The handoff first, then the files it asked for, then the pinned turns.
653 const replacement = [guarded.messages[0], ...restored.messages, ...guarded.messages.slice(1)];
654
655 record.messagesOut = replacement.length;
656 record.restore = restored.restore;
657 record.summaryChars = handoff.text.length;
658 record.analysisChars = handoff.analysisChars ?? 0;
659 record.handoffChars = assembled.text.length;
660 record.handoffCharsBefore = guarded.charsBefore;
661 record.handoffCharsAfter = guarded.charsAfter;
662 record.trimmedTurns = guarded.trimmedTurns;
663 record.overCeiling = guarded.overCeiling;
664 record.parts = assembled.notes;
665 record.elapsedMs = Date.now() - startedAt;
666 record.aborted = next.signal.aborted;
667
668 priceRun(record, {
669 commitmentsUsd: assembled.notes.commitmentsCostUsd,
670 commitmentsBasis: assembled.notes.commitmentsCostBasis,
671 });
672
673 const artifacts = {
674 handoff: assembled.text,
675 summary: handoff.text,
676 transcript: renderTranscript(e.messages),
677 rows: assembled.rows,
678 };
679
680 if (!(await isLive($))) {
681 artifacts.replacement = renderTranscript(replacement);
682
683 await finish($, record, "rehearsed", artifacts);
684
685 const ordinal = await openerOrdinal($, record);
686 const result = await next(e);
687
688 armStock(record, ordinal, e.messages, result);
689
690 return result;
691 }
692
693 // A dispatch the engine has already given up on cannot be trusted to
694 // have its answer applied, and half a compaction is worse than the
695 // engine's own. Measured in the live checks; the row says which path.
696 if (next.signal.aborted) {
697 await finish($, record, "abortedFallback", artifacts);
698
699 const ordinal = await openerOrdinal($, record);
700 const result = await next(e);
701
702 armStock(record, ordinal, e.messages, result);
703
704 return result;
705 }
706
707 const ordinal = await openerOrdinal($, record);
708
709 await finish($, record, "replaced", artifacts);
710
711 if (ordinal !== null) {
712 armMonitor(record, { kind: "handoff", ordinal, from: replacement.length, before: e.messages });
713 restartWindow(record.sessionId);
714 }
715
716 return { messages: replacement };
717 });
718
719 // The noun another plugin reaches this one through. The fold runs once per
720 // load and before any hook of this plugin, and `built` is `$` as the steps
721 // beneath made it. Nothing may be done with `built` but spread it: at
722 // `engine.create` the static scan refuses the value of `next(e)` passed to
723 // a function, and handing it to `pluginVersion` here is what refused 0.5.0
724 // at load ("the value of next(e) at engine.create is passed as an argument").
725 on("engine.create", async (_$, e, next) => {
726 const built = await next(e);
727
728 return { ...built, compactHandoff: { beforeCompact, version } };
729 });
730};
731
732/* ------------------------------------------------------------------ *
733 * The seam: another plugin's work, run beside this plugin's fork.
734 *
735 * The seam carries strings and nothing else. Revision 1 took a callback, and a
736 * callback cannot cross a plugin boundary here at all: every plugin runs in its
737 * own environment and an interface call's arguments go through `cloneInto`,
738 * which throws `DataCloneError` on a function. So a subscriber names a TOOL it
739 * answers, and this plugin raises that tool through `$.tool.call`, which runs
740 * the subscriber's work in the subscriber's own environment.
741 * ------------------------------------------------------------------ */
742
743/**
744 * The version `$.compactHandoff.version()` answers, and a constant because the
745 * `engine.create` fold may not pass `built` to a function and `$` there is the
746 * empty table, so nothing at the fold can read the manifest. `test/module.test.js`
747 * asserts it against `.claude-plugin/plugin.json` so the two cannot drift.
748 */
749export const PLUGIN_VERSION = "0.13.0";
750
751/** How long one subscriber may run before the compaction goes on without it. */
752const DEFAULT_SEAM_TIMEOUT_MS = 90_000;
753
754/**
755 * Every subscriber, keyed by the tool it answers: subscribing twice is once.
756 *
757 * Exported so a test can start from an empty seam. A subscription lives for the
758 * load and nothing in the runtime ever removes one.
759 */
760export const seamSubscribers = new Map();
761
762/**
763 * Registers the tool this plugin raises when a compaction is about to happen,
764 * beside its own fork. Resolves `{ subscribed: true, tool }`.
765 *
766 * This is the whole public API of `$.compactHandoff` beside `version`, and it
767 * exists because a plugin keyed after this one in `enabledPlugins` never sees
768 * `session.compact` at all: this plugin answers that event without calling
769 * `next`. There is no unsubscribe, and a subscription lives for the load.
770 */
771const beforeCompact = async (options) => {
772 const tool = typeof options?.tool === "string" ? options.tool.trim() : "";
773
774 if (tool === "") {
775 throw new TypeError("compactHandoff.beforeCompact takes { tool }, the full name of the tool to raise");
776 }
777
778 if (seamSubscribers.has(tool)) {
779 return { subscribed: true, tool };
780 }
781
782 const name = typeof options?.name === "string" && options.name.trim() !== "" ? options.name.trim() : tool;
783
784 seamSubscribers.set(tool, { tool, name });
785
786 return { subscribed: true, tool };
787};
788
789/** This plugin's version, for a subscriber that wants to know what it got. */
790const version = async () => PLUGIN_VERSION;
791
792/**
793 * Every subscriber, raised at once and waited for together.
794 *
795 * Nothing a subscriber does can change what this plugin answers the compaction
796 * with: a throw and a deny are row fields, a subscriber still running after the
797 * timeout is left running and its result ignored. The compaction waits for the
798 * slower of this and the fork, which is the price of both reading the same
799 * transcript while its cache is warm.
800 */
801const dispatchSeam = ($, e) => {
802 const waiting = [...seamSubscribers.values()];
803
804 if (waiting.length === 0) {
805 return Promise.resolve({ subscribers: 0 });
806 }
807
808 const started = waiting.map((entry) => ({
809 name: entry.name,
810 tool: entry.tool,
811 startedAt: Date.now(),
812 work: settledCall($, entry.tool, e),
813 }));
814
815 // A host op failing while the seam is collected is a field on the row like
816 // any other reading here, never a compaction this plugin drops.
817 return collectSeam($, started).catch((error) => ({
818 subscribers: started.length,
819 failed: String(error).slice(0, 300),
820 }));
821};
822
823/**
824 * The subscriber's tool, raised now, as a promise that resolves however it ends.
825 *
826 * The raise carries three small values and no transcript. The subscriber reads
827 * the conversation the way this plugin does, with `$.model.fork` over the live
828 * session, which is the same pre-compaction transcript and the same warm cache;
829 * sending the messages through would clone a megabyte to say what a fork reads
830 * for nothing.
831 */
832const settledCall = ($, tool, e) => {
833 const threw = (error) => ({ outcome: "threw", detail: String(error).slice(0, 300) });
834
835 try {
836 return Promise.resolve(
837 $.tool.call({ tool, trigger: e.trigger, messageCount: e.messages.length }),
838 ).then(seamOutcome, threw);
839 } catch (error) {
840 return Promise.resolve(threw(error));
841 }
842};
843
844/** What the raise resolved to, read as an outcome: an answer, or a refusal. */
845const seamOutcome = (answer) => {
846 if (answer !== null && typeof answer === "object" && typeof answer.deny === "string") {
847 return { outcome: "denied", detail: answer.deny.slice(0, 300) };
848 }
849
850 if (answer === null || typeof answer !== "object" || answer.result === undefined) {
851 return { outcome: "ok", detail: "the raise was answered with no result" };
852 }
853
854 return { outcome: "ok" };
855};
856
857const collectSeam = async ($, started) => {
858 const capMs = capOf(opt("seamTimeoutMs") ?? (await $.env.get("COMPACT_HANDOFF_SEAM_TIMEOUT_MS")), DEFAULT_SEAM_TIMEOUT_MS);
859
860 const results = await Promise.all(
861 started.map(async (entry) => {
862 const ended = await Promise.race([
863 entry.work,
864 $.clock.sleep(capMs).then(() => ({ outcome: "timedOut" })),
865 ]);
866
867 return { name: entry.name, tool: entry.tool, ...ended, elapsedMs: Date.now() - entry.startedAt };
868 }),
869 );
870
871 return { subscribers: started.length, results };
872};
873
874/** How long one restore read may take before the file is read from the transcript instead. */
875const RESTORE_MS = 20_000;
876
877/**
878 * The files the summariser asked for, read through the real Read tool and
879 * shaped as the tool blocks a live read produces.
880 *
881 * Fresh first: the engine's own restore re-reads from disk, so a file edited
882 * mid-session comes back current, and `$.tool.call` is a host op, so its time
883 * in flight does not count against the dispatch budget. A read that is denied,
884 * errors or runs past `RESTORE_MS` falls back to the text of the last Read in
885 * the transcript, which is stale at worst and absent at best; the row names
886 * which of the three it was. Every cap is an env var, and every number that
887 * would let the caps be tuned is on the row.
888 */
889const restoreFiles = async ($, e, requests, { depth, totalChars }) => {
890 const startedAt = Date.now();
891 // Three literal reads rather than one helper: the static scan lists the
892 // variables a module reads, and it can only do that off a literal name.
893 const caps = {
894 maxFiles: capOf(opt("restoreFiles") ?? (await $.env.get("COMPACT_HANDOFF_RESTORE_FILES")), RESTORE_MAX_FILES),
895 fileChars: capOf(opt("restoreFileChars") ?? (await $.env.get("COMPACT_HANDOFF_RESTORE_FILE_CHARS")), RESTORE_FILE_CHARS),
896 totalChars: Math.min(capOf(opt("restoreTotalChars") ?? (await $.env.get("COMPACT_HANDOFF_RESTORE_TOTAL_CHARS")), RESTORE_TOTAL_CHARS), totalChars),
897 };
898 const chosen = await chooseRestores({
899 requests,
900 candidates: restoreCandidates(e.messages),
901 shellMentions: shellMentions(e.messages),
902 maxFiles: caps.maxFiles,
903 exists: (path) => existsOnDisk($, path),
904 });
905 const read = [];
906
907 for (const file of chosen.files) {
908 read.push(await readRestore($, file, caps.fileChars));
909 }
910
911 const fitted = fitRestores(read, caps.totalChars);
912 const messages = fitted
913 .filter((file) => file.text !== null)
914 .flatMap((file, index) => restorePair({ ...file, id: `toolu_handoff_${depth}_${index}` }));
915 const rows = fitted.map(restoreRow);
916
917 return {
918 messages,
919 restore: {
920 source: chosen.source,
921 requested: requests.length,
922 restored: messages.length / 2,
923 rejected: chosen.rejected,
924 files: rows,
925 chars: rows.reduce((total, row) => total + row.chars, 0),
926 approxTokens: rows.reduce((total, row) => total + row.approxTokens, 0),
927 caps,
928 ms: Date.now() - startedAt,
929 },
930 };
931};
932
933/**
934 * Whether a file is still there; an unanswerable check counts as yes, so the read decides.
935 *
936 * @param {import('claude-code').EngineInterface} $
937 * @param {string} path
938 */
939const existsOnDisk = async ($, path) => {
940 try {
941 return await $.fs.exists(path);
942 } catch {
943 return true;
944 }
945};
946
947const readRestore = async ($, file, fileChars) => {
948 const startedAt = Date.now();
949 const row = { ...file, text: null, source: "failed", detail: "" };
950
951 try {
952 const fresh = await withTimeout($, freshRead($, file), RESTORE_MS, `restore ${file.path}`);
953
954 if (fresh !== null) {
955 Object.assign(row, { text: fresh, source: "fresh" });
956 } else {
957 row.detail = "the Read tool denied or errored";
958 }
959 } catch (error) {
960 row.detail = String(error).slice(0, 200);
961 }
962
963 if (row.text === null && typeof file.stored === "string") {
964 Object.assign(row, { text: file.stored, source: "stored" });
965 }
966
967 if (row.text !== null) {
968 const clipped = clipRestore(row.text, fileChars);
969
970 row.text = clipped.text;
971 row.clippedChars = clipped.clippedChars;
972 }
973
974 row.ms = Date.now() - startedAt;
975
976 return row;
977};
978
979/** The Read tool's own text for the file, or null when it would not read it. */
980const freshRead = async ($, { path, from, to }) => {
981 const input = { tool: "Read", file_path: path };
982
983 if (from !== null && to !== null) {
984 input.offset = from;
985 input.limit = to - from + 1;
986 }
987
988 const answer = await $.tool.call(input);
989
990 if (answer === null || typeof answer !== "object" || typeof answer.deny === "string" || answer.isError === true) {
991 return null;
992 }
993
994 return typeof answer.text === "string" && answer.text !== "" ? answer.text : null;
995};
996
997const capOf = (value, fallback) => {
998 const raw = Number.parseInt(value ?? "", 10);
999
1000 return Number.isFinite(raw) && raw > 0 ? raw : fallback;
1001};
1002
1003/**
1004 * The handoff, written inside the event by a fork of this very session.
1005 *
1006 * `$.model.fork` runs one completion over the main thread's own
1007 * cache-safe transcript snapshot, which is how the engine's own compaction
1008 * reads a conversation, so nothing has to be marshalled in and the prompt
1009 * cache is already warm. It is a host op, and a host op's time in flight does
1010 * not count against the ten second dispatch budget: measured at 21768 ms from
1011 * `turn.complete`, not aborted.
1012 *
1013 * The fork resolves `ModelForkResult`, the union engine 2.1.280 introduced.
1014 * Engines before it resolved the reply or null, and that shape is not read any
1015 * more: the plugin targets 2.1.280 and later, and a dual-shape shim would hide
1016 * the next change the way the null check hid this one. An unanswered fork is a
1017 * named outcome (`nothing-to-fork`, which is no warm transcript and what the
1018 * handoff on disk is kept for, then `api-error`, `empty-reply` and `aborted`),
1019 * and the caller falls back on it like any other. So is a fork that runs past
1020 * `FORK_TIMEOUT_MS`, as `timeout`.
1021 *
1022 * A reply is not proof that the fork read this conversation. Some forks come
1023 * back having been charged for a fraction of the session's context (#690), and
1024 * the reply reads like any other summary, so `record.forkInput` is compared
1025 * against the context and a short one is refused here rather than handed up.
1026 *
1027 * @param {import('claude-code').EngineInterface} $
1028 */
1029const forkHandoff = async ($, e, record) => {
1030 const asked = typeof e.instructions === "string" && e.instructions.trim() !== ""
1031 ? `${FORK_PROMPT}\n\nThe person asked for this compaction with these instructions, and they outrank everything above: ${e.instructions.trim()}`
1032 : FORK_PROMPT;
1033
1034 record.forkContext = await forkContext($, e);
1035
1036 let reply;
1037
1038 try {
1039 reply = await withTimeout($, $.model.fork({ prompt: asked }), FORK_TIMEOUT_MS, "the fork");
1040 } catch (error) {
1041 if (error instanceof TimeoutError) {
1042 return { outcome: "timeout", detail: error.message };
1043 }
1044
1045 throw error;
1046 }
1047
1048 if (!reply.isAnswered) {
1049 // Every arm but `nothing-to-fork` made a request, and what it spent
1050 // stays on the row after the fallback replaces its answer.
1051 if (reply.reason !== "nothing-to-fork") {
1052 record.usage = reply.usage;
1053 }
1054
1055 return { outcome: reply.reason, detail: unansweredForkDetail(reply) };
1056 }
1057
1058 record.forkInput = forkInputOf(reply.usage, record.forkContext?.context ?? null);
1059
1060 // A fork charged for a fraction of the context answered over a transcript
1061 // that is not this conversation, and a summary of the wrong conversation is
1062 // worse than the engine's own. The tokens are spent either way, so the row
1063 // keeps the usage and is priced on it.
1064 if (record.forkInput.matchesContext === false) {
1065 record.usage = reply.usage;
1066
1067 return {
1068 outcome: "mismatch",
1069 detail:
1070 `the fork was charged for ${record.forkInput.sent} input tokens against a ` +
1071 `${record.forkInput.contextTokens} token context, so it did not read this conversation`,
1072 };
1073 }
1074
1075 const parsed = parseRestoreRequests(reply.text.trim());
1076 const kept = withoutScratchpad(parsed.summary);
1077
1078 if (kept.text === "") {
1079 return { outcome: "empty", detail: "the fork answered nothing" };
1080 }
1081
1082 return {
1083 outcome: "handoff",
1084 text: kept.text,
1085 analysisChars: kept.analysisChars,
1086 requests: parsed.requests,
1087 usage: reply.usage,
1088 detail: "",
1089 };
1090};
1091
1092/**
1093 * How long a compaction waits on its fork before it falls back to the handoff
1094 * on disk.
1095 *
1096 * Five minutes, from the index this plugin writes: across 58 warm forks since
1097 * 0.6.0 the median took about two minutes, the 95th percentile under four, and
1098 * one alone ran past five (about 25 minutes, with the person waiting on it the
1099 * whole time). A fork that is still going at five minutes is far more likely
1100 * to be the stuck one than a slow good one, and the forks this bound cuts most
1101 * often are the short ones the `forkInput` check refuses anyway.
1102 *
1103 * The bound gives the person their session back. It does not stop the spend:
1104 * `$.model.fork` takes a `prompt` and nothing else (no signal, no `timeoutMs`,
1105 * unlike `$.model.complete`), and its `aborted` arm fires only when the turn's
1106 * own dispatch is aborted, so there is no way to cancel the request from here.
1107 * The abandoned fork runs to its end, and whatever it answers is dropped.
1108 */
1109const FORK_TIMEOUT_MS = 300_000;
1110
1111/**
1112 * Why a fork came back without a reply, in the words the row keeps.
1113 *
1114 * @param {Exclude<import('claude-code').ModelForkResult, { isAnswered: true }>} reply
1115 * @returns {string}
1116 */
1117const unansweredForkDetail = (reply) => {
1118 switch (reply.reason) {
1119 case "nothing-to-fork":
1120 return "the fork found no warm main-thread transcript";
1121 case "api-error":
1122 return `the fork's request failed: ${reply.error}, status ${reply.status ?? "none, no response arrived"}`;
1123 case "empty-reply":
1124 return "the fork replied with no text";
1125 case "aborted":
1126 return "the fork was cut off by the turn's abort";
1127 default:
1128 return unreachable(reply);
1129 }
1130};
1131
1132/**
1133 * A union arm no case above handles. The `never` makes a new arm in the
1134 * engine's declarations a type error here rather than a silent fallthrough.
1135 *
1136 * @param {never} value
1137 * @returns {never}
1138 */
1139const unreachable = (value) => {
1140 throw new Error(`an arm nothing handles: ${JSON.stringify(value)}`);
1141};
1142
1143/**
1144 * What the session looked like the instant before it forked, logged and nothing
1145 * else. Every reading is allowed to fail on its own: a missing row here must
1146 * never cost the fork.
1147 */
1148const forkContext = async ($, e) => {
1149 const now = Date.now();
1150 const usage = await safely($, () => $.session.usage());
1151 const lastForkAt = await safely($, () => $.store.get(LAST_FORK_KEY));
1152
1153 await safely($, () => $.store.set(LAST_FORK_KEY, now));
1154
1155 return {
1156 context: usage?.context ?? null,
1157 model: await safely($, () => $.session.model()),
1158 messages: e.messages.length,
1159 msSinceLastFork: typeof lastForkAt === "number" ? now - lastForkAt : null,
1160 subagentRanThisTurn: subagentRanThisTurn(e.messages),
1161 };
1162};
1163
1164/**
1165 * The read half of a compaction, and all of it that runs inside the event: a
1166 * handoff already on disk, or a named reason the engine should take this one.
1167 */
1168const readHandoff = async ($, e) => {
1169 // A `/compact <instructions>` asks for something this handoff was written
1170 // before anyone asked for, and nothing here can rewrite it in time.
1171 if (typeof e.instructions === "string" && e.instructions.trim() !== "") {
1172 return { outcome: "instructed", detail: "the compaction carried instructions a precomputed handoff cannot honour" };
1173 }
1174
1175 if (!(await $.fs.exists(atRoot($, LATEST)))) {
1176 return { outcome: "noHandoff", detail: "no handoff has been written yet" };
1177 }
1178
1179 const text = (await $.fs.read(atRoot($, LATEST))).trim();
1180
1181 if (text === "") {
1182 return { outcome: "noHandoff", detail: "the handoff on disk is empty" };
1183 }
1184
1185 const ready = await $.store.get(READY_KEY);
1186 const staleMessages = stalenessOf(e.messages.length, ready);
1187
1188 if (staleMessages === null || staleMessages > MAX_STALE_MESSAGES) {
1189 return { outcome: "tooStale", staleMessages, detail: "the handoff predates too much of this conversation" };
1190 }
1191
1192 return { outcome: "handoff", text, staleMessages, detail: "" };
1193};
1194
1195/**
1196 * How much conversation happened after the handoff was written, which is the
1197 * only thing that decides whether it can still stand in for one.
1198 */
1199const stalenessOf = (messagesNow, ready) =>
1200 typeof ready?.messages === "number" ? Math.max(0, messagesNow - ready.messages) : null;hooks/restore.js 346 lines1/**
2 * Restoring files after a compaction, chosen by the summariser.
3 *
4 * Claude Code's own compaction hands the next window up to five of the most
5 * recently read files, and it renders each as a meta user message narrating a
6 * Read call inside a system-reminder. That path never runs when a hook
7 * replaces the messages, so until 0.3.0 a handoff came back with no files at
8 * all.
9 *
10 * This is the pure half of the replacement. The summariser, which has read the
11 * whole conversation, names the files the next window should have open, in a
12 * fenced block at the end of its summary; the code here parses that block,
13 * checks every path against the files the session actually touched, falls
14 * back to recency when the model named nothing usable, and shapes each file
15 * as a real Read tool_use and its tool_result, which is what a live read looks
16 * like. (Operator, 2026-09-15: "Why wouldn't we let the model that handles the
17 * compaction summary decide which files are re-read and injected", and the
18 * caps are theirs to move: "if the size guard needs expanded we can expand
19 * it".)
20 *
21 * Nothing in this file touches the host; the reads happen in `module.js`.
22 */
23
24import { PATH_SHAPED, SHELL_EXPANDED, nameOf } from "./lib.js";
25
26/** How many files may come back; Claude Code's own restore stops at five. */
27export const RESTORE_MAX_FILES = 5;
28
29/** The most of one file that comes back, at four characters per token. */
30export const RESTORE_FILE_CHARS = 20_000;
31
32/** The most all restored files may add to the window together. */
33export const RESTORE_TOTAL_CHARS = 100_000;
34
35export const RESTORE_OPEN = "<restore-files>";
36export const RESTORE_CLOSE = "</restore-files>";
37
38/** The tools whose target is a file the session has seen. */
39export const RESTORE_TOOLS = new Set(["Read", "Edit", "Write", "MultiEdit", "NotebookEdit"]);
40
41/** Memory files come back through the engine's own path; restoring them doubles them. */
42export const MEMORY_FILE = /(^|\/)(CLAUDE\.md|CLAUDE\.local\.md|AGENTS\.md)$/u;
43
44/** One request line: an absolute path, `all` or a line range, and a reason. */
45const REQUEST_LINE = /^\s*`?([^`|]+?)`?\s*\|\s*([^|]*?)\s*(?:\|\s*(.*?))?\s*$/u;
46
47const LINE_RANGE = /^(\d+)\s*[-–]\s*(\d+)$/u;
48
49/**
50 * The block out of the summary, and the summary without it.
51 *
52 * A summary with no block, or an empty one, asks for nothing; the caller then
53 * falls back to recency and the row says so.
54 */
55export const parseRestoreRequests = (text) => {
56 const open = text.lastIndexOf(RESTORE_OPEN);
57 const close = open === -1 ? -1 : text.indexOf(RESTORE_CLOSE, open);
58
59 if (open === -1 || close === -1) {
60 return { summary: text, requests: [] };
61 }
62
63 const body = text.slice(open + RESTORE_OPEN.length, close);
64 const summary = `${text.slice(0, open)}${text.slice(close + RESTORE_CLOSE.length)}`.trim();
65
66 return { summary, requests: body.split("\n").map(parseRequest).filter((request) => request !== null) };
67};
68
69const parseRequest = (raw) => {
70 const line = raw.trim();
71
72 if (line === "" || line.startsWith("#") || line.startsWith("-")) {
73 return null;
74 }
75
76 const match = REQUEST_LINE.exec(line);
77
78 if (match === null) {
79 return null;
80 }
81
82 const range = LINE_RANGE.exec(match[2] ?? "");
83 const from = range === null ? null : Number(range[1]);
84 const to = range === null ? null : Number(range[2]);
85
86 if (from !== null && (from < 1 || to < from)) {
87 return { path: match[1].trim(), from: null, to: null, reason: (match[3] ?? "").trim() };
88 }
89
90 return { path: match[1].trim(), from, to, reason: (match[3] ?? "").trim() };
91};
92
93/**
94 * Every file the conversation touched, most recently touched first, each with
95 * the text of its last successful Read if the transcript still holds one.
96 */
97export const restoreCandidates = (messages) => {
98 const seen = new Map();
99
100 messages.forEach((message, index) => {
101 for (const use of message.toolUses ?? []) {
102 const tool = nameOf(use);
103
104 if (!RESTORE_TOOLS.has(tool)) {
105 continue;
106 }
107
108 const path = pathOf(use);
109
110 if (path === null) {
111 continue;
112 }
113
114 const prior = seen.get(path);
115 const stored = tool === "Read" && use.isError !== true && typeof use.text === "string" && use.text !== ""
116 ? use.text
117 : (prior?.stored ?? null);
118
119 seen.set(path, { path, at: index, stored });
120 }
121 });
122
123 return [...seen.values()].sort((left, right) => right.at - left.at);
124};
125
126/**
127 * Every path-shaped word a shell command named, newest command first.
128 *
129 * A file read with `cat` or rewritten by a script was touched as surely as one
130 * that went through Read or Edit, so the model may ask for it back. These are
131 * words, not resolved paths: a relative one only ever matches the tail of an
132 * absolute path the model names, and none of them feeds the recency fallback,
133 * which would otherwise restore whatever log a command last redirected into.
134 */
135export const shellMentions = (messages) => {
136 const mentions = [];
137
138 messages.forEach((message, index) => {
139 for (const use of message.toolUses ?? []) {
140 const command = use.input?.command;
141
142 if (nameOf(use) !== "Bash" || typeof command !== "string") {
143 continue;
144 }
145
146 for (const word of command.split(SHELL_WORD_BREAK)) {
147 const token = word.slice(word.lastIndexOf("=") + 1).replace(/^\.\//u, "");
148
149 if (PATH_SHAPED.test(token) && !SHELL_EXPANDED.test(token)) {
150 mentions.push({ token, at: index });
151 }
152 }
153 }
154 });
155
156 return mentions.sort((left, right) => right.at - left.at);
157};
158
159const SHELL_WORD_BREAK = /[\s"'`;|&()<>,]+/u;
160
161/** The file a shell command named, if any word it used ends the requested path. */
162const shellCandidateFor = (path, mentions) => {
163 const wanted = path.replace(/^\.\//u, "");
164
165 for (const { token, at } of mentions) {
166 if (token.startsWith("/") && (token === wanted || token.endsWith(`/${wanted}`))) {
167 return { path: token, at, stored: null };
168 }
169
170 if (wanted.startsWith("/") && wanted.endsWith(`/${token}`)) {
171 return { path: wanted, at, stored: null };
172 }
173 }
174
175 return null;
176};
177
178const pathOf = (use) => {
179 const input = use.input ?? {};
180 const path = input.file_path ?? input.notebook_path;
181
182 return typeof path === "string" && path.trim() !== "" ? path.trim() : null;
183};
184
185/**
186 * Which files go back, and why each one the model asked for did not.
187 *
188 * The model chooses; the code only refuses. A path the session never touched,
189 * through a file tool or by a shell command naming it, is refused because the
190 * model cannot have read it, a memory file because the
191 * engine re-emits those itself, and anything past the cap because the cap is
192 * the operator's. A file that is gone from disk is refused too: a worktree
193 * removed after its files were edited once left the fallback restoring nothing.
194 * When nothing survives, the most recently touched files that still exist go
195 * instead, and `source` says which of the two happened so the rate can be
196 * measured. `exists` is the host's check, injected so this stays pure.
197 *
198 * @param {object} options
199 * @param {Array<{ path: string, from: number | null, to: number | null, reason: string }>} options.requests
200 * @param {Array<{ path: string, at: number, stored: string | null }>} options.candidates
201 * @param {Array<{ token: string, at: number }>} [options.shellMentions] what shell commands named; see `shellMentions`
202 * @param {number} [options.maxFiles]
203 * @param {(path: string) => Promise<boolean>} [options.exists]
204 */
205export const chooseRestores = async ({ requests, candidates, shellMentions: mentions = [], maxFiles = RESTORE_MAX_FILES, exists = async () => true }) => {
206 const files = [];
207 const rejected = [];
208
209 for (const request of requests) {
210 const candidate = candidateFor(request.path, candidates) ?? shellCandidateFor(request.path, mentions);
211 const why = refusal(request.path, candidate, files, maxFiles) ?? (await goneFromDisk(candidate, exists));
212
213 if (why !== null) {
214 rejected.push({ path: request.path, why });
215 continue;
216 }
217
218 files.push({ ...request, path: candidate.path, stored: candidate.stored });
219 }
220
221 if (files.length > 0) {
222 return { files, source: "model", rejected };
223 }
224
225 const fallback = [];
226
227 for (const candidate of candidates) {
228 if (fallback.length >= maxFiles) {
229 break;
230 }
231
232 if (MEMORY_FILE.test(candidate.path) || !(await exists(candidate.path))) {
233 continue;
234 }
235
236 fallback.push({
237 path: candidate.path,
238 from: null,
239 to: null,
240 reason: "most recently touched; the summary named no usable file",
241 stored: candidate.stored,
242 });
243 }
244
245 return { files: fallback, source: fallback.length === 0 ? "none" : "recency", rejected };
246};
247
248/**
249 * @param {{ path: string }} candidate
250 * @param {(path: string) => Promise<boolean>} exists
251 */
252const goneFromDisk = async (candidate, exists) => ((await exists(candidate.path)) ? null : "no longer exists on disk");
253
254const candidateFor = (path, candidates) =>
255 candidates.find((candidate) => candidate.path === path) ??
256 candidates.find((candidate) => candidate.path.endsWith(`/${path.replace(/^\.\//u, "")}`)) ??
257 null;
258
259const refusal = (path, candidate, chosen, maxFiles) => {
260 if (candidate === null) {
261 return "not read or written in this conversation";
262 }
263
264 if (MEMORY_FILE.test(candidate.path)) {
265 return "memory file; the engine re-emits it on the next turn";
266 }
267
268 if (chosen.some((file) => file.path === candidate.path)) {
269 return "named twice";
270 }
271
272 if (chosen.length >= maxFiles) {
273 return `over the cap of ${maxFiles} files`;
274 }
275
276 return null;
277};
278
279/**
280 * The two messages the model sees for one restored file: a Read it made, and
281 * what the file said. Real tool blocks, not a narration of them, so the next
282 * window treats the content exactly as it treats any file it has read.
283 */
284export const restorePair = ({ id, path, from, to, text }) => {
285 const input = { file_path: path };
286
287 if (from !== null && to !== null) {
288 input.offset = from;
289 input.limit = to - from + 1;
290 }
291
292 return [
293 { role: "assistant", text: "", toolUses: [{ tool_use_id: id, tool: "Read", input }] },
294 { role: "user", text: "", toolUses: [], toolResults: [{ tool_use_id: id, text }] },
295 ];
296};
297
298/** What replaces the tail of a file too large to carry whole. */
299export const clipRestore = (text, maxChars) => {
300 if (text.length <= maxChars) {
301 return { text, clippedChars: 0 };
302 }
303
304 const kept = text.slice(0, maxChars);
305
306 return {
307 text: `${kept}\n[compact-handoff clipped ${text.length - maxChars} characters here; Read the file with an offset for the rest.]`,
308 clippedChars: text.length - maxChars,
309 };
310};
311
312/**
313 * Fits the read files under the total budget, in the order the model gave
314 * them, so the first file named is the last one dropped.
315 */
316export const fitRestores = (files, totalChars) => {
317 let used = 0;
318
319 return files.map((file) => {
320 if (file.text === null) {
321 return file;
322 }
323
324 if (used + file.text.length > totalChars) {
325 return { ...file, text: null, source: "dropped", detail: `over the total budget of ${totalChars} characters` };
326 }
327
328 used += file.text.length;
329
330 return file;
331 });
332};
333
334/** The metrics row for one file, whatever happened to it. */
335export const restoreRow = ({ path, from, to, reason, source, text, clippedChars, detail, ms }) => ({
336 path,
337 lines: from === null ? "all" : `${from}-${to}`,
338 reason,
339 source,
340 chars: text === null ? 0 : text.length,
341 approxTokens: text === null ? 0 : Math.ceil(text.length / 4),
342 clippedChars: clippedChars ?? 0,
343 ...(detail === undefined || detail === "" ? {} : { detail }),
344 ms,
345});
346hooks/lib.js 1755 lines1/**
2 * compact-handoff: everything a handoff is built out of that needs no `$`.
3 *
4 * Split out of `module.js` so it can be tested. The runtime loads a sibling
5 * import: `import { x } from "./lib.js"` inside a plugin's `hooks/module.js`
6 * resolves and runs in a real session, verified 2026-09-14 against build
7 * 2.1.270 with a probe plugin that wrote the imported value to disk. The static
8 * scan `claude plugin validate` runs accepts it too, which is the cheaper half
9 * of that check and on its own would have proved nothing.
10 *
11 * Everything here is a pure function of its arguments. Nothing reads the clock,
12 * the filesystem, the environment or the session. That is the whole point: the
13 * parts of a compaction that can be wrong in a way a test can catch live here,
14 * and the parts that can only be wrong against a live engine live next door.
15 */
16
17/** Longest tool argument kept in the ledger; enough for a real command line. */
18export const LEDGER_ARG_CHARS = 300;
19
20/** Longest shell command the trimmed ledger shows; the full one keeps LEDGER_ARG_CHARS. */
21export const LEDGER_SHELL_CHARS = 150;
22
23/** The newest shell commands the trimmed ledger always shows, read-only or not. */
24export const LEDGER_RECENT_COMMANDS = 10;
25
26/** The newest failures the trimmed ledger shows. */
27export const LEDGER_FAILURES = 20;
28
29/** Budget for the trimmed ledger's body; older commands that changed something drop first. */
30export const LEDGER_CHARS = 6000;
31
32/** Longest error line quoted back. */
33export const LEDGER_ERROR_CHARS = 200;
34
35/** How much of a tool result is scanned for an error shape. */
36export const ERROR_SCAN_CHARS = 4000;
37
38/** Longest assistant turn fed to the commitments pass. */
39export const COMMITMENT_TURN_CHARS = 6000;
40
41/** Ceiling on the whole commitments prompt; oldest turns drop first. */
42export const COMMITMENT_PROMPT_CHARS = 400_000;
43
44/**
45 * What an error looks like in a tool result that carries no error flag.
46 *
47 * A missing `isError` does not mean a command succeeded. The one genuinely
48 * failed command in the Round 3 fixture wrote `ugrep: warning: ...: No such
49 * file or directory` to stdout, with empty stderr and no flag anywhere, and
50 * every arm that trusted the flag reported it as a success.
51 */
52export const ERROR_SHAPE =
53 /^.*\b(?:no such file or directory|command not found|permission denied|fatal:|error:|warning:|traceback \(most recent call last\)|cannot access)\b.*$/imu;
54
55/** The argument worth naming for a tool call, in the order worth trying. */
56export const LEDGER_KEYS = ["command", "file_path", "path", "pattern", "url", "description"];
57
58/** One parsed row of the commitments pass. */
59export const COMMITMENT_ROW = /^\s*(UNKEPT|CORRECTED|UNANSWERED)\s*\|\s*T(\d+)\s*\|\s*(.+?)\s*\|\s*(.+?)\s*$/iu;
60
61export const COMMITMENT_HEADING = {
62 unkept: "Said it would, no evidence it did",
63 corrected: "Claims corrected or withdrawn",
64 unanswered: "Questions put to the user and never answered",
65};
66
67/* ------------------------------------------------------------------ *
68 * Which user turns a person actually typed.
69 * ------------------------------------------------------------------ */
70
71/**
72 * The user-role turns nobody typed, by the shape they actually arrive in.
73 *
74 * These are counted, not guessed. A sweep over every session JSONL on this box
75 * (2026-09-14) classified the opening of every user turn that carries no tool
76 * result: 2199 continuation prompts, 1993 task notifications, 613 slash-command
77 * envelopes, 410 local command outputs, 31 cross-session messages and 14 idle
78 * notices. A filter written from memory would have missed the two that matter
79 * most here, because a cross-session message never starts with its own tag: the
80 * harness puts "Another Claude session sent a message:" in front of it.
81 *
82 * The /loop wakeup is deliberately absent. A wakeup re-fires the user's own
83 * `/loop` prompt verbatim, so the turn it produces IS the person's words and
84 * pinning it is right.
85 */
86export const NON_PERSON = [
87 { kind: "cross-session", test: (text) => text.includes("<cross-session-message") },
88 { kind: "task-notification", test: (text) => text.startsWith("<task-notification") },
89 { kind: "idle-notice", test: (text) => text.startsWith("[Cross-session idle notice]") },
90 { kind: "command-output", test: (text) => text.startsWith("<local-command-stdout>") },
91 { kind: "system-reminder", test: (text) => text.startsWith("<system-reminder>") },
92 {
93 kind: "continuation",
94 test: (text) => text.startsWith("This session is being continued from a previous conversation"),
95 },
96 { kind: "interrupt", test: (text) => text.startsWith("[Request interrupted by user") },
97];
98
99/**
100 * Who produced a user turn: the person at the keyboard, or the harness.
101 *
102 * A slash-command envelope (`<command-name>/clear</command-name>`) counts as the
103 * person. They typed it, and `/loop keep going` carries the whole instruction in
104 * its arguments; dropping those would lose real intent to save a line.
105 */
106export const originOf = (message) => {
107 const text = (message.text ?? "").trim();
108
109 for (const { kind, test } of NON_PERSON) {
110 if (test(text)) {
111 return kind;
112 }
113 }
114
115 return "person";
116};
117
118/**
119 * Whether a message is a user turn worth keeping verbatim: a person said it,
120 * the engine will vouch for it, and it is not the transcript's way of carrying
121 * a tool result back to the model.
122 */
123export const isPinnable = (message) =>
124 message.role === "user" &&
125 typeof message.handle === "string" &&
126 (message.text ?? "").trim() !== "" &&
127 (message.toolResults === undefined || message.toolResults.length === 0) &&
128 originOf(message) === "person";
129
130/**
131 * The user turns the engine would vouch for that a person did not type, by kind.
132 *
133 * Recorded per compaction so a filter that starts dropping real turns shows up
134 * as a number moving rather than as a session quietly losing its instructions.
135 */
136export const pinSkipCounts = (messages) => {
137 const counts = {};
138
139 for (const message of messages) {
140 if (message.role !== "user" || typeof message.handle !== "string") {
141 continue;
142 }
143
144 if ((message.text ?? "").trim() === "") {
145 continue;
146 }
147
148 if (message.toolResults !== undefined && message.toolResults.length > 0) {
149 continue;
150 }
151
152 const kind = originOf(message);
153
154 if (kind !== "person") {
155 counts[kind] = (counts[kind] ?? 0) + 1;
156 }
157 }
158
159 return counts;
160};
161
162/**
163 * The conversation the session carries on with: the handoff first, then the
164 * words of every user turn the engine would vouch for, without the handle.
165 *
166 * The handle selects the turn; it is not handed back. A message returned with
167 * it stands as the engine has it, and the engine's copy carries every
168 * attachment the turn arrived with: the instruction bundle (CLAUDE.md, every
169 * rule file, AGENTS.md, MEMORY.md), the hook outputs, the skill and agent
170 * listings. Measured on 2026-09-17 (session 7c6495a3, depth 8): a 22.7k-char
171 * handoff came back as a 124k-token first turn, of which 218k chars were four
172 * copies of the instruction bundle riding on 19 pinned turns. The engine
173 * re-emits that bundle on its own once a compaction has run
174 * (`InstructionsLoaded` has a `compact` load reason), so those copies said
175 * nothing twice. A turn rebuilt from `role` and `text` is the person's words
176 * and nothing else.
177 */
178export const replacementFor = (messages, handoff) => [
179 { role: "assistant", text: handoff, toolUses: [] },
180 ...messages.filter(isPinnable).map(wordsOnly),
181];
182
183/** A user turn as the person typed it: the engine's attachments stay behind. */
184const wordsOnly = (message) => ({ role: message.role, text: message.text, toolUses: message.toolUses ?? [] });
185
186/* ------------------------------------------------------------------ *
187 * The size guard.
188 * ------------------------------------------------------------------ */
189
190/**
191 * Keep the replacement small enough that the engine does not compact it again.
192 *
193 * A handoff plus every pinned user turn is not bounded by anything. One pasted
194 * 30k-character blob is a normal thing for a person to do and several of them
195 * is a normal week, so the replacement can plausibly come back larger than the
196 * conversation it replaced, and the engine's answer to that is to compact
197 * immediately: a second summary, of a summary, and the real user turns gone.
198 *
199 * The largest pinned turns are replaced with a pointer at the stored copy until
200 * the whole thing fits. **The summary is never trimmed** — it is the part that
201 * scored 67.3% on its own, and a session that has lost it has lost everything
202 * the compaction was for. If the summary alone is over the ceiling the guard
203 * gives up and says so rather than cutting into it.
204 */
205export const applySizeGuard = (messages, { maxChars, pointerFor }) => {
206 const charsOf = (list) => list.reduce((total, message) => total + (message.text ?? "").length, 0);
207 const charsBefore = charsOf(messages);
208
209 if (charsBefore <= maxChars) {
210 return { messages, trimmedTurns: 0, charsBefore, charsAfter: charsBefore, overCeiling: false };
211 }
212
213 // Index 0 is the summary and is not a candidate at any size.
214 const order = messages
215 .map((message, index) => ({ index, length: (message.text ?? "").length }))
216 .slice(1)
217 .sort((left, right) => right.length - left.length);
218
219 const out = [...messages];
220 let trimmedTurns = 0;
221 let charsAfter = charsBefore;
222
223 for (const { index, length } of order) {
224 if (charsAfter <= maxChars) {
225 break;
226 }
227
228 const pointer = pointerFor({ index, chars: length });
229
230 // A pointer longer than what it replaces would make things worse.
231 if (pointer.length >= length) {
232 continue;
233 }
234
235 out[index] = { ...out[index], text: pointer, handle: undefined };
236 charsAfter = charsAfter - length + pointer.length;
237 trimmedTurns += 1;
238 }
239
240 return { messages: out, trimmedTurns, charsBefore, charsAfter, overCeiling: charsAfter > maxChars };
241};
242
243/* ------------------------------------------------------------------ *
244 * Lineage: how a second compaction knows about the first.
245 * ------------------------------------------------------------------ */
246
247/**
248 * The machine-readable first line every handoff carries.
249 *
250 * Without it the second compaction sees the first one's handoff as ordinary
251 * assistant prose and re-summarises it, which is how a tool ledger erodes into
252 * "the session ran some commands" over three compactions. With it the ledger
253 * rows are merged from the stored JSON instead of being re-read from English.
254 */
255export const lineageLine = ({ session, n, prev }) =>
256 `<!-- compact-handoff: session=${session} n=${n} prev=${prev ?? "none"} -->`;
257
258export const LINEAGE_SHAPE = /<!--\s*compact-handoff:\s*session=(\S+)\s+n=(\d+)\s+prev=(\d+|none)\s*-->/u;
259
260/** The lineage a piece of text carries, or null if it carries none. */
261export const lineageOf = (text) => {
262 const match = LINEAGE_SHAPE.exec(text ?? "");
263
264 if (match === null) {
265 return null;
266 }
267
268 const [, session, n, prev] = match;
269
270 return { session, n: Number(n), prev: prev === "none" ? null : Number(prev) };
271};
272
273/**
274 * The lineage of a message that IS a handoff, or null.
275 *
276 * A handoff opens with its marker. The same marker also turns up inside a
277 * tool result (`handoff_lookup`, a Read of a stored `.md`, a grep, the
278 * README's own example), and counting one of those as a handoff put a
279 * session's depth back to whatever the quoted handoff said and ended the
280 * post-compaction watch the moment a session read its history.
281 */
282export const handoffLineageOf = (message) => {
283 const text = (message.text ?? "").trimStart();
284
285 if ((message.toolResults ?? []).length > 0 || LINEAGE_SHAPE.exec(text)?.index !== 0) {
286 return null;
287 }
288
289 return lineageOf(text);
290};
291
292/** The most recent handoff already sitting in the conversation, if any. */
293export const priorHandoffIn = (messages) => {
294 for (let index = messages.length - 1; index >= 0; index -= 1) {
295 const lineage = handoffLineageOf(messages[index]);
296
297 if (lineage !== null) {
298 return { index, lineage };
299 }
300 }
301
302 return null;
303};
304
305/**
306 * Every place a compaction started a new conversation, oldest first: a handoff
307 * this plugin wrote (`kind: "handoff"`, `n` from its lineage) or the engine's
308 * own continuation summary (`kind: "stock"`, `n` its position among openers).
309 *
310 * `$.session.messages()` answers the whole session, every window since the
311 * first, so this is how a reading finds the window it is actually in.
312 */
313export const openersIn = (messages) => {
314 const openers = [];
315
316 messages.forEach((message, index) => {
317 const lineage = handoffLineageOf(message);
318 const chars = (message.text ?? "").length;
319
320 if (lineage !== null) {
321 openers.push({ kind: "handoff", index, n: lineage.n, chars });
322 } else if (
323 message.role === "user" &&
324 (message.toolResults ?? []).length === 0 &&
325 originOf(message) === "continuation"
326 ) {
327 openers.push({ kind: "stock", index, n: openers.length + 1, chars });
328 }
329 });
330
331 return openers;
332};
333
334/* ------------------------------------------------------------------ *
335 * The tool ledger.
336 * ------------------------------------------------------------------ */
337
338/**
339 * Every tool call the session made, what it was aimed at, and what the
340 * transcript says came back. Never an exit code: there isn't one.
341 */
342export const ledgerRows = (messages) => {
343 const rows = [];
344
345 for (const message of messages) {
346 for (const use of message.toolUses ?? []) {
347 const { outcome, detail } = outcomeOf(use);
348
349 const name = nameOf(use);
350 const command = name === "Bash" && typeof use.input?.command === "string" ? use.input.command : "";
351
352 rows.push({
353 name,
354 target: targetOf(use),
355 outcome,
356 detail,
357 writes: shellWrites(command),
358 readOnly: command !== "" && isReadOnlyCommand(command),
359 });
360 }
361 }
362
363 return rows;
364};
365
366/**
367 * The ledger as the next session reads it: this compaction's rows and nothing
368 * older.
369 *
370 * Until 0.2.0 every earlier compaction's rows were merged in under their own
371 * heading, and by the tenth compaction of one session the ledger was 65k of an
372 * 82k-character handoff (measured 2026-09-15). The earlier rows are still on
373 * disk in each compaction's own JSON; `handoff_lookup` reads them back, and
374 * the note above the summary says so.
375 *
376 * Within one compaction the shell list is trimmed too: a window of 90 commands
377 * rendered 21k characters, most of it read-only probes and the long bodies of
378 * commands that are also in the summary's inventory. The trimmed list keeps the
379 * newest LEDGER_RECENT_COMMANDS whatever they were, then every older command
380 * that changed something while LEDGER_CHARS lasts, and counts the rest. `full`
381 * renders every row, for `handoff_lookup section=ledger`, which reads the rows
382 * back from the compaction's JSON.
383 */
384export const renderLedger = (rows, { full = false, n = null } = {}) => {
385 if (rows.length === 0) {
386 return "";
387 }
388
389 return ["## Tool ledger", "", ...ledgerBody(rows, { full, n })].join("\n").trimEnd();
390};
391
392/**
393 * Four characters per token is the estimate the size guard already uses. It is
394 * not a tokenizer, and the field is named for what it is.
395 */
396export const approxTokens = (text) => Math.ceil(text.length / 4);
397
398/**
399 * How full the window is on one turn, and how much of that is the handoff.
400 *
401 * The question this answers cannot be answered from a compaction row alone: a
402 * compaction row says what the handoff cost to write, not what carrying it
403 * costs to read. Turn 1 of a session with no handoff in it is the floor - the
404 * system prompt, the rules files, the tool declarations and the first message,
405 * everything a window pays before any work happens - and turn 1 of a
406 * post-compact window is that same floor plus the handoff and its restored
407 * files. Logging both, per turn, is what makes the difference measurable
408 * instead of argued. Operator, 2026-09-15: "This tells us how much context is
409 * our compaction summary vs claude rules and similar."
410 *
411 * `percent` is computed here to one decimal rather than taken from the engine,
412 * which reports whole numbers; the engine's own figure is the fallback for a
413 * reading that carries no window to divide by.
414 */
415export const windowReading = ({ at, session, turn, context, messages, opener, arm = null, design = null }) => {
416 const tokens = context?.tokens ?? null;
417 const window = context?.window ?? null;
418 const share = (of) => (of === null || window === null || window === 0 ? null : Math.round((of / window) * 1000) / 10);
419 // The same four-characters-a-token estimate `approxTokens` uses, over a
420 // count rather than the text itself.
421 const handoffTokens = opener === null ? null : Math.ceil(opener.chars / 4);
422
423 return {
424 at,
425 session,
426 turn,
427 first: turn === 1,
428 phase: PHASE_OF[opener?.kind ?? "none"],
429 arm,
430 design,
431 compaction: opener?.n ?? 0,
432 tokens,
433 window,
434 percent: share(tokens) ?? context?.percent ?? null,
435 messages,
436 handoffChars: opener?.chars ?? null,
437 handoffTokens,
438 handoffPercent: share(handoffTokens),
439 };
440};
441
442/** A window's phase by the kind of compaction that opened it, if any. */
443const PHASE_OF = { none: "fresh", handoff: "post-compact", stock: "stock-compact" };
444
445/** A whole scratchpad, the common case: the model closed the tag. */
446const CLOSED_ANALYSIS = /<analysis>([\s\S]*?)<\/analysis>[ \t]*\n?/gu;
447
448/** The tags the summary itself is wrapped in, which carry nothing. */
449const SUMMARY_TAGS = /[ \t]*<\/?summary>[ \t]*\n?/gu;
450
451/**
452 * The summary without the thinking that produced it.
453 *
454 * Through 0.10.0 the fork was asked to write its inventory in <analysis> tags
455 * first and then copy it into the summary, and the block was dropped here. The
456 * prompt no longer asks for it, so this strip is a defensive one now: a model
457 * that writes the block anyway still has it removed. When one does, it is
458 * first-person deliberation rather than findings: it plans the summary, and it
459 * argues with itself and corrects mid-paragraph, which the next window reads
460 * as prose. Operator, 2026-09-15: "I think analysis just bloats it without much
461 * value add."
462 *
463 * Two shapes are deliberate. A block the model never closed is left alone
464 * unless a summary follows it, because a reply cut off inside the scratchpad
465 * has nothing else in it and dropping to the end would hand the next window an
466 * empty handoff. And anything outside the tags is kept, wherever it sits: a
467 * preamble before the analysis and a section written after </summary> are both
468 * content, and only the tags themselves are noise.
469 */
470export const withoutScratchpad = (text) => {
471 let analysisChars = 0;
472 let out = text.replace(CLOSED_ANALYSIS, (_whole, body) => {
473 analysisChars += body.length;
474
475 return "";
476 });
477
478 const open = out.indexOf("<analysis>");
479 const summary = out.indexOf("<summary>");
480
481 if (open !== -1 && summary > open) {
482 analysisChars += summary - open;
483 out = `${out.slice(0, open)}${out.slice(summary)}`;
484 }
485
486 const unwrapped = SUMMARY_TAGS.test(out);
487
488 return { text: out.replace(SUMMARY_TAGS, "").trim(), analysisChars, unwrapped };
489};
490
491/**
492 * One line of the lookup log: which tool read what back, for which session, and
493 * what it cost the window. Logged so a session can be charged for its history
494 * reads the way it is charged for its compactions.
495 */
496export const lookupRecord = ({ at, sessionId, tool, args, text }) => ({
497 at,
498 sessionId: sessionId ?? "unknown",
499 tool,
500 args: args ?? {},
501 chars: text.length,
502 lines: text === "" ? 0 : text.split("\n").length,
503 approxTokens: approxTokens(text),
504});
505
506/** The subsections a set of rows renders to: files written, failures, shell commands. */
507const ledgerBody = (rows, { full, n }) => {
508 const written = writersByFile(rows);
509 const shells = rows.filter((row) => row.name === "Bash");
510 const failed = rows.filter((row) => row.outcome !== "ok" && row.outcome !== "no result in transcript");
511 const out = [
512 `${count(rows.length, "tool call")} (${count(shells.length, "shell command")}), ${count(written.size, "file")} written.`,
513 "Read off the transcript rather than recalled. **The transcript stores no exit",
514 "code**, so the outcome of a command is what its output supports and no more.",
515 "",
516 ];
517
518 if (written.size > 0) {
519 out.push(`### Files written (${written.size})`, "");
520 out.push(...[...written].map(([path, tools]) => `- ${[...tools].join(", ")} \`${path}\``));
521 out.push("");
522 }
523
524 out.push(...failureLines(failed, full));
525
526 const fixed = out.join("\n").length;
527
528 if (shells.length > 0) {
529 const budget = Math.max(0, LEDGER_CHARS - fixed - SHELL_LIST_FRAME);
530
531 out.push(...(full ? everyShellLine(shells) : trimmedShellLines(shells, budget, n)));
532 }
533
534 return out;
535};
536
537const WRITE_TOOLS = new Set(["Write", "Edit", "NotebookEdit"]);
538
539/**
540 * Every file written, once, in the order first written, with each tool that
541 * wrote it. A file is counted once however many calls touched it: a count of
542 * calls beside a list of paths read as two numbers for one thing.
543 */
544const writersByFile = (rows) => {
545 const written = new Map();
546 const note = (path, tool) => written.set(path, (written.get(path) ?? new Set()).add(tool));
547
548 for (const row of rows) {
549 if (WRITE_TOOLS.has(row.name)) {
550 note(row.target, row.name);
551 }
552
553 for (const path of row.writes ?? []) {
554 note(path, "Bash");
555 }
556 }
557
558 return foldRelativePaths(written);
559};
560
561/**
562 * A shell command names a file relative to wherever it ran, and Write and Edit
563 * name the same file absolutely, so one file was listed twice. A relative path
564 * folds into the one absolute path it is the tail of; when none or several
565 * are, it stays as written, because the cwd it ran in is not on the row.
566 */
567const foldRelativePaths = (written) => {
568 const absolute = [...written.keys()].filter((path) => path.startsWith("/") || path.startsWith("~"));
569
570 for (const [path, tools] of written) {
571 const tail = `/${path.replace(/^\.\//u, "")}`;
572 const matches = absolute.filter((candidate) => candidate.endsWith(tail));
573
574 if (!absolute.includes(path) && matches.length === 1) {
575 tools.forEach((tool) => written.get(matches[0]).add(tool));
576 written.delete(path);
577 }
578 }
579
580 return written;
581};
582
583/** Room kept for the ledger's own heading, the list's heading and the count line. */
584const SHELL_LIST_FRAME = 200;
585
586const failureLines = (failed, full) => {
587 if (failed.length === 0) {
588 return [];
589 }
590
591 const shown = full ? failed : failed.slice(-LEDGER_FAILURES);
592 const limit = full ? LEDGER_ARG_CHARS : LEDGER_SHELL_CHARS;
593 const out = [`### Output that reads as a failure (${failed.length})`, ""];
594
595 if (shown.length < failed.length) {
596 out.push(`- ${failed.length - shown.length} older failures omitted`);
597 }
598
599 for (const row of shown) {
600 out.push(`- ${callLabel({ ...row, target: clipTo(row.target, limit) })}`);
601 out.push(` - ${row.outcome}${row.detail === "" ? "" : `: ${row.detail}`}`);
602 }
603
604 return [...out, ""];
605};
606
607const shellLine = (row, limit) => `- \`${clipTo(row.target, limit)}\`${row.outcome === "ok" ? "" : ` — ${row.outcome}`}`;
608
609const everyShellLine = (shells) => [
610 `### Every shell command, in order (${shells.length})`,
611 "",
612 ...shells.map((row) => shellLine(row, LEDGER_ARG_CHARS)),
613 "",
614];
615
616/**
617 * The newest commands always, then older ones that changed something, newest
618 * first, while the budget lasts; shown in the order they ran.
619 */
620const trimmedShellLines = (shells, budget, n) => {
621 const recentFrom = Math.max(0, shells.length - LEDGER_RECENT_COMMANDS);
622 const kept = new Set();
623 let used = 0;
624
625 for (let index = shells.length - 1; index >= 0; index -= 1) {
626 const size = shellLine(shells[index], LEDGER_SHELL_CHARS).length + 1;
627 const isRecent = index >= recentFrom;
628
629 if (isRecent || (!shells[index].readOnly && used + size <= budget)) {
630 kept.add(index);
631 used += size;
632 }
633 }
634
635 const omitted = shells.filter((_, index) => !kept.has(index));
636 const out = [
637 kept.size === shells.length
638 ? `### Every shell command, in order (${shells.length})`
639 : `### Shell commands, in order (${kept.size} of ${shells.length})`,
640 "",
641 ...shells.flatMap((row, index) => (kept.has(index) ? [shellLine(row, LEDGER_SHELL_CHARS)] : [])),
642 ];
643
644 if (omitted.length > 0) {
645 const readOnly = omitted.filter((row) => row.readOnly).length;
646 const where = `handoff_lookup ${n === null ? "" : `n=${n} `}section=ledger`;
647
648 out.push(`- ${count(omitted.length, "older command")} omitted (${readOnly} read-only); ${where} lists every one.`);
649 }
650
651 return [...out, ""];
652};
653
654/**
655 * A failed call as the ledger lists it. A shell command is its own label; any
656 * other tool is named, with its target when it has one, so an EnterWorktree
657 * error does not render as an empty code span.
658 */
659const callLabel = (row) => {
660 if (row.target === "") {
661 return row.name;
662 }
663
664 return row.name === "Bash" ? `\`${row.target}\`` : `${row.name} \`${row.target}\``;
665};
666
667/** What the record supports about how a call ended, and nothing beyond it. */
668export const outcomeOf = (use) => {
669 if (use.isError === true) {
670 return { outcome: "errored", detail: clipTo(firstLine(use.text ?? ""), LEDGER_ERROR_CHARS) };
671 }
672
673 if (typeof use.text !== "string" || use.text === "") {
674 return { outcome: "no result in transcript", detail: "" };
675 }
676
677 const shaped = ERROR_SHAPE.exec(use.text.slice(0, ERROR_SCAN_CHARS));
678
679 if (shaped === null) {
680 return { outcome: "ok", detail: "" };
681 }
682
683 return { outcome: "unflagged, output reads as an error", detail: clipTo(shaped[0].trim(), LEDGER_ERROR_CHARS) };
684};
685
686/**
687 * Which tool was called, under whichever key this build spells it.
688 *
689 * `types/claude-code.d.ts` out of build 2.1.269 declares `ToolUseSummary` as
690 * `{id, name, input}` plus `{result, text, isError}`. Build 2.1.270 hands a hook
691 * `{tool_use_id, tool, input, result, text}` instead, so reading `use.name`
692 * alone yields undefined for every call. That is not hypothetical: the first
693 * live compaction rendered "2 tool calls: 0 file writes, 0 shell commands" over
694 * two Bash calls, and dropped the one that succeeded, because every row fell
695 * through to the placeholder. Read both, and let the declared name win if a
696 * later build restores it.
697 */
698export const nameOf = (use) => use.name ?? use.tool ?? "?";
699
700/** The tools that run a subagent, under either name the engine has used. */
701const SUBAGENT_TOOLS = new Set(["Agent", "Task"]);
702
703/**
704 * Whether a subagent ran since the last turn a person typed.
705 *
706 * Logged next to the fork's context because a subagent's own transcript is
707 * something the fork may or may not see, and a compaction landing right after
708 * one is the case worth being able to tell apart later.
709 */
710export const subagentRanThisTurn = (messages) => {
711 let start = 0;
712
713 for (let index = messages.length - 1; index >= 0; index -= 1) {
714 if (isPinnable(messages[index])) {
715 start = index;
716 break;
717 }
718 }
719
720 return messages
721 .slice(start)
722 .some((message) => (message.toolUses ?? []).some((use) => SUBAGENT_TOOLS.has(nameOf(use))));
723};
724
725/** What a call was aimed at, from whichever of its arguments names a target. */
726export const targetOf = (use) => {
727 const input = use.input ?? {};
728
729 for (const key of LEDGER_KEYS) {
730 const value = input[key];
731
732 if (typeof value === "string" && value.trim() !== "") {
733 return clipTo(value.replace(/\n/gu, " ; ").trim(), LEDGER_ARG_CHARS);
734 }
735 }
736
737 return "";
738};
739
740/* ------------------------------------------------------------------ *
741 * What a shell command did, read off its text.
742 * ------------------------------------------------------------------ */
743
744/** Has a directory separator, or ends in an extension. */
745export const PATH_SHAPED = /\/|\.[A-Za-z0-9]+$/u;
746
747/** Anything the shell would have rewritten before the file was opened. */
748export const SHELL_EXPANDED = /[$*?{}~]/u;
749
750/** Where one command in a compound ends: a newline, `&&`, `||`, `;` or a pipe. */
751const SEGMENT_BREAK = /\n|&&|\|\||;|\|/u;
752
753/** Prefixes that run the command after them unchanged. */
754const RUNS_THE_REST = /^(?:\w+=\S*\s+|sudo\s+(?:-\S+\s+)*(?:-u\s+\S+\s+)?|timeout\s+\S+\s+|time\s+|then\s+|do\s+|else\s+)+/u;
755
756/** Verbs that only look. Anything not here counts as having changed something. */
757const READ_ONLY =
758 /^(?:cd|ls|cat|head|tail|wc|grep|rg|find|echo|printf|pwd|stat|readlink|realpath|file|which|type|date|env|ps|pgrep|du|df|tree|sort|uniq|cut|awk|diff|jq|test|true|\[|sed|python3?\s+-c|git\s+(?:-C\s+\S+\s+)?(?:status|log|diff|show|rev-parse|fetch|ls-files|worktree\s+list|remote(?:\s+-v)?|branch\s+--show-current|config\s+--get)|gh\s+(?:pr|issue|run)\s+(?:view|list|checks|diff)|gh\s+api(?!.*-X\s*(?:POST|PATCH|PUT|DELETE)))(?:\s|$)/u;
759
760/** A redirect into a file: `>`, `>>`, `2>`, never `2>&1` or a `=>` arrow. */
761const REDIRECT = /(?:^|[^<>&=\-\d])\d?>>?[ \t]*([^\s;&|<>()'"`]+)/gu;
762
763/** `open("x", "w")` in an inline script. */
764const OPEN_FOR_WRITE = /open\(\s*(['"])([^'"]+)\1\s*,\s*(['"])[wax]/gu;
765
766/** A redirect that is only plumbing: the bit bucket, or a gate's own log. */
767const NOT_WORK = /^\/dev\/|\.log$/u;
768
769/** `sed -i`, which writes even when its target is a variable this cannot read. */
770const IN_PLACE = /^sed\s(?:.*\s)?-i/u;
771
772/** One-line quoted strings emptied, so a `|` or `>` inside one is not read as the shell's. */
773const blankQuotes = (command) => command.replace(/'[^'\n]*'|"[^"\n]*"/gu, "''");
774
775/** Every command in a compound, with the prefixes that just run it stripped. */
776const segmentsOf = (command) =>
777 blankQuotes(command)
778 .split(SEGMENT_BREAK)
779 .map((segment) => segment.trim().replace(RUNS_THE_REST, ""))
780 .filter((segment) => segment !== "");
781
782/** The unquoted, path-shaped words of one command after its verb. */
783const pathWords = (segment) =>
784 segment
785 .split(/\s+/u)
786 .slice(1)
787 .filter((word) => word !== "" && !word.startsWith("-") && PATH_SHAPED.test(word) && !SHELL_EXPANDED.test(word));
788
789/**
790 * The files a shell command wrote, as far as its text says: redirects, `tee`,
791 * `sed -i`, the destination of `mv` and `cp`, and `open(..., "w")` in an inline
792 * script. A path built from a variable cannot be read off the text and is left
793 * out, so this can miss a write but should not invent one.
794 */
795export const shellWrites = (command) => {
796 if (command === "") {
797 return [];
798 }
799
800 const found = [];
801
802 for (const match of blankQuotes(command).matchAll(REDIRECT)) {
803 found.push(match[1]);
804 }
805
806 for (const match of command.matchAll(OPEN_FOR_WRITE)) {
807 if (isInterpreted(command, match.index) && isRunAsCode(command, match.index)) {
808 found.push(match[2]);
809 }
810 }
811
812 for (const segment of segmentsOf(command)) {
813 const words = pathWords(segment);
814
815 if (/^tee\s/u.test(segment) || IN_PLACE.test(segment)) {
816 found.push(...words);
817 } else if (/^(?:mv|cp)\s/u.test(segment) && words.length >= 2) {
818 found.push(words[words.length - 1]);
819 }
820 }
821
822 return unique(found.filter((path) => PATH_SHAPED.test(path) && !SHELL_EXPANDED.test(path) && !NOT_WORK.test(path)));
823};
824
825/** A program that runs the script it is handed. */
826const INTERPRETER = /(?:^|[\s;&|({])(?:python[\d.]*|node|bun|deno|ruby|perl)(?:\s|$)/u;
827
828/** A heredoc opener, `<<'EOF'` or `<<-EOF`, capturing its delimiter. */
829const HEREDOC = /<<-?[ \t]*(['"]?)(\w+)\1/gu;
830
831/**
832 * Whether the text at `index` is handed to an interpreter at all. Inside a
833 * heredoc still open at that point, it is when the line that opened the heredoc
834 * runs one; outside, when its own line does. A PR body written with
835 * `cat <<'EOF'` that quotes `open('x','w')` is prose going into a file, and
836 * reading it as a write listed a file nothing wrote.
837 */
838const isInterpreted = (command, index) => {
839 const before = command.slice(0, index);
840 const opener = [...before.matchAll(HEREDOC)]
841 .filter((match) => !new RegExp(`\\n[ \\t]*${match[2]}[ \\t]*(?:\\n|$)`, "u").test(before.slice(match.index)))
842 .at(-1);
843 const upTo = opener === undefined ? before.length : opener.index;
844
845 return INTERPRETER.test(before.slice(before.lastIndexOf("\n", upTo - 1) + 1, upTo));
846};
847
848/** A flag whose argument is a script the interpreter runs: `python3 -c`, `node -e`. */
849const SCRIPT_FLAG = /(?:^|\s)(?:-c|-e|--eval)\s*$/u;
850
851/**
852 * Whether the text at `index` is code that runs rather than a string some code
853 * only holds. On its own line it is code when no quote is left open before it,
854 * or when the one left open is the argument of `-c` or `-e`. A script that
855 * writes test cases holds `open('x','w')` inside a string, and reading that as
856 * a write listed a file nothing wrote.
857 */
858const isRunAsCode = (command, index) => {
859 const line = command.slice(command.lastIndexOf("\n", index - 1) + 1, index);
860 const opened = openQuoteIn(line);
861
862 return opened === -1 || SCRIPT_FLAG.test(line.slice(0, opened));
863};
864
865/** Where the quote still open at the end of `line` begins, or -1 when every quote closed. */
866const openQuoteIn = (line) => {
867 let quote = "";
868 let opened = -1;
869
870 for (let index = 0; index < line.length; index += 1) {
871 const char = line[index];
872
873 if (char === "\\" && quote !== "") {
874 index += 1;
875 } else if (quote === "" && (char === "'" || char === '"')) {
876 quote = char;
877 opened = index;
878 } else if (char === quote) {
879 quote = "";
880 opened = -1;
881 }
882 }
883
884 return opened;
885};
886
887/** Whether every command in a compound only looked, and none of them wrote. */
888export const isReadOnlyCommand = (command) => {
889 const redirects = blankQuotes(command).replace(HARMLESS_REDIRECT, "");
890
891 if (shellWrites(command).length > 0 || />/u.test(redirects)) {
892 return false;
893 }
894
895 return segmentsOf(command).every((segment) => READ_ONLY.test(segment) && !IN_PLACE.test(segment));
896};
897
898/** `2>&1` and `>/dev/null`, which send output nowhere that lasts. */
899const HARMLESS_REDIRECT = /\d?>&\d|\d?>\s*\/dev\/null/gu;
900
901/* ------------------------------------------------------------------ *
902 * The commitments pass: what goes in, and what comes back out.
903 * ------------------------------------------------------------------ */
904
905/**
906 * The assistant's own words, thinking included, newest kept when space is short.
907 *
908 * The last turns before a conversation ends carry the most unkept commitments,
909 * because they had the least time to be acted on, so the prompt is trimmed from
910 * the front and the trim is stated in it rather than hidden.
911 */
912export const assistantTurns = (messages) => {
913 const turns = [];
914
915 messages.forEach((message, index) => {
916 if (message.role !== "assistant" || (message.text ?? "").trim() === "") {
917 return;
918 }
919
920 turns.push(`===== turn ${index} =====\n${clipTo(message.text.trim(), COMMITMENT_TURN_CHARS)}`);
921 });
922
923 let kept = turns;
924 let size = kept.join("\n\n").length;
925
926 while (size > COMMITMENT_PROMPT_CHARS && kept.length > 1) {
927 kept = kept.slice(1);
928 size = kept.join("\n\n").length;
929 }
930
931 if (kept.length < turns.length) {
932 kept = [`===== ${turns.length - kept.length} earlier turns omitted for length =====`, ...kept];
933 }
934
935 return kept.join("\n\n");
936};
937
938/** What was actually run, one line each. No results: the absence is the point. */
939export const toolIndex = (messages) => {
940 const rows = [];
941
942 messages.forEach((message, index) => {
943 for (const use of message.toolUses ?? []) {
944 rows.push(`T${index} | ${nameOf(use)} | ${targetOf(use)}`);
945 }
946 });
947
948 return rows.join("\n");
949};
950
951/** The most a handoff may be, in tokens, whatever the window: past this the next window is mostly handoff. */
952export const SAFE_CAP_TOKENS = 200_000;
953export const DEFAULT_CAP_TOKENS = 150_000;
954export const DEFAULT_FRACTION = 0.25;
955export const DEFAULT_MAX_CHARS = 400_000;
956const CHARS_PER_TOKEN = 4;
957
958const usable = (value) => typeof value === "number" && Number.isFinite(value) && value > 0;
959
960/**
961 * How large a handoff may be, and why: min(fraction × window, capTokens) × 4
962 * characters. A hand-set override wins outright and is never clamped, because
963 * setting it was a deliberate act; the row says by how much it passed the safe
964 * cap instead. The cap setting is a default, so it is clamped.
965 */
966export const handoffCeiling = ({ window, override, fraction, capTokens } = {}) => {
967 const knownWindow = usable(window) ? window : null;
968
969 if (usable(override)) {
970 const over = override - SAFE_CAP_TOKENS * CHARS_PER_TOKEN;
971 const overrun = over > 0
972 ? { overSafeCapChars: over, overSafeCapTokens: Math.ceil(over / CHARS_PER_TOKEN) }
973 : {};
974
975 return { window: knownWindow, fraction: null, capTokens: null, chars: override, decider: "override", ...overrun };
976 }
977
978 const usedFraction = usable(fraction) && fraction <= 1 ? fraction : DEFAULT_FRACTION;
979 const usedCap = Math.min(usable(capTokens) ? capTokens : DEFAULT_CAP_TOKENS, SAFE_CAP_TOKENS);
980 const base = { window: knownWindow, fraction: usedFraction, capTokens: usedCap };
981
982 if (knownWindow === null) {
983 const chars = Math.min(DEFAULT_MAX_CHARS, usedCap * CHARS_PER_TOKEN);
984
985 return { ...base, chars, decider: "default: window unknown" };
986 }
987
988 const fromWindow = Math.floor(knownWindow * usedFraction);
989 const tokens = Math.min(fromWindow, usedCap);
990
991 return { ...base, chars: tokens * CHARS_PER_TOKEN, decider: tokens === fromWindow ? "fraction" : "cap" };
992};
993
994/** A row the pass started and did not finish: it names a kind, and stops before the shape closes. */
995const ROW_OPENING = /^\s*(UNKEPT|CORRECTED|UNANSWERED)\s*\|/iu;
996
997/**
998 * Whether the commitments reply stopped at its output cap, and how many rows
999 * came back. Near the ceiling in length, or a final row cut mid-shape, both
1000 * count; a trailing line of prose does not.
1001 */
1002export const commitmentsOutcome = (reply, maxTokens) => {
1003 const text = reply ?? "";
1004 const rows = commitmentsFrom(text).length;
1005 const lines = text.split("\n").map((line) => line.trim()).filter((line) => line !== "");
1006 const last = lines.at(-1) ?? "";
1007 const nearCeiling = text.length >= maxTokens * CHARS_PER_TOKEN * 0.9;
1008 const truncatedRow = ROW_OPENING.test(last) && !COMMITMENT_ROW.test(last);
1009 const hitCapReason = nearCeiling ? "length" : truncatedRow ? "truncated row" : null;
1010
1011 return { rows, hitCap: hitCapReason !== null, hitCapReason };
1012};
1013
1014/** The rows the pass emitted, in the order it emitted them. */
1015export const commitmentsFrom = (reply) => {
1016 const found = [];
1017
1018 for (const line of (reply ?? "").split("\n")) {
1019 const match = COMMITMENT_ROW.exec(line);
1020
1021 if (match === null) {
1022 continue;
1023 }
1024
1025 const [, kind, turn, quote, note] = match;
1026
1027 found.push({
1028 kind: kind.toLowerCase(),
1029 turn: Number(turn),
1030 quote: quote.trim().replace(/^"|"$/gu, "").trim(),
1031 note: note.trim(),
1032 });
1033 }
1034
1035 return dedupeFindings(found.filter((row) => row.quote.toLowerCase() !== "none"));
1036};
1037
1038/**
1039 * Which kind wins when the pass files one sentence under two. A question the
1040 * user never answered is also, loosely, something the assistant still owes,
1041 * and a live 0.11.0 handoff listed it both ways; the narrower kind says more.
1042 */
1043const KIND_RANK = { corrected: 0, unanswered: 1, unkept: 2 };
1044
1045/** @param {string} quote */
1046const quoteKey = (quote) => quote.toLowerCase().replace(/\s+/gu, " ").replace(/[\s.,;:!?]+$/u, "").trim();
1047
1048/** One row per quoted sentence, in first-seen order, under its most specific kind. */
1049const dedupeFindings = (found) => {
1050 const kept = new Map();
1051
1052 for (const row of found) {
1053 const key = quoteKey(row.quote);
1054 const prior = kept.get(key);
1055
1056 if (prior === undefined || KIND_RANK[row.kind] < KIND_RANK[prior.kind]) {
1057 kept.set(key, row);
1058 }
1059 }
1060
1061 return [...kept.values()];
1062};
1063
1064export const renderCommitments = (findings) => {
1065 if (findings.length === 0) {
1066 return "";
1067 }
1068
1069 const out = [
1070 "## Commitments and open questions",
1071 "",
1072 "Written by one model pass over the assistant's own turns and an index of what",
1073 "it ran. Unlike the two sections above this is a judgement rather than a",
1074 "reading: each item claims something was promised and that no evidence of it",
1075 "appears later, which is an absence and cannot be proved from a conversation",
1076 "that was cut off mid-flight. Check before acting, and do not re-do work.",
1077 "",
1078 ];
1079
1080 for (const kind of ["unkept", "corrected", "unanswered"]) {
1081 const rows = findings.filter((row) => row.kind === kind);
1082
1083 if (rows.length === 0) {
1084 continue;
1085 }
1086
1087 out.push(`### ${COMMITMENT_HEADING[kind]} (${rows.length})`, "");
1088
1089 for (const row of rows) {
1090 out.push(`- "${row.quote}"`);
1091 out.push(` - ${row.note}`);
1092 }
1093
1094 out.push("");
1095 }
1096
1097 return out.join("\n").trimEnd();
1098};
1099
1100/* ------------------------------------------------------------------ *
1101 * What a compaction costs.
1102 * ------------------------------------------------------------------ */
1103
1104/**
1105 * List prices per million tokens, read off
1106 * https://platform.claude.com/docs/en/about-claude/pricing on 2026-09-14.
1107 * `cacheWrite` is the 5-minute tier, which is what a fork and a `claude -p`
1108 * pass actually use.
1109 *
1110 * **Matched most specific first, and a family name is never enough.** Sonnet 5
1111 * is $2/$10 and Sonnet 4.6 is $3/$15; Opus 5 is $5/$25 and the retired Opus 4.1
1112 * is $15/$75. A table keyed on the word "sonnet" would have overstated every
1113 * compaction on this box by 50%, and the first draft of this file did exactly
1114 * that from memory. The numbers below were read off the page, not recalled.
1115 *
1116 * A row whose model matches nothing records `cost: null` with a reason.
1117 * Pricing a token at zero because the table is stale is worse than admitting
1118 * the number is unknown: one is a gap, the other is a wrong number in a
1119 * document the operator is going to compare against their bill.
1120 */
1121/*
1122 * Dollars per million tokens, read on 2026-09-14 from the "Model pricing" table at
1123 * https://platform.claude.com/docs/en/about-claude/pricing
1124 *
1125 * Every row below was read off that table, not recalled: all four numbers for all
1126 * seven rows were checked against it on 2026-09-14. Nothing here is from memory.
1127 * `cacheWrite` is the 5-minute write column, which is what a fork and a
1128 * `claude -p` actually pay; the 1-hour column is not modelled because nothing
1129 * here asks for a 1-hour cache.
1130 *
1131 * The Fable/Mythos split is not cosmetic. Cache hits are 0.1x base input on every
1132 * model EXCEPT Fable 5.1 and Mythos 5.1, which the page prices at 0.025x, so one
1133 * pattern over both generations under-charged a Fable 5 or Mythos 5 cache read by
1134 * four times. The specific row has to come first; `priceRowFor` takes the first
1135 * match.
1136 *
1137 * A wrong number here is silent: it mis-states every cost row and the row still
1138 * looks like a measurement. Re-read the page before changing PRICES_TAKEN, and
1139 * change them together.
1140 *
1141 * Cache reads are priced here and never charged. On a subscription a cache
1142 * read costs nothing, so `priceUsage` keeps the token count and records what
1143 * the reads would have cost at list as `cacheReadWaivedUsd`, and leaves them
1144 * out of `usd`. Operator, 2026-09-15: "Cache reads are FREE for subscriptions."
1145 */
1146export const COST_BASIS = "subscription: cache reads free";
1147
1148export const PRICES_TAKEN = "2026-09-14";
1149
1150export const PRICES_SOURCE = "https://platform.claude.com/docs/en/about-claude/pricing";
1151
1152export const PRICES = [
1153 { match: /fable-5-1|mythos-5-1|fable5-1|mythos5-1/u, name: "Fable/Mythos 5.1", input: 10, output: 50, cacheRead: 0.25, cacheWrite: 12.5 },
1154 { match: /fable|mythos/u, name: "Fable/Mythos 5", input: 10, output: 50, cacheRead: 1, cacheWrite: 12.5 },
1155 { match: /opus-4-1|opus-4(?!\d)/u, name: "Opus 4.1", input: 15, output: 75, cacheRead: 1.5, cacheWrite: 18.75 },
1156 { match: /opus/u, name: "Opus 5 / 4.5-4.8", input: 5, output: 25, cacheRead: 0.5, cacheWrite: 6.25 },
1157 { match: /sonnet-5|sonnet5/u, name: "Sonnet 5", input: 2, output: 10, cacheRead: 0.2, cacheWrite: 2.5 },
1158 { match: /sonnet/u, name: "Sonnet 4.6 and earlier", input: 3, output: 15, cacheRead: 0.3, cacheWrite: 3.75 },
1159 { match: /haiku-3|haiku3/u, name: "Haiku 3.5", input: 0.8, output: 4, cacheRead: 0.08, cacheWrite: 1 },
1160 { match: /haiku/u, name: "Haiku 4.5", input: 1, output: 5, cacheRead: 0.1, cacheWrite: 1.25 },
1161];
1162
1163/** Which price row a model id falls under, or null if the table cannot say. */
1164export const priceRowFor = (model) => {
1165 const id = (model ?? "").toLowerCase();
1166
1167 if (id === "") {
1168 return null;
1169 }
1170
1171 return PRICES.find((row) => row.match.test(id)) ?? null;
1172};
1173
1174/** The four token counts a usage record carries, under either spelling. */
1175export const usageOf = (usage) => {
1176 if (usage === null || usage === undefined || typeof usage !== "object") {
1177 return null;
1178 }
1179
1180 const pick = (...keys) => {
1181 for (const key of keys) {
1182 if (typeof usage[key] === "number") {
1183 return usage[key];
1184 }
1185 }
1186
1187 return 0;
1188 };
1189
1190 return {
1191 input: pick("input_tokens", "inputTokens"),
1192 output: pick("output_tokens", "outputTokens"),
1193 cacheRead: pick("cache_read_input_tokens", "cacheReadInputTokens"),
1194 cacheWrite: pick("cache_creation_input_tokens", "cacheCreationInputTokens"),
1195 };
1196};
1197
1198/**
1199 * What a fork's usage cost, or null with the reason it cannot be said.
1200 *types/compact-handoff.d.ts 57 lines1// The type contract of `$.compactHandoff`, the noun this plugin's hooks module
2// adds at `engine.create`.
3//
4// `plugin.json`'s `types` field points here. `/plugin-types` copies this file
5// to `.claude/types/claude-code-plugins/compact-handoff.d.ts` and indexes it in
6// `claude-code-plugins.d.ts`, so a plugin that subscribes to the seam types
7// against the real shape instead of a README. `claude plugin validate` checks
8// it. Self-contained by rule: no import, export-from, require or reference.
9
10/**
11 * What `beforeCompact` takes: the tool this plugin raises when a compaction is
12 * about to happen, beside its own fork.
13 *
14 * The seam carries strings and nothing else. Every plugin runs in its own
15 * environment and an interface call's arguments cross through `cloneInto`,
16 * which throws `DataCloneError` on a function, so a subscriber names a TOOL it
17 * answers rather than handing over a callback.
18 */
19export interface CompactHandoffSubscribeOptions {
20 /** The full name of the tool to raise, as `$.tool.call` spells it. */
21 tool: string;
22 /** What to call the subscriber in this plugin's own records; the tool name when omitted. */
23 name?: string;
24}
25
26/** What `beforeCompact` resolves: the subscription, and the tool it is keyed by. */
27export interface CompactHandoffSubscribeResult {
28 subscribed: true;
29 tool: string;
30}
31
32/**
33 * The noun `$.compactHandoff`: how another plugin runs its own work beside this
34 * plugin's compaction fork.
35 *
36 * It exists because a plugin keyed after this one in `enabledPlugins` never
37 * sees `session.compact` at all — this plugin answers that event without
38 * calling `next`. There is no unsubscribe, and a subscription lives for the
39 * load; subscribing twice with one tool is once.
40 */
41export interface CompactHandoff {
42 /**
43 * Registers the tool raised when a compaction is about to happen.
44 *
45 * Rejects with a `TypeError` when `tool` is missing or blank.
46 */
47 beforeCompact: (options: CompactHandoffSubscribeOptions) => Promise<CompactHandoffSubscribeResult>;
48 /** This plugin's version, for a subscriber that wants to know what it got. */
49 version: () => Promise<string>;
50}
51
52declare module 'claude-code' {
53 interface EngineInterface {
54 compactHandoff: CompactHandoff;
55 }
56}
57