SLOPSHOPPER

compact-handoff

Replaces the engine's compaction with a subagent's handoff, and measures whether it can be one

newguardtoolmodelprocessagents
v0.13.0NOASSERTIONupdated 2026-10-05AnExiledDev/compact-handoff-plugin
A shopper browsing a rack in a slop shop
README

AI-written. No human has read this. Every requirement below is an agent's inference. session: e0b89429 | 2026-09-14

compact-handoff

When Claude Code runs out of context it writes a summary of the conversation and throws the conversation away. This plugin writes the summary instead, and staples three things to it that no summary can be trusted to remember: what was actually run, what the repository actually looks like right now, and what was promised and never finished.

It also keeps the conversation it replaced, so anything the handoff left out can still be searched for afterwards.

By the same author: changelogs.core-directive.com, a changelog for every Claude Code release, written from what changed in the build.

What a compaction hands up

Six parts. Only the first is a model summarising a conversation.

  1. The summary, written by a fork of the session itself. The fork is asked to enumerate before it narrates: a tagged one-line-per-fact inventory of the whole conversation first, the prose reading second. That is the whole trick. Prose makes the writer choose what is interesting, and the fact nobody finds interesting is the one the next session needed. Through 0.4.x the fork was asked with Claude Code's own nine-section summariser instruction plus a Work ledger section; against the same 82-fact answer key, graded blind, the inventory prompt carries 70.0% where that one carried 63.0%, and it is the only arm whose worst run beat the old prompt's best. Through 0.10.0 the fork wrote that inventory twice, first in <analysis> tags that were dropped before the handoff was assembled, then again as the summary's first part. Now it writes it once, straight into the summary. Over the same answer key that carried 69.4% against 67.7%, which is inside the grader's noise, and it took about a third fewer output tokens (16,777 against 24,756) and 182 s per fork against 257 s. The fork is also told to answer in one reply, call no tools and leave the session's task alone, because it inherits the session's tools and a denied tool call doesn't end it, it turns into another request that gets billed too. One fork that kept working on the session's task billed 144,745 output tokens over 24 minutes. An <analysis> block a model writes anyway is still dropped, along with the <summary> wrapper tags, and analysisChars on the row says how much was dropped. A block the model never closed is left alone unless a summary follows it, because a reply cut off inside the scratchpad has nothing else in it. Since 0.11.4 the fork does not copy the user's messages into the inventory: every typed turn already comes back verbatim (item 5), and quoting each one on its [ask] line put every prompt in the window twice. An [ask] line now names the ask in a few words and gives its status; [constraint] and [rejected] lines still quote, but only the sentence that sets the rule, and an ask that survives only inside an earlier handoff keeps its line as it stands.
  2. The tool ledger, read off the messages, not recalled: files written (through Write and Edit, and through the shell where the command's text says so: redirects, tee, sed -i, the destination of mv and cp, open(..., "w") in an inline script, unless it sits inside a string the script only holds or the heredoc it sits in feeds something other than an interpreter, such as a cat <<'EOF' PR body (0.11.5); each file once, with every tool that wrote it, and counted as files rather than calls; since 0.11.5 a relative path the shell wrote folds into the one absolute path Write or Edit named that it is the tail of, and stays as written when none or several match), the newest 20 calls whose output reads as a failure, and the shell commands in order. Nothing here is an exit code, because the transcript does not store one, so "no error flag" does not mean "succeeded". Only the rows since the last compaction. Until 0.2.0 every earlier compaction's rows were merged in too, and by the tenth compaction of one session that was 65k of an 82k-character handoff, a quarter of a 200k window spent re-reading history every turn. Since 0.11.3 the shell list is trimmed as well: the newest 10 commands always, then older commands that changed something, clipped to 150 characters, while the ledger stays under 6,000 characters; read-only probes (ls, grep, git status, gh pr view and the like) go first. A count line says how many were left out and names the handoff_lookup n=N section=ledger call that lists every one at full length, rebuilt from the rows the compaction's JSON kept. Over the 78 stored compactions with rows, the largest list (87 commands) went from 24,833 characters to 5,894.
  3. Session state, read live from git and gh as the handoff is written: branch, working tree, uncommitted files, the repository's open PRs by number and title (the whole repo's, not only this session's), the last five commits, the session's model, and any agents still running. A submodule holding nothing but untracked files is not listed as a change. It describes the moment of compaction and nothing else. A row that cannot be read is left out rather than guessed at.
  4. Commitments and open questions, from one model pass over the assistant's own turns, on claude-opus-5-5 since 0.11.1 (Sonnet 5 before). The fact it looks for is an absence: "let me check X" is easy to find, but what matters is that nothing after it ever checked X, and no pattern sees that. Two rounds of prompt wording could not move that class of fact; one model call moved it from 35.7% to 73.8%. A step handed to the user is not counted as the assistant's, and one sentence is reported once, under its most specific kind.
  5. Every user turn, verbatim, selected by the engine's own handle and handed back as its words alone, so nothing is rebuilt from a paraphrase. Until 0.7.0 the turn went back with its handle, which hands the engine's own copy up whole, and that copy carries every attachment the turn arrived with: the instruction bundle (CLAUDE.md, every rule file, AGENTS.md, MEMORY.md), the hook outputs, the skill and agent listings. Measured on 2026-09-17 (session 7c6495a3, depth 8): a 22.7k-character handoff came back as a 124k-token first turn, and 218k characters of it were four copies of the instruction bundle riding on 19 pinned turns, one of them a 44k-character AGENTS.md. The engine re-emits that bundle on its own after a compaction, so every copy said the same thing twice. A turn without its handle is the person's words and nothing else; the hook_additional_context lines those turns carried (the intent ledger's op: ids here) go with the attachments. Measured live on 2026-09-17 (engine 2.1.274, the bench's tools check): the transcript after the boundary held the handoff and two pinned turns of 2,155 and 577 bytes with nothing attached to them, the engine re-emitted its instruction bundle once, and the first real turn cost 58k tokens where 0.6.0 had cost 124k. (withHandles is a field of the forced-run record compact_force writes to runs.jsonl, not of the index row.)
  6. The files the next window should have open, chosen by the summariser (0.3.0). Claude Code's own compaction re-attaches up to five of the most recently read files, and that path never runs when a hook answers the event, so until 0.3.0 a handoff came back with none. Now the fork ends its summary with a <restore-files> block naming up to five files, each by absolute path with all or a line range and a reason. The code only refuses: a path the session never read or wrote (since 0.11.2 a file a shell command named counts, so one read with cat or rewritten by a script can come back; the command's word only has to end the requested path, and shell-named files never feed the recency fallback), a CLAUDE.md or AGENTS.md (the engine re-emits those itself), a duplicate, anything past the cap, and since 0.11.1 a file no longer on disk (no longer exists on disk). When nothing usable was named, the most recently touched files that still exist go instead and the row says source: "recency". Each file is re-read through the real Read tool, so a file edited mid-session comes back current; a read that is denied, errors or takes over 20 seconds falls back to the text of the transcript's last Read of it (source: "stored"). Each one is handed up as a real Read tool_use and its tool_result, not a narration of one, so the next window treats it exactly as a file it read. Per file 20,000 characters, 100,000 in all, both clipped rather than dropped past the file cap and dropped past the total, and never past the size guard's ceiling. Every cap is a setting below, and record.restore carries what would be needed to move one: source, requested, restored, every rejection and why, and per file the lines, chars, approximate tokens, clipped chars and milliseconds. handoff_status sums them under restores.

Parts 2 to 4 are gathered concurrently and every one of them may fail. A part that throws, times out or comes back empty is left out and named in the row; the summary alone is still the arm that scored 67.3%.

Measured paired against byte-identical model text, 82 atoms from one real session, three grading passes each, on claude-sonnet-5:

armappendedrecallnetspread
armAnothing67.3%66.1%3.0%
armBtool ledger68.7%65.0%4.9%
armDledger + commitments71.1%67.5%1.2%

It rehearses by default. Without COMPACT_HANDOFF_LIVE it does the whole thing, writes down what it would have handed up, and then calls next(e) anyway, so the engine compacts exactly as it does today. Every failure path does the same. The worst case of installing it is the behaviour you already have, plus a log.

Where the data lives

~/.claude/compact-handoff/, outside any repository, because the plugin's own root is a worktree somebody may delete. COMPACT_HANDOFF_DATA_DIR moves it.

~/.claude/compact-handoff/
  index.jsonl                        every compaction on this box, one line each
  window.jsonl                       one occupancy reading per turn, every session
  sessions/<sessionId>/
    NNN-<iso>.md                     the handoff that was handed up
    NNN-<iso>.summary.md             just the model's part of it
    NNN-<iso>.transcript.md          the conversation as it stood before
    NNN-<iso>.json                   the row, the ledger, feedback, lineage
    NNN-<iso>.post.json              what the session did in its next ten person turns
    runs.jsonl                       this session's compactions, append-only
    lookups.jsonl                    every history read this session made, with its size
    diagnostics.jsonl                probes, forced compactions, A/B arms
  rehearsals/<sessionId>/            the same, for runs that did not go live

Nothing is ever rewritten. Every file is written once and every log is appended to, because two sessions compacting in the same minute must not be able to lose each other's rows. index.jsonl is appended through sh -c 'cat >> ...', because the host filesystem API has no append.

Each handoff's first line is a machine-readable lineage marker:

<!-- compact-handoff: session=<id> n=003 prev=002 -->

A later compaction reads it, records depth, and prepends an "Earlier compactions" note naming every earlier pass and how to read it back. Only a message that opens with the marker and is not a tool result counts as a handoff: since 0.11.5 a marker quoted by handoff_lookup, a Read of a stored handoff or this README's example no longer sets the depth or ends a watch. Only the newest handoff travels in the window. Everything an earlier compaction wrote stays on disk behind handoff_lookup and handoff_search, so history costs context only when a session asks for it, and every such read is logged:

~/.claude/compact-handoff/lookups.jsonl            every history read on this box
~/.claude/compact-handoff/sessions/<id>/lookups.jsonl   this session's reads

One row per handoff_lookup, handoff_search or handoff_list call: at, sessionId, tool, args, chars, lines and approxTokens. The token figure is characters over four, the same estimate the size guard uses, not a tokenizer; the name says so. handoff_status sums them under lookups.

What carrying a handoff costs to read

A compaction row says what a handoff cost to write. window.jsonl says what it costs to carry, and it is the only file here written by sessions that never compact at all:

~/.claude/compact-handoff/window.jsonl

One row per completed turn, whatever the session: at, session, turn, first, phase (fresh, post-compact or stock-compact), compaction, tokens, window, percent, messages, handoffChars, handoffTokens, handoffPercent. Nothing is sampled and nothing is conditional, because the comparison only works if both arms are there.

turn counts this session's turns in its current window: 1 is the first turn after the session started or after a compaction. phase and compaction come from the last opener in the transcript, a handoff this plugin wrote (post-compact, compaction its lineage n) or the engine's own continuation summary (stock-compact, compaction its position, handoffChars the summary's size), because $.session.messages() answers the whole session, every window since the first. The count lives in module memory, keyed by session. Through 0.11.4 it lived in $.store, one file every session on the box shares, so every session's turns ran up one counter and first: true fired once in 1,374 rows; nothing before 0.11.5 is usable for the floor comparison below. A reload the module never saw start (/reload-plugins mid-session) writes turn: null until the next compaction, rather than a number that means something else.

The reading that matters is first: true. On a fresh session that is the floor every window pays before any work happens: system prompt, CLAUDE.md and AGENTS.md, the tool declarations, the first message. On a post-compact session it is that same floor plus the handoff and its restored files. The difference between the two is the handoff's real price, and handoffPercent is the handoff's own share of the window, so the two together separate what this plugin costs from what the rules files cost. Added in 0.4.2 at the operator's direction: "This tells us how much context is our compaction summary vs claude rules and similar."

percent is computed to one decimal from tokens / window; the engine's own whole-number figure is the fallback when a reading carries no window.

It wrote nothing at all until 0.4.3. The reading was taken off the turn.complete event's own messages, and that event carries no transcript: it is answer, durationMs, aborted, turnId and reason, whatever the declarations imply. So every turn on engine 2.1.273 threw a TypeError, the engine printed turn.complete hook skipped: threw and no row was ever appended. The transcript comes from $.session.messages() now, which costs a host round trip per turn, and a turn whose transcript cannot be read still writes its row with messages and the handoff fields null.

What the session did after a compaction

For ten person turns after a compaction, turn.complete rewrites NNN-<iso>.post.json beside the compaction's row. No model call: kind (handoff, or stock for a main-session compaction the engine did itself, a rehearsal, a fallback or an aborted dispatch), anchored, turnsObserved (turns the person typed; tool results and harness messages are not turns), firstUserMessage, turnsToFirstToolCall (0 is a tool call before the person said anything), reRunCommands and reReadFiles (byte-identical commands and whole-file reads the pre-compaction window already did), handoffToolsCalled, compactedAgain and done. A compaction that arrives before any turn of the previous window ends writes that window's file first, with compactedAgain: true, so the quickest re-compactions are counted rather than overwritten. The stock arm is the baseline: the same watch over the engine's own compactions, which bench/summarise_runs.py prints beside the handoff arm. It is not a random sample, since a fallback happens for a reason, so a gap between the arms is a lead and not a result. A subagent's compaction is not watched.

The watch finds where the new conversation starts by the opener the compaction put in, counted in the whole transcript, and stops anchored: false rather than guess when that opener is not there. Every record written before 0.11.5 is unusable and the summariser sets them aside: the watch sliced the whole session from the replacement's length, so a record said 510 to 894 turns, 32 to 63 re-run commands and an empty first message. The monitor lives in module memory like the turn count, so a reload mid-watch ends it, and the file keeps what it last had.

The live A/B split

Set abStockShare in /plugin (or COMPACT_HANDOFF_AB_STOCK_SHARE) to a fraction, with live on, and that share of compactions is left to the engine's own compaction while the rest get the handoff. 0.5 gives the most data per week; 0.2 keeps most on the handoff. Clear it to end the split.

  • Two designs run at once, half the sessions each (since 0.13.0). A hash of the session id, salted so it says nothing about the arm, picks the design:
  • session: the whole session is one arm, from a hash of its id. A handoff always builds on a handoff, so the arm is the only difference between two sessions. This is the only design before 0.13.0, and a row with no design is one.
  • compaction: every compaction flips its own coin, a hash of the session id and the compaction's at. The same session's work is seen under both arms, at the price that a handoff can build on the engine's summary or the reverse; priorArm says which. The coin is never seeded from depth, which restarts when a session is resumed.

Both are hashes, so they survive a reload or restart with nothing stored and can be recomputed from the row. Raising the share only ever moves a session or a compaction from handoff to stock.

  • The stock arm is logged like a handoff. Its row (disposition: abStock, outcome: stock) is filed under sessions/<id>/ with the pre-compaction transcript, the engine's summary as .summary.md, tokensBefore, tokensAfter, stockSummaryChars, the engine's usage and a priced cost. The fork never runs in that arm; the seam is still raised, so a subscriber behaves the same in both.
  • Every record says its arm and its design. Compaction rows carry ab (null when the split is off): design, arm, share, bucket (the session's arm hash), designBucket, coin (the per-compaction hash, null under session), priorArm (handoff, stock or none: what the compaction before this one left in the conversation, read off the transcript, so it survives a resume) and handoffRun (handoffs in a row immediately before this one). A resumed transcript only shows its last opener, so the run is kept in the plugin's store after every compaction; runFrom is store, or transcript when the store had nothing that agreed and the run was counted from the openers still visible, which can undercount. post.json carries arm and design. Every window.jsonl reading carries design, and arm under the session design (before the first compaction as well as after); under compaction a turn has no arm of its own, so its arm is null and the window takes its arm from the compaction that opened it.
  • Read it with python3 bench/summarise_runs.py --ab: per arm, both designs together and then each design alone, sessions, compactions, how many got their assigned arm, elapsed, cost, summary size, context before and at turn 1 after, turns until the next compaction, and the ten-turn watch (turns to first tool, re-ran commands, re-read files, handoff tools used, compacted again). Below it, the compaction design is broken down by priorArm. It groups by assigned arm, so a handoff that fell back still counts against the handoff. --ab-export ab.jsonl writes one joined line per compaction, with the file paths and every ab field, for digging past the table.
  • Windows are joined by time. A window.jsonl reading belongs to the latest compaction of its session at or before it (two in one clock tick are ordered by file stem), and a window is closed only when a later compaction of that session exists. Through 0.12.1 the join was on (session, depth), and depth restarts when a session is resumed in a new process, so a resumed session's windows were merged.

A rehearsal (live off) has no split: every compaction there is already the engine's.

Reading one row

A row says what the compaction did and, when it did not do it, why. The fields that matter for that second case:

fieldwhenwhat it says
dispositionalwaysreplaced, rehearsed, abStock (the A/B split left this compaction to the engine), fellBack, passedThrough, abortedFallback, skipped (the engine tried to compact this plugin's own fork loop and was declined). Historical rows only: overBudget, written by 0.10.0 and earlier when a session had spent past its USD ceiling. That ceiling is gone, and those rows still read back through every tool
enginealwaysthe Claude Code version, from $.session.version() since 0.11.5, with CLAUDE_CODE_VERSION the fallback. Null on nearly every earlier row, because that variable is unset in an ordinary session
fallbackReasonevery disposition but replacedone line naming why, e.g. noHandoff: no handoff has been written yet (fork nothing-to-fork: the fork found no warm main-thread transcript)
forkOutcome, forkDetailevery row whose fork was not handed upwhy the fork's own answer was not used, and the fallback went to the handoff on disk. On engine 2.1.280 and later an unanswered fork records the engine's own reason: nothing-to-fork (no warm transcript yet), api-error (the detail carries the HTTP status and error kind), empty-reply or aborted, and the spend of any that made a request stays in usage. timeout is a fork that had not answered after five minutes: the compaction stops waiting and falls back, but the engine gives a plugin no way to cancel a fork, so the request runs on and still spends, and its answer is dropped when it arrives (that spend is on no row). mismatch is an answer refused by the forkInput check below, threw anything unexpected. Rows from 0.10.0 and earlier say cold where they now say nothing-to-fork (or threw, on engine 2.1.280 and later)
costalways{forkUsd, commitmentsUsd, commitmentsBasis, totalUsd, cacheReadWaivedUsd, basis, forkUsage, model, priced, pricesTaken}, or null. Since 0.4.1 cache reads are stored in forkUsage.cacheRead and priced into cacheReadWaivedUsd at list, and never added to forkUsd or totalUsd, because on a subscription they cost nothing; basis says so. Rows from 0.4.0 and earlier charged them. commitmentsBasis is measured: the usage the result reported after 0.10.0, estimate: chars/4, no cache from 0.4.0 (and since, on a result that arrives with no usage), none when no commitments pass ran, measured on rows from 0.3.0 and earlier

| maxHandoff | every replaced and rehearsed row | the handoff ceiling the size guard used and how it was arrived at: {window, fraction, capTokens, chars, decider}, where decider is fraction, cap, default: window unknown or override; an override past the safe cap adds overSafeCapChars and `overSafeC

Source 4 files
hooks/module.js 3155 lines
1/**
2 * compact-handoff: the engine's compaction, replaced by a subagent's handoff.
3 *
4 * A `session.compact` hook is handed the whole transcript and its answer *is*
5 * the conversation afterwards, so returning `{ messages }` without calling
6 * `next` means the engine's summariser never runs. That much is real and this
7 * module does it. What the module cannot do is write the handoff while the
8 * event is waiting, and finding that out is what shaped everything below.
9 *
10 * **A hook dispatch is cut off after about ten seconds.** Measured at 2.1.269:
11 * a hook sitting on a file nobody writes is aborted at 10052 ms, and a real
12 * `/compact` was aborted at 10525 ms. A subagent takes thirty seconds to two
13 * minutes. So the handoff cannot be written inside the compaction, and any
14 * design that waits for one there falls back to the engine every single time.
15 *
16 * So the handoff is written *between* turns and only read at compaction. Every
17 * turn's end promotes a finished handoff into place and, when the conversation
18 * has moved enough, starts the subagent that writes the next one; `$.agent.spawn`
19 * returns in about 400 ms without waiting for it, which is a defect everywhere
20 * else and exactly what is wanted here. `session.compact` then does nothing but
21 * read a file that is already on disk, which costs milliseconds.
22 *
23 * The cost is staleness: the handoff describes the conversation as of the last
24 * refresh, not as of the compaction. `MAX_STALE_MESSAGES` bounds it, and past
25 * that bound the hook hands the compaction back to the engine rather than hand
26 * up a handoff that is missing the last hour.
27 *
28 * It rehearses by default. Unless `COMPACT_HANDOFF_LIVE` is on it does the whole
29 * thing, writes down the conversation it would have handed up, and then calls
30 * `next(e)` anyway, which is the engine compacting exactly as it does today.
31 * Every failure path does the same. The worst case of running this is the
32 * behaviour you already have, plus a log.
33 */
34
35const SCRATCH = ".runs";
36
37/** The handoff a compaction reads. Only ever written whole, by a promotion. */
38const LATEST = `${SCRATCH}/latest.md`;
39
40/**
41 * A subagent's Write lands in its own time, so it writes somewhere else and a
42 * later turn moves the finished text into `LATEST`. Without that, a compaction
43 * landing mid-write reads half a handoff and cannot tell.
44 */
45const SENTINEL = "<!-- handoff-complete -->";
46
47/** Set by the `force_compact` tool, read and cleared at the turn's end. */
48const ARMED_KEY = "armed";
49
50/** The refresh that is in flight: when it started and what it is reading. */
51const PENDING_KEY = "pending";
52
53/** The handoff on disk: when it landed and how long the transcript was then. */
54const READY_KEY = "ready";
55
56/** When this session last forked for a handoff, wall clock ms, for the fork's context row. */
57const LAST_FORK_KEY = "lastForkAt";
58
59/**
60 * What the last A/B compaction of a session left in its conversation and the
61 * handoffs in a row it ended, under `abChain:<session>`. One small key per
62 * session that ever compacted under the split; a resume reads it back, because
63 * a resumed transcript only holds its last opener.
64 */
65const abChainKey = (session) => `abChain:${session}`;
66
67/** The agent type this plugin defines for the handoff writer, `<plugin>:<name>`. */
68const HANDOFF_AGENT = "compact-handoff:handoff";
69
70/** Whether the type took this session, so a spawn knows which one to name. */
71const HANDOFF_AGENT_KEY = "handoffAgent";
72
73/**
74 * The A/B bench. A variant is a file rather than a tool argument so the
75 * instruction never enters the transcript: every arm is then asked over a
76 * byte-identical context, which is the only way two arms are comparable.
77 */
78const AB_PROMPTS = `${SCRATCH}/ab/prompts`;
79
80/** Where each arm's answer lands, one file per replicate, for scoring later. */
81const AB_OUT = `${SCRATCH}/ab/out`;
82
83/**
84 * Three environment variables steer this, and the static scan will only take
85 * them spelled out at the call site, so they are named here and nowhere else:
86 *
87 * - `COMPACT_HANDOFF_LIVE` off, the hook rehearses and lets the engine
88 *   compact; on, it answers the event itself and the summariser never runs.
89 * - `COMPACT_HANDOFF_MODEL` an alias (`haiku`) or a full id for the subagent;
90 *   unset lets the agent's own model, then the parent's, decide.
91 * - `COMPACT_HANDOFF_REFRESH_MS` the shortest gap between two refreshes.
92 */
93const DEFAULT_REFRESH_MS = 10 * 60 * 1000;
94
95/** No handoff at all until the conversation is long enough to need one. */
96const MIN_MESSAGES = 20;
97
98/** And no refresh until it has moved enough to be worth paying to re-read. */
99const MIN_NEW_MESSAGES = 20;
100
101/**
102 * How far behind the handoff may be and still be used. Past this the engine
103 * takes the compaction back, because a handoff this stale would drop the work
104 * the session is in the middle of.
105 */
106const MAX_STALE_MESSAGES = 60;
107
108/** A refresh still unfinished after this is treated as dead and restarted. */
109const ABANDON_AFTER_MS = 15 * 60 * 1000;
110
111/** How often the dispatch-budget probe looks for the file it is waiting on. */
112const POLL_MS = 1000;
113
114/**
115 * `$.fs` rejects a write over 4 MiB, and a prompt pointing at an enormous file
116 * is its own problem, so the transcript is clamped well under that.
117 */
118const MAX_TRANSCRIPT_CHARS = 1_500_000;
119
120/** How much of one tool result is worth keeping for a handoff to read. */
121const MAX_TOOL_TEXT_CHARS = 600;
122
123/** What a fork is asked when the probe names nothing: the handoff itself. */
124/**
125 * The manifest's `userConfig`, as `register` was handed it. Every setting here
126 * was an environment variable first and both spellings still work, the
127 * manifest's winning: a config-menu row is the discoverable half, and the
128 * variable is what a cron line already exports around the session.
129 */
130let options = {};
131
132/**
133 * Per-session watch state, keyed by session id: the compaction being watched,
134 * and each session's turn count for `window.jsonl`. Module memory rather than
135 * `$.store`, because the store is one file every session on the box shares:
136 * a counter kept there counted every session's turns as one, and a monitor
137 * kept there could be read and ended by whichever session ticked first. A hot
138 * reload forgets both; the post file keeps what it had and the next rows say
139 * `turn: null`.
140 */
141let watches = new Map();
142let windows = new Map();
143
144/**
145 * One `userConfig` value as a non-empty string, or `undefined` so that `??`
146 * falls through to the variable. Stringified, because every parser below takes
147 * the text `$.env.get` returns and a declared number or boolean has to read
148 * the same way to them. A number row left at `0` counts as unset, which is why
149 * each numeric default is `0` and the real default lives in the constant.
150 * A boolean row declares no default at all: the engine fills a declared default
151 * in as though the person had set it, and `false` is a setting, so a declared
152 * `false` would keep the variable from ever being read.
153 */
154const opt = (key) => {
155    const value = options[key];
156
157    if (value === undefined || value === null || value === "" || value === 0) return undefined;
158
159    return String(value).trim() || undefined;
160};
161
162/**
163 * The compaction instruction, and the bench's winner over five rounds.
164 *
165 * Rounds 1 to 3 varied the wording of the engine's own summariser prompt, which
166 * asks for nine numbered prose sections. Round 4 asked whether that inherited
167 * shape is the right one at all, and it is not. Prose makes the writer choose
168 * what is interesting, and a fact nobody finds interesting is exactly the fact
169 * the next session needed. So this prompt makes it enumerate before it narrates:
170 * a tagged one-line-per-fact inventory first, the prose reading second.
171 *
172 * Measured over the same 82-atom answer key, graded blind, two forks per arm:
173 * 70.0% carried against 63.0% for the nine-section prompt we shipped before, and
174 * it is the only arm whose worst run beat that prompt's best. Two arms that also
175 * restructured (an answer sheet of lettered sections, and the successor's ten
176 * questions as headings) scored 63.4% and 66.5% with spreads of 18.3 and 15.9,
177 * against this one's 7.3. High variance is disqualifying for a prompt that gets
178 * one attempt per compaction.
179 *
180 * In round 4 it cost about 2,600 more output tokens than the prompt it
181 * replaced, roughly 40% more, against a post-compaction floor measured at
182 * 54,600 to 66,058 tokens in a real session.
183 *
184 * Round 5 took out the <analysis> block that round 4's version wrote the
185 * inventory in first. PART 1 was a copy of it, so every compaction wrote the
186 * inventory twice and paid for both. Written once, straight into the summary,
187 * over the same fixture and answer key, two forks per arm graded three times
188 * each: 69.4% recall (spread 5.5) against 67.7% (spread 6.8) for the two-copy
189 * prompt, inside the grader's 3.5-point noise, for 16,777 output tokens against
190 * 24,756 and 182 s against 257 s per fork. `withoutScratchpad` still strips the
191 * block if a model writes one anyway.
192 *
193 * The <restore-files> block below is appended mechanically and was not part of
194 * the benched arm, the same way the Work ledger was appended to the old one.
195 *
196 * The one-reply, no-tools paragraph is there because a fork inherits the
197 * session's tools. A denied tool call does not end the fork, it becomes another
198 * request, and the engine sums the usage across them; a transcript that stops
199 * mid-task reads to the model like a turn in that task. One fork took the
200 * prompt that way, carried on the session's work and billed 144,745 output
201 * tokens over 24 minutes before it wrote a handoff.
202 */
203const FORK_PROMPT = `Your task is to compact this conversation into a handoff for the next session. It is the only thing that survives.
204
205Answer in a single reply and call no tools; everything you need is already in the conversation. Do not continue, finish or act on the task the conversation was working on, even if it stopped mid-step. Your only job now is to write the handoff.
206
207Summaries lose facts because they are written as prose, and prose makes the writer choose what is interesting. You will not choose. You will enumerate first, then narrate.
208
209In <summary> tags, produce the handoff in two parts:
210
211PART 1 — THE INVENTORY. Sweep the conversation from the first message to the last and emit an inventory: one line per discrete fact, no prose, no commentary. A discrete fact is anything a successor could be wrong about — a request, a constraint, a rejection, a decision, a reason, a file, a command run and its result, a number, an identifier, an error, a promise, an unfinished item, a correction. Aim for completeness over elegance; a hundred lines is normal and a short inventory means you skipped. Mark each line with one tag in brackets at the start: [ask] [constraint] [rejected] [decision] [file] [command] [identifier] [error] [promise] [pending] [state]. Group the lines by tag, in that tag order, keeping conversation order within each group. Do not drop a line because it seems minor, and do not merge two lines into one. Every message the user typed comes back word for word right after this handoff, so do not copy one onto an [ask] line: name the ask in a dozen words or fewer, then its status and any op id. An ask that appears only inside an earlier handoff, never as a message the user typed, does not come back; keep its line as it stands. On a [constraint] or [rejected] line quote the user verbatim, the sentence that sets the rule and not the whole message.
212
213PART 2 — THE READING. Now, and only now, write the prose a successor needs to make sense of Part 1: what the work is, what has been done, what is being done right now, and what the next step is, naming the user's most recent request without quoting it. Keep this short. It explains the inventory; it does not replace it.
214
215Two rules that override any instinct toward brevity. Never write "various", "several", "etc.", "and similar", or any phrase that stands in for items you could have named. Never state a fact the conversation did not establish — if you do not know the branch, the file or the number, write "not established", because a plausible invention is worse to a successor than a gap.
216
217There may be additional summarization instructions in the included context; follow them too.
218
219After the closing </summary> tag, add a <restore-files> block naming up to 5 files the next window should have open before it does anything: the file being edited, the spec or test it is being written against, the file the current step depends on. Only files this conversation actually read or wrote, through any tool including shell commands, by absolute path. One file per line, three fields separated by |: the absolute path, then either all or a line range like 120-260, then a one-line reason. Prefer a line range when only part of a large file matters. Do not name CLAUDE.md or AGENTS.md files; they come back on their own. Leave the block empty if nothing qualifies.
220<restore-files>
221/abs/path/to/file.ts | 40-120 | the function being changed
222</restore-files>`;
223
224import {
225    RESTORE_FILE_CHARS,
226    RESTORE_MAX_FILES,
227    RESTORE_TOTAL_CHARS,
228    chooseRestores,
229    clipRestore,
230    fitRestores,
231    parseRestoreRequests,
232    restoreCandidates,
233    shellMentions,
234    restorePair,
235    restoreRow,
236} from "./restore.js";
237import {
238    abArmFor,
239    abChainAfter,
240    abPriorOf,
241    applySizeGuard,
242    assistantTurns,
243    closedByCompaction,
244    commitmentsFrom,
245    commitmentsOutcome,
246    costOf,
247    costOfNothing,
248    estimatedUsage,
249    forkInputOf,
250    handoffCeiling,
251    fallbackReasonFor,
252    isPinnable,
253    ledgerRows,
254    lineageLine,
255    nameOf,
256    monitorSeed,
257    observePost,
258    openersIn,
259    originOf,
260    pinSkipCounts,
261    priorHandoffIn,
262    sectionOf,
263    stockSummaryOf,
264    priceUsage,
265    renderCommitments,
266    lookupRecord,
267    renderLedger,
268    replacementFor,
269    subagentRanThisTurn,
270    summarisePrs,
271    toolIndex,
272    windowReading,
273    withoutScratchpad,
274} from "./lib.js";
275
276export const register = (on, pluginOptions) => {
277    options = pluginOptions ?? {};
278    watches = new Map();
279    windows = new Map();
280
281    on("session.start", async ($, e, next) => {
282        await safely($, () => startWindow($));
283
284        await $.tool.register({
285            name: "force_compact",
286            description:
287                "Force a compaction of this conversation now, and report what it resolved to. " +
288                "For testing the compact-handoff plugin without waiting for the context to fill.",
289            inputSchema: {
290                type: "object",
291                properties: {
292                    instructions: {
293                        type: "string",
294                        description: "What the handoff should keep or stress.",
295                    },
296                },
297            },
298        });
299
300        await $.tool.register({
301            name: "handoff_status",
302            description:
303                "What compact-handoff would do if this conversation compacted right now: whether it is live, " +
304                "where its data lives, how many compactions this session has already been through, what the " +
305                "last one did, and what this session has spent on handoffs so far.",
306            inputSchema: {
307                type: "object",
308                properties: {
309                    limit: { type: "number", description: "How many recent rows, newest last. Default 10." },
310                },
311            },
312        });
313
314        await $.tool.register({
315            name: "handoff_list",
316            description:
317                "List the compactions this session has been through, oldest first: when each ran, what it cost, " +
318                "how large its handoff was, and which files it left behind.",
319            inputSchema: { type: "object", properties: {} },
320        });
321
322        await $.tool.register({
323            name: "handoff_lookup",
324            description:
325                "Read back a stored compaction. Without arguments, the whole of the most recent handoff. " +
326                "`section` reads one part: summary, ledger, state, commitments, or transcript for the whole " +
327                "conversation as it stood before that compaction. Long sections page with offset and limit.",
328            inputSchema: {
329                type: "object",
330                properties: {
331                    n: { type: "number", description: "Which compaction. Default the most recent." },
332                    section: {
333                        type: "string",
334                        description: "full, summary, ledger, state, commitments, or transcript. Default full.",
335                    },
336                    offset: { type: "number", description: "First line returned, 0-based. Default 0." },
337                    limit: { type: "number", description: "How many lines. Default 400." },
338                },
339            },
340        });
341
342        await $.tool.register({
343            name: "handoff_search",
344            description:
345                "Search every stored handoff and pre-compaction transcript of this session for a pattern and " +
346                "return the matching lines with the compaction they came from and their line numbers. This is " +
347                "how to find something a handoff did not carry up.",
348            inputSchema: {
349                type: "object",
350                properties: {
351                    pattern: { type: "string", description: "A basic regular expression, as grep reads it." },
352                    limit: { type: "number", description: "How many matching lines. Default 50." },
353                },
354                required: ["pattern"],
355            },
356        });
357
358        await $.tool.register({
359            name: "handoff_feedback",
360            description:
361                "Record that a handoff was missing something or wrong about something, against the compaction " +
362                "it came from. Written beside that compaction's own record, where the bench reads it.",
363            inputSchema: {
364                type: "object",
365                properties: {
366                    note: { type: "string", description: "What was missing or wrong, in your own words." },
367                    n: { type: "number", description: "Which compaction. Default the most recent." },
368                },
369                required: ["note"],
370            },
371        });
372
373        if (await isDev($)) {
374            await registerBenchTools($);
375        }
376
377        await registerHandoffAgent($);
378
379        return next(e);
380    });
381
382    // Nothing should delegate to this agent but this plugin. `agent.offer`
383    // hides it from the model's listing without hiding it from
384    // `$.agent.spawn`, which is what a runner-only type is for.
385    on("agent.offer", { agent: HANDOFF_AGENT }, () => ({ isOffered: false }));
386
387
388    // The tool only arms it. `$.session.compact` refuses to run under a hook
389    // that is holding the turn, and says so: "called from a tool.call hook, it
390    // would compact under the turn this hook is holding; call it from a later
391    // event (turn.complete)". So the tool writes a flag and the turn's end
392    // reads it.
393    on("tool.call", { tool: "mcp__compact-handoff__force_compact" }, async ($, e) => {
394        const instructions = typeof e.instructions === "string" ? e.instructions.trim() : "";
395
396        await $.store.set(ARMED_KEY, { instructions });
397
398        return { result: "Armed. The compaction runs when this turn ends; read it back with handoff_status." };
399    });
400
401    on("tool.call", { tool: "mcp__compact-handoff__refresh_handoff" }, async ($) => {
402        await $.store.set(ARMED_KEY, { refresh: true });
403
404        return { result: "Armed. The refresh starts when this turn ends; watch it with handoff_status." };
405    });
406
407    on("tool.call", { tool: "mcp__compact-handoff__probe_spawn" }, async ($, e, next) => {
408        const seconds = typeof e.seconds === "number" && e.seconds > 0 ? Math.floor(e.seconds) : 30;
409
410        // Two hooks can hold a spawn and they need not behave alike, so the
411        // probe runs in whichever the caller names.
412        if (e.where === "wait") {
413            await probeWait($, seconds, next.signal);
414
415            return { result: "Waited inside the tool.call hook; read it back with handoff_status." };
416        }
417
418        if (e.where === "tool") {
419            await probeSpawn($, seconds, next.signal);
420
421            return { result: "Probed from inside the tool.call hook; read it back with handoff_status." };
422        }
423
424        await $.store.set(ARMED_KEY, { probeSeconds: seconds });
425
426        return { result: `Armed. A subagent sleeping ${seconds}s runs when this turn ends; read it back with handoff_status.` };
427    });
428
429    on("tool.call", { tool: "mcp__compact-handoff__probe_budget" }, async ($, e, next) => {
430        // A fork reads the main thread's transcript, and a tool.call hook is
431        // holding the turn that transcript belongs to, so the probe can also
432        // be run from the turn's end instead.
433        if (e.at === "turn") {
434            await $.store.set(ARMED_KEY, { forkProbe: true });
435
436            return { result: "Armed. The probe runs when this turn ends; read it back with handoff_status." };
437        }
438
439        return { result: JSON.stringify(await probeBudget($, e, next.signal), null, 2) };
440    });
441
442    // Same reason `force_compact` only arms: a fork reads the main thread's
443    // transcript and this hook is holding the turn that transcript belongs to,
444    // so from here it waits out the turn and answers null.
445    on("tool.call", { tool: "mcp__compact-handoff__ab_fork" }, async ($, e) => {
446        const label = typeof e.label === "string" ? e.label.trim() : "";
447
448        if (label === "" || !/^[\w.-]+$/.test(label)) {
449            return { result: "A label must be a plain filename: letters, digits, dot, dash, underscore." };
450        }
451
452        const replicates = typeof e.replicates === "number" && e.replicates > 0 ? Math.min(Math.floor(e.replicates), 10) : 1;
453
454        if (!(await $.fs.exists(atRoot($, `${AB_PROMPTS}/${label}.txt`)))) {
455            return { result: `No variant at ${AB_PROMPTS}/${label}.txt.` };
456        }
457
458        await $.store.set(ARMED_KEY, { ab: { label, replicates } });
459
460        return { result: `Armed ${label} x${replicates}. It runs when this turn ends; read it back with handoff_status.` };
461    });
462
463    on("tool.call", { tool: "mcp__compact-handoff__handoff_status" }, async ($, e) => {
464        return { result: JSON.stringify(await readRuns($, e), null, 2) };
465    });
466
467    on("tool.call", { tool: "mcp__compact-handoff__handoff_list" }, async ($, e) => {
468        return { result: await logged($, "handoff_list", e, JSON.stringify(await listHandoffs($), null, 2)) };
469    });
470
471    on("tool.call", { tool: "mcp__compact-handoff__handoff_lookup" }, async ($, e) => {
472        return { result: await logged($, "handoff_lookup", e, await lookupHandoff($, e)) };
473    });
474
475    on("tool.call", { tool: "mcp__compact-handoff__handoff_search" }, async ($, e) => {
476        return { result: await logged($, "handoff_search", e, await searchHandoffs($, e)) };
477    });
478
479    on("tool.call", { tool: "mcp__compact-handoff__handoff_feedback" }, async ($, e) => {
480        return { result: await recordFeedback($, e) };
481    });
482
483    // Everything expensive happens here, between turns, where a dispatch that
484    // runs long delays nothing the user is waiting on.
485    //
486    // One step that throws used to cost the four after it and `next(e)` with
487    // them, which is how a single `TypeError` in the watch took the whole
488    // dispatch down on every turn at 2.1.273. Each step is wrapped now, so a
489    // failing one costs its own reading and nothing else.
490    on("turn.complete", async ($, e, next) => {
491        const armed = await safely($, () => $.store.get(ARMED_KEY));
492
493        if (armed !== undefined && armed !== null) {
494            await safely($, () => $.store.delete(ARMED_KEY));
495            await safely($, () => runArmed($, armed, next.signal));
496        }
497
498        await safely($, () => promotePending($));
499        await safely($, () => maybeRefresh($, armed?.refresh === true));
500        await safely($, () => watchPost($));
501        await safely($, () => watchWindow($));
502
503        return next(e);
504    });
505
506    // `precompute` is the engine building a compaction before it is needed, so
507    // that the real one is instant. Answering it with `next(e)` was a hole: the
508    // declarations say a precomputed result "is kept for the compaction that
509    // comes, if the conversation it ran over still leads", which is exactly the
510    // automatic path, so an auto-compaction could be served an engine summary
511    // this plugin had waved through and no row would ever say so. Declining the
512    // precompute forces the engine to dispatch a real event when it actually
513    // compacts. This is the declarations' own example for the case.
514    on("session.compact", { trigger: "precompute" }, () => ({
515        skip: "compact-handoff answers compactions live",
516    }));
517
518    on("session.compact", async ($, e, next) => {
519        const startedAt = Date.now();
520        const record = {
521            at: new Date().toISOString(),
522            sessionId: await safely($, () => $.session.id()),
523            cwd: await safely($, () => $.session.cwd()),
524            model: await safely($, () => $.session.model()),
525            trigger: e.trigger,
526            agentId: e.agentId ?? null,
527            messagesIn: e.messages.length,
528            // What the fork was charged to read, against what this session
529            // holds; null on every row that never reached a fork. Issue #690:
530            // a fork answering over a transcript that is not this conversation
531            // is invisible in every other field on the row.
532            forkInput: null,
533            pinnable: e.messages.filter(isPinnable).length,
534            pinnedSkipped: pinSkipCounts(e.messages),
535            plugin: await pluginVersion($),
536            engine: await engineVersion($),
537            // The arm of the live A/B split this compaction is in, the design
538            // that picked it and what it was decided on, or null when the
539            // split is off. Set below, once the compaction is known to be the
540            // main conversation's own.
541            ab: /** @type {AbRow | null} */ (null),
542        };
543
544        // The fork this plugin is waiting on is itself a query loop, and the
545        // engine checks it for auto-compaction like any other. Passing it
546        // through lets the engine summarise it, and the fork then answers over
547        // that summary instead of this conversation: a cold fork. Declining it
548        // lets the fork read the real transcript.
549        if (isOwnForkLoop(e)) {
550            record.outcome = "ownFork";
551            record.detail = "the engine tried to compact this plugin's own handoff fork; declined so the fork reads the conversation";
552            record.elapsedMs = Date.now() - startedAt;
553
554            await finish($, record, "skipped");
555
556            return { skip: "compact-handoff's own handoff fork reads the conversation uncompacted" };
557        }
558
559        // A subagent's compaction is a different conversation with a different
560        // owner, and nothing here has been measured against one. Pass it
561        // through with a row rather than reshape a transcript this plugin has
562        // never been graded on.
563        if (e.agentId !== null && e.agentId !== undefined && !(await handlesSubagents($))) {
564            record.outcome = "subagent";
565            record.elapsedMs = Date.now() - startedAt;
566
567            await finish($, record, "passedThrough");
568
569            return next(e);
570        }
571
572        record.ab = await abRowFor($, record);
573
574        // A subagent's compaction is another conversation; the main session's
575        // watch is not ended by it.
576        if (record.agentId === null) {
577            await safely($, () => closeWatch($, record.sessionId));
578        }
579
580        if (record.ab?.arm === "stock") {
581            return compactStock($, e, next, record, startedAt);
582        }
583
584        // Another plugin's work starts here, before anything is awaited, so it
585        // runs beside this plugin's fork over the same pre-compaction
586        // transcript and both read one warm cache. With nobody subscribed this
587        // is one array copy.
588        const seam = dispatchSeam($, e);
589
590        const context = await lineageContext($, e, record);
591
592        let handoff = { outcome: "threw", detail: "" };
593
594        try {
595            // A fork answers over this session's own transcript, so it needs
596            // nothing written in advance and it can honour the instructions a
597            // `/compact <instructions>` carries. The handoff on disk is only
598            // what stands when the fork has no warm transcript to read.
599            handoff = await forkHandoff($, e, record);
600
601            if (handoff.outcome !== "handoff") {
602                record.forkOutcome = handoff.outcome;
603                record.forkDetail = handoff.detail;
604                handoff = await readHandoff($, e);
605            }
606        } catch (error) {
607            handoff = { outcome: "threw", detail: String(error) };
608        }
609
610        record.seam = await seam;
611        record.outcome = handoff.outcome;
612        // A fork that ran and was refused still spent its tokens, so a usage
613        // already on the record outlives the fallback that replaced it.
614        record.usage = handoff.usage ?? record.usage ?? null;
615        record.staleMessages = handoff.staleMessages ?? null;
616        record.elapsedMs = Date.now() - startedAt;
617        record.aborted = next.signal.aborted;
618
619        if (handoff.detail !== "") {
620            record.detail = handoff.detail.slice(0, 2000);
621        }
622
623        if (handoff.outcome !== "handoff") {
624            priceRun(record, { commitmentsUsd: 0 });
625
626            await finish($, record, "fellBack", { transcript: renderTranscript(e.messages) });
627
628            const ordinal = await openerOrdinal($, record);
629            const result = await next(e);
630
631            armStock(record, ordinal, e.messages, result);
632
633            return result;
634        }
635
636        // The model's summary is only the first of four parts. The rest are read
637        // or gathered rather than recalled, and each one is allowed to fail: a
638        // summary on its own is still the arm that scored 67.3%.
639        const assembled = await assembledHandoff($, e, handoff.text, context, record.depth ?? null);
640        const maxHandoff = await handoffCeilingFor($);
641        const guarded = applySizeGuard(replacementFor(e.messages, assembled.text), {
642            maxChars: maxHandoff.chars,
643            pointerFor: ({ index, chars }) => trimPointer(record, index, chars),
644        });
645
646        record.maxHandoff = maxHandoff;
647
648        const restored = await restoreFiles($, e, handoff.requests ?? [], {
649            depth: record.depth ?? 0,
650            totalChars: Math.max(0, maxHandoff.chars - guarded.charsAfter),
651        });
652        // The handoff first, then the files it asked for, then the pinned turns.
653        const replacement = [guarded.messages[0], ...restored.messages, ...guarded.messages.slice(1)];
654
655        record.messagesOut = replacement.length;
656        record.restore = restored.restore;
657        record.summaryChars = handoff.text.length;
658        record.analysisChars = handoff.analysisChars ?? 0;
659        record.handoffChars = assembled.text.length;
660        record.handoffCharsBefore = guarded.charsBefore;
661        record.handoffCharsAfter = guarded.charsAfter;
662        record.trimmedTurns = guarded.trimmedTurns;
663        record.overCeiling = guarded.overCeiling;
664        record.parts = assembled.notes;
665        record.elapsedMs = Date.now() - startedAt;
666        record.aborted = next.signal.aborted;
667
668        priceRun(record, {
669            commitmentsUsd: assembled.notes.commitmentsCostUsd,
670            commitmentsBasis: assembled.notes.commitmentsCostBasis,
671        });
672
673        const artifacts = {
674            handoff: assembled.text,
675            summary: handoff.text,
676            transcript: renderTranscript(e.messages),
677            rows: assembled.rows,
678        };
679
680        if (!(await isLive($))) {
681            artifacts.replacement = renderTranscript(replacement);
682
683            await finish($, record, "rehearsed", artifacts);
684
685            const ordinal = await openerOrdinal($, record);
686            const result = await next(e);
687
688            armStock(record, ordinal, e.messages, result);
689
690            return result;
691        }
692
693        // A dispatch the engine has already given up on cannot be trusted to
694        // have its answer applied, and half a compaction is worse than the
695        // engine's own. Measured in the live checks; the row says which path.
696        if (next.signal.aborted) {
697            await finish($, record, "abortedFallback", artifacts);
698
699            const ordinal = await openerOrdinal($, record);
700            const result = await next(e);
701
702            armStock(record, ordinal, e.messages, result);
703
704            return result;
705        }
706
707        const ordinal = await openerOrdinal($, record);
708
709        await finish($, record, "replaced", artifacts);
710
711        if (ordinal !== null) {
712            armMonitor(record, { kind: "handoff", ordinal, from: replacement.length, before: e.messages });
713            restartWindow(record.sessionId);
714        }
715
716        return { messages: replacement };
717    });
718
719    // The noun another plugin reaches this one through. The fold runs once per
720    // load and before any hook of this plugin, and `built` is `$` as the steps
721    // beneath made it. Nothing may be done with `built` but spread it: at
722    // `engine.create` the static scan refuses the value of `next(e)` passed to
723    // a function, and handing it to `pluginVersion` here is what refused 0.5.0
724    // at load ("the value of next(e) at engine.create is passed as an argument").
725    on("engine.create", async (_$, e, next) => {
726        const built = await next(e);
727
728        return { ...built, compactHandoff: { beforeCompact, version } };
729    });
730};
731
732/* ------------------------------------------------------------------ *
733 * The seam: another plugin's work, run beside this plugin's fork.
734 *
735 * The seam carries strings and nothing else. Revision 1 took a callback, and a
736 * callback cannot cross a plugin boundary here at all: every plugin runs in its
737 * own environment and an interface call's arguments go through `cloneInto`,
738 * which throws `DataCloneError` on a function. So a subscriber names a TOOL it
739 * answers, and this plugin raises that tool through `$.tool.call`, which runs
740 * the subscriber's work in the subscriber's own environment.
741 * ------------------------------------------------------------------ */
742
743/**
744 * The version `$.compactHandoff.version()` answers, and a constant because the
745 * `engine.create` fold may not pass `built` to a function and `$` there is the
746 * empty table, so nothing at the fold can read the manifest. `test/module.test.js`
747 * asserts it against `.claude-plugin/plugin.json` so the two cannot drift.
748 */
749export const PLUGIN_VERSION = "0.13.0";
750
751/** How long one subscriber may run before the compaction goes on without it. */
752const DEFAULT_SEAM_TIMEOUT_MS = 90_000;
753
754/**
755 * Every subscriber, keyed by the tool it answers: subscribing twice is once.
756 *
757 * Exported so a test can start from an empty seam. A subscription lives for the
758 * load and nothing in the runtime ever removes one.
759 */
760export const seamSubscribers = new Map();
761
762/**
763 * Registers the tool this plugin raises when a compaction is about to happen,
764 * beside its own fork. Resolves `{ subscribed: true, tool }`.
765 *
766 * This is the whole public API of `$.compactHandoff` beside `version`, and it
767 * exists because a plugin keyed after this one in `enabledPlugins` never sees
768 * `session.compact` at all: this plugin answers that event without calling
769 * `next`. There is no unsubscribe, and a subscription lives for the load.
770 */
771const beforeCompact = async (options) => {
772    const tool = typeof options?.tool === "string" ? options.tool.trim() : "";
773
774    if (tool === "") {
775        throw new TypeError("compactHandoff.beforeCompact takes { tool }, the full name of the tool to raise");
776    }
777
778    if (seamSubscribers.has(tool)) {
779        return { subscribed: true, tool };
780    }
781
782    const name = typeof options?.name === "string" && options.name.trim() !== "" ? options.name.trim() : tool;
783
784    seamSubscribers.set(tool, { tool, name });
785
786    return { subscribed: true, tool };
787};
788
789/** This plugin's version, for a subscriber that wants to know what it got. */
790const version = async () => PLUGIN_VERSION;
791
792/**
793 * Every subscriber, raised at once and waited for together.
794 *
795 * Nothing a subscriber does can change what this plugin answers the compaction
796 * with: a throw and a deny are row fields, a subscriber still running after the
797 * timeout is left running and its result ignored. The compaction waits for the
798 * slower of this and the fork, which is the price of both reading the same
799 * transcript while its cache is warm.
800 */
801const dispatchSeam = ($, e) => {
802    const waiting = [...seamSubscribers.values()];
803
804    if (waiting.length === 0) {
805        return Promise.resolve({ subscribers: 0 });
806    }
807
808    const started = waiting.map((entry) => ({
809        name: entry.name,
810        tool: entry.tool,
811        startedAt: Date.now(),
812        work: settledCall($, entry.tool, e),
813    }));
814
815    // A host op failing while the seam is collected is a field on the row like
816    // any other reading here, never a compaction this plugin drops.
817    return collectSeam($, started).catch((error) => ({
818        subscribers: started.length,
819        failed: String(error).slice(0, 300),
820    }));
821};
822
823/**
824 * The subscriber's tool, raised now, as a promise that resolves however it ends.
825 *
826 * The raise carries three small values and no transcript. The subscriber reads
827 * the conversation the way this plugin does, with `$.model.fork` over the live
828 * session, which is the same pre-compaction transcript and the same warm cache;
829 * sending the messages through would clone a megabyte to say what a fork reads
830 * for nothing.
831 */
832const settledCall = ($, tool, e) => {
833    const threw = (error) => ({ outcome: "threw", detail: String(error).slice(0, 300) });
834
835    try {
836        return Promise.resolve(
837            $.tool.call({ tool, trigger: e.trigger, messageCount: e.messages.length }),
838        ).then(seamOutcome, threw);
839    } catch (error) {
840        return Promise.resolve(threw(error));
841    }
842};
843
844/** What the raise resolved to, read as an outcome: an answer, or a refusal. */
845const seamOutcome = (answer) => {
846    if (answer !== null && typeof answer === "object" && typeof answer.deny === "string") {
847        return { outcome: "denied", detail: answer.deny.slice(0, 300) };
848    }
849
850    if (answer === null || typeof answer !== "object" || answer.result === undefined) {
851        return { outcome: "ok", detail: "the raise was answered with no result" };
852    }
853
854    return { outcome: "ok" };
855};
856
857const collectSeam = async ($, started) => {
858    const capMs = capOf(opt("seamTimeoutMs") ?? (await $.env.get("COMPACT_HANDOFF_SEAM_TIMEOUT_MS")), DEFAULT_SEAM_TIMEOUT_MS);
859
860    const results = await Promise.all(
861        started.map(async (entry) => {
862            const ended = await Promise.race([
863                entry.work,
864                $.clock.sleep(capMs).then(() => ({ outcome: "timedOut" })),
865            ]);
866
867            return { name: entry.name, tool: entry.tool, ...ended, elapsedMs: Date.now() - entry.startedAt };
868        }),
869    );
870
871    return { subscribers: started.length, results };
872};
873
874/** How long one restore read may take before the file is read from the transcript instead. */
875const RESTORE_MS = 20_000;
876
877/**
878 * The files the summariser asked for, read through the real Read tool and
879 * shaped as the tool blocks a live read produces.
880 *
881 * Fresh first: the engine's own restore re-reads from disk, so a file edited
882 * mid-session comes back current, and `$.tool.call` is a host op, so its time
883 * in flight does not count against the dispatch budget. A read that is denied,
884 * errors or runs past `RESTORE_MS` falls back to the text of the last Read in
885 * the transcript, which is stale at worst and absent at best; the row names
886 * which of the three it was. Every cap is an env var, and every number that
887 * would let the caps be tuned is on the row.
888 */
889const restoreFiles = async ($, e, requests, { depth, totalChars }) => {
890    const startedAt = Date.now();
891    // Three literal reads rather than one helper: the static scan lists the
892    // variables a module reads, and it can only do that off a literal name.
893    const caps = {
894        maxFiles: capOf(opt("restoreFiles") ?? (await $.env.get("COMPACT_HANDOFF_RESTORE_FILES")), RESTORE_MAX_FILES),
895        fileChars: capOf(opt("restoreFileChars") ?? (await $.env.get("COMPACT_HANDOFF_RESTORE_FILE_CHARS")), RESTORE_FILE_CHARS),
896        totalChars: Math.min(capOf(opt("restoreTotalChars") ?? (await $.env.get("COMPACT_HANDOFF_RESTORE_TOTAL_CHARS")), RESTORE_TOTAL_CHARS), totalChars),
897    };
898    const chosen = await chooseRestores({
899        requests,
900        candidates: restoreCandidates(e.messages),
901        shellMentions: shellMentions(e.messages),
902        maxFiles: caps.maxFiles,
903        exists: (path) => existsOnDisk($, path),
904    });
905    const read = [];
906
907    for (const file of chosen.files) {
908        read.push(await readRestore($, file, caps.fileChars));
909    }
910
911    const fitted = fitRestores(read, caps.totalChars);
912    const messages = fitted
913        .filter((file) => file.text !== null)
914        .flatMap((file, index) => restorePair({ ...file, id: `toolu_handoff_${depth}_${index}` }));
915    const rows = fitted.map(restoreRow);
916
917    return {
918        messages,
919        restore: {
920            source: chosen.source,
921            requested: requests.length,
922            restored: messages.length / 2,
923            rejected: chosen.rejected,
924            files: rows,
925            chars: rows.reduce((total, row) => total + row.chars, 0),
926            approxTokens: rows.reduce((total, row) => total + row.approxTokens, 0),
927            caps,
928            ms: Date.now() - startedAt,
929        },
930    };
931};
932
933/**
934 * Whether a file is still there; an unanswerable check counts as yes, so the read decides.
935 *
936 * @param {import('claude-code').EngineInterface} $
937 * @param {string} path
938 */
939const existsOnDisk = async ($, path) => {
940    try {
941        return await $.fs.exists(path);
942    } catch {
943        return true;
944    }
945};
946
947const readRestore = async ($, file, fileChars) => {
948    const startedAt = Date.now();
949    const row = { ...file, text: null, source: "failed", detail: "" };
950
951    try {
952        const fresh = await withTimeout($, freshRead($, file), RESTORE_MS, `restore ${file.path}`);
953
954        if (fresh !== null) {
955            Object.assign(row, { text: fresh, source: "fresh" });
956        } else {
957            row.detail = "the Read tool denied or errored";
958        }
959    } catch (error) {
960        row.detail = String(error).slice(0, 200);
961    }
962
963    if (row.text === null && typeof file.stored === "string") {
964        Object.assign(row, { text: file.stored, source: "stored" });
965    }
966
967    if (row.text !== null) {
968        const clipped = clipRestore(row.text, fileChars);
969
970        row.text = clipped.text;
971        row.clippedChars = clipped.clippedChars;
972    }
973
974    row.ms = Date.now() - startedAt;
975
976    return row;
977};
978
979/** The Read tool's own text for the file, or null when it would not read it. */
980const freshRead = async ($, { path, from, to }) => {
981    const input = { tool: "Read", file_path: path };
982
983    if (from !== null && to !== null) {
984        input.offset = from;
985        input.limit = to - from + 1;
986    }
987
988    const answer = await $.tool.call(input);
989
990    if (answer === null || typeof answer !== "object" || typeof answer.deny === "string" || answer.isError === true) {
991        return null;
992    }
993
994    return typeof answer.text === "string" && answer.text !== "" ? answer.text : null;
995};
996
997const capOf = (value, fallback) => {
998    const raw = Number.parseInt(value ?? "", 10);
999
1000    return Number.isFinite(raw) && raw > 0 ? raw : fallback;
1001};
1002
1003/**
1004 * The handoff, written inside the event by a fork of this very session.
1005 *
1006 * `$.model.fork` runs one completion over the main thread's own
1007 * cache-safe transcript snapshot, which is how the engine's own compaction
1008 * reads a conversation, so nothing has to be marshalled in and the prompt
1009 * cache is already warm. It is a host op, and a host op's time in flight does
1010 * not count against the ten second dispatch budget: measured at 21768 ms from
1011 * `turn.complete`, not aborted.
1012 *
1013 * The fork resolves `ModelForkResult`, the union engine 2.1.280 introduced.
1014 * Engines before it resolved the reply or null, and that shape is not read any
1015 * more: the plugin targets 2.1.280 and later, and a dual-shape shim would hide
1016 * the next change the way the null check hid this one. An unanswered fork is a
1017 * named outcome (`nothing-to-fork`, which is no warm transcript and what the
1018 * handoff on disk is kept for, then `api-error`, `empty-reply` and `aborted`),
1019 * and the caller falls back on it like any other. So is a fork that runs past
1020 * `FORK_TIMEOUT_MS`, as `timeout`.
1021 *
1022 * A reply is not proof that the fork read this conversation. Some forks come
1023 * back having been charged for a fraction of the session's context (#690), and
1024 * the reply reads like any other summary, so `record.forkInput` is compared
1025 * against the context and a short one is refused here rather than handed up.
1026 *
1027 * @param {import('claude-code').EngineInterface} $
1028 */
1029const forkHandoff = async ($, e, record) => {
1030    const asked = typeof e.instructions === "string" && e.instructions.trim() !== ""
1031        ? `${FORK_PROMPT}\n\nThe person asked for this compaction with these instructions, and they outrank everything above: ${e.instructions.trim()}`
1032        : FORK_PROMPT;
1033
1034    record.forkContext = await forkContext($, e);
1035
1036    let reply;
1037
1038    try {
1039        reply = await withTimeout($, $.model.fork({ prompt: asked }), FORK_TIMEOUT_MS, "the fork");
1040    } catch (error) {
1041        if (error instanceof TimeoutError) {
1042            return { outcome: "timeout", detail: error.message };
1043        }
1044
1045        throw error;
1046    }
1047
1048    if (!reply.isAnswered) {
1049        // Every arm but `nothing-to-fork` made a request, and what it spent
1050        // stays on the row after the fallback replaces its answer.
1051        if (reply.reason !== "nothing-to-fork") {
1052            record.usage = reply.usage;
1053        }
1054
1055        return { outcome: reply.reason, detail: unansweredForkDetail(reply) };
1056    }
1057
1058    record.forkInput = forkInputOf(reply.usage, record.forkContext?.context ?? null);
1059
1060    // A fork charged for a fraction of the context answered over a transcript
1061    // that is not this conversation, and a summary of the wrong conversation is
1062    // worse than the engine's own. The tokens are spent either way, so the row
1063    // keeps the usage and is priced on it.
1064    if (record.forkInput.matchesContext === false) {
1065        record.usage = reply.usage;
1066
1067        return {
1068            outcome: "mismatch",
1069            detail:
1070                `the fork was charged for ${record.forkInput.sent} input tokens against a ` +
1071                `${record.forkInput.contextTokens} token context, so it did not read this conversation`,
1072        };
1073    }
1074
1075    const parsed = parseRestoreRequests(reply.text.trim());
1076    const kept = withoutScratchpad(parsed.summary);
1077
1078    if (kept.text === "") {
1079        return { outcome: "empty", detail: "the fork answered nothing" };
1080    }
1081
1082    return {
1083        outcome: "handoff",
1084        text: kept.text,
1085        analysisChars: kept.analysisChars,
1086        requests: parsed.requests,
1087        usage: reply.usage,
1088        detail: "",
1089    };
1090};
1091
1092/**
1093 * How long a compaction waits on its fork before it falls back to the handoff
1094 * on disk.
1095 *
1096 * Five minutes, from the index this plugin writes: across 58 warm forks since
1097 * 0.6.0 the median took about two minutes, the 95th percentile under four, and
1098 * one alone ran past five (about 25 minutes, with the person waiting on it the
1099 * whole time). A fork that is still going at five minutes is far more likely
1100 * to be the stuck one than a slow good one, and the forks this bound cuts most
1101 * often are the short ones the `forkInput` check refuses anyway.
1102 *
1103 * The bound gives the person their session back. It does not stop the spend:
1104 * `$.model.fork` takes a `prompt` and nothing else (no signal, no `timeoutMs`,
1105 * unlike `$.model.complete`), and its `aborted` arm fires only when the turn's
1106 * own dispatch is aborted, so there is no way to cancel the request from here.
1107 * The abandoned fork runs to its end, and whatever it answers is dropped.
1108 */
1109const FORK_TIMEOUT_MS = 300_000;
1110
1111/**
1112 * Why a fork came back without a reply, in the words the row keeps.
1113 *
1114 * @param {Exclude<import('claude-code').ModelForkResult, { isAnswered: true }>} reply
1115 * @returns {string}
1116 */
1117const unansweredForkDetail = (reply) => {
1118    switch (reply.reason) {
1119        case "nothing-to-fork":
1120            return "the fork found no warm main-thread transcript";
1121        case "api-error":
1122            return `the fork's request failed: ${reply.error}, status ${reply.status ?? "none, no response arrived"}`;
1123        case "empty-reply":
1124            return "the fork replied with no text";
1125        case "aborted":
1126            return "the fork was cut off by the turn's abort";
1127        default:
1128            return unreachable(reply);
1129    }
1130};
1131
1132/**
1133 * A union arm no case above handles. The `never` makes a new arm in the
1134 * engine's declarations a type error here rather than a silent fallthrough.
1135 *
1136 * @param {never} value
1137 * @returns {never}
1138 */
1139const unreachable = (value) => {
1140    throw new Error(`an arm nothing handles: ${JSON.stringify(value)}`);
1141};
1142
1143/**
1144 * What the session looked like the instant before it forked, logged and nothing
1145 * else. Every reading is allowed to fail on its own: a missing row here must
1146 * never cost the fork.
1147 */
1148const forkContext = async ($, e) => {
1149    const now = Date.now();
1150    const usage = await safely($, () => $.session.usage());
1151    const lastForkAt = await safely($, () => $.store.get(LAST_FORK_KEY));
1152
1153    await safely($, () => $.store.set(LAST_FORK_KEY, now));
1154
1155    return {
1156        context: usage?.context ?? null,
1157        model: await safely($, () => $.session.model()),
1158        messages: e.messages.length,
1159        msSinceLastFork: typeof lastForkAt === "number" ? now - lastForkAt : null,
1160        subagentRanThisTurn: subagentRanThisTurn(e.messages),
1161    };
1162};
1163
1164/**
1165 * The read half of a compaction, and all of it that runs inside the event: a
1166 * handoff already on disk, or a named reason the engine should take this one.
1167 */
1168const readHandoff = async ($, e) => {
1169    // A `/compact <instructions>` asks for something this handoff was written
1170    // before anyone asked for, and nothing here can rewrite it in time.
1171    if (typeof e.instructions === "string" && e.instructions.trim() !== "") {
1172        return { outcome: "instructed", detail: "the compaction carried instructions a precomputed handoff cannot honour" };
1173    }
1174
1175    if (!(await $.fs.exists(atRoot($, LATEST)))) {
1176        return { outcome: "noHandoff", detail: "no handoff has been written yet" };
1177    }
1178
1179    const text = (await $.fs.read(atRoot($, LATEST))).trim();
1180
1181    if (text === "") {
1182        return { outcome: "noHandoff", detail: "the handoff on disk is empty" };
1183    }
1184
1185    const ready = await $.store.get(READY_KEY);
1186    const staleMessages = stalenessOf(e.messages.length, ready);
1187
1188    if (staleMessages === null || staleMessages > MAX_STALE_MESSAGES) {
1189        return { outcome: "tooStale", staleMessages, detail: "the handoff predates too much of this conversation" };
1190    }
1191
1192    return { outcome: "handoff", text, staleMessages, detail: "" };
1193};
1194
1195/**
1196 * How much conversation happened after the handoff was written, which is the
1197 * only thing that decides whether it can still stand in for one.
1198 */
1199const stalenessOf = (messagesNow, ready) =>
1200    typeof ready?.messages === "number" ? Math.max(0, messagesNow - ready.messages) : null;
hooks/restore.js 346 lines
1/**
2 * Restoring files after a compaction, chosen by the summariser.
3 *
4 * Claude Code's own compaction hands the next window up to five of the most
5 * recently read files, and it renders each as a meta user message narrating a
6 * Read call inside a system-reminder. That path never runs when a hook
7 * replaces the messages, so until 0.3.0 a handoff came back with no files at
8 * all.
9 *
10 * This is the pure half of the replacement. The summariser, which has read the
11 * whole conversation, names the files the next window should have open, in a
12 * fenced block at the end of its summary; the code here parses that block,
13 * checks every path against the files the session actually touched, falls
14 * back to recency when the model named nothing usable, and shapes each file
15 * as a real Read tool_use and its tool_result, which is what a live read looks
16 * like. (Operator, 2026-09-15: "Why wouldn't we let the model that handles the
17 * compaction summary decide which files are re-read and injected", and the
18 * caps are theirs to move: "if the size guard needs expanded we can expand
19 * it".)
20 *
21 * Nothing in this file touches the host; the reads happen in `module.js`.
22 */
23
24import { PATH_SHAPED, SHELL_EXPANDED, nameOf } from "./lib.js";
25
26/** How many files may come back; Claude Code's own restore stops at five. */
27export const RESTORE_MAX_FILES = 5;
28
29/** The most of one file that comes back, at four characters per token. */
30export const RESTORE_FILE_CHARS = 20_000;
31
32/** The most all restored files may add to the window together. */
33export const RESTORE_TOTAL_CHARS = 100_000;
34
35export const RESTORE_OPEN = "<restore-files>";
36export const RESTORE_CLOSE = "</restore-files>";
37
38/** The tools whose target is a file the session has seen. */
39export const RESTORE_TOOLS = new Set(["Read", "Edit", "Write", "MultiEdit", "NotebookEdit"]);
40
41/** Memory files come back through the engine's own path; restoring them doubles them. */
42export const MEMORY_FILE = /(^|\/)(CLAUDE\.md|CLAUDE\.local\.md|AGENTS\.md)$/u;
43
44/** One request line: an absolute path, `all` or a line range, and a reason. */
45const REQUEST_LINE = /^\s*`?([^`|]+?)`?\s*\|\s*([^|]*?)\s*(?:\|\s*(.*?))?\s*$/u;
46
47const LINE_RANGE = /^(\d+)\s*[-–]\s*(\d+)$/u;
48
49/**
50 * The block out of the summary, and the summary without it.
51 *
52 * A summary with no block, or an empty one, asks for nothing; the caller then
53 * falls back to recency and the row says so.
54 */
55export const parseRestoreRequests = (text) => {
56    const open = text.lastIndexOf(RESTORE_OPEN);
57    const close = open === -1 ? -1 : text.indexOf(RESTORE_CLOSE, open);
58
59    if (open === -1 || close === -1) {
60        return { summary: text, requests: [] };
61    }
62
63    const body = text.slice(open + RESTORE_OPEN.length, close);
64    const summary = `${text.slice(0, open)}${text.slice(close + RESTORE_CLOSE.length)}`.trim();
65
66    return { summary, requests: body.split("\n").map(parseRequest).filter((request) => request !== null) };
67};
68
69const parseRequest = (raw) => {
70    const line = raw.trim();
71
72    if (line === "" || line.startsWith("#") || line.startsWith("-")) {
73        return null;
74    }
75
76    const match = REQUEST_LINE.exec(line);
77
78    if (match === null) {
79        return null;
80    }
81
82    const range = LINE_RANGE.exec(match[2] ?? "");
83    const from = range === null ? null : Number(range[1]);
84    const to = range === null ? null : Number(range[2]);
85
86    if (from !== null && (from < 1 || to < from)) {
87        return { path: match[1].trim(), from: null, to: null, reason: (match[3] ?? "").trim() };
88    }
89
90    return { path: match[1].trim(), from, to, reason: (match[3] ?? "").trim() };
91};
92
93/**
94 * Every file the conversation touched, most recently touched first, each with
95 * the text of its last successful Read if the transcript still holds one.
96 */
97export const restoreCandidates = (messages) => {
98    const seen = new Map();
99
100    messages.forEach((message, index) => {
101        for (const use of message.toolUses ?? []) {
102            const tool = nameOf(use);
103
104            if (!RESTORE_TOOLS.has(tool)) {
105                continue;
106            }
107
108            const path = pathOf(use);
109
110            if (path === null) {
111                continue;
112            }
113
114            const prior = seen.get(path);
115            const stored = tool === "Read" && use.isError !== true && typeof use.text === "string" && use.text !== ""
116                ? use.text
117                : (prior?.stored ?? null);
118
119            seen.set(path, { path, at: index, stored });
120        }
121    });
122
123    return [...seen.values()].sort((left, right) => right.at - left.at);
124};
125
126/**
127 * Every path-shaped word a shell command named, newest command first.
128 *
129 * A file read with `cat` or rewritten by a script was touched as surely as one
130 * that went through Read or Edit, so the model may ask for it back. These are
131 * words, not resolved paths: a relative one only ever matches the tail of an
132 * absolute path the model names, and none of them feeds the recency fallback,
133 * which would otherwise restore whatever log a command last redirected into.
134 */
135export const shellMentions = (messages) => {
136    const mentions = [];
137
138    messages.forEach((message, index) => {
139        for (const use of message.toolUses ?? []) {
140            const command = use.input?.command;
141
142            if (nameOf(use) !== "Bash" || typeof command !== "string") {
143                continue;
144            }
145
146            for (const word of command.split(SHELL_WORD_BREAK)) {
147                const token = word.slice(word.lastIndexOf("=") + 1).replace(/^\.\//u, "");
148
149                if (PATH_SHAPED.test(token) && !SHELL_EXPANDED.test(token)) {
150                    mentions.push({ token, at: index });
151                }
152            }
153        }
154    });
155
156    return mentions.sort((left, right) => right.at - left.at);
157};
158
159const SHELL_WORD_BREAK = /[\s"'`;|&()<>,]+/u;
160
161/** The file a shell command named, if any word it used ends the requested path. */
162const shellCandidateFor = (path, mentions) => {
163    const wanted = path.replace(/^\.\//u, "");
164
165    for (const { token, at } of mentions) {
166        if (token.startsWith("/") && (token === wanted || token.endsWith(`/${wanted}`))) {
167            return { path: token, at, stored: null };
168        }
169
170        if (wanted.startsWith("/") && wanted.endsWith(`/${token}`)) {
171            return { path: wanted, at, stored: null };
172        }
173    }
174
175    return null;
176};
177
178const pathOf = (use) => {
179    const input = use.input ?? {};
180    const path = input.file_path ?? input.notebook_path;
181
182    return typeof path === "string" && path.trim() !== "" ? path.trim() : null;
183};
184
185/**
186 * Which files go back, and why each one the model asked for did not.
187 *
188 * The model chooses; the code only refuses. A path the session never touched,
189 * through a file tool or by a shell command naming it, is refused because the
190 * model cannot have read it, a memory file because the
191 * engine re-emits those itself, and anything past the cap because the cap is
192 * the operator's. A file that is gone from disk is refused too: a worktree
193 * removed after its files were edited once left the fallback restoring nothing.
194 * When nothing survives, the most recently touched files that still exist go
195 * instead, and `source` says which of the two happened so the rate can be
196 * measured. `exists` is the host's check, injected so this stays pure.
197 *
198 * @param {object} options
199 * @param {Array<{ path: string, from: number | null, to: number | null, reason: string }>} options.requests
200 * @param {Array<{ path: string, at: number, stored: string | null }>} options.candidates
201 * @param {Array<{ token: string, at: number }>} [options.shellMentions] what shell commands named; see `shellMentions`
202 * @param {number} [options.maxFiles]
203 * @param {(path: string) => Promise<boolean>} [options.exists]
204 */
205export const chooseRestores = async ({ requests, candidates, shellMentions: mentions = [], maxFiles = RESTORE_MAX_FILES, exists = async () => true }) => {
206    const files = [];
207    const rejected = [];
208
209    for (const request of requests) {
210        const candidate = candidateFor(request.path, candidates) ?? shellCandidateFor(request.path, mentions);
211        const why = refusal(request.path, candidate, files, maxFiles) ?? (await goneFromDisk(candidate, exists));
212
213        if (why !== null) {
214            rejected.push({ path: request.path, why });
215            continue;
216        }
217
218        files.push({ ...request, path: candidate.path, stored: candidate.stored });
219    }
220
221    if (files.length > 0) {
222        return { files, source: "model", rejected };
223    }
224
225    const fallback = [];
226
227    for (const candidate of candidates) {
228        if (fallback.length >= maxFiles) {
229            break;
230        }
231
232        if (MEMORY_FILE.test(candidate.path) || !(await exists(candidate.path))) {
233            continue;
234        }
235
236        fallback.push({
237            path: candidate.path,
238            from: null,
239            to: null,
240            reason: "most recently touched; the summary named no usable file",
241            stored: candidate.stored,
242        });
243    }
244
245    return { files: fallback, source: fallback.length === 0 ? "none" : "recency", rejected };
246};
247
248/**
249 * @param {{ path: string }} candidate
250 * @param {(path: string) => Promise<boolean>} exists
251 */
252const goneFromDisk = async (candidate, exists) => ((await exists(candidate.path)) ? null : "no longer exists on disk");
253
254const candidateFor = (path, candidates) =>
255    candidates.find((candidate) => candidate.path === path) ??
256    candidates.find((candidate) => candidate.path.endsWith(`/${path.replace(/^\.\//u, "")}`)) ??
257    null;
258
259const refusal = (path, candidate, chosen, maxFiles) => {
260    if (candidate === null) {
261        return "not read or written in this conversation";
262    }
263
264    if (MEMORY_FILE.test(candidate.path)) {
265        return "memory file; the engine re-emits it on the next turn";
266    }
267
268    if (chosen.some((file) => file.path === candidate.path)) {
269        return "named twice";
270    }
271
272    if (chosen.length >= maxFiles) {
273        return `over the cap of ${maxFiles} files`;
274    }
275
276    return null;
277};
278
279/**
280 * The two messages the model sees for one restored file: a Read it made, and
281 * what the file said. Real tool blocks, not a narration of them, so the next
282 * window treats the content exactly as it treats any file it has read.
283 */
284export const restorePair = ({ id, path, from, to, text }) => {
285    const input = { file_path: path };
286
287    if (from !== null && to !== null) {
288        input.offset = from;
289        input.limit = to - from + 1;
290    }
291
292    return [
293        { role: "assistant", text: "", toolUses: [{ tool_use_id: id, tool: "Read", input }] },
294        { role: "user", text: "", toolUses: [], toolResults: [{ tool_use_id: id, text }] },
295    ];
296};
297
298/** What replaces the tail of a file too large to carry whole. */
299export const clipRestore = (text, maxChars) => {
300    if (text.length <= maxChars) {
301        return { text, clippedChars: 0 };
302    }
303
304    const kept = text.slice(0, maxChars);
305
306    return {
307        text: `${kept}\n[compact-handoff clipped ${text.length - maxChars} characters here; Read the file with an offset for the rest.]`,
308        clippedChars: text.length - maxChars,
309    };
310};
311
312/**
313 * Fits the read files under the total budget, in the order the model gave
314 * them, so the first file named is the last one dropped.
315 */
316export const fitRestores = (files, totalChars) => {
317    let used = 0;
318
319    return files.map((file) => {
320        if (file.text === null) {
321            return file;
322        }
323
324        if (used + file.text.length > totalChars) {
325            return { ...file, text: null, source: "dropped", detail: `over the total budget of ${totalChars} characters` };
326        }
327
328        used += file.text.length;
329
330        return file;
331    });
332};
333
334/** The metrics row for one file, whatever happened to it. */
335export const restoreRow = ({ path, from, to, reason, source, text, clippedChars, detail, ms }) => ({
336    path,
337    lines: from === null ? "all" : `${from}-${to}`,
338    reason,
339    source,
340    chars: text === null ? 0 : text.length,
341    approxTokens: text === null ? 0 : Math.ceil(text.length / 4),
342    clippedChars: clippedChars ?? 0,
343    ...(detail === undefined || detail === "" ? {} : { detail }),
344    ms,
345});
346
hooks/lib.js 1755 lines
1/**
2 * compact-handoff: everything a handoff is built out of that needs no `$`.
3 *
4 * Split out of `module.js` so it can be tested. The runtime loads a sibling
5 * import: `import { x } from "./lib.js"` inside a plugin's `hooks/module.js`
6 * resolves and runs in a real session, verified 2026-09-14 against build
7 * 2.1.270 with a probe plugin that wrote the imported value to disk. The static
8 * scan `claude plugin validate` runs accepts it too, which is the cheaper half
9 * of that check and on its own would have proved nothing.
10 *
11 * Everything here is a pure function of its arguments. Nothing reads the clock,
12 * the filesystem, the environment or the session. That is the whole point: the
13 * parts of a compaction that can be wrong in a way a test can catch live here,
14 * and the parts that can only be wrong against a live engine live next door.
15 */
16
17/** Longest tool argument kept in the ledger; enough for a real command line. */
18export const LEDGER_ARG_CHARS = 300;
19
20/** Longest shell command the trimmed ledger shows; the full one keeps LEDGER_ARG_CHARS. */
21export const LEDGER_SHELL_CHARS = 150;
22
23/** The newest shell commands the trimmed ledger always shows, read-only or not. */
24export const LEDGER_RECENT_COMMANDS = 10;
25
26/** The newest failures the trimmed ledger shows. */
27export const LEDGER_FAILURES = 20;
28
29/** Budget for the trimmed ledger's body; older commands that changed something drop first. */
30export const LEDGER_CHARS = 6000;
31
32/** Longest error line quoted back. */
33export const LEDGER_ERROR_CHARS = 200;
34
35/** How much of a tool result is scanned for an error shape. */
36export const ERROR_SCAN_CHARS = 4000;
37
38/** Longest assistant turn fed to the commitments pass. */
39export const COMMITMENT_TURN_CHARS = 6000;
40
41/** Ceiling on the whole commitments prompt; oldest turns drop first. */
42export const COMMITMENT_PROMPT_CHARS = 400_000;
43
44/**
45 * What an error looks like in a tool result that carries no error flag.
46 *
47 * A missing `isError` does not mean a command succeeded. The one genuinely
48 * failed command in the Round 3 fixture wrote `ugrep: warning: ...: No such
49 * file or directory` to stdout, with empty stderr and no flag anywhere, and
50 * every arm that trusted the flag reported it as a success.
51 */
52export const ERROR_SHAPE =
53    /^.*\b(?:no such file or directory|command not found|permission denied|fatal:|error:|warning:|traceback \(most recent call last\)|cannot access)\b.*$/imu;
54
55/** The argument worth naming for a tool call, in the order worth trying. */
56export const LEDGER_KEYS = ["command", "file_path", "path", "pattern", "url", "description"];
57
58/** One parsed row of the commitments pass. */
59export const COMMITMENT_ROW = /^\s*(UNKEPT|CORRECTED|UNANSWERED)\s*\|\s*T(\d+)\s*\|\s*(.+?)\s*\|\s*(.+?)\s*$/iu;
60
61export const COMMITMENT_HEADING = {
62    unkept: "Said it would, no evidence it did",
63    corrected: "Claims corrected or withdrawn",
64    unanswered: "Questions put to the user and never answered",
65};
66
67/* ------------------------------------------------------------------ *
68 * Which user turns a person actually typed.
69 * ------------------------------------------------------------------ */
70
71/**
72 * The user-role turns nobody typed, by the shape they actually arrive in.
73 *
74 * These are counted, not guessed. A sweep over every session JSONL on this box
75 * (2026-09-14) classified the opening of every user turn that carries no tool
76 * result: 2199 continuation prompts, 1993 task notifications, 613 slash-command
77 * envelopes, 410 local command outputs, 31 cross-session messages and 14 idle
78 * notices. A filter written from memory would have missed the two that matter
79 * most here, because a cross-session message never starts with its own tag: the
80 * harness puts "Another Claude session sent a message:" in front of it.
81 *
82 * The /loop wakeup is deliberately absent. A wakeup re-fires the user's own
83 * `/loop` prompt verbatim, so the turn it produces IS the person's words and
84 * pinning it is right.
85 */
86export const NON_PERSON = [
87    { kind: "cross-session", test: (text) => text.includes("<cross-session-message") },
88    { kind: "task-notification", test: (text) => text.startsWith("<task-notification") },
89    { kind: "idle-notice", test: (text) => text.startsWith("[Cross-session idle notice]") },
90    { kind: "command-output", test: (text) => text.startsWith("<local-command-stdout>") },
91    { kind: "system-reminder", test: (text) => text.startsWith("<system-reminder>") },
92    {
93        kind: "continuation",
94        test: (text) => text.startsWith("This session is being continued from a previous conversation"),
95    },
96    { kind: "interrupt", test: (text) => text.startsWith("[Request interrupted by user") },
97];
98
99/**
100 * Who produced a user turn: the person at the keyboard, or the harness.
101 *
102 * A slash-command envelope (`<command-name>/clear</command-name>`) counts as the
103 * person. They typed it, and `/loop keep going` carries the whole instruction in
104 * its arguments; dropping those would lose real intent to save a line.
105 */
106export const originOf = (message) => {
107    const text = (message.text ?? "").trim();
108
109    for (const { kind, test } of NON_PERSON) {
110        if (test(text)) {
111            return kind;
112        }
113    }
114
115    return "person";
116};
117
118/**
119 * Whether a message is a user turn worth keeping verbatim: a person said it,
120 * the engine will vouch for it, and it is not the transcript's way of carrying
121 * a tool result back to the model.
122 */
123export const isPinnable = (message) =>
124    message.role === "user" &&
125    typeof message.handle === "string" &&
126    (message.text ?? "").trim() !== "" &&
127    (message.toolResults === undefined || message.toolResults.length === 0) &&
128    originOf(message) === "person";
129
130/**
131 * The user turns the engine would vouch for that a person did not type, by kind.
132 *
133 * Recorded per compaction so a filter that starts dropping real turns shows up
134 * as a number moving rather than as a session quietly losing its instructions.
135 */
136export const pinSkipCounts = (messages) => {
137    const counts = {};
138
139    for (const message of messages) {
140        if (message.role !== "user" || typeof message.handle !== "string") {
141            continue;
142        }
143
144        if ((message.text ?? "").trim() === "") {
145            continue;
146        }
147
148        if (message.toolResults !== undefined && message.toolResults.length > 0) {
149            continue;
150        }
151
152        const kind = originOf(message);
153
154        if (kind !== "person") {
155            counts[kind] = (counts[kind] ?? 0) + 1;
156        }
157    }
158
159    return counts;
160};
161
162/**
163 * The conversation the session carries on with: the handoff first, then the
164 * words of every user turn the engine would vouch for, without the handle.
165 *
166 * The handle selects the turn; it is not handed back. A message returned with
167 * it stands as the engine has it, and the engine's copy carries every
168 * attachment the turn arrived with: the instruction bundle (CLAUDE.md, every
169 * rule file, AGENTS.md, MEMORY.md), the hook outputs, the skill and agent
170 * listings. Measured on 2026-09-17 (session 7c6495a3, depth 8): a 22.7k-char
171 * handoff came back as a 124k-token first turn, of which 218k chars were four
172 * copies of the instruction bundle riding on 19 pinned turns. The engine
173 * re-emits that bundle on its own once a compaction has run
174 * (`InstructionsLoaded` has a `compact` load reason), so those copies said
175 * nothing twice. A turn rebuilt from `role` and `text` is the person's words
176 * and nothing else.
177 */
178export const replacementFor = (messages, handoff) => [
179    { role: "assistant", text: handoff, toolUses: [] },
180    ...messages.filter(isPinnable).map(wordsOnly),
181];
182
183/** A user turn as the person typed it: the engine's attachments stay behind. */
184const wordsOnly = (message) => ({ role: message.role, text: message.text, toolUses: message.toolUses ?? [] });
185
186/* ------------------------------------------------------------------ *
187 * The size guard.
188 * ------------------------------------------------------------------ */
189
190/**
191 * Keep the replacement small enough that the engine does not compact it again.
192 *
193 * A handoff plus every pinned user turn is not bounded by anything. One pasted
194 * 30k-character blob is a normal thing for a person to do and several of them
195 * is a normal week, so the replacement can plausibly come back larger than the
196 * conversation it replaced, and the engine's answer to that is to compact
197 * immediately: a second summary, of a summary, and the real user turns gone.
198 *
199 * The largest pinned turns are replaced with a pointer at the stored copy until
200 * the whole thing fits. **The summary is never trimmed** — it is the part that
201 * scored 67.3% on its own, and a session that has lost it has lost everything
202 * the compaction was for. If the summary alone is over the ceiling the guard
203 * gives up and says so rather than cutting into it.
204 */
205export const applySizeGuard = (messages, { maxChars, pointerFor }) => {
206    const charsOf = (list) => list.reduce((total, message) => total + (message.text ?? "").length, 0);
207    const charsBefore = charsOf(messages);
208
209    if (charsBefore <= maxChars) {
210        return { messages, trimmedTurns: 0, charsBefore, charsAfter: charsBefore, overCeiling: false };
211    }
212
213    // Index 0 is the summary and is not a candidate at any size.
214    const order = messages
215        .map((message, index) => ({ index, length: (message.text ?? "").length }))
216        .slice(1)
217        .sort((left, right) => right.length - left.length);
218
219    const out = [...messages];
220    let trimmedTurns = 0;
221    let charsAfter = charsBefore;
222
223    for (const { index, length } of order) {
224        if (charsAfter <= maxChars) {
225            break;
226        }
227
228        const pointer = pointerFor({ index, chars: length });
229
230        // A pointer longer than what it replaces would make things worse.
231        if (pointer.length >= length) {
232            continue;
233        }
234
235        out[index] = { ...out[index], text: pointer, handle: undefined };
236        charsAfter = charsAfter - length + pointer.length;
237        trimmedTurns += 1;
238    }
239
240    return { messages: out, trimmedTurns, charsBefore, charsAfter, overCeiling: charsAfter > maxChars };
241};
242
243/* ------------------------------------------------------------------ *
244 * Lineage: how a second compaction knows about the first.
245 * ------------------------------------------------------------------ */
246
247/**
248 * The machine-readable first line every handoff carries.
249 *
250 * Without it the second compaction sees the first one's handoff as ordinary
251 * assistant prose and re-summarises it, which is how a tool ledger erodes into
252 * "the session ran some commands" over three compactions. With it the ledger
253 * rows are merged from the stored JSON instead of being re-read from English.
254 */
255export const lineageLine = ({ session, n, prev }) =>
256    `<!-- compact-handoff: session=${session} n=${n} prev=${prev ?? "none"} -->`;
257
258export const LINEAGE_SHAPE = /<!--\s*compact-handoff:\s*session=(\S+)\s+n=(\d+)\s+prev=(\d+|none)\s*-->/u;
259
260/** The lineage a piece of text carries, or null if it carries none. */
261export const lineageOf = (text) => {
262    const match = LINEAGE_SHAPE.exec(text ?? "");
263
264    if (match === null) {
265        return null;
266    }
267
268    const [, session, n, prev] = match;
269
270    return { session, n: Number(n), prev: prev === "none" ? null : Number(prev) };
271};
272
273/**
274 * The lineage of a message that IS a handoff, or null.
275 *
276 * A handoff opens with its marker. The same marker also turns up inside a
277 * tool result (`handoff_lookup`, a Read of a stored `.md`, a grep, the
278 * README's own example), and counting one of those as a handoff put a
279 * session's depth back to whatever the quoted handoff said and ended the
280 * post-compaction watch the moment a session read its history.
281 */
282export const handoffLineageOf = (message) => {
283    const text = (message.text ?? "").trimStart();
284
285    if ((message.toolResults ?? []).length > 0 || LINEAGE_SHAPE.exec(text)?.index !== 0) {
286        return null;
287    }
288
289    return lineageOf(text);
290};
291
292/** The most recent handoff already sitting in the conversation, if any. */
293export const priorHandoffIn = (messages) => {
294    for (let index = messages.length - 1; index >= 0; index -= 1) {
295        const lineage = handoffLineageOf(messages[index]);
296
297        if (lineage !== null) {
298            return { index, lineage };
299        }
300    }
301
302    return null;
303};
304
305/**
306 * Every place a compaction started a new conversation, oldest first: a handoff
307 * this plugin wrote (`kind: "handoff"`, `n` from its lineage) or the engine's
308 * own continuation summary (`kind: "stock"`, `n` its position among openers).
309 *
310 * `$.session.messages()` answers the whole session, every window since the
311 * first, so this is how a reading finds the window it is actually in.
312 */
313export const openersIn = (messages) => {
314    const openers = [];
315
316    messages.forEach((message, index) => {
317        const lineage = handoffLineageOf(message);
318        const chars = (message.text ?? "").length;
319
320        if (lineage !== null) {
321            openers.push({ kind: "handoff", index, n: lineage.n, chars });
322        } else if (
323            message.role === "user" &&
324            (message.toolResults ?? []).length === 0 &&
325            originOf(message) === "continuation"
326        ) {
327            openers.push({ kind: "stock", index, n: openers.length + 1, chars });
328        }
329    });
330
331    return openers;
332};
333
334/* ------------------------------------------------------------------ *
335 * The tool ledger.
336 * ------------------------------------------------------------------ */
337
338/**
339 * Every tool call the session made, what it was aimed at, and what the
340 * transcript says came back. Never an exit code: there isn't one.
341 */
342export const ledgerRows = (messages) => {
343    const rows = [];
344
345    for (const message of messages) {
346        for (const use of message.toolUses ?? []) {
347            const { outcome, detail } = outcomeOf(use);
348
349            const name = nameOf(use);
350            const command = name === "Bash" && typeof use.input?.command === "string" ? use.input.command : "";
351
352            rows.push({
353                name,
354                target: targetOf(use),
355                outcome,
356                detail,
357                writes: shellWrites(command),
358                readOnly: command !== "" && isReadOnlyCommand(command),
359            });
360        }
361    }
362
363    return rows;
364};
365
366/**
367 * The ledger as the next session reads it: this compaction's rows and nothing
368 * older.
369 *
370 * Until 0.2.0 every earlier compaction's rows were merged in under their own
371 * heading, and by the tenth compaction of one session the ledger was 65k of an
372 * 82k-character handoff (measured 2026-09-15). The earlier rows are still on
373 * disk in each compaction's own JSON; `handoff_lookup` reads them back, and
374 * the note above the summary says so.
375 *
376 * Within one compaction the shell list is trimmed too: a window of 90 commands
377 * rendered 21k characters, most of it read-only probes and the long bodies of
378 * commands that are also in the summary's inventory. The trimmed list keeps the
379 * newest LEDGER_RECENT_COMMANDS whatever they were, then every older command
380 * that changed something while LEDGER_CHARS lasts, and counts the rest. `full`
381 * renders every row, for `handoff_lookup section=ledger`, which reads the rows
382 * back from the compaction's JSON.
383 */
384export const renderLedger = (rows, { full = false, n = null } = {}) => {
385    if (rows.length === 0) {
386        return "";
387    }
388
389    return ["## Tool ledger", "", ...ledgerBody(rows, { full, n })].join("\n").trimEnd();
390};
391
392/**
393 * Four characters per token is the estimate the size guard already uses. It is
394 * not a tokenizer, and the field is named for what it is.
395 */
396export const approxTokens = (text) => Math.ceil(text.length / 4);
397
398/**
399 * How full the window is on one turn, and how much of that is the handoff.
400 *
401 * The question this answers cannot be answered from a compaction row alone: a
402 * compaction row says what the handoff cost to write, not what carrying it
403 * costs to read. Turn 1 of a session with no handoff in it is the floor - the
404 * system prompt, the rules files, the tool declarations and the first message,
405 * everything a window pays before any work happens - and turn 1 of a
406 * post-compact window is that same floor plus the handoff and its restored
407 * files. Logging both, per turn, is what makes the difference measurable
408 * instead of argued. Operator, 2026-09-15: "This tells us how much context is
409 * our compaction summary vs claude rules and similar."
410 *
411 * `percent` is computed here to one decimal rather than taken from the engine,
412 * which reports whole numbers; the engine's own figure is the fallback for a
413 * reading that carries no window to divide by.
414 */
415export const windowReading = ({ at, session, turn, context, messages, opener, arm = null, design = null }) => {
416    const tokens = context?.tokens ?? null;
417    const window = context?.window ?? null;
418    const share = (of) => (of === null || window === null || window === 0 ? null : Math.round((of / window) * 1000) / 10);
419    // The same four-characters-a-token estimate `approxTokens` uses, over a
420    // count rather than the text itself.
421    const handoffTokens = opener === null ? null : Math.ceil(opener.chars / 4);
422
423    return {
424        at,
425        session,
426        turn,
427        first: turn === 1,
428        phase: PHASE_OF[opener?.kind ?? "none"],
429        arm,
430        design,
431        compaction: opener?.n ?? 0,
432        tokens,
433        window,
434        percent: share(tokens) ?? context?.percent ?? null,
435        messages,
436        handoffChars: opener?.chars ?? null,
437        handoffTokens,
438        handoffPercent: share(handoffTokens),
439    };
440};
441
442/** A window's phase by the kind of compaction that opened it, if any. */
443const PHASE_OF = { none: "fresh", handoff: "post-compact", stock: "stock-compact" };
444
445/** A whole scratchpad, the common case: the model closed the tag. */
446const CLOSED_ANALYSIS = /<analysis>([\s\S]*?)<\/analysis>[ \t]*\n?/gu;
447
448/** The tags the summary itself is wrapped in, which carry nothing. */
449const SUMMARY_TAGS = /[ \t]*<\/?summary>[ \t]*\n?/gu;
450
451/**
452 * The summary without the thinking that produced it.
453 *
454 * Through 0.10.0 the fork was asked to write its inventory in <analysis> tags
455 * first and then copy it into the summary, and the block was dropped here. The
456 * prompt no longer asks for it, so this strip is a defensive one now: a model
457 * that writes the block anyway still has it removed. When one does, it is
458 * first-person deliberation rather than findings: it plans the summary, and it
459 * argues with itself and corrects mid-paragraph, which the next window reads
460 * as prose. Operator, 2026-09-15: "I think analysis just bloats it without much
461 * value add."
462 *
463 * Two shapes are deliberate. A block the model never closed is left alone
464 * unless a summary follows it, because a reply cut off inside the scratchpad
465 * has nothing else in it and dropping to the end would hand the next window an
466 * empty handoff. And anything outside the tags is kept, wherever it sits: a
467 * preamble before the analysis and a section written after </summary> are both
468 * content, and only the tags themselves are noise.
469 */
470export const withoutScratchpad = (text) => {
471    let analysisChars = 0;
472    let out = text.replace(CLOSED_ANALYSIS, (_whole, body) => {
473        analysisChars += body.length;
474
475        return "";
476    });
477
478    const open = out.indexOf("<analysis>");
479    const summary = out.indexOf("<summary>");
480
481    if (open !== -1 && summary > open) {
482        analysisChars += summary - open;
483        out = `${out.slice(0, open)}${out.slice(summary)}`;
484    }
485
486    const unwrapped = SUMMARY_TAGS.test(out);
487
488    return { text: out.replace(SUMMARY_TAGS, "").trim(), analysisChars, unwrapped };
489};
490
491/**
492 * One line of the lookup log: which tool read what back, for which session, and
493 * what it cost the window. Logged so a session can be charged for its history
494 * reads the way it is charged for its compactions.
495 */
496export const lookupRecord = ({ at, sessionId, tool, args, text }) => ({
497    at,
498    sessionId: sessionId ?? "unknown",
499    tool,
500    args: args ?? {},
501    chars: text.length,
502    lines: text === "" ? 0 : text.split("\n").length,
503    approxTokens: approxTokens(text),
504});
505
506/** The subsections a set of rows renders to: files written, failures, shell commands. */
507const ledgerBody = (rows, { full, n }) => {
508    const written = writersByFile(rows);
509    const shells = rows.filter((row) => row.name === "Bash");
510    const failed = rows.filter((row) => row.outcome !== "ok" && row.outcome !== "no result in transcript");
511    const out = [
512        `${count(rows.length, "tool call")} (${count(shells.length, "shell command")}), ${count(written.size, "file")} written.`,
513        "Read off the transcript rather than recalled. **The transcript stores no exit",
514        "code**, so the outcome of a command is what its output supports and no more.",
515        "",
516    ];
517
518    if (written.size > 0) {
519        out.push(`### Files written (${written.size})`, "");
520        out.push(...[...written].map(([path, tools]) => `- ${[...tools].join(", ")} \`${path}\``));
521        out.push("");
522    }
523
524    out.push(...failureLines(failed, full));
525
526    const fixed = out.join("\n").length;
527
528    if (shells.length > 0) {
529        const budget = Math.max(0, LEDGER_CHARS - fixed - SHELL_LIST_FRAME);
530
531        out.push(...(full ? everyShellLine(shells) : trimmedShellLines(shells, budget, n)));
532    }
533
534    return out;
535};
536
537const WRITE_TOOLS = new Set(["Write", "Edit", "NotebookEdit"]);
538
539/**
540 * Every file written, once, in the order first written, with each tool that
541 * wrote it. A file is counted once however many calls touched it: a count of
542 * calls beside a list of paths read as two numbers for one thing.
543 */
544const writersByFile = (rows) => {
545    const written = new Map();
546    const note = (path, tool) => written.set(path, (written.get(path) ?? new Set()).add(tool));
547
548    for (const row of rows) {
549        if (WRITE_TOOLS.has(row.name)) {
550            note(row.target, row.name);
551        }
552
553        for (const path of row.writes ?? []) {
554            note(path, "Bash");
555        }
556    }
557
558    return foldRelativePaths(written);
559};
560
561/**
562 * A shell command names a file relative to wherever it ran, and Write and Edit
563 * name the same file absolutely, so one file was listed twice. A relative path
564 * folds into the one absolute path it is the tail of; when none or several
565 * are, it stays as written, because the cwd it ran in is not on the row.
566 */
567const foldRelativePaths = (written) => {
568    const absolute = [...written.keys()].filter((path) => path.startsWith("/") || path.startsWith("~"));
569
570    for (const [path, tools] of written) {
571        const tail = `/${path.replace(/^\.\//u, "")}`;
572        const matches = absolute.filter((candidate) => candidate.endsWith(tail));
573
574        if (!absolute.includes(path) && matches.length === 1) {
575            tools.forEach((tool) => written.get(matches[0]).add(tool));
576            written.delete(path);
577        }
578    }
579
580    return written;
581};
582
583/** Room kept for the ledger's own heading, the list's heading and the count line. */
584const SHELL_LIST_FRAME = 200;
585
586const failureLines = (failed, full) => {
587    if (failed.length === 0) {
588        return [];
589    }
590
591    const shown = full ? failed : failed.slice(-LEDGER_FAILURES);
592    const limit = full ? LEDGER_ARG_CHARS : LEDGER_SHELL_CHARS;
593    const out = [`### Output that reads as a failure (${failed.length})`, ""];
594
595    if (shown.length < failed.length) {
596        out.push(`- ${failed.length - shown.length} older failures omitted`);
597    }
598
599    for (const row of shown) {
600        out.push(`- ${callLabel({ ...row, target: clipTo(row.target, limit) })}`);
601        out.push(`  - ${row.outcome}${row.detail === "" ? "" : `: ${row.detail}`}`);
602    }
603
604    return [...out, ""];
605};
606
607const shellLine = (row, limit) => `- \`${clipTo(row.target, limit)}\`${row.outcome === "ok" ? "" : ` — ${row.outcome}`}`;
608
609const everyShellLine = (shells) => [
610    `### Every shell command, in order (${shells.length})`,
611    "",
612    ...shells.map((row) => shellLine(row, LEDGER_ARG_CHARS)),
613    "",
614];
615
616/**
617 * The newest commands always, then older ones that changed something, newest
618 * first, while the budget lasts; shown in the order they ran.
619 */
620const trimmedShellLines = (shells, budget, n) => {
621    const recentFrom = Math.max(0, shells.length - LEDGER_RECENT_COMMANDS);
622    const kept = new Set();
623    let used = 0;
624
625    for (let index = shells.length - 1; index >= 0; index -= 1) {
626        const size = shellLine(shells[index], LEDGER_SHELL_CHARS).length + 1;
627        const isRecent = index >= recentFrom;
628
629        if (isRecent || (!shells[index].readOnly && used + size <= budget)) {
630            kept.add(index);
631            used += size;
632        }
633    }
634
635    const omitted = shells.filter((_, index) => !kept.has(index));
636    const out = [
637        kept.size === shells.length
638            ? `### Every shell command, in order (${shells.length})`
639            : `### Shell commands, in order (${kept.size} of ${shells.length})`,
640        "",
641        ...shells.flatMap((row, index) => (kept.has(index) ? [shellLine(row, LEDGER_SHELL_CHARS)] : [])),
642    ];
643
644    if (omitted.length > 0) {
645        const readOnly = omitted.filter((row) => row.readOnly).length;
646        const where = `handoff_lookup ${n === null ? "" : `n=${n} `}section=ledger`;
647
648        out.push(`- ${count(omitted.length, "older command")} omitted (${readOnly} read-only); ${where} lists every one.`);
649    }
650
651    return [...out, ""];
652};
653
654/**
655 * A failed call as the ledger lists it. A shell command is its own label; any
656 * other tool is named, with its target when it has one, so an EnterWorktree
657 * error does not render as an empty code span.
658 */
659const callLabel = (row) => {
660    if (row.target === "") {
661        return row.name;
662    }
663
664    return row.name === "Bash" ? `\`${row.target}\`` : `${row.name} \`${row.target}\``;
665};
666
667/** What the record supports about how a call ended, and nothing beyond it. */
668export const outcomeOf = (use) => {
669    if (use.isError === true) {
670        return { outcome: "errored", detail: clipTo(firstLine(use.text ?? ""), LEDGER_ERROR_CHARS) };
671    }
672
673    if (typeof use.text !== "string" || use.text === "") {
674        return { outcome: "no result in transcript", detail: "" };
675    }
676
677    const shaped = ERROR_SHAPE.exec(use.text.slice(0, ERROR_SCAN_CHARS));
678
679    if (shaped === null) {
680        return { outcome: "ok", detail: "" };
681    }
682
683    return { outcome: "unflagged, output reads as an error", detail: clipTo(shaped[0].trim(), LEDGER_ERROR_CHARS) };
684};
685
686/**
687 * Which tool was called, under whichever key this build spells it.
688 *
689 * `types/claude-code.d.ts` out of build 2.1.269 declares `ToolUseSummary` as
690 * `{id, name, input}` plus `{result, text, isError}`. Build 2.1.270 hands a hook
691 * `{tool_use_id, tool, input, result, text}` instead, so reading `use.name`
692 * alone yields undefined for every call. That is not hypothetical: the first
693 * live compaction rendered "2 tool calls: 0 file writes, 0 shell commands" over
694 * two Bash calls, and dropped the one that succeeded, because every row fell
695 * through to the placeholder. Read both, and let the declared name win if a
696 * later build restores it.
697 */
698export const nameOf = (use) => use.name ?? use.tool ?? "?";
699
700/** The tools that run a subagent, under either name the engine has used. */
701const SUBAGENT_TOOLS = new Set(["Agent", "Task"]);
702
703/**
704 * Whether a subagent ran since the last turn a person typed.
705 *
706 * Logged next to the fork's context because a subagent's own transcript is
707 * something the fork may or may not see, and a compaction landing right after
708 * one is the case worth being able to tell apart later.
709 */
710export const subagentRanThisTurn = (messages) => {
711    let start = 0;
712
713    for (let index = messages.length - 1; index >= 0; index -= 1) {
714        if (isPinnable(messages[index])) {
715            start = index;
716            break;
717        }
718    }
719
720    return messages
721        .slice(start)
722        .some((message) => (message.toolUses ?? []).some((use) => SUBAGENT_TOOLS.has(nameOf(use))));
723};
724
725/** What a call was aimed at, from whichever of its arguments names a target. */
726export const targetOf = (use) => {
727    const input = use.input ?? {};
728
729    for (const key of LEDGER_KEYS) {
730        const value = input[key];
731
732        if (typeof value === "string" && value.trim() !== "") {
733            return clipTo(value.replace(/\n/gu, " ; ").trim(), LEDGER_ARG_CHARS);
734        }
735    }
736
737    return "";
738};
739
740/* ------------------------------------------------------------------ *
741 * What a shell command did, read off its text.
742 * ------------------------------------------------------------------ */
743
744/** Has a directory separator, or ends in an extension. */
745export const PATH_SHAPED = /\/|\.[A-Za-z0-9]+$/u;
746
747/** Anything the shell would have rewritten before the file was opened. */
748export const SHELL_EXPANDED = /[$*?{}~]/u;
749
750/** Where one command in a compound ends: a newline, `&&`, `||`, `;` or a pipe. */
751const SEGMENT_BREAK = /\n|&&|\|\||;|\|/u;
752
753/** Prefixes that run the command after them unchanged. */
754const RUNS_THE_REST = /^(?:\w+=\S*\s+|sudo\s+(?:-\S+\s+)*(?:-u\s+\S+\s+)?|timeout\s+\S+\s+|time\s+|then\s+|do\s+|else\s+)+/u;
755
756/** Verbs that only look. Anything not here counts as having changed something. */
757const READ_ONLY =
758    /^(?:cd|ls|cat|head|tail|wc|grep|rg|find|echo|printf|pwd|stat|readlink|realpath|file|which|type|date|env|ps|pgrep|du|df|tree|sort|uniq|cut|awk|diff|jq|test|true|\[|sed|python3?\s+-c|git\s+(?:-C\s+\S+\s+)?(?:status|log|diff|show|rev-parse|fetch|ls-files|worktree\s+list|remote(?:\s+-v)?|branch\s+--show-current|config\s+--get)|gh\s+(?:pr|issue|run)\s+(?:view|list|checks|diff)|gh\s+api(?!.*-X\s*(?:POST|PATCH|PUT|DELETE)))(?:\s|$)/u;
759
760/** A redirect into a file: `>`, `>>`, `2>`, never `2>&1` or a `=>` arrow. */
761const REDIRECT = /(?:^|[^<>&=\-\d])\d?>>?[ \t]*([^\s;&|<>()'"`]+)/gu;
762
763/** `open("x", "w")` in an inline script. */
764const OPEN_FOR_WRITE = /open\(\s*(['"])([^'"]+)\1\s*,\s*(['"])[wax]/gu;
765
766/** A redirect that is only plumbing: the bit bucket, or a gate's own log. */
767const NOT_WORK = /^\/dev\/|\.log$/u;
768
769/** `sed -i`, which writes even when its target is a variable this cannot read. */
770const IN_PLACE = /^sed\s(?:.*\s)?-i/u;
771
772/** One-line quoted strings emptied, so a `|` or `>` inside one is not read as the shell's. */
773const blankQuotes = (command) => command.replace(/'[^'\n]*'|"[^"\n]*"/gu, "''");
774
775/** Every command in a compound, with the prefixes that just run it stripped. */
776const segmentsOf = (command) =>
777    blankQuotes(command)
778        .split(SEGMENT_BREAK)
779        .map((segment) => segment.trim().replace(RUNS_THE_REST, ""))
780        .filter((segment) => segment !== "");
781
782/** The unquoted, path-shaped words of one command after its verb. */
783const pathWords = (segment) =>
784    segment
785        .split(/\s+/u)
786        .slice(1)
787        .filter((word) => word !== "" && !word.startsWith("-") && PATH_SHAPED.test(word) && !SHELL_EXPANDED.test(word));
788
789/**
790 * The files a shell command wrote, as far as its text says: redirects, `tee`,
791 * `sed -i`, the destination of `mv` and `cp`, and `open(..., "w")` in an inline
792 * script. A path built from a variable cannot be read off the text and is left
793 * out, so this can miss a write but should not invent one.
794 */
795export const shellWrites = (command) => {
796    if (command === "") {
797        return [];
798    }
799
800    const found = [];
801
802    for (const match of blankQuotes(command).matchAll(REDIRECT)) {
803        found.push(match[1]);
804    }
805
806    for (const match of command.matchAll(OPEN_FOR_WRITE)) {
807        if (isInterpreted(command, match.index) && isRunAsCode(command, match.index)) {
808            found.push(match[2]);
809        }
810    }
811
812    for (const segment of segmentsOf(command)) {
813        const words = pathWords(segment);
814
815        if (/^tee\s/u.test(segment) || IN_PLACE.test(segment)) {
816            found.push(...words);
817        } else if (/^(?:mv|cp)\s/u.test(segment) && words.length >= 2) {
818            found.push(words[words.length - 1]);
819        }
820    }
821
822    return unique(found.filter((path) => PATH_SHAPED.test(path) && !SHELL_EXPANDED.test(path) && !NOT_WORK.test(path)));
823};
824
825/** A program that runs the script it is handed. */
826const INTERPRETER = /(?:^|[\s;&|({])(?:python[\d.]*|node|bun|deno|ruby|perl)(?:\s|$)/u;
827
828/** A heredoc opener, `<<'EOF'` or `<<-EOF`, capturing its delimiter. */
829const HEREDOC = /<<-?[ \t]*(['"]?)(\w+)\1/gu;
830
831/**
832 * Whether the text at `index` is handed to an interpreter at all. Inside a
833 * heredoc still open at that point, it is when the line that opened the heredoc
834 * runs one; outside, when its own line does. A PR body written with
835 * `cat <<'EOF'` that quotes `open('x','w')` is prose going into a file, and
836 * reading it as a write listed a file nothing wrote.
837 */
838const isInterpreted = (command, index) => {
839    const before = command.slice(0, index);
840    const opener = [...before.matchAll(HEREDOC)]
841        .filter((match) => !new RegExp(`\\n[ \\t]*${match[2]}[ \\t]*(?:\\n|$)`, "u").test(before.slice(match.index)))
842        .at(-1);
843    const upTo = opener === undefined ? before.length : opener.index;
844
845    return INTERPRETER.test(before.slice(before.lastIndexOf("\n", upTo - 1) + 1, upTo));
846};
847
848/** A flag whose argument is a script the interpreter runs: `python3 -c`, `node -e`. */
849const SCRIPT_FLAG = /(?:^|\s)(?:-c|-e|--eval)\s*$/u;
850
851/**
852 * Whether the text at `index` is code that runs rather than a string some code
853 * only holds. On its own line it is code when no quote is left open before it,
854 * or when the one left open is the argument of `-c` or `-e`. A script that
855 * writes test cases holds `open('x','w')` inside a string, and reading that as
856 * a write listed a file nothing wrote.
857 */
858const isRunAsCode = (command, index) => {
859    const line = command.slice(command.lastIndexOf("\n", index - 1) + 1, index);
860    const opened = openQuoteIn(line);
861
862    return opened === -1 || SCRIPT_FLAG.test(line.slice(0, opened));
863};
864
865/** Where the quote still open at the end of `line` begins, or -1 when every quote closed. */
866const openQuoteIn = (line) => {
867    let quote = "";
868    let opened = -1;
869
870    for (let index = 0; index < line.length; index += 1) {
871        const char = line[index];
872
873        if (char === "\\" && quote !== "") {
874            index += 1;
875        } else if (quote === "" && (char === "'" || char === '"')) {
876            quote = char;
877            opened = index;
878        } else if (char === quote) {
879            quote = "";
880            opened = -1;
881        }
882    }
883
884    return opened;
885};
886
887/** Whether every command in a compound only looked, and none of them wrote. */
888export const isReadOnlyCommand = (command) => {
889    const redirects = blankQuotes(command).replace(HARMLESS_REDIRECT, "");
890
891    if (shellWrites(command).length > 0 || />/u.test(redirects)) {
892        return false;
893    }
894
895    return segmentsOf(command).every((segment) => READ_ONLY.test(segment) && !IN_PLACE.test(segment));
896};
897
898/** `2>&1` and `>/dev/null`, which send output nowhere that lasts. */
899const HARMLESS_REDIRECT = /\d?>&\d|\d?>\s*\/dev\/null/gu;
900
901/* ------------------------------------------------------------------ *
902 * The commitments pass: what goes in, and what comes back out.
903 * ------------------------------------------------------------------ */
904
905/**
906 * The assistant's own words, thinking included, newest kept when space is short.
907 *
908 * The last turns before a conversation ends carry the most unkept commitments,
909 * because they had the least time to be acted on, so the prompt is trimmed from
910 * the front and the trim is stated in it rather than hidden.
911 */
912export const assistantTurns = (messages) => {
913    const turns = [];
914
915    messages.forEach((message, index) => {
916        if (message.role !== "assistant" || (message.text ?? "").trim() === "") {
917            return;
918        }
919
920        turns.push(`===== turn ${index} =====\n${clipTo(message.text.trim(), COMMITMENT_TURN_CHARS)}`);
921    });
922
923    let kept = turns;
924    let size = kept.join("\n\n").length;
925
926    while (size > COMMITMENT_PROMPT_CHARS && kept.length > 1) {
927        kept = kept.slice(1);
928        size = kept.join("\n\n").length;
929    }
930
931    if (kept.length < turns.length) {
932        kept = [`===== ${turns.length - kept.length} earlier turns omitted for length =====`, ...kept];
933    }
934
935    return kept.join("\n\n");
936};
937
938/** What was actually run, one line each. No results: the absence is the point. */
939export const toolIndex = (messages) => {
940    const rows = [];
941
942    messages.forEach((message, index) => {
943        for (const use of message.toolUses ?? []) {
944            rows.push(`T${index} | ${nameOf(use)} | ${targetOf(use)}`);
945        }
946    });
947
948    return rows.join("\n");
949};
950
951/** The most a handoff may be, in tokens, whatever the window: past this the next window is mostly handoff. */
952export const SAFE_CAP_TOKENS = 200_000;
953export const DEFAULT_CAP_TOKENS = 150_000;
954export const DEFAULT_FRACTION = 0.25;
955export const DEFAULT_MAX_CHARS = 400_000;
956const CHARS_PER_TOKEN = 4;
957
958const usable = (value) => typeof value === "number" && Number.isFinite(value) && value > 0;
959
960/**
961 * How large a handoff may be, and why: min(fraction × window, capTokens) × 4
962 * characters. A hand-set override wins outright and is never clamped, because
963 * setting it was a deliberate act; the row says by how much it passed the safe
964 * cap instead. The cap setting is a default, so it is clamped.
965 */
966export const handoffCeiling = ({ window, override, fraction, capTokens } = {}) => {
967    const knownWindow = usable(window) ? window : null;
968
969    if (usable(override)) {
970        const over = override - SAFE_CAP_TOKENS * CHARS_PER_TOKEN;
971        const overrun = over > 0
972            ? { overSafeCapChars: over, overSafeCapTokens: Math.ceil(over / CHARS_PER_TOKEN) }
973            : {};
974
975        return { window: knownWindow, fraction: null, capTokens: null, chars: override, decider: "override", ...overrun };
976    }
977
978    const usedFraction = usable(fraction) && fraction <= 1 ? fraction : DEFAULT_FRACTION;
979    const usedCap = Math.min(usable(capTokens) ? capTokens : DEFAULT_CAP_TOKENS, SAFE_CAP_TOKENS);
980    const base = { window: knownWindow, fraction: usedFraction, capTokens: usedCap };
981
982    if (knownWindow === null) {
983        const chars = Math.min(DEFAULT_MAX_CHARS, usedCap * CHARS_PER_TOKEN);
984
985        return { ...base, chars, decider: "default: window unknown" };
986    }
987
988    const fromWindow = Math.floor(knownWindow * usedFraction);
989    const tokens = Math.min(fromWindow, usedCap);
990
991    return { ...base, chars: tokens * CHARS_PER_TOKEN, decider: tokens === fromWindow ? "fraction" : "cap" };
992};
993
994/** A row the pass started and did not finish: it names a kind, and stops before the shape closes. */
995const ROW_OPENING = /^\s*(UNKEPT|CORRECTED|UNANSWERED)\s*\|/iu;
996
997/**
998 * Whether the commitments reply stopped at its output cap, and how many rows
999 * came back. Near the ceiling in length, or a final row cut mid-shape, both
1000 * count; a trailing line of prose does not.
1001 */
1002export const commitmentsOutcome = (reply, maxTokens) => {
1003    const text = reply ?? "";
1004    const rows = commitmentsFrom(text).length;
1005    const lines = text.split("\n").map((line) => line.trim()).filter((line) => line !== "");
1006    const last = lines.at(-1) ?? "";
1007    const nearCeiling = text.length >= maxTokens * CHARS_PER_TOKEN * 0.9;
1008    const truncatedRow = ROW_OPENING.test(last) && !COMMITMENT_ROW.test(last);
1009    const hitCapReason = nearCeiling ? "length" : truncatedRow ? "truncated row" : null;
1010
1011    return { rows, hitCap: hitCapReason !== null, hitCapReason };
1012};
1013
1014/** The rows the pass emitted, in the order it emitted them. */
1015export const commitmentsFrom = (reply) => {
1016    const found = [];
1017
1018    for (const line of (reply ?? "").split("\n")) {
1019        const match = COMMITMENT_ROW.exec(line);
1020
1021        if (match === null) {
1022            continue;
1023        }
1024
1025        const [, kind, turn, quote, note] = match;
1026
1027        found.push({
1028            kind: kind.toLowerCase(),
1029            turn: Number(turn),
1030            quote: quote.trim().replace(/^"|"$/gu, "").trim(),
1031            note: note.trim(),
1032        });
1033    }
1034
1035    return dedupeFindings(found.filter((row) => row.quote.toLowerCase() !== "none"));
1036};
1037
1038/**
1039 * Which kind wins when the pass files one sentence under two. A question the
1040 * user never answered is also, loosely, something the assistant still owes,
1041 * and a live 0.11.0 handoff listed it both ways; the narrower kind says more.
1042 */
1043const KIND_RANK = { corrected: 0, unanswered: 1, unkept: 2 };
1044
1045/** @param {string} quote */
1046const quoteKey = (quote) => quote.toLowerCase().replace(/\s+/gu, " ").replace(/[\s.,;:!?]+$/u, "").trim();
1047
1048/** One row per quoted sentence, in first-seen order, under its most specific kind. */
1049const dedupeFindings = (found) => {
1050    const kept = new Map();
1051
1052    for (const row of found) {
1053        const key = quoteKey(row.quote);
1054        const prior = kept.get(key);
1055
1056        if (prior === undefined || KIND_RANK[row.kind] < KIND_RANK[prior.kind]) {
1057            kept.set(key, row);
1058        }
1059    }
1060
1061    return [...kept.values()];
1062};
1063
1064export const renderCommitments = (findings) => {
1065    if (findings.length === 0) {
1066        return "";
1067    }
1068
1069    const out = [
1070        "## Commitments and open questions",
1071        "",
1072        "Written by one model pass over the assistant's own turns and an index of what",
1073        "it ran. Unlike the two sections above this is a judgement rather than a",
1074        "reading: each item claims something was promised and that no evidence of it",
1075        "appears later, which is an absence and cannot be proved from a conversation",
1076        "that was cut off mid-flight. Check before acting, and do not re-do work.",
1077        "",
1078    ];
1079
1080    for (const kind of ["unkept", "corrected", "unanswered"]) {
1081        const rows = findings.filter((row) => row.kind === kind);
1082
1083        if (rows.length === 0) {
1084            continue;
1085        }
1086
1087        out.push(`### ${COMMITMENT_HEADING[kind]} (${rows.length})`, "");
1088
1089        for (const row of rows) {
1090            out.push(`- "${row.quote}"`);
1091            out.push(`  - ${row.note}`);
1092        }
1093
1094        out.push("");
1095    }
1096
1097    return out.join("\n").trimEnd();
1098};
1099
1100/* ------------------------------------------------------------------ *
1101 * What a compaction costs.
1102 * ------------------------------------------------------------------ */
1103
1104/**
1105 * List prices per million tokens, read off
1106 * https://platform.claude.com/docs/en/about-claude/pricing on 2026-09-14.
1107 * `cacheWrite` is the 5-minute tier, which is what a fork and a `claude -p`
1108 * pass actually use.
1109 *
1110 * **Matched most specific first, and a family name is never enough.** Sonnet 5
1111 * is $2/$10 and Sonnet 4.6 is $3/$15; Opus 5 is $5/$25 and the retired Opus 4.1
1112 * is $15/$75. A table keyed on the word "sonnet" would have overstated every
1113 * compaction on this box by 50%, and the first draft of this file did exactly
1114 * that from memory. The numbers below were read off the page, not recalled.
1115 *
1116 * A row whose model matches nothing records `cost: null` with a reason.
1117 * Pricing a token at zero because the table is stale is worse than admitting
1118 * the number is unknown: one is a gap, the other is a wrong number in a
1119 * document the operator is going to compare against their bill.
1120 */
1121/*
1122 * Dollars per million tokens, read on 2026-09-14 from the "Model pricing" table at
1123 * https://platform.claude.com/docs/en/about-claude/pricing
1124 *
1125 * Every row below was read off that table, not recalled: all four numbers for all
1126 * seven rows were checked against it on 2026-09-14. Nothing here is from memory.
1127 * `cacheWrite` is the 5-minute write column, which is what a fork and a
1128 * `claude -p` actually pay; the 1-hour column is not modelled because nothing
1129 * here asks for a 1-hour cache.
1130 *
1131 * The Fable/Mythos split is not cosmetic. Cache hits are 0.1x base input on every
1132 * model EXCEPT Fable 5.1 and Mythos 5.1, which the page prices at 0.025x, so one
1133 * pattern over both generations under-charged a Fable 5 or Mythos 5 cache read by
1134 * four times. The specific row has to come first; `priceRowFor` takes the first
1135 * match.
1136 *
1137 * A wrong number here is silent: it mis-states every cost row and the row still
1138 * looks like a measurement. Re-read the page before changing PRICES_TAKEN, and
1139 * change them together.
1140 *
1141 * Cache reads are priced here and never charged. On a subscription a cache
1142 * read costs nothing, so `priceUsage` keeps the token count and records what
1143 * the reads would have cost at list as `cacheReadWaivedUsd`, and leaves them
1144 * out of `usd`. Operator, 2026-09-15: "Cache reads are FREE for subscriptions."
1145 */
1146export const COST_BASIS = "subscription: cache reads free";
1147
1148export const PRICES_TAKEN = "2026-09-14";
1149
1150export const PRICES_SOURCE = "https://platform.claude.com/docs/en/about-claude/pricing";
1151
1152export const PRICES = [
1153    { match: /fable-5-1|mythos-5-1|fable5-1|mythos5-1/u, name: "Fable/Mythos 5.1", input: 10, output: 50, cacheRead: 0.25, cacheWrite: 12.5 },
1154    { match: /fable|mythos/u, name: "Fable/Mythos 5", input: 10, output: 50, cacheRead: 1, cacheWrite: 12.5 },
1155    { match: /opus-4-1|opus-4(?!\d)/u, name: "Opus 4.1", input: 15, output: 75, cacheRead: 1.5, cacheWrite: 18.75 },
1156    { match: /opus/u, name: "Opus 5 / 4.5-4.8", input: 5, output: 25, cacheRead: 0.5, cacheWrite: 6.25 },
1157    { match: /sonnet-5|sonnet5/u, name: "Sonnet 5", input: 2, output: 10, cacheRead: 0.2, cacheWrite: 2.5 },
1158    { match: /sonnet/u, name: "Sonnet 4.6 and earlier", input: 3, output: 15, cacheRead: 0.3, cacheWrite: 3.75 },
1159    { match: /haiku-3|haiku3/u, name: "Haiku 3.5", input: 0.8, output: 4, cacheRead: 0.08, cacheWrite: 1 },
1160    { match: /haiku/u, name: "Haiku 4.5", input: 1, output: 5, cacheRead: 0.1, cacheWrite: 1.25 },
1161];
1162
1163/** Which price row a model id falls under, or null if the table cannot say. */
1164export const priceRowFor = (model) => {
1165    const id = (model ?? "").toLowerCase();
1166
1167    if (id === "") {
1168        return null;
1169    }
1170
1171    return PRICES.find((row) => row.match.test(id)) ?? null;
1172};
1173
1174/** The four token counts a usage record carries, under either spelling. */
1175export const usageOf = (usage) => {
1176    if (usage === null || usage === undefined || typeof usage !== "object") {
1177        return null;
1178    }
1179
1180    const pick = (...keys) => {
1181        for (const key of keys) {
1182            if (typeof usage[key] === "number") {
1183                return usage[key];
1184            }
1185        }
1186
1187        return 0;
1188    };
1189
1190    return {
1191        input: pick("input_tokens", "inputTokens"),
1192        output: pick("output_tokens", "outputTokens"),
1193        cacheRead: pick("cache_read_input_tokens", "cacheReadInputTokens"),
1194        cacheWrite: pick("cache_creation_input_tokens", "cacheCreationInputTokens"),
1195    };
1196};
1197
1198/**
1199 * What a fork's usage cost, or null with the reason it cannot be said.
1200 *
types/compact-handoff.d.ts 57 lines
1// The type contract of `$.compactHandoff`, the noun this plugin's hooks module
2// adds at `engine.create`.
3//
4// `plugin.json`'s `types` field points here. `/plugin-types` copies this file
5// to `.claude/types/claude-code-plugins/compact-handoff.d.ts` and indexes it in
6// `claude-code-plugins.d.ts`, so a plugin that subscribes to the seam types
7// against the real shape instead of a README. `claude plugin validate` checks
8// it. Self-contained by rule: no import, export-from, require or reference.
9
10/**
11 * What `beforeCompact` takes: the tool this plugin raises when a compaction is
12 * about to happen, beside its own fork.
13 *
14 * The seam carries strings and nothing else. Every plugin runs in its own
15 * environment and an interface call's arguments cross through `cloneInto`,
16 * which throws `DataCloneError` on a function, so a subscriber names a TOOL it
17 * answers rather than handing over a callback.
18 */
19export interface CompactHandoffSubscribeOptions {
20    /** The full name of the tool to raise, as `$.tool.call` spells it. */
21    tool: string;
22    /** What to call the subscriber in this plugin's own records; the tool name when omitted. */
23    name?: string;
24}
25
26/** What `beforeCompact` resolves: the subscription, and the tool it is keyed by. */
27export interface CompactHandoffSubscribeResult {
28    subscribed: true;
29    tool: string;
30}
31
32/**
33 * The noun `$.compactHandoff`: how another plugin runs its own work beside this
34 * plugin's compaction fork.
35 *
36 * It exists because a plugin keyed after this one in `enabledPlugins` never
37 * sees `session.compact` at all — this plugin answers that event without
38 * calling `next`. There is no unsubscribe, and a subscription lives for the
39 * load; subscribing twice with one tool is once.
40 */
41export interface CompactHandoff {
42    /**
43     * Registers the tool raised when a compaction is about to happen.
44     *
45     * Rejects with a `TypeError` when `tool` is missing or blank.
46     */
47    beforeCompact: (options: CompactHandoffSubscribeOptions) => Promise<CompactHandoffSubscribeResult>;
48    /** This plugin's version, for a subscriber that wants to know what it got. */
49    version: () => Promise<string>;
50}
51
52declare module 'claude-code' {
53    interface EngineInterface {
54        compactHandoff: CompactHandoff;
55    }
56}
57