Diagnoses wrong processes, workflows and habits in Claude Code sessions: fires on plan acceptance, workflow launches, pace complaints and the clock, names the…

<h1><img src="assets/logo.png" alt="" height="40" align="top"> ContextSaver</h1>
Diagnoses wrong processes, wrong workflows and wrong habits in Claude Code sessions — and lets you fix them in one click.
Sessions bog down when the process is wrong. Review and fix rounds per lane, stacked on a whole-branch review that already covers them. The premium model on mechanical stages. Agents spawned for jobs a single command would do. The full check suite after every merge. Tool-level waste — the same file read again, the full suite after a one-line edit — compounds on top. ContextSaver catches both: the wrong process at the moments that matter, and the wrong habit as it builds. It gives you the measured cost and one click to change it.
<img src="docs/screenshot.png" width="880" alt="The ContextSaver pane beside the transcript, showing two cards with their cost and the Fix, Fix… and Ignore buttons">
Write.[!NOTE] Needs Claude Code 2.1.273 or newer. ContextSaver is a Claude Mod, built on function hooks, which are early access — so enable the flag first.
export CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 # your shell, or ~/.claude/settings.json under "env"
Then, in Claude Code:
/plugin marketplace add AlmogBaku/ContextSaver
/plugin install contextsaver@contextsaver
That is the whole setup: no config, no API key, no dependencies, no build step. Installing mid-session works too — the plugin reads what already happened out of the transcript and checks it right away, without waiting for your next prompt.
i for the calls behind it.Fix sends the suggested fix, Fix… lets you reword it first, Ignore drops it for the session.Write saves it, Try uses it for this session only, Skip drops it.[!TIP] The pane docks beside your transcript in a wide terminal and sits above the prompt otherwise.
/savertoggles it at any width,ctrl+x tabmoves the keyboard into it,Tabcycles the buttons,Enterpresses, andEschands the keys back.
| Command | What it does |
|---|---|
/saver | Show or hide the pane. |
/saver check | Check now, instead of waiting for the next automatic check. |
/saver fix <n> | Send card n's suggested fix. |
/saver fix [n] <text> | Send your own instruction instead. Without a number: the open card, else card 1. |
/saver ignore <n> | Drop card n for the rest of the session. |
/saver debug | Print the session state: ledger, findings, decisions, what the audit cost, savings. |
/saver reset | Clear this session's ledger and decisions. Learned patterns survive. |
Every tool call becomes a row in a session ledger: what ran, how long it took, how much it added to context, which files it touched. The plugin runs two judges.
The process judge fires at moments that matter — an accepted plan, a Workflow launch, a pace complaint you type, every thirty minutes. It builds a digest of what the user asked, how the work was organised, what models ran which roles, how many merges got a check, what was re-read and what was slow. From that, it asks: given what the user asked, how would a lean expert run this work, and where does this session diverge? At most three findings, each with a measured cost.
The habit judge runs continuously — about every 30k tokens and three turns, or every 40 calls and five minutes inside a long turn. It asks what has repeated and bogged things down, and whether there was a shorter path. Whatever it finds becomes a card, with the costs computed from the ledger.
Fix on a process card sends a one-time re-plan to the main loop. Fix on a habit card sends a standing instruction that rides every subsequent prompt and is also appended to every new subagent's prompt — the card shows "sent ×N" to confirm it landed. Neither ever blocks a tool.
The architecture, the judge's prompt and the design brief are in docs/SPEC.md; the product spec is docs/PRD.md.
[!IMPORTANT] The audit runs on your session's model, so a session on Opus pays Opus for it. It keeps itself to a few percent of the session's tokens, and
/saver debugshows exactly what it spent.
✕ Last turn ended in … · type anything to continue until your next prompt; while a workflow runs and nothing is found it names the run, its stage, its agents and their calls.git clone https://github.com/AlmogBaku/ContextSaver && cd ContextSaver
./scripts/check.sh # validate --strict, typecheck, tests
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir . # run Claude Code with this folder loaded
All logic is pure functions over a single State in hooks/core/. hooks/register.ts is the only file that touches events, and every hook falls back to next(e) on any path it does not own. hooks/ui.tsx renders two view models and never reads State. With CONTEXTSAVER_DEBUG=1, /saver demo fills the pane with sample cards, so the drawing can be worked on without waiting for a real finding.
Under --plugin-dir, editing a file hot-reloads the plugin and resets session state. Learned patterns persist in the plugin store.
MIT © Almog Baku. See LICENSE.
hooks/register.ts 900 lines1import type { ModelForkResult, On, PaneOpenArgs, RenderElement } from 'claude-code'
2
3import { adoptRows } from './core/adopt'
4import { demoForkUsage, demoPatterns, demoRows, demoTurns, demoUsage } from './core/demo'
5import { buildPrompt, judgeAliases, merge, parseReply, shouldRun, spentOf, usageOf } from './core/judge'
6import { rowOf } from './core/ledger'
7import type { ToolEvent } from './core/ledger'
8import { bandModel, debugDump, fromStored, mergeStored, paneModel, parseRegistry, processLine, reduce, storableOf, usageLine } from './core/patterns'
9import { buildProcessPrompt, isPace, mergeProcess, parseProcessReply, processAliases, shouldProcess } from './core/process'
10import { appendedTo, bulletOnly, mergeSettings, propose } from './core/rules'
11import { activeRuns, agentOf, journalPath, parseJournal, phasesOf, runOf } from './core/spawns'
12import { collapseWs, duration, fit, instructionOf, isRecord, pctOf } from './core/text'
13import {
14 ASK_HEAD_MAX, AUTO_OPEN_MIN_COLUMNS, CLAUDE_MD_HEADING, COMMAND, DEBUG_MAX_DROPPED, DELIVERY_BURST_MS, JUDGE_MIN_ROWS, MAX_PATTERNS, PANE_ID,
15 PANE_INLINE_ROWS, PANE_TITLE, PLUGIN_NAME, RUN_REFRESH_MS, STEER_RING_TRIES, STEER_RING_WAIT_MS, initialState, storeKey,
16} from './core/types'
17import type { Action, Actions, Artifact, Choice, ProcessTrigger, Run, State, Tokens, Ui } from './core/types'
18import type { Host } from './host'
19import { Band, Pane } from './ui'
20
21const FIX_USAGE = 'Usage: /saver fix [n] [instruction] (a leading number is the card the pane draws; without one: the card whose Fix… field is open, else card 1)'
22const SAVER_USAGE = 'Usage: /saver [check | fix [n] [text] | ignore <n> | debug | reset]'
23const NOTHING_TEXT = 'ContextSaver: nothing to decide on'
24const CHECKING_TEXT = 'ContextSaver: checking this session for waste…'
25const ALREADY_TEXT = 'ContextSaver: already checking'
26const ANSWER_HEAD = 100 // characters of the turn's answer kept as an evidence quote
27const CARD_KIND = 60 // characters of a card's behaviour quoted back in a command's reply
28const PLAN_TOOL = 'ExitPlanMode' // its accepted result is a plan the process judge can weigh
29const TYPED: readonly string[] = ['composer', 'bridge'] // the origins a person typed: only these count as asks
30const DEMO_CONTEXT = [120_000, 190_000, 250_000, 320_000] // `/saver demo`: the window filling up to the sample's own 32%, so the trend draws
31
32// What set one judge run going, as the debug log names it: the mid-turn cadence, the turn's end, the
33// person, or a check armed at load over a transcript this plugin joined late.
34type JudgeReason = 'tool.call' | 'turn.complete' | '/saver check' | 'load'
35
36// The lanes the run answers out loud: the check the person typed, and the one the load armed for them.
37const REQUESTED: readonly JudgeReason[] = ['/saver check', 'load']
38
39/**
40 * Registers ContextSaver: the ledger of every tool call, the judge that names wasteful
41 * behaviours, the band and the pane that let the user fix or ignore them.
42 *
43 * @param on the engine's registrar
44 */
45export function register(on: On): void {
46 let state: State = initialState('', 0)
47 let host: Host | null = null
48 let isDebug = false
49 // `judge.start` lands one clock read after the decision to run, and a storm of tool calls decides
50 // inside that window: this flag is what stops a second fork of the same session.
51 let forking = false
52 // A check asked for while a run is in flight is answered by that run: cadence runs are frequent now,
53 // and the person who pressed Check now would otherwise be told `already checking` and never told more.
54 let asked = false
55 // The load lane's one failure toast. Its arming survives every failure, so a toast per retry would be a
56 // storm — but total silence reads exactly like a check that never fired, so the first failure speaks.
57 let armedSpoke = false
58 // The process judge's own lock, beside `forking`: `process.start` lands one clock read after the decision.
59 let processing = false
60 // The loops a journal was re-read for once already: an id no journal names after that is the engine's own.
61 const probed = new Set<string>()
62
63 const messageOf = (err: unknown): string => (err instanceof Error ? err.message : String(err))
64
65 // `focus` is a request, not a grant (d.ts 4915-4922): the surface hands the pane the keyboard only while
66 // the prompt holds them over an empty composer, and refuses it otherwise — the pane opens either way.
67 const paneArgs = (focus?: true): PaneOpenArgs =>
68 focus === undefined
69 ? { id: PANE_ID, title: PANE_TITLE, rows: PANE_INLINE_ROWS }
70 : { id: PANE_ID, title: PANE_TITLE, rows: PANE_INLINE_ROWS, focus }
71
72 const savedToast = (before: State['saved']): void => {
73 const ms = state.saved.ms - before.ms
74 const pct = pctOf(state.saved.chars - before.chars, state.usage.window)
75 const grew = [...(ms > 0 ? [`+${duration(ms)}`] : []), ...(pct > 0 ? [`+~${pct}%`] : [])]
76 if (grew.length === 0) return
77 host?.toast(`${grew.join(' · ')} context saved`)
78 }
79
80 const openIds = (): string[] => state.patterns.filter(p => p.openedAtTurn !== null).map(p => p.id)
81
82 // Outside a render hook: fold the action in, redraw, and credit an instruction that settled in it.
83 const dispatch = (action: Action): void => {
84 const before = state.saved
85 const open = openIds()
86 state = reduce(state, action)
87 host?.invalidate()
88 // Only a settled instruction is a saving to announce; the per-turn accrual behind it stays quiet.
89 if (open.some(id => state.patterns.find(p => p.id === id)?.openedAtTurn === null)) savedToast(before)
90 }
91
92 // A `/clear`, a `/saver reset` or a resume starts the session over: the closure flags that belong to
93 // the session go with its state, so a new one may speak for its armed check again.
94 const resetSession = (): void => {
95 dispatch({ type: 'reset' })
96 armedSpoke = false
97 probed.clear()
98 }
99
100 // Inside a render hook: fold the action in with no redraw, since a redraw loops.
101 const observe = (action: Action): void => {
102 state = reduce(state, action)
103 }
104
105 const persist = (): void => {
106 const engine = host
107 if (engine === null) return
108 const key = storeKey(state.cwd)
109 const mine = storableOf(state)
110 void engine
111 .storeGet(key)
112 .then(value => engine.storeSet(key, mergeStored(parseRegistry(value), mine)))
113 .catch(() => undefined)
114 }
115
116 // An open the person asked for asks for their keyboard too, so the pane they just called up is the pane
117 // they can type in; an unasked one interrupts whatever they were doing and never asks.
118 const openPane = async (auto?: true): Promise<void> => {
119 const engine = host
120 if (engine === null) return
121 await engine.openPane(auto === undefined ? paneArgs(true) : paneArgs())
122 dispatch(auto === undefined ? { type: 'pane', open: true } : { type: 'pane', open: true, auto })
123 }
124
125 // Only after fresh cards arrived, once a session, and only where the surface would draw it.
126 const autoOpen = async (fresh: readonly string[]): Promise<void> => {
127 const queued = fresh.some(id => state.cards.includes(id))
128 if (!queued || state.paneOpen || state.autoOpened || (state.columns ?? 0) < AUTO_OPEN_MIN_COLUMNS) return
129 try {
130 await openPane(true)
131 } catch {
132 // an unasked open the surface or another plugin refused is no error of the user's
133 }
134 }
135
136 // A check the person asked for: they are waiting for the answer, so the pane opens at any width, every time.
137 const openForCheck = async (queued: readonly string[]): Promise<void> => {
138 if (queued.length === 0 || state.paneOpen) return
139 try {
140 await openPane()
141 } catch {
142 // the surface or another plugin refused: the band still says what was found
143 }
144 }
145
146 // The four counts a turn was billed, zero where no response came back to bill.
147 const tokensOf = (u: { input_tokens: number; output_tokens: number; cache_read_input_tokens: number; cache_creation_input_tokens: number } | undefined): Tokens =>
148 ({ input: u?.input_tokens ?? 0, output: u?.output_tokens ?? 0, cacheRead: u?.cache_read_input_tokens ?? 0, cacheCreate: u?.cache_creation_input_tokens ?? 0 })
149
150 // An `Agent` result names its loop, but the description it was given is the call's own argument, so the
151 // value is handed over with it; every other tool's result is read as it came.
152 const spawnValue = (e: ToolEvent, value: unknown): unknown =>
153 e.tool === 'Agent' && typeof e.description === 'string' && typeof value === 'object' && value !== null
154 ? { ...value, description: e.description }
155 : value
156
157 // One run's journal, read and folded in; a read that fails still stamps the run, so it is not retried at once.
158 const readJournal = async (engine: Host, run: Run, now: number): Promise<void> => {
159 const path = journalPath(run)
160 if (path === null) return
161 const entries = await engine.readFile(path).then(parseJournal).catch(() => [])
162 dispatch({ type: 'run.journal', runId: run.id, entries, now })
163 }
164
165 // The journals of the runs still going, at most every RUN_REFRESH_MS each — every run's when forced: a
166 // launch reads its own at once, and the judge reads them all before it asks. Never awaited from a hook.
167 const refreshRuns = async (force: boolean): Promise<void> => {
168 const engine = host
169 if (engine === null || state.runs.length === 0) return
170 const now = await engine.now()
171 const due = force ? state.runs : activeRuns(state, now).filter(run => now - run.refreshedAt >= RUN_REFRESH_MS)
172 await Promise.all(due.map(run => readJournal(engine, run, now)))
173 }
174
175 // What the run put in front of the user: a recurrence is news too, and it is the D4 moment this exists for.
176 const checkedText = (queued: number): string =>
177 queued === 0 ? 'ContextSaver: nothing new' : `ContextSaver: ${queued} new waster${queued === 1 ? '' : 's'}`
178
179 // A run that reported nothing: a cold snapshot, a refusal, or a failure of ours.
180 const judgedNothing = (error: string): Action => ({
181 type: 'judge.done', patterns: state.patterns, fresh: [], recurred: [], focus: null, time: null, context: null, spent: 0, error,
182 returned: 0, kept: 0, dropped: [], usage: null,
183 })
184
185 // A run the person asked for answers them, whatever it found: silence is what a check must never be.
186 // An armed check nobody asked for says it once, in its own words, since it will be retried: the first
187 // failure names itself and the retries stay quiet.
188 const failedToast = (reason: JudgeReason, failure: string): void => {
189 if (reason === 'load' && !asked) {
190 if (armedSpoke) return
191 armedSpoke = true
192 host?.toast(`ContextSaver: could not check yet — ${failure}`)
193 return
194 }
195 if (REQUESTED.includes(reason) || asked) host?.toast(`ContextSaver: check failed — ${failure}`)
196 }
197
198 const judgeOnce = async (engine: Host, reason: JudgeReason): Promise<void> => {
199 const requested = REQUESTED.includes(reason)
200 const seq = state.seq
201 const now = await engine.now()
202 dispatch({ type: 'judge.start', now, seq })
203 // AGENTS is read off the journals, so they are brought up to date once, here, before the prompt is built.
204 await refreshRuns(true).catch(() => undefined)
205 // One alias table for the run: the loops keep spawning while the fork thinks, and a new agent's first
206 // row would renumber the `agent:aN` handles the reply cites against the AGENTS the prompt printed.
207 const aliases = judgeAliases(state)
208 let reply: ModelForkResult | null = null
209 let failed: string | null = null
210 try {
211 reply = await engine.fork(buildPrompt(state, aliases))
212 } catch (err) {
213 failed = messageOf(err)
214 }
215 if (reply === null) {
216 dispatch(judgedNothing(failed ?? 'cold snapshot'))
217 failedToast(reason, failed ?? 'cold snapshot')
218 return
219 }
220 try {
221 const { findings, focus, time, context, dropped, returned } = parseReply(reply.text, state, aliases)
222 const merged = merge(state, findings)
223 // A finding the registry cap evicted never becomes a card, so it is dropped, not kept.
224 const reasons = [...dropped, ...merged.evicted.map(id => `${id}: evicted, over MAX_PATTERNS (${MAX_PATTERNS})`)]
225 const kept = findings.length - merged.evicted.length
226 const usage = usageOf(reply.usage)
227 dispatch({
228 type: 'judge.done', patterns: merged.patterns, fresh: merged.fresh, recurred: merged.recurred, focus,
229 time, context, spent: spentOf(usage), error: null, returned, kept, dropped: reasons, usage,
230 })
231 try {
232 if (isDebug) {
233 engine.log(`ContextSaver judge: ${returned} returned · ${kept} kept · ${reasons.length} dropped · from ${reason}`)
234 for (const line of reasons.slice(0, DEBUG_MAX_DROPPED)) engine.log(line)
235 // A cold cache is what makes a run expensive, and only the four counts say which it was.
236 engine.log(usageLine(usage))
237 }
238 } catch {
239 // a log we could not write is not a failed run: the findings are already in the registry
240 }
241 persist()
242 // Every card this run queued: a fresh finding, or a steered behaviour that came back.
243 const queued = [...merged.fresh, ...merged.recurred].filter(id => state.cards.includes(id))
244 // A check asked for mid-run is answered by the run it arrived in, whichever run that was.
245 if (!(requested || asked)) return autoOpen(queued)
246 engine.toast(checkedText(queued.length))
247 return openForCheck(queued)
248 } catch (err) {
249 // Whatever went wrong, the run is over: `running` may never stay true.
250 dispatch(judgedNothing(messageOf(err)))
251 failedToast(reason, messageOf(err))
252 }
253 }
254
255 /**
256 * Judges the session once, if no run is already in flight.
257 *
258 * @param reason what set this run going: the mid-turn cadence, the turn's end, the person, or the
259 * check a load armed. The two REQUESTED lanes are answered with a toast and open the pane wherever
260 * it can be drawn; a cadence run stays quiet and keeps the once-a-session, wide-terminal rule for
261 * opening itself — unless someone asks while it is in flight, in which case that run answers them.
262 */
263 async function runJudge(reason: JudgeReason): Promise<void> {
264 const engine = host
265 if (engine === null || forking || state.judge.running) return
266 forking = true
267 try {
268 await judgeOnce(engine, reason)
269 } finally {
270 forking = false
271 asked = false
272 }
273 }
274
275 // A check armed at load consults no gate — `shouldRun` counts new work, and a session joined late has
276 // all of its work behind it. The load fires this itself; every opportunity after it is a retry of a run
277 // that came back with nothing. Answers whether it took this opportunity.
278 const armedCheck = (): boolean => {
279 if (!state.pendingCheck || state.judge.running) return false
280 void runJudge('load').catch(() => undefined)
281 return true
282 }
283
284 // The opportunities a run can start at: the armed check first, else the cadence's own count of new work.
285 const judgeAt = (now: number, cadence: JudgeReason): void => {
286 if (!armedCheck() && shouldRun(state, now)) void runJudge(cadence).catch(() => undefined)
287 }
288
289 // A process run that reported nothing: the failure is said, since a process judge that went quiet looks fine.
290 const processFailed = (failure: string): void => {
291 dispatch({ type: 'process.done', patterns: state.patterns, fresh: [], recurred: [], spent: 0, error: failure, returned: 0, kept: 0, dropped: [], usage: null })
292 host?.toast(`ContextSaver: process check failed — ${failure}`)
293 }
294
295 const processOnce = async (engine: Host, trigger: ProcessTrigger): Promise<void> => {
296 const now = await engine.now()
297 const { lastAtMs, lastAtSeq, lastAtTurn } = state.process
298 dispatch({ type: 'process.start', now, seq: state.seq, trigger })
299 // One alias table for the run, as the habit judge keeps: the reply cites the handles this digest printed.
300 const aliases = processAliases(state)
301 let reply: ModelForkResult | null = null
302 let failed: string | null = null
303 try {
304 reply = await engine.fork(buildProcessPrompt(state, aliases))
305 } catch (err) {
306 failed = messageOf(err)
307 }
308 // Every trigger is automatic, none asked for: a cold snapshot is no run and no failure, and the next trigger tries again.
309 if (reply === null) return failed === null ? dispatch({ type: 'process.cold', was: { lastAtMs, lastAtSeq, lastAtTurn } }) : processFailed(failed)
310 try {
311 const { findings, dropped, returned } = parseProcessReply(reply.text, state, aliases)
312 const merged = mergeProcess(state, findings)
313 const reasons = [...dropped, ...merged.evicted.map(id => `${id}: evicted, over MAX_PATTERNS (${MAX_PATTERNS})`)]
314 const kept = findings.length - merged.evicted.length
315 const usage = usageOf(reply.usage)
316 dispatch({
317 type: 'process.done', patterns: merged.patterns, fresh: merged.fresh, recurred: merged.recurred,
318 spent: spentOf(usage), error: null, returned, kept, dropped: reasons, usage,
319 })
320 if (isDebug) engine.log(`ContextSaver process: ${returned} returned · ${kept} kept · ${reasons.length} dropped · from ${trigger}`)
321 persist()
322 return autoOpen(merged.fresh.filter(id => state.cards.includes(id)))
323 } catch (err) {
324 // Whatever went wrong, the run is over: `running` may never stay true.
325 processFailed(messageOf(err))
326 }
327 }
328
329 /**
330 * Starts one detached process run when its gate lets it, answering whether it did.
331 *
332 * @param now the clock the gate reads
333 * @param trigger what asked: a plan, a launch, a pace complaint or the clock
334 * @param before what the run waits for before the digest is built (a launch's declared phases)
335 */
336 const processAt = (now: number, trigger: ProcessTrigger, before?: () => Promise<void>): boolean => {
337 const engine = host
338 if (engine === null || processing || !shouldProcess(state, now, trigger)) return false
339 processing = true
340 void (async () => {
341 try {
342 await before?.()
343 await processOnce(engine, trigger)
344 } finally {
345 processing = false
346 }
347 })().catch(() => undefined)
348 return true
349 }
350
351 // The phases a launch's script declares, read as a literal and never run; a script we cannot read declares none.
352 const readPhases = async (engine: Host, runId: string, value: unknown): Promise<void> => {
353 const path = isRecord(value) && typeof value['scriptPath'] === 'string' ? value['scriptPath'] : null
354 if (path === null) return
355 const script = await engine.readFile(path).catch(() => null)
356 if (script !== null) dispatch({ type: 'run.phases', runId, phases: phasesOf(script) })
357 }
358
359 // A burst of deliveries is said once: a workflow's agents start in a crowd, and a toast each is a storm.
360 const deliveredToast = (now: number): void => {
361 if (now - state.delivery.lastToastAt < DELIVERY_BURST_MS) return
362 host?.toast('ContextSaver: your fixes reached a new subagent')
363 dispatch({ type: 'delivery.toasted', now })
364 }
365
366 // A loop a journal or an `Agent` result named: the engine's own forks (compaction, memory) are named by neither.
367 const isNamed = (agentId: string): boolean => state.loops.some(l => l.id === agentId && (l.run !== null || l.label !== null))
368
369 // The standing fixes an agent the spawn rewrite missed is owed on its first call; a spawn in flight counts as reached.
370 const owedTo = async (engine: Host, agentId: string, now: number): Promise<string[]> => {
371 if (state.standing.length === 0 || state.delivery.pending > 0 || state.delivery.agents.includes(agentId)) return []
372 // A workflow agent calls before the journal that names it was re-read: it is read once more for that id.
373 if (!isNamed(agentId) && !probed.has(agentId) && activeRuns(state, now).length > 0) {
374 probed.add(agentId)
375 await refreshRuns(true).catch(() => undefined)
376 }
377 if (!isNamed(agentId) || state.delivery.agents.includes(agentId)) return []
378 const standing = [...state.standing]
379 dispatch({ type: 'delivery.sent', agentId, now })
380 if (isDebug) engine.log(`ContextSaver delivery: first call → ${agentId}`)
381 deliveredToast(now)
382 return standing
383 }
384
385 // What a typed prompt is kept as for the digest: one line, no control characters, at most ASK_HEAD_MAX.
386 const headOf = (text: string): string => collapseWs(text.replace(/\p{Cc}/gu, ' ')).slice(0, ASK_HEAD_MAX)
387
388 const checkNow = (): string => {
389 if (forking || state.judge.running) {
390 // The run already going answers this ask: nothing is forked, and nobody is left without a reply.
391 asked = true
392 return ALREADY_TEXT
393 }
394 void runJudge('/saver check').catch(() => undefined)
395 return CHECKING_TEXT
396 }
397
398 const togglePane = async (): Promise<void> => {
399 const engine = host
400 if (engine === null) return
401 try {
402 // The `ui.close` hook records the close, as it records the person's own.
403 if (state.paneOpen) await engine.closePane({ id: PANE_ID })
404 else await openPane()
405 } catch {
406 // the surface or another plugin refused: the pane stays as it was, and `/saver` says so
407 }
408 }
409
410 const firstLine = (text: string): string => {
411 const [head = ''] = text.split('\n')
412 return head === text ? text : `${head} …`
413 }
414
415 const decide = (patternId: string, choice: Choice, text?: string): void => {
416 const p = state.patterns.find(q => q.id === patternId)
417 if (p === undefined) return
418 dispatch({ type: 'decide', patternId, choice, text })
419 if (state.patterns.find(q => q.id === patternId)?.decision !== choice) return
420 if (choice === 'keep') host?.toast(`ContextSaver: ignored "${p.kind}"`)
421 if (choice === 'kill') host?.toast(`ContextSaver: fixed — ${p.alternative}`)
422 if (choice === 'steer') host?.toast(`ContextSaver: fixed with your note — ${firstLine(text ?? '')}`)
423 persist()
424 }
425
426 // The number the pane draws beside a card is its seat in `cards`; 0 means the card is no longer listed.
427 const seatOf = (patternId: string): number => state.cards.indexOf(patternId) + 1
428
429 const cardReply = (patternId: string, seat: number, tail: string): string => {
430 const kind = state.patterns.find(q => q.id === patternId)?.kind ?? patternId
431 return `ContextSaver: card ${seat} — "${fit(kind, CARD_KIND)}" · ${tail}`
432 }
433
434 const numberOf = (token: string): number | null => (/^\d+$/.test(token) ? Number(token) : null)
435
436 // A number no card wears is a numbering mistake, whichever verb typed it: it is refused, never obeyed.
437 const noCardText = (n: number): string => `ContextSaver: no card ${n} (1–${state.cards.length})`
438
439 // `/saver ignore 2` and `/saver fix 2` decide the card the pane numbers 2, and say which one they took.
440 const decideByNumber = (choice: Choice, token: string): string => {
441 if (state.cards.length === 0) return NOTHING_TEXT
442 const n = numberOf(token)
443 if (n === null) return SAVER_USAGE
444 const patternId = state.cards[n - 1]
445 if (patternId === undefined) return noCardText(n)
446 const reply = cardReply(patternId, n, choice === 'keep' ? 'ignored' : 'fixed')
447 decide(patternId, choice)
448 return reply
449 }
450
451 const steerSubmit = (patternId: string, text: string): void => {
452 const wanted = text.trim()
453 if (wanted === '') {
454 host?.toast('ContextSaver: write the instruction first')
455 return
456 }
457 decide(patternId, 'steer', wanted)
458 }
459
460 // The artifact's own words without the file's furniture: a bullet, or the prose under a frontmatter.
461 const bodyOf = (content: string): string =>
462 content.includes(CLAUDE_MD_HEADING) ? bulletOnly(content).replace(/^- /, '') : (content.split('---\n').at(-1) ?? content)
463
464 const writeArtifact = async (a: Artifact): Promise<void> => {
465 const engine = host
466 if (engine === null) return
467 try {
468 // A whole-file write reads nothing; an append and a settings merge need what is there.
469 const existing = a.mode === 'write' || !(await engine.exists(a.path)) ? null : await engine.readFile(a.path)
470 if (a.mode === 'append') await engine.writeFile(a.path, appendedTo(existing, a.content))
471 if (a.mode === 'write') await engine.writeFile(a.path, a.content)
472 if (a.mode === 'merge-settings') await engine.writeFile(a.path, mergeSettings(existing, a.content))
473 dispatch({ type: 'artifact.done', patternId: a.patternId, kind: a.kind, written: true })
474 engine.toast(`Wrote ${a.path}`)
475 } catch (err) {
476 engine.toast(messageOf(err))
477 }
478 }
479
480 const tryArtifact = (a: Artifact): void => {
481 dispatch({ type: 'standing.add', text: instructionOf(collapseWs(bodyOf(a.content)) || a.title) })
482 dispatch({ type: 'artifact.done', patternId: a.patternId, kind: a.kind, written: true })
483 host?.toast(`Trying "${a.title}" for this session`)
484 }
485
486 // The key the pane draws a card's Fix… field under.
487 const steerFieldKey = (patternId: string): string => `card:${patternId}:text`
488
489 // The ring lands only on an element the drawn tree already holds ('no element of its own is drawn under that
490 // key', d.ts 8846-8853), and a tree lands after the render hook that built it returns. The press that opens
491 // the field asks for a redraw and nothing more, so the ask waits for the frame the field is drawn in, and
492 // asks again while the engine answers that nothing is drawn under the key — a few frames, then it stops.
493 const steerRing = async (patternId: string): Promise<void> => {
494 const engine = host
495 if (engine === null) return
496 // `focus` is a request, not a grant (d.ts 4915-4922): the surface refuses it while the person holds an
497 // element of ours, which the press that opened the field is — but where the composer holds the keys over
498 // an empty line it is granted, and then `autoFocus` lands the ring on the field by itself.
499 void engine.openPane(paneArgs(true)).catch(() => undefined)
500 let denied = 'the ask was never answered'
501 for (let tries = STEER_RING_TRIES; tries > 0; tries -= 1) {
502 await engine.sleep(STEER_RING_WAIT_MS).catch(() => undefined)
503 // The field was closed again, or another card's opened: this ring is nobody's now.
504 if (state.steering !== patternId) return
505 const deny = await engine
506 .focusElement({ requestId: PANE_ID, key: steerFieldKey(patternId) })
507 .then(result => result.deny ?? null)
508 .catch(err => messageOf(err))
509 if (deny === null) return
510 denied = deny
511 }
512 // The ring stayed put, so the keystrokes are the composer's: the way in is the line the person is owed.
513 engine.toast(`ContextSaver: the composer has your keys — type /saver fix ${seatOf(patternId)} <your note>`)
514 if (isDebug) engine.log(`${PLUGIN_NAME}: the ring never reached ${steerFieldKey(patternId)} — ${denied}`)
515 }
516
517 const actions: Actions = {
518 keep: patternId => decide(patternId, 'keep'),
519 steer: patternId => {
520 dispatch({ type: 'steer.begin', patternId })
521 // A second press closed the field; only the press that opened one goes looking for the keyboard.
522 if (state.steering === patternId) void steerRing(patternId)
523 },
524 // The pane's body is this hook's tree, so the redraw is what paints the keystroke; the text it draws
525 // back is this one, which is also what `/saver fix` sends when the keyboard never reaches the field.
526 steerDraft: text => dispatch({ type: 'steer.draft', text }),
527 steerSubmit: (patternId, text) => steerSubmit(patternId, text),
528 kill: patternId => decide(patternId, 'kill'),
529 info: patternId => dispatch({ type: 'expand', patternId }),
530 togglePane: () => {
531 void togglePane()
532 },
533 // The band and the pane have no reply to write in, so the press is answered with a toast.
534 check: () => {
535 host?.toast(checkNow())
536 },
537 write: a => {
538 void writeArtifact(a)
539 },
540 tryOnce: a => tryArtifact(a),
541 // A skipped rule is handled: `state.written` is the set the pane never offers again.
542 skip: a => dispatch({ type: 'artifact.done', patternId: a.patternId, kind: a.kind, written: true }),
543 }
544
545 on('session.start', async ($, e, next) => {
546 try {
547 const engine: Host = {
548 now: () => $.clock.now(),
549 sleep: ms => $.clock.sleep(ms),
550 invalidate: () => $.ui.invalidate('ui.render'),
551 toast: text => $.ui.toast(text),
552 log: text => $.ui.log(text),
553 openPane: args => $.ui.open(args),
554 closePane: args => $.ui.close(args),
555 focusElement: args => $.ui.focus(args),
556 registerCommand: spec => $.command.register(spec),
557 usage: args => $.session.usage(args),
558 messages: () => $.session.messages(),
559 storeGet: key => $.store.get(key),
560 storeSet: (key, value) => $.store.set(key, value),
561 fork: prompt => $.model.fork({ prompt }),
562 readFile: path => $.fs.read(path),
563 writeFile: (path, text) => $.fs.write(path, text),
564 exists: path => $.fs.exists(path),
565 debugFlag: () => $.env.get('CONTEXTSAVER_DEBUG'),
566 }
567 host = engine
568 const u = await engine.usage({ breakdown: 'summary' })
569 const now = await engine.now()
570 const stored = parseRegistry(await engine.storeGet(storeKey(e.cwd)))
571 state = { ...initialState(e.cwd, u.context.window), patterns: stored.map(fromStored) }
572 dispatch({
573 type: 'usage',
574 usage: { window: u.context.window, compactAt: u.context.breakdown?.autoCompactThreshold, tokens: u.context.tokens, percent: u.context.percent },
575 now,
576 })
577 const tokensOf = (items: readonly { tokens: number }[] | undefined): number => (items ?? []).reduce((n, i) => n + i.tokens, 0)
578 const breakdown = u.context.breakdown
579 dispatch({
580 type: 'overhead',
581 overhead: { memory: tokensOf(breakdown?.memoryFiles), mcp: tokensOf(breakdown?.mcpTools), agents: tokensOf(breakdown?.agents) },
582 })
583 try {
584 await engine.registerCommand(COMMAND)
585 } catch (err) {
586 engine.log(`${PLUGIN_NAME}: /${COMMAND.name} is taken — ${messageOf(err)}`)
587 }
588 try {
589 const flag = await engine.debugFlag()
590 isDebug = flag !== undefined && flag !== '' && flag !== '0'
591 } catch {
592 isDebug = false
593 }
594 // Last, so a transcript we cannot read costs the session nothing it already has.
595 const adopted = adoptRows(await engine.messages())
596 if (adopted.length === 0) return next(e)
597 dispatch({ type: 'adopt', rows: adopted })
598 // Too few rows to judge: a fresh session, nothing armed and nothing forked — which the log says too,
599 // since a debug line that claims a check on three rows is worse than no line at all.
600 const enough = state.rows.length >= JUDGE_MIN_ROWS
601 if (isDebug) {
602 const tail = enough ? 'checking them now' : `under the ${JUDGE_MIN_ROWS}-row floor, nothing to check`
603 engine.log(`ContextSaver adopted ${adopted.length} rows from the transcript · ${tail}`)
604 }
605 if (!enough) return next(e)
606 dispatch({ type: 'check.arm' })
607 if (isDebug) engine.log(`ContextSaver fired a check over ${state.rows.length} adopted rows · armed, so a cold answer retries`)
608 // Fired here, not left for the person's next keystroke: a session with this much history behind it has
609 // run turns, and `$.model.fork` reads its last turn's cache-safe snapshot (d.ts 2019-2034) — which is
610 // exactly what a `/reload-plugins`, an edit under `--plugin-dir` or a resume hands us (d.ts 3106-3111).
611 // Detached, never awaited: this hook is awaited by the engine and has a budget, and the run speaks in a
612 // toast, not in a return value. A snapshot that really is cold answers null, and the arming stays up
613 // (§5.2) for the first warm opportunity below — the next prompt, the next tool call, or the turn's end.
614 armedCheck()
615 return next(e)
616 } catch {
617 return next(e)
618 }
619 })
620
621 on('turn.start', async ($, e, next) => {
622 try {
623 // The clock dates the wait since the last answer, so TURNS can say the person was away.
624 const now = host === null ? 0 : await host.now()
625 dispatch({ type: 'turn.start', now })
626 return next(e)
627 } catch {
628 return next(e)
629 }
630 })
631
632 on('tool.call', async ($, e, next) => {
633 const engine = host
634 let started = 0
635 try {
636 if (engine === null || next.origin.plugin === PLUGIN_NAME) return next(e)
637 started = await engine.now()
638 } catch {
639 return next(e)
640 }
641 // `next(e)` is called exactly once: a rejection is the engine's to report — calling it again would run the tool twice.
642 const result = await next(e)
643 try {
644 const ended = await engine.now()
645 dispatch({ type: 'row', row: rowOf(e, result, ended - started, state.turn) })
646 const row = state.rows[state.rows.length - 1]
647 if (isDebug && row !== undefined) engine.log(`ContextSaver row r${row.seq} ${row.tool} ${row.key} ${row.ms}ms ${row.chars}ch`)
648 // A `Workflow` result is a run launched, an `Agent` result a loop named; a launch reads its journal at
649 // once, and any run still going is re-read on the plugin's own cadence — detached, the call is answered.
650 const run = runOf(result.result)
651 if (run !== null) dispatch({ type: 'run.start', run, now: ended })
652 const agent = agentOf(spawnValue(e, result.result))
653 if (agent !== null) dispatch({ type: 'agent.start', ...agent })
654 void refreshRuns(run !== null).catch(() => undefined)
655 // One agentic turn can run for hours, so the cadence is judged here too, not only between turns.
656 judgeAt(ended, 'tool.call')
657 // A launch waits for the phases its script declares; a plan counts once the person accepted it.
658 const phases = (): Promise<void> => (run === null ? Promise.resolve() : readPhases(engine, run.id, result.result))
659 const planned = (e.tool as string) === PLAN_TOOL && result.deny === undefined && result.isError !== true
660 const launched = run !== null && processAt(ended, 'workflow', phases)
661 if (run !== null && !launched) void phases().catch(() => undefined)
662 if (!launched && !(planned && processAt(ended, 'plan'))) processAt(ended, 'clock')
663 if (result.deny !== undefined) return result
664 // A one-time note is the main loop's: the orchestrator is who re-plans, and a subagent is a random reader.
665 const main = e.agentId === undefined
666 const pending = main ? state.notes : await owedTo(engine, e.agentId ?? '', ended)
667 if (pending.length === 0) return result
668 if (main) dispatch({ type: 'notes.drained' })
669 return { ...result, context: [...(result.context ?? []), ...pending] }
670 } catch {
671 return result
672 }
673 })
674
675 on('turn.complete', async ($, e, next) => {
676 try {
677 const engine = host
678 if (engine === null) return next(e)
679 const u = e.usage
680 // A subagent's turn is its loop's: what it cost and how it ended go onto the loop, and nothing else
681 // moves — the window is the main loop's to sample, and so is the cadence.
682 if (e.agentId !== undefined) {
683 dispatch({ type: 'loop.turn', agentId: e.agentId, model: u?.model ?? null, ms: e.durationMs, tokens: tokensOf(u), ended: e.reason, turn: state.turn })
684 return next(e)
685 }
686 // The window is sampled before the turn is recorded: how full it is after this turn is the turn's
687 // own figure, and its growth over the last turns is the pace compaction actually runs at. The
688 // sample is optional, though: a refused `session.usage` costs this turn its context reading, never
689 // the turn itself — without the stat the trend, the pace and every token gate go with it.
690 const seen = await engine.usage().catch(() => null)
691 const now = await engine.now()
692 dispatch({
693 type: 'turn.complete',
694 stat: {
695 ...tokensOf(u),
696 ms: e.durationMs,
697 answerChars: e.answer.length,
698 answerHead: e.answer.slice(0, ANSWER_HEAD),
699 aborted: e.isAborted,
700 ended: e.reason,
701 at: now,
702 idleMs: 0,
703 context: seen?.context.tokens ?? null,
704 },
705 })
706 if (seen !== null) dispatch({ type: 'usage', usage: { window: seen.context.window, tokens: seen.context.tokens, percent: seen.context.percent }, now })
707 judgeAt(now, 'turn.complete')
708 processAt(now, 'clock')
709 return next(e)
710 } catch {
711 return next(e)
712 }
713 })
714
715 // The standing habit fixes ride a new subagent's prompt; the engine's own forks inherit the parent's context.
716 on('agent.spawn', async ($, e, next) => {
717 const standing = [...state.standing]
718 if (host === null || e.fork || standing.length === 0) return next(e)
719 dispatch({ type: 'delivery.pending', delta: 1 })
720 try {
721 const started = await next({ ...e, prompt: `${e.prompt}\n\n${standing.join('\n')}` })
722 try {
723 const now = await host.now()
724 if (started.agentId === undefined) return started
725 dispatch({ type: 'delivery.sent', agentId: started.agentId, now })
726 if (isDebug) host.log(`ContextSaver delivery: spawn rewrite → ${started.agentId}`)
727 deliveredToast(now)
728 } catch {
729 // the subagent started with the fixes either way
730 }
731 return started
732 } finally {
733 dispatch({ type: 'delivery.pending', delta: -1 })
734 }
735 })
736
737 on('session.compact', async ($, e, next) => {
738 const result = await next(e)
739 try {
740 dispatch({ type: 'compact' })
741 } catch {
742 // a compaction we failed to record is still the compaction the engine performed
743 }
744 return result
745 })
746
747 on('prompt.submit', async ($, e, next) => {
748 try {
749 // `origin` is the engine's to stamp; read it defensively so an unstamped submission still carries the texts.
750 if (e.origin?.kind === 'plugin' || e.text.trimStart().startsWith(`/${COMMAND.name}`)) return next(e)
751 // Only what the person typed is an ask: a task notification in the same words is nobody complaining.
752 if (host !== null && TYPED.includes(e.origin?.kind ?? '')) {
753 const now = await host.now()
754 const pace = isPace(e.text)
755 dispatch({ type: 'ask', ask: { turn: e.turnId !== undefined ? state.turn : state.turn + 1, at: now, head: headOf(e.text), pace } })
756 if (pace) processAt(now, 'pace')
757 }
758 const extra = [...state.notes, ...state.standing.filter(text => !state.notes.includes(text))]
759 if (extra.length > 0) dispatch({ type: 'notes.drained' })
760 const carried = extra.length === 0 ? e : { ...e, context: [...(e.context ?? []), ...extra] }
761 // A prompt is no new work, so there is no cadence lane here: only a check still armed fires — the load
762 // fired its own, so this is the retry of one that came back cold — and it never changes what the
763 // prompt carries.
764 armedCheck()
765 return next(carried)
766 } catch {
767 return next(e)
768 }
769 })
770
771 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
772 const engine = host
773 try {
774 if (engine === null || e.props.hasSurvey || e.surface === 'mobile') return next(e)
775 if (e.props.bodyColumns !== state.columns) observe({ type: 'columns', columns: e.props.bodyColumns })
776 } catch {
777 return next(e)
778 }
779 // Drawn once: a band we cannot build answers with what is beneath it, never with a second dispatch.
780 const below: RenderElement = await next(e)
781 // The clock, read once per draw: the newest reading the state holds is a turn's end or a run's launch, so
782 // a loopless run would read as fresh for the rest of a quiet session. A host that will not say the time
783 // leaves the band that newest reading of its own rather than undrawn.
784 const now = await engine.now().catch(() => null)
785 try {
786 const { Box, Text, Button, Input, Raster } = $.ui.resolve(e) as unknown as Ui
787 const band = Band({
788 ui: { Box, Text, Button, Input, Raster },
789 model: now === null ? bandModel(state) : bandModel(state, now),
790 site: { bodyColumns: e.props.bodyColumns, maxRows: e.props.maxRows },
791 actions,
792 })
793 return Box({ flexDirection: 'column', children: [below, band] })
794 } catch {
795 return below
796 }
797 })
798
799 on('ui.render', { component: 'Pane', requestId: PANE_ID }, ($, e, next) => {
800 try {
801 if (host === null || e.surface === 'mobile') return next(e)
802 const { Box, Text, Button, Input, Raster } = $.ui.resolve(e) as unknown as Ui
803 return Pane({
804 ui: { Box, Text, Button, Input, Raster },
805 model: paneModel(state, propose(state)),
806 site: { bodyColumns: e.props.bodyColumns, maxRows: e.props.scroll.bodyRows },
807 placement: e.props.placement,
808 actions,
809 })
810 } catch {
811 return next(e)
812 }
813 })
814
815 on('command.run', { command: COMMAND.name }, async ($, e, next) => {
816 try {
817 if (host === null) return next(e)
818 const args = e.args.trim()
819 const [sub = ''] = args.split(/\s+/)
820 if (sub === '' || sub === 'rules') {
821 await togglePane()
822 const process = processLine(state)
823 const shown = state.paneOpen ? 'ContextSaver pane shown' : 'ContextSaver pane hidden'
824 return { text: process === null ? shown : `${shown} · ${process}` }
825 }
826 if (sub === 'check') return { text: checkNow() }
827 if (sub === 'fix') {
828 const rest = args.slice(sub.length).trim() // newlines inside the instruction survive
829 const [first = ''] = rest.split(/\s+/)
830 // A leading number is always the card: folding a mistyped one back into the instruction would
831 // fix the wrong card with a garbled sentence, and `standing` keeps it for the whole session.
832 const n = numberOf(first)
833 if (state.cards.length === 0) return { text: NOTHING_TEXT }
834 if (n !== null && (n < 1 || n > state.cards.length)) return { text: noCardText(n) }
835 const text = (n === null ? rest : rest.slice(first.length)).trim()
836 // Nothing after the number sends the fix the card already offers; a note sends the note instead.
837 if (text === '') return { text: n === null ? FIX_USAGE : decideByNumber('kill', first) }
838 const patternId = n === null ? (state.steering ?? state.cards[0]) : state.cards[n - 1]
839 const seat = patternId === undefined ? 0 : seatOf(patternId)
840 if (patternId === undefined || seat === 0) return { text: FIX_USAGE }
841 steerSubmit(patternId, text)
842 return { text: cardReply(patternId, seat, `fixed with your note: ${text}`) }
843 }
844 if (sub === 'ignore') return { text: decideByNumber('keep', args.slice(sub.length).trim()) }
845 if (sub === 'demo' && isDebug) {
846 // Debug-only: the pane's own look, without waiting for a real finding. The header is part of that
847 // look, so a usage sample and its turns come first — without them the hero row draws its empty
848 // state above cards that state a percentage of context. (`now` is the reducer's to ignore.)
849 dispatch({ type: 'usage', usage: demoUsage(), now: 0 })
850 // Four turns, each with the context it left behind, so the header's trend and its run to
851 // compaction draw from a window filling up rather than from what a turn was billed.
852 const samples = demoTurns()
853 for (const [at, context] of DEMO_CONTEXT.entries()) {
854 const stat = samples[at % samples.length]
855 if (stat !== undefined) dispatch({ type: 'turn.complete', stat: { ...stat, context } })
856 }
857 for (const row of demoRows(state.turn)) dispatch({ type: 'row', row })
858 const patterns = demoPatterns(state.turn)
859 const fresh = patterns.filter(p => p.decision === null).map(p => p.id)
860 const usage = demoForkUsage()
861 dispatch({
862 type: 'judge.done', patterns, fresh, recurred: [], focus: null, time: null, context: null,
863 spent: spentOf(usage), error: null, returned: patterns.length, kept: patterns.length, dropped: [], usage,
864 })
865 await openPane()
866 return { text: 'ContextSaver: demo wasters loaded' }
867 }
868 if (sub === 'debug') return { text: debugDump(state, armedSpoke) }
869 if (sub === 'reset') {
870 resetSession()
871 return { text: 'ContextSaver: session state reset' }
872 }
873 return { text: SAVER_USAGE }
874 } catch {
875 return next(e)
876 }
877 })
878
879 on('command.run', { command: ['clear', 'resume'] }, async ($, e, next) => {
880 const result = await next(e)
881 try {
882 resetSession()
883 } catch {
884 // the command ran either way: it is not ours to run a second time
885 }
886 return result
887 })
888
889 on('ui.close', { id: PANE_ID }, async ($, e, next) => {
890 const result = await next(e)
891 try {
892 // A close another plugin refused leaves the pane open, so the band still reads `Close`.
893 if (result.deny === undefined) dispatch({ type: 'pane', open: false })
894 } catch {
895 // the pane stays as the surface left it
896 }
897 return result
898 })
899}
900hooks/core/adopt.ts 35 lines1import type { SessionMessage, ToolCallResult, ToolUseSummary } from 'claude-code'
2
3import { rowOf } from './ledger'
4import { RECOVERED_FLAG, ROW_CAP } from './types'
5import type { Row } from './types'
6
7// Only a request starts a turn: a user message with no typed text is the tool loop or a message the engine queued itself.
8const isPrompt = (m: SessionMessage): boolean =>
9 m.role === 'user' && m.text.trim() !== '' && (m.toolResults ?? []).length === 0
10
11// The transcript kept what the call returned; a refusal was stored as its error text, never as a `deny`.
12const resultOf = (u: ToolUseSummary): ToolCallResult =>
13 u.isError === true ? { isError: true, result: u.result, text: u.text ?? '' } : { result: u.result, text: u.text }
14
15// Appended last, so a flag the call itself earned (`err`, `dedup`) still reads first.
16const flagged = (row: Omit<Row, 'seq'>): Omit<Row, 'seq'> => ({ ...row, flags: [...row.flags, RECOVERED_FLAG] })
17
18// The transcript records no duration and no agent, so `ms` is 0 and `agent` reads `main`.
19const rowFrom = (u: ToolUseSummary, turn: number): Omit<Row, 'seq'> =>
20 flagged(rowOf({ ...u.input, tool: u.tool, tool_use_id: u.tool_use_id }, resultOf(u), 0, turn))
21
22/** Rebuilds the newest ROW_CAP ledger rows from the transcript of a session this plugin joined late. */
23export const adoptRows = (messages: readonly SessionMessage[]): readonly Omit<Row, 'seq'>[] => {
24 const rows: Omit<Row, 'seq'>[] = []
25 let prompts = 0
26 for (const m of messages) {
27 if (isPrompt(m)) prompts += 1
28 if (m.role !== 'assistant') continue
29 // A call with neither a result nor a text is still in flight: it has no size and no outcome to record.
30 const settled = m.toolUses.filter(u => u.result !== undefined || u.text !== undefined)
31 for (const u of settled) rows.push(rowFrom(u, Math.max(1, prompts)))
32 }
33 return rows.slice(-ROW_CAP)
34}
35hooks/core/demo.ts 131 lines1import { MAIN_AGENT } from './types'
2import type { JudgeUsage, Pattern, Row, TurnStat, Usage } from './types'
3
4// A sample session, three turns wide: what the demo's rows and patterns are dated against. A session
5// younger than the sample is dated from turn 8, so the cards read 'turns 5–8' rather than 'turn 1'.
6const DEMO_TURN = 8
7const at = (turn: number, back: number): number => Math.max(1, Math.max(turn, DEMO_TURN) - back)
8
9const SUITE_KEY = 'test:bun test'
10const LOG_KEY = 'read:cat logs/api.log'
11const AGENT_KEY = 'agent:explore'
12const EXPLORE_AGENT = 'sub-explore-1' // one sample read ran inside the explore loop, so the alias column draws
13
14const sample = (
15 id: string,
16 tool: string,
17 key: string,
18 cls: Row['cls'],
19 turn: number,
20 ms: number,
21 chars: number,
22 head: string,
23 spawn: Row['spawn'] = null,
24): Omit<Row, 'seq'> => ({ id, tool, key, cls, agent: MAIN_AGENT, turn, ms, chars, head, flags: [], lines: null, paths: [], spawn })
25
26// The same call, made inside the explore agent rather than the main loop: the pane names that loop `a1`.
27const inExplore = (row: Omit<Row, 'seq'>): Omit<Row, 'seq'> => ({ ...row, agent: EXPLORE_AGENT })
28
29const turnSample = (input: number, output: number, ms: number, answer: string): Omit<TurnStat, 'turn' | 'calls'> => ({
30 input, output, cacheRead: 180_000, cacheCreate: 7_000, ms, answerChars: answer.length, answerHead: answer, aborted: false, ended: 'answer', at: 0, idleMs: 0, context: null,
31})
32
33/** The usage the demo's header draws from: a third of a million-token window spent, with a compaction threshold. */
34export const demoUsage = (): Usage => ({ window: 1_000_000, compactAt: 900_000, tokens: 320_000, percent: 32 })
35
36/** What the demo's one judge run cost: a cold fork, whose cache creation is most of the bill. */
37export const demoForkUsage = (): JudgeUsage => ({ input: 96_000, output: 1_200, cacheRead: 0, cacheCreate: 221_000 })
38
39/** Three sample turns, so the header's run to compaction has a pace to state it in turns. */
40export const demoTurns = (): Omit<TurnStat, 'turn' | 'calls'>[] => [
41 turnSample(42_000, 3_000, 96_000, 'Ran the suite: 212 pass. Reading the api log for the 500 next.'),
42 turnSample(41_000, 4_000, 132_000, 'The refresh flow lives in src/auth/refresh.ts; the token TTL is the bug.'),
43 turnSample(43_000, 2_000, 88_000, 'Suite is green again after the token fix.'),
44]
45
46/** The ledger rows the demo's wasters cite, so their cost and their evidence quotes are real rows. */
47export const demoRows = (turn: number): Omit<Row, 'seq'>[] => [
48 sample('demo-suite-1', 'Bash', SUITE_KEY, 'test', at(turn, 3), 62_000, 24_000, '212 pass · 0 fail · ran 1284 expect() calls in 61.98s'),
49 sample('demo-log-1', 'Bash', LOG_KEY, 'read', at(turn, 3), 4_000, 80_000, 'GET /health 200 12ms · GET /v1/users 200 41ms · POST /v1/tokens 500 88ms'),
50 sample('demo-agent-1', 'Agent', AGENT_KEY, 'other', at(turn, 2), 90_000, 36_000, 'Explored src/auth: 6 files read, the refresh flow lives in src/auth/refresh.ts', {
51 type: 'explore', requested: null, resolved: 'claude-sonnet-4-6', status: 'completed', tokens: 9_000, edits: 0, promptChars: 420,
52 }),
53 sample('demo-suite-2', 'Bash', SUITE_KEY, 'test', at(turn, 2), 59_000, 24_000, '212 pass · 0 fail · ran 1284 expect() calls in 58.71s'),
54 inExplore(sample('demo-log-2', 'Bash', LOG_KEY, 'read', at(turn, 1), 11_000, 80_000, 'GET /health 200 11ms · GET /v1/users 200 39ms · POST /v1/tokens 500 91ms')),
55 sample('demo-agent-2', 'Agent', AGENT_KEY, 'other', at(turn, 1), 90_000, 36_000, 'Explored src/auth: 6 files read, the refresh flow lives in src/auth/refresh.ts', {
56 type: 'explore', requested: null, resolved: 'claude-sonnet-4-6', status: 'completed', tokens: 9_000, edits: 0, promptChars: 430,
57 }),
58 sample('demo-suite-3', 'Bash', SUITE_KEY, 'test', at(turn, 0), 71_000, 24_000, '212 pass · 0 fail · ran 1284 expect() calls in 70.44s'),
59]
60
61/**
62 * Three sample wasters for `/saver demo`: two awaiting a decision — the spawn one carrying the judge's
63 * own rule — and one already steered, whose rule the pane derives from the instruction that was sent.
64 */
65export const demoPatterns = (turn: number): Pattern[] => [
66 {
67 id: 'execution:full-suite',
68 category: 'execution',
69 kind: 'Claude keeps running the whole bun test suite after every single-file edit',
70 signature: { tool: 'Bash', key: SUITE_KEY },
71 why: 'the suite ran in full three times while only src/auth.ts changed between runs; the user asked for a fix, not full verification',
72 alternative: 'run only the tests covering the files you changed; run the full suite once when the phase is done',
73 confidence: 0.9,
74 proposal: null,
75 estTokensPerTurn: null,
76 lastDecision: null,
77 lean: null,
78 hits: ['demo-suite-1', 'demo-suite-2', 'demo-suite-3'],
79 decision: 'steer',
80 decidedAtTurn: at(turn, 1),
81 instruction: 'run only the tests for the file you just edited; the full suite once at the end of the phase',
82 openedAtTurn: at(turn, 1),
83 ignored: 0,
84 sent: 0,
85 },
86 {
87 id: 'reading:api-logs',
88 category: 'reading',
89 kind: 'Claude keeps reading 2000 lines of api logs instead of grepping for the error',
90 signature: { tool: 'Bash', key: LOG_KEY },
91 why: 'the whole log was read twice when one grep would have shown the traceback',
92 alternative: "grep -nE 'ERROR|Traceback' and read only the 50 lines around the match",
93 confidence: 0.8,
94 proposal: null,
95 estTokensPerTurn: null,
96 lastDecision: null,
97 lean: null,
98 hits: ['demo-log-1', 'demo-log-2'],
99 decision: null,
100 decidedAtTurn: null,
101 instruction: null,
102 openedAtTurn: null,
103 ignored: 0,
104 sent: 0,
105 },
106 {
107 id: 'multi-agent:re-explores',
108 category: 'multi-agent',
109 kind: 'Claude keeps spawning a fresh explore agent that re-reads what the last one already reported',
110 signature: { tool: 'Agent', key: AGENT_KEY },
111 why: 'two explore agents read the same six files in src/auth, and the second brief carried none of the first report',
112 alternative: "pass the previous explore agent's findings into the next brief instead of asking a new agent to read the same files",
113 confidence: 0.8,
114 proposal: {
115 kind: 'claude-md',
116 title: "Reuse an explore agent's findings",
117 body: "Give a new explore agent the previous agent's findings; never ask it to re-read files another agent has already reported on.",
118 },
119 estTokensPerTurn: null,
120 lastDecision: null,
121 lean: null,
122 hits: ['demo-agent-1', 'demo-agent-2'],
123 decision: null,
124 decidedAtTurn: null,
125 instruction: null,
126 openedAtTurn: null,
127 ignored: 0,
128 sent: 0,
129 },
130]
131hooks/core/judge.ts 381 lines1import type { ModelForkUsage } from 'claude-code'
2
3import { agentsBlock, citableRows, decisionsBlock, knownPatternsBlock, ledgerBlock, sinksBlock, statsLines, turnsBlock } from './blocks'
4import { agentAliases } from './evidence'
5import { totalTokens } from './patterns'
6import { cell, citedTurns, collapseWs, evidenceOf, groundedCap, isRecord, median, parseObject, proposalOf, short, str, unique } from './text'
7import {
8 ALTERNATIVE_MAX, JUDGE_MIN_GAP_MS, JUDGE_MIN_NEW_ROWS, JUDGE_MIN_NEW_TOKENS,
9 JUDGE_MIN_ROWS, JUDGE_MIN_TURNS, KIND_MAX, MAX_BEHAVIORAL_FINDINGS, MAX_FINDINGS, MAX_PATTERNS,
10} from './types'
11import type { Category, Finding, JudgeUsage, Pattern, Row, Signature, State } from './types'
12
13/** The judge prompt (build spec Appendix A, verbatim) with the eight evidence placeholders. */
14export const JUDGE_PROMPT = `You are auditing THIS session for wasted context and wasted time. The transcript above is your own: read it for intent — what the user asked for, what you were told, what you already decided. The blocks below are the only evidence of what actually ran; nothing outside them exists for this audit.
15
16Answer four narrow questions. What repeated: which behaviours have already happened more than once, separated by other work, or have you said you will keep doing — and what should be done instead? Where the time and the context went: which of the largest sinks below are repetition or work nobody asked for rather than the work this session needed. What is going in circles: the same failing command retried with no diagnostic step between, read/edit/read on one path with nothing finished, an edit failing on one path over and over. What was out of proportion: which spawned work — an agent, a workflow stage, a review or verification round — cost far more than what it produced, or did what a shell step would have done (AGENTS below).
17
18You are writing an interruption. Every finding can put a card in front of the user mid-work and can become a standing instruction that constrains you for the rest of the session. A wrong finding costs more than a missed one: it interrupts correct work, teaches a bad rule, and makes the user distrust the next card. Prefer silence to a guess. \`"findings": []\` is a correct and common answer.
19
20## Rules
211. Report behaviours, not incidents. A finding needs two unexcused occurrences of the same behaviour separated by other work (see Counting), or one occurrence plus your own stated intent to keep doing it ("I'll re-run the suite after each fix"). A single expensive call is never a finding.
222. \`evidence\`: row ids copied from the LEDGER \`id\` column (\`r12\`), or \`turn:<n>\` where \`<n>\` is a \`turn\` number printed in TURNS, or \`agent:<alias>\` where \`<alias>\` is an \`alias\` printed in AGENTS — turn and agent handles only for findings whose \`signature\` is null. Copy ids exactly; never renumber, abbreviate or reformat one. A finding with an id that is not in these blocks is discarded whole. STATS lines and \`~\` summary lines carry no id: use them for counts and history in \`why\` (they are the whole session, counted for you), never cite them as evidence, and never assume the oldest full row is the first occurrence.
233. \`signature\`: copy \`key\` character-for-character from one LEDGER row and \`tool\` from that same row. Keys are cut at 200 characters — copy the cut, never complete a command from memory. A pair that does not match a row discards the finding. Choose a key that only the wasteful version of the call carries; when no single recurring call carries the behaviour, use \`"signature": null\`.
244. Reuse ids. If KNOWN PATTERNS already names the behaviour, return that exact \`id\` with fresh evidence. The same command, file or lens under a different slug or category is the same waste: scan KNOWN PATTERNS and DECISIONS before minting an id.
255. Respect DECISIONS. A \`keep\` silences that behaviour under any id, category or wording for this session, and a kept occurrence may not be cited as evidence inside another finding. A behaviour kept in a previous session may be reported only with three or more occurrences and confidence 0.8 or higher. A \`steer\` or \`kill\` may be reported again only if it recurred after that turn: cite rows whose \`turn\` is greater and say so in \`why\`.
266. \`kind\`: one sentence of at most 120 characters, starting exactly \`Claude keeps \`, naming the concrete thing — the command, the file, the agent.
277. \`alternative\`: one imperative sentence of at most 200 characters addressed to Claude. It is sent to Claude verbatim, may be re-sent with every prompt for the rest of the session, and may be written into CLAUDE.md, so it must be safe to obey in situations you did not see: scope it ("run only the tests covering the files you changed, then the full suite once per phase"), never ban a capability outright ("never run the test suite"). If you cannot phrase the fix without forbidding something legitimate, drop the finding.
288. \`why\`: one or two sentences of evidence — how many times, what changed between occurrences, what the transcript shows — and the legitimate explanation you considered and what rules it out. Never a count or a cost these blocks do not contain.
299. At most six findings, at most three with \`signature: null\`, ordered by the sum of \`chars\` over the rows cited, largest first. Six is a ceiling, not a target; never split one behaviour into two findings.
30
31## Counting — what makes two occurrences a repeat
32- Separated by other work. Two occurrences of the same behaviour count as two decisions when at least one other row sits between them — an edit, a read, another command — whether in the same turn or a later one; a run after an edit is a second decision, not a second call in one batch — whether it is excused is the category's own question. Calls issued together with nothing between them are one batch and count once: a parallel set of Reads, or a fan-out of subagents launched at once, is one choice however wide. Breadth is one decision; weight is not: every stage of a workflow (each \`label\` in AGENTS) is a decision of its own, and the same role recurring stage after stage — a check loop per chunk, a review round after every fix — is a repeat.
33- Both unexcused. An occurrence the "Never report" list excuses does not count and may not be cited. Subtract the excused ones first; if fewer than two remain, there is no finding. A baseline suite run at the start plus the check before a commit is zero findings.
34- Same side of a compaction, for reading only. TURNS lists the turns where a compaction happened; what it dropped must be re-acquired, so a read or a re-derivation after one is not a repeat of one before it. Repeated work is: a check, a build, a review round or an agent run again across a compaction counts on both sides.
35- A correction is the first occurrence. When the user corrected a behaviour this session, or a saved feedback memory in the transcript names it, that correction counts as its first occurrence: one unexcused occurrence after it is a finding, and \`why\` names the correction.
36- Same behaviour, not the same shape. For a signature finding that means the same \`key\`; two Read keys differing only in \`:offset-limit\` are different slices, not a repeat. For a null-signature finding you must name one behaviour and show it in each cited turn; do not staple unrelated expensive turns together.
37- Agents are loops of their own. The \`agent\` column names the loop; a repeat inside one agent's rows counts exactly like a repeat in the main loop, and the main loop re-doing after an agent returns what that agent's rows show it already did (the same Read key, the same check) is a repeat across loops.
38- Short ledgers. With fewer than about 12 rows or fewer than 4 turns, report only behaviours with three or more surviving occurrences, or one plus explicit stated intent.
39- The legitimacy ladder. A sink needed once is nothing, however large. A sink repeated because its inputs changed between the runs — an edit, an install, a migration — is nothing. A sink repeated with nothing changed between, or work the transcript shows nobody asked for, is a finding, and the excuse you considered is written into \`why\`.
40
41## Categories — the nine names are the whole enum; the cues are examples and \`kind\` is free text
42- execution — the whole suite/build/typecheck after each edit; re-running a check with nothing edited since it last passed; the same failing command retried with no diagnostic step between; \`sleep\` polling or a watch/dev server run as a blocking call (\`bg\` absent, large \`ms\`). Not: a run after any intervening edit, install, migration or config change; the session's baseline run; broad verification after shared code changed; the last check before a commit; one retry of a transient failure. The error text is not in these blocks, so you cannot claim two failures were the same failure.
43- reading — the same path read again with nothing having changed it; whole-file reads where a range would do (\`trunc\`); unfiltered log/diff/verbose dumps (large \`chars\`, \`persist=\`); a Bash read of a file that is then Read again; wide greps with no path scope and no edit after them. Not: a read after your own edit or after any command that could have rewritten the file; a read after a compaction; paging (a second read at a new offset, especially after \`trunc\`); the first look at an unfamiliar file or log; a row flagged \`dedup\` (core charged nothing).
44- production — whole-file rewrites for small changes (\`+a/-d\` near the file's size); edits that cancel out; tests or docs nobody asked for; the same Edit failing on one path over and over. Not: a new file; a rewrite the user asked for; call-site updates the change requires.
45- behavior — read/edit/read oscillation with nothing finished; approach flip-flops; repeating what the user already corrected. Not: read-edit-verify cycles that are the working method, or a step that depends on the previous result.
46- communication — turns with \`calls 0\` and large \`answerChars\` that restate the plan or recap finished work; stopping to ask what the transcript, the repo or your instructions already answer; an \`ask\` row that held the turn for minutes while no agent ran (no rows between it and the next prompt) and whose options carried a recommended default (\`recommended\`) — proceed on the default and ask beside the work. Not: the turn that answers a question the user asked; a plan or explanation you were asked for; the session's last turn; plan mode, where making no tool call is required; a question whose answer the transcript shows changed the plan. \`out\` includes thinking, so point at the restated content, not the token shape.
47- multi-agent — parallel agents each re-reading the same large file the parent already had; agents with a thin brief (small \`promptChars\`, large \`tokens\`); results never read; agents spawned again after a limit error; mechanical agents on the premium model (\`agent=\` flag shows the resolved model and \`edits\`); the main loop re-reading files or re-running checks an agent's rows already covered, after it returned; a brief that pastes in whole files (large \`promptChars\`) to an agent whose rows then Read the same paths anyway; a workflow whose later agents re-read what earlier agents read (the same Read keys under successive \`agent\` values, spread over turns); a loop whose rows are only checks that passed, with \`edits 0\` and a report as its outcome (AGENTS \`checks\` > 0) — a shell step given a model; review or verify loops with \`edits 0\` whose outcome carries no medium, high or critical finding, recurring stage after stage; two or more verifier loops per finding; a run whose tokens per edit are several times the others'; a fix loop followed by a review whose outcome carries a new high in the same stage, twice (regression chasing: stop the loop and re-plan). Not: agents with disjoint file sets each reading one shared spec; two agents touching one path unless the ledger shows a conflict (an errored edit right after another agent's edit, or a re-edit in the main loop after they returned); a re-read whose brief the transcript shows is a review or verification pass; every agent reading the one spec its brief names; the parent reading an agent's result; one review per stage that found a medium or higher; a loop that edited; a fan-out's breadth on its own.
48- environment — installs repeated with no manifest edit; Bash used where Read/Grep/Edit is cheaper; fixed per-turn overhead (memory files, agent descriptions, MCP schemas in the facts line) larger than the work. Not: the user's own denials (\`denied\`); a single approval prompt.
49- process — many tiny commits or amends on one change; work declared done with no check run. Not: docs- or config-only changes with no check to run, or a check the environment cannot run.
50- other — a repetition none of the above names. Name it plainly.
51
52## Never report
53- A first occurrence, or anything with fewer than two unexcused occurrences after the Counting rules.
54- A single long call that was needed once, however long it ran: the largest row in TIME or CONTEXT is a fact to explain, never a finding on its own.
55- Orientation: the first look at any file, directory or log, an unfamiliar area, or a scope the user left open ("audit every call site", "review the repo").
56- Parallelism: calls issued together with nothing between them are one decision, and agents on disjoint scopes launched at once are one decision — one, not none: their weight is judged under multi-agent.
57- A re-read, or re-deriving what was settled, that a compaction made necessary. Compaction, prompt-cache reads and the host's own truncation are the harness working as designed; work repeated across a compaction is not excused by it.
58- A denied call (\`denied\`): the user or a policy said no, never your waste. The only reportable version is re-running an unchanged command already declined twice, and then the fix is a \`settings-allow\` proposal, not a rebuke.
59- A file change you cannot see: the \`paths\` column records only Edit/Write and Bash calls the host diffed, and nothing for a staged edit. Treat an intervening formatter, codegen, migration, install, \`git checkout|stash|pull|apply\`, \`sed -i\`, MCP edit or another agent's edit as having changed the file.
60- Volume alone. A large read is waste only when a cheaper call would have answered the same question for the same purpose; if the output was the deliverable (the diff under review, the log you were asked to explain, a file about to be rewritten) it is not a finding.
61- Turns spent thinking on a genuinely hard decision.
62- A cue whose evidence is not in these columns (an error message, a file's true size, worktree isolation): if you cannot see it, you cannot evidence it.
63- The duration of a \`recovered\` row, or which agent ran it: neither was recorded, and its \`err\` may be a refusal the transcript stored as error text, so treat a \`recovered\` \`err\` row as \`denied\` and never as waste.
64- Anything the user asked for this session, however wasteful it looks. Read the transcript before you accuse.
65
66## Confidence
67\`confidence\` runs 0.5 to 1.0. 0.9+: the same key three or more times with other work between each, nothing changed between, no request for it in the transcript. 0.7-0.9: the repetition is plain and the transcript offers no legitimate reason. 0.5-0.7: the repetition is real but a legitimate reason is plausible — report here only if you looked for that reason and \`why\` names what rules it out; if it could plausibly have been the right call, drop it. Below 0.5: say nothing.
68
69\`est_tokens_per_turn\`: null whenever \`signature\` is an object. For a null signature it is an integer grounded in the \`answerChars\` of the cited turns divided by four, conservative end, or 0 when you cannot ground it; the user sees it multiplied into a savings figure every turn after a decision. For agent handles it is the tokens one avoided loop would have cost, grounded in AGENTS \`tok\`, conservative end.
70
71\`proposal\`: null unless the fix should outlive the session. Otherwise \`{"kind","title","body"}\` where body is, per kind: \`claude-md\` one imperative rule line; \`skill\` the workflow as the body of a SKILL.md; \`agent-brief\` the brief, whose first line may be \`model: haiku\` or \`model: sonnet\`; \`settings-allow\` nothing but a permission rule such as \`Bash(bun test:*)\`.
72
73## Contract — the shape of your reply, stated once (documentation, not a template to echo)
74\`\`\`json
75{"focus": "<one line: what this session is doing>",
76 "time": "<one sentence, at most 200 chars: where the wall-clock went>",
77 "context": "<one sentence, at most 200 chars: where the context went>",
78 "findings": [{"id": "<category>:<kebab-slug, at most 40 chars>",
79 "category": "execution|reading|production|behavior|communication|multi-agent|environment|process|other",
80 "kind": "<one sentence, at most 120 chars, starts 'Claude keeps '>",
81 "evidence": ["<row id>", "turn:<n>"],
82 "signature": {"tool": "<the row's tool cell>", "key": "<the row's key cell, verbatim>"},
83 "why": "<one or two sentences>",
84 "alternative": "<one imperative sentence, at most 200 chars>",
85 "confidence": 0.85,
86 "est_tokens_per_turn": null,
87 "proposal": null}]}
88\`\`\`
89\`time\` and \`context\` explain where each went, for the user to read and in the words of the work: "45 min per chunk: the full proxy suite runs after every fix round and each chunk gets two review rounds". Neither is an accusation and neither is a finding by itself, so write both even when \`findings\` is \`[]\` — the examples below leave them out where they are not the point, your reply never does.
90
91\`findings\` may be \`[]\`. \`signature\` is that object or \`null\`. No other keys, and never null where a string is specified. Reply with one JSON object: first character \`{\`, last character \`}\`, no prose before or after, no code fence.
92
93## Examples — evidence, then what it justifies
94Rows \`r41\`, \`r45\`, \`r50\` carry \`test:bun test\` in turns 7, 8, 9 while only \`/src/auth.ts\` was edited between them, \`r41\` being the session's baseline run: \`{"focus":"fixing the auth token refresh in /src/auth.ts","findings":[{"id":"execution:full-suite-after-each-edit","category":"execution","kind":"Claude keeps running the whole bun test suite after every single-file edit","evidence":["r45","r50"],"signature":{"tool":"Bash","key":"test:bun test"},"why":"r41 was the baseline and is excused; the suite then ran in full at turns 8 and 9 after single-file edits to /src/auth.ts alone, about a minute and 9.7k characters each. Nothing shared changed, and the user asked for a fix, not full verification.","alternative":"Run only the test files covering the files you changed, then the whole suite once when the phase is done.","confidence":0.92,"est_tokens_per_turn":null,"proposal":{"kind":"claude-md","title":"Targeted tests","body":"Run only the tests covering the files you changed; run the full suite at the end of a phase."}}]}\`
95Rows \`r12\` Read \`/src/api.ts:-\`, \`r15\` Edit \`/src/api.ts\`, \`r16\` Read \`/src/api.ts:-\` with \`dedup\`, all in turn 4: \`{"focus":"a one-file change in /src/api.ts","findings":[]}\` — the second read follows your own edit and the third was deduped, so no unexcused occurrence remains.
96Rows \`r08\` \`test:bun test\` (turn 2, baseline), \`r23\` \`test:bun test test/db.test.ts\` (turn 6, after an edit), \`r40\` \`test:bun test\` (turn 11) followed by \`r41\` \`git:git commit …\`: \`{"focus":"a db pool fix, verified narrowly then once before the commit","findings":[]}\` — both full runs are excused, so nothing survives the Counting rules.
97Rows \`r61\` and \`r72\` both \`read:docker compose logs api --tail 2000\` in turns 11 and 13, each ~40k \`chars\` with \`persist=\`: the same shape as the first example with \`"id":"reading:unfiltered-log-dump"\`, \`"kind":"Claude keeps reading 2000 lines of api logs instead of grepping for the error"\`, \`"alternative":"Pipe log commands through grep -nE 'ERROR|Traceback' and tail -50 instead of reading the whole tail."\`, \`"confidence":0.85\`, \`"proposal":null\`.
98TURNS shows turns 14 and 15 with \`calls 0\` and \`answerChars\` 5400 and 6100 after a single edit at turn 13, neither answering a question: \`"id":"communication:restates-plan-each-turn"\`, \`"evidence":["turn:14","turn:15"]\`, \`"signature":null\`, \`"est_tokens_per_turn":1200\`, \`"alternative":"State the result in one or two lines and take the next action; do not restate the plan or recap completed steps."\`.
99Rows \`r80\`-\`r83\` under \`agent\` \`a1\` Read four files, then \`r84\` Agent \`agent:general-purpose\` flagged \`agent=general-purpose/opus/completed/41000tok/0edits/300pch\` closes that loop (a spawn row lands after the rows it caused), then \`r90\`-\`r93\` in the main loop Read the same four keys in the next turn: \`"id":"multi-agent:re-reads-what-the-agent-read"\`, \`"kind":"Claude keeps re-reading the files a subagent already read for it"\`, \`"evidence":["r80","r83","r90","r93"]\` (a row from each loop is the repeat; the four main-loop reads together are one batch), \`"signature":null\` (no single key carries it), \`"alternative":"Use the subagent's report; re-read a file it covered only to edit it."\`, \`"confidence":0.8\`, \`"est_tokens_per_turn"\` grounded in the cited turns' \`answerChars\` or 0.
100KNOWN PATTERNS lists \`execution:full-suite-after-each-edit | … | steer @ 9\` and rows \`r70\` (turn 12) and \`r76\` (turn 14) carry \`test:bun test\` again: return that same id with \`"evidence":["r70","r76"]\` and a \`why\` that names turns 12 and 14 as after the steer at turn 9.
101Agent rows \`r30\`, \`r58\` and \`r91\` run about 40 minutes each, a review agent follows each, every loop reads a different module and no key repeats: \`{"focus":"rewriting three modules, one agent each","time":"2h 40m in three module rewrites of about 40 minutes each, plus one review pass per module; nothing ran twice.","context":"1.1M chars, three quarters of it the agents' own reads of the modules they rewrote.","findings":[]}\` — a long session is not a wasteful one.
102Four suite runs with an install or a migration between every pair, and \`agents | ×3 | Σ2100000ms | apart\` above them in TIME: \`{"focus":"a schema migration and the call sites it broke","time":"68m, most of it four suite runs, each after a migration or an install changed what the suite covers.","context":"620k chars, over half the migration diff and the failures it produced.","findings":[]}\` — every repeat had changed inputs, so the ladder stops at nothing.
103AGENTS lists \`a3 | w3 | check:C3 | sonnet | 1 | 1.7m | 48k | edits 0 | checks 4 | reads 0 | report 1600ch | answer\` and \`a7 | w3 | check:C4 | …\` alike, and no row of theirs carries \`err\`: \`{"id":"multi-agent:check-loops-for-shell-steps","kind":"Claude keeps spawning an agent per chunk whose only job is to run four passing check commands","evidence":["agent:a3","agent:a7"],"signature":null,"alternative":"Run a stage's checks as a shell step of the workflow script or inside the reviewer; spawn an agent only for work that needs judgment.","confidence":0.85,"est_tokens_per_turn":48000}\`
104Three \`review:*\` loops of one run, each \`edits 0\` with outcome \`3 low\`, a fix loop after each: \`{"id":"multi-agent:review-rounds-that-find-only-lows","kind":"Claude keeps running a full review round after every fix although the last two found only low findings","evidence":["agent:a9","agent:a12"],"signature":null,"alternative":"Route only medium-or-higher review findings to a fix round; end the stage when a review returns lows alone.","confidence":0.8}\`
105
106## KNOWN PATTERNS — \`id | kind | decision @ turn | previous\`. Reuse these ids; never mint a second id or signature for waste listed here.
107{{KNOWN_PATTERNS}}
108
109## DECISIONS — \`id | key | keep|steer|kill @ turn\`, then previous-session keeps. A \`keep\` key is off limits under any id this session.
110{{DECISIONS}}
111
112## STATS — the whole session, counted for you. Per call: \`tool | key | cls | ×count | Σms | Σchars | turns first-last | edits-between | agents\` (edits-between: median number of files edited between consecutive runs; 0 means it re-ran with nothing changed), then \`waits:\` — every AskUserQuestion this session, its total wait and how many carried a recommended default; not citable, but the time it held the session is a fact to explain. Then per class and per agent. No ids here; cite LEDGER rows.
113{{STATS}}
114
115## TIME — where the wall-clock went. \`total\`, then \`label | ×count | Σms | share%\` for the largest sinks, then the five longest rows as \`r<seq> | tool | key | Σms\`. An Agent row holds its own loop's rows, so it is listed apart and never added in; \`ms\` includes any wait on a permission prompt.
116{{TIME}}
117
118## CONTEXT — where the context went. The same shape measured in \`chars\`: the total, the largest sinks with their share, then the five largest rows. An Agent row is listed apart and never added in.
119{{CONTEXT}}
120
121## AGENTS — the loops this session spawned. First one line per run: \`name | id | loops | Σmin | Σtok | edits | turn\`. Then one line per loop, oldest first: \`alias | run | label | model | turns | min | tok | edits | checks | reads | outcome | ended\` — \`alias\` is the LEDGER's \`agent\` name for that loop, \`label\` the stage the workflow gave it (\`impl:C3\`, \`check:C3\`, \`review:C6-r1\`), \`outcome\` what it returned (\`1 high 4 low\`, \`report 5900ch\`), \`ended\` \`answer\`, \`error\`, \`aborted\`, \`refusal\` or \`running\`. \`tok\` counts new tokens (input, cache creation, output) in thousands. Loops older than the window fold into \`~ run | ×loops | Σmin | Σtok\` lines. Cite a loop as \`agent:<alias>\`; a line here is a loop's whole cost, so its weight against its \`edits\` and \`outcome\` is the proportion question's evidence.
122{{AGENTS}}
123
124## TURNS — \`turn | in | out | cacheCreate | calls | ms | answerChars\`, then \`| aborted\`, \`| error\` or \`| refusal\` when the turn ended that way and \`| idle <m>m\` when the next prompt came a minute or more later, then the facts line (context window, fixed per-turn overhead, turns where a compaction happened)
125{{TURNS}}
126
127## LEDGER — \`id | tool | key | cls | agent | turn | ms | chars | flags | paths\`, oldest first (a spawn row lands after the rows it caused: an agent's own calls finish before its Agent row does). Agents are named \`a1\`, \`a2\`… in order of first appearance; \`main\` is the main loop. \`ms\` is wall time and includes any wait on a permission prompt, so a long \`ms\` alone is not machine cost. A row flagged \`recovered\` was rebuilt from the transcript before this plugin joined the session: its \`ms\` is 0 and its agent reads \`main\`, so never reason about its duration or which loop ran it. flags: \`err\` \`denied\` \`dedup\` \`trunc\` \`bg\` \`timeout\` \`persist=<bytes>\` \`ask\` (an AskUserQuestion: its \`ms\` is the wait for the person) \`recommended\` (its options named a default) \`+adds/-dels\` \`agent=<type>/<model>/<status>/<tokens>tok/<edits>edits/<promptChars>pch\`, or \`-\`. Rows older than the window are folded into \`~ | tool | key | ×count | Σchars\` lines: no id, never citable, key usable as a signature only if it also appears in a full row.
128{{LEDGER}}
129
130Return the JSON object only.
131`
132
133// Turns and tokens: the ordinary session, where the judge runs between turns.
134const turnGate = (state: State): boolean =>
135 totalTokens(state) - state.judge.lastAtTokens >= JUDGE_MIN_NEW_TOKENS * state.judge.backoff &&
136 state.turn - state.judge.lastAtTurn >= JUDGE_MIN_TURNS
137
138// Rows and wall time: one agentic turn can run for hours, and `turn.complete` is no cadence inside it.
139const rowGate = (state: State, now: number): boolean =>
140 state.seq - state.judge.lastAtSeq >= JUDGE_MIN_NEW_ROWS * state.judge.backoff &&
141 now - state.judge.lastAtMs >= JUDGE_MIN_GAP_MS
142
143/**
144 * True when the cadence gates allow another judge run.
145 *
146 * @param state the session so far
147 * @param now the clock, for the gap the mid-turn gate keeps between runs
148 */
149export const shouldRun = (state: State, now: number): boolean =>
150 !state.judge.running &&
151 state.rows.length >= JUDGE_MIN_ROWS &&
152 (turnGate(state) || rowGate(state, now))
153
154/** The four counts the fork reported, under our own names: what `/saver debug` and the debug log print. */
155export const usageOf = (u: ModelForkUsage): JudgeUsage => ({
156 input: u.input_tokens,
157 output: u.output_tokens,
158 cacheRead: u.cache_read_input_tokens,
159 cacheCreate: u.cache_creation_input_tokens,
160})
161
162/** What a run's counts cost us: input, output and cache creation (cache reads are free). */
163export const spentOf = (u: JudgeUsage): number => u.input + u.output + u.cacheCreate
164
165/** What one judge fork cost us, straight from the usage the API reported. */
166export const costOf = (u: ModelForkUsage): number => spentOf(usageOf(u))
167
168/** Names this session's loops the way one judge run sees them: built once before the prompt, read again when the reply comes back. */
169export const judgeAliases = (state: State): ReadonlyMap<string, string> => agentAliases(state.rows, state.loops)
170
171/** Fills the judge prompt with this session's evidence blocks; `aliases` is the naming AGENTS and LEDGER print, so the caller can hand the same one to `parseReply`. */
172export const buildPrompt = (state: State, aliases: ReadonlyMap<string, string> = judgeAliases(state)): string =>
173 ([
174 ['{{KNOWN_PATTERNS}}', knownPatternsBlock(state)],
175 ['{{DECISIONS}}', decisionsBlock(state)],
176 ['{{STATS}}', statsLines(state.rows, state.folded, aliases).join('\n')],
177 ['{{TIME}}', sinksBlock(state.rows, 'ms', state.loops, state.folded)],
178 ['{{CONTEXT}}', sinksBlock(state.rows, 'chars', [], state.folded)],
179 ['{{AGENTS}}', agentsBlock(state, aliases)],
180 ['{{TURNS}}', turnsBlock(state)],
181 ['{{LEDGER}}', ledgerBlock(state, aliases)],
182 ] as const).reduce((text, [placeholder, value]) => text.split(placeholder).join(value), JUDGE_PROMPT)
183
184const ID_SHAPE = /^[a-z-]+:[a-z0-9-]{1,40}$/
185
186// A slug naming the thing as the work spells it (`multi-agent:C6-review-rounds`) is the same waste as its
187// lowercase self, and case is not a judgement: the slug is lowered before the shape is tested, and the
188// finding is stored under the lowered id, so one behaviour keeps one id across runs and sessions.
189const lowerSlug = (id: string): string => {
190 const at = id.indexOf(':')
191 return at < 0 ? id : `${id.slice(0, at)}:${id.slice(at + 1).toLowerCase()}`
192}
193
194const isCategory = (v: unknown): v is Category =>
195 typeof v === 'string' &&
196 ['execution', 'reading', 'production', 'behavior', 'communication', 'multi-agent', 'environment', 'process', 'other'].includes(v)
197
198// undefined = the key is absent or malformed, or no visible row carries the pair: the finding is discarded.
199const signatureOf = (value: unknown, visible: Row[]): Signature | null | undefined => {
200 if (value === null) return null
201 if (!isRecord(value)) return undefined
202 const tool = str(value['tool'])
203 const key = str(value['key'])
204 return visible.some(r => r.tool === tool && r.key === key) ? { tool, key } : undefined
205}
206
207const signatureDrop = (value: unknown): string =>
208 isRecord(value)
209 ? `signature (${cell(value['tool'], 20)}, ${cell(value['key'], 40)}) matches no row`
210 : 'signature must be a (tool, key) pair or null'
211
212const estOf = (value: unknown, state: State, signature: Signature | null, evidence: string[]): number | null => {
213 if (signature !== null) return null
214 const claimed = typeof value === 'number' && Number.isFinite(value) ? Math.max(0, Math.round(value)) : 0
215 return Math.min(claimed, groundedCap(state, evidence))
216}
217
218const sameSignature = (a: Signature | null, b: Signature | null): boolean =>
219 a !== null && b !== null && a.tool === b.tool && a.key === b.key
220
221const isKept = (state: State, id: string, signature: Signature | null): boolean =>
222 state.patterns.some(p =>
223 p.decision === 'keep' && (p.id === id || sameSignature(p.signature, signature)))
224
225// One cited occurrence is a finding only after a first the ledger cannot show: stated intent, a user correction or a saved feedback memory.
226const FIRST_OCCURRENCE = /intent|\bcorrect(?:ed|ion)\b|\bfeedback\b/i
227
228// A string is the one short reason the finding was dropped; the object is the finding itself.
229const findingOf = (value: unknown, state: State, visible: Row[], aliases: ReadonlyMap<string, string>): Finding | string => {
230 if (!isRecord(value)) return 'not an object'
231 const given = str(value['id'])
232 const id = lowerSlug(given)
233 const category = value['category']
234 const kind = str(value['kind'])
235 const why = str(value['why'])
236 const alternative = str(value['alternative'])
237 const confidence = value['confidence']
238 if (!ID_SHAPE.test(id)) return `id ${cell(given, 40) || '(missing)'} is not <category>:<kebab-slug>`
239 if (!isCategory(category)) return 'category is not one of the nine'
240 if (id.slice(0, id.indexOf(':')) !== category) return `category ${category} does not match the id`
241 if (!kind.startsWith('Claude keeps ')) return 'kind must start with "Claude keeps "'
242 // A cap missed by a few characters is a sentence that ran long, not a wrong finding: the pane wraps, so it
243 // is kept as returned and the run report notes the length. Twice the cap is a different answer, and dropped.
244 if (kind.length > KIND_MAX * 2) return `kind is ${kind.length} chars, over twice ${KIND_MAX}`
245 if (alternative.length === 0) return 'alternative is empty'
246 if (alternative.length > ALTERNATIVE_MAX * 2) return `alternative is ${alternative.length} chars, over twice ${ALTERNATIVE_MAX}`
247 if (typeof confidence !== 'number' || confidence < 0.5 || confidence > 1) return 'confidence is not a number in 0.5..1'
248 const signature = signatureOf(value['signature'], visible)
249 if (signature === undefined) return signatureDrop(value['signature'])
250 if (isKept(state, id, signature)) return 'kept this session'
251 const evidence = evidenceOf(value['evidence'], state, visible, aliases, signature)
252 if (typeof evidence === 'string') return evidence
253 if (evidence.length < 2 && !FIRST_OCCURRENCE.test(why)) return 'one handle and no stated intent or correction in why'
254 return {
255 id, category, kind, evidence, signature, why, alternative, confidence,
256 estTokensPerTurn: estOf(value['est_tokens_per_turn'], state, signature, evidence),
257 proposal: proposalOf(value['proposal']),
258 lean: null,
259 }
260}
261
262// Each finding under the label the drop report names it by: its own id, or its place in the reply.
263type Reviewed = { label: string; finding: Finding } | { label: string; reason: string }
264type Sifted = { findings: Finding[]; dropped: string[] }
265
266const labelOf = (value: unknown, index: number): string => {
267 const id = lowerSlug(str(isRecord(value) ? value['id'] : ''))
268 return ID_SHAPE.test(id) ? id : `#${index + 1}`
269}
270
271const capFindings = (reviewed: readonly Reviewed[]): Sifted =>
272 reviewed.reduce<Sifted>((kept, item) => {
273 const dropped = (reason: string): Sifted => ({ findings: kept.findings, dropped: [...kept.dropped, `${item.label}: ${reason}`] })
274 if (!('finding' in item)) return dropped(item.reason)
275 if (kept.findings.length >= MAX_FINDINGS) return dropped(`over MAX_FINDINGS (${MAX_FINDINGS})`)
276 if (item.finding.signature === null && kept.findings.filter(k => k.signature === null).length >= MAX_BEHAVIORAL_FINDINGS) {
277 return dropped(`over MAX_BEHAVIORAL_FINDINGS (${MAX_BEHAVIORAL_FINDINGS})`)
278 }
279 if (kept.findings.some(k => k.id === item.finding.id)) return dropped('the reply already reported this id')
280 return { findings: [...kept.findings, item.finding], dropped: kept.dropped }
281 }, { findings: [], dropped: [] })
282
283const EXPLAIN_MAX = 200 // characters of the judge's `time` and `context` sentences kept; the rest is cut off
284
285// One sentence written for the user: an essay is cut to the sentence's length rather than thrown away,
286// because its first 200 characters still say where the time went. Anything but text is nothing to draw.
287const explanationOf = (value: unknown): string | null => {
288 const text = collapseWs(str(value))
289 return text.length > 0 ? short(text, EXPLAIN_MAX) : null
290}
291
292// An essay where a sentence was asked for is a prompt problem, so the run report says the cut happened:
293// without it a trimmed sentence reads like a sentence the judge chose to end there.
294const explanationDrop = (label: string, value: unknown): string[] => {
295 const text = collapseWs(str(value))
296 return text.length > EXPLAIN_MAX ? [`${label}: ${text.length} chars, trimmed to ${EXPLAIN_MAX}`] : []
297}
298
299// The caps a kept finding overran, reported the same way: nothing was cut, so the note is the whole story.
300const capNotes = (f: Finding): string[] => [
301 ...(f.kind.length > KIND_MAX ? [`${f.id}: kind: ${f.kind.length} chars, over ${KIND_MAX}`] : []),
302 ...(f.alternative.length > ALTERNATIVE_MAX ? [`${f.id}: alternative: ${f.alternative.length} chars, over ${ALTERNATIVE_MAX}`] : []),
303]
304
305/** What the judge said: the valid findings, the focus, its time and context sentences, one line per drop and per cap a kept finding missed, and how many it returned; never throws. `aliases` must be the table `buildPrompt` printed. */
306export const parseReply = (text: string, state: State, aliases: ReadonlyMap<string, string> = judgeAliases(state)): { findings: Finding[]; focus: string | null; time: string | null; context: string | null; dropped: string[]; returned: number } => {
307 const root = parseObject(text)
308 if (root === null) return { findings: [], focus: null, time: null, context: null, dropped: ['reply was not JSON'], returned: 0 }
309 const raw = root['findings']
310 const focusText = collapseWs(str(root['focus']))
311 const focus = focusText.length > 0 ? focusText : null
312 const said = { time: explanationOf(root['time']), context: explanationOf(root['context']) }
313 const overlong = [...explanationDrop('time', root['time']), ...explanationDrop('context', root['context'])]
314 if (!Array.isArray(raw)) return { findings: [], focus, ...said, dropped: [...overlong, 'findings was not an array'], returned: 0 }
315 const visible = citableRows(state)
316 const reviewed: Reviewed[] = raw.map((value, i) => {
317 const label = labelOf(value, i)
318 const result = findingOf(value, state, visible, aliases)
319 return typeof result === 'string' ? { label, reason: result } : { label, finding: result }
320 })
321 const sifted = capFindings(reviewed)
322 const noted = sifted.findings.flatMap(capNotes)
323 return { ...sifted, dropped: [...overlong, ...noted, ...sifted.dropped], focus, ...said, returned: raw.length }
324}
325
326const patternOf = (f: Finding): Pattern => ({
327 id: f.id, category: f.category, kind: f.kind, signature: f.signature, why: f.why,
328 alternative: f.alternative, confidence: f.confidence, proposal: f.proposal,
329 estTokensPerTurn: f.estTokensPerTurn, lastDecision: null, lean: f.lean,
330 hits: [...f.evidence], decision: null, decidedAtTurn: null, instruction: null, openedAtTurn: null, ignored: 0, sent: 0,
331})
332
333const updatedWith = (p: Pattern, f: Finding): Pattern => ({
334 ...p, why: f.why, alternative: f.alternative, proposal: f.proposal, confidence: f.confidence,
335 estTokensPerTurn: f.estTokensPerTurn, hits: unique([...p.hits, ...f.evidence]),
336})
337
338const hasRecurred = (state: State, p: Pattern, f: Finding): boolean =>
339 (p.decision === 'steer' || p.decision === 'kill') && p.decidedAtTurn !== null &&
340 citedTurns(state, f.evidence).some(turn => turn > (p.decidedAtTurn ?? 0))
341
342const rank = (p: Pattern): number => (p.decision === null ? 0 : 1000) + p.confidence
343
344// A cited pattern nobody has decided belongs in front of the user, whether the id is new, remembered by
345// the store or found by an earlier run: the only undecided pattern that is not news is one already queued.
346const isFresh = (state: State, patterns: Pattern[], id: string): boolean =>
347 patterns.find(p => p.id === id)?.decision === null && !state.cards.includes(id)
348
349const capPatterns = (patterns: Pattern[]): Pattern[] => {
350 if (patterns.length <= MAX_PATTERNS) return patterns
351 const dropped = [...patterns]
352 .sort((a, b) => rank(a) - rank(b))
353 .slice(0, patterns.length - MAX_PATTERNS)
354 .map(p => p.id)
355 return patterns.filter(p => !dropped.includes(p.id))
356}
357
358/** Folds findings into the complete registry, naming the fresh, the recurred, and the ids the cap evicted. */
359export const merge = (state: State, findings: Finding[]): { patterns: Pattern[]; fresh: string[]; recurred: string[]; evicted: string[] } => {
360 const added = findings.filter(f => !state.patterns.some(p => p.id === f.id))
361 const patterns = capPatterns([
362 ...state.patterns.map(p => {
363 const f = findings.find(x => x.id === p.id)
364 return f !== undefined ? updatedWith(p, f) : p
365 }),
366 ...added.map(patternOf),
367 ])
368 const kept = (id: string): boolean => patterns.some(p => p.id === id)
369 const recurred = state.patterns.filter(p => {
370 const f = findings.find(x => x.id === p.id)
371 return f !== undefined && hasRecurred(state, p, f)
372 })
373 return {
374 patterns,
375 fresh: findings.map(f => f.id).filter(id => isFresh(state, patterns, id)),
376 recurred: recurred.map(p => p.id).filter(kept),
377 // A validated finding the cap pushed out has no card: the shell reports it as dropped, not kept.
378 evicted: findings.map(f => f.id).filter(id => !kept(id)),
379 }
380}
381hooks/core/ledger.ts 221 lines1import type { ToolCallResult } from 'claude-code'
2
3import { collapseWs, stableJson } from './text'
4import { FILE_TOOLS, KEY_MAX, MAIN_AGENT } from './types'
5import type { CommandClass, Row } from './types'
6
7/** The flat tool event a row is built from: the tool, this call's id, the loop, and the tool's arguments beside them. */
8export type ToolEvent = { tool: string; tool_use_id: string; agentId?: string } & Record<string, unknown>
9
10const RESERVED = ['tool', 'tool_use_id', 'agentId', 'consent'] as const
11
12const HEAD_MAX = 80 // characters of result.text quoted as evidence (Row.head)
13
14// A leading `cd <dir> &&`, `VAR=value`, `timeout <duration>` or `time` is noise in front of the command that matters.
15const NOISE = /^(?:cd\s+[^\s&|;]+\s*&&\s*|[A-Za-z_][A-Za-z0-9_]*=(?:"[^"]*"|'[^']*'|\S*)\s+|timeout\s+\d+[smhd]?\s+|time\s+)/
16
17// `a && b`, `a; b`, `a || b`: one call can carry several commands, and quoting is not worth parsing.
18const CHAIN = /&&|\|\||;/
19
20// From the first pipe or redirect on (an fd digit like `2>` goes with it): how the output was filtered, not what ran.
21const FILTER = /\s*(?:\d?[<>]|\|).*$/
22
23// Script runners: what follows them is the command that matters (longest first).
24const RUNNERS = ['npm run', 'bun run', 'bun x', 'pnpm run', 'yarn run', 'python3 -m', 'python -m', 'npx', 'bunx', 'pnpm', 'yarn'] as const
25
26// Script names seen after a runner (`bun run lint`), which no binary in the table covers.
27const SCRIPTS: Readonly<Record<string, CommandClass>> = { test: 'test', lint: 'lint', format: 'format', typecheck: 'typecheck', build: 'build' }
28
29// Command heads, matched on a token boundary; two-token heads never shadow one another.
30const COMMANDS: readonly (readonly [string, CommandClass])[] = [
31 ['npm test', 'test'], ['bun test', 'test'], ['go test', 'test'], ['cargo test', 'test'],
32 ['jest', 'test'], ['vitest', 'test'], ['pytest', 'test'], ['mocha', 'test'],
33 ['biome lint', 'lint'], ['eslint', 'lint'], ['ruff', 'lint'], ['flake8', 'lint'], ['golangci-lint', 'lint'],
34 ['biome format', 'format'], ['prettier', 'format'], ['black', 'format'], ['gofmt', 'format'], ['rustfmt', 'format'],
35 ['tsc', 'typecheck'], ['mypy', 'typecheck'], ['pyright', 'typecheck'],
36 ['make build', 'build'], ['cargo build', 'build'], ['go build', 'build'], ['docker build', 'build'], ['vite', 'build'], ['webpack', 'build'],
37 ['npm install', 'install'], ['npm i', 'install'], ['bun install', 'install'], ['bun add', 'install'],
38 ['pnpm install', 'install'], ['pnpm add', 'install'], ['yarn install', 'install'], ['yarn add', 'install'],
39 ['pip install', 'install'], ['uv pip', 'install'], ['cargo fetch', 'install'], ['apt-get install', 'install'], ['brew install', 'install'],
40 ['git', 'git'], ['gh', 'git'],
41 ['cat', 'read'], ['head', 'read'], ['tail', 'read'], ['less', 'read'], ['ls', 'read'], ['tree', 'read'], ['sed', 'read'],
42 ['grep', 'search'], ['rg', 'search'], ['ag', 'search'], ['find', 'search'], ['fd', 'search'], ['ast-grep', 'search'],
43]
44
45const asRecord = (v: unknown): Record<string, unknown> => (typeof v === 'object' && v !== null ? (v as Record<string, unknown>) : {})
46
47const asString = (v: unknown): string | null => (typeof v === 'string' ? v : null)
48
49const asNumber = (v: unknown): number | null => (typeof v === 'number' && Number.isFinite(v) ? v : null)
50
51const pathArg = (args: Record<string, unknown>): string =>
52 asString(args.file_path) ?? asString(args.notebook_path) ?? ''
53
54const stripNoise = (command: string): string => {
55 let rest = command
56 while (NOISE.test(rest)) rest = rest.replace(NOISE, '')
57 return rest
58}
59
60const headToken = (command: string): string => command.split(' ')[0] ?? ''
61
62// The head by its basename: `venv/bin/python -m pytest` is `python -m pytest` to the table.
63const baseHead = (command: string): string => {
64 const head = headToken(command)
65 return `${head.slice(head.lastIndexOf('/') + 1)}${command.slice(head.length)}`
66}
67
68const tableClass = (command: string): CommandClass | null =>
69 COMMANDS.find(([head]) => command === head || command.startsWith(`${head} `))?.[1] ?? null
70
71const bareOf = (segment: string): string => stripNoise(collapseWs(segment))
72
73const segmentClass = (segment: string): CommandClass => {
74 const bare = baseHead(bareOf(segment))
75 const direct = tableClass(bare)
76 if (direct !== null) return direct
77 const runner = RUNNERS.find(r => bare === r || bare.startsWith(`${r} `))
78 if (runner === undefined) return 'other'
79 const script = bare.slice(runner.length).trim()
80 return tableClass(script) ?? SCRIPTS[headToken(script)] ?? 'other'
81}
82
83const segmentsOf = (command: string): string[] => collapseWs(command).split(CHAIN)
84
85/** Classifies a shell command by what it does, seeing through cd, env, timeout, runner prefixes and `&&`/`;`/`||` chains. */
86export const classOf = (command: string): CommandClass =>
87 segmentsOf(command)
88 .map(segmentClass)
89 .find(cls => cls !== 'other') ?? 'other'
90
91// The key names what ran: the segment that classified, as typed, minus the noise before it and the filtering after it.
92const canonicalOf = (command: string): string =>
93 bareOf(segmentsOf(command).find(segment => segmentClass(segment) !== 'other') ?? '').replace(FILTER, '').trim()
94
95/** Computes the ledger key and command class of a call from its arguments alone. */
96export const normalize = (tool: string, input: unknown): { key: string; cls: CommandClass } => {
97 const args = asRecord(input)
98 if (tool === 'Bash') {
99 const command = collapseWs(asString(args.command) ?? '')
100 const cls = classOf(command)
101 // An `other` command has no segment that named something: the whole pipeline is the key.
102 return { key: `${cls}:${cls === 'other' ? command : canonicalOf(command)}`.slice(0, KEY_MAX), cls }
103 }
104 if (FILE_TOOLS.includes(tool)) {
105 const path = pathArg(args)
106 const range = tool === 'Read' ? `:${asNumber(args.offset) ?? ''}-${asNumber(args.limit) ?? ''}` : ''
107 return { key: `${path}${range}`.slice(0, KEY_MAX), cls: tool === 'Read' ? 'read' : 'other' }
108 }
109 if (tool === 'Grep' || tool === 'Glob') {
110 return { key: `${tool}:${asString(args.pattern) ?? ''}:${asString(args.path) ?? ''}`.slice(0, KEY_MAX), cls: 'search' }
111 }
112 if (tool === 'Agent') return { key: `agent:${asString(args.subagent_type) ?? 'general'}`.slice(0, KEY_MAX), cls: 'other' }
113 return { key: `${tool}:${stableJson(input, RESERVED)}`.slice(0, KEY_MAX), cls: 'other' }
114}
115
116const printable = (ch: string): boolean => {
117 const code = ch.codePointAt(0) ?? 0
118 return code >= 0x20 && code !== 0x7f
119}
120
121// Code point by code point, stopping before the one that would not fit whole: never half a surrogate pair.
122const headOf = (text: string | undefined): string => {
123 let head = ''
124 for (const ch of text ?? '') {
125 const kept = ch === '\t' || ch === '\n' ? ' ' : ch
126 if (!printable(kept)) continue
127 if (head.length + kept.length > HEAD_MAX) break
128 head += kept
129 }
130 return head
131}
132
133const flagsOf = (e: ToolEvent, result: ToolCallResult, res: Record<string, unknown>): string[] => {
134 const persisted = e.tool === 'Bash' ? asNumber(res.persistedOutputSize) : null
135 // `bg` describes a call that ran: a denied or errored Bash started no background task.
136 const isBackground = answered(result) && (e.run_in_background === true || asString(res.backgroundTaskId) !== null)
137 // An ask's `ms` is the wait for the person; `recommended` says the model marked an option for them.
138 const isAsk = e.tool === 'AskUserQuestion'
139 return [
140 result.isError === true ? 'err' : '',
141 result.deny !== undefined ? 'denied' : '',
142 e.tool === 'Read' && res.type === 'file_unchanged' ? 'dedup' : '',
143 e.tool === 'Read' && asRecord(res.file).truncatedByTokenCap === true ? 'trunc' : '',
144 e.tool === 'Bash' && isBackground ? 'bg' : '',
145 e.tool === 'Bash' && asNumber(res.timedOutAfterMs) !== null ? 'timeout' : '',
146 persisted !== null ? `persist=${persisted}` : '',
147 isAsk ? 'ask' : '',
148 isAsk && JSON.stringify(e.questions ?? '').includes('(Recommended)') ? 'recommended' : '',
149 ].filter(flag => flag !== '')
150}
151
152const answered = (result: ToolCallResult): boolean => result.deny === undefined && result.isError !== true
153
154const patchLines = (patch: unknown): { add: number; del: number } => {
155 const hunks = Array.isArray(patch) ? patch : []
156 const lines = hunks.flatMap(hunk => {
157 const own = asRecord(hunk).lines
158 return Array.isArray(own) ? own.filter((line): line is string => typeof line === 'string') : []
159 })
160 return { add: lines.filter(line => line.startsWith('+')).length, del: lines.filter(line => line.startsWith('-')).length }
161}
162
163const lineCount = (content: string): number => (content === '' ? 0 : content.replace(/\n$/, '').split('\n').length)
164
165const linesOf = (e: ToolEvent, res: Record<string, unknown>): { add: number; del: number } | null => {
166 if (e.tool === 'Edit') {
167 const diff = asRecord(res.gitDiff)
168 const add = asNumber(diff.additions)
169 const del = asNumber(diff.deletions)
170 return add !== null && del !== null ? { add, del } : patchLines(res.structuredPatch)
171 }
172 if (e.tool === 'Write') return { add: lineCount(asString(res.content) ?? asString(e.content) ?? ''), del: 0 }
173 return null
174}
175
176const pathsOf = (e: ToolEvent, res: Record<string, unknown>): string[] => {
177 if (e.tool === 'Edit' || e.tool === 'Write' || e.tool === 'NotebookEdit') {
178 const path = asString(res.filePath) ?? (pathArg(e) || null)
179 return path !== null && res.staged !== true ? [path] : []
180 }
181 if (e.tool === 'Bash') {
182 const changed = asRecord(res.bashEditDiff).changedFiles
183 return Array.isArray(changed) ? changed.filter((file): file is string => typeof file === 'string') : []
184 }
185 return []
186}
187
188const spawnOf = (e: ToolEvent, res: Record<string, unknown>): Row['spawn'] =>
189 e.tool !== 'Agent'
190 ? null
191 : {
192 type: asString(e.subagent_type) ?? 'general',
193 requested: asString(e.model),
194 resolved: asString(res.resolvedModel),
195 status: asString(res.status),
196 tokens: asNumber(res.totalTokens),
197 edits: asNumber(asRecord(res.toolStats).editFileCount),
198 promptChars: (asString(e.prompt) ?? '').length,
199 }
200
201/** Builds the ledger row of a finished tool call, reading every result field defensively. */
202export const rowOf = (e: ToolEvent, result: ToolCallResult, ms: number, turn: number): Omit<Row, 'seq'> => {
203 const res = asRecord(result.result)
204 const { key, cls } = normalize(e.tool, e)
205 return {
206 id: e.tool_use_id,
207 tool: e.tool,
208 key,
209 cls,
210 agent: asString(e.agentId) ?? MAIN_AGENT,
211 turn,
212 ms,
213 chars: result.text?.length ?? 0,
214 head: headOf(result.text),
215 flags: flagsOf(e, result, res),
216 lines: answered(result) ? linesOf(e, res) : null,
217 paths: answered(result) ? pathsOf(e, res) : [],
218 spawn: spawnOf(e, res),
219 }
220}
221hooks/core/patterns.ts 872 lines1import { agentAliases, aliasOf, baseline, foldRows, rowsOf, sinks, sumOf } from './evidence'
2import { measuredCost } from './process'
3import { activeRuns, countRow } from './spawns'
4import { collapseWs, duration, instructionOf, killPrompt, kilo, median, pctOf } from './text'
5import {
6 ALTERNATIVE_MAX, CARD_EVIDENCE, DEBUG_MAX_DROPPED, DEBUG_MAX_LINES, DEBUG_MAX_PATTERNS, FILE_TOOLS,
7 JUDGE_BUDGET_SHARE, JUDGE_MAX_BACKOFF, KEY_MAX, KIND_MAX, LOOP_CAP, MAIN_AGENT, MAX_PATTERNS, NO_CALLS, ROW_CAP,
8 SETTLE_TURNS, TREND_TURNS, initialState,
9} from './types'
10import type {
11 Action, Artifact, BandModel, Card, Choice, CommandClass, DecidedRow, Evidence, Header, JournalEntry, JudgeRun,
12 JudgeUsage, Loop, PaneModel, Pattern, Proposal, Row, Signature, Sinks, State, StoredPattern, Tokens, TurnStat,
13} from './types'
14
15// The `:offset-limit` slice `normalize` appends to a Read key: the path is what the details name.
16const READ_RANGE = /:\d*-\d*$/
17
18// The loop an `agent:<id>` evidence handle names.
19const AGENT_HANDLE = /^agent:(.+)$/
20
21const GROWTH_TURNS = 5 // turns with a context sample the pace to compaction is read from
22const GROWTH_SAMPLES = 3 // growth samples below which no pace is stated at all
23
24type Decided = Pattern & { decision: Choice }
25type Settle = { pattern: Pattern; ms: number; chars: number; requeue: string | null }
26
27const isSent = (c: Choice | null): boolean => c === 'steer' || c === 'kill'
28
29const isDecided = (p: Pattern): p is Decided => p.decision !== null
30
31const pushUnique = (xs: readonly string[], x: string): string[] => (xs.includes(x) ? [...xs] : [...xs, x])
32
33const queueCard = (cards: readonly string[], id: string): string[] => (cards.includes(id) ? [...cards] : [id, ...cards])
34
35const patternById = (patterns: readonly Pattern[], id: string): Pattern | undefined => patterns.find(p => p.id === id)
36
37const signatureHit = (p: Pattern, row: Pick<Row, 'tool' | 'key'>): boolean =>
38 p.signature !== null && p.signature.tool === row.tool && p.signature.key === row.key
39
40const classOfPattern = (state: State, p: Pattern): CommandClass | null => rowsOf(state, p)[0]?.cls ?? null
41
42const turnHandle = (handle: string): number | null => {
43 const n = /^turn:(\d+)$/.exec(handle)?.[1]
44 return n === undefined ? null : Number(n)
45}
46
47const loopsCited = (p: Pattern, state: State): Loop[] =>
48 p.hits.flatMap(handle => {
49 const id = AGENT_HANDLE.exec(handle)?.[1]
50 const loop = id === undefined ? undefined : state.loops.find(l => l.id === id)
51 return loop === undefined ? [] : [loop]
52 })
53
54const turnsCited = (p: Pattern, state: State, rows: readonly Row[]): number[] => {
55 const fromHandles = p.hits.map(turnHandle).filter((n): n is number => n !== null)
56 return [...rows.map(r => r.turn), ...fromHandles, ...loopsCited(p, state).map(l => l.firstTurn)].sort((a, b) => a - b)
57}
58
59const newTokens = (t: TurnStat): number => t.input + t.output + t.cacheCreate
60
61/** Sums every new (non-cache-read) token this session's turns reported. */
62export const totalTokens = (state: State): number => state.turns.reduce((n, t) => n + newTokens(t), 0)
63
64/** Tokens left before auto-compaction; null while the session's token count is unknown. */
65export const tokensToCompaction = (state: State): number | null => {
66 const tokens = state.usage.tokens
67 if (tokens === undefined) return null
68 return (state.usage.compactAt ?? Math.round(state.usage.window * 0.9)) - tokens
69}
70
71/** The context after each turn that reported one, oldest first: how full the window was, turn by turn. */
72const contexts = (state: State): number[] =>
73 state.turns.map(t => t.context).filter((tokens): tokens is number => tokens !== null)
74
75/**
76 * Turns left before compaction at the recent pace; null under three growth samples or a zero median.
77 *
78 * The pace is how fast the window fills, not what a turn is billed: a turn can spend 60k tokens and
79 * grow the window by 8k, so the tokens of a turn would have claimed compaction was three turns away.
80 * A turn that shrank the window (a compaction, a `/clear`) is no pace at all, so drops are skipped.
81 */
82export const turnsToCompaction = (state: State): number | null => {
83 const left = tokensToCompaction(state)
84 const seen = contexts(state).slice(-GROWTH_TURNS)
85 const growth = seen.slice(1).map((tokens, at) => tokens - (seen[at] ?? 0)).filter(step => step > 0)
86 if (left === null || growth.length < GROWTH_SAMPLES) return null
87 const perTurn = median(growth)
88 return perTurn === 0 ? null : Math.round(left / perTurn)
89}
90
91const applySettlements = (state: State, list: readonly Settle[]): Pick<State, 'patterns' | 'cards' | 'saved'> => ({
92 patterns: list.map(s => s.pattern),
93 cards: list.reduce<string[]>((cards, s) => (s.requeue === null ? cards : queueCard(cards, s.requeue)), [...state.cards]),
94 saved: {
95 ms: list.reduce((ms, s) => ms + s.ms, state.saved.ms),
96 chars: list.reduce((chars, s) => chars + s.chars, state.saved.chars),
97 },
98})
99
100const settleWithRow = (state: State, p: Pattern, row: Omit<Row, 'seq'>): Settle => {
101 const grown: Pattern = signatureHit(p, row) ? { ...p, hits: pushUnique(p.hits, row.id) } : p
102 const still = { pattern: grown, ms: 0, chars: 0, requeue: null }
103 if (row.agent !== 'main' || p.openedAtTurn === null || !isSent(p.decision)) return still
104 // A row in the decision's own turn was already in flight before Claude could read the instruction.
105 if (row.turn <= p.openedAtTurn) return still
106 // D4: every ignored instruction brings the card back, so the user can Ignore it or say something else.
107 if (p.signature !== null && row.key === p.signature.key) {
108 return { pattern: { ...grown, ignored: p.ignored + 1, openedAtTurn: null }, ms: 0, chars: 0, requeue: p.id }
109 }
110 if (row.cls !== classOfPattern(state, p)) return still
111 const base = baseline(state, p)
112 return {
113 pattern: { ...grown, openedAtTurn: null },
114 ms: Math.max(0, base.ms - row.ms),
115 chars: Math.max(0, base.chars - row.chars),
116 requeue: null,
117 }
118}
119
120const settleAtTurn = (state: State, p: Pattern): Settle => {
121 // D4: an ignored instruction saves nothing, so a behavioural pattern accrues only while it is still believed.
122 const accrued = p.signature === null && isSent(p.decision) && p.ignored === 0 ? (p.estTokensPerTurn ?? 0) * 4 : 0
123 if (p.openedAtTurn === null || state.turn - p.openedAtTurn < SETTLE_TURNS) {
124 return { pattern: p, ms: 0, chars: accrued, requeue: null }
125 }
126 const base = baseline(state, p)
127 return { pattern: { ...p, openedAtTurn: null }, ms: base.ms, chars: base.chars + accrued, requeue: null }
128}
129
130const NO_TOKENS: Tokens = { input: 0, output: 0, cacheRead: 0, cacheCreate: 0 }
131
132const addTokens = (a: Tokens, b: Tokens): Tokens =>
133 ({ input: a.input + b.input, output: a.output + b.output, cacheRead: a.cacheRead + b.cacheRead, cacheCreate: a.cacheCreate + b.cacheCreate })
134
135// A loop known by nothing but its id yet: dated by the turn and the row that first showed it.
136const bareLoop = (id: string, firstTurn: number, firstSeq: number): Loop => ({
137 id, run: null, label: null, phase: null, model: null, turns: 0, ms: 0, tokens: NO_TOKENS, ended: null, firstTurn, firstSeq, outcome: null,
138 calls: 0, edits: 0, checks: 0, reads: 0,
139})
140
141// Every change to a loop: created where unknown, grown where known, the oldest dropped past the cap.
142const withLoop = (loops: readonly Loop[], id: string, grow: (l: Loop) => Loop, first: { turn: number; seq: number }): Loop[] => {
143 const known = loops.some(l => l.id === id)
144 const grown = known ? loops.map(l => (l.id === id ? grow(l) : l)) : [...loops, grow(bareLoop(id, first.turn, first.seq))]
145 return grown.slice(-LOOP_CAP)
146}
147
148const applyLoopTurn = (state: State, a: Extract<Action, { type: 'loop.turn' }>): State => ({
149 ...state,
150 loops: withLoop(state.loops, a.agentId, l => ({
151 ...l, turns: l.turns + 1, ms: l.ms + a.ms, tokens: addTokens(l.tokens, a.tokens), ended: a.ended, model: a.model ?? l.model,
152 }), { turn: a.turn, seq: state.seq }),
153})
154
155const applyAgentStart = (state: State, a: Extract<Action, { type: 'agent.start' }>): State => ({
156 ...state,
157 loops: withLoop(state.loops, a.agentId, l => ({
158 ...l, label: a.description.length > 0 ? a.description : l.label, model: a.model ?? l.model,
159 }), { turn: state.turn, seq: state.seq }),
160})
161
162const applyRunStart = (state: State, run: { id: string; name: string; dir: string | null }, now: number): State =>
163 state.runs.some(r => r.id === run.id)
164 ? state // a resume: the run is the one already known, dated at its first launch
165 : { ...state, runs: [...state.runs, { ...run, turn: state.turn, seq: state.seq, at: now, refreshedAt: 0, phases: [] }] }
166
167// A `started` entry names the loop's run and stage; a `result` entry what it returned. Either creates the loop.
168const applyEntry = (state: State, runId: string, loops: readonly Loop[], entry: JournalEntry): Loop[] =>
169 withLoop(loops, entry.agentId, l =>
170 entry.kind === 'started'
171 ? { ...l, run: runId, label: entry.label ?? l.label, phase: entry.phase ?? l.phase }
172 : { ...l, outcome: entry.outcome }, { turn: state.turn, seq: state.seq })
173
174const applyRunJournal = (state: State, a: Extract<Action, { type: 'run.journal' }>): State => ({
175 ...state,
176 loops: a.entries.reduce((loops, entry) => applyEntry(state, a.runId, loops, entry), state.loops),
177 runs: state.runs.map(r => (r.id === a.runId ? { ...r, refreshedAt: a.now } : r)),
178})
179
180// The wait before this prompt is written onto the turn it followed, so TURNS can say the person was away.
181const applyTurnStart = (state: State, now: number): State => {
182 const last = state.turns[state.turns.length - 1]
183 const dated = last !== undefined && last.at > 0 ? { ...last, idleMs: Math.max(0, now - last.at) } : last
184 return { ...state, turn: state.turn + 1, turns: dated === undefined ? state.turns : [...state.turns.slice(0, -1), dated] }
185}
186
187// The ledger with the cap applied: a row it pushes off the front is folded into its pair, never simply lost,
188// so STATS, the LEDGER's `~` lines and the sinks still count a call the judge can no longer cite.
189const capped = (state: State, rows: readonly Row[]): Pick<State, 'rows' | 'folded'> => {
190 const over = rows.length - ROW_CAP
191 return over <= 0
192 ? { rows: [...rows], folded: state.folded }
193 : { rows: rows.slice(over), folded: foldRows(state.folded, rows.slice(0, over)) }
194}
195
196const applyRow = (state: State, row: Omit<Row, 'seq'>): State => {
197 const seq = state.seq + 1
198 return {
199 ...state,
200 seq,
201 ...capped(state, [...state.rows, { ...row, seq }]),
202 // A row of a loop nobody has named yet names it here, so the aliases stay in the ledger's order; every
203 // row of a loop is counted on it here, so its line reads whole after ROW_CAP has dropped the row.
204 loops: row.agent === MAIN_AGENT ? state.loops : withLoop(state.loops, row.agent, l => countRow(l, row), { turn: row.turn, seq }),
205 ...applySettlements(state, state.patterns.map(p => settleWithRow(state, p, row))),
206 }
207}
208
209const grownWith = (p: Pattern, rows: readonly Row[]): Pattern =>
210 rows.filter(row => signatureHit(p, row)).reduce((q, row) => ({ ...q, hits: pushUnique(q.hits, row.id) }), p)
211
212// History, not a live call: no turn stats exist for it, nothing settles on it, and nothing was saved by it.
213const applyAdopt = (state: State, rows: readonly Omit<Row, 'seq'>[]): State => {
214 const seeded = rows.map((row, i) => ({ ...row, seq: state.seq + i + 1 }))
215 const seq = state.seq + seeded.length
216 return {
217 ...state,
218 seq,
219 turn: Math.max(state.turn, ...seeded.map(row => row.turn)),
220 ...capped(state, [...state.rows, ...seeded]),
221 patterns: state.patterns.map(p => grownWith(p, seeded)),
222 // The mid-turn cadence starts where the history ends: adopted rows are not new work, so a session
223 // joined late waits for JUDGE_MIN_NEW_ROWS of its own. `/saver check` still judges them on request.
224 judge: { ...state.judge, lastAtSeq: seq },
225 }
226}
227
228const applyTurnComplete = (state: State, stat: Omit<TurnStat, 'turn' | 'calls'>): State => ({
229 ...state,
230 turns: [...state.turns, { ...stat, turn: state.turn, calls: state.rows.filter(r => r.turn === state.turn).length }],
231 ...applySettlements(state, state.patterns.map(p => settleAtTurn(state, p))),
232})
233
234// Present values win; `window`/`compactAt` sticky (§5.2, widened so a partial usage dispatch never blanks the header).
235// The first sample also dates the mid-turn cadence: without it `lastAtMs` is 0 against a clock reading ms
236// since the epoch, so the five-minute floor would be no floor at all. `/saver demo` passes 0 and seeds nothing.
237const applyUsage = (state: State, usage: State['usage'], now: number): State => ({
238 ...state,
239 usage: {
240 window: usage.window || state.usage.window,
241 compactAt: usage.compactAt ?? state.usage.compactAt,
242 tokens: usage.tokens ?? state.usage.tokens,
243 percent: usage.percent ?? state.usage.percent,
244 },
245 judge: state.judge.lastAtMs === 0 && now > 0 ? { ...state.judge, lastAtMs: now } : state.judge,
246})
247
248const REPLAN = 'Change the remaining work: '
249
250// The one-time re-plan a process Fix sends, said once even when the finding's own words already say it.
251const replanOf = (text: string): string => (text.trimStart().toLowerCase().startsWith(REPLAN.toLowerCase()) ? text.trim() : `${REPLAN}${text.trim()}`)
252
253const applyDecide = (state: State, id: string, choice: Choice, text: string | undefined): State => {
254 const p = patternById(state.patterns, id)
255 if (p === undefined) return state
256 const process = p.lean !== null
257 // A process Fix sends the lean alternative itself: `killPrompt` stops a behaviour, a re-plan changes the work.
258 const instruction = choice === 'kill' ? (process ? p.alternative : killPrompt(p)) : (text ?? '')
259 if (choice !== 'keep' && instruction.trim() === '') return state
260 const sending = choice !== 'keep'
261 // Fix rides the same wrapper as Fix… (§5.2 "same with killPrompt(p)", Appendix C 5c "the same way"):
262 // Claude reads `Instruction from the user (via ContextSaver): Stop this behaviour …`.
263 const note = instructionOf(process ? replanOf(instruction) : instruction)
264 const decided: Pattern = {
265 ...p,
266 decision: choice,
267 decidedAtTurn: state.turn,
268 lastDecision: choice,
269 instruction: sending ? instruction : p.instruction,
270 openedAtTurn: sending ? state.turn : p.openedAtTurn,
271 }
272 return {
273 ...state,
274 patterns: state.patterns.map(q => (q.id === p.id ? decided : q)),
275 cards: state.cards.filter(c => c !== p.id),
276 expanded: state.expanded === p.id ? null : state.expanded,
277 steering: null,
278 steerDraft: null,
279 notes: sending ? pushUnique(state.notes, note) : [...state.notes],
280 // A re-plan is the orchestrator's, once (D3): standing would hand it to every subagent spawned after.
281 standing: sending && !process ? pushUnique(state.standing, note) : [...state.standing],
282 }
283}
284
285// A process run's patterns join the registry beside the habit judge's; its cadence and cost are its own.
286const applyProcessDone = (state: State, a: Extract<Action, { type: 'process.done' }>): State => {
287 const patterns = [...a.patterns, ...state.patterns.filter(p => patternById(a.patterns, p.id) === undefined)]
288 const wanted = a.fresh.filter(id => patternById(patterns, id)?.decision === null && !state.cards.includes(id))
289 const total = totalTokens(state)
290 const spent = state.process.spent + a.spent
291 return {
292 ...state,
293 patterns,
294 cards: [...wanted, ...state.cards],
295 process: {
296 ...state.process,
297 running: false,
298 runs: state.process.runs + 1,
299 spent,
300 backoff: total > 0 && spent > JUDGE_BUDGET_SHARE * total ? state.process.backoff * 2 : state.process.backoff,
301 error: a.error,
302 last: { returned: a.returned, kept: a.kept, dropped: a.dropped, usage: a.usage },
303 },
304 }
305}
306
307// One agent reached by the standing instructions: counted once, on every pattern whose instruction is standing.
308const applyDeliverySent = (state: State, agentId: string): State =>
309 state.delivery.agents.includes(agentId)
310 ? state
311 : {
312 ...state,
313 delivery: { ...state.delivery, agents: [...state.delivery.agents, agentId] },
314 patterns: state.patterns.map(p => (p.lean === null && p.instruction !== null && state.standing.some(t => t.includes(p.instruction ?? '')) ? { ...p, sent: p.sent + 1 } : p)),
315 }
316
317const applyJudgeDone = (state: State, a: Extract<Action, { type: 'judge.done' }>): State => {
318 const recurred = (p: Pattern): boolean => a.recurred.includes(p.id) && isSent(p.decision)
319 const marked = a.patterns.filter(recurred).map(p => p.id)
320 const reported = a.patterns.map(p => (recurred(p) ? { ...p, ignored: p.ignored + 1, openedAtTurn: null } : p))
321 // A session decision is never lost, even when the judge's registry omits it.
322 const patterns = [...reported, ...state.patterns.filter(p => isDecided(p) && patternById(reported, p.id) === undefined)]
323 const wanted = [...a.fresh.filter(id => patternById(patterns, id)?.decision === null), ...marked]
324 const added = wanted.filter((id, i) => wanted.indexOf(id) === i && !state.cards.includes(id))
325 const total = totalTokens(state)
326 const spent = state.judge.spent + a.spent
327 return {
328 ...state,
329 patterns,
330 cards: [...added, ...state.cards].filter(id => patternById(patterns, id) !== undefined),
331 judge: {
332 lastAtTokens: total,
333 lastAtTurn: state.turn,
334 lastAtSeq: state.judge.lastAtSeq,
335 lastAtMs: state.judge.lastAtMs,
336 running: false,
337 runs: state.judge.runs + 1,
338 spent,
339 // No completed turn is no budget to be over: the first run of a reloaded session doubles nothing.
340 backoff: total > 0 && spent > JUDGE_BUDGET_SHARE * total ? Math.min(state.judge.backoff * 2, JUDGE_MAX_BACKOFF) : state.judge.backoff,
341 error: a.error,
342 focus: a.focus,
343 time: a.time,
344 context: a.context,
345 last: { returned: a.returned, kept: a.kept, dropped: [...a.dropped], usage: a.usage },
346 },
347 // Only a run that answered spends the arming: a cold snapshot or a refusal leaves it for the next opportunity.
348 pendingCheck: a.error === null ? false : state.pendingCheck,
349 }
350}
351
352const applyReset = (state: State): State => ({
353 ...initialState(state.cwd, state.usage.window),
354 overhead: state.overhead,
355 columns: state.columns,
356 paneOpen: state.paneOpen,
357 patterns: state.patterns.map(p => fromStored(toStored(p))),
358})
359
360/** Applies one action to the state, returning a new state (never mutates its input). */
361export const reduce = (state: State, action: Action): State => {
362 switch (action.type) {
363 case 'turn.start':
364 return applyTurnStart(state, action.now)
365 case 'loop.turn':
366 return applyLoopTurn(state, action)
367 case 'agent.start':
368 return applyAgentStart(state, action)
369 case 'run.start':
370 return applyRunStart(state, action.run, action.now)
371 case 'run.journal':
372 return applyRunJournal(state, action)
373 case 'row':
374 return applyRow(state, action.row)
375 case 'adopt':
376 return applyAdopt(state, action.rows)
377 case 'turn.complete':
378 return applyTurnComplete(state, action.stat)
379 case 'usage':
380 return applyUsage(state, action.usage, action.now)
381 case 'overhead':
382 return { ...state, overhead: action.overhead }
383 case 'compact':
384 // The fill a compaction invalidated is forgotten: `tokens` is sticky and a just-compacted window
385 // reports none until its next response (d.ts 7023-7025), so the header awaits the next turn instead
386 // of announcing two turns to a compaction that just happened.
387 return { ...state, compactions: [...state.compactions, state.turn], usage: { ...state.usage, tokens: undefined, percent: undefined } }
388 case 'expand':
389 return { ...state, expanded: action.patternId === state.expanded ? null : action.patternId }
390 case 'steer.begin':
391 return { ...state, steering: state.steering === action.patternId ? null : action.patternId, steerDraft: null }
392 case 'steer.draft':
393 return { ...state, steerDraft: action.text }
394 case 'decide':
395 return applyDecide(state, action.patternId, action.choice, action.text)
396 case 'judge.start':
397 return { ...state, judge: { ...state.judge, running: true, lastAtMs: action.now, lastAtSeq: action.seq } }
398 case 'judge.done':
399 return applyJudgeDone(state, action)
400 case 'ask':
401 return { ...state, asks: [...state.asks, action.ask] }
402 case 'run.phases':
403 return { ...state, runs: state.runs.map(r => (r.id === action.runId ? { ...r, phases: [...action.phases] } : r)) }
404 case 'process.start':
405 return { ...state, process: { ...state.process, running: true, lastAtMs: action.now, lastAtSeq: action.seq, lastAtTurn: state.turn } }
406 case 'process.done':
407 return applyProcessDone(state, action)
408 case 'process.cold':
409 return { ...state, process: { ...state.process, ...action.was, running: false } }
410 case 'delivery.pending':
411 return { ...state, delivery: { ...state.delivery, pending: Math.max(0, state.delivery.pending + action.delta) } }
412 case 'delivery.sent':
413 return applyDeliverySent(state, action.agentId)
414 case 'delivery.toasted':
415 return { ...state, delivery: { ...state.delivery, lastToastAt: action.now } }
416 case 'check.arm':
417 return { ...state, pendingCheck: true }
418 case 'notes.drained':
419 return { ...state, notes: [] }
420 case 'standing.add':
421 return { ...state, standing: pushUnique(state.standing, action.text) }
422 case 'artifact.done':
423 return {
424 ...state,
425 patterns: state.patterns.map(p => (p.id === action.patternId ? { ...p, proposal: null } : p)),
426 written: action.written ? pushUnique(state.written, `${action.patternId}:${action.kind}`) : [...state.written],
427 }
428 case 'pane':
429 return { ...state, paneOpen: action.open, autoOpened: action.auto === true ? true : state.autoOpened }
430 case 'columns':
431 return { ...state, columns: action.columns }
432 case 'reset':
433 return applyReset(state)
434 }
435}
436
437const statsOf = (p: Pattern, state: State, rows: readonly Row[]): string => {
438 const cost = sumOf(rows)
439 const pct = pctOf(cost.chars, state.usage.window)
440 const turns = turnsCited(p, state, rows)
441 const first = turns[0]
442 const last = turns[turns.length - 1]
443 // Zero segments are dropped whole: turn handles and rebuilt rows carry no duration, and `0s` would claim a suite that ran for minutes cost nothing.
444 return [
445 `${p.hits.length}×`,
446 ...(pct > 0 ? [`~${pct}% of context`] : []),
447 ...(cost.ms > 0 ? [duration(cost.ms)] : []),
448 // One turn is not a range: 'turns 1–1' reads as a bug.
449 ...(first === undefined || last === undefined ? [] : [first === last ? `turn ${first}` : `turns ${first}–${last}`]),
450 ].join(' · ')
451}
452
453// What ran, in the words the person typed or read: the command, the file, or the tool and its key.
454const whatRan = (r: Row): string => {
455 if (r.tool === 'Bash') return r.key.startsWith(`${r.cls}:`) ? r.key.slice(r.cls.length + 1) : r.key
456 if (FILE_TOOLS.includes(r.tool)) return r.key.replace(READ_RANGE, '')
457 return `${r.tool} ${r.key}`
458}
459
460const citedRows = (rows: readonly Row[], aliases: ReadonlyMap<string, string>): Evidence[] =>
461 rows.map(r => ({
462 turn: r.turn,
463 what: whatRan(r),
464 // The loop is named only when it was not the main one: an alias on every row would be noise.
465 agent: r.agent === MAIN_AGENT ? null : aliasOf(aliases, r.agent),
466 ms: r.ms,
467 chars: r.chars,
468 head: r.head,
469 }))
470
471const citedTurnStats = (p: Pattern, state: State): TurnStat[] =>
472 p.hits
473 .map(turnHandle)
474 .map(n => (n === null ? undefined : state.turns.find(t => t.turn === n)))
475 .filter((t): t is TurnStat => t !== undefined)
476
477const ktokOf = (l: Loop): number => Math.round((l.tokens.input + l.tokens.cacheCreate + l.tokens.output) / 1000)
478
479// A cited loop, quoted as one line of evidence: its stage and model, dated by the turn that spawned it.
480const citedLoop = (state: State, aliases: ReadonlyMap<string, string>, l: Loop): Evidence => ({
481 turn: l.firstTurn,
482 what: `${l.label ?? 'agent'} · ${l.model ?? '?'}`,
483 agent: aliasOf(aliases, l.id),
484 ms: l.ms,
485 chars: 0,
486 head: `${ktokOf(l)}k tokens · ${l.edits} edits`,
487})
488
489const evidenceOf = (p: Pattern, state: State, rows: readonly Row[], aliases: ReadonlyMap<string, string>): Evidence[] => {
490 const turns: Evidence[] = citedTurnStats(p, state)
491 .map(t => ({ turn: t.turn, what: NO_CALLS, agent: null, ms: 0, chars: t.answerChars, head: t.answerHead }))
492 const loops: Evidence[] = loopsCited(p, state).map(l => citedLoop(state, aliases, l))
493 // Reversed first, so two calls inside one turn also read newest-first once the stable sort has run.
494 return [...citedRows(rows, aliases).reverse(), ...loops.reverse(), ...turns]
495 .sort((a, b) => b.turn - a.turn)
496 .slice(0, CARD_EVIDENCE)
497}
498
499// A pattern with no row in hand cites loops or turns, not calls: it counts them, and its context cost is the
500// judge's per-turn estimate. The unit is stated here, so the drawing never has to guess it back.
501const totalOf = (p: Pattern, state: State, rows: readonly Row[]): Card['total'] => {
502 if (rows.length > 0) return { unit: 'calls', calls: rows.length, ...sumOf(rows) }
503 const loops = loopsCited(p, state)
504 const chars = (p.estTokensPerTurn ?? 0) * 4
505 if (loops.length > 0) return { unit: 'agents', calls: loops.length, ms: loops.reduce((ms, l) => ms + l.ms, 0), chars }
506 return { unit: 'turns', calls: citedTurnStats(p, state).length, ms: 0, chars }
507}
508
509// A process card's Cost so far: the newest run's claim, clamped to its cited handles; what they measured when none is held.
510const costSoFar = (state: State, p: Pattern): string => {
511 const cost = p.cost ?? measuredCost(state, p.hits)
512 return `${duration(cost.ms)} · ${kilo(cost.tokens)} tokens`
513}
514
515/** Derives the waster card in seat `n` from a pattern, the evidence it cites and the ledger's loop aliases. */
516export const cardOf = (p: Pattern, state: State, n: number, aliases: ReadonlyMap<string, string>): Card => {
517 const rows = rowsOf(state, p)
518 return {
519 patternId: p.id,
520 n,
521 category: p.category,
522 kind: p.ignored > 0 ? `ignored · ${p.kind}` : p.kind,
523 stats: p.lean === null ? statsOf(p, state, rows) : costSoFar(state, p),
524 why: p.why,
525 fix: p.alternative,
526 lean: p.lean,
527 total: totalOf(p, state, rows),
528 evidence: evidenceOf(p, state, rows, aliases),
529 }
530}
531
532// Null before the first turn completes: the session's own tokens are 0 then, and a share of nothing is
533// not a figure — a run that cost 24k printed `2394600%` of a denominator that had not been measured yet.
534const judgeShare = (state: State): number | null => {
535 const total = totalTokens(state)
536 return total === 0 ? null : Math.round((state.judge.spent / total) * 1000) / 10
537}
538
539// How full the window was after each of the last turns that reported it: the shape the header draws.
540const trendOf = (state: State): number[] =>
541 contexts(state)
542 .slice(-TREND_TURNS)
543 .map(tokens => Math.round((tokens / Math.max(1, state.usage.window)) * 1000) / 10)
544
545// Before the first row nothing was measured, and a total of zero would read as a session that cost nothing.
546const sinksOf = (state: State, measure: 'ms' | 'chars'): Sinks | null =>
547 state.rows.length === 0 ? null : sinks(state.rows, measure, state.loops, state.folded)
548
549const headerOf = (state: State): Header => ({
550 percent: state.usage.percent ?? null,
551 tokensToCompaction: tokensToCompaction(state),
552 turnsToCompaction: turnsToCompaction(state),
553 trend: trendOf(state),
554 time: sinksOf(state, 'ms'),
555 context: sinksOf(state, 'chars'),
556 judgeTime: state.judge.time,
557 judgeContext: state.judge.context,
558 judgeRuns: state.judge.runs,
559 judgeTokens: state.judge.spent,
560 judgeShare: judgeShare(state) ?? 0,
561 judgeRunning: state.judge.running,
562 savedPct: pctOf(state.saved.chars, state.usage.window),
563 savedMs: state.saved.ms,
564})
565
566const decidedRowOf = (state: State, p: Decided): DecidedRow => ({
567 patternId: p.id,
568 choice: p.decision,
569 kind: p.kind,
570 savedPct: isSent(p.decision) ? pctOf(baseline(state, p).chars, state.usage.window) : null,
571 // D4: the figure is a projection of one avoided repeat until the instruction settled — nothing ignored
572 // it and nothing is still in flight — so the drawing can say `per repeat` before it says `saved`.
573 settled: isSent(p.decision) && p.openedAtTurn === null && p.ignored === 0,
574 // What Fix sends is the kind and the fix the row already carries; only a note of the user's own is news.
575 instruction: p.decision === 'steer' && p.instruction !== null && collapseWs(p.instruction) !== collapseWs(p.alternative)
576 ? p.instruction
577 : null,
578 ignored: p.ignored,
579 sent: p.sent,
580})
581
582/** Builds everything the pane renders: header, wasters, decisions and rules. */
583export const paneModel = (state: State, artifacts: Artifact[]): PaneModel => {
584 // One alias table for the whole draw: the naming is the ledger's, not a card's, and the pane
585 // redraws on every ledger row and every keystroke in the Fix… field.
586 const aliases = agentAliases(state.rows, state.loops)
587 return {
588 header: headerOf(state),
589 // Numbered as they are drawn, so `/saver keep 2` names the card the person is looking at.
590 wasters: state.cards
591 .map(id => patternById(state.patterns, id))
592 .filter((p): p is Pattern => p !== undefined)
593 .map((p, at) => cardOf(p, state, at + 1, aliases)),
594 expanded: state.expanded,
595 steering: state.steering,
596 steerDraft: state.steerDraft,
597 decided: state.patterns
598 .filter(isDecided)
599 .sort((a, b) => (b.decidedAtTurn ?? 0) - (a.decidedAtTurn ?? 0))
600 .map(p => decidedRowOf(state, p)),
601 artifacts,
602 }
603}
604
605// What the cards still awaiting a decision have already cost: the figures the band's teaser states.
606const waitingCost = (state: State): { ms: number; chars: number } =>
607 state.cards
608 .map(id => patternById(state.patterns, id))
609 .filter((p): p is Pattern => p !== undefined)
610 .reduce(
611 (sum, p) => {
612 const total = totalOf(p, state, rowsOf(state, p))
613 return { ms: sum.ms + total.ms, chars: sum.chars + total.chars }
614 },
615 { ms: 0, chars: 0 },
616 )
617
618// The last turn's end when it was an error or a refusal and no prompt has followed it: the session stopped.
619const diedOf = (state: State): BandModel['died'] => {
620 const last = state.turns[state.turns.length - 1]
621 if (last === undefined || last.turn !== state.turn) return null
622 return last.ended === 'error' || last.ended === 'refusal' ? last.ended : null
623}
624
625// The newest run still going: how many loops it has spawned, how many calls they made, which stage is running.
626const runningOf = (state: State, now: number): BandModel['running'] => {
627 const active = activeRuns(state, now)
628 const run = active[active.length - 1]
629 if (run === undefined) return null
630 const loops = state.loops.filter(l => l.run === run.id)
631 const ids = new Set(loops.map(l => l.id))
632 const going = loops.filter(l => l.ended === null)
633 return { name: run.name, loops: loops.length, calls: state.rows.filter(r => ids.has(r.agent)).length, label: going[going.length - 1]?.label ?? null }
634}
635
636// The first state that applies wins: a dead turn, a run in flight, then cards waiting, then a saving to show off.
637const bandState = (state: State, died: BandModel['died']): BandModel['state'] => {
638 if (died !== null) return 'died'
639 if (state.judge.running) return 'checking'
640 if (state.cards.length > 0) return 'found'
641 if (state.saved.chars > 0 || state.saved.ms > 0) return 'saved'
642 return 'watching'
643}
644
645// The latest clock the state has seen: a turn's completion or a run's launch. Without a clock a loopless run
646// would read as fresh forever, so the band takes the newest reading it holds when the caller has none.
647const clockOf = (state: State): number =>
648 Math.max(0, ...state.turns.map(t => t.at), ...state.runs.map(r => r.at))
649
650/** Builds the one teaser line the band shows above the prompt; `now` dates a loopless run's freshness. */
651export const bandModel = (state: State, now: number = clockOf(state)): BandModel => {
652 const cost = waitingCost(state)
653 const died = diedOf(state)
654 return {
655 state: bandState(state, died),
656 died,
657 running: runningOf(state, now),
658 fresh: state.cards.length,
659 costPct: pctOf(cost.chars, state.usage.window),
660 costMs: cost.ms,
661 savedPct: pctOf(state.saved.chars, state.usage.window),
662 savedMs: state.saved.ms,
663 calls: state.rows.length,
664 paneOpen: state.paneOpen,
665 }
666}
667
668const isFilled = (v: unknown): v is string => typeof v === 'string' && v.trim().length > 0
669
670const isText = (v: unknown, max: number): v is string => isFilled(v) && v.length <= max
671
672const isOneOf = <T extends string>(v: unknown, options: readonly T[]): v is T =>
673 typeof v === 'string' && (options as readonly string[]).includes(v)
674
675const fields = (v: unknown): Record<string, unknown> | null =>
676 v !== null && typeof v === 'object' && !Array.isArray(v) ? (v as Record<string, unknown>) : null
677
678const signatureOf = (v: unknown): Signature | null | undefined => {
679 if (v === null) return null
680 const o = fields(v)
681 const tool = o?.['tool']
682 const key = o?.['key']
683 return isText(tool, KEY_MAX) && isText(key, KEY_MAX) ? { tool, key } : undefined
684}
685
686const proposalOf = (v: unknown): Proposal | null | undefined => {
687 if (v === null) return null
688 const o = fields(v)
689 const kind = o?.['kind']
690 const title = o?.['title']
691 const body = o?.['body']
692 if (!isOneOf(kind, ['claude-md', 'skill', 'agent-brief', 'settings-allow'])) return undefined
693 return isFilled(title) && isFilled(body) ? { kind, title, body } : undefined
694}
695
696const storedOf = (v: unknown): StoredPattern | null => {
697 const o = fields(v)
698 if (o === null) return null
699 const id = o['id']
700 const category = o['category']
701 const kind = o['kind']
702 const why = o['why']
703 const alternative = o['alternative']
704 const confidence = o['confidence']
705 const estTokensPerTurn = o['estTokensPerTurn']
706 const lastDecision = o['lastDecision']
707 const signature = signatureOf(o['signature'])
708 const proposal = proposalOf(o['proposal'])
709 if (typeof id !== 'string' || !/^[a-z-]+:[a-z0-9-]{1,40}$/.test(id)) return null
710 if (!isOneOf(category, ['execution', 'reading', 'production', 'behavior', 'communication', 'multi-agent', 'environment', 'process', 'other'])) return null
711 if (!isText(kind, KIND_MAX) || !isText(alternative, ALTERNATIVE_MAX) || typeof why !== 'string') return null
712 if (typeof confidence !== 'number' || !(confidence >= 0.5) || !(confidence <= 1)) return null
713 if (estTokensPerTurn !== null && !(typeof estTokensPerTurn === 'number' && Number.isFinite(estTokensPerTurn) && estTokensPerTurn >= 0)) return null
714 if (lastDecision !== null && !isOneOf(lastDecision, ['keep', 'steer', 'kill'])) return null
715 if (signature === undefined || proposal === undefined) return null
716 const lean = o['lean'] ?? null
717 if (lean !== null && !isText(lean, ALTERNATIVE_MAX)) return null
718 return { id, category, kind, signature, why, alternative, confidence, proposal, estTokensPerTurn, lastDecision, lean }
719}
720
721/** Reads a stored registry from the plugin store, dropping every entry that does not validate. */
722export const parseRegistry = (value: unknown): StoredPattern[] => {
723 if (!Array.isArray(value)) return []
724 const byId = new Map<string, StoredPattern>()
725 for (const item of value) {
726 const stored = storedOf(item)
727 if (stored !== null) byId.set(stored.id, stored)
728 }
729 return [...byId.values()]
730}
731
732/** Strips a pattern's session fields, leaving what is persisted per project. */
733export const toStored = (p: Pattern): StoredPattern => ({
734 id: p.id,
735 category: p.category,
736 kind: p.kind,
737 signature: p.signature,
738 why: p.why,
739 alternative: p.alternative,
740 confidence: p.confidence,
741 proposal: p.proposal,
742 estTokensPerTurn: p.estTokensPerTurn,
743 lastDecision: p.lastDecision,
744 lean: p.lean,
745})
746
747const QUOTE_MIN = 24 // consecutive characters of a typed ask that make a stored text a quote of it
748
749// Whether `text` carries QUOTE_MIN consecutive characters of any head, case and whitespace aside.
750const quotesAsk = (text: string, heads: readonly string[]): boolean => {
751 const t = collapseWs(text).toLowerCase()
752 return heads.some(h => Array.from({ length: Math.max(0, h.length - QUOTE_MIN + 1) }, (_, i) => h.slice(i, i + QUOTE_MIN)).some(q => t.includes(q)))
753}
754
755/**
756 * What this session persists: every pattern stored, never a typed prompt. A process `why` is not stored at
757 * all, a habit `why` quoting an ask is blanked, and a pattern whose kind, fix or lean quotes one stays in the session.
758 */
759export const storableOf = (state: State): StoredPattern[] => {
760 const heads = state.asks.map(a => collapseWs(a.head).toLowerCase())
761 return state.patterns.map(toStored).flatMap(s =>
762 [s.kind, s.alternative, s.lean ?? ''].some(t => quotesAsk(t, heads))
763 ? []
764 : [{ ...s, why: s.lean !== null || quotesAsk(s.why, heads) ? '' : s.why }])
765}
766
767/** Revives a stored pattern with empty session fields. */
768export const fromStored = (s: StoredPattern): Pattern => ({
769 ...s, hits: [], decision: null, decidedAtTurn: null, instruction: null, openedAtTurn: null, ignored: 0, sent: 0,
770})
771
772// What a stored entry is worth when the registry overflows: a decision outranks any confidence.
773const rankOf = (p: StoredPattern): number => (p.lastDecision === null ? p.confidence : 2 + p.confidence)
774
775const capStored = (entries: readonly StoredPattern[]): StoredPattern[] => {
776 if (entries.length <= MAX_PATTERNS) return [...entries]
777 return [...entries.entries()]
778 .sort(([atA, a], [atB, b]) => rankOf(b) - rankOf(a) || atA - atB)
779 .slice(0, MAX_PATTERNS)
780 .sort(([atA], [atB]) => atA - atB)
781 .map(([, p]) => p)
782}
783
784/** Merges two stored registries by id (`b` wins), dropping the weakest undecided entries past the cap. */
785export const mergeStored = (a: readonly StoredPattern[], b: readonly StoredPattern[]): StoredPattern[] => {
786 const byId = new Map<string, StoredPattern>(a.map(p => [p.id, p]))
787 for (const p of b) byId.set(p.id, p)
788 return capStored([...byId.values()])
789}
790
791const classCounts = (rows: readonly Row[]): string => {
792 const counts = new Map<string, number>()
793 for (const r of rows) counts.set(r.cls, (counts.get(r.cls) ?? 0) + 1)
794 return [...counts.entries()].sort((a, b) => b[1] - a[1]).map(([cls, n]) => `${cls}×${n}`).join(' ') || '(none)'
795}
796
797const oneLine = (text: string | null): string => (text === null ? '-' : `"${text.replace(/\n/g, '\\n').slice(0, 120)}"`)
798
799const patternLine = (p: Pattern): string =>
800 ` ${p.id} · hits ${p.hits.length} [${p.hits.slice(0, 5).join(' ')}] · ${p.decision ?? '-'} @ ${p.decidedAtTurn ?? '-'} · previous ${p.lastDecision ?? '-'} · ignored ${p.ignored} · opened ${p.openedAtTurn ?? '-'} · sent ${oneLine(p.instruction)}`
801
802const patternLines = (state: State): string[] => {
803 const shown = state.patterns.slice(0, DEBUG_MAX_PATTERNS).map(patternLine)
804 const rest = state.patterns.length - shown.length
805 return rest > 0 ? [...shown, ` … ${rest} more patterns`] : shown
806}
807
808/** What one judge fork cost, one line: the same text `/saver debug` and the debug log both print. */
809export const usageLine = (u: JudgeUsage, label = 'judge'): string =>
810 `${label} usage: in ${u.input} · out ${u.output} · cache read ${u.cacheRead} · cache create ${u.cacheCreate}`
811
812// What the last run reported, and why anything it returned never reached the user.
813const judgeRunLines = (run: JudgeRun | null, label = 'judge'): string[] =>
814 run === null
815 ? []
816 : [
817 `${label} last: ${run.returned} returned · ${run.kept} kept · ${run.dropped.length} dropped`,
818 ...run.dropped.slice(0, DEBUG_MAX_DROPPED).map(reason => ` ${reason}`),
819 ...(run.usage === null ? [] : [usageLine(run.usage, label)]),
820 ]
821
822// The audit a load owes a session it joined late: fired at `session.start` over the adopted rows, retried at
823// every warm opportunity while it keeps coming back with nothing, and never armed at all under the row floor.
824const loadCheckLine = (state: State, spoke: boolean): string => {
825 const lane = state.pendingCheck ? 'retrying' : state.judge.runs === 0 ? 'not armed' : 'answered'
826 return `load check: ${lane} · ${spoke ? 'reported' : 'not yet reported'}`
827}
828
829// Where a budget went, as the pane's Time and Context rows no longer spell out: the total and the largest sinks.
830const sinkLine = (label: string, unit: string, budget: Sinks | null): string =>
831 `${label} sinks: ${budget === null ? '-' : [`${budget.total}${unit} total`, ...budget.sinks.map(s => `${s.label} ${s.amount} ×${s.count}`)].join(' · ')}`
832
833// The process judge's cadence and the asks it counts: how many and how many read as pace, never what was typed.
834const processDebugLine = (state: State): string => {
835 const p = state.process
836 return `process runs ${p.runs} · spent ${p.spent} tokens · backoff ${p.backoff} · running ${p.running} · lastAt turn ${p.lastAtTurn} / row ${p.lastAtSeq} / ${p.lastAtMs}ms · error ${p.error ?? '-'} · asks ${state.asks.length} · pace ${state.asks.filter(a => a.pace).length}`
837}
838
839/** Renders the whole state, and whether the load lane has reported yet, for `/saver debug` in ≤ 40 lines. */
840export const debugDump = (state: State, spoke = false): string => {
841 const j = state.judge
842 const u = state.usage
843 const o = state.overhead
844 const share = judgeShare(state)
845 return [
846 `ContextSaver · turn ${state.turn} · seq ${state.seq} · rows ${state.rows.length} · turns ${state.turns.length} · patterns ${state.patterns.length}`,
847 `rows ${classCounts(state.rows)}`,
848 `runs ${state.runs.length} · loops ${state.loops.length} · active ${activeRuns(state, clockOf(state)).length}`,
849 sinkLine('time', 'ms', sinksOf(state, 'ms')),
850 sinkLine('context', 'ch', sinksOf(state, 'chars')),
851 ...patternLines(state),
852 `cards ${state.cards.length}${state.cards.length === 0 ? '' : `: ${state.cards.join(', ')}`}`,
853 `notes ${state.notes.length} · standing ${state.standing.length} · written ${state.written.length}${state.written.length === 0 ? '' : `: ${state.written.join(', ')}`}`,
854 `judge runs ${j.runs} · spent ${j.spent} tokens (${share === null ? '-' : `${share}%`} of the session) · backoff ${j.backoff} · running ${j.running} · lastAt ${j.lastAtTokens} tokens / turn ${j.lastAtTurn} / row ${j.lastAtSeq} / ${j.lastAtMs}ms · error ${j.error ?? '-'} · focus ${oneLine(j.focus)}`,
855 `judge time: ${oneLine(j.time)}`,
856 `judge context: ${oneLine(j.context)}`,
857 processDebugLine(state),
858 ...judgeRunLines(j.last),
859 ...judgeRunLines(state.process.last, 'process'),
860 loadCheckLine(state, spoke),
861 `usage ${u.percent ?? '-'}% · ${u.tokens ?? '-'} / ${u.window} tokens · compactAt ${u.compactAt ?? '-'} · toCompaction ${tokensToCompaction(state) ?? '-'} · turnsLeft ${turnsToCompaction(state) ?? '-'} · session ${totalTokens(state)} new`,
862 `overhead ${o === null ? '-' : `memory ${o.memory} · mcp ${o.mcp} · agents ${o.agents}`}`,
863 `compactions ${state.compactions.length === 0 ? 'none' : state.compactions.join(', ')}`,
864 `pane ${state.paneOpen ? 'open' : 'closed'} · autoOpened ${state.autoOpened} · columns ${state.columns ?? '-'} · expanded ${state.expanded ?? '-'} · steering ${state.steering ?? '-'}`,
865 `saved ${duration(state.saved.ms)} · ~${pctOf(state.saved.chars, u.window)}% · ${state.saved.chars} chars`,
866 ].slice(0, DEBUG_MAX_LINES).join('\n')
867}
868
869/** What `/saver` says of the process judge: its runs and their spend; null before its first run. */
870export const processLine = (state: State): string | null =>
871 state.process.runs === 0 ? null : `process ${state.process.runs} run${state.process.runs === 1 ? '' : 's'} · ${kilo(state.process.spent)} tokens`
872hooks/core/process.ts 413 lines1import { agentAliases } from './evidence'
2import { merge } from './judge'
3import { isCheck, isEdit } from './spawns'
4import { cell, isRecord, loopOf, loopTokens, parseObject, proposalOf, slug, str, unique } from './text'
5import {
6 ALTERNATIVE_MAX, DIGEST_MAX_CHARS, KIND_MAX, MAIN_AGENT, MAX_PROCESS_FINDINGS, PROCESS_CLOCK_MS, PROCESS_MIN_CONFIDENCE,
7} from './types'
8import type { Finding, Loop, Outcome, Pattern, ProcessTrigger, Row, Run, State, TurnStat } from './types'
9
10/** A process finding with the cost it claims, clamped to what its cited handles measured. */
11export type ProcessFinding = Finding & { cost: { tokens: number; ms: number } }
12
13/** The process judge prompt (build spec Appendix E, verbatim) with its one placeholder. */
14export const PROCESS_PROMPT = `This message is not a step of the task. The work is paused while you answer one review question: you cannot call tools, and do not continue, retry or resume anything the transcript was doing. Your whole reply is the JSON object described below.
15
16You are reviewing HOW this session is being run, not what it produced. The transcript above is your own: read it for what the user asked for and what was decided. The digest below is the measured shape of the work so far; nothing outside it is evidence for this review.
17
18Answer one question: given what the user asked, how would a lean expert run this work? Then say where this session's process diverges from that, what each divergence has cost so far, and what the one change to the remaining work is. Look at the shape, not at single calls: review and fix rounds stacked on a review that already covers them; the heaviest model on mechanical stages; an agent spawned for a job one command would do; the full check suite after every merge where once per batch would do; the same plan or file re-read over and over; the pace the user complained about.
19
20Prefer silence to a guess. A wrong "lean" suggestion is worse than none: report a divergence only when the digest measures its cost and the lean alternative plainly reaches the same result for less. Anything the user asked for is not a divergence, and a long session is not a wasteful one. At most three findings; fewer and bigger is better, and \`"findings": []\` is a correct and common answer. INTERVENTIONS lists every process finding already raised. The same divergence under another id is the same finding: return its id. One the person ignored is never reported again.
21
22## Evidence
23\`evidence\`: handles the digest prints, copied exactly — \`run:<id>\` (a SHAPE run row), \`agent:<alias>\` (a loop SHAPE lists, as \`agent:a3\`), \`turn:<n>\` (a turn ASKS or CADENCE names). A finding citing any other handle is discarded whole. Cite every run, loop or turn the cost comes from.
24
25## Cost
26\`cost_tokens\` and \`cost_ms\`: what the divergence has cost so far, measured from the cited handles — a run's or a loop's tokens and minutes, a turn's. A claim above what the cited handles total is cut down to it. Never a projection: what it would cost if unchanged belongs in \`why\`.
27
28## Confidence
29\`confidence\` runs 0.75 to 1.0. 0.9+: the digest shows the divergence repeated and costed, and nothing in the transcript asked for it. 0.75-0.9: plain and costed, and no legitimate reason is visible. Below 0.75: say nothing; it is discarded.
30
31## Contract — the shape of your reply, stated once (documentation, not a template to echo)
32\`\`\`json
33{"findings": [{"id": "process:<kebab-slug, at most 40 chars>",
34 "kind": "<one sentence, at most 120 chars: what this session does>",
35 "lean": "<one sentence, at most 200 chars: how a lean expert runs it>",
36 "alternative": "<one imperative sentence, at most 200 chars, addressed to Claude: the one change to the remaining work>",
37 "why": "<one or two sentences: the measured cost, and the reason for it you considered and ruled out>",
38 "evidence": ["run:<id>", "agent:<alias>", "turn:<n>"],
39 "cost_tokens": 0,
40 "cost_ms": 0,
41 "confidence": 0.85,
42 "proposal": null}]}
43\`\`\`
44\`proposal\`: null unless the change should outlive the session. Otherwise \`{"kind","title","body"}\` where body is, per kind: \`claude-md\` one imperative rule line; \`skill\` the workflow as the body of a SKILL.md; \`agent-brief\` the brief, whose first line may be \`model: haiku\` or \`model: sonnet\`.
45
46Reply with one JSON object: first character \`{\`, last character \`}\`, no prose before or after, no code fence.
47
48## PROCESS DIGEST
49{{DIGEST}}
50
51Return the JSON object only.
52`
53
54/** Names this session's loops the way one process run sees them: built once before the prompt, read again when the reply comes back. */
55export const processAliases = (state: State): ReadonlyMap<string, string> => agentAliases(state.rows, state.loops)
56
57// ---- the digest ----------------------------------------------------------------------------------------
58
59type Section = { name: 'asks' | 'shape' | 'cadence' | 'artifacts' | 'interventions'; header: string; lines: string[]; oldestFirst: boolean; pinned: number } // pinned: leading lines a cut never drops
60
61// The sections cut first when the digest runs past its cap: the old asks, then the small sections; Shape is the last resort.
62const CUT_ORDER: readonly Section['name'][] = ['asks', 'artifacts', 'interventions', 'cadence', 'shape']
63const CUT_MARK_MAX = 24 // room for the `~ ×N lines cut` line a cut section ends with
64const LIST_MAX = 10 // paths, re-reads and gaps each section names before it stops
65const REREAD_MIN = 3 // reads of one path before it is named a re-read
66const GAP_MIN_MS = 60_000 // a wait shorter than a minute before the next prompt is the person reading
67const READ_RANGE = /:\d*-\d*$/
68const MERGE = /\bmerge\b/
69const COMMIT = /\bcommit\b/
70const MEMORY_FILE = /(\/memory\/[^/]+\.md|\/(CLAUDE|MEMORY)\.md)$/
71
72const minutes = (ms: number): string => `${(ms / 60_000).toFixed(1)}m`
73
74const ktok = (tokens: number): string => `${Math.round(tokens / 1000)}k`
75
76// The family word of a model id (`claude-opus-4-1` → `opus`), or the id itself when it names no family.
77const modelShort = (model: string | null): string => model === null ? '?' : (/(opus|sonnet|haiku|fable)/.exec(model)?.[1] ?? model)
78
79// The role a loop plays: its label's first word (`review:C3-r1` → `review`), else its phase, else `agent`.
80const roleOf = (l: Loop): string =>
81 /^[A-Za-z][\w-]*/.exec(l.label ?? '')?.[0]?.toLowerCase() ?? l.phase?.toLowerCase() ?? 'agent'
82
83// `name ×n` counts, largest first then by name.
84const tally = (names: readonly string[]): string => {
85 const counts = names.reduce((m, n) => m.set(n, (m.get(n) ?? 0) + 1), new Map<string, number>())
86 return [...counts].sort((a, b) => b[1] - a[1] || (a[0] < b[0] ? -1 : a[0] > b[0] ? 1 : 0)).map(([n, c]) => `${n} ×${c}`).join(' · ')
87}
88
89const SEVERITIES = ['critical', 'high', 'medium', 'low'] as const
90
91// What a group of loops returned: severities summed across their reviews, reports counted.
92const outcomesCell = (outcomes: readonly (Outcome | null)[]): string => {
93 const found = outcomes.filter((o): o is Extract<Outcome, { kind: 'findings' }> => o?.kind === 'findings')
94 const severities = SEVERITIES.map(s => [s, found.reduce((n, o) => n + o[s], 0)] as const).filter(([, n]) => n > 0).map(([s, n]) => `${n} ${s}`)
95 const reports = outcomes.filter(o => o?.kind === 'report').length
96 const parts = [...(severities.length > 0 ? [severities.join(' ')] : found.length > 0 ? ['0 findings'] : []), ...(reports > 0 ? [`report ×${reports}`] : [])]
97 return parts.length > 0 ? parts.join(' · ') : '-'
98}
99
100const sumOf = (xs: readonly number[]): number => xs.reduce((a, b) => a + b, 0)
101
102// The run a row belongs to: an agent's row by its loop's run; a main-loop row to the newest run launched before it.
103const runOfRow = (state: State, row: Row): string | null => {
104 if (row.agent !== MAIN_AGENT) return state.loops.find(l => l.id === row.agent)?.run ?? null
105 return [...state.runs].reverse().find(r => r.seq <= row.seq)?.id ?? null
106}
107
108const isMerge = (r: Row): boolean => r.cls === 'git' && MERGE.test(r.key)
109
110// Merges, and the merges a check run followed before the next merge.
111const mergeChecks = (rows: readonly Row[]): { merges: number; checked: number } =>
112 [...rows].sort((a, b) => a.seq - b.seq).reduce((acc, r) => {
113 if (isMerge(r)) return { merges: acc.merges + 1, checked: acc.checked, open: true }
114 if (acc.open && isCheck(r)) return { ...acc, checked: acc.checked + 1, open: false }
115 return acc
116 }, { merges: 0, checked: 0, open: false })
117
118const handleOf = (aliases: ReadonlyMap<string, string>, l: Loop): string => `agent:${aliases.get(l.id) ?? l.id}`
119
120// One declared (or seen) phase of a run: its loops, their models, time, tokens, edits and what they returned.
121const phaseLine = (aliases: ReadonlyMap<string, string>, phase: string, loops: readonly Loop[]): string =>
122 [
123 ` ${phase}`, `×${loops.length}`, tally(loops.map(l => modelShort(l.model))) || '-', minutes(sumOf(loops.map(l => l.ms))),
124 `${ktok(sumOf(loops.map(loopTokens)))} tok`, `edits ${sumOf(loops.map(l => l.edits))}`, outcomesCell(loops.map(l => l.outcome)),
125 loops.map(l => handleOf(aliases, l)).join(' ') || '-',
126 ].join(' | ')
127
128const runLines = (state: State, aliases: ReadonlyMap<string, string>, run: Run): string[] => {
129 const loops = state.loops.filter(l => l.run === run.id)
130 const { merges, checked } = mergeChecks(state.rows.filter(r => runOfRow(state, r) === run.id))
131 const phases = unique([...run.phases, ...loops.map(l => l.phase ?? '-')])
132 return [
133 [
134 `run:${run.id}`, run.name, `phases ${run.phases.length > 0 ? run.phases.join(' → ') : '(none declared)'}`, `loops ${loops.length}`,
135 `Σ${minutes(sumOf(loops.map(l => l.ms)))}`, `Σ${ktok(sumOf(loops.map(loopTokens)))} tok`, `edits ${sumOf(loops.map(l => l.edits))}`,
136 `merges ${merges} · checks after ${checked}`, `turn ${run.turn}`,
137 ].join(' | '),
138 ...phases.map(p => phaseLine(aliases, p, loops.filter(l => (l.phase ?? '-') === p))),
139 ]
140}
141
142const directLine = (aliases: ReadonlyMap<string, string>, l: Loop): string =>
143 [` ${handleOf(aliases, l)}`, l.label ?? '-', modelShort(l.model), minutes(l.ms), `${ktok(loopTokens(l))} tok`, `calls ${l.calls}`, `edits ${l.edits}`, l.ended ?? 'running'].join(' | ')
144
145const shapeLines = (state: State, aliases: ReadonlyMap<string, string>): string[] => {
146 const direct = state.loops.filter(l => l.run === null || !state.runs.some(r => r.id === l.run))
147 return [
148 ...state.runs.flatMap(run => runLines(state, aliases, run)),
149 ...(direct.length > 0 ? [`direct agents: ×${direct.length}`, ...direct.map(l => directLine(aliases, l))] : []),
150 `roles × model: ${tally(state.loops.map(l => `${roleOf(l)} ${modelShort(l.model)}`)) || '(no agents)'}`,
151 ]
152}
153
154const firstAt = (state: State): number | null => state.asks[0]?.at ?? state.runs[0]?.at ?? null
155
156const askLines = (state: State): string[] => {
157 const first = state.asks[0]?.at ?? 0
158 return state.asks.map(a => `turn:${a.turn} | +${Math.round((a.at - first) / 60_000)}m | ${a.head}${a.pace ? ' | pace' : ''}`)
159}
160
161const gapLine = (turns: readonly TurnStat[]): string => {
162 const gaps = [...turns].filter(t => t.idleMs >= GAP_MIN_MS).sort((a, b) => b.idleMs - a.idleMs).slice(0, LIST_MAX)
163 return `gaps: ${gaps.length > 0 ? gaps.map(t => `turn:${t.turn} idle ${Math.round(t.idleMs / 60_000)}m`).join(', ') : 'none'}`
164}
165
166const cadenceLines = (state: State): string[] => {
167 const pace = state.asks.filter(a => a.pace)
168 const { merges, checked } = mergeChecks(state.rows)
169 const last = state.turns[state.turns.length - 1]?.at ?? 0
170 const since = firstAt(state)
171 return [
172 [
173 `turns ${state.turns.length}`,
174 ...(since !== null && last > since ? [`span ${minutes(last - since)}`] : []),
175 `compactions ${state.compactions.length}${state.compactions.length > 0 ? ` (turn ${state.compactions.join(', ')})` : ''}`,
176 `pace complaints ${pace.length}${pace.length > 0 ? ` (${pace.map(a => `turn:${a.turn}`).join(', ')})` : ''}`,
177 ].join(' · '),
178 `checks ${state.rows.filter(isCheck).length} · merges ${merges} · checks after a merge ${checked} · commits ${state.rows.filter(r => r.cls === 'git' && COMMIT.test(r.key)).length}`,
179 gapLine(state.turns),
180 ]
181}
182
183// The path a file row names: a Read key without its slice, else the paths it changed.
184const pathsOf = (r: Row): string[] => (r.tool === 'Read' ? [r.key.replace(READ_RANGE, '')] : r.paths)
185
186const topCounts = (entries: readonly string[]): [string, number][] =>
187 [...entries.reduce((m, p) => m.set(p, (m.get(p) ?? 0) + 1), new Map<string, number>())].sort((a, b) => b[1] - a[1]).slice(0, LIST_MAX)
188
189const artifactLines = (state: State): string[] => {
190 const edits = state.rows.filter(isEdit)
191 const written = edits.flatMap(r => r.paths)
192 const reads = state.rows.filter(r => r.tool === 'Read')
193 const reread = topCounts(reads.map(r => r.key.replace(READ_RANGE, ''))).filter(([, n]) => n >= REREAD_MIN)
194 const loopsOn = (path: string): number => unique(reads.filter(r => r.key.replace(READ_RANGE, '') === path).map(r => r.agent)).length
195 // Memory files by name and path only: what they say is never read.
196 const memory = unique(state.rows.flatMap(pathsOf).filter(p => MEMORY_FILE.test(p))).map(p =>
197 `memory: ${p.slice(p.lastIndexOf('/') + 1)} ${p} ${written.includes(p) ? 'written' : 'read'}`)
198 return [
199 `edits ${edits.length} on ${unique(written).length} paths · agents' own edits ${sumOf(state.loops.map(l => l.edits))}`,
200 ...(written.length > 0 ? [`written: ${topCounts(written).map(([p, n]) => `${p} ×${n}`).join(' · ')}`] : []),
201 ...(reread.length > 0 ? [`re-read: ${reread.map(([p, n]) => `${p} ×${n} by ${loopsOn(p)} loops`).join(' · ')}`] : []),
202 ...memory,
203 ]
204}
205
206// Every process finding already raised, open ones too, in its own words: without them the next run mints the
207// same divergence under a fresh id, and a second card undoes the person's Ignore.
208const interventionLines = (state: State): string[] => {
209 const listed = state.patterns.filter(p => p.decision !== null || p.category === 'process')
210 return [
211 ...listed.map(p => [
212 p.id,
213 p.decision === null ? 'open' : `${p.decision} @ ${p.decidedAtTurn ?? '-'}`,
214 `sent ×${p.sent}${p.ignored > 0 ? ` | ignored ×${p.ignored}` : ''}`,
215 ...(p.category === 'process' ? [p.kind] : []),
216 ].join(' | ')),
217 `cards waiting ${state.cards.length} · standing instructions ${state.standing.length} · one-time notes pending ${state.notes.length} · subagents reached ${state.delivery.agents.length}`,
218 ]
219}
220
221const render = (sections: readonly Section[]): string =>
222 sections.map(s => [s.header, ...(s.lines.length > 0 ? s.lines : ['(none)'])].join('\n')).join('\n\n')
223
224// A section with enough lines dropped to give back `over` characters: the oldest first (after the pinned) where order is time, else the last.
225const cut = (s: Section, over: number): Section => {
226 const pinned = s.lines.slice(0, s.pinned)
227 const order = s.oldestFirst ? s.lines.slice(s.pinned) : s.lines.slice(s.pinned).reverse()
228 let removed = 0
229 let n = 0
230 while (n < order.length && removed < over + CUT_MARK_MAX) { removed += (order[n]?.length ?? 0) + 1; n += 1 }
231 const rest = order.slice(n)
232 const mark = `~ ×${n} lines cut`
233 return { ...s, lines: s.oldestFirst ? [...pinned, mark, ...rest] : [...pinned, ...rest.reverse(), mark] }
234}
235
236/**
237 * The PROCESS digest: what the person typed, the shape of the spawned work, the cadence, the artifacts and the
238 * interventions, at most DIGEST_MAX_CHARS; over the cap the oldest asks (never the first) go first, then the small sections, and Shape last.
239 *
240 * @param state the session so far
241 * @param aliases the loop naming the prompt is built with, so a cited `agent:<alias>` reads back to its loop
242 */
243export const digestOf = (state: State, aliases: ReadonlyMap<string, string> = processAliases(state)): string => {
244 const sections: Section[] = [
245 { name: 'asks', header: '## ASKS — what the person typed, oldest first: `turn:<n> | +<min>m since the first | what they typed`, then `| pace` when it complains about pace', lines: askLines(state), oldestFirst: true, pinned: 1 },
246 { name: 'shape', header: '## SHAPE — one row per workflow run: `run:<id> | name | declared phases | loops | Σmin | Σtok | edits | merges · checks after | launch turn`, then per phase `phase | ×loops | models | min | tok | edits | outcomes | agent handles`; then agents outside any run, and every role by model', lines: shapeLines(state, aliases), oldestFirst: true, pinned: 0 },
247 { name: 'cadence', header: '## CADENCE — turns, span, compactions, pace complaints; check runs, merges, merges a check followed, commits; the longest waits for the person', lines: cadenceLines(state), oldestFirst: false, pinned: 0 },
248 { name: 'artifacts', header: '## ARTIFACTS — edits, the paths written most, paths re-read, memory files by name and path', lines: artifactLines(state), oldestFirst: false, pinned: 0 },
249 { name: 'interventions', header: '## INTERVENTIONS — process findings already raised and cards decided `id | open or decision @ turn | sent ×N`, a process finding\'s words after it; then what is waiting and what was sent', lines: interventionLines(state), oldestFirst: false, pinned: 0 },
250 ]
251 const fitted = CUT_ORDER.reduce((all, name) => {
252 const over = render(all).length - DIGEST_MAX_CHARS
253 return over <= 0 ? all : all.map(s => (s.name === name ? cut(s, over) : s))
254 }, sections)
255 return render(fitted).slice(0, DIGEST_MAX_CHARS)
256}
257
258// ---- triggers ------------------------------------------------------------------------------------------
259
260// Complaints about pace, as the person types them; lowercased first.
261const PACE: readonly RegExp[] = [
262 /\b(taking|takes|took|take) (so |way |far )?(forever|ages)\b/,
263 /\b(taking|takes|took) (so |way |far )?too long\b/,
264 /\b(this|it|that)('s| is| was)? (taking )?(way |far )?too long\b/,
265 /\bwhy (is (this|it) )?(so |this )?slow\b/,
266 /\b(so|too) slow\b/,
267 /\bhurry( up)?\b/,
268 /\bspeed (this|it|things) up\b/,
269 /\bbeen ages\b/,
270 /\bwhat('s| is) taking so long\b/,
271 /\bstill (not done|waiting)\b/,
272 /\b(take|takes|taking|took) (you |it |this )?so long\b/,
273 /\b(should|ought to) (have been|be) (much |way |a lot )?faster\b/,
274 /\b(expensive|pricey|costly) and slow\b/,
275 /\bslow and (expensive|pricey|costly)\b/,
276]
277
278/** True when a typed prompt reads as a complaint about pace: 'this is taking forever', 'why so slow'. */
279export const isPace = (text: string): boolean => {
280 const lower = text.toLowerCase()
281 return PACE.some(re => re.test(lower))
282}
283
284/**
285 * True when a process run may start: none is running, something happened since the last one, and the clock
286 * has waited PROCESS_CLOCK_MS × backoff — the clock always, an event trigger only once over the budget share.
287 *
288 * @param state the session so far
289 * @param now the clock
290 * @param trigger what asked for the run
291 */
292export const shouldProcess = (state: State, now: number, trigger: ProcessTrigger): boolean => {
293 const p = state.process
294 if (p.running || state.seq <= p.lastAtSeq) return false
295 const since = p.lastAtMs > 0 ? p.lastAtMs : (firstAt(state) ?? now)
296 const waited = now - since >= PROCESS_CLOCK_MS * p.backoff
297 return trigger === 'clock' || p.backoff > 1 ? waited : true
298}
299
300/** Fills the process prompt with this session's digest; `aliases` is the naming the digest prints, for `parseProcessReply`. */
301export const buildProcessPrompt = (state: State, aliases: ReadonlyMap<string, string> = processAliases(state)): string =>
302 PROCESS_PROMPT.split('{{DIGEST}}').join(digestOf(state, aliases))
303
304// ---- the reply -----------------------------------------------------------------------------------------
305
306const RUN_HANDLE = /^run:(.+)$/
307const TURN_HANDLE = /^turn:(\d+)$/
308const AGENT = /^agent:(.+)$/
309
310/** What the cited handles measured: the tokens and wall time of every loop they name (a run's loops, each once) and of every turn. */
311export const measuredCost = (state: State, handles: readonly string[]): { tokens: number; ms: number } => {
312 const runs = handles.flatMap(h => RUN_HANDLE.exec(h)?.[1] ?? [])
313 const agents = handles.flatMap(h => AGENT.exec(h)?.[1] ?? [])
314 const turns = handles.flatMap(h => { const n = TURN_HANDLE.exec(h)?.[1]; return n === undefined ? [] : [Number(n)] })
315 const loops = state.loops.filter(l => agents.includes(l.id) || (l.run !== null && runs.includes(l.run)))
316 const cited = state.turns.filter(t => turns.includes(t.turn))
317 return {
318 tokens: sumOf(loops.map(loopTokens)) + sumOf(cited.map(t => t.input + t.cacheCreate + t.output)),
319 ms: sumOf(loops.map(l => l.ms)) + sumOf(cited.map(t => t.ms)),
320 }
321}
322
323// A handle the digest printed, stored the way the registry keeps it (`agent:<agentId>`), or null.
324const processHandle = (value: unknown, state: State, aliases: ReadonlyMap<string, string>): string | null => {
325 const handle = str(value)
326 const turn = TURN_HANDLE.exec(handle)
327 // The digest prints completed turns (CADENCE) and every ask's turn (ASKS), the one still running included.
328 if (turn !== null) {
329 const n = Number(turn[1])
330 return state.turns.some(t => t.turn === n) || state.asks.some(a => a.turn === n) ? `turn:${n}` : null
331 }
332 const run = RUN_HANDLE.exec(handle)
333 if (run !== null) return state.runs.some(r => r.id === run[1]) ? handle : null
334 const agent = AGENT.exec(handle)
335 const loop = agent === null ? undefined : loopOf(state, aliases, agent[1] ?? '')
336 return loop !== undefined ? `agent:${loop.id}` : null
337}
338
339const evidenceOf = (value: unknown, state: State, aliases: ReadonlyMap<string, string>): string[] | string => {
340 if (!Array.isArray(value) || value.length === 0) return 'evidence must be a non-empty array of handles'
341 const handles = value.map(h => processHandle(h, state, aliases))
342 const badAt = handles.indexOf(null)
343 if (badAt >= 0) return `evidence ${cell(value[badAt], 40) || '(not a string)'} is not a run, agent or turn the digest printed`
344 return unique(handles.filter((h): h is string => h !== null))
345}
346
347const idOf = (value: unknown): string => {
348 const given = str(value)
349 const s = slug(given.slice(given.indexOf(':') + 1))
350 return s.length > 0 ? `process:${s}` : ''
351}
352
353const claimed = (value: unknown): number => (typeof value === 'number' && Number.isFinite(value) ? Math.max(0, Math.round(value)) : 0)
354
355const textDrop = (name: string, text: string, max: number): string | null =>
356 text.length === 0 ? `${name} is empty` : text.length > max * 2 ? `${name} is ${text.length} chars, over twice ${max}` : null
357
358// A string is the one short reason the finding was dropped; the object is the finding itself.
359const findingOf = (value: unknown, state: State, aliases: ReadonlyMap<string, string>): ProcessFinding | string => {
360 if (!isRecord(value)) return 'not an object'
361 const id = idOf(value['id'])
362 const kind = str(value['kind']).trim()
363 const lean = str(value['lean']).trim()
364 const alternative = str(value['alternative']).trim()
365 const confidence = value['confidence']
366 if (id.length === 0) return 'id is missing'
367 const bad = textDrop('kind', kind, KIND_MAX) ?? textDrop('lean', lean, ALTERNATIVE_MAX) ?? textDrop('alternative', alternative, ALTERNATIVE_MAX)
368 if (bad !== null) return bad
369 if (typeof confidence !== 'number' || confidence > 1) return 'confidence is not a number up to 1'
370 if (confidence < PROCESS_MIN_CONFIDENCE) return `confidence ${confidence} is below ${PROCESS_MIN_CONFIDENCE}`
371 if (state.patterns.some(p => p.id === id && p.decision === 'keep')) return 'kept this session'
372 const evidence = evidenceOf(value['evidence'], state, aliases)
373 if (typeof evidence === 'string') return evidence
374 const measured = measuredCost(state, evidence)
375 return {
376 id, category: 'process', kind, evidence, signature: null, why: str(value['why']).trim(), alternative, confidence,
377 estTokensPerTurn: null, proposal: proposalOf(value['proposal']), lean,
378 cost: { tokens: Math.min(claimed(value['cost_tokens']), measured.tokens), ms: Math.min(claimed(value['cost_ms']), measured.ms) },
379 }
380}
381
382type Sifted = { findings: ProcessFinding[]; dropped: string[] }
383
384/** What the process judge said: at most MAX_PROCESS_FINDINGS valid findings, their cost clamped to the cited handles, and one line per drop; never throws. `aliases` must be the table `buildProcessPrompt` printed. */
385export const parseProcessReply = (text: string, state: State, aliases: ReadonlyMap<string, string> = processAliases(state)): { findings: ProcessFinding[]; dropped: string[]; returned: number } => {
386 const root = parseObject(text)
387 if (root === null) return { findings: [], dropped: ['reply was not JSON'], returned: 0 }
388 const raw = root['findings']
389 if (!Array.isArray(raw)) return { findings: [], dropped: ['findings was not an array'], returned: 0 }
390 const sifted = raw.reduce<Sifted>((kept, value, i) => {
391 const label = idOf(isRecord(value) ? value['id'] : '') || `#${i + 1}`
392 const dropped = (reason: string): Sifted => ({ findings: kept.findings, dropped: [...kept.dropped, `${label}: ${reason}`] })
393 const f = findingOf(value, state, aliases)
394 if (typeof f === 'string') return dropped(f)
395 if (kept.findings.some(k => k.id === f.id)) return dropped('the reply already reported this id')
396 if (kept.findings.length >= MAX_PROCESS_FINDINGS) return dropped(`over MAX_PROCESS_FINDINGS (${MAX_PROCESS_FINDINGS})`)
397 return { findings: [...kept.findings, f], dropped: kept.dropped }
398 }, { findings: [], dropped: [] })
399 return { ...sifted, returned: raw.length }
400}
401
402/** Folds process findings into the registry as the habit judge's `merge` does, the newest run's `kind`, `lean` and clamped `cost` kept; names the fresh, the recurred and the ids the cap evicted, for `process.done`. */
403export const mergeProcess = (state: State, findings: readonly ProcessFinding[]): { patterns: Pattern[]; fresh: string[]; recurred: string[]; evicted: string[] } => {
404 const merged = merge(state, [...findings])
405 return {
406 ...merged,
407 patterns: merged.patterns.map(p => {
408 const f = findings.find(x => x.id === p.id)
409 return f !== undefined ? { ...p, kind: f.kind, lean: f.lean, cost: f.cost } : p
410 }),
411 }
412}
413hooks/core/rules.ts 119 lines1import { baseline } from './evidence'
2import { collapseWs, pctOf, slug } from './text'
3import { BRIEF_TOOLS, CLAUDE_MD_HEADING } from './types'
4import type { Artifact, ArtifactKind, Pattern, Proposal, State } from './types'
5
6/** Renders one single-line CLAUDE.md bullet; the shell appends only this when the heading already exists. */
7export const bulletOf = (body: string): string => `- ${collapseWs(body).replace(/^[-*] /, '')}\n`
8
9/** Returns the bullet alone from a claude-md artifact's content, for appending under an existing heading. */
10export const bulletOnly = (content: string): string => content.slice(Math.max(0, content.indexOf('\n- ') + 1))
11
12/** Appends an artifact's content to a file's text: the bullet alone under an existing heading, never glued to a line. */
13export const appendedTo = (existing: string | null, content: string): string => {
14 const base = existing ?? ''
15 const glue = base === '' || base.endsWith('\n') ? '' : '\n'
16 if (base.includes(CLAUDE_MD_HEADING)) return base + glue + bulletOnly(content)
17 // A brand-new file opens on the heading itself; an existing one keeps the blank line before it.
18 return base + glue + (base === '' ? content.replace(/^\n+/, '') : content)
19}
20
21/** Renders the file an artifact kind writes: where it goes, what it says, how it lands. */
22export const render = (kind: ArtifactKind, p: Proposal, cwd: string): Pick<Artifact, 'path' | 'content' | 'mode'> => {
23 if (kind === 'skill') return { path: `${cwd}/.claude/skills/${slug(p.title)}/SKILL.md`, content: skillDoc(p), mode: 'write' }
24 if (kind === 'agent-brief') return { path: `${cwd}/.claude/agents/${slug(p.title)}.md`, content: briefDoc(p), mode: 'write' }
25 if (kind === 'settings-allow') return { path: `${cwd}/.claude/settings.json`, content: p.body, mode: 'merge-settings' }
26 return { path: `${cwd}/CLAUDE.md`, content: `\n${CLAUDE_MD_HEADING}\n${bulletOf(p.body)}`, mode: 'append' }
27}
28
29/** Returns the artifacts that make this session's decisions permanent, largest saving first. */
30export const propose = (state: State): Artifact[] =>
31 state.patterns.flatMap(p => artifactsOf(state, p)).sort((a, b) => b.savingPct - a.savingPct)
32
33/** Adds a permission rule to permissions.allow, keeping every other setting; idempotent. */
34export const mergeSettings = (existing: string | null, rule: string): string => {
35 const root = objectOf(existing)
36 const permissions = recordOf(root['permissions'])
37 const allow = Array.isArray(permissions['allow']) ? permissions['allow'] : []
38 const merged = { ...root, permissions: { ...permissions, allow: allow.includes(rule) ? allow : [...allow, rule] } }
39 return `${JSON.stringify(merged, null, 2)}\n`
40}
41
42const artifactsOf = (state: State, p: Pattern): Artifact[] => {
43 const proposal = proposalOf(p)
44 // An artifact written, tried or skipped this session is done: the pane never offers it twice.
45 if (proposal === null || state.written.includes(`${p.id}:${proposal.kind}`)) return []
46 return [{
47 patternId: p.id,
48 kind: proposal.kind,
49 title: proposal.title,
50 savingPct: pctOf(baseline(state, p).chars * 3, state.usage.window),
51 ...render(proposal.kind, proposal, state.cwd),
52 }]
53}
54
55// Two decisions where D2 wins over 5.4's wording: an Ignore is final for the session, so it proposes nothing even
56// when the judge attached a proposal; and a kill's permanent line is the scoped alternative, never its kill prompt.
57const proposalOf = (p: Pattern): Proposal | null => {
58 if (p.decision === null || p.decision === 'keep') return null
59 // A process finding has no permission to grant, and its lasting line is the lean way, not the one-time re-plan.
60 if (p.proposal !== null && (p.lean === null || p.proposal.kind !== 'settings-allow')) return p.proposal
61 const body = p.lean ?? (p.decision === 'kill' ? p.alternative : (p.instruction ?? p.alternative))
62 // The label is the rule, never the waste: a row offering `Write` reads as what would be written.
63 return { kind: 'claude-md', title: titleOf(body) || p.id, body }
64}
65
66// The rule's first clause as a label: what it tells Claude to do, capitalised and without its full stop.
67const titleOf = (body: string): string => {
68 const s = collapseWs(body).split(/[;,]/)[0]?.replace(/\.$/, '') ?? ''
69 return s.charAt(0).toUpperCase() + s.slice(1)
70}
71
72const skillDoc = (p: Proposal): string =>
73 `${frontmatter([field('name', slug(p.title)), field('description', p.title)])}${prose(p.body)}`
74
75const briefDoc = (p: Proposal): string => {
76 const lines = p.body.split('\n')
77 const head = leading(lines)
78 const rest = lines.slice(head.length)
79 const model = valueOf(head, 'model')
80 const tools = valueOf(head, 'tools') ?? BRIEF_TOOLS
81 const fields = [field('name', slug(p.title)), field('description', p.title), ...(model === null ? [] : [field('model', model)]), field('tools', tools)]
82 return `${frontmatter(fields)}${prose([...without(without(head, 'model'), 'tools'), ...rest].join('\n'))}`
83}
84
85const leading = (lines: string[]): string[] => {
86 const blank = lines.findIndex(l => l.trim() === '')
87 return blank === -1 ? lines : lines.slice(0, blank)
88}
89
90const frontmatter = (fields: string[]): string => `---\n${fields.join('\n')}\n---\n\n`
91
92const field = (key: string, value: string): string => `${key}: ${JSON.stringify(collapseWs(value))}`
93
94const prose = (body: string): string => `${body.trim()}\n`
95
96const isField = (line: string, key: string): boolean => line.trim().toLowerCase().startsWith(`${key}:`)
97
98const valueOf = (lines: string[], key: string): string | null => {
99 const line = lines.find(l => isField(l, key))
100 return line === undefined ? null : collapseWs(line.trim().slice(key.length + 1)) || null
101}
102
103const without = (lines: string[], key: string): string[] => {
104 const i = lines.findIndex(l => isField(l, key))
105 return i === -1 ? lines : [...lines.slice(0, i), ...lines.slice(i + 1)]
106}
107
108const objectOf = (text: string | null): Record<string, unknown> => {
109 if (text === null) return {}
110 try {
111 return recordOf(JSON.parse(text))
112 } catch {
113 return {}
114 }
115}
116
117const recordOf = (v: unknown): Record<string, unknown> =>
118 v !== null && typeof v === 'object' && !Array.isArray(v) ? { ...(v as Record<string, unknown>) } : {}
119hooks/core/spawns.ts 229 lines1import { RUN_FRESH_MS } from './types'
2import type { CommandClass, JournalEntry, Loop, Outcome, Row, Run, State } from './types'
3
4// The tools whose row is an edit whatever its paths say, the classes that are a check, the tools that are a read.
5const EDIT_TOOLS: readonly string[] = ['Edit', 'Write', 'NotebookEdit']
6const CHECK_CLASSES: readonly CommandClass[] = ['test', 'lint', 'typecheck', 'format', 'build']
7const READ_TOOLS: readonly string[] = ['Read', 'Grep', 'Glob']
8
9type Severity = 'critical' | 'high' | 'medium' | 'low'
10
11const isRecord = (v: unknown): v is Record<string, unknown> =>
12 typeof v === 'object' && v !== null && !Array.isArray(v)
13
14const str = (v: unknown): string | null => (typeof v === 'string' ? v : null)
15
16const isSeverity = (v: unknown): v is Severity =>
17 v === 'critical' || v === 'high' || v === 'medium' || v === 'low'
18
19// A reviewer's findings, counted by severity; a finding with none of the four names is not counted at all.
20const findingsOf = (findings: readonly unknown[]): Outcome =>
21 findings.reduce<Extract<Outcome, { kind: 'findings' }>>((counts, f) => {
22 const severity = isRecord(f) ? f['severity'] : undefined
23 return isSeverity(severity) ? { ...counts, [severity]: counts[severity] + 1 } : counts
24 }, { kind: 'findings', critical: 0, high: 0, medium: 0, low: 0 })
25
26// The words a prose review labels a finding with, under the four severities the outcome counts.
27const PROSE_SEVERITY: Readonly<Record<string, Severity>> = {
28 critical: 'critical', blocker: 'critical', important: 'high', high: 'high', major: 'high',
29 medium: 'medium', moderate: 'medium', minor: 'low', low: 'low', nit: 'low',
30}
31
32// A label, not a word in a sentence: `(Important)`, `[Minor]`, `**Critical**`, `Severity: high`, or `Minor:` opening a line or a bullet.
33const PROSE_LABEL = /\(\s*(\w+)\s*\)|\[\s*(\w+)\s*\]|\*\*(\w+)\*\*|severity:\s*(\w+)|^[\s>*\-\d.)#]*(\w+):/gim
34
35// What follows a label that found nothing under it: `Critical: none`, `**Minor:** 0`, `High: -`.
36const PROSE_NONE = /^[\s*_:]*(none|nothing|0|n\/a|na|—|-|no findings)[\s.,;:!*_]*$/i
37
38// A prose reply's labelled severities, counted; null when it labels none, so it stays a report.
39const proseOf = (text: string): Outcome | null => {
40 const counts = { kind: 'findings' as const, critical: 0, high: 0, medium: 0, low: 0 }
41 let labelled = 0
42 for (const m of text.matchAll(PROSE_LABEL)) {
43 const severity = PROSE_SEVERITY[(m.slice(1).find(g => g !== undefined) ?? '').toLowerCase()]
44 if (severity === undefined) continue
45 if (PROSE_NONE.test(text.slice(m.index + m[0].length).split('\n')[0] ?? '')) continue
46 counts[severity] += 1
47 labelled += 1
48 }
49 return labelled > 0 ? counts : null
50}
51
52// A findings list is counted, a prose review's labels are counted; anything else is a report, sized by its text or by the JSON it would print as.
53const outcomeOf = (result: unknown): Outcome => {
54 if (isRecord(result) && Array.isArray(result['findings'])) return findingsOf(result['findings'])
55 const prose = typeof result === 'string' ? proseOf(result) : null
56 if (prose !== null) return prose
57 return { kind: 'report', chars: typeof result === 'string' ? result.length : (JSON.stringify(result) ?? '').length }
58}
59
60const entryOf = (line: string): JournalEntry | null => {
61 let value: unknown
62 try {
63 value = JSON.parse(line)
64 } catch {
65 return null
66 }
67 if (!isRecord(value)) return null
68 const agentId = str(value['agentId'])
69 if (agentId === null) return null
70 if (value['type'] === 'started') return { kind: 'started', agentId, label: str(value['label']), phase: str(value['phase']) }
71 if (value['type'] === 'result') return { kind: 'result', agentId, outcome: outcomeOf(value['result']) }
72 return null
73}
74
75/** Reads a workflow journal: one JSON object per line, `started` and `result` entries kept, anything else skipped. */
76export const parseJournal = (text: string): JournalEntry[] =>
77 text.split('\n').map(entryOf).filter((entry): entry is JournalEntry => entry !== null)
78
79/** The run a `Workflow` result launched locally — its id, name and transcript dir — or null for anything else. */
80export const runOf = (value: unknown): { id: string; name: string; dir: string | null } | null => {
81 if (!isRecord(value) || value['taskType'] !== 'local_workflow') return null
82 const id = str(value['runId'])
83 const name = str(value['workflowName'])
84 return id === null || name === null ? null : { id, name, dir: str(value['transcriptDir']) }
85}
86
87/** The loop an `Agent` result closed — its id, the description it was given and the model that ran it — or null. */
88export const agentOf = (value: unknown): { agentId: string; description: string; model: string | null } | null => {
89 if (!isRecord(value)) return null
90 const agentId = str(value['agentId'])
91 return agentId === null ? null : { agentId, description: str(value['description']) ?? '', model: str(value['resolvedModel']) }
92}
93
94type Work = Pick<Row, 'tool' | 'cls' | 'paths'>
95
96/** An edit whatever its paths say, or any call that changed a path. */
97export const isEdit = (r: Work): boolean => EDIT_TOOLS.includes(r.tool) || r.paths.length > 0
98
99/** A test, lint, typecheck, format or build run. */
100export const isCheck = (r: Work): boolean => CHECK_CLASSES.includes(r.cls)
101
102/** A read or a search, by tool or by command class. */
103export const isRead = (r: Work): boolean => r.cls === 'read' || r.cls === 'search' || READ_TOOLS.includes(r.tool)
104
105const one = (yes: boolean): number => (yes ? 1 : 0)
106
107/** The loop with one more of its own rows counted: a call, and an edit, a check or a read when the row was one — counted as it lands, so the count outlives the row. */
108export const countRow = (loop: Loop, row: Work): Loop => ({
109 ...loop, calls: loop.calls + 1, edits: loop.edits + one(isEdit(row)), checks: loop.checks + one(isCheck(row)), reads: loop.reads + one(isRead(row)),
110})
111
112/** What a loop's own surviving rows did: how many calls, and how many of them edited, checked or read; the loop's own counters are the whole figure. */
113export const loopStats = (rows: readonly Row[], loop: Loop): { calls: number; edits: number; checks: number; reads: number } => {
114 const own = rows.filter(r => r.agent === loop.id)
115 return { calls: own.length, edits: own.filter(isEdit).length, checks: own.filter(isCheck).length, reads: own.filter(isRead).length }
116}
117
118// A run is going while one of its loops is, or while it is young enough that its first loop may not have reported yet.
119const isActive = (state: State, run: Run, now: number): boolean => {
120 const loops = state.loops.filter(l => l.run === run.id)
121 return loops.length === 0 ? now - run.at < RUN_FRESH_MS : loops.some(l => l.ended === null)
122}
123
124/** The runs still going: one with an unended loop, or one launched less than RUN_FRESH_MS ago with no loop yet. */
125export const activeRuns = (state: State, now: number): Run[] => state.runs.filter(run => isActive(state, run, now))
126
127/** Where a run's journal is written; null for a run that reported no transcript dir. */
128export const journalPath = (run: Run): string | null => (run.dir === null ? null : `${run.dir}/journal.jsonl`)
129
130// One token of a script's source: a string literal's value, or a single punctuation or word character run.
131type Token = { kind: 'string' | 'word' | 'punct'; text: string }
132
133// Reads the source as tokens, comments skipped, stopping at the first string or comment left open.
134const lex = (source: string): Token[] => {
135 const tokens: Token[] = []
136 let i = 0
137 while (i < source.length) {
138 const c = source[i] ?? ''
139 if (/\s/.test(c)) { i += 1; continue }
140 if (c === '/' && source[i + 1] === '/') { const end = source.indexOf('\n', i); i = end < 0 ? source.length : end; continue }
141 if (c === '/' && source[i + 1] === '*') { const end = source.indexOf('*/', i + 2); if (end < 0) return tokens; i = end + 2; continue }
142 if (c === "'" || c === '"' || c === '`') {
143 let j = i + 1
144 let text = ''
145 while (j < source.length && source[j] !== c) {
146 if (source[j] === '\\') j += 1
147 text += source[j] ?? ''
148 j += 1
149 }
150 if (j >= source.length) return tokens
151 tokens.push({ kind: 'string', text })
152 i = j + 1
153 continue
154 }
155 const word = /^[\w$]+/.exec(source.slice(i, i + 200))
156 if (word !== null) { tokens.push({ kind: 'word', text: word[0] }); i += word[0].length; continue }
157 tokens.push({ kind: 'punct', text: c })
158 i += 1
159 }
160 return tokens
161}
162
163const OPEN: Readonly<Record<string, string>> = { '{': '}', '[': ']', '(': ')' }
164
165// The index just past the bracket that closes the one at `at`, or -1 when the literal never closes.
166const closeOf = (tokens: readonly Token[], at: number): number => {
167 const stack: string[] = []
168 for (let i = at; i < tokens.length; i += 1) {
169 const t = tokens[i]
170 if (t === undefined || t.kind !== 'punct') continue
171 const close = OPEN[t.text]
172 if (close !== undefined) stack.push(close)
173 else if (t.text === stack[stack.length - 1]) {
174 stack.pop()
175 if (stack.length === 0) return i + 1
176 }
177 }
178 return -1
179}
180
181// A key of the object at depth one: a word or a quoted name followed by a colon.
182const isKey = (tokens: readonly Token[], i: number, name: string): boolean =>
183 tokens[i]?.text === name && tokens[i]?.kind !== 'punct' && tokens[i + 1]?.text === ':'
184
185// The value token range of `name` inside the object whose `{` is at `at`, or null.
186const valueAt = (tokens: readonly Token[], at: number, end: number, name: string): number | null => {
187 for (let i = at + 1, depth = 0; i < end - 1; i += 1) {
188 const t = tokens[i]
189 if (t?.kind === 'punct' && OPEN[t.text] !== undefined) depth += 1
190 else if (t?.kind === 'punct' && Object.values(OPEN).includes(t.text)) depth -= 1
191 else if (depth === 0 && isKey(tokens, i, name)) return i + 2
192 }
193 return null
194}
195
196// The titles of a phases array literal: a bare string element or an object's `title` string; null when an element is anything else.
197const titlesOf = (tokens: readonly Token[], at: number, end: number): string[] | null => {
198 const titles: string[] = []
199 let i = at + 1
200 while (i < end - 1) {
201 const t = tokens[i]
202 if (t?.kind === 'string') { titles.push(t.text); i += 1 }
203 else if (t?.text === '{') {
204 const close = closeOf(tokens, i)
205 const value = valueAt(tokens, i, close, 'title')
206 const title = value === null ? undefined : tokens[value]
207 if (title?.kind !== 'string') return null
208 titles.push(title.text)
209 i = close
210 } else return null
211 if (tokens[i]?.text === ',') i += 1
212 }
213 return titles
214}
215
216/** The phase titles a Workflow script's `meta` literal declares, read as a literal and never run; [] when there is no readable `meta.phases`. */
217export const phasesOf = (script: string): string[] => {
218 const tokens = lex(script)
219 const meta = tokens.findIndex((t, i) => t.text === 'meta' && tokens[i + 1]?.text === '=' && tokens[i + 2]?.text === '{')
220 if (meta < 0) return []
221 const open = meta + 2
222 const close = closeOf(tokens, open)
223 if (close < 0) return []
224 const value = valueAt(tokens, open, close, 'phases')
225 if (value === null || tokens[value]?.text !== '[') return []
226 const end = closeOf(tokens, value)
227 return end < 0 ? [] : (titlesOf(tokens, value, end) ?? [])
228}
229hooks/core/text.ts 194 lines1import type { ArtifactKind, Loop, Proposal, Row, Signature, State, StoredPattern, Usage } from './types'
2
3/** Returns the median of the numbers; 0 when there are none. */
4export const median = (xs: number[]): number => {
5 if (xs.length === 0) return 0
6 const sorted = [...xs].sort((a, b) => a - b)
7 const mid = Math.floor(sorted.length / 2)
8 return sorted.length % 2 !== 0 ? (sorted[mid] ?? 0) : (((sorted[mid - 1] ?? 0) + (sorted[mid] ?? 0)) / 2)
9}
10
11/** Returns chars/4/window*100 rounded to 1 decimal. */
12export const pctOf = (chars: number, window: number): number =>
13 Math.round((chars / 4 / window) * 1000) / 10
14
15/** Returns 100 - percent, or null when percent is undefined. */
16export const pctLeft = (u: Usage): number | null =>
17 u.percent !== undefined ? Math.round((100 - u.percent) * 10) / 10 : null
18
19/** Formats milliseconds as '11m', '3m 50s', or '12s'. */
20export const duration = (ms: number): string => {
21 const s = Math.round(ms / 1000)
22 if (s < 60) return `${s}s`
23 const mins = Math.floor(s / 60)
24 const secs = s % 60
25 return secs === 0 ? `${mins}m` : `${mins}m ${secs}s`
26}
27
28/** Returns what a length of text costs in tokens, four characters to one. */
29export const tokensOf = (chars: number): number => Math.round(chars / 4)
30
31/** Formats a count short and rounded: '9.9k', '41k', '800'. */
32export const kilo = (n: number): string =>
33 n >= 10_000 ? `${Math.round(n / 1000)}k` : n >= 1000 ? `${Math.round(n / 100) / 10}k` : `${n}`
34
35/** Truncates text to `cells` characters, ending it with '…' when it is cut and never a space before it. */
36export const fit = (text: string, cells: number): string =>
37 text.length <= cells ? text : `${text.slice(0, Math.max(0, cells - 1)).trimEnd()}…`
38
39/** Returns a kebab-case slug of at most forty characters. */
40export const slug = (s: string): string =>
41 s
42 .toLowerCase()
43 .replace(/[^a-z0-9]+/g, '-')
44 .replace(/^-+|-+$/g, '')
45 .slice(0, 40)
46
47/** Collapses runs of whitespace (including newlines) to a single space and trims. */
48export const collapseWs = (s: string): string => s.replace(/\s+/g, ' ').trim()
49
50/** Returns JSON with keys sorted, top-level keys in omit removed. */
51export const stableJson = (v: unknown, omit: readonly string[]): string => {
52 const replacer = (_key: string, val: unknown): unknown => {
53 if (val !== null && typeof val === 'object' && !Array.isArray(val)) {
54 const obj = val as Record<string, unknown>
55 return Object.keys(obj)
56 .filter(k => !omit.includes(k))
57 .sort()
58 .reduce<Record<string, unknown>>((acc, k) => { acc[k] = obj[k]; return acc }, {})
59 }
60 return val
61 }
62 return JSON.stringify(v, replacer)
63}
64
65/** Renders a filled/empty gauge of `width` cells for a percentage 0..100. */
66export const gauge = (percent: number, width: number): string => {
67 const filled = Math.round(Math.max(0, Math.min(100, percent)) / 100 * width)
68 return '█'.repeat(filled) + '░'.repeat(width - filled)
69}
70
71/** Wraps a user instruction in the standard ContextSaver prefix. */
72export const instructionOf = (text: string): string =>
73 `Instruction from the user (via ContextSaver): ${text}`
74
75/** Produces the kill prompt for a stored pattern. */
76export const killPrompt = (p: StoredPattern): string =>
77 `Stop this behaviour for the rest of the session: ${p.kind}. From now on: ${p.alternative}`
78
79// Reply parsing shared by the habit and process judges: the JSON object, the evidence handles each may
80// cite, and the grounded ceiling on what a finding may claim.
81
82export const unique = (xs: string[]): string[] => [...new Set(xs)]
83
84export const str = (v: unknown): string => (typeof v === 'string' ? v : '')
85
86export const isRecord = (v: unknown): v is Record<string, unknown> =>
87 typeof v === 'object' && v !== null && !Array.isArray(v)
88
89export const isArtifactKind = (v: unknown): v is ArtifactKind =>
90 typeof v === 'string' && ['claude-md', 'skill', 'agent-brief', 'settings-allow'].includes(v)
91
92export const parseObject = (text: string): Record<string, unknown> | null => {
93 const start = text.indexOf('{')
94 const end = text.lastIndexOf('}')
95 if (start < 0 || end < start) return null
96 try {
97 const value: unknown = JSON.parse(text.slice(start, end + 1))
98 return isRecord(value) ? value : null
99 } catch {
100 return null
101 }
102}
103
104export const short = (text: string, max: number): string => (text.length <= max ? text : `${text.slice(0, max)}…`)
105
106// The model's own text quoted inside a one-line drop reason: whitespace collapsed first, so a
107// heredoc key or a multi-line id can never break the one-reason-per-line contract `debugDump` keeps.
108export const cell = (value: unknown, max: number): string => short(collapseWs(str(value)), max)
109
110export const AGENT_HANDLE = /^agent:(.+)$/
111
112// Handles that name no single row: a turn of the main loop, or a whole agent loop.
113export const isWide = (handle: string): boolean => /^turn:\d+$/.test(handle) || AGENT_HANDLE.test(handle)
114
115export const handleDrop = (value: unknown, signature: Signature | null): string => {
116 const handle = cell(value, 40)
117 if (handle.length === 0) return 'evidence handle is not a string'
118 if (isWide(handle) && signature !== null) return `evidence ${handle} needs signature null`
119 return AGENT_HANDLE.test(handle) ? `evidence ${handle} not in AGENTS` : `evidence ${handle} not in the ledger`
120}
121
122// The loop an `agent:<alias>` handle names, under the alias table AGENTS and LEDGER printed — the one the
123// prompt was built with, not today's: a judge run takes minutes, and a new agent's first row landing in
124// that time (or an old agent's last row falling off ROW_CAP) renumbers every loop the rows never showed.
125export const loopOf = (state: State, aliases: ReadonlyMap<string, string>, alias: string): Loop | undefined =>
126 state.loops.find(l => aliases.get(l.id) === alias)
127
128export const handleOf = (value: unknown, state: State, visible: Row[], aliases: ReadonlyMap<string, string>, signature: Signature | null): string | null => {
129 const handle = str(value)
130 const alias = /^r(\d+)$/.exec(handle)
131 if (alias !== null) {
132 const row = visible.find(r => r.seq === Number(alias[1] ?? ''))
133 return row !== undefined ? row.id : null
134 }
135 if (signature !== null) return null
136 const turn = /^turn:(\d+)$/.exec(handle)
137 if (turn !== null) {
138 const n = Number(turn[1] ?? '')
139 return state.turns.some(t => t.turn === n) ? `turn:${n}` : null
140 }
141 const agent = AGENT_HANDLE.exec(handle)
142 if (agent !== null) {
143 const loop = loopOf(state, aliases, agent[1] ?? '')
144 // Stored by id, not alias: the alias is a naming of the rows in hand and can renumber between runs.
145 return loop !== undefined ? `agent:${loop.id}` : null
146 }
147 return null
148}
149
150// A string is the reason the evidence cannot be used; the array is the handles it maps to.
151export const evidenceOf = (value: unknown, state: State, visible: Row[], aliases: ReadonlyMap<string, string>, signature: Signature | null): string[] | string => {
152 if (!Array.isArray(value) || value.length === 0) return 'evidence must be a non-empty array of row ids'
153 const handles = value.map(h => handleOf(h, state, visible, aliases, signature))
154 const badAt = handles.indexOf(null)
155 if (badAt >= 0) return handleDrop(value[badAt], signature)
156 return unique(handles.filter((h): h is string => h !== null))
157}
158
159export const citedLoops = (state: State, evidence: string[]): Loop[] =>
160 evidence.flatMap(handle => {
161 const agent = AGENT_HANDLE.exec(handle)
162 const loop = agent === null ? undefined : state.loops.find(l => l.id === agent[1])
163 return loop !== undefined ? [loop] : []
164 })
165
166export const citedTurns = (state: State, evidence: string[]): number[] =>
167 evidence.flatMap(handle => {
168 const turn = /^turn:(\d+)$/.exec(handle)
169 if (turn !== null) return [Number(turn[1] ?? '')]
170 if (AGENT_HANDLE.test(handle)) return citedLoops(state, [handle]).map(l => l.firstTurn)
171 const row = state.rows.find(r => r.id === handle)
172 return row !== undefined ? [row.turn] : []
173 })
174
175export const loopTokens = (l: Loop): number => l.tokens.input + l.tokens.cacheCreate + l.tokens.output
176
177// What one avoided loop would have cost when the finding cites loops alone; else the cited turns' answers.
178export const groundedCap = (state: State, evidence: string[]): number => {
179 const loops = citedLoops(state, evidence)
180 if (loops.length > 0 && !evidence.some(h => /^turn:\d+$/.test(h))) return Math.floor(median(loops.map(loopTokens)))
181 const turns = citedTurns(state, evidence)
182 return Math.floor(median(state.turns.filter(t => turns.includes(t.turn)).map(t => t.answerChars)) / 4)
183}
184
185export const proposalOf = (value: unknown): Proposal | null => {
186 if (!isRecord(value)) return null
187 const kind = value['kind']
188 const title = collapseWs(str(value['title']))
189 const body = str(value['body']).trim()
190 if (!isArtifactKind(kind) || title.length === 0 || body.length === 0) return null
191 if (kind === 'settings-allow' && !/^[A-Za-z][A-Za-z0-9_]*\(.+\)$/.test(body)) return null
192 return { kind, title, body }
193}
194hooks/core/types.ts 286 lines1import type { Elements } from 'claude-code'
2
3export const PLUGIN_NAME = 'contextsaver'
4export const PANE_ID = 'saver'
5export const PANE_TITLE = 'ContextSaver'
6export const PANE_INLINE_ROWS = 18 // body rows requested when seated inline above the prompt (the compact card is framed)
7export const AUTO_OPEN_MIN_COLUMNS = 144 // unasked opens wait undrawn below this width (d.ts 1943-1945)
8export const STEER_RING_TRIES = 8 // frames the Fix… field's ring is asked for before the composer route is said
9export const STEER_RING_WAIT_MS = 40 // a frame and a little: the shown pane redraws at most thirty times a second
10export const COMMAND = { name: 'saver', description: 'ContextSaver: toggle the pane · check | fix [n] [text] | ignore <n> | debug | reset', argumentHint: '[check | fix [n] [text] | ignore <n> | debug | reset]' } as const
11export const SETTLE_TURNS = 2 // an instruction not ignored for this many turns is credited
12export const JUDGE_MIN_NEW_TOKENS = 30_000
13export const JUDGE_MIN_TURNS = 3
14export const JUDGE_MIN_ROWS = 8
15export const JUDGE_MAX_BACKOFF = 4
16export const JUDGE_BUDGET_SHARE = 0.03
17export const JUDGE_LEDGER_ROWS = 150 // full rows rendered; older rows are folded into `~` summary lines
18export const MAX_FINDINGS = 6
19export const MAX_BEHAVIORAL_FINDINGS = 3 // findings with signature: null per judge run (agent findings are signature-null too)
20export const MAX_PATTERNS = 50
21export const ROW_CAP = 2000
22export const KIND_MAX = 120
23export const ALTERNATIVE_MAX = 200
24export const KEY_MAX = 200
25export const DEBUG_MAX_LINES = 40 // `/saver debug` ceiling
26export const DEBUG_MAX_PATTERNS = 16 // pattern lines `/saver debug` prints before folding the rest
27export const DEBUG_MAX_DROPPED = 6 // dropped-finding reasons `/saver debug` and the debug log print
28export const BRIEF_TOOLS = 'Read, Grep, Glob' // the tools an agent brief allows when the proposal names none
29export const FILE_TOOLS: readonly string[] = ['Read', 'Edit', 'Write', 'NotebookEdit'] // tools whose ledger key is the path they touched
30export const CLAUDE_MD_HEADING = '## ContextSaver'
31export const RECOVERED_FLAG = 'recovered' // `Row.flags` marker for a row rebuilt from the transcript: its `ms` is 0 and its agent reads `main`
32export const MAIN_AGENT = 'main' // `Row.agent` of the main loop; the alias table leaves it as it is
33export const NO_CALLS = 'no tool calls' // `Evidence.what` of a turn handle: that turn ran none
34export const CARD_EVIDENCE = 3 // cited calls one card's details show, newest first
35export const JUDGE_MIN_NEW_ROWS = 40 // mid-turn cadence: ledger rows since the last run (a turn can last hours)
36export const JUDGE_MIN_GAP_MS = 300_000 // mid-turn cadence: at least five minutes between runs
37export const TREND_TURNS = 10 // context samples the header's trend draws
38export const SINKS = 3 // named sinks the Time and Context rows show
39export const LOOP_CAP = 400 // loops kept (oldest dropped)
40export const AGENTS_ROWS = 60 // loop lines the AGENTS block renders in full; older ones fold per run
41export const RUN_REFRESH_MS = 10_000 // a running workflow's journal is re-read at most this often
42export const RUN_FRESH_MS = 600_000 // a run with no loop yet counts as active this long after its launch
43export const STORE_KEY_VERSION = 'v5' // 0.5 starts fresh: a 0.4 `patterns:<cwd>` entry is never read
44export const storeKey = (cwd: string): string => `patterns.${STORE_KEY_VERSION}:${cwd}`
45export const FOLD_LINES = 40 // `~` lines the LEDGER prints, largest by characters first; the rest become one `~ ×N more` line
46export const JUDGE_AGENT_ROWS = 40 // rows one agent loop may hold in the citable window, so no single loop fills it
47export const ASK_HEAD_MAX = 200 // characters of a typed prompt kept for the process digest; session-only, never stored
48export const DIGEST_MAX_CHARS = 40_000 // the PROCESS digest's ceiling; Shape is never the section cut
49export const PROCESS_CLOCK_MS = 1_800_000 // the process judge's clock: at most every thirty minutes, doubled past the budget share
50export const PROCESS_MIN_CONFIDENCE = 0.75 // a process finding below this is dropped
51export const MAX_PROCESS_FINDINGS = 3 // process findings kept per run
52export const DELIVERY_BURST_MS = 5_000 // subagent deliveries inside this window share one toast
53
54export type CommandClass = 'test' | 'lint' | 'format' | 'typecheck' | 'build' | 'install' | 'git' | 'read' | 'search' | 'other'
55export type Category = 'execution' | 'reading' | 'production' | 'behavior' | 'communication' | 'multi-agent' | 'environment' | 'process' | 'other'
56export type Choice = 'keep' | 'steer' | 'kill'
57export type ProcessTrigger = 'plan' | 'workflow' | 'pace' | 'clock' // an accepted ExitPlanMode, a Workflow launch, a typed pace complaint, the clock
58/** One prompt the person typed (`composer` or `bridge` origin): the process judge's view of what was asked. Session-only. */
59export type Ask = { turn: number; at: number; head: string; pace: boolean } // head: ≤ ASK_HEAD_MAX chars, control characters stripped; pace: it reads as a complaint about pace
60
61export type Row = {
62 seq: number // 1-based, monotonically increasing; the judge sees `r${seq}`
63 id: string // tool_use_id
64 tool: string; key: string; cls: CommandClass
65 agent: string // e.agentId ?? 'main'
66 turn: number
67 ms: number; chars: number
68 head: string // first 80 chars of result.text, control characters stripped; quoted as evidence in the pane, never sent to the judge
69 flags: string[] // 'err' (tool reported an error) | 'denied' (result.deny: the user or a policy said no) | 'dedup' (Read type 'file_unchanged') | 'trunc' (truncatedByTokenCap) | 'bg' (run_in_background or backgroundTaskId) | 'timeout' (timedOutAfterMs) | `persist=${persistedOutputSize}` | 'ask' (AskUserQuestion: ms is the wait for the person) | 'recommended' (an ask whose questions carry '(Recommended)') | 'recovered' (rebuilt from the transcript at load: ms is 0 and agent reads 'main')
70 lines: { add: number; del: number } | null // Edit: gitDiff.additions/deletions else counted from structuredPatch; Write: content line count as add
71 paths: string[] // absolute paths this call edited (Edit/Write filePath unless staged; Bash bashEditDiff.changedFiles)
72 spawn: { type: string; requested: string | null; resolved: string | null; status: string | null; tokens: number | null; edits: number | null; promptChars: number } | null // Agent rows only
73}
74export type TurnStat = { turn: number; input: number; output: number; cacheRead: number; cacheCreate: number; calls: number; ms: number; answerChars: number; answerHead: string; aborted: boolean; ended: TurnEnd; at: number; idleMs: number; context: number | null } // answerHead: first 100 chars of e.answer, for evidence quotes; ended: how the turn ended; at: clock at completion; idleMs: wait until the next turn started (0 until it does); context: tokens in the window after the turn (usage), null when unknown — growth between turns is the pace compaction runs at
75export type TurnEnd = 'answer' | 'aborted' | 'refusal' | 'error'
76export type Tokens = { input: number; output: number; cacheRead: number; cacheCreate: number }
77export type Outcome = { kind: 'findings'; critical: number; high: number; medium: number; low: number } | { kind: 'report'; chars: number }
78export type Loop = { id: string; run: string | null; label: string | null; phase: string | null; model: string | null; turns: number; ms: number; tokens: Tokens; ended: TurnEnd | null; firstTurn: number; firstSeq: number; outcome: Outcome | null; calls: number; edits: number; checks: number; reads: number } // one spawned agent loop: `id` is its agentId (the ledger's `agent`), `run` the workflow run that launched it, `ended` null while it runs; calls/edits/checks/reads count its own rows as they land, so the line stays whole once ROW_CAP drops them
79export type Folded = { tool: string; key: string; cls: CommandClass; agent: string; count: number; ms: number; chars: number; firstTurn: number; lastTurn: number; flags: { ask: number; recommended: number; err: number } } // one (tool, key) pair's rows dropped past ROW_CAP: nothing citable, everything counted; `agent` is 'main' or the first loop seen
80export type Run = { id: string; name: string; dir: string | null; turn: number; seq: number; at: number; refreshedAt: number; phases: string[] } // at: clock at launch; refreshedAt: last journal read (0 never); phases: the script's declared `meta.phases` titles, [] when unread or undeclared
81export type JournalEntry = { kind: 'started'; agentId: string; label: string | null; phase: string | null } | { kind: 'result'; agentId: string; outcome: Outcome }
82export type Signature = { tool: string; key: string }
83export type ArtifactKind = 'claude-md' | 'skill' | 'agent-brief' | 'settings-allow'
84export type Proposal = { kind: ArtifactKind; title: string; body: string }
85
86/** Persisted across sessions, per project. */
87export type StoredPattern = {
88 id: string // `${category}:${slug}`, slug ≤ 40
89 category: Category
90 kind: string // the behaviour, one sentence ≤ 120 chars: "Claude keeps running `bun test` after every step"
91 signature: Signature | null // null = behavioural, no single command carries it
92 why: string
93 alternative: string // the fix, one imperative sentence ≤ 200 chars written for Claude; pre-fills Fix…, completes Fix
94 confidence: number // 0.5..1
95 proposal: Proposal | null
96 estTokensPerTurn: number | null // judge's estimate for behavioural patterns; null when a signature exists
97 lastDecision: Choice | null // the most recent session's decision, for the judge's calibration
98 lean: string | null // process findings only: how a lean expert would run it (the card's Lean line); `kind` is what the session does; null for a habit
99}
100/** Session-only fields. */
101export type Pattern = StoredPattern & {
102 hits: string[] // evidence handles: row ids (tool_use_id), `turn:<n>` or `agent:<agentId>`; grows with matching rows; cost/baseline from rows only
103 decision: Choice | null
104 decidedAtTurn: number | null
105 instruction: string | null // the text sent for steer/kill
106 openedAtTurn: number | null // turn of the last steer/kill; settles saved (credited or ignored)
107 ignored: number // times the instruction was ignored
108 sent: number // subagents its standing instruction reached this session (the card's "sent ×N")
109 cost?: { tokens: number; ms: number } // process only: the newest run's claimed cost, clamped to its cited handles; absent, the card shows what they measured
110}
111
112/** What one fork of the judge cost, in the four token counts the API reports. */
113export type JudgeUsage = { input: number; output: number; cacheRead: number; cacheCreate: number }
114/** What one judge run reported: findings returned, findings kept, one short line per drop and per cap a kept finding missed, and what the fork cost (null when it returned nothing). */
115export type JudgeRun = { returned: number; kept: number; dropped: readonly string[]; usage: JudgeUsage | null }
116
117/** One cited call (or turn) as the details render it: what ran, in which loop, what it cost, what it answered. */
118export type Evidence = {
119 turn: number
120 what: string // the command for Bash, the path for a file tool, `tool key` otherwise; NO_CALLS for a turn handle
121 agent: string | null // the loop's alias (a1, a2…), null when it was the main loop's own call
122 ms: number // 0 when nothing measured it (a turn handle, a recovered row)
123 chars: number // in-context size of the result, or of the turn's answer
124 head: string // the first line of what came back, quoted under the call; '' when there is none
125}
126/** One waster as the pane draws it: the behaviour, the stats, the fix, and the receipts behind `i`. */
127export type Card = {
128 patternId: string
129 n: number // 1-based seat in the pane's list, top to bottom
130 category: Category // the dim tag on the title row
131 kind: string
132 stats: string
133 why: string
134 fix: string
135 lean: string | null // process cards: the Lean line; null for a habit card
136 total: { unit: 'calls' | 'turns' | 'agents'; calls: number; ms: number; chars: number } // cited calls, their wall time and their context; `unit: 'turns'` when the pattern cites turns instead, so `calls` counts turns and `chars` is the per-turn estimate; `unit: 'agents'` when it cites loops, so `calls` counts loops and `ms` is their sum
137 evidence: readonly Evidence[] // ≤ CARD_EVIDENCE cited calls, newest first
138}
139export type Artifact = { patternId: string; kind: ArtifactKind; title: string; path: string; content: string; savingPct: number; mode: 'append' | 'write' | 'merge-settings' }
140export type Usage = { tokens?: number; window: number; percent?: number; compactAt?: number }
141
142export type State = {
143 cwd: string
144 turn: number
145 seq: number // last Row.seq issued
146 rows: Row[] // capped at ROW_CAP (oldest dropped)
147 folded: Record<string, Folded> // the rows ROW_CAP dropped, folded per `${tool}\t${key}`: STATS, the LEDGER's `~` lines and the sinks count them, so a long session's totals stay whole
148 turns: TurnStat[]
149 loops: Loop[] // every agent loop seen, oldest first, capped at LOOP_CAP (oldest dropped)
150 runs: Run[] // every Workflow launched this session, oldest first
151 usage: Usage
152 overhead: { memory: number; mcp: number; agents: number } | null
153 compactions: number[] // turn indices at which session.compact fired
154 patterns: Pattern[]
155 cards: string[] // pattern ids awaiting a decision, newest first (the pane's WASTERS list)
156 expanded: string | null // pattern id whose (i) details are open; one at a time
157 steering: string | null // pattern id whose Fix… field is open
158 steerDraft: string | null // what the person has typed so far: every render draws it back into the field, so a redraw for any other reason never wipes it
159 notes: string[] // one-shot texts: drained into the next tool result or prompt
160 standing: string[] // texts re-sent with every prompt this session
161 written: string[] // `${patternId}:${kind}` of artifacts written, tried or skipped this session; propose() omits them
162 judge: { lastAtTokens: number; lastAtTurn: number; lastAtSeq: number; lastAtMs: number; running: boolean; runs: number; spent: number; backoff: number; error: string | null; focus: string | null; time: string | null; context: string | null; last: JudgeRun | null } // lastAtSeq/lastAtMs: the mid-turn cadence; time/context: the judge's one-line explanations of where they went
163 asks: Ask[] // prompts the person typed, oldest first; session-only
164 process: { lastAtMs: number; lastAtSeq: number; lastAtTurn: number; running: boolean; runs: number; spent: number; backoff: number; error: string | null; last: JudgeRun | null } // the process judge's own cadence and cost, beside `judge`
165 delivery: { agents: string[]; pending: number; lastToastAt: number } // agent ids the standing instructions reached; spawns in flight (counted as reached); when the last delivery toast showed
166 pendingCheck: boolean // a check armed at load and not yet answered: a plugin that joined a session with history fires one there and retries it at every warm opportunity until a run answers (§6)
167 paneOpen: boolean
168 autoOpened: boolean // the pane auto-opened once this session (like /diff on the first edit)
169 columns: number | null // last band width seen (e.props.bodyColumns), for the auto-open decision
170 saved: { ms: number; chars: number }
171}
172
173export const initialState = (cwd: string, window: number): State => ({
174 cwd, turn: 0, seq: 0, rows: [], folded: {}, turns: [], loops: [], runs: [], usage: { window }, overhead: null, compactions: [], patterns: [], cards: [], expanded: null, steering: null, steerDraft: null, notes: [], standing: [], written: [],
175 judge: { lastAtTokens: 0, lastAtTurn: 0, lastAtSeq: 0, lastAtMs: 0, running: false, runs: 0, spent: 0, backoff: 1, error: null, focus: null, time: null, context: null, last: null },
176 asks: [], process: { lastAtMs: 0, lastAtSeq: 0, lastAtTurn: 0, running: false, runs: 0, spent: 0, backoff: 1, error: null, last: null }, delivery: { agents: [], pending: 0, lastToastAt: 0 },
177 pendingCheck: false, paneOpen: false, autoOpened: false, columns: null, saved: { ms: 0, chars: 0 },
178})
179
180export type Action =
181 | { type: 'turn.start'; now: number } // now: dates the idle wait since the last completed turn
182 | { type: 'loop.turn'; agentId: string; model: string | null; ms: number; tokens: Tokens; ended: TurnEnd; turn: number } // one turn of an agent loop completed
183 | { type: 'agent.start'; agentId: string; description: string; model: string | null } // an Agent tool result: names the loop
184 | { type: 'run.start'; run: { id: string; name: string; dir: string | null }; now: number } // a Workflow tool result: a run launched (an id already present is a resume)
185 | { type: 'run.journal'; runId: string; entries: JournalEntry[]; now: number } // a run's journal read: stages and outcomes of its loops
186 | { type: 'row'; row: Omit<Row, 'seq'> }
187 | { type: 'adopt'; rows: readonly Omit<Row, 'seq'>[] } // rows rebuilt from the transcript of a session joined late
188 | { type: 'turn.complete'; stat: Omit<TurnStat, 'turn' | 'calls'> }
189 | { type: 'usage'; usage: Usage; now: number }
190 | { type: 'overhead'; overhead: { memory: number; mcp: number; agents: number } }
191 | { type: 'compact' }
192 | { type: 'expand'; patternId: string | null } // (i) toggled; null collapses
193 | { type: 'steer.begin'; patternId: string }
194 | { type: 'steer.draft'; text: string }
195 | { type: 'decide'; patternId: string; choice: Choice; text?: string } // text required for steer
196 | { type: 'judge.start'; now: number; seq: number } // when and at which ledger row the run began, for the mid-turn cadence
197 | { type: 'judge.done'; patterns: Pattern[]; fresh: string[]; recurred: string[]; focus: string | null; time: string | null; context: string | null; spent: number; error: string | null; returned: number; kept: number; dropped: readonly string[]; usage: JudgeUsage | null }
198 | { type: 'ask'; ask: Ask } // a prompt the person typed
199 | { type: 'run.phases'; runId: string; phases: string[] } // the run's script declared these phases
200 | { type: 'process.start'; now: number; seq: number; trigger: ProcessTrigger }
201 | { type: 'process.cold'; was: Pick<State['process'], 'lastAtMs' | 'lastAtSeq' | 'lastAtTurn'> } // the fork had no warm snapshot: the run never happened, so the cadence it spent is handed back
202 | { type: 'process.done'; patterns: Pattern[]; fresh: string[]; recurred: string[]; spent: number; error: string | null; returned: number; kept: number; dropped: readonly string[]; usage: JudgeUsage | null }
203 | { type: 'delivery.pending'; delta: 1 | -1 } // an agent.spawn rewrite began (1) or finished (-1)
204 | { type: 'delivery.sent'; agentId: string; now: number } // the standing instructions reached this agent, by spawn rewrite or on its first call
205 | { type: 'delivery.toasted'; now: number }
206 | { type: 'check.arm' } // the ledger a session.start adopted already passes the row floor: judge it there, and again at the next warm opportunity if that run answers nothing
207 | { type: 'notes.drained' }
208 | { type: 'standing.add'; text: string }
209 | { type: 'artifact.done'; patternId: string; kind: ArtifactKind; written: boolean } // written: true once the rule is handled — written, tried or skipped — and recorded in state.written
210 | { type: 'pane'; open: boolean; auto?: true }
211 | { type: 'columns'; columns: number }
212 | { type: 'reset' }
213
214/** Judge output after validation (section 5.3). */
215export type Finding = { id: string; category: Category; kind: string; evidence: string[]; signature: Signature | null; why: string; alternative: string; confidence: number; estTokensPerTurn: number | null; proposal: Proposal | null; lean: string | null } // lean: non-null for a process finding
216
217/** UI ↔ shell interface. */
218export type Ui = Pick<Elements['terminal'], 'Box' | 'Text' | 'Button' | 'Input' | 'Raster'>
219export type Site = { bodyColumns: number; maxRows: number }
220export type Actions = {
221 keep(patternId: string): void
222 steer(patternId: string): void // toggles the Fix… field under the waster's verbs
223 steerDraft(text: string): void // every keystroke: the state keeps the text and the redraw is what paints it
224 steerSubmit(patternId: string, text: string): void // Enter in the field, or /saver fix <text>
225 kill(patternId: string): void
226 info(patternId: string): void // toggles the (i) details
227 togglePane(): void
228 check(): void // run the judge now
229 write(a: Artifact): void
230 tryOnce(a: Artifact): void
231 skip(a: Artifact): void
232}
233/** View models: computed by patterns.ts from State, rendered by ui.tsx. Keeps the UI free of state logic. */
234export type Sink = { label: string; amount: number; count: number } // one named consumer: `tests`, `reads`, `agents`, `git`…; amount in ms (time) or chars (context)
235export type Sinks = { total: number; sinks: readonly Sink[] } // total over every ledger row of the main loop plus the agents' own rows (Agent spawn rows excluded: they contain their loop's rows); the SINKS largest named
236export type Header = {
237 percent: number | null // context used, 0..100
238 tokensToCompaction: number | null // exact: threshold - tokens
239 turnsToCompaction: number | null // estimate at the recent pace: tokens to compaction / median context growth per turn
240 trend: readonly number[] // context percent after each of the last TREND_TURNS turns, oldest first; [] before the first
241 time: Sinks | null // where the wall-clock went, from the ledger; null before the first row
242 context: Sinks | null // where the context went, from the ledger; null before the first row
243 judgeTime: string | null // the judge's one-line explanation of the time, verbatim; null until it has run
244 judgeContext: string | null // the judge's one-line explanation of the context
245 judgeRuns: number
246 judgeTokens: number // tokens the judge has spent this session; the pane's JUDGE row
247 judgeShare: number // those tokens as a percentage of the session's, 1 decimal; 0 while the session has none; `/saver debug` only
248 judgeRunning: boolean // a run is in flight: Check now reads `Checking…`, dims, and ignores presses
249 savedPct: number
250 savedMs: number
251}
252export type DecidedRow = {
253 patternId: string
254 choice: Choice
255 kind: string
256 savedPct: number | null // what one avoided repeat is worth; a rate until `settled`, a credit after it
257 settled: boolean // the instruction was neither ignored nor still in flight, so the saving is real
258 instruction: string | null // the sentence the user sent, when it is not the fix the card offered
259 ignored: number
260 sent: number // subagents its standing instruction reached (the row's "sent ×N")
261}
262export type PaneModel = {
263 header: Header
264 wasters: Card[] // undecided patterns, newest first
265 expanded: string | null
266 steering: string | null
267 steerDraft: string | null
268 decided: DecidedRow[] // newest first
269 artifacts: Artifact[]
270}
271/** The band's one teaser line: which of the four states the session is in, and the figures that state names. */
272export type BandModel = {
273 state: 'died' | 'checking' | 'found' | 'saved' | 'watching' // the first that applies: the last turn died, a judge run in flight, cards waiting, a saving credited, else watching
274 died: TurnEnd | null // the last turn's `ended` when it was an error or a refusal and no turn has started since
275 running: { name: string; loops: number; calls: number; label: string | null } | null // the newest active workflow run: its loops, their rows, the newest unended loop's stage
276 fresh: number // cards awaiting a decision
277 costPct: number // what those cards have already cost, as a share of the window
278 costMs: number // and in wall time
279 savedPct: number
280 savedMs: number
281 calls: number // ledger rows watched this session
282 paneOpen: boolean
283}
284export type BandProps = { ui: Ui; model: BandModel; site: Site; actions: Actions }
285export type PaneProps = { ui: Ui; model: PaneModel; site: Site; placement: 'dock' | 'inline'; actions: Actions }
286hooks/host.ts 52 lines1import type {
2 CommandSpec,
3 ModelForkResult,
4 PaneCloseArgs,
5 PaneOpenArgs,
6 SessionMessage,
7 SessionUsage,
8 SessionUsageArgs,
9 UiFocusArgs,
10 UiFocusResult,
11} from 'claude-code'
12
13/** Host table: one lambda per `$.noun.verb` call, bound in session.start. */
14export type Host = {
15 /** $.clock.now() — current time in milliseconds. */
16 now(): Promise<number>
17 /** $.clock.sleep(ms) — resolve after ms milliseconds. */
18 sleep(ms: number): Promise<void>
19 /** $.ui.invalidate('ui.render') — request a redraw (fire-and-forget). */
20 invalidate(): void
21 /** $.ui.toast(text) — show a transient notification (fire-and-forget). */
22 toast(text: string): void
23 /** $.ui.log(text) — emit a log line (fire-and-forget). */
24 log(text: string): void
25 /** $.ui.open(args) — open the named pane. */
26 openPane(args: PaneOpenArgs): Promise<void>
27 /** $.ui.close(args) — close the named pane. */
28 closePane(args: PaneCloseArgs): Promise<void>
29 /** $.ui.focus(args) — move a site's focus ring onto one of this plugin's elements. */
30 focusElement(args: UiFocusArgs): Promise<UiFocusResult>
31 /** $.command.register(spec) — register a slash command. */
32 registerCommand(spec: CommandSpec): Promise<{ command: string }>
33 /** $.session.usage(args?) — read context window usage. */
34 usage(args?: SessionUsageArgs): Promise<SessionUsage>
35 /** $.session.messages() — read the transcript so far, newest 4096 messages. */
36 messages(): Promise<SessionMessage[]>
37 /** $.store.get(key) — read a value from the plugin store. */
38 storeGet(key: string): Promise<unknown>
39 /** $.store.set(key, v) — write a value to the plugin store. */
40 storeSet(key: string, v: unknown): Promise<void>
41 /** $.model.fork({ prompt }) — run a detached model completion over the session transcript. */
42 fork(prompt: string): Promise<ModelForkResult | null>
43 /** $.fs.read(p) — read a file as a string. */
44 readFile(p: string): Promise<string>
45 /** $.fs.write(p, t) — write a string to a file. */
46 writeFile(p: string, t: string): Promise<void>
47 /** $.fs.exists(p) — check if a path exists. */
48 exists(p: string): Promise<boolean>
49 /** $.env.get('CONTEXTSAVER_DEBUG') — read the debug flag. */
50 debugFlag(): Promise<string | undefined>
51}
52