SLOPSHOPPER

contextsaver

Diagnoses wrong processes, workflows and habits in Claude Code sessions: fires on plan acceptance, workflow launches, pace complaints and the clock, names the…

newpanebandguardcommandtoast
★ 26v0.5.0MITupdated 2026-09-27AlmogBaku/ContextSaver
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · contextsaver
│ ┃ ContextSaver ✕ › fix the failing auth test and add an audit log call │ ┃ ▄██▀██▄ ContextSaver ↻ C │ ┃ ▄▄ ▀ 49% of context · 83k to compa ⏺ Read(src/auth.ts) │ ┃ ██ ▄ ███████████████░░░░░░░░░░░░░░ ⎿ Read 6 lines │ ┃ ▀██▄██▀ ⏺ Update(src/auth.ts) │ ┃ ⎿ Added 2 lines, removed 1 line │ ┃ Time 0s in tools · nothing stands o ⏺ Bash(bun test) │ ┃ Context 676 from tools · nothing stand ⎿ 3 pass, 1 fail │ ┃ │ ┃ ╭────────────────────────────────────────╮ ● Done. refresh now rejects expired claims and logs an audit event. │ ┃ │ Watching quietly. Nothing repeating… │ │ ┃ ╰────────────────────────────────────────╯ ✻ Worked for 42s · done 4:20 PM │ ┃ │ ┃ ctrl+x tab focuses this pane › /saver │ ⎿ contextsaver: ContextSaver pane shown │ │ ⟨Claude Code's own drawing⟩ ContextSaver ◌ 9 calls watched · nothing wasteful yet Close ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Band
⟨Claude Code's own drawing⟩ ContextSaver ◌ 9 calls watched · nothing wasteful yet Close
Pane · ContextSaver
▄██▀██▄ ContextSaver ↻ Check now ▄▄ ▀ 49% of context · 83k to compaction ██ ▄ ███████████████░░░░░░░░░░░░░░░ ▀██▄██▀ Time 0s in tools · nothing stands out yet Context 676 from tools · nothing stands out y… ╭────────────────────────────────────────────────────────╮ │ Watching quietly. Nothing repeating yet. │ ╰────────────────────────────────────────────────────────╯ ctrl+x tab focuses this pane
README

<h1><img src="assets/logo.png" alt="" height="40" align="top"> ContextSaver</h1>

Diagnoses wrong processes, wrong workflows and wrong habits in Claude Code sessions — and lets you fix them in one click.

Claude Code plugin tests dependencies license

Sessions bog down when the process is wrong. Review and fix rounds per lane, stacked on a whole-branch review that already covers them. The premium model on mechanical stages. Agents spawned for jobs a single command would do. The full check suite after every merge. Tool-level waste — the same file read again, the full suite after a one-line edit — compounds on top. ContextSaver catches both: the wrong process at the moments that matter, and the wrong habit as it builds. It gives you the measured cost and one click to change it.

<img src="docs/screenshot.png" width="880" alt="The ContextSaver pane beside the transcript, showing two cards with their cost and the Fix, Fix… and Ignore buttons">

Features

  • It diagnoses the process, not just the calls. "Per-lane review rounds duplicate the whole-branch review — 3 rounds × 3 lanes, ~4h so far. One whole-branch review reaches the same result." Process findings fire at moments that matter: an accepted plan, a Workflow launch, a pace complaint, every thirty minutes.
  • It names tool-level habits too. "Claude keeps running the whole suite after every one-file edit — 5×, ~9% of context, 4m each." The calls behind the claim are one keypress away.
  • Nothing to configure. No thresholds, no rules to tune. The plugin gathers the evidence and the model asks the right question; that catches things nobody wrote a rule for.
  • Fixes reach every subagent. When you fix a habit, the instruction rides the prompt of every new agent the session spawns. The card shows "sent ×N" so you can see it landed.
  • It works mid-turn. Long agentic turns are checked while they run, and your fix reaches Claude on its next tool result, then rides every prompt after it, so it survives compaction.
  • It shows where your session went. One line each for time and context, and a plain sentence on what those minutes and tokens actually bought.
  • It stays quiet. A single occurrence is never a finding. A wrong card costs more than a missed one.
  • Fixes outlive the session. Turn a decision into a CLAUDE.md rule, a skill, an agent brief or a permission rule, written only when you click Write.
  • It never touches your work. No tool denied, no output trimmed, no error hidden. If a hook throws, your session carries on as if the plugin weren't there.

Install

[!NOTE] Needs Claude Code 2.1.273 or newer. ContextSaver is a Claude Mod, built on function hooks, which are early access — so enable the flag first.

export CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1   # your shell, or ~/.claude/settings.json under "env"

Then, in Claude Code:

/plugin marketplace add AlmogBaku/ContextSaver
/plugin install contextsaver@contextsaver

That is the whole setup: no config, no API key, no dependencies, no build step. Installing mid-session works too — the plugin reads what already happened out of the transcript and checks it right away, without waiting for your next prompt.

Usage

  1. Work as usual. One line sits above your prompt and stays quiet until something repeats.
  2. A card appears in the pane beside your transcript: what Claude keeps doing, how often, what it has cost you in context and minutes, and what to do instead. Press i for the calls behind it.
  3. Choose. Fix sends the suggested fix, Fix… lets you reword it first, Ignore drops it for the session.
  4. Claude changes course in the turn already running, and the pane credits what you saved — or tells you the instruction was ignored.
  5. Keep what worked. Each decision is offered as a CLAUDE.md rule, a skill, an agent brief or a permission rule: Write saves it, Try uses it for this session only, Skip drops it.

[!TIP] The pane docks beside your transcript in a wide terminal and sits above the prompt otherwise. /saver toggles it at any width, ctrl+x tab moves the keyboard into it, Tab cycles the buttons, Enter presses, and Esc hands the keys back.

Commands

CommandWhat it does
/saverShow or hide the pane.
/saver checkCheck now, instead of waiting for the next automatic check.
/saver fix <n>Send card n's suggested fix.
/saver fix [n] <text>Send your own instruction instead. Without a number: the open card, else card 1.
/saver ignore <n>Drop card n for the rest of the session.
/saver debugPrint the session state: ledger, findings, decisions, what the audit cost, savings.
/saver resetClear this session's ledger and decisions. Learned patterns survive.

How it works

Every tool call becomes a row in a session ledger: what ran, how long it took, how much it added to context, which files it touched. The plugin runs two judges.

The process judge fires at moments that matter — an accepted plan, a Workflow launch, a pace complaint you type, every thirty minutes. It builds a digest of what the user asked, how the work was organised, what models ran which roles, how many merges got a check, what was re-read and what was slow. From that, it asks: given what the user asked, how would a lean expert run this work, and where does this session diverge? At most three findings, each with a measured cost.

The habit judge runs continuously — about every 30k tokens and three turns, or every 40 calls and five minutes inside a long turn. It asks what has repeated and bogged things down, and whether there was a shorter path. Whatever it finds becomes a card, with the costs computed from the ledger.

Fix on a process card sends a one-time re-plan to the main loop. Fix on a habit card sends a standing instruction that rides every subsequent prompt and is also appended to every new subagent's prompt — the card shows "sent ×N" to confirm it landed. Neither ever blocks a tool.

The architecture, the judge's prompt and the design brief are in docs/SPEC.md; the product spec is docs/PRD.md.

Notes and limits

[!IMPORTANT] The audit runs on your session's model, so a session on Opus pays Opus for it. It keeps itself to a few percent of the session's tokens, and /saver debug shows exactly what it spent.

  • Early access. Function hooks are new and the API underneath can still change. Every module is tested and the main flow is verified live, but expect rough edges.
  • Short sessions stay quiet. Under ~15 tool calls or 5 turns, only behaviour seen three or more times is reported.
  • Durations are wall time. They include the time a permission prompt spent waiting for you, and the model is told as much.
  • Terminal and desktop only. On mobile surfaces the plugin keeps its ledger and draws nothing.
  • The band speaks for dead turns and running workflows. After a turn that ended in an API error or a refusal it reads ✕ Last turn ended in … · type anything to continue until your next prompt; while a workflow runs and nothing is found it names the run, its stage, its agents and their calls.

Development

git clone https://github.com/AlmogBaku/ContextSaver && cd ContextSaver
./scripts/check.sh                                          # validate --strict, typecheck, tests
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir .    # run Claude Code with this folder loaded

All logic is pure functions over a single State in hooks/core/. hooks/register.ts is the only file that touches events, and every hook falls back to next(e) on any path it does not own. hooks/ui.tsx renders two view models and never reads State. With CONTEXTSAVER_DEBUG=1, /saver demo fills the pane with sample cards, so the drawing can be worked on without waiting for a real finding.

Under --plugin-dir, editing a file hot-reloads the plugin and resets session state. Learned patterns persist in the plugin store.

License

MIT © Almog Baku. See LICENSE.

Source 17 files
hooks/register.ts 900 lines
1import type { ModelForkResult, On, PaneOpenArgs, RenderElement } from 'claude-code'
2
3import { adoptRows } from './core/adopt'
4import { demoForkUsage, demoPatterns, demoRows, demoTurns, demoUsage } from './core/demo'
5import { buildPrompt, judgeAliases, merge, parseReply, shouldRun, spentOf, usageOf } from './core/judge'
6import { rowOf } from './core/ledger'
7import type { ToolEvent } from './core/ledger'
8import { bandModel, debugDump, fromStored, mergeStored, paneModel, parseRegistry, processLine, reduce, storableOf, usageLine } from './core/patterns'
9import { buildProcessPrompt, isPace, mergeProcess, parseProcessReply, processAliases, shouldProcess } from './core/process'
10import { appendedTo, bulletOnly, mergeSettings, propose } from './core/rules'
11import { activeRuns, agentOf, journalPath, parseJournal, phasesOf, runOf } from './core/spawns'
12import { collapseWs, duration, fit, instructionOf, isRecord, pctOf } from './core/text'
13import {
14  ASK_HEAD_MAX, AUTO_OPEN_MIN_COLUMNS, CLAUDE_MD_HEADING, COMMAND, DEBUG_MAX_DROPPED, DELIVERY_BURST_MS, JUDGE_MIN_ROWS, MAX_PATTERNS, PANE_ID,
15  PANE_INLINE_ROWS, PANE_TITLE, PLUGIN_NAME, RUN_REFRESH_MS, STEER_RING_TRIES, STEER_RING_WAIT_MS, initialState, storeKey,
16} from './core/types'
17import type { Action, Actions, Artifact, Choice, ProcessTrigger, Run, State, Tokens, Ui } from './core/types'
18import type { Host } from './host'
19import { Band, Pane } from './ui'
20
21const FIX_USAGE = 'Usage: /saver fix [n] [instruction] (a leading number is the card the pane draws; without one: the card whose Fix… field is open, else card 1)'
22const SAVER_USAGE = 'Usage: /saver [check | fix [n] [text] | ignore <n> | debug | reset]'
23const NOTHING_TEXT = 'ContextSaver: nothing to decide on'
24const CHECKING_TEXT = 'ContextSaver: checking this session for waste…'
25const ALREADY_TEXT = 'ContextSaver: already checking'
26const ANSWER_HEAD = 100   // characters of the turn's answer kept as an evidence quote
27const CARD_KIND = 60      // characters of a card's behaviour quoted back in a command's reply
28const PLAN_TOOL = 'ExitPlanMode'   // its accepted result is a plan the process judge can weigh
29const TYPED: readonly string[] = ['composer', 'bridge']   // the origins a person typed: only these count as asks
30const DEMO_CONTEXT = [120_000, 190_000, 250_000, 320_000]   // `/saver demo`: the window filling up to the sample's own 32%, so the trend draws
31
32// What set one judge run going, as the debug log names it: the mid-turn cadence, the turn's end, the
33// person, or a check armed at load over a transcript this plugin joined late.
34type JudgeReason = 'tool.call' | 'turn.complete' | '/saver check' | 'load'
35
36// The lanes the run answers out loud: the check the person typed, and the one the load armed for them.
37const REQUESTED: readonly JudgeReason[] = ['/saver check', 'load']
38
39/**
40 * Registers ContextSaver: the ledger of every tool call, the judge that names wasteful
41 * behaviours, the band and the pane that let the user fix or ignore them.
42 *
43 * @param on the engine's registrar
44 */
45export function register(on: On): void {
46  let state: State = initialState('', 0)
47  let host: Host | null = null
48  let isDebug = false
49  // `judge.start` lands one clock read after the decision to run, and a storm of tool calls decides
50  // inside that window: this flag is what stops a second fork of the same session.
51  let forking = false
52  // A check asked for while a run is in flight is answered by that run: cadence runs are frequent now,
53  // and the person who pressed Check now would otherwise be told `already checking` and never told more.
54  let asked = false
55  // The load lane's one failure toast. Its arming survives every failure, so a toast per retry would be a
56  // storm — but total silence reads exactly like a check that never fired, so the first failure speaks.
57  let armedSpoke = false
58  // The process judge's own lock, beside `forking`: `process.start` lands one clock read after the decision.
59  let processing = false
60  // The loops a journal was re-read for once already: an id no journal names after that is the engine's own.
61  const probed = new Set<string>()
62
63  const messageOf = (err: unknown): string => (err instanceof Error ? err.message : String(err))
64
65  // `focus` is a request, not a grant (d.ts 4915-4922): the surface hands the pane the keyboard only while
66  // the prompt holds them over an empty composer, and refuses it otherwise — the pane opens either way.
67  const paneArgs = (focus?: true): PaneOpenArgs =>
68    focus === undefined
69      ? { id: PANE_ID, title: PANE_TITLE, rows: PANE_INLINE_ROWS }
70      : { id: PANE_ID, title: PANE_TITLE, rows: PANE_INLINE_ROWS, focus }
71
72  const savedToast = (before: State['saved']): void => {
73    const ms = state.saved.ms - before.ms
74    const pct = pctOf(state.saved.chars - before.chars, state.usage.window)
75    const grew = [...(ms > 0 ? [`+${duration(ms)}`] : []), ...(pct > 0 ? [`+~${pct}%`] : [])]
76    if (grew.length === 0) return
77    host?.toast(`${grew.join(' · ')} context saved`)
78  }
79
80  const openIds = (): string[] => state.patterns.filter(p => p.openedAtTurn !== null).map(p => p.id)
81
82  // Outside a render hook: fold the action in, redraw, and credit an instruction that settled in it.
83  const dispatch = (action: Action): void => {
84    const before = state.saved
85    const open = openIds()
86    state = reduce(state, action)
87    host?.invalidate()
88    // Only a settled instruction is a saving to announce; the per-turn accrual behind it stays quiet.
89    if (open.some(id => state.patterns.find(p => p.id === id)?.openedAtTurn === null)) savedToast(before)
90  }
91
92  // A `/clear`, a `/saver reset` or a resume starts the session over: the closure flags that belong to
93  // the session go with its state, so a new one may speak for its armed check again.
94  const resetSession = (): void => {
95    dispatch({ type: 'reset' })
96    armedSpoke = false
97    probed.clear()
98  }
99
100  // Inside a render hook: fold the action in with no redraw, since a redraw loops.
101  const observe = (action: Action): void => {
102    state = reduce(state, action)
103  }
104
105  const persist = (): void => {
106    const engine = host
107    if (engine === null) return
108    const key = storeKey(state.cwd)
109    const mine = storableOf(state)
110    void engine
111      .storeGet(key)
112      .then(value => engine.storeSet(key, mergeStored(parseRegistry(value), mine)))
113      .catch(() => undefined)
114  }
115
116  // An open the person asked for asks for their keyboard too, so the pane they just called up is the pane
117  // they can type in; an unasked one interrupts whatever they were doing and never asks.
118  const openPane = async (auto?: true): Promise<void> => {
119    const engine = host
120    if (engine === null) return
121    await engine.openPane(auto === undefined ? paneArgs(true) : paneArgs())
122    dispatch(auto === undefined ? { type: 'pane', open: true } : { type: 'pane', open: true, auto })
123  }
124
125  // Only after fresh cards arrived, once a session, and only where the surface would draw it.
126  const autoOpen = async (fresh: readonly string[]): Promise<void> => {
127    const queued = fresh.some(id => state.cards.includes(id))
128    if (!queued || state.paneOpen || state.autoOpened || (state.columns ?? 0) < AUTO_OPEN_MIN_COLUMNS) return
129    try {
130      await openPane(true)
131    } catch {
132      // an unasked open the surface or another plugin refused is no error of the user's
133    }
134  }
135
136  // A check the person asked for: they are waiting for the answer, so the pane opens at any width, every time.
137  const openForCheck = async (queued: readonly string[]): Promise<void> => {
138    if (queued.length === 0 || state.paneOpen) return
139    try {
140      await openPane()
141    } catch {
142      // the surface or another plugin refused: the band still says what was found
143    }
144  }
145
146  // The four counts a turn was billed, zero where no response came back to bill.
147  const tokensOf = (u: { input_tokens: number; output_tokens: number; cache_read_input_tokens: number; cache_creation_input_tokens: number } | undefined): Tokens =>
148    ({ input: u?.input_tokens ?? 0, output: u?.output_tokens ?? 0, cacheRead: u?.cache_read_input_tokens ?? 0, cacheCreate: u?.cache_creation_input_tokens ?? 0 })
149
150  // An `Agent` result names its loop, but the description it was given is the call's own argument, so the
151  // value is handed over with it; every other tool's result is read as it came.
152  const spawnValue = (e: ToolEvent, value: unknown): unknown =>
153    e.tool === 'Agent' && typeof e.description === 'string' && typeof value === 'object' && value !== null
154      ? { ...value, description: e.description }
155      : value
156
157  // One run's journal, read and folded in; a read that fails still stamps the run, so it is not retried at once.
158  const readJournal = async (engine: Host, run: Run, now: number): Promise<void> => {
159    const path = journalPath(run)
160    if (path === null) return
161    const entries = await engine.readFile(path).then(parseJournal).catch(() => [])
162    dispatch({ type: 'run.journal', runId: run.id, entries, now })
163  }
164
165  // The journals of the runs still going, at most every RUN_REFRESH_MS each — every run's when forced: a
166  // launch reads its own at once, and the judge reads them all before it asks. Never awaited from a hook.
167  const refreshRuns = async (force: boolean): Promise<void> => {
168    const engine = host
169    if (engine === null || state.runs.length === 0) return
170    const now = await engine.now()
171    const due = force ? state.runs : activeRuns(state, now).filter(run => now - run.refreshedAt >= RUN_REFRESH_MS)
172    await Promise.all(due.map(run => readJournal(engine, run, now)))
173  }
174
175  // What the run put in front of the user: a recurrence is news too, and it is the D4 moment this exists for.
176  const checkedText = (queued: number): string =>
177    queued === 0 ? 'ContextSaver: nothing new' : `ContextSaver: ${queued} new waster${queued === 1 ? '' : 's'}`
178
179  // A run that reported nothing: a cold snapshot, a refusal, or a failure of ours.
180  const judgedNothing = (error: string): Action => ({
181    type: 'judge.done', patterns: state.patterns, fresh: [], recurred: [], focus: null, time: null, context: null, spent: 0, error,
182    returned: 0, kept: 0, dropped: [], usage: null,
183  })
184
185  // A run the person asked for answers them, whatever it found: silence is what a check must never be.
186  // An armed check nobody asked for says it once, in its own words, since it will be retried: the first
187  // failure names itself and the retries stay quiet.
188  const failedToast = (reason: JudgeReason, failure: string): void => {
189    if (reason === 'load' && !asked) {
190      if (armedSpoke) return
191      armedSpoke = true
192      host?.toast(`ContextSaver: could not check yet — ${failure}`)
193      return
194    }
195    if (REQUESTED.includes(reason) || asked) host?.toast(`ContextSaver: check failed — ${failure}`)
196  }
197
198  const judgeOnce = async (engine: Host, reason: JudgeReason): Promise<void> => {
199    const requested = REQUESTED.includes(reason)
200    const seq = state.seq
201    const now = await engine.now()
202    dispatch({ type: 'judge.start', now, seq })
203    // AGENTS is read off the journals, so they are brought up to date once, here, before the prompt is built.
204    await refreshRuns(true).catch(() => undefined)
205    // One alias table for the run: the loops keep spawning while the fork thinks, and a new agent's first
206    // row would renumber the `agent:aN` handles the reply cites against the AGENTS the prompt printed.
207    const aliases = judgeAliases(state)
208    let reply: ModelForkResult | null = null
209    let failed: string | null = null
210    try {
211      reply = await engine.fork(buildPrompt(state, aliases))
212    } catch (err) {
213      failed = messageOf(err)
214    }
215    if (reply === null) {
216      dispatch(judgedNothing(failed ?? 'cold snapshot'))
217      failedToast(reason, failed ?? 'cold snapshot')
218      return
219    }
220    try {
221      const { findings, focus, time, context, dropped, returned } = parseReply(reply.text, state, aliases)
222      const merged = merge(state, findings)
223      // A finding the registry cap evicted never becomes a card, so it is dropped, not kept.
224      const reasons = [...dropped, ...merged.evicted.map(id => `${id}: evicted, over MAX_PATTERNS (${MAX_PATTERNS})`)]
225      const kept = findings.length - merged.evicted.length
226      const usage = usageOf(reply.usage)
227      dispatch({
228        type: 'judge.done', patterns: merged.patterns, fresh: merged.fresh, recurred: merged.recurred, focus,
229        time, context, spent: spentOf(usage), error: null, returned, kept, dropped: reasons, usage,
230      })
231      try {
232        if (isDebug) {
233          engine.log(`ContextSaver judge: ${returned} returned · ${kept} kept · ${reasons.length} dropped · from ${reason}`)
234          for (const line of reasons.slice(0, DEBUG_MAX_DROPPED)) engine.log(line)
235          // A cold cache is what makes a run expensive, and only the four counts say which it was.
236          engine.log(usageLine(usage))
237        }
238      } catch {
239        // a log we could not write is not a failed run: the findings are already in the registry
240      }
241      persist()
242      // Every card this run queued: a fresh finding, or a steered behaviour that came back.
243      const queued = [...merged.fresh, ...merged.recurred].filter(id => state.cards.includes(id))
244      // A check asked for mid-run is answered by the run it arrived in, whichever run that was.
245      if (!(requested || asked)) return autoOpen(queued)
246      engine.toast(checkedText(queued.length))
247      return openForCheck(queued)
248    } catch (err) {
249      // Whatever went wrong, the run is over: `running` may never stay true.
250      dispatch(judgedNothing(messageOf(err)))
251      failedToast(reason, messageOf(err))
252    }
253  }
254
255  /**
256   * Judges the session once, if no run is already in flight.
257   *
258   * @param reason what set this run going: the mid-turn cadence, the turn's end, the person, or the
259   *   check a load armed. The two REQUESTED lanes are answered with a toast and open the pane wherever
260   *   it can be drawn; a cadence run stays quiet and keeps the once-a-session, wide-terminal rule for
261   *   opening itself — unless someone asks while it is in flight, in which case that run answers them.
262   */
263  async function runJudge(reason: JudgeReason): Promise<void> {
264    const engine = host
265    if (engine === null || forking || state.judge.running) return
266    forking = true
267    try {
268      await judgeOnce(engine, reason)
269    } finally {
270      forking = false
271      asked = false
272    }
273  }
274
275  // A check armed at load consults no gate — `shouldRun` counts new work, and a session joined late has
276  // all of its work behind it. The load fires this itself; every opportunity after it is a retry of a run
277  // that came back with nothing. Answers whether it took this opportunity.
278  const armedCheck = (): boolean => {
279    if (!state.pendingCheck || state.judge.running) return false
280    void runJudge('load').catch(() => undefined)
281    return true
282  }
283
284  // The opportunities a run can start at: the armed check first, else the cadence's own count of new work.
285  const judgeAt = (now: number, cadence: JudgeReason): void => {
286    if (!armedCheck() && shouldRun(state, now)) void runJudge(cadence).catch(() => undefined)
287  }
288
289  // A process run that reported nothing: the failure is said, since a process judge that went quiet looks fine.
290  const processFailed = (failure: string): void => {
291    dispatch({ type: 'process.done', patterns: state.patterns, fresh: [], recurred: [], spent: 0, error: failure, returned: 0, kept: 0, dropped: [], usage: null })
292    host?.toast(`ContextSaver: process check failed — ${failure}`)
293  }
294
295  const processOnce = async (engine: Host, trigger: ProcessTrigger): Promise<void> => {
296    const now = await engine.now()
297    const { lastAtMs, lastAtSeq, lastAtTurn } = state.process
298    dispatch({ type: 'process.start', now, seq: state.seq, trigger })
299    // One alias table for the run, as the habit judge keeps: the reply cites the handles this digest printed.
300    const aliases = processAliases(state)
301    let reply: ModelForkResult | null = null
302    let failed: string | null = null
303    try {
304      reply = await engine.fork(buildProcessPrompt(state, aliases))
305    } catch (err) {
306      failed = messageOf(err)
307    }
308    // Every trigger is automatic, none asked for: a cold snapshot is no run and no failure, and the next trigger tries again.
309    if (reply === null) return failed === null ? dispatch({ type: 'process.cold', was: { lastAtMs, lastAtSeq, lastAtTurn } }) : processFailed(failed)
310    try {
311      const { findings, dropped, returned } = parseProcessReply(reply.text, state, aliases)
312      const merged = mergeProcess(state, findings)
313      const reasons = [...dropped, ...merged.evicted.map(id => `${id}: evicted, over MAX_PATTERNS (${MAX_PATTERNS})`)]
314      const kept = findings.length - merged.evicted.length
315      const usage = usageOf(reply.usage)
316      dispatch({
317        type: 'process.done', patterns: merged.patterns, fresh: merged.fresh, recurred: merged.recurred,
318        spent: spentOf(usage), error: null, returned, kept, dropped: reasons, usage,
319      })
320      if (isDebug) engine.log(`ContextSaver process: ${returned} returned · ${kept} kept · ${reasons.length} dropped · from ${trigger}`)
321      persist()
322      return autoOpen(merged.fresh.filter(id => state.cards.includes(id)))
323    } catch (err) {
324      // Whatever went wrong, the run is over: `running` may never stay true.
325      processFailed(messageOf(err))
326    }
327  }
328
329  /**
330   * Starts one detached process run when its gate lets it, answering whether it did.
331   *
332   * @param now the clock the gate reads
333   * @param trigger what asked: a plan, a launch, a pace complaint or the clock
334   * @param before what the run waits for before the digest is built (a launch's declared phases)
335   */
336  const processAt = (now: number, trigger: ProcessTrigger, before?: () => Promise<void>): boolean => {
337    const engine = host
338    if (engine === null || processing || !shouldProcess(state, now, trigger)) return false
339    processing = true
340    void (async () => {
341      try {
342        await before?.()
343        await processOnce(engine, trigger)
344      } finally {
345        processing = false
346      }
347    })().catch(() => undefined)
348    return true
349  }
350
351  // The phases a launch's script declares, read as a literal and never run; a script we cannot read declares none.
352  const readPhases = async (engine: Host, runId: string, value: unknown): Promise<void> => {
353    const path = isRecord(value) && typeof value['scriptPath'] === 'string' ? value['scriptPath'] : null
354    if (path === null) return
355    const script = await engine.readFile(path).catch(() => null)
356    if (script !== null) dispatch({ type: 'run.phases', runId, phases: phasesOf(script) })
357  }
358
359  // A burst of deliveries is said once: a workflow's agents start in a crowd, and a toast each is a storm.
360  const deliveredToast = (now: number): void => {
361    if (now - state.delivery.lastToastAt < DELIVERY_BURST_MS) return
362    host?.toast('ContextSaver: your fixes reached a new subagent')
363    dispatch({ type: 'delivery.toasted', now })
364  }
365
366  // A loop a journal or an `Agent` result named: the engine's own forks (compaction, memory) are named by neither.
367  const isNamed = (agentId: string): boolean => state.loops.some(l => l.id === agentId && (l.run !== null || l.label !== null))
368
369  // The standing fixes an agent the spawn rewrite missed is owed on its first call; a spawn in flight counts as reached.
370  const owedTo = async (engine: Host, agentId: string, now: number): Promise<string[]> => {
371    if (state.standing.length === 0 || state.delivery.pending > 0 || state.delivery.agents.includes(agentId)) return []
372    // A workflow agent calls before the journal that names it was re-read: it is read once more for that id.
373    if (!isNamed(agentId) && !probed.has(agentId) && activeRuns(state, now).length > 0) {
374      probed.add(agentId)
375      await refreshRuns(true).catch(() => undefined)
376    }
377    if (!isNamed(agentId) || state.delivery.agents.includes(agentId)) return []
378    const standing = [...state.standing]
379    dispatch({ type: 'delivery.sent', agentId, now })
380    if (isDebug) engine.log(`ContextSaver delivery: first call → ${agentId}`)
381    deliveredToast(now)
382    return standing
383  }
384
385  // What a typed prompt is kept as for the digest: one line, no control characters, at most ASK_HEAD_MAX.
386  const headOf = (text: string): string => collapseWs(text.replace(/\p{Cc}/gu, ' ')).slice(0, ASK_HEAD_MAX)
387
388  const checkNow = (): string => {
389    if (forking || state.judge.running) {
390      // The run already going answers this ask: nothing is forked, and nobody is left without a reply.
391      asked = true
392      return ALREADY_TEXT
393    }
394    void runJudge('/saver check').catch(() => undefined)
395    return CHECKING_TEXT
396  }
397
398  const togglePane = async (): Promise<void> => {
399    const engine = host
400    if (engine === null) return
401    try {
402      // The `ui.close` hook records the close, as it records the person's own.
403      if (state.paneOpen) await engine.closePane({ id: PANE_ID })
404      else await openPane()
405    } catch {
406      // the surface or another plugin refused: the pane stays as it was, and `/saver` says so
407    }
408  }
409
410  const firstLine = (text: string): string => {
411    const [head = ''] = text.split('\n')
412    return head === text ? text : `${head} …`
413  }
414
415  const decide = (patternId: string, choice: Choice, text?: string): void => {
416    const p = state.patterns.find(q => q.id === patternId)
417    if (p === undefined) return
418    dispatch({ type: 'decide', patternId, choice, text })
419    if (state.patterns.find(q => q.id === patternId)?.decision !== choice) return
420    if (choice === 'keep') host?.toast(`ContextSaver: ignored "${p.kind}"`)
421    if (choice === 'kill') host?.toast(`ContextSaver: fixed — ${p.alternative}`)
422    if (choice === 'steer') host?.toast(`ContextSaver: fixed with your note — ${firstLine(text ?? '')}`)
423    persist()
424  }
425
426  // The number the pane draws beside a card is its seat in `cards`; 0 means the card is no longer listed.
427  const seatOf = (patternId: string): number => state.cards.indexOf(patternId) + 1
428
429  const cardReply = (patternId: string, seat: number, tail: string): string => {
430    const kind = state.patterns.find(q => q.id === patternId)?.kind ?? patternId
431    return `ContextSaver: card ${seat} — "${fit(kind, CARD_KIND)}" · ${tail}`
432  }
433
434  const numberOf = (token: string): number | null => (/^\d+$/.test(token) ? Number(token) : null)
435
436  // A number no card wears is a numbering mistake, whichever verb typed it: it is refused, never obeyed.
437  const noCardText = (n: number): string => `ContextSaver: no card ${n} (1–${state.cards.length})`
438
439  // `/saver ignore 2` and `/saver fix 2` decide the card the pane numbers 2, and say which one they took.
440  const decideByNumber = (choice: Choice, token: string): string => {
441    if (state.cards.length === 0) return NOTHING_TEXT
442    const n = numberOf(token)
443    if (n === null) return SAVER_USAGE
444    const patternId = state.cards[n - 1]
445    if (patternId === undefined) return noCardText(n)
446    const reply = cardReply(patternId, n, choice === 'keep' ? 'ignored' : 'fixed')
447    decide(patternId, choice)
448    return reply
449  }
450
451  const steerSubmit = (patternId: string, text: string): void => {
452    const wanted = text.trim()
453    if (wanted === '') {
454      host?.toast('ContextSaver: write the instruction first')
455      return
456    }
457    decide(patternId, 'steer', wanted)
458  }
459
460  // The artifact's own words without the file's furniture: a bullet, or the prose under a frontmatter.
461  const bodyOf = (content: string): string =>
462    content.includes(CLAUDE_MD_HEADING) ? bulletOnly(content).replace(/^- /, '') : (content.split('---\n').at(-1) ?? content)
463
464  const writeArtifact = async (a: Artifact): Promise<void> => {
465    const engine = host
466    if (engine === null) return
467    try {
468      // A whole-file write reads nothing; an append and a settings merge need what is there.
469      const existing = a.mode === 'write' || !(await engine.exists(a.path)) ? null : await engine.readFile(a.path)
470      if (a.mode === 'append') await engine.writeFile(a.path, appendedTo(existing, a.content))
471      if (a.mode === 'write') await engine.writeFile(a.path, a.content)
472      if (a.mode === 'merge-settings') await engine.writeFile(a.path, mergeSettings(existing, a.content))
473      dispatch({ type: 'artifact.done', patternId: a.patternId, kind: a.kind, written: true })
474      engine.toast(`Wrote ${a.path}`)
475    } catch (err) {
476      engine.toast(messageOf(err))
477    }
478  }
479
480  const tryArtifact = (a: Artifact): void => {
481    dispatch({ type: 'standing.add', text: instructionOf(collapseWs(bodyOf(a.content)) || a.title) })
482    dispatch({ type: 'artifact.done', patternId: a.patternId, kind: a.kind, written: true })
483    host?.toast(`Trying "${a.title}" for this session`)
484  }
485
486  // The key the pane draws a card's Fix… field under.
487  const steerFieldKey = (patternId: string): string => `card:${patternId}:text`
488
489  // The ring lands only on an element the drawn tree already holds ('no element of its own is drawn under that
490  // key', d.ts 8846-8853), and a tree lands after the render hook that built it returns. The press that opens
491  // the field asks for a redraw and nothing more, so the ask waits for the frame the field is drawn in, and
492  // asks again while the engine answers that nothing is drawn under the key — a few frames, then it stops.
493  const steerRing = async (patternId: string): Promise<void> => {
494    const engine = host
495    if (engine === null) return
496    // `focus` is a request, not a grant (d.ts 4915-4922): the surface refuses it while the person holds an
497    // element of ours, which the press that opened the field is — but where the composer holds the keys over
498    // an empty line it is granted, and then `autoFocus` lands the ring on the field by itself.
499    void engine.openPane(paneArgs(true)).catch(() => undefined)
500    let denied = 'the ask was never answered'
501    for (let tries = STEER_RING_TRIES; tries > 0; tries -= 1) {
502      await engine.sleep(STEER_RING_WAIT_MS).catch(() => undefined)
503      // The field was closed again, or another card's opened: this ring is nobody's now.
504      if (state.steering !== patternId) return
505      const deny = await engine
506        .focusElement({ requestId: PANE_ID, key: steerFieldKey(patternId) })
507        .then(result => result.deny ?? null)
508        .catch(err => messageOf(err))
509      if (deny === null) return
510      denied = deny
511    }
512    // The ring stayed put, so the keystrokes are the composer's: the way in is the line the person is owed.
513    engine.toast(`ContextSaver: the composer has your keys — type /saver fix ${seatOf(patternId)} <your note>`)
514    if (isDebug) engine.log(`${PLUGIN_NAME}: the ring never reached ${steerFieldKey(patternId)} — ${denied}`)
515  }
516
517  const actions: Actions = {
518    keep: patternId => decide(patternId, 'keep'),
519    steer: patternId => {
520      dispatch({ type: 'steer.begin', patternId })
521      // A second press closed the field; only the press that opened one goes looking for the keyboard.
522      if (state.steering === patternId) void steerRing(patternId)
523    },
524    // The pane's body is this hook's tree, so the redraw is what paints the keystroke; the text it draws
525    // back is this one, which is also what `/saver fix` sends when the keyboard never reaches the field.
526    steerDraft: text => dispatch({ type: 'steer.draft', text }),
527    steerSubmit: (patternId, text) => steerSubmit(patternId, text),
528    kill: patternId => decide(patternId, 'kill'),
529    info: patternId => dispatch({ type: 'expand', patternId }),
530    togglePane: () => {
531      void togglePane()
532    },
533    // The band and the pane have no reply to write in, so the press is answered with a toast.
534    check: () => {
535      host?.toast(checkNow())
536    },
537    write: a => {
538      void writeArtifact(a)
539    },
540    tryOnce: a => tryArtifact(a),
541    // A skipped rule is handled: `state.written` is the set the pane never offers again.
542    skip: a => dispatch({ type: 'artifact.done', patternId: a.patternId, kind: a.kind, written: true }),
543  }
544
545  on('session.start', async ($, e, next) => {
546    try {
547      const engine: Host = {
548        now: () => $.clock.now(),
549        sleep: ms => $.clock.sleep(ms),
550        invalidate: () => $.ui.invalidate('ui.render'),
551        toast: text => $.ui.toast(text),
552        log: text => $.ui.log(text),
553        openPane: args => $.ui.open(args),
554        closePane: args => $.ui.close(args),
555        focusElement: args => $.ui.focus(args),
556        registerCommand: spec => $.command.register(spec),
557        usage: args => $.session.usage(args),
558        messages: () => $.session.messages(),
559        storeGet: key => $.store.get(key),
560        storeSet: (key, value) => $.store.set(key, value),
561        fork: prompt => $.model.fork({ prompt }),
562        readFile: path => $.fs.read(path),
563        writeFile: (path, text) => $.fs.write(path, text),
564        exists: path => $.fs.exists(path),
565        debugFlag: () => $.env.get('CONTEXTSAVER_DEBUG'),
566      }
567      host = engine
568      const u = await engine.usage({ breakdown: 'summary' })
569      const now = await engine.now()
570      const stored = parseRegistry(await engine.storeGet(storeKey(e.cwd)))
571      state = { ...initialState(e.cwd, u.context.window), patterns: stored.map(fromStored) }
572      dispatch({
573        type: 'usage',
574        usage: { window: u.context.window, compactAt: u.context.breakdown?.autoCompactThreshold, tokens: u.context.tokens, percent: u.context.percent },
575        now,
576      })
577      const tokensOf = (items: readonly { tokens: number }[] | undefined): number => (items ?? []).reduce((n, i) => n + i.tokens, 0)
578      const breakdown = u.context.breakdown
579      dispatch({
580        type: 'overhead',
581        overhead: { memory: tokensOf(breakdown?.memoryFiles), mcp: tokensOf(breakdown?.mcpTools), agents: tokensOf(breakdown?.agents) },
582      })
583      try {
584        await engine.registerCommand(COMMAND)
585      } catch (err) {
586        engine.log(`${PLUGIN_NAME}: /${COMMAND.name} is taken — ${messageOf(err)}`)
587      }
588      try {
589        const flag = await engine.debugFlag()
590        isDebug = flag !== undefined && flag !== '' && flag !== '0'
591      } catch {
592        isDebug = false
593      }
594      // Last, so a transcript we cannot read costs the session nothing it already has.
595      const adopted = adoptRows(await engine.messages())
596      if (adopted.length === 0) return next(e)
597      dispatch({ type: 'adopt', rows: adopted })
598      // Too few rows to judge: a fresh session, nothing armed and nothing forked — which the log says too,
599      // since a debug line that claims a check on three rows is worse than no line at all.
600      const enough = state.rows.length >= JUDGE_MIN_ROWS
601      if (isDebug) {
602        const tail = enough ? 'checking them now' : `under the ${JUDGE_MIN_ROWS}-row floor, nothing to check`
603        engine.log(`ContextSaver adopted ${adopted.length} rows from the transcript · ${tail}`)
604      }
605      if (!enough) return next(e)
606      dispatch({ type: 'check.arm' })
607      if (isDebug) engine.log(`ContextSaver fired a check over ${state.rows.length} adopted rows · armed, so a cold answer retries`)
608      // Fired here, not left for the person's next keystroke: a session with this much history behind it has
609      // run turns, and `$.model.fork` reads its last turn's cache-safe snapshot (d.ts 2019-2034) — which is
610      // exactly what a `/reload-plugins`, an edit under `--plugin-dir` or a resume hands us (d.ts 3106-3111).
611      // Detached, never awaited: this hook is awaited by the engine and has a budget, and the run speaks in a
612      // toast, not in a return value. A snapshot that really is cold answers null, and the arming stays up
613      // (§5.2) for the first warm opportunity below — the next prompt, the next tool call, or the turn's end.
614      armedCheck()
615      return next(e)
616    } catch {
617      return next(e)
618    }
619  })
620
621  on('turn.start', async ($, e, next) => {
622    try {
623      // The clock dates the wait since the last answer, so TURNS can say the person was away.
624      const now = host === null ? 0 : await host.now()
625      dispatch({ type: 'turn.start', now })
626      return next(e)
627    } catch {
628      return next(e)
629    }
630  })
631
632  on('tool.call', async ($, e, next) => {
633    const engine = host
634    let started = 0
635    try {
636      if (engine === null || next.origin.plugin === PLUGIN_NAME) return next(e)
637      started = await engine.now()
638    } catch {
639      return next(e)
640    }
641    // `next(e)` is called exactly once: a rejection is the engine's to report — calling it again would run the tool twice.
642    const result = await next(e)
643    try {
644      const ended = await engine.now()
645      dispatch({ type: 'row', row: rowOf(e, result, ended - started, state.turn) })
646      const row = state.rows[state.rows.length - 1]
647      if (isDebug && row !== undefined) engine.log(`ContextSaver row r${row.seq} ${row.tool} ${row.key} ${row.ms}ms ${row.chars}ch`)
648      // A `Workflow` result is a run launched, an `Agent` result a loop named; a launch reads its journal at
649      // once, and any run still going is re-read on the plugin's own cadence — detached, the call is answered.
650      const run = runOf(result.result)
651      if (run !== null) dispatch({ type: 'run.start', run, now: ended })
652      const agent = agentOf(spawnValue(e, result.result))
653      if (agent !== null) dispatch({ type: 'agent.start', ...agent })
654      void refreshRuns(run !== null).catch(() => undefined)
655      // One agentic turn can run for hours, so the cadence is judged here too, not only between turns.
656      judgeAt(ended, 'tool.call')
657      // A launch waits for the phases its script declares; a plan counts once the person accepted it.
658      const phases = (): Promise<void> => (run === null ? Promise.resolve() : readPhases(engine, run.id, result.result))
659      const planned = (e.tool as string) === PLAN_TOOL && result.deny === undefined && result.isError !== true
660      const launched = run !== null && processAt(ended, 'workflow', phases)
661      if (run !== null && !launched) void phases().catch(() => undefined)
662      if (!launched && !(planned && processAt(ended, 'plan'))) processAt(ended, 'clock')
663      if (result.deny !== undefined) return result
664      // A one-time note is the main loop's: the orchestrator is who re-plans, and a subagent is a random reader.
665      const main = e.agentId === undefined
666      const pending = main ? state.notes : await owedTo(engine, e.agentId ?? '', ended)
667      if (pending.length === 0) return result
668      if (main) dispatch({ type: 'notes.drained' })
669      return { ...result, context: [...(result.context ?? []), ...pending] }
670    } catch {
671      return result
672    }
673  })
674
675  on('turn.complete', async ($, e, next) => {
676    try {
677      const engine = host
678      if (engine === null) return next(e)
679      const u = e.usage
680      // A subagent's turn is its loop's: what it cost and how it ended go onto the loop, and nothing else
681      // moves — the window is the main loop's to sample, and so is the cadence.
682      if (e.agentId !== undefined) {
683        dispatch({ type: 'loop.turn', agentId: e.agentId, model: u?.model ?? null, ms: e.durationMs, tokens: tokensOf(u), ended: e.reason, turn: state.turn })
684        return next(e)
685      }
686      // The window is sampled before the turn is recorded: how full it is after this turn is the turn's
687      // own figure, and its growth over the last turns is the pace compaction actually runs at. The
688      // sample is optional, though: a refused `session.usage` costs this turn its context reading, never
689      // the turn itself — without the stat the trend, the pace and every token gate go with it.
690      const seen = await engine.usage().catch(() => null)
691      const now = await engine.now()
692      dispatch({
693        type: 'turn.complete',
694        stat: {
695          ...tokensOf(u),
696          ms: e.durationMs,
697          answerChars: e.answer.length,
698          answerHead: e.answer.slice(0, ANSWER_HEAD),
699          aborted: e.isAborted,
700          ended: e.reason,
701          at: now,
702          idleMs: 0,
703          context: seen?.context.tokens ?? null,
704        },
705      })
706      if (seen !== null) dispatch({ type: 'usage', usage: { window: seen.context.window, tokens: seen.context.tokens, percent: seen.context.percent }, now })
707      judgeAt(now, 'turn.complete')
708      processAt(now, 'clock')
709      return next(e)
710    } catch {
711      return next(e)
712    }
713  })
714
715  // The standing habit fixes ride a new subagent's prompt; the engine's own forks inherit the parent's context.
716  on('agent.spawn', async ($, e, next) => {
717    const standing = [...state.standing]
718    if (host === null || e.fork || standing.length === 0) return next(e)
719    dispatch({ type: 'delivery.pending', delta: 1 })
720    try {
721      const started = await next({ ...e, prompt: `${e.prompt}\n\n${standing.join('\n')}` })
722      try {
723        const now = await host.now()
724        if (started.agentId === undefined) return started
725        dispatch({ type: 'delivery.sent', agentId: started.agentId, now })
726        if (isDebug) host.log(`ContextSaver delivery: spawn rewrite → ${started.agentId}`)
727        deliveredToast(now)
728      } catch {
729        // the subagent started with the fixes either way
730      }
731      return started
732    } finally {
733      dispatch({ type: 'delivery.pending', delta: -1 })
734    }
735  })
736
737  on('session.compact', async ($, e, next) => {
738    const result = await next(e)
739    try {
740      dispatch({ type: 'compact' })
741    } catch {
742      // a compaction we failed to record is still the compaction the engine performed
743    }
744    return result
745  })
746
747  on('prompt.submit', async ($, e, next) => {
748    try {
749      // `origin` is the engine's to stamp; read it defensively so an unstamped submission still carries the texts.
750      if (e.origin?.kind === 'plugin' || e.text.trimStart().startsWith(`/${COMMAND.name}`)) return next(e)
751      // Only what the person typed is an ask: a task notification in the same words is nobody complaining.
752      if (host !== null && TYPED.includes(e.origin?.kind ?? '')) {
753        const now = await host.now()
754        const pace = isPace(e.text)
755        dispatch({ type: 'ask', ask: { turn: e.turnId !== undefined ? state.turn : state.turn + 1, at: now, head: headOf(e.text), pace } })
756        if (pace) processAt(now, 'pace')
757      }
758      const extra = [...state.notes, ...state.standing.filter(text => !state.notes.includes(text))]
759      if (extra.length > 0) dispatch({ type: 'notes.drained' })
760      const carried = extra.length === 0 ? e : { ...e, context: [...(e.context ?? []), ...extra] }
761      // A prompt is no new work, so there is no cadence lane here: only a check still armed fires — the load
762      // fired its own, so this is the retry of one that came back cold — and it never changes what the
763      // prompt carries.
764      armedCheck()
765      return next(carried)
766    } catch {
767      return next(e)
768    }
769  })
770
771  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
772    const engine = host
773    try {
774      if (engine === null || e.props.hasSurvey || e.surface === 'mobile') return next(e)
775      if (e.props.bodyColumns !== state.columns) observe({ type: 'columns', columns: e.props.bodyColumns })
776    } catch {
777      return next(e)
778    }
779    // Drawn once: a band we cannot build answers with what is beneath it, never with a second dispatch.
780    const below: RenderElement = await next(e)
781    // The clock, read once per draw: the newest reading the state holds is a turn's end or a run's launch, so
782    // a loopless run would read as fresh for the rest of a quiet session. A host that will not say the time
783    // leaves the band that newest reading of its own rather than undrawn.
784    const now = await engine.now().catch(() => null)
785    try {
786      const { Box, Text, Button, Input, Raster } = $.ui.resolve(e) as unknown as Ui
787      const band = Band({
788        ui: { Box, Text, Button, Input, Raster },
789        model: now === null ? bandModel(state) : bandModel(state, now),
790        site: { bodyColumns: e.props.bodyColumns, maxRows: e.props.maxRows },
791        actions,
792      })
793      return Box({ flexDirection: 'column', children: [below, band] })
794    } catch {
795      return below
796    }
797  })
798
799  on('ui.render', { component: 'Pane', requestId: PANE_ID }, ($, e, next) => {
800    try {
801      if (host === null || e.surface === 'mobile') return next(e)
802      const { Box, Text, Button, Input, Raster } = $.ui.resolve(e) as unknown as Ui
803      return Pane({
804        ui: { Box, Text, Button, Input, Raster },
805        model: paneModel(state, propose(state)),
806        site: { bodyColumns: e.props.bodyColumns, maxRows: e.props.scroll.bodyRows },
807        placement: e.props.placement,
808        actions,
809      })
810    } catch {
811      return next(e)
812    }
813  })
814
815  on('command.run', { command: COMMAND.name }, async ($, e, next) => {
816    try {
817      if (host === null) return next(e)
818      const args = e.args.trim()
819      const [sub = ''] = args.split(/\s+/)
820      if (sub === '' || sub === 'rules') {
821        await togglePane()
822        const process = processLine(state)
823        const shown = state.paneOpen ? 'ContextSaver pane shown' : 'ContextSaver pane hidden'
824        return { text: process === null ? shown : `${shown} · ${process}` }
825      }
826      if (sub === 'check') return { text: checkNow() }
827      if (sub === 'fix') {
828        const rest = args.slice(sub.length).trim()   // newlines inside the instruction survive
829        const [first = ''] = rest.split(/\s+/)
830        // A leading number is always the card: folding a mistyped one back into the instruction would
831        // fix the wrong card with a garbled sentence, and `standing` keeps it for the whole session.
832        const n = numberOf(first)
833        if (state.cards.length === 0) return { text: NOTHING_TEXT }
834        if (n !== null && (n < 1 || n > state.cards.length)) return { text: noCardText(n) }
835        const text = (n === null ? rest : rest.slice(first.length)).trim()
836        // Nothing after the number sends the fix the card already offers; a note sends the note instead.
837        if (text === '') return { text: n === null ? FIX_USAGE : decideByNumber('kill', first) }
838        const patternId = n === null ? (state.steering ?? state.cards[0]) : state.cards[n - 1]
839        const seat = patternId === undefined ? 0 : seatOf(patternId)
840        if (patternId === undefined || seat === 0) return { text: FIX_USAGE }
841        steerSubmit(patternId, text)
842        return { text: cardReply(patternId, seat, `fixed with your note: ${text}`) }
843      }
844      if (sub === 'ignore') return { text: decideByNumber('keep', args.slice(sub.length).trim()) }
845      if (sub === 'demo' && isDebug) {
846        // Debug-only: the pane's own look, without waiting for a real finding. The header is part of that
847        // look, so a usage sample and its turns come first — without them the hero row draws its empty
848        // state above cards that state a percentage of context. (`now` is the reducer's to ignore.)
849        dispatch({ type: 'usage', usage: demoUsage(), now: 0 })
850        // Four turns, each with the context it left behind, so the header's trend and its run to
851        // compaction draw from a window filling up rather than from what a turn was billed.
852        const samples = demoTurns()
853        for (const [at, context] of DEMO_CONTEXT.entries()) {
854          const stat = samples[at % samples.length]
855          if (stat !== undefined) dispatch({ type: 'turn.complete', stat: { ...stat, context } })
856        }
857        for (const row of demoRows(state.turn)) dispatch({ type: 'row', row })
858        const patterns = demoPatterns(state.turn)
859        const fresh = patterns.filter(p => p.decision === null).map(p => p.id)
860        const usage = demoForkUsage()
861        dispatch({
862          type: 'judge.done', patterns, fresh, recurred: [], focus: null, time: null, context: null,
863          spent: spentOf(usage), error: null, returned: patterns.length, kept: patterns.length, dropped: [], usage,
864        })
865        await openPane()
866        return { text: 'ContextSaver: demo wasters loaded' }
867      }
868      if (sub === 'debug') return { text: debugDump(state, armedSpoke) }
869      if (sub === 'reset') {
870        resetSession()
871        return { text: 'ContextSaver: session state reset' }
872      }
873      return { text: SAVER_USAGE }
874    } catch {
875      return next(e)
876    }
877  })
878
879  on('command.run', { command: ['clear', 'resume'] }, async ($, e, next) => {
880    const result = await next(e)
881    try {
882      resetSession()
883    } catch {
884      // the command ran either way: it is not ours to run a second time
885    }
886    return result
887  })
888
889  on('ui.close', { id: PANE_ID }, async ($, e, next) => {
890    const result = await next(e)
891    try {
892      // A close another plugin refused leaves the pane open, so the band still reads `Close`.
893      if (result.deny === undefined) dispatch({ type: 'pane', open: false })
894    } catch {
895      // the pane stays as the surface left it
896    }
897    return result
898  })
899}
900
hooks/core/adopt.ts 35 lines
1import type { SessionMessage, ToolCallResult, ToolUseSummary } from 'claude-code'
2
3import { rowOf } from './ledger'
4import { RECOVERED_FLAG, ROW_CAP } from './types'
5import type { Row } from './types'
6
7// Only a request starts a turn: a user message with no typed text is the tool loop or a message the engine queued itself.
8const isPrompt = (m: SessionMessage): boolean =>
9  m.role === 'user' && m.text.trim() !== '' && (m.toolResults ?? []).length === 0
10
11// The transcript kept what the call returned; a refusal was stored as its error text, never as a `deny`.
12const resultOf = (u: ToolUseSummary): ToolCallResult =>
13  u.isError === true ? { isError: true, result: u.result, text: u.text ?? '' } : { result: u.result, text: u.text }
14
15// Appended last, so a flag the call itself earned (`err`, `dedup`) still reads first.
16const flagged = (row: Omit<Row, 'seq'>): Omit<Row, 'seq'> => ({ ...row, flags: [...row.flags, RECOVERED_FLAG] })
17
18// The transcript records no duration and no agent, so `ms` is 0 and `agent` reads `main`.
19const rowFrom = (u: ToolUseSummary, turn: number): Omit<Row, 'seq'> =>
20  flagged(rowOf({ ...u.input, tool: u.tool, tool_use_id: u.tool_use_id }, resultOf(u), 0, turn))
21
22/** Rebuilds the newest ROW_CAP ledger rows from the transcript of a session this plugin joined late. */
23export const adoptRows = (messages: readonly SessionMessage[]): readonly Omit<Row, 'seq'>[] => {
24  const rows: Omit<Row, 'seq'>[] = []
25  let prompts = 0
26  for (const m of messages) {
27    if (isPrompt(m)) prompts += 1
28    if (m.role !== 'assistant') continue
29    // A call with neither a result nor a text is still in flight: it has no size and no outcome to record.
30    const settled = m.toolUses.filter(u => u.result !== undefined || u.text !== undefined)
31    for (const u of settled) rows.push(rowFrom(u, Math.max(1, prompts)))
32  }
33  return rows.slice(-ROW_CAP)
34}
35
hooks/core/demo.ts 131 lines
1import { MAIN_AGENT } from './types'
2import type { JudgeUsage, Pattern, Row, TurnStat, Usage } from './types'
3
4// A sample session, three turns wide: what the demo's rows and patterns are dated against. A session
5// younger than the sample is dated from turn 8, so the cards read 'turns 5–8' rather than 'turn 1'.
6const DEMO_TURN = 8
7const at = (turn: number, back: number): number => Math.max(1, Math.max(turn, DEMO_TURN) - back)
8
9const SUITE_KEY = 'test:bun test'
10const LOG_KEY = 'read:cat logs/api.log'
11const AGENT_KEY = 'agent:explore'
12const EXPLORE_AGENT = 'sub-explore-1'   // one sample read ran inside the explore loop, so the alias column draws
13
14const sample = (
15  id: string,
16  tool: string,
17  key: string,
18  cls: Row['cls'],
19  turn: number,
20  ms: number,
21  chars: number,
22  head: string,
23  spawn: Row['spawn'] = null,
24): Omit<Row, 'seq'> => ({ id, tool, key, cls, agent: MAIN_AGENT, turn, ms, chars, head, flags: [], lines: null, paths: [], spawn })
25
26// The same call, made inside the explore agent rather than the main loop: the pane names that loop `a1`.
27const inExplore = (row: Omit<Row, 'seq'>): Omit<Row, 'seq'> => ({ ...row, agent: EXPLORE_AGENT })
28
29const turnSample = (input: number, output: number, ms: number, answer: string): Omit<TurnStat, 'turn' | 'calls'> => ({
30  input, output, cacheRead: 180_000, cacheCreate: 7_000, ms, answerChars: answer.length, answerHead: answer, aborted: false, ended: 'answer', at: 0, idleMs: 0, context: null,
31})
32
33/** The usage the demo's header draws from: a third of a million-token window spent, with a compaction threshold. */
34export const demoUsage = (): Usage => ({ window: 1_000_000, compactAt: 900_000, tokens: 320_000, percent: 32 })
35
36/** What the demo's one judge run cost: a cold fork, whose cache creation is most of the bill. */
37export const demoForkUsage = (): JudgeUsage => ({ input: 96_000, output: 1_200, cacheRead: 0, cacheCreate: 221_000 })
38
39/** Three sample turns, so the header's run to compaction has a pace to state it in turns. */
40export const demoTurns = (): Omit<TurnStat, 'turn' | 'calls'>[] => [
41  turnSample(42_000, 3_000, 96_000, 'Ran the suite: 212 pass. Reading the api log for the 500 next.'),
42  turnSample(41_000, 4_000, 132_000, 'The refresh flow lives in src/auth/refresh.ts; the token TTL is the bug.'),
43  turnSample(43_000, 2_000, 88_000, 'Suite is green again after the token fix.'),
44]
45
46/** The ledger rows the demo's wasters cite, so their cost and their evidence quotes are real rows. */
47export const demoRows = (turn: number): Omit<Row, 'seq'>[] => [
48  sample('demo-suite-1', 'Bash', SUITE_KEY, 'test', at(turn, 3), 62_000, 24_000, '212 pass · 0 fail · ran 1284 expect() calls in 61.98s'),
49  sample('demo-log-1', 'Bash', LOG_KEY, 'read', at(turn, 3), 4_000, 80_000, 'GET /health 200 12ms · GET /v1/users 200 41ms · POST /v1/tokens 500 88ms'),
50  sample('demo-agent-1', 'Agent', AGENT_KEY, 'other', at(turn, 2), 90_000, 36_000, 'Explored src/auth: 6 files read, the refresh flow lives in src/auth/refresh.ts', {
51    type: 'explore', requested: null, resolved: 'claude-sonnet-4-6', status: 'completed', tokens: 9_000, edits: 0, promptChars: 420,
52  }),
53  sample('demo-suite-2', 'Bash', SUITE_KEY, 'test', at(turn, 2), 59_000, 24_000, '212 pass · 0 fail · ran 1284 expect() calls in 58.71s'),
54  inExplore(sample('demo-log-2', 'Bash', LOG_KEY, 'read', at(turn, 1), 11_000, 80_000, 'GET /health 200 11ms · GET /v1/users 200 39ms · POST /v1/tokens 500 91ms')),
55  sample('demo-agent-2', 'Agent', AGENT_KEY, 'other', at(turn, 1), 90_000, 36_000, 'Explored src/auth: 6 files read, the refresh flow lives in src/auth/refresh.ts', {
56    type: 'explore', requested: null, resolved: 'claude-sonnet-4-6', status: 'completed', tokens: 9_000, edits: 0, promptChars: 430,
57  }),
58  sample('demo-suite-3', 'Bash', SUITE_KEY, 'test', at(turn, 0), 71_000, 24_000, '212 pass · 0 fail · ran 1284 expect() calls in 70.44s'),
59]
60
61/**
62 * Three sample wasters for `/saver demo`: two awaiting a decision — the spawn one carrying the judge's
63 * own rule — and one already steered, whose rule the pane derives from the instruction that was sent.
64 */
65export const demoPatterns = (turn: number): Pattern[] => [
66  {
67    id: 'execution:full-suite',
68    category: 'execution',
69    kind: 'Claude keeps running the whole bun test suite after every single-file edit',
70    signature: { tool: 'Bash', key: SUITE_KEY },
71    why: 'the suite ran in full three times while only src/auth.ts changed between runs; the user asked for a fix, not full verification',
72    alternative: 'run only the tests covering the files you changed; run the full suite once when the phase is done',
73    confidence: 0.9,
74    proposal: null,
75    estTokensPerTurn: null,
76    lastDecision: null,
77    lean: null,
78    hits: ['demo-suite-1', 'demo-suite-2', 'demo-suite-3'],
79    decision: 'steer',
80    decidedAtTurn: at(turn, 1),
81    instruction: 'run only the tests for the file you just edited; the full suite once at the end of the phase',
82    openedAtTurn: at(turn, 1),
83    ignored: 0,
84    sent: 0,
85  },
86  {
87    id: 'reading:api-logs',
88    category: 'reading',
89    kind: 'Claude keeps reading 2000 lines of api logs instead of grepping for the error',
90    signature: { tool: 'Bash', key: LOG_KEY },
91    why: 'the whole log was read twice when one grep would have shown the traceback',
92    alternative: "grep -nE 'ERROR|Traceback' and read only the 50 lines around the match",
93    confidence: 0.8,
94    proposal: null,
95    estTokensPerTurn: null,
96    lastDecision: null,
97    lean: null,
98    hits: ['demo-log-1', 'demo-log-2'],
99    decision: null,
100    decidedAtTurn: null,
101    instruction: null,
102    openedAtTurn: null,
103    ignored: 0,
104    sent: 0,
105  },
106  {
107    id: 'multi-agent:re-explores',
108    category: 'multi-agent',
109    kind: 'Claude keeps spawning a fresh explore agent that re-reads what the last one already reported',
110    signature: { tool: 'Agent', key: AGENT_KEY },
111    why: 'two explore agents read the same six files in src/auth, and the second brief carried none of the first report',
112    alternative: "pass the previous explore agent's findings into the next brief instead of asking a new agent to read the same files",
113    confidence: 0.8,
114    proposal: {
115      kind: 'claude-md',
116      title: "Reuse an explore agent's findings",
117      body: "Give a new explore agent the previous agent's findings; never ask it to re-read files another agent has already reported on.",
118    },
119    estTokensPerTurn: null,
120    lastDecision: null,
121    lean: null,
122    hits: ['demo-agent-1', 'demo-agent-2'],
123    decision: null,
124    decidedAtTurn: null,
125    instruction: null,
126    openedAtTurn: null,
127    ignored: 0,
128    sent: 0,
129  },
130]
131
hooks/core/judge.ts 381 lines
1import type { ModelForkUsage } from 'claude-code'
2
3import { agentsBlock, citableRows, decisionsBlock, knownPatternsBlock, ledgerBlock, sinksBlock, statsLines, turnsBlock } from './blocks'
4import { agentAliases } from './evidence'
5import { totalTokens } from './patterns'
6import { cell, citedTurns, collapseWs, evidenceOf, groundedCap, isRecord, median, parseObject, proposalOf, short, str, unique } from './text'
7import {
8  ALTERNATIVE_MAX, JUDGE_MIN_GAP_MS, JUDGE_MIN_NEW_ROWS, JUDGE_MIN_NEW_TOKENS,
9  JUDGE_MIN_ROWS, JUDGE_MIN_TURNS, KIND_MAX, MAX_BEHAVIORAL_FINDINGS, MAX_FINDINGS, MAX_PATTERNS,
10} from './types'
11import type { Category, Finding, JudgeUsage, Pattern, Row, Signature, State } from './types'
12
13/** The judge prompt (build spec Appendix A, verbatim) with the eight evidence placeholders. */
14export const JUDGE_PROMPT = `You are auditing THIS session for wasted context and wasted time. The transcript above is your own: read it for intent — what the user asked for, what you were told, what you already decided. The blocks below are the only evidence of what actually ran; nothing outside them exists for this audit.
15
16Answer four narrow questions. What repeated: which behaviours have already happened more than once, separated by other work, or have you said you will keep doing — and what should be done instead? Where the time and the context went: which of the largest sinks below are repetition or work nobody asked for rather than the work this session needed. What is going in circles: the same failing command retried with no diagnostic step between, read/edit/read on one path with nothing finished, an edit failing on one path over and over. What was out of proportion: which spawned work — an agent, a workflow stage, a review or verification round — cost far more than what it produced, or did what a shell step would have done (AGENTS below).
17
18You are writing an interruption. Every finding can put a card in front of the user mid-work and can become a standing instruction that constrains you for the rest of the session. A wrong finding costs more than a missed one: it interrupts correct work, teaches a bad rule, and makes the user distrust the next card. Prefer silence to a guess. \`"findings": []\` is a correct and common answer.
19
20## Rules
211. Report behaviours, not incidents. A finding needs two unexcused occurrences of the same behaviour separated by other work (see Counting), or one occurrence plus your own stated intent to keep doing it ("I'll re-run the suite after each fix"). A single expensive call is never a finding.
222. \`evidence\`: row ids copied from the LEDGER \`id\` column (\`r12\`), or \`turn:<n>\` where \`<n>\` is a \`turn\` number printed in TURNS, or \`agent:<alias>\` where \`<alias>\` is an \`alias\` printed in AGENTS — turn and agent handles only for findings whose \`signature\` is null. Copy ids exactly; never renumber, abbreviate or reformat one. A finding with an id that is not in these blocks is discarded whole. STATS lines and \`~\` summary lines carry no id: use them for counts and history in \`why\` (they are the whole session, counted for you), never cite them as evidence, and never assume the oldest full row is the first occurrence.
233. \`signature\`: copy \`key\` character-for-character from one LEDGER row and \`tool\` from that same row. Keys are cut at 200 characters — copy the cut, never complete a command from memory. A pair that does not match a row discards the finding. Choose a key that only the wasteful version of the call carries; when no single recurring call carries the behaviour, use \`"signature": null\`.
244. Reuse ids. If KNOWN PATTERNS already names the behaviour, return that exact \`id\` with fresh evidence. The same command, file or lens under a different slug or category is the same waste: scan KNOWN PATTERNS and DECISIONS before minting an id.
255. Respect DECISIONS. A \`keep\` silences that behaviour under any id, category or wording for this session, and a kept occurrence may not be cited as evidence inside another finding. A behaviour kept in a previous session may be reported only with three or more occurrences and confidence 0.8 or higher. A \`steer\` or \`kill\` may be reported again only if it recurred after that turn: cite rows whose \`turn\` is greater and say so in \`why\`.
266. \`kind\`: one sentence of at most 120 characters, starting exactly \`Claude keeps \`, naming the concrete thing — the command, the file, the agent.
277. \`alternative\`: one imperative sentence of at most 200 characters addressed to Claude. It is sent to Claude verbatim, may be re-sent with every prompt for the rest of the session, and may be written into CLAUDE.md, so it must be safe to obey in situations you did not see: scope it ("run only the tests covering the files you changed, then the full suite once per phase"), never ban a capability outright ("never run the test suite"). If you cannot phrase the fix without forbidding something legitimate, drop the finding.
288. \`why\`: one or two sentences of evidence — how many times, what changed between occurrences, what the transcript shows — and the legitimate explanation you considered and what rules it out. Never a count or a cost these blocks do not contain.
299. At most six findings, at most three with \`signature: null\`, ordered by the sum of \`chars\` over the rows cited, largest first. Six is a ceiling, not a target; never split one behaviour into two findings.
30
31## Counting — what makes two occurrences a repeat
32- Separated by other work. Two occurrences of the same behaviour count as two decisions when at least one other row sits between them — an edit, a read, another command — whether in the same turn or a later one; a run after an edit is a second decision, not a second call in one batch — whether it is excused is the category's own question. Calls issued together with nothing between them are one batch and count once: a parallel set of Reads, or a fan-out of subagents launched at once, is one choice however wide. Breadth is one decision; weight is not: every stage of a workflow (each \`label\` in AGENTS) is a decision of its own, and the same role recurring stage after stage — a check loop per chunk, a review round after every fix — is a repeat.
33- Both unexcused. An occurrence the "Never report" list excuses does not count and may not be cited. Subtract the excused ones first; if fewer than two remain, there is no finding. A baseline suite run at the start plus the check before a commit is zero findings.
34- Same side of a compaction, for reading only. TURNS lists the turns where a compaction happened; what it dropped must be re-acquired, so a read or a re-derivation after one is not a repeat of one before it. Repeated work is: a check, a build, a review round or an agent run again across a compaction counts on both sides.
35- A correction is the first occurrence. When the user corrected a behaviour this session, or a saved feedback memory in the transcript names it, that correction counts as its first occurrence: one unexcused occurrence after it is a finding, and \`why\` names the correction.
36- Same behaviour, not the same shape. For a signature finding that means the same \`key\`; two Read keys differing only in \`:offset-limit\` are different slices, not a repeat. For a null-signature finding you must name one behaviour and show it in each cited turn; do not staple unrelated expensive turns together.
37- Agents are loops of their own. The \`agent\` column names the loop; a repeat inside one agent's rows counts exactly like a repeat in the main loop, and the main loop re-doing after an agent returns what that agent's rows show it already did (the same Read key, the same check) is a repeat across loops.
38- Short ledgers. With fewer than about 12 rows or fewer than 4 turns, report only behaviours with three or more surviving occurrences, or one plus explicit stated intent.
39- The legitimacy ladder. A sink needed once is nothing, however large. A sink repeated because its inputs changed between the runs — an edit, an install, a migration — is nothing. A sink repeated with nothing changed between, or work the transcript shows nobody asked for, is a finding, and the excuse you considered is written into \`why\`.
40
41## Categories — the nine names are the whole enum; the cues are examples and \`kind\` is free text
42- execution — the whole suite/build/typecheck after each edit; re-running a check with nothing edited since it last passed; the same failing command retried with no diagnostic step between; \`sleep\` polling or a watch/dev server run as a blocking call (\`bg\` absent, large \`ms\`). Not: a run after any intervening edit, install, migration or config change; the session's baseline run; broad verification after shared code changed; the last check before a commit; one retry of a transient failure. The error text is not in these blocks, so you cannot claim two failures were the same failure.
43- reading — the same path read again with nothing having changed it; whole-file reads where a range would do (\`trunc\`); unfiltered log/diff/verbose dumps (large \`chars\`, \`persist=\`); a Bash read of a file that is then Read again; wide greps with no path scope and no edit after them. Not: a read after your own edit or after any command that could have rewritten the file; a read after a compaction; paging (a second read at a new offset, especially after \`trunc\`); the first look at an unfamiliar file or log; a row flagged \`dedup\` (core charged nothing).
44- production — whole-file rewrites for small changes (\`+a/-d\` near the file's size); edits that cancel out; tests or docs nobody asked for; the same Edit failing on one path over and over. Not: a new file; a rewrite the user asked for; call-site updates the change requires.
45- behavior — read/edit/read oscillation with nothing finished; approach flip-flops; repeating what the user already corrected. Not: read-edit-verify cycles that are the working method, or a step that depends on the previous result.
46- communication — turns with \`calls 0\` and large \`answerChars\` that restate the plan or recap finished work; stopping to ask what the transcript, the repo or your instructions already answer; an \`ask\` row that held the turn for minutes while no agent ran (no rows between it and the next prompt) and whose options carried a recommended default (\`recommended\`) — proceed on the default and ask beside the work. Not: the turn that answers a question the user asked; a plan or explanation you were asked for; the session's last turn; plan mode, where making no tool call is required; a question whose answer the transcript shows changed the plan. \`out\` includes thinking, so point at the restated content, not the token shape.
47- multi-agent — parallel agents each re-reading the same large file the parent already had; agents with a thin brief (small \`promptChars\`, large \`tokens\`); results never read; agents spawned again after a limit error; mechanical agents on the premium model (\`agent=\` flag shows the resolved model and \`edits\`); the main loop re-reading files or re-running checks an agent's rows already covered, after it returned; a brief that pastes in whole files (large \`promptChars\`) to an agent whose rows then Read the same paths anyway; a workflow whose later agents re-read what earlier agents read (the same Read keys under successive \`agent\` values, spread over turns); a loop whose rows are only checks that passed, with \`edits 0\` and a report as its outcome (AGENTS \`checks\` > 0) — a shell step given a model; review or verify loops with \`edits 0\` whose outcome carries no medium, high or critical finding, recurring stage after stage; two or more verifier loops per finding; a run whose tokens per edit are several times the others'; a fix loop followed by a review whose outcome carries a new high in the same stage, twice (regression chasing: stop the loop and re-plan). Not: agents with disjoint file sets each reading one shared spec; two agents touching one path unless the ledger shows a conflict (an errored edit right after another agent's edit, or a re-edit in the main loop after they returned); a re-read whose brief the transcript shows is a review or verification pass; every agent reading the one spec its brief names; the parent reading an agent's result; one review per stage that found a medium or higher; a loop that edited; a fan-out's breadth on its own.
48- environment — installs repeated with no manifest edit; Bash used where Read/Grep/Edit is cheaper; fixed per-turn overhead (memory files, agent descriptions, MCP schemas in the facts line) larger than the work. Not: the user's own denials (\`denied\`); a single approval prompt.
49- process — many tiny commits or amends on one change; work declared done with no check run. Not: docs- or config-only changes with no check to run, or a check the environment cannot run.
50- other — a repetition none of the above names. Name it plainly.
51
52## Never report
53- A first occurrence, or anything with fewer than two unexcused occurrences after the Counting rules.
54- A single long call that was needed once, however long it ran: the largest row in TIME or CONTEXT is a fact to explain, never a finding on its own.
55- Orientation: the first look at any file, directory or log, an unfamiliar area, or a scope the user left open ("audit every call site", "review the repo").
56- Parallelism: calls issued together with nothing between them are one decision, and agents on disjoint scopes launched at once are one decision — one, not none: their weight is judged under multi-agent.
57- A re-read, or re-deriving what was settled, that a compaction made necessary. Compaction, prompt-cache reads and the host's own truncation are the harness working as designed; work repeated across a compaction is not excused by it.
58- A denied call (\`denied\`): the user or a policy said no, never your waste. The only reportable version is re-running an unchanged command already declined twice, and then the fix is a \`settings-allow\` proposal, not a rebuke.
59- A file change you cannot see: the \`paths\` column records only Edit/Write and Bash calls the host diffed, and nothing for a staged edit. Treat an intervening formatter, codegen, migration, install, \`git checkout|stash|pull|apply\`, \`sed -i\`, MCP edit or another agent's edit as having changed the file.
60- Volume alone. A large read is waste only when a cheaper call would have answered the same question for the same purpose; if the output was the deliverable (the diff under review, the log you were asked to explain, a file about to be rewritten) it is not a finding.
61- Turns spent thinking on a genuinely hard decision.
62- A cue whose evidence is not in these columns (an error message, a file's true size, worktree isolation): if you cannot see it, you cannot evidence it.
63- The duration of a \`recovered\` row, or which agent ran it: neither was recorded, and its \`err\` may be a refusal the transcript stored as error text, so treat a \`recovered\` \`err\` row as \`denied\` and never as waste.
64- Anything the user asked for this session, however wasteful it looks. Read the transcript before you accuse.
65
66## Confidence
67\`confidence\` runs 0.5 to 1.0. 0.9+: the same key three or more times with other work between each, nothing changed between, no request for it in the transcript. 0.7-0.9: the repetition is plain and the transcript offers no legitimate reason. 0.5-0.7: the repetition is real but a legitimate reason is plausible — report here only if you looked for that reason and \`why\` names what rules it out; if it could plausibly have been the right call, drop it. Below 0.5: say nothing.
68
69\`est_tokens_per_turn\`: null whenever \`signature\` is an object. For a null signature it is an integer grounded in the \`answerChars\` of the cited turns divided by four, conservative end, or 0 when you cannot ground it; the user sees it multiplied into a savings figure every turn after a decision. For agent handles it is the tokens one avoided loop would have cost, grounded in AGENTS \`tok\`, conservative end.
70
71\`proposal\`: null unless the fix should outlive the session. Otherwise \`{"kind","title","body"}\` where body is, per kind: \`claude-md\` one imperative rule line; \`skill\` the workflow as the body of a SKILL.md; \`agent-brief\` the brief, whose first line may be \`model: haiku\` or \`model: sonnet\`; \`settings-allow\` nothing but a permission rule such as \`Bash(bun test:*)\`.
72
73## Contract — the shape of your reply, stated once (documentation, not a template to echo)
74\`\`\`json
75{"focus": "<one line: what this session is doing>",
76 "time": "<one sentence, at most 200 chars: where the wall-clock went>",
77 "context": "<one sentence, at most 200 chars: where the context went>",
78 "findings": [{"id": "<category>:<kebab-slug, at most 40 chars>",
79   "category": "execution|reading|production|behavior|communication|multi-agent|environment|process|other",
80   "kind": "<one sentence, at most 120 chars, starts 'Claude keeps '>",
81   "evidence": ["<row id>", "turn:<n>"],
82   "signature": {"tool": "<the row's tool cell>", "key": "<the row's key cell, verbatim>"},
83   "why": "<one or two sentences>",
84   "alternative": "<one imperative sentence, at most 200 chars>",
85   "confidence": 0.85,
86   "est_tokens_per_turn": null,
87   "proposal": null}]}
88\`\`\`
89\`time\` and \`context\` explain where each went, for the user to read and in the words of the work: "45 min per chunk: the full proxy suite runs after every fix round and each chunk gets two review rounds". Neither is an accusation and neither is a finding by itself, so write both even when \`findings\` is \`[]\` — the examples below leave them out where they are not the point, your reply never does.
90
91\`findings\` may be \`[]\`. \`signature\` is that object or \`null\`. No other keys, and never null where a string is specified. Reply with one JSON object: first character \`{\`, last character \`}\`, no prose before or after, no code fence.
92
93## Examples — evidence, then what it justifies
94Rows \`r41\`, \`r45\`, \`r50\` carry \`test:bun test\` in turns 7, 8, 9 while only \`/src/auth.ts\` was edited between them, \`r41\` being the session's baseline run: \`{"focus":"fixing the auth token refresh in /src/auth.ts","findings":[{"id":"execution:full-suite-after-each-edit","category":"execution","kind":"Claude keeps running the whole bun test suite after every single-file edit","evidence":["r45","r50"],"signature":{"tool":"Bash","key":"test:bun test"},"why":"r41 was the baseline and is excused; the suite then ran in full at turns 8 and 9 after single-file edits to /src/auth.ts alone, about a minute and 9.7k characters each. Nothing shared changed, and the user asked for a fix, not full verification.","alternative":"Run only the test files covering the files you changed, then the whole suite once when the phase is done.","confidence":0.92,"est_tokens_per_turn":null,"proposal":{"kind":"claude-md","title":"Targeted tests","body":"Run only the tests covering the files you changed; run the full suite at the end of a phase."}}]}\`
95Rows \`r12\` Read \`/src/api.ts:-\`, \`r15\` Edit \`/src/api.ts\`, \`r16\` Read \`/src/api.ts:-\` with \`dedup\`, all in turn 4: \`{"focus":"a one-file change in /src/api.ts","findings":[]}\` — the second read follows your own edit and the third was deduped, so no unexcused occurrence remains.
96Rows \`r08\` \`test:bun test\` (turn 2, baseline), \`r23\` \`test:bun test test/db.test.ts\` (turn 6, after an edit), \`r40\` \`test:bun test\` (turn 11) followed by \`r41\` \`git:git commit …\`: \`{"focus":"a db pool fix, verified narrowly then once before the commit","findings":[]}\` — both full runs are excused, so nothing survives the Counting rules.
97Rows \`r61\` and \`r72\` both \`read:docker compose logs api --tail 2000\` in turns 11 and 13, each ~40k \`chars\` with \`persist=\`: the same shape as the first example with \`"id":"reading:unfiltered-log-dump"\`, \`"kind":"Claude keeps reading 2000 lines of api logs instead of grepping for the error"\`, \`"alternative":"Pipe log commands through grep -nE 'ERROR|Traceback' and tail -50 instead of reading the whole tail."\`, \`"confidence":0.85\`, \`"proposal":null\`.
98TURNS shows turns 14 and 15 with \`calls 0\` and \`answerChars\` 5400 and 6100 after a single edit at turn 13, neither answering a question: \`"id":"communication:restates-plan-each-turn"\`, \`"evidence":["turn:14","turn:15"]\`, \`"signature":null\`, \`"est_tokens_per_turn":1200\`, \`"alternative":"State the result in one or two lines and take the next action; do not restate the plan or recap completed steps."\`.
99Rows \`r80\`-\`r83\` under \`agent\` \`a1\` Read four files, then \`r84\` Agent \`agent:general-purpose\` flagged \`agent=general-purpose/opus/completed/41000tok/0edits/300pch\` closes that loop (a spawn row lands after the rows it caused), then \`r90\`-\`r93\` in the main loop Read the same four keys in the next turn: \`"id":"multi-agent:re-reads-what-the-agent-read"\`, \`"kind":"Claude keeps re-reading the files a subagent already read for it"\`, \`"evidence":["r80","r83","r90","r93"]\` (a row from each loop is the repeat; the four main-loop reads together are one batch), \`"signature":null\` (no single key carries it), \`"alternative":"Use the subagent's report; re-read a file it covered only to edit it."\`, \`"confidence":0.8\`, \`"est_tokens_per_turn"\` grounded in the cited turns' \`answerChars\` or 0.
100KNOWN PATTERNS lists \`execution:full-suite-after-each-edit | … | steer @ 9\` and rows \`r70\` (turn 12) and \`r76\` (turn 14) carry \`test:bun test\` again: return that same id with \`"evidence":["r70","r76"]\` and a \`why\` that names turns 12 and 14 as after the steer at turn 9.
101Agent rows \`r30\`, \`r58\` and \`r91\` run about 40 minutes each, a review agent follows each, every loop reads a different module and no key repeats: \`{"focus":"rewriting three modules, one agent each","time":"2h 40m in three module rewrites of about 40 minutes each, plus one review pass per module; nothing ran twice.","context":"1.1M chars, three quarters of it the agents' own reads of the modules they rewrote.","findings":[]}\` — a long session is not a wasteful one.
102Four suite runs with an install or a migration between every pair, and \`agents | ×3 | Σ2100000ms | apart\` above them in TIME: \`{"focus":"a schema migration and the call sites it broke","time":"68m, most of it four suite runs, each after a migration or an install changed what the suite covers.","context":"620k chars, over half the migration diff and the failures it produced.","findings":[]}\` — every repeat had changed inputs, so the ladder stops at nothing.
103AGENTS lists \`a3 | w3 | check:C3 | sonnet | 1 | 1.7m | 48k | edits 0 | checks 4 | reads 0 | report 1600ch | answer\` and \`a7 | w3 | check:C4 | …\` alike, and no row of theirs carries \`err\`: \`{"id":"multi-agent:check-loops-for-shell-steps","kind":"Claude keeps spawning an agent per chunk whose only job is to run four passing check commands","evidence":["agent:a3","agent:a7"],"signature":null,"alternative":"Run a stage's checks as a shell step of the workflow script or inside the reviewer; spawn an agent only for work that needs judgment.","confidence":0.85,"est_tokens_per_turn":48000}\`
104Three \`review:*\` loops of one run, each \`edits 0\` with outcome \`3 low\`, a fix loop after each: \`{"id":"multi-agent:review-rounds-that-find-only-lows","kind":"Claude keeps running a full review round after every fix although the last two found only low findings","evidence":["agent:a9","agent:a12"],"signature":null,"alternative":"Route only medium-or-higher review findings to a fix round; end the stage when a review returns lows alone.","confidence":0.8}\`
105
106## KNOWN PATTERNS — \`id | kind | decision @ turn | previous\`. Reuse these ids; never mint a second id or signature for waste listed here.
107{{KNOWN_PATTERNS}}
108
109## DECISIONS — \`id | key | keep|steer|kill @ turn\`, then previous-session keeps. A \`keep\` key is off limits under any id this session.
110{{DECISIONS}}
111
112## STATS — the whole session, counted for you. Per call: \`tool | key | cls | ×count | Σms | Σchars | turns first-last | edits-between | agents\` (edits-between: median number of files edited between consecutive runs; 0 means it re-ran with nothing changed), then \`waits:\` — every AskUserQuestion this session, its total wait and how many carried a recommended default; not citable, but the time it held the session is a fact to explain. Then per class and per agent. No ids here; cite LEDGER rows.
113{{STATS}}
114
115## TIME — where the wall-clock went. \`total\`, then \`label | ×count | Σms | share%\` for the largest sinks, then the five longest rows as \`r<seq> | tool | key | Σms\`. An Agent row holds its own loop's rows, so it is listed apart and never added in; \`ms\` includes any wait on a permission prompt.
116{{TIME}}
117
118## CONTEXT — where the context went. The same shape measured in \`chars\`: the total, the largest sinks with their share, then the five largest rows. An Agent row is listed apart and never added in.
119{{CONTEXT}}
120
121## AGENTS — the loops this session spawned. First one line per run: \`name | id | loops | Σmin | Σtok | edits | turn\`. Then one line per loop, oldest first: \`alias | run | label | model | turns | min | tok | edits | checks | reads | outcome | ended\` — \`alias\` is the LEDGER's \`agent\` name for that loop, \`label\` the stage the workflow gave it (\`impl:C3\`, \`check:C3\`, \`review:C6-r1\`), \`outcome\` what it returned (\`1 high 4 low\`, \`report 5900ch\`), \`ended\` \`answer\`, \`error\`, \`aborted\`, \`refusal\` or \`running\`. \`tok\` counts new tokens (input, cache creation, output) in thousands. Loops older than the window fold into \`~ run | ×loops | Σmin | Σtok\` lines. Cite a loop as \`agent:<alias>\`; a line here is a loop's whole cost, so its weight against its \`edits\` and \`outcome\` is the proportion question's evidence.
122{{AGENTS}}
123
124## TURNS — \`turn | in | out | cacheCreate | calls | ms | answerChars\`, then \`| aborted\`, \`| error\` or \`| refusal\` when the turn ended that way and \`| idle <m>m\` when the next prompt came a minute or more later, then the facts line (context window, fixed per-turn overhead, turns where a compaction happened)
125{{TURNS}}
126
127## LEDGER — \`id | tool | key | cls | agent | turn | ms | chars | flags | paths\`, oldest first (a spawn row lands after the rows it caused: an agent's own calls finish before its Agent row does). Agents are named \`a1\`, \`a2\`… in order of first appearance; \`main\` is the main loop. \`ms\` is wall time and includes any wait on a permission prompt, so a long \`ms\` alone is not machine cost. A row flagged \`recovered\` was rebuilt from the transcript before this plugin joined the session: its \`ms\` is 0 and its agent reads \`main\`, so never reason about its duration or which loop ran it. flags: \`err\` \`denied\` \`dedup\` \`trunc\` \`bg\` \`timeout\` \`persist=<bytes>\` \`ask\` (an AskUserQuestion: its \`ms\` is the wait for the person) \`recommended\` (its options named a default) \`+adds/-dels\` \`agent=<type>/<model>/<status>/<tokens>tok/<edits>edits/<promptChars>pch\`, or \`-\`. Rows older than the window are folded into \`~ | tool | key | ×count | Σchars\` lines: no id, never citable, key usable as a signature only if it also appears in a full row.
128{{LEDGER}}
129
130Return the JSON object only.
131`
132
133// Turns and tokens: the ordinary session, where the judge runs between turns.
134const turnGate = (state: State): boolean =>
135  totalTokens(state) - state.judge.lastAtTokens >= JUDGE_MIN_NEW_TOKENS * state.judge.backoff &&
136  state.turn - state.judge.lastAtTurn >= JUDGE_MIN_TURNS
137
138// Rows and wall time: one agentic turn can run for hours, and `turn.complete` is no cadence inside it.
139const rowGate = (state: State, now: number): boolean =>
140  state.seq - state.judge.lastAtSeq >= JUDGE_MIN_NEW_ROWS * state.judge.backoff &&
141  now - state.judge.lastAtMs >= JUDGE_MIN_GAP_MS
142
143/**
144 * True when the cadence gates allow another judge run.
145 *
146 * @param state the session so far
147 * @param now the clock, for the gap the mid-turn gate keeps between runs
148 */
149export const shouldRun = (state: State, now: number): boolean =>
150  !state.judge.running &&
151  state.rows.length >= JUDGE_MIN_ROWS &&
152  (turnGate(state) || rowGate(state, now))
153
154/** The four counts the fork reported, under our own names: what `/saver debug` and the debug log print. */
155export const usageOf = (u: ModelForkUsage): JudgeUsage => ({
156  input: u.input_tokens,
157  output: u.output_tokens,
158  cacheRead: u.cache_read_input_tokens,
159  cacheCreate: u.cache_creation_input_tokens,
160})
161
162/** What a run's counts cost us: input, output and cache creation (cache reads are free). */
163export const spentOf = (u: JudgeUsage): number => u.input + u.output + u.cacheCreate
164
165/** What one judge fork cost us, straight from the usage the API reported. */
166export const costOf = (u: ModelForkUsage): number => spentOf(usageOf(u))
167
168/** Names this session's loops the way one judge run sees them: built once before the prompt, read again when the reply comes back. */
169export const judgeAliases = (state: State): ReadonlyMap<string, string> => agentAliases(state.rows, state.loops)
170
171/** Fills the judge prompt with this session's evidence blocks; `aliases` is the naming AGENTS and LEDGER print, so the caller can hand the same one to `parseReply`. */
172export const buildPrompt = (state: State, aliases: ReadonlyMap<string, string> = judgeAliases(state)): string =>
173  ([
174    ['{{KNOWN_PATTERNS}}', knownPatternsBlock(state)],
175    ['{{DECISIONS}}', decisionsBlock(state)],
176    ['{{STATS}}', statsLines(state.rows, state.folded, aliases).join('\n')],
177    ['{{TIME}}', sinksBlock(state.rows, 'ms', state.loops, state.folded)],
178    ['{{CONTEXT}}', sinksBlock(state.rows, 'chars', [], state.folded)],
179    ['{{AGENTS}}', agentsBlock(state, aliases)],
180    ['{{TURNS}}', turnsBlock(state)],
181    ['{{LEDGER}}', ledgerBlock(state, aliases)],
182  ] as const).reduce((text, [placeholder, value]) => text.split(placeholder).join(value), JUDGE_PROMPT)
183
184const ID_SHAPE = /^[a-z-]+:[a-z0-9-]{1,40}$/
185
186// A slug naming the thing as the work spells it (`multi-agent:C6-review-rounds`) is the same waste as its
187// lowercase self, and case is not a judgement: the slug is lowered before the shape is tested, and the
188// finding is stored under the lowered id, so one behaviour keeps one id across runs and sessions.
189const lowerSlug = (id: string): string => {
190  const at = id.indexOf(':')
191  return at < 0 ? id : `${id.slice(0, at)}:${id.slice(at + 1).toLowerCase()}`
192}
193
194const isCategory = (v: unknown): v is Category =>
195  typeof v === 'string' &&
196  ['execution', 'reading', 'production', 'behavior', 'communication', 'multi-agent', 'environment', 'process', 'other'].includes(v)
197
198// undefined = the key is absent or malformed, or no visible row carries the pair: the finding is discarded.
199const signatureOf = (value: unknown, visible: Row[]): Signature | null | undefined => {
200  if (value === null) return null
201  if (!isRecord(value)) return undefined
202  const tool = str(value['tool'])
203  const key = str(value['key'])
204  return visible.some(r => r.tool === tool && r.key === key) ? { tool, key } : undefined
205}
206
207const signatureDrop = (value: unknown): string =>
208  isRecord(value)
209    ? `signature (${cell(value['tool'], 20)}, ${cell(value['key'], 40)}) matches no row`
210    : 'signature must be a (tool, key) pair or null'
211
212const estOf = (value: unknown, state: State, signature: Signature | null, evidence: string[]): number | null => {
213  if (signature !== null) return null
214  const claimed = typeof value === 'number' && Number.isFinite(value) ? Math.max(0, Math.round(value)) : 0
215  return Math.min(claimed, groundedCap(state, evidence))
216}
217
218const sameSignature = (a: Signature | null, b: Signature | null): boolean =>
219  a !== null && b !== null && a.tool === b.tool && a.key === b.key
220
221const isKept = (state: State, id: string, signature: Signature | null): boolean =>
222  state.patterns.some(p =>
223    p.decision === 'keep' && (p.id === id || sameSignature(p.signature, signature)))
224
225// One cited occurrence is a finding only after a first the ledger cannot show: stated intent, a user correction or a saved feedback memory.
226const FIRST_OCCURRENCE = /intent|\bcorrect(?:ed|ion)\b|\bfeedback\b/i
227
228// A string is the one short reason the finding was dropped; the object is the finding itself.
229const findingOf = (value: unknown, state: State, visible: Row[], aliases: ReadonlyMap<string, string>): Finding | string => {
230  if (!isRecord(value)) return 'not an object'
231  const given = str(value['id'])
232  const id = lowerSlug(given)
233  const category = value['category']
234  const kind = str(value['kind'])
235  const why = str(value['why'])
236  const alternative = str(value['alternative'])
237  const confidence = value['confidence']
238  if (!ID_SHAPE.test(id)) return `id ${cell(given, 40) || '(missing)'} is not <category>:<kebab-slug>`
239  if (!isCategory(category)) return 'category is not one of the nine'
240  if (id.slice(0, id.indexOf(':')) !== category) return `category ${category} does not match the id`
241  if (!kind.startsWith('Claude keeps ')) return 'kind must start with "Claude keeps "'
242  // A cap missed by a few characters is a sentence that ran long, not a wrong finding: the pane wraps, so it
243  // is kept as returned and the run report notes the length. Twice the cap is a different answer, and dropped.
244  if (kind.length > KIND_MAX * 2) return `kind is ${kind.length} chars, over twice ${KIND_MAX}`
245  if (alternative.length === 0) return 'alternative is empty'
246  if (alternative.length > ALTERNATIVE_MAX * 2) return `alternative is ${alternative.length} chars, over twice ${ALTERNATIVE_MAX}`
247  if (typeof confidence !== 'number' || confidence < 0.5 || confidence > 1) return 'confidence is not a number in 0.5..1'
248  const signature = signatureOf(value['signature'], visible)
249  if (signature === undefined) return signatureDrop(value['signature'])
250  if (isKept(state, id, signature)) return 'kept this session'
251  const evidence = evidenceOf(value['evidence'], state, visible, aliases, signature)
252  if (typeof evidence === 'string') return evidence
253  if (evidence.length < 2 && !FIRST_OCCURRENCE.test(why)) return 'one handle and no stated intent or correction in why'
254  return {
255    id, category, kind, evidence, signature, why, alternative, confidence,
256    estTokensPerTurn: estOf(value['est_tokens_per_turn'], state, signature, evidence),
257    proposal: proposalOf(value['proposal']),
258    lean: null,
259  }
260}
261
262// Each finding under the label the drop report names it by: its own id, or its place in the reply.
263type Reviewed = { label: string; finding: Finding } | { label: string; reason: string }
264type Sifted = { findings: Finding[]; dropped: string[] }
265
266const labelOf = (value: unknown, index: number): string => {
267  const id = lowerSlug(str(isRecord(value) ? value['id'] : ''))
268  return ID_SHAPE.test(id) ? id : `#${index + 1}`
269}
270
271const capFindings = (reviewed: readonly Reviewed[]): Sifted =>
272  reviewed.reduce<Sifted>((kept, item) => {
273    const dropped = (reason: string): Sifted => ({ findings: kept.findings, dropped: [...kept.dropped, `${item.label}: ${reason}`] })
274    if (!('finding' in item)) return dropped(item.reason)
275    if (kept.findings.length >= MAX_FINDINGS) return dropped(`over MAX_FINDINGS (${MAX_FINDINGS})`)
276    if (item.finding.signature === null && kept.findings.filter(k => k.signature === null).length >= MAX_BEHAVIORAL_FINDINGS) {
277      return dropped(`over MAX_BEHAVIORAL_FINDINGS (${MAX_BEHAVIORAL_FINDINGS})`)
278    }
279    if (kept.findings.some(k => k.id === item.finding.id)) return dropped('the reply already reported this id')
280    return { findings: [...kept.findings, item.finding], dropped: kept.dropped }
281  }, { findings: [], dropped: [] })
282
283const EXPLAIN_MAX = 200   // characters of the judge's `time` and `context` sentences kept; the rest is cut off
284
285// One sentence written for the user: an essay is cut to the sentence's length rather than thrown away,
286// because its first 200 characters still say where the time went. Anything but text is nothing to draw.
287const explanationOf = (value: unknown): string | null => {
288  const text = collapseWs(str(value))
289  return text.length > 0 ? short(text, EXPLAIN_MAX) : null
290}
291
292// An essay where a sentence was asked for is a prompt problem, so the run report says the cut happened:
293// without it a trimmed sentence reads like a sentence the judge chose to end there.
294const explanationDrop = (label: string, value: unknown): string[] => {
295  const text = collapseWs(str(value))
296  return text.length > EXPLAIN_MAX ? [`${label}: ${text.length} chars, trimmed to ${EXPLAIN_MAX}`] : []
297}
298
299// The caps a kept finding overran, reported the same way: nothing was cut, so the note is the whole story.
300const capNotes = (f: Finding): string[] => [
301  ...(f.kind.length > KIND_MAX ? [`${f.id}: kind: ${f.kind.length} chars, over ${KIND_MAX}`] : []),
302  ...(f.alternative.length > ALTERNATIVE_MAX ? [`${f.id}: alternative: ${f.alternative.length} chars, over ${ALTERNATIVE_MAX}`] : []),
303]
304
305/** What the judge said: the valid findings, the focus, its time and context sentences, one line per drop and per cap a kept finding missed, and how many it returned; never throws. `aliases` must be the table `buildPrompt` printed. */
306export const parseReply = (text: string, state: State, aliases: ReadonlyMap<string, string> = judgeAliases(state)): { findings: Finding[]; focus: string | null; time: string | null; context: string | null; dropped: string[]; returned: number } => {
307  const root = parseObject(text)
308  if (root === null) return { findings: [], focus: null, time: null, context: null, dropped: ['reply was not JSON'], returned: 0 }
309  const raw = root['findings']
310  const focusText = collapseWs(str(root['focus']))
311  const focus = focusText.length > 0 ? focusText : null
312  const said = { time: explanationOf(root['time']), context: explanationOf(root['context']) }
313  const overlong = [...explanationDrop('time', root['time']), ...explanationDrop('context', root['context'])]
314  if (!Array.isArray(raw)) return { findings: [], focus, ...said, dropped: [...overlong, 'findings was not an array'], returned: 0 }
315  const visible = citableRows(state)
316  const reviewed: Reviewed[] = raw.map((value, i) => {
317    const label = labelOf(value, i)
318    const result = findingOf(value, state, visible, aliases)
319    return typeof result === 'string' ? { label, reason: result } : { label, finding: result }
320  })
321  const sifted = capFindings(reviewed)
322  const noted = sifted.findings.flatMap(capNotes)
323  return { ...sifted, dropped: [...overlong, ...noted, ...sifted.dropped], focus, ...said, returned: raw.length }
324}
325
326const patternOf = (f: Finding): Pattern => ({
327  id: f.id, category: f.category, kind: f.kind, signature: f.signature, why: f.why,
328  alternative: f.alternative, confidence: f.confidence, proposal: f.proposal,
329  estTokensPerTurn: f.estTokensPerTurn, lastDecision: null, lean: f.lean,
330  hits: [...f.evidence], decision: null, decidedAtTurn: null, instruction: null, openedAtTurn: null, ignored: 0, sent: 0,
331})
332
333const updatedWith = (p: Pattern, f: Finding): Pattern => ({
334  ...p, why: f.why, alternative: f.alternative, proposal: f.proposal, confidence: f.confidence,
335  estTokensPerTurn: f.estTokensPerTurn, hits: unique([...p.hits, ...f.evidence]),
336})
337
338const hasRecurred = (state: State, p: Pattern, f: Finding): boolean =>
339  (p.decision === 'steer' || p.decision === 'kill') && p.decidedAtTurn !== null &&
340  citedTurns(state, f.evidence).some(turn => turn > (p.decidedAtTurn ?? 0))
341
342const rank = (p: Pattern): number => (p.decision === null ? 0 : 1000) + p.confidence
343
344// A cited pattern nobody has decided belongs in front of the user, whether the id is new, remembered by
345// the store or found by an earlier run: the only undecided pattern that is not news is one already queued.
346const isFresh = (state: State, patterns: Pattern[], id: string): boolean =>
347  patterns.find(p => p.id === id)?.decision === null && !state.cards.includes(id)
348
349const capPatterns = (patterns: Pattern[]): Pattern[] => {
350  if (patterns.length <= MAX_PATTERNS) return patterns
351  const dropped = [...patterns]
352    .sort((a, b) => rank(a) - rank(b))
353    .slice(0, patterns.length - MAX_PATTERNS)
354    .map(p => p.id)
355  return patterns.filter(p => !dropped.includes(p.id))
356}
357
358/** Folds findings into the complete registry, naming the fresh, the recurred, and the ids the cap evicted. */
359export const merge = (state: State, findings: Finding[]): { patterns: Pattern[]; fresh: string[]; recurred: string[]; evicted: string[] } => {
360  const added = findings.filter(f => !state.patterns.some(p => p.id === f.id))
361  const patterns = capPatterns([
362    ...state.patterns.map(p => {
363      const f = findings.find(x => x.id === p.id)
364      return f !== undefined ? updatedWith(p, f) : p
365    }),
366    ...added.map(patternOf),
367  ])
368  const kept = (id: string): boolean => patterns.some(p => p.id === id)
369  const recurred = state.patterns.filter(p => {
370    const f = findings.find(x => x.id === p.id)
371    return f !== undefined && hasRecurred(state, p, f)
372  })
373  return {
374    patterns,
375    fresh: findings.map(f => f.id).filter(id => isFresh(state, patterns, id)),
376    recurred: recurred.map(p => p.id).filter(kept),
377    // A validated finding the cap pushed out has no card: the shell reports it as dropped, not kept.
378    evicted: findings.map(f => f.id).filter(id => !kept(id)),
379  }
380}
381
hooks/core/ledger.ts 221 lines
1import type { ToolCallResult } from 'claude-code'
2
3import { collapseWs, stableJson } from './text'
4import { FILE_TOOLS, KEY_MAX, MAIN_AGENT } from './types'
5import type { CommandClass, Row } from './types'
6
7/** The flat tool event a row is built from: the tool, this call's id, the loop, and the tool's arguments beside them. */
8export type ToolEvent = { tool: string; tool_use_id: string; agentId?: string } & Record<string, unknown>
9
10const RESERVED = ['tool', 'tool_use_id', 'agentId', 'consent'] as const
11
12const HEAD_MAX = 80   // characters of result.text quoted as evidence (Row.head)
13
14// A leading `cd <dir> &&`, `VAR=value`, `timeout <duration>` or `time` is noise in front of the command that matters.
15const NOISE = /^(?:cd\s+[^\s&|;]+\s*&&\s*|[A-Za-z_][A-Za-z0-9_]*=(?:"[^"]*"|'[^']*'|\S*)\s+|timeout\s+\d+[smhd]?\s+|time\s+)/
16
17// `a && b`, `a; b`, `a || b`: one call can carry several commands, and quoting is not worth parsing.
18const CHAIN = /&&|\|\||;/
19
20// From the first pipe or redirect on (an fd digit like `2>` goes with it): how the output was filtered, not what ran.
21const FILTER = /\s*(?:\d?[<>]|\|).*$/
22
23// Script runners: what follows them is the command that matters (longest first).
24const RUNNERS = ['npm run', 'bun run', 'bun x', 'pnpm run', 'yarn run', 'python3 -m', 'python -m', 'npx', 'bunx', 'pnpm', 'yarn'] as const
25
26// Script names seen after a runner (`bun run lint`), which no binary in the table covers.
27const SCRIPTS: Readonly<Record<string, CommandClass>> = { test: 'test', lint: 'lint', format: 'format', typecheck: 'typecheck', build: 'build' }
28
29// Command heads, matched on a token boundary; two-token heads never shadow one another.
30const COMMANDS: readonly (readonly [string, CommandClass])[] = [
31  ['npm test', 'test'], ['bun test', 'test'], ['go test', 'test'], ['cargo test', 'test'],
32  ['jest', 'test'], ['vitest', 'test'], ['pytest', 'test'], ['mocha', 'test'],
33  ['biome lint', 'lint'], ['eslint', 'lint'], ['ruff', 'lint'], ['flake8', 'lint'], ['golangci-lint', 'lint'],
34  ['biome format', 'format'], ['prettier', 'format'], ['black', 'format'], ['gofmt', 'format'], ['rustfmt', 'format'],
35  ['tsc', 'typecheck'], ['mypy', 'typecheck'], ['pyright', 'typecheck'],
36  ['make build', 'build'], ['cargo build', 'build'], ['go build', 'build'], ['docker build', 'build'], ['vite', 'build'], ['webpack', 'build'],
37  ['npm install', 'install'], ['npm i', 'install'], ['bun install', 'install'], ['bun add', 'install'],
38  ['pnpm install', 'install'], ['pnpm add', 'install'], ['yarn install', 'install'], ['yarn add', 'install'],
39  ['pip install', 'install'], ['uv pip', 'install'], ['cargo fetch', 'install'], ['apt-get install', 'install'], ['brew install', 'install'],
40  ['git', 'git'], ['gh', 'git'],
41  ['cat', 'read'], ['head', 'read'], ['tail', 'read'], ['less', 'read'], ['ls', 'read'], ['tree', 'read'], ['sed', 'read'],
42  ['grep', 'search'], ['rg', 'search'], ['ag', 'search'], ['find', 'search'], ['fd', 'search'], ['ast-grep', 'search'],
43]
44
45const asRecord = (v: unknown): Record<string, unknown> => (typeof v === 'object' && v !== null ? (v as Record<string, unknown>) : {})
46
47const asString = (v: unknown): string | null => (typeof v === 'string' ? v : null)
48
49const asNumber = (v: unknown): number | null => (typeof v === 'number' && Number.isFinite(v) ? v : null)
50
51const pathArg = (args: Record<string, unknown>): string =>
52  asString(args.file_path) ?? asString(args.notebook_path) ?? ''
53
54const stripNoise = (command: string): string => {
55  let rest = command
56  while (NOISE.test(rest)) rest = rest.replace(NOISE, '')
57  return rest
58}
59
60const headToken = (command: string): string => command.split(' ')[0] ?? ''
61
62// The head by its basename: `venv/bin/python -m pytest` is `python -m pytest` to the table.
63const baseHead = (command: string): string => {
64  const head = headToken(command)
65  return `${head.slice(head.lastIndexOf('/') + 1)}${command.slice(head.length)}`
66}
67
68const tableClass = (command: string): CommandClass | null =>
69  COMMANDS.find(([head]) => command === head || command.startsWith(`${head} `))?.[1] ?? null
70
71const bareOf = (segment: string): string => stripNoise(collapseWs(segment))
72
73const segmentClass = (segment: string): CommandClass => {
74  const bare = baseHead(bareOf(segment))
75  const direct = tableClass(bare)
76  if (direct !== null) return direct
77  const runner = RUNNERS.find(r => bare === r || bare.startsWith(`${r} `))
78  if (runner === undefined) return 'other'
79  const script = bare.slice(runner.length).trim()
80  return tableClass(script) ?? SCRIPTS[headToken(script)] ?? 'other'
81}
82
83const segmentsOf = (command: string): string[] => collapseWs(command).split(CHAIN)
84
85/** Classifies a shell command by what it does, seeing through cd, env, timeout, runner prefixes and `&&`/`;`/`||` chains. */
86export const classOf = (command: string): CommandClass =>
87  segmentsOf(command)
88    .map(segmentClass)
89    .find(cls => cls !== 'other') ?? 'other'
90
91// The key names what ran: the segment that classified, as typed, minus the noise before it and the filtering after it.
92const canonicalOf = (command: string): string =>
93  bareOf(segmentsOf(command).find(segment => segmentClass(segment) !== 'other') ?? '').replace(FILTER, '').trim()
94
95/** Computes the ledger key and command class of a call from its arguments alone. */
96export const normalize = (tool: string, input: unknown): { key: string; cls: CommandClass } => {
97  const args = asRecord(input)
98  if (tool === 'Bash') {
99    const command = collapseWs(asString(args.command) ?? '')
100    const cls = classOf(command)
101    // An `other` command has no segment that named something: the whole pipeline is the key.
102    return { key: `${cls}:${cls === 'other' ? command : canonicalOf(command)}`.slice(0, KEY_MAX), cls }
103  }
104  if (FILE_TOOLS.includes(tool)) {
105    const path = pathArg(args)
106    const range = tool === 'Read' ? `:${asNumber(args.offset) ?? ''}-${asNumber(args.limit) ?? ''}` : ''
107    return { key: `${path}${range}`.slice(0, KEY_MAX), cls: tool === 'Read' ? 'read' : 'other' }
108  }
109  if (tool === 'Grep' || tool === 'Glob') {
110    return { key: `${tool}:${asString(args.pattern) ?? ''}:${asString(args.path) ?? ''}`.slice(0, KEY_MAX), cls: 'search' }
111  }
112  if (tool === 'Agent') return { key: `agent:${asString(args.subagent_type) ?? 'general'}`.slice(0, KEY_MAX), cls: 'other' }
113  return { key: `${tool}:${stableJson(input, RESERVED)}`.slice(0, KEY_MAX), cls: 'other' }
114}
115
116const printable = (ch: string): boolean => {
117  const code = ch.codePointAt(0) ?? 0
118  return code >= 0x20 && code !== 0x7f
119}
120
121// Code point by code point, stopping before the one that would not fit whole: never half a surrogate pair.
122const headOf = (text: string | undefined): string => {
123  let head = ''
124  for (const ch of text ?? '') {
125    const kept = ch === '\t' || ch === '\n' ? ' ' : ch
126    if (!printable(kept)) continue
127    if (head.length + kept.length > HEAD_MAX) break
128    head += kept
129  }
130  return head
131}
132
133const flagsOf = (e: ToolEvent, result: ToolCallResult, res: Record<string, unknown>): string[] => {
134  const persisted = e.tool === 'Bash' ? asNumber(res.persistedOutputSize) : null
135  // `bg` describes a call that ran: a denied or errored Bash started no background task.
136  const isBackground = answered(result) && (e.run_in_background === true || asString(res.backgroundTaskId) !== null)
137  // An ask's `ms` is the wait for the person; `recommended` says the model marked an option for them.
138  const isAsk = e.tool === 'AskUserQuestion'
139  return [
140    result.isError === true ? 'err' : '',
141    result.deny !== undefined ? 'denied' : '',
142    e.tool === 'Read' && res.type === 'file_unchanged' ? 'dedup' : '',
143    e.tool === 'Read' && asRecord(res.file).truncatedByTokenCap === true ? 'trunc' : '',
144    e.tool === 'Bash' && isBackground ? 'bg' : '',
145    e.tool === 'Bash' && asNumber(res.timedOutAfterMs) !== null ? 'timeout' : '',
146    persisted !== null ? `persist=${persisted}` : '',
147    isAsk ? 'ask' : '',
148    isAsk && JSON.stringify(e.questions ?? '').includes('(Recommended)') ? 'recommended' : '',
149  ].filter(flag => flag !== '')
150}
151
152const answered = (result: ToolCallResult): boolean => result.deny === undefined && result.isError !== true
153
154const patchLines = (patch: unknown): { add: number; del: number } => {
155  const hunks = Array.isArray(patch) ? patch : []
156  const lines = hunks.flatMap(hunk => {
157    const own = asRecord(hunk).lines
158    return Array.isArray(own) ? own.filter((line): line is string => typeof line === 'string') : []
159  })
160  return { add: lines.filter(line => line.startsWith('+')).length, del: lines.filter(line => line.startsWith('-')).length }
161}
162
163const lineCount = (content: string): number => (content === '' ? 0 : content.replace(/\n$/, '').split('\n').length)
164
165const linesOf = (e: ToolEvent, res: Record<string, unknown>): { add: number; del: number } | null => {
166  if (e.tool === 'Edit') {
167    const diff = asRecord(res.gitDiff)
168    const add = asNumber(diff.additions)
169    const del = asNumber(diff.deletions)
170    return add !== null && del !== null ? { add, del } : patchLines(res.structuredPatch)
171  }
172  if (e.tool === 'Write') return { add: lineCount(asString(res.content) ?? asString(e.content) ?? ''), del: 0 }
173  return null
174}
175
176const pathsOf = (e: ToolEvent, res: Record<string, unknown>): string[] => {
177  if (e.tool === 'Edit' || e.tool === 'Write' || e.tool === 'NotebookEdit') {
178    const path = asString(res.filePath) ?? (pathArg(e) || null)
179    return path !== null && res.staged !== true ? [path] : []
180  }
181  if (e.tool === 'Bash') {
182    const changed = asRecord(res.bashEditDiff).changedFiles
183    return Array.isArray(changed) ? changed.filter((file): file is string => typeof file === 'string') : []
184  }
185  return []
186}
187
188const spawnOf = (e: ToolEvent, res: Record<string, unknown>): Row['spawn'] =>
189  e.tool !== 'Agent'
190    ? null
191    : {
192        type: asString(e.subagent_type) ?? 'general',
193        requested: asString(e.model),
194        resolved: asString(res.resolvedModel),
195        status: asString(res.status),
196        tokens: asNumber(res.totalTokens),
197        edits: asNumber(asRecord(res.toolStats).editFileCount),
198        promptChars: (asString(e.prompt) ?? '').length,
199      }
200
201/** Builds the ledger row of a finished tool call, reading every result field defensively. */
202export const rowOf = (e: ToolEvent, result: ToolCallResult, ms: number, turn: number): Omit<Row, 'seq'> => {
203  const res = asRecord(result.result)
204  const { key, cls } = normalize(e.tool, e)
205  return {
206    id: e.tool_use_id,
207    tool: e.tool,
208    key,
209    cls,
210    agent: asString(e.agentId) ?? MAIN_AGENT,
211    turn,
212    ms,
213    chars: result.text?.length ?? 0,
214    head: headOf(result.text),
215    flags: flagsOf(e, result, res),
216    lines: answered(result) ? linesOf(e, res) : null,
217    paths: answered(result) ? pathsOf(e, res) : [],
218    spawn: spawnOf(e, res),
219  }
220}
221
hooks/core/patterns.ts 872 lines
1import { agentAliases, aliasOf, baseline, foldRows, rowsOf, sinks, sumOf } from './evidence'
2import { measuredCost } from './process'
3import { activeRuns, countRow } from './spawns'
4import { collapseWs, duration, instructionOf, killPrompt, kilo, median, pctOf } from './text'
5import {
6  ALTERNATIVE_MAX, CARD_EVIDENCE, DEBUG_MAX_DROPPED, DEBUG_MAX_LINES, DEBUG_MAX_PATTERNS, FILE_TOOLS,
7  JUDGE_BUDGET_SHARE, JUDGE_MAX_BACKOFF, KEY_MAX, KIND_MAX, LOOP_CAP, MAIN_AGENT, MAX_PATTERNS, NO_CALLS, ROW_CAP,
8  SETTLE_TURNS, TREND_TURNS, initialState,
9} from './types'
10import type {
11  Action, Artifact, BandModel, Card, Choice, CommandClass, DecidedRow, Evidence, Header, JournalEntry, JudgeRun,
12  JudgeUsage, Loop, PaneModel, Pattern, Proposal, Row, Signature, Sinks, State, StoredPattern, Tokens, TurnStat,
13} from './types'
14
15// The `:offset-limit` slice `normalize` appends to a Read key: the path is what the details name.
16const READ_RANGE = /:\d*-\d*$/
17
18// The loop an `agent:<id>` evidence handle names.
19const AGENT_HANDLE = /^agent:(.+)$/
20
21const GROWTH_TURNS = 5     // turns with a context sample the pace to compaction is read from
22const GROWTH_SAMPLES = 3   // growth samples below which no pace is stated at all
23
24type Decided = Pattern & { decision: Choice }
25type Settle = { pattern: Pattern; ms: number; chars: number; requeue: string | null }
26
27const isSent = (c: Choice | null): boolean => c === 'steer' || c === 'kill'
28
29const isDecided = (p: Pattern): p is Decided => p.decision !== null
30
31const pushUnique = (xs: readonly string[], x: string): string[] => (xs.includes(x) ? [...xs] : [...xs, x])
32
33const queueCard = (cards: readonly string[], id: string): string[] => (cards.includes(id) ? [...cards] : [id, ...cards])
34
35const patternById = (patterns: readonly Pattern[], id: string): Pattern | undefined => patterns.find(p => p.id === id)
36
37const signatureHit = (p: Pattern, row: Pick<Row, 'tool' | 'key'>): boolean =>
38  p.signature !== null && p.signature.tool === row.tool && p.signature.key === row.key
39
40const classOfPattern = (state: State, p: Pattern): CommandClass | null => rowsOf(state, p)[0]?.cls ?? null
41
42const turnHandle = (handle: string): number | null => {
43  const n = /^turn:(\d+)$/.exec(handle)?.[1]
44  return n === undefined ? null : Number(n)
45}
46
47const loopsCited = (p: Pattern, state: State): Loop[] =>
48  p.hits.flatMap(handle => {
49    const id = AGENT_HANDLE.exec(handle)?.[1]
50    const loop = id === undefined ? undefined : state.loops.find(l => l.id === id)
51    return loop === undefined ? [] : [loop]
52  })
53
54const turnsCited = (p: Pattern, state: State, rows: readonly Row[]): number[] => {
55  const fromHandles = p.hits.map(turnHandle).filter((n): n is number => n !== null)
56  return [...rows.map(r => r.turn), ...fromHandles, ...loopsCited(p, state).map(l => l.firstTurn)].sort((a, b) => a - b)
57}
58
59const newTokens = (t: TurnStat): number => t.input + t.output + t.cacheCreate
60
61/** Sums every new (non-cache-read) token this session's turns reported. */
62export const totalTokens = (state: State): number => state.turns.reduce((n, t) => n + newTokens(t), 0)
63
64/** Tokens left before auto-compaction; null while the session's token count is unknown. */
65export const tokensToCompaction = (state: State): number | null => {
66  const tokens = state.usage.tokens
67  if (tokens === undefined) return null
68  return (state.usage.compactAt ?? Math.round(state.usage.window * 0.9)) - tokens
69}
70
71/** The context after each turn that reported one, oldest first: how full the window was, turn by turn. */
72const contexts = (state: State): number[] =>
73  state.turns.map(t => t.context).filter((tokens): tokens is number => tokens !== null)
74
75/**
76 * Turns left before compaction at the recent pace; null under three growth samples or a zero median.
77 *
78 * The pace is how fast the window fills, not what a turn is billed: a turn can spend 60k tokens and
79 * grow the window by 8k, so the tokens of a turn would have claimed compaction was three turns away.
80 * A turn that shrank the window (a compaction, a `/clear`) is no pace at all, so drops are skipped.
81 */
82export const turnsToCompaction = (state: State): number | null => {
83  const left = tokensToCompaction(state)
84  const seen = contexts(state).slice(-GROWTH_TURNS)
85  const growth = seen.slice(1).map((tokens, at) => tokens - (seen[at] ?? 0)).filter(step => step > 0)
86  if (left === null || growth.length < GROWTH_SAMPLES) return null
87  const perTurn = median(growth)
88  return perTurn === 0 ? null : Math.round(left / perTurn)
89}
90
91const applySettlements = (state: State, list: readonly Settle[]): Pick<State, 'patterns' | 'cards' | 'saved'> => ({
92  patterns: list.map(s => s.pattern),
93  cards: list.reduce<string[]>((cards, s) => (s.requeue === null ? cards : queueCard(cards, s.requeue)), [...state.cards]),
94  saved: {
95    ms: list.reduce((ms, s) => ms + s.ms, state.saved.ms),
96    chars: list.reduce((chars, s) => chars + s.chars, state.saved.chars),
97  },
98})
99
100const settleWithRow = (state: State, p: Pattern, row: Omit<Row, 'seq'>): Settle => {
101  const grown: Pattern = signatureHit(p, row) ? { ...p, hits: pushUnique(p.hits, row.id) } : p
102  const still = { pattern: grown, ms: 0, chars: 0, requeue: null }
103  if (row.agent !== 'main' || p.openedAtTurn === null || !isSent(p.decision)) return still
104  // A row in the decision's own turn was already in flight before Claude could read the instruction.
105  if (row.turn <= p.openedAtTurn) return still
106  // D4: every ignored instruction brings the card back, so the user can Ignore it or say something else.
107  if (p.signature !== null && row.key === p.signature.key) {
108    return { pattern: { ...grown, ignored: p.ignored + 1, openedAtTurn: null }, ms: 0, chars: 0, requeue: p.id }
109  }
110  if (row.cls !== classOfPattern(state, p)) return still
111  const base = baseline(state, p)
112  return {
113    pattern: { ...grown, openedAtTurn: null },
114    ms: Math.max(0, base.ms - row.ms),
115    chars: Math.max(0, base.chars - row.chars),
116    requeue: null,
117  }
118}
119
120const settleAtTurn = (state: State, p: Pattern): Settle => {
121  // D4: an ignored instruction saves nothing, so a behavioural pattern accrues only while it is still believed.
122  const accrued = p.signature === null && isSent(p.decision) && p.ignored === 0 ? (p.estTokensPerTurn ?? 0) * 4 : 0
123  if (p.openedAtTurn === null || state.turn - p.openedAtTurn < SETTLE_TURNS) {
124    return { pattern: p, ms: 0, chars: accrued, requeue: null }
125  }
126  const base = baseline(state, p)
127  return { pattern: { ...p, openedAtTurn: null }, ms: base.ms, chars: base.chars + accrued, requeue: null }
128}
129
130const NO_TOKENS: Tokens = { input: 0, output: 0, cacheRead: 0, cacheCreate: 0 }
131
132const addTokens = (a: Tokens, b: Tokens): Tokens =>
133  ({ input: a.input + b.input, output: a.output + b.output, cacheRead: a.cacheRead + b.cacheRead, cacheCreate: a.cacheCreate + b.cacheCreate })
134
135// A loop known by nothing but its id yet: dated by the turn and the row that first showed it.
136const bareLoop = (id: string, firstTurn: number, firstSeq: number): Loop => ({
137  id, run: null, label: null, phase: null, model: null, turns: 0, ms: 0, tokens: NO_TOKENS, ended: null, firstTurn, firstSeq, outcome: null,
138  calls: 0, edits: 0, checks: 0, reads: 0,
139})
140
141// Every change to a loop: created where unknown, grown where known, the oldest dropped past the cap.
142const withLoop = (loops: readonly Loop[], id: string, grow: (l: Loop) => Loop, first: { turn: number; seq: number }): Loop[] => {
143  const known = loops.some(l => l.id === id)
144  const grown = known ? loops.map(l => (l.id === id ? grow(l) : l)) : [...loops, grow(bareLoop(id, first.turn, first.seq))]
145  return grown.slice(-LOOP_CAP)
146}
147
148const applyLoopTurn = (state: State, a: Extract<Action, { type: 'loop.turn' }>): State => ({
149  ...state,
150  loops: withLoop(state.loops, a.agentId, l => ({
151    ...l, turns: l.turns + 1, ms: l.ms + a.ms, tokens: addTokens(l.tokens, a.tokens), ended: a.ended, model: a.model ?? l.model,
152  }), { turn: a.turn, seq: state.seq }),
153})
154
155const applyAgentStart = (state: State, a: Extract<Action, { type: 'agent.start' }>): State => ({
156  ...state,
157  loops: withLoop(state.loops, a.agentId, l => ({
158    ...l, label: a.description.length > 0 ? a.description : l.label, model: a.model ?? l.model,
159  }), { turn: state.turn, seq: state.seq }),
160})
161
162const applyRunStart = (state: State, run: { id: string; name: string; dir: string | null }, now: number): State =>
163  state.runs.some(r => r.id === run.id)
164    ? state   // a resume: the run is the one already known, dated at its first launch
165    : { ...state, runs: [...state.runs, { ...run, turn: state.turn, seq: state.seq, at: now, refreshedAt: 0, phases: [] }] }
166
167// A `started` entry names the loop's run and stage; a `result` entry what it returned. Either creates the loop.
168const applyEntry = (state: State, runId: string, loops: readonly Loop[], entry: JournalEntry): Loop[] =>
169  withLoop(loops, entry.agentId, l =>
170    entry.kind === 'started'
171      ? { ...l, run: runId, label: entry.label ?? l.label, phase: entry.phase ?? l.phase }
172      : { ...l, outcome: entry.outcome }, { turn: state.turn, seq: state.seq })
173
174const applyRunJournal = (state: State, a: Extract<Action, { type: 'run.journal' }>): State => ({
175  ...state,
176  loops: a.entries.reduce((loops, entry) => applyEntry(state, a.runId, loops, entry), state.loops),
177  runs: state.runs.map(r => (r.id === a.runId ? { ...r, refreshedAt: a.now } : r)),
178})
179
180// The wait before this prompt is written onto the turn it followed, so TURNS can say the person was away.
181const applyTurnStart = (state: State, now: number): State => {
182  const last = state.turns[state.turns.length - 1]
183  const dated = last !== undefined && last.at > 0 ? { ...last, idleMs: Math.max(0, now - last.at) } : last
184  return { ...state, turn: state.turn + 1, turns: dated === undefined ? state.turns : [...state.turns.slice(0, -1), dated] }
185}
186
187// The ledger with the cap applied: a row it pushes off the front is folded into its pair, never simply lost,
188// so STATS, the LEDGER's `~` lines and the sinks still count a call the judge can no longer cite.
189const capped = (state: State, rows: readonly Row[]): Pick<State, 'rows' | 'folded'> => {
190  const over = rows.length - ROW_CAP
191  return over <= 0
192    ? { rows: [...rows], folded: state.folded }
193    : { rows: rows.slice(over), folded: foldRows(state.folded, rows.slice(0, over)) }
194}
195
196const applyRow = (state: State, row: Omit<Row, 'seq'>): State => {
197  const seq = state.seq + 1
198  return {
199    ...state,
200    seq,
201    ...capped(state, [...state.rows, { ...row, seq }]),
202    // A row of a loop nobody has named yet names it here, so the aliases stay in the ledger's order; every
203    // row of a loop is counted on it here, so its line reads whole after ROW_CAP has dropped the row.
204    loops: row.agent === MAIN_AGENT ? state.loops : withLoop(state.loops, row.agent, l => countRow(l, row), { turn: row.turn, seq }),
205    ...applySettlements(state, state.patterns.map(p => settleWithRow(state, p, row))),
206  }
207}
208
209const grownWith = (p: Pattern, rows: readonly Row[]): Pattern =>
210  rows.filter(row => signatureHit(p, row)).reduce((q, row) => ({ ...q, hits: pushUnique(q.hits, row.id) }), p)
211
212// History, not a live call: no turn stats exist for it, nothing settles on it, and nothing was saved by it.
213const applyAdopt = (state: State, rows: readonly Omit<Row, 'seq'>[]): State => {
214  const seeded = rows.map((row, i) => ({ ...row, seq: state.seq + i + 1 }))
215  const seq = state.seq + seeded.length
216  return {
217    ...state,
218    seq,
219    turn: Math.max(state.turn, ...seeded.map(row => row.turn)),
220    ...capped(state, [...state.rows, ...seeded]),
221    patterns: state.patterns.map(p => grownWith(p, seeded)),
222    // The mid-turn cadence starts where the history ends: adopted rows are not new work, so a session
223    // joined late waits for JUDGE_MIN_NEW_ROWS of its own. `/saver check` still judges them on request.
224    judge: { ...state.judge, lastAtSeq: seq },
225  }
226}
227
228const applyTurnComplete = (state: State, stat: Omit<TurnStat, 'turn' | 'calls'>): State => ({
229  ...state,
230  turns: [...state.turns, { ...stat, turn: state.turn, calls: state.rows.filter(r => r.turn === state.turn).length }],
231  ...applySettlements(state, state.patterns.map(p => settleAtTurn(state, p))),
232})
233
234// Present values win; `window`/`compactAt` sticky (§5.2, widened so a partial usage dispatch never blanks the header).
235// The first sample also dates the mid-turn cadence: without it `lastAtMs` is 0 against a clock reading ms
236// since the epoch, so the five-minute floor would be no floor at all. `/saver demo` passes 0 and seeds nothing.
237const applyUsage = (state: State, usage: State['usage'], now: number): State => ({
238  ...state,
239  usage: {
240    window: usage.window || state.usage.window,
241    compactAt: usage.compactAt ?? state.usage.compactAt,
242    tokens: usage.tokens ?? state.usage.tokens,
243    percent: usage.percent ?? state.usage.percent,
244  },
245  judge: state.judge.lastAtMs === 0 && now > 0 ? { ...state.judge, lastAtMs: now } : state.judge,
246})
247
248const REPLAN = 'Change the remaining work: '
249
250// The one-time re-plan a process Fix sends, said once even when the finding's own words already say it.
251const replanOf = (text: string): string => (text.trimStart().toLowerCase().startsWith(REPLAN.toLowerCase()) ? text.trim() : `${REPLAN}${text.trim()}`)
252
253const applyDecide = (state: State, id: string, choice: Choice, text: string | undefined): State => {
254  const p = patternById(state.patterns, id)
255  if (p === undefined) return state
256  const process = p.lean !== null
257  // A process Fix sends the lean alternative itself: `killPrompt` stops a behaviour, a re-plan changes the work.
258  const instruction = choice === 'kill' ? (process ? p.alternative : killPrompt(p)) : (text ?? '')
259  if (choice !== 'keep' && instruction.trim() === '') return state
260  const sending = choice !== 'keep'
261  // Fix rides the same wrapper as Fix… (§5.2 "same with killPrompt(p)", Appendix C 5c "the same way"):
262  // Claude reads `Instruction from the user (via ContextSaver): Stop this behaviour …`.
263  const note = instructionOf(process ? replanOf(instruction) : instruction)
264  const decided: Pattern = {
265    ...p,
266    decision: choice,
267    decidedAtTurn: state.turn,
268    lastDecision: choice,
269    instruction: sending ? instruction : p.instruction,
270    openedAtTurn: sending ? state.turn : p.openedAtTurn,
271  }
272  return {
273    ...state,
274    patterns: state.patterns.map(q => (q.id === p.id ? decided : q)),
275    cards: state.cards.filter(c => c !== p.id),
276    expanded: state.expanded === p.id ? null : state.expanded,
277    steering: null,
278    steerDraft: null,
279    notes: sending ? pushUnique(state.notes, note) : [...state.notes],
280    // A re-plan is the orchestrator's, once (D3): standing would hand it to every subagent spawned after.
281    standing: sending && !process ? pushUnique(state.standing, note) : [...state.standing],
282  }
283}
284
285// A process run's patterns join the registry beside the habit judge's; its cadence and cost are its own.
286const applyProcessDone = (state: State, a: Extract<Action, { type: 'process.done' }>): State => {
287  const patterns = [...a.patterns, ...state.patterns.filter(p => patternById(a.patterns, p.id) === undefined)]
288  const wanted = a.fresh.filter(id => patternById(patterns, id)?.decision === null && !state.cards.includes(id))
289  const total = totalTokens(state)
290  const spent = state.process.spent + a.spent
291  return {
292    ...state,
293    patterns,
294    cards: [...wanted, ...state.cards],
295    process: {
296      ...state.process,
297      running: false,
298      runs: state.process.runs + 1,
299      spent,
300      backoff: total > 0 && spent > JUDGE_BUDGET_SHARE * total ? state.process.backoff * 2 : state.process.backoff,
301      error: a.error,
302      last: { returned: a.returned, kept: a.kept, dropped: a.dropped, usage: a.usage },
303    },
304  }
305}
306
307// One agent reached by the standing instructions: counted once, on every pattern whose instruction is standing.
308const applyDeliverySent = (state: State, agentId: string): State =>
309  state.delivery.agents.includes(agentId)
310    ? state
311    : {
312        ...state,
313        delivery: { ...state.delivery, agents: [...state.delivery.agents, agentId] },
314        patterns: state.patterns.map(p => (p.lean === null && p.instruction !== null && state.standing.some(t => t.includes(p.instruction ?? '')) ? { ...p, sent: p.sent + 1 } : p)),
315      }
316
317const applyJudgeDone = (state: State, a: Extract<Action, { type: 'judge.done' }>): State => {
318  const recurred = (p: Pattern): boolean => a.recurred.includes(p.id) && isSent(p.decision)
319  const marked = a.patterns.filter(recurred).map(p => p.id)
320  const reported = a.patterns.map(p => (recurred(p) ? { ...p, ignored: p.ignored + 1, openedAtTurn: null } : p))
321  // A session decision is never lost, even when the judge's registry omits it.
322  const patterns = [...reported, ...state.patterns.filter(p => isDecided(p) && patternById(reported, p.id) === undefined)]
323  const wanted = [...a.fresh.filter(id => patternById(patterns, id)?.decision === null), ...marked]
324  const added = wanted.filter((id, i) => wanted.indexOf(id) === i && !state.cards.includes(id))
325  const total = totalTokens(state)
326  const spent = state.judge.spent + a.spent
327  return {
328    ...state,
329    patterns,
330    cards: [...added, ...state.cards].filter(id => patternById(patterns, id) !== undefined),
331    judge: {
332      lastAtTokens: total,
333      lastAtTurn: state.turn,
334      lastAtSeq: state.judge.lastAtSeq,
335      lastAtMs: state.judge.lastAtMs,
336      running: false,
337      runs: state.judge.runs + 1,
338      spent,
339      // No completed turn is no budget to be over: the first run of a reloaded session doubles nothing.
340      backoff: total > 0 && spent > JUDGE_BUDGET_SHARE * total ? Math.min(state.judge.backoff * 2, JUDGE_MAX_BACKOFF) : state.judge.backoff,
341      error: a.error,
342      focus: a.focus,
343      time: a.time,
344      context: a.context,
345      last: { returned: a.returned, kept: a.kept, dropped: [...a.dropped], usage: a.usage },
346    },
347    // Only a run that answered spends the arming: a cold snapshot or a refusal leaves it for the next opportunity.
348    pendingCheck: a.error === null ? false : state.pendingCheck,
349  }
350}
351
352const applyReset = (state: State): State => ({
353  ...initialState(state.cwd, state.usage.window),
354  overhead: state.overhead,
355  columns: state.columns,
356  paneOpen: state.paneOpen,
357  patterns: state.patterns.map(p => fromStored(toStored(p))),
358})
359
360/** Applies one action to the state, returning a new state (never mutates its input). */
361export const reduce = (state: State, action: Action): State => {
362  switch (action.type) {
363    case 'turn.start':
364      return applyTurnStart(state, action.now)
365    case 'loop.turn':
366      return applyLoopTurn(state, action)
367    case 'agent.start':
368      return applyAgentStart(state, action)
369    case 'run.start':
370      return applyRunStart(state, action.run, action.now)
371    case 'run.journal':
372      return applyRunJournal(state, action)
373    case 'row':
374      return applyRow(state, action.row)
375    case 'adopt':
376      return applyAdopt(state, action.rows)
377    case 'turn.complete':
378      return applyTurnComplete(state, action.stat)
379    case 'usage':
380      return applyUsage(state, action.usage, action.now)
381    case 'overhead':
382      return { ...state, overhead: action.overhead }
383    case 'compact':
384      // The fill a compaction invalidated is forgotten: `tokens` is sticky and a just-compacted window
385      // reports none until its next response (d.ts 7023-7025), so the header awaits the next turn instead
386      // of announcing two turns to a compaction that just happened.
387      return { ...state, compactions: [...state.compactions, state.turn], usage: { ...state.usage, tokens: undefined, percent: undefined } }
388    case 'expand':
389      return { ...state, expanded: action.patternId === state.expanded ? null : action.patternId }
390    case 'steer.begin':
391      return { ...state, steering: state.steering === action.patternId ? null : action.patternId, steerDraft: null }
392    case 'steer.draft':
393      return { ...state, steerDraft: action.text }
394    case 'decide':
395      return applyDecide(state, action.patternId, action.choice, action.text)
396    case 'judge.start':
397      return { ...state, judge: { ...state.judge, running: true, lastAtMs: action.now, lastAtSeq: action.seq } }
398    case 'judge.done':
399      return applyJudgeDone(state, action)
400    case 'ask':
401      return { ...state, asks: [...state.asks, action.ask] }
402    case 'run.phases':
403      return { ...state, runs: state.runs.map(r => (r.id === action.runId ? { ...r, phases: [...action.phases] } : r)) }
404    case 'process.start':
405      return { ...state, process: { ...state.process, running: true, lastAtMs: action.now, lastAtSeq: action.seq, lastAtTurn: state.turn } }
406    case 'process.done':
407      return applyProcessDone(state, action)
408    case 'process.cold':
409      return { ...state, process: { ...state.process, ...action.was, running: false } }
410    case 'delivery.pending':
411      return { ...state, delivery: { ...state.delivery, pending: Math.max(0, state.delivery.pending + action.delta) } }
412    case 'delivery.sent':
413      return applyDeliverySent(state, action.agentId)
414    case 'delivery.toasted':
415      return { ...state, delivery: { ...state.delivery, lastToastAt: action.now } }
416    case 'check.arm':
417      return { ...state, pendingCheck: true }
418    case 'notes.drained':
419      return { ...state, notes: [] }
420    case 'standing.add':
421      return { ...state, standing: pushUnique(state.standing, action.text) }
422    case 'artifact.done':
423      return {
424        ...state,
425        patterns: state.patterns.map(p => (p.id === action.patternId ? { ...p, proposal: null } : p)),
426        written: action.written ? pushUnique(state.written, `${action.patternId}:${action.kind}`) : [...state.written],
427      }
428    case 'pane':
429      return { ...state, paneOpen: action.open, autoOpened: action.auto === true ? true : state.autoOpened }
430    case 'columns':
431      return { ...state, columns: action.columns }
432    case 'reset':
433      return applyReset(state)
434  }
435}
436
437const statsOf = (p: Pattern, state: State, rows: readonly Row[]): string => {
438  const cost = sumOf(rows)
439  const pct = pctOf(cost.chars, state.usage.window)
440  const turns = turnsCited(p, state, rows)
441  const first = turns[0]
442  const last = turns[turns.length - 1]
443  // Zero segments are dropped whole: turn handles and rebuilt rows carry no duration, and `0s` would claim a suite that ran for minutes cost nothing.
444  return [
445    `${p.hits.length}×`,
446    ...(pct > 0 ? [`~${pct}% of context`] : []),
447    ...(cost.ms > 0 ? [duration(cost.ms)] : []),
448    // One turn is not a range: 'turns 1–1' reads as a bug.
449    ...(first === undefined || last === undefined ? [] : [first === last ? `turn ${first}` : `turns ${first}–${last}`]),
450  ].join(' · ')
451}
452
453// What ran, in the words the person typed or read: the command, the file, or the tool and its key.
454const whatRan = (r: Row): string => {
455  if (r.tool === 'Bash') return r.key.startsWith(`${r.cls}:`) ? r.key.slice(r.cls.length + 1) : r.key
456  if (FILE_TOOLS.includes(r.tool)) return r.key.replace(READ_RANGE, '')
457  return `${r.tool} ${r.key}`
458}
459
460const citedRows = (rows: readonly Row[], aliases: ReadonlyMap<string, string>): Evidence[] =>
461  rows.map(r => ({
462    turn: r.turn,
463    what: whatRan(r),
464    // The loop is named only when it was not the main one: an alias on every row would be noise.
465    agent: r.agent === MAIN_AGENT ? null : aliasOf(aliases, r.agent),
466    ms: r.ms,
467    chars: r.chars,
468    head: r.head,
469  }))
470
471const citedTurnStats = (p: Pattern, state: State): TurnStat[] =>
472  p.hits
473    .map(turnHandle)
474    .map(n => (n === null ? undefined : state.turns.find(t => t.turn === n)))
475    .filter((t): t is TurnStat => t !== undefined)
476
477const ktokOf = (l: Loop): number => Math.round((l.tokens.input + l.tokens.cacheCreate + l.tokens.output) / 1000)
478
479// A cited loop, quoted as one line of evidence: its stage and model, dated by the turn that spawned it.
480const citedLoop = (state: State, aliases: ReadonlyMap<string, string>, l: Loop): Evidence => ({
481  turn: l.firstTurn,
482  what: `${l.label ?? 'agent'} · ${l.model ?? '?'}`,
483  agent: aliasOf(aliases, l.id),
484  ms: l.ms,
485  chars: 0,
486  head: `${ktokOf(l)}k tokens · ${l.edits} edits`,
487})
488
489const evidenceOf = (p: Pattern, state: State, rows: readonly Row[], aliases: ReadonlyMap<string, string>): Evidence[] => {
490  const turns: Evidence[] = citedTurnStats(p, state)
491    .map(t => ({ turn: t.turn, what: NO_CALLS, agent: null, ms: 0, chars: t.answerChars, head: t.answerHead }))
492  const loops: Evidence[] = loopsCited(p, state).map(l => citedLoop(state, aliases, l))
493  // Reversed first, so two calls inside one turn also read newest-first once the stable sort has run.
494  return [...citedRows(rows, aliases).reverse(), ...loops.reverse(), ...turns]
495    .sort((a, b) => b.turn - a.turn)
496    .slice(0, CARD_EVIDENCE)
497}
498
499// A pattern with no row in hand cites loops or turns, not calls: it counts them, and its context cost is the
500// judge's per-turn estimate. The unit is stated here, so the drawing never has to guess it back.
501const totalOf = (p: Pattern, state: State, rows: readonly Row[]): Card['total'] => {
502  if (rows.length > 0) return { unit: 'calls', calls: rows.length, ...sumOf(rows) }
503  const loops = loopsCited(p, state)
504  const chars = (p.estTokensPerTurn ?? 0) * 4
505  if (loops.length > 0) return { unit: 'agents', calls: loops.length, ms: loops.reduce((ms, l) => ms + l.ms, 0), chars }
506  return { unit: 'turns', calls: citedTurnStats(p, state).length, ms: 0, chars }
507}
508
509// A process card's Cost so far: the newest run's claim, clamped to its cited handles; what they measured when none is held.
510const costSoFar = (state: State, p: Pattern): string => {
511  const cost = p.cost ?? measuredCost(state, p.hits)
512  return `${duration(cost.ms)} · ${kilo(cost.tokens)} tokens`
513}
514
515/** Derives the waster card in seat `n` from a pattern, the evidence it cites and the ledger's loop aliases. */
516export const cardOf = (p: Pattern, state: State, n: number, aliases: ReadonlyMap<string, string>): Card => {
517  const rows = rowsOf(state, p)
518  return {
519    patternId: p.id,
520    n,
521    category: p.category,
522    kind: p.ignored > 0 ? `ignored · ${p.kind}` : p.kind,
523    stats: p.lean === null ? statsOf(p, state, rows) : costSoFar(state, p),
524    why: p.why,
525    fix: p.alternative,
526    lean: p.lean,
527    total: totalOf(p, state, rows),
528    evidence: evidenceOf(p, state, rows, aliases),
529  }
530}
531
532// Null before the first turn completes: the session's own tokens are 0 then, and a share of nothing is
533// not a figure — a run that cost 24k printed `2394600%` of a denominator that had not been measured yet.
534const judgeShare = (state: State): number | null => {
535  const total = totalTokens(state)
536  return total === 0 ? null : Math.round((state.judge.spent / total) * 1000) / 10
537}
538
539// How full the window was after each of the last turns that reported it: the shape the header draws.
540const trendOf = (state: State): number[] =>
541  contexts(state)
542    .slice(-TREND_TURNS)
543    .map(tokens => Math.round((tokens / Math.max(1, state.usage.window)) * 1000) / 10)
544
545// Before the first row nothing was measured, and a total of zero would read as a session that cost nothing.
546const sinksOf = (state: State, measure: 'ms' | 'chars'): Sinks | null =>
547  state.rows.length === 0 ? null : sinks(state.rows, measure, state.loops, state.folded)
548
549const headerOf = (state: State): Header => ({
550  percent: state.usage.percent ?? null,
551  tokensToCompaction: tokensToCompaction(state),
552  turnsToCompaction: turnsToCompaction(state),
553  trend: trendOf(state),
554  time: sinksOf(state, 'ms'),
555  context: sinksOf(state, 'chars'),
556  judgeTime: state.judge.time,
557  judgeContext: state.judge.context,
558  judgeRuns: state.judge.runs,
559  judgeTokens: state.judge.spent,
560  judgeShare: judgeShare(state) ?? 0,
561  judgeRunning: state.judge.running,
562  savedPct: pctOf(state.saved.chars, state.usage.window),
563  savedMs: state.saved.ms,
564})
565
566const decidedRowOf = (state: State, p: Decided): DecidedRow => ({
567  patternId: p.id,
568  choice: p.decision,
569  kind: p.kind,
570  savedPct: isSent(p.decision) ? pctOf(baseline(state, p).chars, state.usage.window) : null,
571  // D4: the figure is a projection of one avoided repeat until the instruction settled — nothing ignored
572  // it and nothing is still in flight — so the drawing can say `per repeat` before it says `saved`.
573  settled: isSent(p.decision) && p.openedAtTurn === null && p.ignored === 0,
574  // What Fix sends is the kind and the fix the row already carries; only a note of the user's own is news.
575  instruction: p.decision === 'steer' && p.instruction !== null && collapseWs(p.instruction) !== collapseWs(p.alternative)
576    ? p.instruction
577    : null,
578  ignored: p.ignored,
579  sent: p.sent,
580})
581
582/** Builds everything the pane renders: header, wasters, decisions and rules. */
583export const paneModel = (state: State, artifacts: Artifact[]): PaneModel => {
584  // One alias table for the whole draw: the naming is the ledger's, not a card's, and the pane
585  // redraws on every ledger row and every keystroke in the Fix… field.
586  const aliases = agentAliases(state.rows, state.loops)
587  return {
588    header: headerOf(state),
589    // Numbered as they are drawn, so `/saver keep 2` names the card the person is looking at.
590    wasters: state.cards
591      .map(id => patternById(state.patterns, id))
592      .filter((p): p is Pattern => p !== undefined)
593      .map((p, at) => cardOf(p, state, at + 1, aliases)),
594    expanded: state.expanded,
595    steering: state.steering,
596    steerDraft: state.steerDraft,
597    decided: state.patterns
598      .filter(isDecided)
599      .sort((a, b) => (b.decidedAtTurn ?? 0) - (a.decidedAtTurn ?? 0))
600      .map(p => decidedRowOf(state, p)),
601    artifacts,
602  }
603}
604
605// What the cards still awaiting a decision have already cost: the figures the band's teaser states.
606const waitingCost = (state: State): { ms: number; chars: number } =>
607  state.cards
608    .map(id => patternById(state.patterns, id))
609    .filter((p): p is Pattern => p !== undefined)
610    .reduce(
611      (sum, p) => {
612        const total = totalOf(p, state, rowsOf(state, p))
613        return { ms: sum.ms + total.ms, chars: sum.chars + total.chars }
614      },
615      { ms: 0, chars: 0 },
616    )
617
618// The last turn's end when it was an error or a refusal and no prompt has followed it: the session stopped.
619const diedOf = (state: State): BandModel['died'] => {
620  const last = state.turns[state.turns.length - 1]
621  if (last === undefined || last.turn !== state.turn) return null
622  return last.ended === 'error' || last.ended === 'refusal' ? last.ended : null
623}
624
625// The newest run still going: how many loops it has spawned, how many calls they made, which stage is running.
626const runningOf = (state: State, now: number): BandModel['running'] => {
627  const active = activeRuns(state, now)
628  const run = active[active.length - 1]
629  if (run === undefined) return null
630  const loops = state.loops.filter(l => l.run === run.id)
631  const ids = new Set(loops.map(l => l.id))
632  const going = loops.filter(l => l.ended === null)
633  return { name: run.name, loops: loops.length, calls: state.rows.filter(r => ids.has(r.agent)).length, label: going[going.length - 1]?.label ?? null }
634}
635
636// The first state that applies wins: a dead turn, a run in flight, then cards waiting, then a saving to show off.
637const bandState = (state: State, died: BandModel['died']): BandModel['state'] => {
638  if (died !== null) return 'died'
639  if (state.judge.running) return 'checking'
640  if (state.cards.length > 0) return 'found'
641  if (state.saved.chars > 0 || state.saved.ms > 0) return 'saved'
642  return 'watching'
643}
644
645// The latest clock the state has seen: a turn's completion or a run's launch. Without a clock a loopless run
646// would read as fresh forever, so the band takes the newest reading it holds when the caller has none.
647const clockOf = (state: State): number =>
648  Math.max(0, ...state.turns.map(t => t.at), ...state.runs.map(r => r.at))
649
650/** Builds the one teaser line the band shows above the prompt; `now` dates a loopless run's freshness. */
651export const bandModel = (state: State, now: number = clockOf(state)): BandModel => {
652  const cost = waitingCost(state)
653  const died = diedOf(state)
654  return {
655    state: bandState(state, died),
656    died,
657    running: runningOf(state, now),
658    fresh: state.cards.length,
659    costPct: pctOf(cost.chars, state.usage.window),
660    costMs: cost.ms,
661    savedPct: pctOf(state.saved.chars, state.usage.window),
662    savedMs: state.saved.ms,
663    calls: state.rows.length,
664    paneOpen: state.paneOpen,
665  }
666}
667
668const isFilled = (v: unknown): v is string => typeof v === 'string' && v.trim().length > 0
669
670const isText = (v: unknown, max: number): v is string => isFilled(v) && v.length <= max
671
672const isOneOf = <T extends string>(v: unknown, options: readonly T[]): v is T =>
673  typeof v === 'string' && (options as readonly string[]).includes(v)
674
675const fields = (v: unknown): Record<string, unknown> | null =>
676  v !== null && typeof v === 'object' && !Array.isArray(v) ? (v as Record<string, unknown>) : null
677
678const signatureOf = (v: unknown): Signature | null | undefined => {
679  if (v === null) return null
680  const o = fields(v)
681  const tool = o?.['tool']
682  const key = o?.['key']
683  return isText(tool, KEY_MAX) && isText(key, KEY_MAX) ? { tool, key } : undefined
684}
685
686const proposalOf = (v: unknown): Proposal | null | undefined => {
687  if (v === null) return null
688  const o = fields(v)
689  const kind = o?.['kind']
690  const title = o?.['title']
691  const body = o?.['body']
692  if (!isOneOf(kind, ['claude-md', 'skill', 'agent-brief', 'settings-allow'])) return undefined
693  return isFilled(title) && isFilled(body) ? { kind, title, body } : undefined
694}
695
696const storedOf = (v: unknown): StoredPattern | null => {
697  const o = fields(v)
698  if (o === null) return null
699  const id = o['id']
700  const category = o['category']
701  const kind = o['kind']
702  const why = o['why']
703  const alternative = o['alternative']
704  const confidence = o['confidence']
705  const estTokensPerTurn = o['estTokensPerTurn']
706  const lastDecision = o['lastDecision']
707  const signature = signatureOf(o['signature'])
708  const proposal = proposalOf(o['proposal'])
709  if (typeof id !== 'string' || !/^[a-z-]+:[a-z0-9-]{1,40}$/.test(id)) return null
710  if (!isOneOf(category, ['execution', 'reading', 'production', 'behavior', 'communication', 'multi-agent', 'environment', 'process', 'other'])) return null
711  if (!isText(kind, KIND_MAX) || !isText(alternative, ALTERNATIVE_MAX) || typeof why !== 'string') return null
712  if (typeof confidence !== 'number' || !(confidence >= 0.5) || !(confidence <= 1)) return null
713  if (estTokensPerTurn !== null && !(typeof estTokensPerTurn === 'number' && Number.isFinite(estTokensPerTurn) && estTokensPerTurn >= 0)) return null
714  if (lastDecision !== null && !isOneOf(lastDecision, ['keep', 'steer', 'kill'])) return null
715  if (signature === undefined || proposal === undefined) return null
716  const lean = o['lean'] ?? null
717  if (lean !== null && !isText(lean, ALTERNATIVE_MAX)) return null
718  return { id, category, kind, signature, why, alternative, confidence, proposal, estTokensPerTurn, lastDecision, lean }
719}
720
721/** Reads a stored registry from the plugin store, dropping every entry that does not validate. */
722export const parseRegistry = (value: unknown): StoredPattern[] => {
723  if (!Array.isArray(value)) return []
724  const byId = new Map<string, StoredPattern>()
725  for (const item of value) {
726    const stored = storedOf(item)
727    if (stored !== null) byId.set(stored.id, stored)
728  }
729  return [...byId.values()]
730}
731
732/** Strips a pattern's session fields, leaving what is persisted per project. */
733export const toStored = (p: Pattern): StoredPattern => ({
734  id: p.id,
735  category: p.category,
736  kind: p.kind,
737  signature: p.signature,
738  why: p.why,
739  alternative: p.alternative,
740  confidence: p.confidence,
741  proposal: p.proposal,
742  estTokensPerTurn: p.estTokensPerTurn,
743  lastDecision: p.lastDecision,
744  lean: p.lean,
745})
746
747const QUOTE_MIN = 24   // consecutive characters of a typed ask that make a stored text a quote of it
748
749// Whether `text` carries QUOTE_MIN consecutive characters of any head, case and whitespace aside.
750const quotesAsk = (text: string, heads: readonly string[]): boolean => {
751  const t = collapseWs(text).toLowerCase()
752  return heads.some(h => Array.from({ length: Math.max(0, h.length - QUOTE_MIN + 1) }, (_, i) => h.slice(i, i + QUOTE_MIN)).some(q => t.includes(q)))
753}
754
755/**
756 * What this session persists: every pattern stored, never a typed prompt. A process `why` is not stored at
757 * all, a habit `why` quoting an ask is blanked, and a pattern whose kind, fix or lean quotes one stays in the session.
758 */
759export const storableOf = (state: State): StoredPattern[] => {
760  const heads = state.asks.map(a => collapseWs(a.head).toLowerCase())
761  return state.patterns.map(toStored).flatMap(s =>
762    [s.kind, s.alternative, s.lean ?? ''].some(t => quotesAsk(t, heads))
763      ? []
764      : [{ ...s, why: s.lean !== null || quotesAsk(s.why, heads) ? '' : s.why }])
765}
766
767/** Revives a stored pattern with empty session fields. */
768export const fromStored = (s: StoredPattern): Pattern => ({
769  ...s, hits: [], decision: null, decidedAtTurn: null, instruction: null, openedAtTurn: null, ignored: 0, sent: 0,
770})
771
772// What a stored entry is worth when the registry overflows: a decision outranks any confidence.
773const rankOf = (p: StoredPattern): number => (p.lastDecision === null ? p.confidence : 2 + p.confidence)
774
775const capStored = (entries: readonly StoredPattern[]): StoredPattern[] => {
776  if (entries.length <= MAX_PATTERNS) return [...entries]
777  return [...entries.entries()]
778    .sort(([atA, a], [atB, b]) => rankOf(b) - rankOf(a) || atA - atB)
779    .slice(0, MAX_PATTERNS)
780    .sort(([atA], [atB]) => atA - atB)
781    .map(([, p]) => p)
782}
783
784/** Merges two stored registries by id (`b` wins), dropping the weakest undecided entries past the cap. */
785export const mergeStored = (a: readonly StoredPattern[], b: readonly StoredPattern[]): StoredPattern[] => {
786  const byId = new Map<string, StoredPattern>(a.map(p => [p.id, p]))
787  for (const p of b) byId.set(p.id, p)
788  return capStored([...byId.values()])
789}
790
791const classCounts = (rows: readonly Row[]): string => {
792  const counts = new Map<string, number>()
793  for (const r of rows) counts.set(r.cls, (counts.get(r.cls) ?? 0) + 1)
794  return [...counts.entries()].sort((a, b) => b[1] - a[1]).map(([cls, n]) => `${cls}×${n}`).join(' ') || '(none)'
795}
796
797const oneLine = (text: string | null): string => (text === null ? '-' : `"${text.replace(/\n/g, '\\n').slice(0, 120)}"`)
798
799const patternLine = (p: Pattern): string =>
800  `  ${p.id} · hits ${p.hits.length} [${p.hits.slice(0, 5).join(' ')}] · ${p.decision ?? '-'} @ ${p.decidedAtTurn ?? '-'} · previous ${p.lastDecision ?? '-'} · ignored ${p.ignored} · opened ${p.openedAtTurn ?? '-'} · sent ${oneLine(p.instruction)}`
801
802const patternLines = (state: State): string[] => {
803  const shown = state.patterns.slice(0, DEBUG_MAX_PATTERNS).map(patternLine)
804  const rest = state.patterns.length - shown.length
805  return rest > 0 ? [...shown, `  … ${rest} more patterns`] : shown
806}
807
808/** What one judge fork cost, one line: the same text `/saver debug` and the debug log both print. */
809export const usageLine = (u: JudgeUsage, label = 'judge'): string =>
810  `${label} usage: in ${u.input} · out ${u.output} · cache read ${u.cacheRead} · cache create ${u.cacheCreate}`
811
812// What the last run reported, and why anything it returned never reached the user.
813const judgeRunLines = (run: JudgeRun | null, label = 'judge'): string[] =>
814  run === null
815    ? []
816    : [
817        `${label} last: ${run.returned} returned · ${run.kept} kept · ${run.dropped.length} dropped`,
818        ...run.dropped.slice(0, DEBUG_MAX_DROPPED).map(reason => `  ${reason}`),
819        ...(run.usage === null ? [] : [usageLine(run.usage, label)]),
820      ]
821
822// The audit a load owes a session it joined late: fired at `session.start` over the adopted rows, retried at
823// every warm opportunity while it keeps coming back with nothing, and never armed at all under the row floor.
824const loadCheckLine = (state: State, spoke: boolean): string => {
825  const lane = state.pendingCheck ? 'retrying' : state.judge.runs === 0 ? 'not armed' : 'answered'
826  return `load check: ${lane} · ${spoke ? 'reported' : 'not yet reported'}`
827}
828
829// Where a budget went, as the pane's Time and Context rows no longer spell out: the total and the largest sinks.
830const sinkLine = (label: string, unit: string, budget: Sinks | null): string =>
831  `${label} sinks: ${budget === null ? '-' : [`${budget.total}${unit} total`, ...budget.sinks.map(s => `${s.label} ${s.amount} ×${s.count}`)].join(' · ')}`
832
833// The process judge's cadence and the asks it counts: how many and how many read as pace, never what was typed.
834const processDebugLine = (state: State): string => {
835  const p = state.process
836  return `process runs ${p.runs} · spent ${p.spent} tokens · backoff ${p.backoff} · running ${p.running} · lastAt turn ${p.lastAtTurn} / row ${p.lastAtSeq} / ${p.lastAtMs}ms · error ${p.error ?? '-'} · asks ${state.asks.length} · pace ${state.asks.filter(a => a.pace).length}`
837}
838
839/** Renders the whole state, and whether the load lane has reported yet, for `/saver debug` in ≤ 40 lines. */
840export const debugDump = (state: State, spoke = false): string => {
841  const j = state.judge
842  const u = state.usage
843  const o = state.overhead
844  const share = judgeShare(state)
845  return [
846    `ContextSaver · turn ${state.turn} · seq ${state.seq} · rows ${state.rows.length} · turns ${state.turns.length} · patterns ${state.patterns.length}`,
847    `rows ${classCounts(state.rows)}`,
848    `runs ${state.runs.length} · loops ${state.loops.length} · active ${activeRuns(state, clockOf(state)).length}`,
849    sinkLine('time', 'ms', sinksOf(state, 'ms')),
850    sinkLine('context', 'ch', sinksOf(state, 'chars')),
851    ...patternLines(state),
852    `cards ${state.cards.length}${state.cards.length === 0 ? '' : `: ${state.cards.join(', ')}`}`,
853    `notes ${state.notes.length} · standing ${state.standing.length} · written ${state.written.length}${state.written.length === 0 ? '' : `: ${state.written.join(', ')}`}`,
854    `judge runs ${j.runs} · spent ${j.spent} tokens (${share === null ? '-' : `${share}%`} of the session) · backoff ${j.backoff} · running ${j.running} · lastAt ${j.lastAtTokens} tokens / turn ${j.lastAtTurn} / row ${j.lastAtSeq} / ${j.lastAtMs}ms · error ${j.error ?? '-'} · focus ${oneLine(j.focus)}`,
855    `judge time: ${oneLine(j.time)}`,
856    `judge context: ${oneLine(j.context)}`,
857    processDebugLine(state),
858    ...judgeRunLines(j.last),
859    ...judgeRunLines(state.process.last, 'process'),
860    loadCheckLine(state, spoke),
861    `usage ${u.percent ?? '-'}% · ${u.tokens ?? '-'} / ${u.window} tokens · compactAt ${u.compactAt ?? '-'} · toCompaction ${tokensToCompaction(state) ?? '-'} · turnsLeft ${turnsToCompaction(state) ?? '-'} · session ${totalTokens(state)} new`,
862    `overhead ${o === null ? '-' : `memory ${o.memory} · mcp ${o.mcp} · agents ${o.agents}`}`,
863    `compactions ${state.compactions.length === 0 ? 'none' : state.compactions.join(', ')}`,
864    `pane ${state.paneOpen ? 'open' : 'closed'} · autoOpened ${state.autoOpened} · columns ${state.columns ?? '-'} · expanded ${state.expanded ?? '-'} · steering ${state.steering ?? '-'}`,
865    `saved ${duration(state.saved.ms)} · ~${pctOf(state.saved.chars, u.window)}% · ${state.saved.chars} chars`,
866  ].slice(0, DEBUG_MAX_LINES).join('\n')
867}
868
869/** What `/saver` says of the process judge: its runs and their spend; null before its first run. */
870export const processLine = (state: State): string | null =>
871  state.process.runs === 0 ? null : `process ${state.process.runs} run${state.process.runs === 1 ? '' : 's'} · ${kilo(state.process.spent)} tokens`
872
hooks/core/process.ts 413 lines
1import { agentAliases } from './evidence'
2import { merge } from './judge'
3import { isCheck, isEdit } from './spawns'
4import { cell, isRecord, loopOf, loopTokens, parseObject, proposalOf, slug, str, unique } from './text'
5import {
6  ALTERNATIVE_MAX, DIGEST_MAX_CHARS, KIND_MAX, MAIN_AGENT, MAX_PROCESS_FINDINGS, PROCESS_CLOCK_MS, PROCESS_MIN_CONFIDENCE,
7} from './types'
8import type { Finding, Loop, Outcome, Pattern, ProcessTrigger, Row, Run, State, TurnStat } from './types'
9
10/** A process finding with the cost it claims, clamped to what its cited handles measured. */
11export type ProcessFinding = Finding & { cost: { tokens: number; ms: number } }
12
13/** The process judge prompt (build spec Appendix E, verbatim) with its one placeholder. */
14export const PROCESS_PROMPT = `This message is not a step of the task. The work is paused while you answer one review question: you cannot call tools, and do not continue, retry or resume anything the transcript was doing. Your whole reply is the JSON object described below.
15
16You are reviewing HOW this session is being run, not what it produced. The transcript above is your own: read it for what the user asked for and what was decided. The digest below is the measured shape of the work so far; nothing outside it is evidence for this review.
17
18Answer one question: given what the user asked, how would a lean expert run this work? Then say where this session's process diverges from that, what each divergence has cost so far, and what the one change to the remaining work is. Look at the shape, not at single calls: review and fix rounds stacked on a review that already covers them; the heaviest model on mechanical stages; an agent spawned for a job one command would do; the full check suite after every merge where once per batch would do; the same plan or file re-read over and over; the pace the user complained about.
19
20Prefer silence to a guess. A wrong "lean" suggestion is worse than none: report a divergence only when the digest measures its cost and the lean alternative plainly reaches the same result for less. Anything the user asked for is not a divergence, and a long session is not a wasteful one. At most three findings; fewer and bigger is better, and \`"findings": []\` is a correct and common answer. INTERVENTIONS lists every process finding already raised. The same divergence under another id is the same finding: return its id. One the person ignored is never reported again.
21
22## Evidence
23\`evidence\`: handles the digest prints, copied exactly — \`run:<id>\` (a SHAPE run row), \`agent:<alias>\` (a loop SHAPE lists, as \`agent:a3\`), \`turn:<n>\` (a turn ASKS or CADENCE names). A finding citing any other handle is discarded whole. Cite every run, loop or turn the cost comes from.
24
25## Cost
26\`cost_tokens\` and \`cost_ms\`: what the divergence has cost so far, measured from the cited handles — a run's or a loop's tokens and minutes, a turn's. A claim above what the cited handles total is cut down to it. Never a projection: what it would cost if unchanged belongs in \`why\`.
27
28## Confidence
29\`confidence\` runs 0.75 to 1.0. 0.9+: the digest shows the divergence repeated and costed, and nothing in the transcript asked for it. 0.75-0.9: plain and costed, and no legitimate reason is visible. Below 0.75: say nothing; it is discarded.
30
31## Contract — the shape of your reply, stated once (documentation, not a template to echo)
32\`\`\`json
33{"findings": [{"id": "process:<kebab-slug, at most 40 chars>",
34  "kind": "<one sentence, at most 120 chars: what this session does>",
35  "lean": "<one sentence, at most 200 chars: how a lean expert runs it>",
36  "alternative": "<one imperative sentence, at most 200 chars, addressed to Claude: the one change to the remaining work>",
37  "why": "<one or two sentences: the measured cost, and the reason for it you considered and ruled out>",
38  "evidence": ["run:<id>", "agent:<alias>", "turn:<n>"],
39  "cost_tokens": 0,
40  "cost_ms": 0,
41  "confidence": 0.85,
42  "proposal": null}]}
43\`\`\`
44\`proposal\`: null unless the change should outlive the session. Otherwise \`{"kind","title","body"}\` where body is, per kind: \`claude-md\` one imperative rule line; \`skill\` the workflow as the body of a SKILL.md; \`agent-brief\` the brief, whose first line may be \`model: haiku\` or \`model: sonnet\`.
45
46Reply with one JSON object: first character \`{\`, last character \`}\`, no prose before or after, no code fence.
47
48## PROCESS DIGEST
49{{DIGEST}}
50
51Return the JSON object only.
52`
53
54/** Names this session's loops the way one process run sees them: built once before the prompt, read again when the reply comes back. */
55export const processAliases = (state: State): ReadonlyMap<string, string> => agentAliases(state.rows, state.loops)
56
57// ---- the digest ----------------------------------------------------------------------------------------
58
59type Section = { name: 'asks' | 'shape' | 'cadence' | 'artifacts' | 'interventions'; header: string; lines: string[]; oldestFirst: boolean; pinned: number }   // pinned: leading lines a cut never drops
60
61// The sections cut first when the digest runs past its cap: the old asks, then the small sections; Shape is the last resort.
62const CUT_ORDER: readonly Section['name'][] = ['asks', 'artifacts', 'interventions', 'cadence', 'shape']
63const CUT_MARK_MAX = 24        // room for the `~ ×N lines cut` line a cut section ends with
64const LIST_MAX = 10            // paths, re-reads and gaps each section names before it stops
65const REREAD_MIN = 3           // reads of one path before it is named a re-read
66const GAP_MIN_MS = 60_000      // a wait shorter than a minute before the next prompt is the person reading
67const READ_RANGE = /:\d*-\d*$/
68const MERGE = /\bmerge\b/
69const COMMIT = /\bcommit\b/
70const MEMORY_FILE = /(\/memory\/[^/]+\.md|\/(CLAUDE|MEMORY)\.md)$/
71
72const minutes = (ms: number): string => `${(ms / 60_000).toFixed(1)}m`
73
74const ktok = (tokens: number): string => `${Math.round(tokens / 1000)}k`
75
76// The family word of a model id (`claude-opus-4-1` → `opus`), or the id itself when it names no family.
77const modelShort = (model: string | null): string => model === null ? '?' : (/(opus|sonnet|haiku|fable)/.exec(model)?.[1] ?? model)
78
79// The role a loop plays: its label's first word (`review:C3-r1` → `review`), else its phase, else `agent`.
80const roleOf = (l: Loop): string =>
81  /^[A-Za-z][\w-]*/.exec(l.label ?? '')?.[0]?.toLowerCase() ?? l.phase?.toLowerCase() ?? 'agent'
82
83// `name ×n` counts, largest first then by name.
84const tally = (names: readonly string[]): string => {
85  const counts = names.reduce((m, n) => m.set(n, (m.get(n) ?? 0) + 1), new Map<string, number>())
86  return [...counts].sort((a, b) => b[1] - a[1] || (a[0] < b[0] ? -1 : a[0] > b[0] ? 1 : 0)).map(([n, c]) => `${n} ×${c}`).join(' · ')
87}
88
89const SEVERITIES = ['critical', 'high', 'medium', 'low'] as const
90
91// What a group of loops returned: severities summed across their reviews, reports counted.
92const outcomesCell = (outcomes: readonly (Outcome | null)[]): string => {
93  const found = outcomes.filter((o): o is Extract<Outcome, { kind: 'findings' }> => o?.kind === 'findings')
94  const severities = SEVERITIES.map(s => [s, found.reduce((n, o) => n + o[s], 0)] as const).filter(([, n]) => n > 0).map(([s, n]) => `${n} ${s}`)
95  const reports = outcomes.filter(o => o?.kind === 'report').length
96  const parts = [...(severities.length > 0 ? [severities.join(' ')] : found.length > 0 ? ['0 findings'] : []), ...(reports > 0 ? [`report ×${reports}`] : [])]
97  return parts.length > 0 ? parts.join(' · ') : '-'
98}
99
100const sumOf = (xs: readonly number[]): number => xs.reduce((a, b) => a + b, 0)
101
102// The run a row belongs to: an agent's row by its loop's run; a main-loop row to the newest run launched before it.
103const runOfRow = (state: State, row: Row): string | null => {
104  if (row.agent !== MAIN_AGENT) return state.loops.find(l => l.id === row.agent)?.run ?? null
105  return [...state.runs].reverse().find(r => r.seq <= row.seq)?.id ?? null
106}
107
108const isMerge = (r: Row): boolean => r.cls === 'git' && MERGE.test(r.key)
109
110// Merges, and the merges a check run followed before the next merge.
111const mergeChecks = (rows: readonly Row[]): { merges: number; checked: number } =>
112  [...rows].sort((a, b) => a.seq - b.seq).reduce((acc, r) => {
113    if (isMerge(r)) return { merges: acc.merges + 1, checked: acc.checked, open: true }
114    if (acc.open && isCheck(r)) return { ...acc, checked: acc.checked + 1, open: false }
115    return acc
116  }, { merges: 0, checked: 0, open: false })
117
118const handleOf = (aliases: ReadonlyMap<string, string>, l: Loop): string => `agent:${aliases.get(l.id) ?? l.id}`
119
120// One declared (or seen) phase of a run: its loops, their models, time, tokens, edits and what they returned.
121const phaseLine = (aliases: ReadonlyMap<string, string>, phase: string, loops: readonly Loop[]): string =>
122  [
123    `  ${phase}`, `×${loops.length}`, tally(loops.map(l => modelShort(l.model))) || '-', minutes(sumOf(loops.map(l => l.ms))),
124    `${ktok(sumOf(loops.map(loopTokens)))} tok`, `edits ${sumOf(loops.map(l => l.edits))}`, outcomesCell(loops.map(l => l.outcome)),
125    loops.map(l => handleOf(aliases, l)).join(' ') || '-',
126  ].join(' | ')
127
128const runLines = (state: State, aliases: ReadonlyMap<string, string>, run: Run): string[] => {
129  const loops = state.loops.filter(l => l.run === run.id)
130  const { merges, checked } = mergeChecks(state.rows.filter(r => runOfRow(state, r) === run.id))
131  const phases = unique([...run.phases, ...loops.map(l => l.phase ?? '-')])
132  return [
133    [
134      `run:${run.id}`, run.name, `phases ${run.phases.length > 0 ? run.phases.join(' → ') : '(none declared)'}`, `loops ${loops.length}`,
135      `Σ${minutes(sumOf(loops.map(l => l.ms)))}`, `Σ${ktok(sumOf(loops.map(loopTokens)))} tok`, `edits ${sumOf(loops.map(l => l.edits))}`,
136      `merges ${merges} · checks after ${checked}`, `turn ${run.turn}`,
137    ].join(' | '),
138    ...phases.map(p => phaseLine(aliases, p, loops.filter(l => (l.phase ?? '-') === p))),
139  ]
140}
141
142const directLine = (aliases: ReadonlyMap<string, string>, l: Loop): string =>
143  [`  ${handleOf(aliases, l)}`, l.label ?? '-', modelShort(l.model), minutes(l.ms), `${ktok(loopTokens(l))} tok`, `calls ${l.calls}`, `edits ${l.edits}`, l.ended ?? 'running'].join(' | ')
144
145const shapeLines = (state: State, aliases: ReadonlyMap<string, string>): string[] => {
146  const direct = state.loops.filter(l => l.run === null || !state.runs.some(r => r.id === l.run))
147  return [
148    ...state.runs.flatMap(run => runLines(state, aliases, run)),
149    ...(direct.length > 0 ? [`direct agents: ×${direct.length}`, ...direct.map(l => directLine(aliases, l))] : []),
150    `roles × model: ${tally(state.loops.map(l => `${roleOf(l)} ${modelShort(l.model)}`)) || '(no agents)'}`,
151  ]
152}
153
154const firstAt = (state: State): number | null => state.asks[0]?.at ?? state.runs[0]?.at ?? null
155
156const askLines = (state: State): string[] => {
157  const first = state.asks[0]?.at ?? 0
158  return state.asks.map(a => `turn:${a.turn} | +${Math.round((a.at - first) / 60_000)}m | ${a.head}${a.pace ? ' | pace' : ''}`)
159}
160
161const gapLine = (turns: readonly TurnStat[]): string => {
162  const gaps = [...turns].filter(t => t.idleMs >= GAP_MIN_MS).sort((a, b) => b.idleMs - a.idleMs).slice(0, LIST_MAX)
163  return `gaps: ${gaps.length > 0 ? gaps.map(t => `turn:${t.turn} idle ${Math.round(t.idleMs / 60_000)}m`).join(', ') : 'none'}`
164}
165
166const cadenceLines = (state: State): string[] => {
167  const pace = state.asks.filter(a => a.pace)
168  const { merges, checked } = mergeChecks(state.rows)
169  const last = state.turns[state.turns.length - 1]?.at ?? 0
170  const since = firstAt(state)
171  return [
172    [
173      `turns ${state.turns.length}`,
174      ...(since !== null && last > since ? [`span ${minutes(last - since)}`] : []),
175      `compactions ${state.compactions.length}${state.compactions.length > 0 ? ` (turn ${state.compactions.join(', ')})` : ''}`,
176      `pace complaints ${pace.length}${pace.length > 0 ? ` (${pace.map(a => `turn:${a.turn}`).join(', ')})` : ''}`,
177    ].join(' · '),
178    `checks ${state.rows.filter(isCheck).length} · merges ${merges} · checks after a merge ${checked} · commits ${state.rows.filter(r => r.cls === 'git' && COMMIT.test(r.key)).length}`,
179    gapLine(state.turns),
180  ]
181}
182
183// The path a file row names: a Read key without its slice, else the paths it changed.
184const pathsOf = (r: Row): string[] => (r.tool === 'Read' ? [r.key.replace(READ_RANGE, '')] : r.paths)
185
186const topCounts = (entries: readonly string[]): [string, number][] =>
187  [...entries.reduce((m, p) => m.set(p, (m.get(p) ?? 0) + 1), new Map<string, number>())].sort((a, b) => b[1] - a[1]).slice(0, LIST_MAX)
188
189const artifactLines = (state: State): string[] => {
190  const edits = state.rows.filter(isEdit)
191  const written = edits.flatMap(r => r.paths)
192  const reads = state.rows.filter(r => r.tool === 'Read')
193  const reread = topCounts(reads.map(r => r.key.replace(READ_RANGE, ''))).filter(([, n]) => n >= REREAD_MIN)
194  const loopsOn = (path: string): number => unique(reads.filter(r => r.key.replace(READ_RANGE, '') === path).map(r => r.agent)).length
195  // Memory files by name and path only: what they say is never read.
196  const memory = unique(state.rows.flatMap(pathsOf).filter(p => MEMORY_FILE.test(p))).map(p =>
197    `memory: ${p.slice(p.lastIndexOf('/') + 1)} ${p} ${written.includes(p) ? 'written' : 'read'}`)
198  return [
199    `edits ${edits.length} on ${unique(written).length} paths · agents' own edits ${sumOf(state.loops.map(l => l.edits))}`,
200    ...(written.length > 0 ? [`written: ${topCounts(written).map(([p, n]) => `${p} ×${n}`).join(' · ')}`] : []),
201    ...(reread.length > 0 ? [`re-read: ${reread.map(([p, n]) => `${p} ×${n} by ${loopsOn(p)} loops`).join(' · ')}`] : []),
202    ...memory,
203  ]
204}
205
206// Every process finding already raised, open ones too, in its own words: without them the next run mints the
207// same divergence under a fresh id, and a second card undoes the person's Ignore.
208const interventionLines = (state: State): string[] => {
209  const listed = state.patterns.filter(p => p.decision !== null || p.category === 'process')
210  return [
211    ...listed.map(p => [
212      p.id,
213      p.decision === null ? 'open' : `${p.decision} @ ${p.decidedAtTurn ?? '-'}`,
214      `sent ×${p.sent}${p.ignored > 0 ? ` | ignored ×${p.ignored}` : ''}`,
215      ...(p.category === 'process' ? [p.kind] : []),
216    ].join(' | ')),
217    `cards waiting ${state.cards.length} · standing instructions ${state.standing.length} · one-time notes pending ${state.notes.length} · subagents reached ${state.delivery.agents.length}`,
218  ]
219}
220
221const render = (sections: readonly Section[]): string =>
222  sections.map(s => [s.header, ...(s.lines.length > 0 ? s.lines : ['(none)'])].join('\n')).join('\n\n')
223
224// A section with enough lines dropped to give back `over` characters: the oldest first (after the pinned) where order is time, else the last.
225const cut = (s: Section, over: number): Section => {
226  const pinned = s.lines.slice(0, s.pinned)
227  const order = s.oldestFirst ? s.lines.slice(s.pinned) : s.lines.slice(s.pinned).reverse()
228  let removed = 0
229  let n = 0
230  while (n < order.length && removed < over + CUT_MARK_MAX) { removed += (order[n]?.length ?? 0) + 1; n += 1 }
231  const rest = order.slice(n)
232  const mark = `~ ×${n} lines cut`
233  return { ...s, lines: s.oldestFirst ? [...pinned, mark, ...rest] : [...pinned, ...rest.reverse(), mark] }
234}
235
236/**
237 * The PROCESS digest: what the person typed, the shape of the spawned work, the cadence, the artifacts and the
238 * interventions, at most DIGEST_MAX_CHARS; over the cap the oldest asks (never the first) go first, then the small sections, and Shape last.
239 *
240 * @param state the session so far
241 * @param aliases the loop naming the prompt is built with, so a cited `agent:<alias>` reads back to its loop
242 */
243export const digestOf = (state: State, aliases: ReadonlyMap<string, string> = processAliases(state)): string => {
244  const sections: Section[] = [
245    { name: 'asks', header: '## ASKS — what the person typed, oldest first: `turn:<n> | +<min>m since the first | what they typed`, then `| pace` when it complains about pace', lines: askLines(state), oldestFirst: true, pinned: 1 },
246    { name: 'shape', header: '## SHAPE — one row per workflow run: `run:<id> | name | declared phases | loops | Σmin | Σtok | edits | merges · checks after | launch turn`, then per phase `phase | ×loops | models | min | tok | edits | outcomes | agent handles`; then agents outside any run, and every role by model', lines: shapeLines(state, aliases), oldestFirst: true, pinned: 0 },
247    { name: 'cadence', header: '## CADENCE — turns, span, compactions, pace complaints; check runs, merges, merges a check followed, commits; the longest waits for the person', lines: cadenceLines(state), oldestFirst: false, pinned: 0 },
248    { name: 'artifacts', header: '## ARTIFACTS — edits, the paths written most, paths re-read, memory files by name and path', lines: artifactLines(state), oldestFirst: false, pinned: 0 },
249    { name: 'interventions', header: '## INTERVENTIONS — process findings already raised and cards decided `id | open or decision @ turn | sent ×N`, a process finding\'s words after it; then what is waiting and what was sent', lines: interventionLines(state), oldestFirst: false, pinned: 0 },
250  ]
251  const fitted = CUT_ORDER.reduce((all, name) => {
252    const over = render(all).length - DIGEST_MAX_CHARS
253    return over <= 0 ? all : all.map(s => (s.name === name ? cut(s, over) : s))
254  }, sections)
255  return render(fitted).slice(0, DIGEST_MAX_CHARS)
256}
257
258// ---- triggers ------------------------------------------------------------------------------------------
259
260// Complaints about pace, as the person types them; lowercased first.
261const PACE: readonly RegExp[] = [
262  /\b(taking|takes|took|take) (so |way |far )?(forever|ages)\b/,
263  /\b(taking|takes|took) (so |way |far )?too long\b/,
264  /\b(this|it|that)('s| is| was)? (taking )?(way |far )?too long\b/,
265  /\bwhy (is (this|it) )?(so |this )?slow\b/,
266  /\b(so|too) slow\b/,
267  /\bhurry( up)?\b/,
268  /\bspeed (this|it|things) up\b/,
269  /\bbeen ages\b/,
270  /\bwhat('s| is) taking so long\b/,
271  /\bstill (not done|waiting)\b/,
272  /\b(take|takes|taking|took) (you |it |this )?so long\b/,
273  /\b(should|ought to) (have been|be) (much |way |a lot )?faster\b/,
274  /\b(expensive|pricey|costly) and slow\b/,
275  /\bslow and (expensive|pricey|costly)\b/,
276]
277
278/** True when a typed prompt reads as a complaint about pace: 'this is taking forever', 'why so slow'. */
279export const isPace = (text: string): boolean => {
280  const lower = text.toLowerCase()
281  return PACE.some(re => re.test(lower))
282}
283
284/**
285 * True when a process run may start: none is running, something happened since the last one, and the clock
286 * has waited PROCESS_CLOCK_MS × backoff — the clock always, an event trigger only once over the budget share.
287 *
288 * @param state the session so far
289 * @param now the clock
290 * @param trigger what asked for the run
291 */
292export const shouldProcess = (state: State, now: number, trigger: ProcessTrigger): boolean => {
293  const p = state.process
294  if (p.running || state.seq <= p.lastAtSeq) return false
295  const since = p.lastAtMs > 0 ? p.lastAtMs : (firstAt(state) ?? now)
296  const waited = now - since >= PROCESS_CLOCK_MS * p.backoff
297  return trigger === 'clock' || p.backoff > 1 ? waited : true
298}
299
300/** Fills the process prompt with this session's digest; `aliases` is the naming the digest prints, for `parseProcessReply`. */
301export const buildProcessPrompt = (state: State, aliases: ReadonlyMap<string, string> = processAliases(state)): string =>
302  PROCESS_PROMPT.split('{{DIGEST}}').join(digestOf(state, aliases))
303
304// ---- the reply -----------------------------------------------------------------------------------------
305
306const RUN_HANDLE = /^run:(.+)$/
307const TURN_HANDLE = /^turn:(\d+)$/
308const AGENT = /^agent:(.+)$/
309
310/** What the cited handles measured: the tokens and wall time of every loop they name (a run's loops, each once) and of every turn. */
311export const measuredCost = (state: State, handles: readonly string[]): { tokens: number; ms: number } => {
312  const runs = handles.flatMap(h => RUN_HANDLE.exec(h)?.[1] ?? [])
313  const agents = handles.flatMap(h => AGENT.exec(h)?.[1] ?? [])
314  const turns = handles.flatMap(h => { const n = TURN_HANDLE.exec(h)?.[1]; return n === undefined ? [] : [Number(n)] })
315  const loops = state.loops.filter(l => agents.includes(l.id) || (l.run !== null && runs.includes(l.run)))
316  const cited = state.turns.filter(t => turns.includes(t.turn))
317  return {
318    tokens: sumOf(loops.map(loopTokens)) + sumOf(cited.map(t => t.input + t.cacheCreate + t.output)),
319    ms: sumOf(loops.map(l => l.ms)) + sumOf(cited.map(t => t.ms)),
320  }
321}
322
323// A handle the digest printed, stored the way the registry keeps it (`agent:<agentId>`), or null.
324const processHandle = (value: unknown, state: State, aliases: ReadonlyMap<string, string>): string | null => {
325  const handle = str(value)
326  const turn = TURN_HANDLE.exec(handle)
327  // The digest prints completed turns (CADENCE) and every ask's turn (ASKS), the one still running included.
328  if (turn !== null) {
329    const n = Number(turn[1])
330    return state.turns.some(t => t.turn === n) || state.asks.some(a => a.turn === n) ? `turn:${n}` : null
331  }
332  const run = RUN_HANDLE.exec(handle)
333  if (run !== null) return state.runs.some(r => r.id === run[1]) ? handle : null
334  const agent = AGENT.exec(handle)
335  const loop = agent === null ? undefined : loopOf(state, aliases, agent[1] ?? '')
336  return loop !== undefined ? `agent:${loop.id}` : null
337}
338
339const evidenceOf = (value: unknown, state: State, aliases: ReadonlyMap<string, string>): string[] | string => {
340  if (!Array.isArray(value) || value.length === 0) return 'evidence must be a non-empty array of handles'
341  const handles = value.map(h => processHandle(h, state, aliases))
342  const badAt = handles.indexOf(null)
343  if (badAt >= 0) return `evidence ${cell(value[badAt], 40) || '(not a string)'} is not a run, agent or turn the digest printed`
344  return unique(handles.filter((h): h is string => h !== null))
345}
346
347const idOf = (value: unknown): string => {
348  const given = str(value)
349  const s = slug(given.slice(given.indexOf(':') + 1))
350  return s.length > 0 ? `process:${s}` : ''
351}
352
353const claimed = (value: unknown): number => (typeof value === 'number' && Number.isFinite(value) ? Math.max(0, Math.round(value)) : 0)
354
355const textDrop = (name: string, text: string, max: number): string | null =>
356  text.length === 0 ? `${name} is empty` : text.length > max * 2 ? `${name} is ${text.length} chars, over twice ${max}` : null
357
358// A string is the one short reason the finding was dropped; the object is the finding itself.
359const findingOf = (value: unknown, state: State, aliases: ReadonlyMap<string, string>): ProcessFinding | string => {
360  if (!isRecord(value)) return 'not an object'
361  const id = idOf(value['id'])
362  const kind = str(value['kind']).trim()
363  const lean = str(value['lean']).trim()
364  const alternative = str(value['alternative']).trim()
365  const confidence = value['confidence']
366  if (id.length === 0) return 'id is missing'
367  const bad = textDrop('kind', kind, KIND_MAX) ?? textDrop('lean', lean, ALTERNATIVE_MAX) ?? textDrop('alternative', alternative, ALTERNATIVE_MAX)
368  if (bad !== null) return bad
369  if (typeof confidence !== 'number' || confidence > 1) return 'confidence is not a number up to 1'
370  if (confidence < PROCESS_MIN_CONFIDENCE) return `confidence ${confidence} is below ${PROCESS_MIN_CONFIDENCE}`
371  if (state.patterns.some(p => p.id === id && p.decision === 'keep')) return 'kept this session'
372  const evidence = evidenceOf(value['evidence'], state, aliases)
373  if (typeof evidence === 'string') return evidence
374  const measured = measuredCost(state, evidence)
375  return {
376    id, category: 'process', kind, evidence, signature: null, why: str(value['why']).trim(), alternative, confidence,
377    estTokensPerTurn: null, proposal: proposalOf(value['proposal']), lean,
378    cost: { tokens: Math.min(claimed(value['cost_tokens']), measured.tokens), ms: Math.min(claimed(value['cost_ms']), measured.ms) },
379  }
380}
381
382type Sifted = { findings: ProcessFinding[]; dropped: string[] }
383
384/** What the process judge said: at most MAX_PROCESS_FINDINGS valid findings, their cost clamped to the cited handles, and one line per drop; never throws. `aliases` must be the table `buildProcessPrompt` printed. */
385export const parseProcessReply = (text: string, state: State, aliases: ReadonlyMap<string, string> = processAliases(state)): { findings: ProcessFinding[]; dropped: string[]; returned: number } => {
386  const root = parseObject(text)
387  if (root === null) return { findings: [], dropped: ['reply was not JSON'], returned: 0 }
388  const raw = root['findings']
389  if (!Array.isArray(raw)) return { findings: [], dropped: ['findings was not an array'], returned: 0 }
390  const sifted = raw.reduce<Sifted>((kept, value, i) => {
391    const label = idOf(isRecord(value) ? value['id'] : '') || `#${i + 1}`
392    const dropped = (reason: string): Sifted => ({ findings: kept.findings, dropped: [...kept.dropped, `${label}: ${reason}`] })
393    const f = findingOf(value, state, aliases)
394    if (typeof f === 'string') return dropped(f)
395    if (kept.findings.some(k => k.id === f.id)) return dropped('the reply already reported this id')
396    if (kept.findings.length >= MAX_PROCESS_FINDINGS) return dropped(`over MAX_PROCESS_FINDINGS (${MAX_PROCESS_FINDINGS})`)
397    return { findings: [...kept.findings, f], dropped: kept.dropped }
398  }, { findings: [], dropped: [] })
399  return { ...sifted, returned: raw.length }
400}
401
402/** Folds process findings into the registry as the habit judge's `merge` does, the newest run's `kind`, `lean` and clamped `cost` kept; names the fresh, the recurred and the ids the cap evicted, for `process.done`. */
403export const mergeProcess = (state: State, findings: readonly ProcessFinding[]): { patterns: Pattern[]; fresh: string[]; recurred: string[]; evicted: string[] } => {
404  const merged = merge(state, [...findings])
405  return {
406    ...merged,
407    patterns: merged.patterns.map(p => {
408      const f = findings.find(x => x.id === p.id)
409      return f !== undefined ? { ...p, kind: f.kind, lean: f.lean, cost: f.cost } : p
410    }),
411  }
412}
413
hooks/core/rules.ts 119 lines
1import { baseline } from './evidence'
2import { collapseWs, pctOf, slug } from './text'
3import { BRIEF_TOOLS, CLAUDE_MD_HEADING } from './types'
4import type { Artifact, ArtifactKind, Pattern, Proposal, State } from './types'
5
6/** Renders one single-line CLAUDE.md bullet; the shell appends only this when the heading already exists. */
7export const bulletOf = (body: string): string => `- ${collapseWs(body).replace(/^[-*] /, '')}\n`
8
9/** Returns the bullet alone from a claude-md artifact's content, for appending under an existing heading. */
10export const bulletOnly = (content: string): string => content.slice(Math.max(0, content.indexOf('\n- ') + 1))
11
12/** Appends an artifact's content to a file's text: the bullet alone under an existing heading, never glued to a line. */
13export const appendedTo = (existing: string | null, content: string): string => {
14  const base = existing ?? ''
15  const glue = base === '' || base.endsWith('\n') ? '' : '\n'
16  if (base.includes(CLAUDE_MD_HEADING)) return base + glue + bulletOnly(content)
17  // A brand-new file opens on the heading itself; an existing one keeps the blank line before it.
18  return base + glue + (base === '' ? content.replace(/^\n+/, '') : content)
19}
20
21/** Renders the file an artifact kind writes: where it goes, what it says, how it lands. */
22export const render = (kind: ArtifactKind, p: Proposal, cwd: string): Pick<Artifact, 'path' | 'content' | 'mode'> => {
23  if (kind === 'skill') return { path: `${cwd}/.claude/skills/${slug(p.title)}/SKILL.md`, content: skillDoc(p), mode: 'write' }
24  if (kind === 'agent-brief') return { path: `${cwd}/.claude/agents/${slug(p.title)}.md`, content: briefDoc(p), mode: 'write' }
25  if (kind === 'settings-allow') return { path: `${cwd}/.claude/settings.json`, content: p.body, mode: 'merge-settings' }
26  return { path: `${cwd}/CLAUDE.md`, content: `\n${CLAUDE_MD_HEADING}\n${bulletOf(p.body)}`, mode: 'append' }
27}
28
29/** Returns the artifacts that make this session's decisions permanent, largest saving first. */
30export const propose = (state: State): Artifact[] =>
31  state.patterns.flatMap(p => artifactsOf(state, p)).sort((a, b) => b.savingPct - a.savingPct)
32
33/** Adds a permission rule to permissions.allow, keeping every other setting; idempotent. */
34export const mergeSettings = (existing: string | null, rule: string): string => {
35  const root = objectOf(existing)
36  const permissions = recordOf(root['permissions'])
37  const allow = Array.isArray(permissions['allow']) ? permissions['allow'] : []
38  const merged = { ...root, permissions: { ...permissions, allow: allow.includes(rule) ? allow : [...allow, rule] } }
39  return `${JSON.stringify(merged, null, 2)}\n`
40}
41
42const artifactsOf = (state: State, p: Pattern): Artifact[] => {
43  const proposal = proposalOf(p)
44  // An artifact written, tried or skipped this session is done: the pane never offers it twice.
45  if (proposal === null || state.written.includes(`${p.id}:${proposal.kind}`)) return []
46  return [{
47    patternId: p.id,
48    kind: proposal.kind,
49    title: proposal.title,
50    savingPct: pctOf(baseline(state, p).chars * 3, state.usage.window),
51    ...render(proposal.kind, proposal, state.cwd),
52  }]
53}
54
55// Two decisions where D2 wins over 5.4's wording: an Ignore is final for the session, so it proposes nothing even
56// when the judge attached a proposal; and a kill's permanent line is the scoped alternative, never its kill prompt.
57const proposalOf = (p: Pattern): Proposal | null => {
58  if (p.decision === null || p.decision === 'keep') return null
59  // A process finding has no permission to grant, and its lasting line is the lean way, not the one-time re-plan.
60  if (p.proposal !== null && (p.lean === null || p.proposal.kind !== 'settings-allow')) return p.proposal
61  const body = p.lean ?? (p.decision === 'kill' ? p.alternative : (p.instruction ?? p.alternative))
62  // The label is the rule, never the waste: a row offering `Write` reads as what would be written.
63  return { kind: 'claude-md', title: titleOf(body) || p.id, body }
64}
65
66// The rule's first clause as a label: what it tells Claude to do, capitalised and without its full stop.
67const titleOf = (body: string): string => {
68  const s = collapseWs(body).split(/[;,]/)[0]?.replace(/\.$/, '') ?? ''
69  return s.charAt(0).toUpperCase() + s.slice(1)
70}
71
72const skillDoc = (p: Proposal): string =>
73  `${frontmatter([field('name', slug(p.title)), field('description', p.title)])}${prose(p.body)}`
74
75const briefDoc = (p: Proposal): string => {
76  const lines = p.body.split('\n')
77  const head = leading(lines)
78  const rest = lines.slice(head.length)
79  const model = valueOf(head, 'model')
80  const tools = valueOf(head, 'tools') ?? BRIEF_TOOLS
81  const fields = [field('name', slug(p.title)), field('description', p.title), ...(model === null ? [] : [field('model', model)]), field('tools', tools)]
82  return `${frontmatter(fields)}${prose([...without(without(head, 'model'), 'tools'), ...rest].join('\n'))}`
83}
84
85const leading = (lines: string[]): string[] => {
86  const blank = lines.findIndex(l => l.trim() === '')
87  return blank === -1 ? lines : lines.slice(0, blank)
88}
89
90const frontmatter = (fields: string[]): string => `---\n${fields.join('\n')}\n---\n\n`
91
92const field = (key: string, value: string): string => `${key}: ${JSON.stringify(collapseWs(value))}`
93
94const prose = (body: string): string => `${body.trim()}\n`
95
96const isField = (line: string, key: string): boolean => line.trim().toLowerCase().startsWith(`${key}:`)
97
98const valueOf = (lines: string[], key: string): string | null => {
99  const line = lines.find(l => isField(l, key))
100  return line === undefined ? null : collapseWs(line.trim().slice(key.length + 1)) || null
101}
102
103const without = (lines: string[], key: string): string[] => {
104  const i = lines.findIndex(l => isField(l, key))
105  return i === -1 ? lines : [...lines.slice(0, i), ...lines.slice(i + 1)]
106}
107
108const objectOf = (text: string | null): Record<string, unknown> => {
109  if (text === null) return {}
110  try {
111    return recordOf(JSON.parse(text))
112  } catch {
113    return {}
114  }
115}
116
117const recordOf = (v: unknown): Record<string, unknown> =>
118  v !== null && typeof v === 'object' && !Array.isArray(v) ? { ...(v as Record<string, unknown>) } : {}
119
hooks/core/spawns.ts 229 lines
1import { RUN_FRESH_MS } from './types'
2import type { CommandClass, JournalEntry, Loop, Outcome, Row, Run, State } from './types'
3
4// The tools whose row is an edit whatever its paths say, the classes that are a check, the tools that are a read.
5const EDIT_TOOLS: readonly string[] = ['Edit', 'Write', 'NotebookEdit']
6const CHECK_CLASSES: readonly CommandClass[] = ['test', 'lint', 'typecheck', 'format', 'build']
7const READ_TOOLS: readonly string[] = ['Read', 'Grep', 'Glob']
8
9type Severity = 'critical' | 'high' | 'medium' | 'low'
10
11const isRecord = (v: unknown): v is Record<string, unknown> =>
12  typeof v === 'object' && v !== null && !Array.isArray(v)
13
14const str = (v: unknown): string | null => (typeof v === 'string' ? v : null)
15
16const isSeverity = (v: unknown): v is Severity =>
17  v === 'critical' || v === 'high' || v === 'medium' || v === 'low'
18
19// A reviewer's findings, counted by severity; a finding with none of the four names is not counted at all.
20const findingsOf = (findings: readonly unknown[]): Outcome =>
21  findings.reduce<Extract<Outcome, { kind: 'findings' }>>((counts, f) => {
22    const severity = isRecord(f) ? f['severity'] : undefined
23    return isSeverity(severity) ? { ...counts, [severity]: counts[severity] + 1 } : counts
24  }, { kind: 'findings', critical: 0, high: 0, medium: 0, low: 0 })
25
26// The words a prose review labels a finding with, under the four severities the outcome counts.
27const PROSE_SEVERITY: Readonly<Record<string, Severity>> = {
28  critical: 'critical', blocker: 'critical', important: 'high', high: 'high', major: 'high',
29  medium: 'medium', moderate: 'medium', minor: 'low', low: 'low', nit: 'low',
30}
31
32// A label, not a word in a sentence: `(Important)`, `[Minor]`, `**Critical**`, `Severity: high`, or `Minor:` opening a line or a bullet.
33const PROSE_LABEL = /\(\s*(\w+)\s*\)|\[\s*(\w+)\s*\]|\*\*(\w+)\*\*|severity:\s*(\w+)|^[\s>*\-\d.)#]*(\w+):/gim
34
35// What follows a label that found nothing under it: `Critical: none`, `**Minor:** 0`, `High: -`.
36const PROSE_NONE = /^[\s*_:]*(none|nothing|0|n\/a|na|—|-|no findings)[\s.,;:!*_]*$/i
37
38// A prose reply's labelled severities, counted; null when it labels none, so it stays a report.
39const proseOf = (text: string): Outcome | null => {
40  const counts = { kind: 'findings' as const, critical: 0, high: 0, medium: 0, low: 0 }
41  let labelled = 0
42  for (const m of text.matchAll(PROSE_LABEL)) {
43    const severity = PROSE_SEVERITY[(m.slice(1).find(g => g !== undefined) ?? '').toLowerCase()]
44    if (severity === undefined) continue
45    if (PROSE_NONE.test(text.slice(m.index + m[0].length).split('\n')[0] ?? '')) continue
46    counts[severity] += 1
47    labelled += 1
48  }
49  return labelled > 0 ? counts : null
50}
51
52// A findings list is counted, a prose review's labels are counted; anything else is a report, sized by its text or by the JSON it would print as.
53const outcomeOf = (result: unknown): Outcome => {
54  if (isRecord(result) && Array.isArray(result['findings'])) return findingsOf(result['findings'])
55  const prose = typeof result === 'string' ? proseOf(result) : null
56  if (prose !== null) return prose
57  return { kind: 'report', chars: typeof result === 'string' ? result.length : (JSON.stringify(result) ?? '').length }
58}
59
60const entryOf = (line: string): JournalEntry | null => {
61  let value: unknown
62  try {
63    value = JSON.parse(line)
64  } catch {
65    return null
66  }
67  if (!isRecord(value)) return null
68  const agentId = str(value['agentId'])
69  if (agentId === null) return null
70  if (value['type'] === 'started') return { kind: 'started', agentId, label: str(value['label']), phase: str(value['phase']) }
71  if (value['type'] === 'result') return { kind: 'result', agentId, outcome: outcomeOf(value['result']) }
72  return null
73}
74
75/** Reads a workflow journal: one JSON object per line, `started` and `result` entries kept, anything else skipped. */
76export const parseJournal = (text: string): JournalEntry[] =>
77  text.split('\n').map(entryOf).filter((entry): entry is JournalEntry => entry !== null)
78
79/** The run a `Workflow` result launched locally — its id, name and transcript dir — or null for anything else. */
80export const runOf = (value: unknown): { id: string; name: string; dir: string | null } | null => {
81  if (!isRecord(value) || value['taskType'] !== 'local_workflow') return null
82  const id = str(value['runId'])
83  const name = str(value['workflowName'])
84  return id === null || name === null ? null : { id, name, dir: str(value['transcriptDir']) }
85}
86
87/** The loop an `Agent` result closed — its id, the description it was given and the model that ran it — or null. */
88export const agentOf = (value: unknown): { agentId: string; description: string; model: string | null } | null => {
89  if (!isRecord(value)) return null
90  const agentId = str(value['agentId'])
91  return agentId === null ? null : { agentId, description: str(value['description']) ?? '', model: str(value['resolvedModel']) }
92}
93
94type Work = Pick<Row, 'tool' | 'cls' | 'paths'>
95
96/** An edit whatever its paths say, or any call that changed a path. */
97export const isEdit = (r: Work): boolean => EDIT_TOOLS.includes(r.tool) || r.paths.length > 0
98
99/** A test, lint, typecheck, format or build run. */
100export const isCheck = (r: Work): boolean => CHECK_CLASSES.includes(r.cls)
101
102/** A read or a search, by tool or by command class. */
103export const isRead = (r: Work): boolean => r.cls === 'read' || r.cls === 'search' || READ_TOOLS.includes(r.tool)
104
105const one = (yes: boolean): number => (yes ? 1 : 0)
106
107/** The loop with one more of its own rows counted: a call, and an edit, a check or a read when the row was one — counted as it lands, so the count outlives the row. */
108export const countRow = (loop: Loop, row: Work): Loop => ({
109  ...loop, calls: loop.calls + 1, edits: loop.edits + one(isEdit(row)), checks: loop.checks + one(isCheck(row)), reads: loop.reads + one(isRead(row)),
110})
111
112/** What a loop's own surviving rows did: how many calls, and how many of them edited, checked or read; the loop's own counters are the whole figure. */
113export const loopStats = (rows: readonly Row[], loop: Loop): { calls: number; edits: number; checks: number; reads: number } => {
114  const own = rows.filter(r => r.agent === loop.id)
115  return { calls: own.length, edits: own.filter(isEdit).length, checks: own.filter(isCheck).length, reads: own.filter(isRead).length }
116}
117
118// A run is going while one of its loops is, or while it is young enough that its first loop may not have reported yet.
119const isActive = (state: State, run: Run, now: number): boolean => {
120  const loops = state.loops.filter(l => l.run === run.id)
121  return loops.length === 0 ? now - run.at < RUN_FRESH_MS : loops.some(l => l.ended === null)
122}
123
124/** The runs still going: one with an unended loop, or one launched less than RUN_FRESH_MS ago with no loop yet. */
125export const activeRuns = (state: State, now: number): Run[] => state.runs.filter(run => isActive(state, run, now))
126
127/** Where a run's journal is written; null for a run that reported no transcript dir. */
128export const journalPath = (run: Run): string | null => (run.dir === null ? null : `${run.dir}/journal.jsonl`)
129
130// One token of a script's source: a string literal's value, or a single punctuation or word character run.
131type Token = { kind: 'string' | 'word' | 'punct'; text: string }
132
133// Reads the source as tokens, comments skipped, stopping at the first string or comment left open.
134const lex = (source: string): Token[] => {
135  const tokens: Token[] = []
136  let i = 0
137  while (i < source.length) {
138    const c = source[i] ?? ''
139    if (/\s/.test(c)) { i += 1; continue }
140    if (c === '/' && source[i + 1] === '/') { const end = source.indexOf('\n', i); i = end < 0 ? source.length : end; continue }
141    if (c === '/' && source[i + 1] === '*') { const end = source.indexOf('*/', i + 2); if (end < 0) return tokens; i = end + 2; continue }
142    if (c === "'" || c === '"' || c === '`') {
143      let j = i + 1
144      let text = ''
145      while (j < source.length && source[j] !== c) {
146        if (source[j] === '\\') j += 1
147        text += source[j] ?? ''
148        j += 1
149      }
150      if (j >= source.length) return tokens
151      tokens.push({ kind: 'string', text })
152      i = j + 1
153      continue
154    }
155    const word = /^[\w$]+/.exec(source.slice(i, i + 200))
156    if (word !== null) { tokens.push({ kind: 'word', text: word[0] }); i += word[0].length; continue }
157    tokens.push({ kind: 'punct', text: c })
158    i += 1
159  }
160  return tokens
161}
162
163const OPEN: Readonly<Record<string, string>> = { '{': '}', '[': ']', '(': ')' }
164
165// The index just past the bracket that closes the one at `at`, or -1 when the literal never closes.
166const closeOf = (tokens: readonly Token[], at: number): number => {
167  const stack: string[] = []
168  for (let i = at; i < tokens.length; i += 1) {
169    const t = tokens[i]
170    if (t === undefined || t.kind !== 'punct') continue
171    const close = OPEN[t.text]
172    if (close !== undefined) stack.push(close)
173    else if (t.text === stack[stack.length - 1]) {
174      stack.pop()
175      if (stack.length === 0) return i + 1
176    }
177  }
178  return -1
179}
180
181// A key of the object at depth one: a word or a quoted name followed by a colon.
182const isKey = (tokens: readonly Token[], i: number, name: string): boolean =>
183  tokens[i]?.text === name && tokens[i]?.kind !== 'punct' && tokens[i + 1]?.text === ':'
184
185// The value token range of `name` inside the object whose `{` is at `at`, or null.
186const valueAt = (tokens: readonly Token[], at: number, end: number, name: string): number | null => {
187  for (let i = at + 1, depth = 0; i < end - 1; i += 1) {
188    const t = tokens[i]
189    if (t?.kind === 'punct' && OPEN[t.text] !== undefined) depth += 1
190    else if (t?.kind === 'punct' && Object.values(OPEN).includes(t.text)) depth -= 1
191    else if (depth === 0 && isKey(tokens, i, name)) return i + 2
192  }
193  return null
194}
195
196// The titles of a phases array literal: a bare string element or an object's `title` string; null when an element is anything else.
197const titlesOf = (tokens: readonly Token[], at: number, end: number): string[] | null => {
198  const titles: string[] = []
199  let i = at + 1
200  while (i < end - 1) {
201    const t = tokens[i]
202    if (t?.kind === 'string') { titles.push(t.text); i += 1 }
203    else if (t?.text === '{') {
204      const close = closeOf(tokens, i)
205      const value = valueAt(tokens, i, close, 'title')
206      const title = value === null ? undefined : tokens[value]
207      if (title?.kind !== 'string') return null
208      titles.push(title.text)
209      i = close
210    } else return null
211    if (tokens[i]?.text === ',') i += 1
212  }
213  return titles
214}
215
216/** The phase titles a Workflow script's `meta` literal declares, read as a literal and never run; [] when there is no readable `meta.phases`. */
217export const phasesOf = (script: string): string[] => {
218  const tokens = lex(script)
219  const meta = tokens.findIndex((t, i) => t.text === 'meta' && tokens[i + 1]?.text === '=' && tokens[i + 2]?.text === '{')
220  if (meta < 0) return []
221  const open = meta + 2
222  const close = closeOf(tokens, open)
223  if (close < 0) return []
224  const value = valueAt(tokens, open, close, 'phases')
225  if (value === null || tokens[value]?.text !== '[') return []
226  const end = closeOf(tokens, value)
227  return end < 0 ? [] : (titlesOf(tokens, value, end) ?? [])
228}
229
hooks/core/text.ts 194 lines
1import type { ArtifactKind, Loop, Proposal, Row, Signature, State, StoredPattern, Usage } from './types'
2
3/** Returns the median of the numbers; 0 when there are none. */
4export const median = (xs: number[]): number => {
5  if (xs.length === 0) return 0
6  const sorted = [...xs].sort((a, b) => a - b)
7  const mid = Math.floor(sorted.length / 2)
8  return sorted.length % 2 !== 0 ? (sorted[mid] ?? 0) : (((sorted[mid - 1] ?? 0) + (sorted[mid] ?? 0)) / 2)
9}
10
11/** Returns chars/4/window*100 rounded to 1 decimal. */
12export const pctOf = (chars: number, window: number): number =>
13  Math.round((chars / 4 / window) * 1000) / 10
14
15/** Returns 100 - percent, or null when percent is undefined. */
16export const pctLeft = (u: Usage): number | null =>
17  u.percent !== undefined ? Math.round((100 - u.percent) * 10) / 10 : null
18
19/** Formats milliseconds as '11m', '3m 50s', or '12s'. */
20export const duration = (ms: number): string => {
21  const s = Math.round(ms / 1000)
22  if (s < 60) return `${s}s`
23  const mins = Math.floor(s / 60)
24  const secs = s % 60
25  return secs === 0 ? `${mins}m` : `${mins}m ${secs}s`
26}
27
28/** Returns what a length of text costs in tokens, four characters to one. */
29export const tokensOf = (chars: number): number => Math.round(chars / 4)
30
31/** Formats a count short and rounded: '9.9k', '41k', '800'. */
32export const kilo = (n: number): string =>
33  n >= 10_000 ? `${Math.round(n / 1000)}k` : n >= 1000 ? `${Math.round(n / 100) / 10}k` : `${n}`
34
35/** Truncates text to `cells` characters, ending it with '…' when it is cut and never a space before it. */
36export const fit = (text: string, cells: number): string =>
37  text.length <= cells ? text : `${text.slice(0, Math.max(0, cells - 1)).trimEnd()}…`
38
39/** Returns a kebab-case slug of at most forty characters. */
40export const slug = (s: string): string =>
41  s
42    .toLowerCase()
43    .replace(/[^a-z0-9]+/g, '-')
44    .replace(/^-+|-+$/g, '')
45    .slice(0, 40)
46
47/** Collapses runs of whitespace (including newlines) to a single space and trims. */
48export const collapseWs = (s: string): string => s.replace(/\s+/g, ' ').trim()
49
50/** Returns JSON with keys sorted, top-level keys in omit removed. */
51export const stableJson = (v: unknown, omit: readonly string[]): string => {
52  const replacer = (_key: string, val: unknown): unknown => {
53    if (val !== null && typeof val === 'object' && !Array.isArray(val)) {
54      const obj = val as Record<string, unknown>
55      return Object.keys(obj)
56        .filter(k => !omit.includes(k))
57        .sort()
58        .reduce<Record<string, unknown>>((acc, k) => { acc[k] = obj[k]; return acc }, {})
59    }
60    return val
61  }
62  return JSON.stringify(v, replacer)
63}
64
65/** Renders a filled/empty gauge of `width` cells for a percentage 0..100. */
66export const gauge = (percent: number, width: number): string => {
67  const filled = Math.round(Math.max(0, Math.min(100, percent)) / 100 * width)
68  return '█'.repeat(filled) + '░'.repeat(width - filled)
69}
70
71/** Wraps a user instruction in the standard ContextSaver prefix. */
72export const instructionOf = (text: string): string =>
73  `Instruction from the user (via ContextSaver): ${text}`
74
75/** Produces the kill prompt for a stored pattern. */
76export const killPrompt = (p: StoredPattern): string =>
77  `Stop this behaviour for the rest of the session: ${p.kind}. From now on: ${p.alternative}`
78
79// Reply parsing shared by the habit and process judges: the JSON object, the evidence handles each may
80// cite, and the grounded ceiling on what a finding may claim.
81
82export const unique = (xs: string[]): string[] => [...new Set(xs)]
83
84export const str = (v: unknown): string => (typeof v === 'string' ? v : '')
85
86export const isRecord = (v: unknown): v is Record<string, unknown> =>
87  typeof v === 'object' && v !== null && !Array.isArray(v)
88
89export const isArtifactKind = (v: unknown): v is ArtifactKind =>
90  typeof v === 'string' && ['claude-md', 'skill', 'agent-brief', 'settings-allow'].includes(v)
91
92export const parseObject = (text: string): Record<string, unknown> | null => {
93  const start = text.indexOf('{')
94  const end = text.lastIndexOf('}')
95  if (start < 0 || end < start) return null
96  try {
97    const value: unknown = JSON.parse(text.slice(start, end + 1))
98    return isRecord(value) ? value : null
99  } catch {
100    return null
101  }
102}
103
104export const short = (text: string, max: number): string => (text.length <= max ? text : `${text.slice(0, max)}…`)
105
106// The model's own text quoted inside a one-line drop reason: whitespace collapsed first, so a
107// heredoc key or a multi-line id can never break the one-reason-per-line contract `debugDump` keeps.
108export const cell = (value: unknown, max: number): string => short(collapseWs(str(value)), max)
109
110export const AGENT_HANDLE = /^agent:(.+)$/
111
112// Handles that name no single row: a turn of the main loop, or a whole agent loop.
113export const isWide = (handle: string): boolean => /^turn:\d+$/.test(handle) || AGENT_HANDLE.test(handle)
114
115export const handleDrop = (value: unknown, signature: Signature | null): string => {
116  const handle = cell(value, 40)
117  if (handle.length === 0) return 'evidence handle is not a string'
118  if (isWide(handle) && signature !== null) return `evidence ${handle} needs signature null`
119  return AGENT_HANDLE.test(handle) ? `evidence ${handle} not in AGENTS` : `evidence ${handle} not in the ledger`
120}
121
122// The loop an `agent:<alias>` handle names, under the alias table AGENTS and LEDGER printed — the one the
123// prompt was built with, not today's: a judge run takes minutes, and a new agent's first row landing in
124// that time (or an old agent's last row falling off ROW_CAP) renumbers every loop the rows never showed.
125export const loopOf = (state: State, aliases: ReadonlyMap<string, string>, alias: string): Loop | undefined =>
126  state.loops.find(l => aliases.get(l.id) === alias)
127
128export const handleOf = (value: unknown, state: State, visible: Row[], aliases: ReadonlyMap<string, string>, signature: Signature | null): string | null => {
129  const handle = str(value)
130  const alias = /^r(\d+)$/.exec(handle)
131  if (alias !== null) {
132    const row = visible.find(r => r.seq === Number(alias[1] ?? ''))
133    return row !== undefined ? row.id : null
134  }
135  if (signature !== null) return null
136  const turn = /^turn:(\d+)$/.exec(handle)
137  if (turn !== null) {
138    const n = Number(turn[1] ?? '')
139    return state.turns.some(t => t.turn === n) ? `turn:${n}` : null
140  }
141  const agent = AGENT_HANDLE.exec(handle)
142  if (agent !== null) {
143    const loop = loopOf(state, aliases, agent[1] ?? '')
144    // Stored by id, not alias: the alias is a naming of the rows in hand and can renumber between runs.
145    return loop !== undefined ? `agent:${loop.id}` : null
146  }
147  return null
148}
149
150// A string is the reason the evidence cannot be used; the array is the handles it maps to.
151export const evidenceOf = (value: unknown, state: State, visible: Row[], aliases: ReadonlyMap<string, string>, signature: Signature | null): string[] | string => {
152  if (!Array.isArray(value) || value.length === 0) return 'evidence must be a non-empty array of row ids'
153  const handles = value.map(h => handleOf(h, state, visible, aliases, signature))
154  const badAt = handles.indexOf(null)
155  if (badAt >= 0) return handleDrop(value[badAt], signature)
156  return unique(handles.filter((h): h is string => h !== null))
157}
158
159export const citedLoops = (state: State, evidence: string[]): Loop[] =>
160  evidence.flatMap(handle => {
161    const agent = AGENT_HANDLE.exec(handle)
162    const loop = agent === null ? undefined : state.loops.find(l => l.id === agent[1])
163    return loop !== undefined ? [loop] : []
164  })
165
166export const citedTurns = (state: State, evidence: string[]): number[] =>
167  evidence.flatMap(handle => {
168    const turn = /^turn:(\d+)$/.exec(handle)
169    if (turn !== null) return [Number(turn[1] ?? '')]
170    if (AGENT_HANDLE.test(handle)) return citedLoops(state, [handle]).map(l => l.firstTurn)
171    const row = state.rows.find(r => r.id === handle)
172    return row !== undefined ? [row.turn] : []
173  })
174
175export const loopTokens = (l: Loop): number => l.tokens.input + l.tokens.cacheCreate + l.tokens.output
176
177// What one avoided loop would have cost when the finding cites loops alone; else the cited turns' answers.
178export const groundedCap = (state: State, evidence: string[]): number => {
179  const loops = citedLoops(state, evidence)
180  if (loops.length > 0 && !evidence.some(h => /^turn:\d+$/.test(h))) return Math.floor(median(loops.map(loopTokens)))
181  const turns = citedTurns(state, evidence)
182  return Math.floor(median(state.turns.filter(t => turns.includes(t.turn)).map(t => t.answerChars)) / 4)
183}
184
185export const proposalOf = (value: unknown): Proposal | null => {
186  if (!isRecord(value)) return null
187  const kind = value['kind']
188  const title = collapseWs(str(value['title']))
189  const body = str(value['body']).trim()
190  if (!isArtifactKind(kind) || title.length === 0 || body.length === 0) return null
191  if (kind === 'settings-allow' && !/^[A-Za-z][A-Za-z0-9_]*\(.+\)$/.test(body)) return null
192  return { kind, title, body }
193}
194
hooks/core/types.ts 286 lines
1import type { Elements } from 'claude-code'
2
3export const PLUGIN_NAME = 'contextsaver'
4export const PANE_ID = 'saver'
5export const PANE_TITLE = 'ContextSaver'
6export const PANE_INLINE_ROWS = 18            // body rows requested when seated inline above the prompt (the compact card is framed)
7export const AUTO_OPEN_MIN_COLUMNS = 144      // unasked opens wait undrawn below this width (d.ts 1943-1945)
8export const STEER_RING_TRIES = 8             // frames the Fix… field's ring is asked for before the composer route is said
9export const STEER_RING_WAIT_MS = 40          // a frame and a little: the shown pane redraws at most thirty times a second
10export const COMMAND = { name: 'saver', description: 'ContextSaver: toggle the pane · check | fix [n] [text] | ignore <n> | debug | reset', argumentHint: '[check | fix [n] [text] | ignore <n> | debug | reset]' } as const
11export const SETTLE_TURNS = 2                 // an instruction not ignored for this many turns is credited
12export const JUDGE_MIN_NEW_TOKENS = 30_000
13export const JUDGE_MIN_TURNS = 3
14export const JUDGE_MIN_ROWS = 8
15export const JUDGE_MAX_BACKOFF = 4
16export const JUDGE_BUDGET_SHARE = 0.03
17export const JUDGE_LEDGER_ROWS = 150          // full rows rendered; older rows are folded into `~` summary lines
18export const MAX_FINDINGS = 6
19export const MAX_BEHAVIORAL_FINDINGS = 3      // findings with signature: null per judge run (agent findings are signature-null too)
20export const MAX_PATTERNS = 50
21export const ROW_CAP = 2000
22export const KIND_MAX = 120
23export const ALTERNATIVE_MAX = 200
24export const KEY_MAX = 200
25export const DEBUG_MAX_LINES = 40             // `/saver debug` ceiling
26export const DEBUG_MAX_PATTERNS = 16          // pattern lines `/saver debug` prints before folding the rest
27export const DEBUG_MAX_DROPPED = 6            // dropped-finding reasons `/saver debug` and the debug log print
28export const BRIEF_TOOLS = 'Read, Grep, Glob' // the tools an agent brief allows when the proposal names none
29export const FILE_TOOLS: readonly string[] = ['Read', 'Edit', 'Write', 'NotebookEdit']   // tools whose ledger key is the path they touched
30export const CLAUDE_MD_HEADING = '## ContextSaver'
31export const RECOVERED_FLAG = 'recovered'     // `Row.flags` marker for a row rebuilt from the transcript: its `ms` is 0 and its agent reads `main`
32export const MAIN_AGENT = 'main'              // `Row.agent` of the main loop; the alias table leaves it as it is
33export const NO_CALLS = 'no tool calls'       // `Evidence.what` of a turn handle: that turn ran none
34export const CARD_EVIDENCE = 3                // cited calls one card's details show, newest first
35export const JUDGE_MIN_NEW_ROWS = 40          // mid-turn cadence: ledger rows since the last run (a turn can last hours)
36export const JUDGE_MIN_GAP_MS = 300_000       // mid-turn cadence: at least five minutes between runs
37export const TREND_TURNS = 10                 // context samples the header's trend draws
38export const SINKS = 3                        // named sinks the Time and Context rows show
39export const LOOP_CAP = 400                   // loops kept (oldest dropped)
40export const AGENTS_ROWS = 60                 // loop lines the AGENTS block renders in full; older ones fold per run
41export const RUN_REFRESH_MS = 10_000          // a running workflow's journal is re-read at most this often
42export const RUN_FRESH_MS = 600_000           // a run with no loop yet counts as active this long after its launch
43export const STORE_KEY_VERSION = 'v5'         // 0.5 starts fresh: a 0.4 `patterns:<cwd>` entry is never read
44export const storeKey = (cwd: string): string => `patterns.${STORE_KEY_VERSION}:${cwd}`
45export const FOLD_LINES = 40                  // `~` lines the LEDGER prints, largest by characters first; the rest become one `~ ×N more` line
46export const JUDGE_AGENT_ROWS = 40            // rows one agent loop may hold in the citable window, so no single loop fills it
47export const ASK_HEAD_MAX = 200               // characters of a typed prompt kept for the process digest; session-only, never stored
48export const DIGEST_MAX_CHARS = 40_000        // the PROCESS digest's ceiling; Shape is never the section cut
49export const PROCESS_CLOCK_MS = 1_800_000     // the process judge's clock: at most every thirty minutes, doubled past the budget share
50export const PROCESS_MIN_CONFIDENCE = 0.75    // a process finding below this is dropped
51export const MAX_PROCESS_FINDINGS = 3         // process findings kept per run
52export const DELIVERY_BURST_MS = 5_000        // subagent deliveries inside this window share one toast
53
54export type CommandClass = 'test' | 'lint' | 'format' | 'typecheck' | 'build' | 'install' | 'git' | 'read' | 'search' | 'other'
55export type Category = 'execution' | 'reading' | 'production' | 'behavior' | 'communication' | 'multi-agent' | 'environment' | 'process' | 'other'
56export type Choice = 'keep' | 'steer' | 'kill'
57export type ProcessTrigger = 'plan' | 'workflow' | 'pace' | 'clock'   // an accepted ExitPlanMode, a Workflow launch, a typed pace complaint, the clock
58/** One prompt the person typed (`composer` or `bridge` origin): the process judge's view of what was asked. Session-only. */
59export type Ask = { turn: number; at: number; head: string; pace: boolean }   // head: ≤ ASK_HEAD_MAX chars, control characters stripped; pace: it reads as a complaint about pace
60
61export type Row = {
62  seq: number              // 1-based, monotonically increasing; the judge sees `r${seq}`
63  id: string               // tool_use_id
64  tool: string; key: string; cls: CommandClass
65  agent: string            // e.agentId ?? 'main'
66  turn: number
67  ms: number; chars: number
68  head: string             // first 80 chars of result.text, control characters stripped; quoted as evidence in the pane, never sent to the judge
69  flags: string[]          // 'err' (tool reported an error) | 'denied' (result.deny: the user or a policy said no) | 'dedup' (Read type 'file_unchanged') | 'trunc' (truncatedByTokenCap) | 'bg' (run_in_background or backgroundTaskId) | 'timeout' (timedOutAfterMs) | `persist=${persistedOutputSize}` | 'ask' (AskUserQuestion: ms is the wait for the person) | 'recommended' (an ask whose questions carry '(Recommended)') | 'recovered' (rebuilt from the transcript at load: ms is 0 and agent reads 'main')
70  lines: { add: number; del: number } | null   // Edit: gitDiff.additions/deletions else counted from structuredPatch; Write: content line count as add
71  paths: string[]          // absolute paths this call edited (Edit/Write filePath unless staged; Bash bashEditDiff.changedFiles)
72  spawn: { type: string; requested: string | null; resolved: string | null; status: string | null; tokens: number | null; edits: number | null; promptChars: number } | null   // Agent rows only
73}
74export type TurnStat = { turn: number; input: number; output: number; cacheRead: number; cacheCreate: number; calls: number; ms: number; answerChars: number; answerHead: string; aborted: boolean; ended: TurnEnd; at: number; idleMs: number; context: number | null }   // answerHead: first 100 chars of e.answer, for evidence quotes; ended: how the turn ended; at: clock at completion; idleMs: wait until the next turn started (0 until it does); context: tokens in the window after the turn (usage), null when unknown — growth between turns is the pace compaction runs at
75export type TurnEnd = 'answer' | 'aborted' | 'refusal' | 'error'
76export type Tokens = { input: number; output: number; cacheRead: number; cacheCreate: number }
77export type Outcome = { kind: 'findings'; critical: number; high: number; medium: number; low: number } | { kind: 'report'; chars: number }
78export type Loop = { id: string; run: string | null; label: string | null; phase: string | null; model: string | null; turns: number; ms: number; tokens: Tokens; ended: TurnEnd | null; firstTurn: number; firstSeq: number; outcome: Outcome | null; calls: number; edits: number; checks: number; reads: number }   // one spawned agent loop: `id` is its agentId (the ledger's `agent`), `run` the workflow run that launched it, `ended` null while it runs; calls/edits/checks/reads count its own rows as they land, so the line stays whole once ROW_CAP drops them
79export type Folded = { tool: string; key: string; cls: CommandClass; agent: string; count: number; ms: number; chars: number; firstTurn: number; lastTurn: number; flags: { ask: number; recommended: number; err: number } }   // one (tool, key) pair's rows dropped past ROW_CAP: nothing citable, everything counted; `agent` is 'main' or the first loop seen
80export type Run = { id: string; name: string; dir: string | null; turn: number; seq: number; at: number; refreshedAt: number; phases: string[] }   // at: clock at launch; refreshedAt: last journal read (0 never); phases: the script's declared `meta.phases` titles, [] when unread or undeclared
81export type JournalEntry = { kind: 'started'; agentId: string; label: string | null; phase: string | null } | { kind: 'result'; agentId: string; outcome: Outcome }
82export type Signature = { tool: string; key: string }
83export type ArtifactKind = 'claude-md' | 'skill' | 'agent-brief' | 'settings-allow'
84export type Proposal = { kind: ArtifactKind; title: string; body: string }
85
86/** Persisted across sessions, per project. */
87export type StoredPattern = {
88  id: string                       // `${category}:${slug}`, slug ≤ 40
89  category: Category
90  kind: string                     // the behaviour, one sentence ≤ 120 chars: "Claude keeps running `bun test` after every step"
91  signature: Signature | null      // null = behavioural, no single command carries it
92  why: string
93  alternative: string              // the fix, one imperative sentence ≤ 200 chars written for Claude; pre-fills Fix…, completes Fix
94  confidence: number               // 0.5..1
95  proposal: Proposal | null
96  estTokensPerTurn: number | null  // judge's estimate for behavioural patterns; null when a signature exists
97  lastDecision: Choice | null      // the most recent session's decision, for the judge's calibration
98  lean: string | null              // process findings only: how a lean expert would run it (the card's Lean line); `kind` is what the session does; null for a habit
99}
100/** Session-only fields. */
101export type Pattern = StoredPattern & {
102  hits: string[]                   // evidence handles: row ids (tool_use_id), `turn:<n>` or `agent:<agentId>`; grows with matching rows; cost/baseline from rows only
103  decision: Choice | null
104  decidedAtTurn: number | null
105  instruction: string | null       // the text sent for steer/kill
106  openedAtTurn: number | null      // turn of the last steer/kill; settles saved (credited or ignored)
107  ignored: number                  // times the instruction was ignored
108  sent: number                     // subagents its standing instruction reached this session (the card's "sent ×N")
109  cost?: { tokens: number; ms: number }   // process only: the newest run's claimed cost, clamped to its cited handles; absent, the card shows what they measured
110}
111
112/** What one fork of the judge cost, in the four token counts the API reports. */
113export type JudgeUsage = { input: number; output: number; cacheRead: number; cacheCreate: number }
114/** What one judge run reported: findings returned, findings kept, one short line per drop and per cap a kept finding missed, and what the fork cost (null when it returned nothing). */
115export type JudgeRun = { returned: number; kept: number; dropped: readonly string[]; usage: JudgeUsage | null }
116
117/** One cited call (or turn) as the details render it: what ran, in which loop, what it cost, what it answered. */
118export type Evidence = {
119  turn: number
120  what: string             // the command for Bash, the path for a file tool, `tool key` otherwise; NO_CALLS for a turn handle
121  agent: string | null     // the loop's alias (a1, a2…), null when it was the main loop's own call
122  ms: number               // 0 when nothing measured it (a turn handle, a recovered row)
123  chars: number            // in-context size of the result, or of the turn's answer
124  head: string             // the first line of what came back, quoted under the call; '' when there is none
125}
126/** One waster as the pane draws it: the behaviour, the stats, the fix, and the receipts behind `i`. */
127export type Card = {
128  patternId: string
129  n: number                // 1-based seat in the pane's list, top to bottom
130  category: Category       // the dim tag on the title row
131  kind: string
132  stats: string
133  why: string
134  fix: string
135  lean: string | null      // process cards: the Lean line; null for a habit card
136  total: { unit: 'calls' | 'turns' | 'agents'; calls: number; ms: number; chars: number }   // cited calls, their wall time and their context; `unit: 'turns'` when the pattern cites turns instead, so `calls` counts turns and `chars` is the per-turn estimate; `unit: 'agents'` when it cites loops, so `calls` counts loops and `ms` is their sum
137  evidence: readonly Evidence[]                          // ≤ CARD_EVIDENCE cited calls, newest first
138}
139export type Artifact = { patternId: string; kind: ArtifactKind; title: string; path: string; content: string; savingPct: number; mode: 'append' | 'write' | 'merge-settings' }
140export type Usage = { tokens?: number; window: number; percent?: number; compactAt?: number }
141
142export type State = {
143  cwd: string
144  turn: number
145  seq: number                      // last Row.seq issued
146  rows: Row[]                      // capped at ROW_CAP (oldest dropped)
147  folded: Record<string, Folded>   // the rows ROW_CAP dropped, folded per `${tool}\t${key}`: STATS, the LEDGER's `~` lines and the sinks count them, so a long session's totals stay whole
148  turns: TurnStat[]
149  loops: Loop[]                    // every agent loop seen, oldest first, capped at LOOP_CAP (oldest dropped)
150  runs: Run[]                      // every Workflow launched this session, oldest first
151  usage: Usage
152  overhead: { memory: number; mcp: number; agents: number } | null
153  compactions: number[]            // turn indices at which session.compact fired
154  patterns: Pattern[]
155  cards: string[]                  // pattern ids awaiting a decision, newest first (the pane's WASTERS list)
156  expanded: string | null          // pattern id whose (i) details are open; one at a time
157  steering: string | null          // pattern id whose Fix… field is open
158  steerDraft: string | null        // what the person has typed so far: every render draws it back into the field, so a redraw for any other reason never wipes it
159  notes: string[]                  // one-shot texts: drained into the next tool result or prompt
160  standing: string[]               // texts re-sent with every prompt this session
161  written: string[]                // `${patternId}:${kind}` of artifacts written, tried or skipped this session; propose() omits them
162  judge: { lastAtTokens: number; lastAtTurn: number; lastAtSeq: number; lastAtMs: number; running: boolean; runs: number; spent: number; backoff: number; error: string | null; focus: string | null; time: string | null; context: string | null; last: JudgeRun | null }   // lastAtSeq/lastAtMs: the mid-turn cadence; time/context: the judge's one-line explanations of where they went
163  asks: Ask[]                      // prompts the person typed, oldest first; session-only
164  process: { lastAtMs: number; lastAtSeq: number; lastAtTurn: number; running: boolean; runs: number; spent: number; backoff: number; error: string | null; last: JudgeRun | null }   // the process judge's own cadence and cost, beside `judge`
165  delivery: { agents: string[]; pending: number; lastToastAt: number }   // agent ids the standing instructions reached; spawns in flight (counted as reached); when the last delivery toast showed
166  pendingCheck: boolean            // a check armed at load and not yet answered: a plugin that joined a session with history fires one there and retries it at every warm opportunity until a run answers (§6)
167  paneOpen: boolean
168  autoOpened: boolean              // the pane auto-opened once this session (like /diff on the first edit)
169  columns: number | null           // last band width seen (e.props.bodyColumns), for the auto-open decision
170  saved: { ms: number; chars: number }
171}
172
173export const initialState = (cwd: string, window: number): State => ({
174  cwd, turn: 0, seq: 0, rows: [], folded: {}, turns: [], loops: [], runs: [], usage: { window }, overhead: null, compactions: [], patterns: [], cards: [], expanded: null, steering: null, steerDraft: null, notes: [], standing: [], written: [],
175  judge: { lastAtTokens: 0, lastAtTurn: 0, lastAtSeq: 0, lastAtMs: 0, running: false, runs: 0, spent: 0, backoff: 1, error: null, focus: null, time: null, context: null, last: null },
176  asks: [], process: { lastAtMs: 0, lastAtSeq: 0, lastAtTurn: 0, running: false, runs: 0, spent: 0, backoff: 1, error: null, last: null }, delivery: { agents: [], pending: 0, lastToastAt: 0 },
177  pendingCheck: false, paneOpen: false, autoOpened: false, columns: null, saved: { ms: 0, chars: 0 },
178})
179
180export type Action =
181  | { type: 'turn.start'; now: number }                        // now: dates the idle wait since the last completed turn
182  | { type: 'loop.turn'; agentId: string; model: string | null; ms: number; tokens: Tokens; ended: TurnEnd; turn: number }   // one turn of an agent loop completed
183  | { type: 'agent.start'; agentId: string; description: string; model: string | null }   // an Agent tool result: names the loop
184  | { type: 'run.start'; run: { id: string; name: string; dir: string | null }; now: number }   // a Workflow tool result: a run launched (an id already present is a resume)
185  | { type: 'run.journal'; runId: string; entries: JournalEntry[]; now: number }   // a run's journal read: stages and outcomes of its loops
186  | { type: 'row'; row: Omit<Row, 'seq'> }
187  | { type: 'adopt'; rows: readonly Omit<Row, 'seq'>[] }      // rows rebuilt from the transcript of a session joined late
188  | { type: 'turn.complete'; stat: Omit<TurnStat, 'turn' | 'calls'> }
189  | { type: 'usage'; usage: Usage; now: number }
190  | { type: 'overhead'; overhead: { memory: number; mcp: number; agents: number } }
191  | { type: 'compact' }
192  | { type: 'expand'; patternId: string | null }              // (i) toggled; null collapses
193  | { type: 'steer.begin'; patternId: string }
194  | { type: 'steer.draft'; text: string }
195  | { type: 'decide'; patternId: string; choice: Choice; text?: string }   // text required for steer
196  | { type: 'judge.start'; now: number; seq: number }        // when and at which ledger row the run began, for the mid-turn cadence
197  | { type: 'judge.done'; patterns: Pattern[]; fresh: string[]; recurred: string[]; focus: string | null; time: string | null; context: string | null; spent: number; error: string | null; returned: number; kept: number; dropped: readonly string[]; usage: JudgeUsage | null }
198  | { type: 'ask'; ask: Ask }                                 // a prompt the person typed
199  | { type: 'run.phases'; runId: string; phases: string[] }   // the run's script declared these phases
200  | { type: 'process.start'; now: number; seq: number; trigger: ProcessTrigger }
201  | { type: 'process.cold'; was: Pick<State['process'], 'lastAtMs' | 'lastAtSeq' | 'lastAtTurn'> }   // the fork had no warm snapshot: the run never happened, so the cadence it spent is handed back
202  | { type: 'process.done'; patterns: Pattern[]; fresh: string[]; recurred: string[]; spent: number; error: string | null; returned: number; kept: number; dropped: readonly string[]; usage: JudgeUsage | null }
203  | { type: 'delivery.pending'; delta: 1 | -1 }               // an agent.spawn rewrite began (1) or finished (-1)
204  | { type: 'delivery.sent'; agentId: string; now: number }   // the standing instructions reached this agent, by spawn rewrite or on its first call
205  | { type: 'delivery.toasted'; now: number }
206  | { type: 'check.arm' }                                      // the ledger a session.start adopted already passes the row floor: judge it there, and again at the next warm opportunity if that run answers nothing
207  | { type: 'notes.drained' }
208  | { type: 'standing.add'; text: string }
209  | { type: 'artifact.done'; patternId: string; kind: ArtifactKind; written: boolean }   // written: true once the rule is handled — written, tried or skipped — and recorded in state.written
210  | { type: 'pane'; open: boolean; auto?: true }
211  | { type: 'columns'; columns: number }
212  | { type: 'reset' }
213
214/** Judge output after validation (section 5.3). */
215export type Finding = { id: string; category: Category; kind: string; evidence: string[]; signature: Signature | null; why: string; alternative: string; confidence: number; estTokensPerTurn: number | null; proposal: Proposal | null; lean: string | null }   // lean: non-null for a process finding
216
217/** UI ↔ shell interface. */
218export type Ui = Pick<Elements['terminal'], 'Box' | 'Text' | 'Button' | 'Input' | 'Raster'>
219export type Site = { bodyColumns: number; maxRows: number }
220export type Actions = {
221  keep(patternId: string): void
222  steer(patternId: string): void          // toggles the Fix… field under the waster's verbs
223  steerDraft(text: string): void          // every keystroke: the state keeps the text and the redraw is what paints it
224  steerSubmit(patternId: string, text: string): void  // Enter in the field, or /saver fix <text>
225  kill(patternId: string): void
226  info(patternId: string): void           // toggles the (i) details
227  togglePane(): void
228  check(): void                           // run the judge now
229  write(a: Artifact): void
230  tryOnce(a: Artifact): void
231  skip(a: Artifact): void
232}
233/** View models: computed by patterns.ts from State, rendered by ui.tsx. Keeps the UI free of state logic. */
234export type Sink = { label: string; amount: number; count: number }   // one named consumer: `tests`, `reads`, `agents`, `git`…; amount in ms (time) or chars (context)
235export type Sinks = { total: number; sinks: readonly Sink[] }         // total over every ledger row of the main loop plus the agents' own rows (Agent spawn rows excluded: they contain their loop's rows); the SINKS largest named
236export type Header = {
237  percent: number | null            // context used, 0..100
238  tokensToCompaction: number | null // exact: threshold - tokens
239  turnsToCompaction: number | null  // estimate at the recent pace: tokens to compaction / median context growth per turn
240  trend: readonly number[]          // context percent after each of the last TREND_TURNS turns, oldest first; [] before the first
241  time: Sinks | null                // where the wall-clock went, from the ledger; null before the first row
242  context: Sinks | null             // where the context went, from the ledger; null before the first row
243  judgeTime: string | null          // the judge's one-line explanation of the time, verbatim; null until it has run
244  judgeContext: string | null       // the judge's one-line explanation of the context
245  judgeRuns: number
246  judgeTokens: number               // tokens the judge has spent this session; the pane's JUDGE row
247  judgeShare: number                // those tokens as a percentage of the session's, 1 decimal; 0 while the session has none; `/saver debug` only
248  judgeRunning: boolean             // a run is in flight: Check now reads `Checking…`, dims, and ignores presses
249  savedPct: number
250  savedMs: number
251}
252export type DecidedRow = {
253  patternId: string
254  choice: Choice
255  kind: string
256  savedPct: number | null   // what one avoided repeat is worth; a rate until `settled`, a credit after it
257  settled: boolean          // the instruction was neither ignored nor still in flight, so the saving is real
258  instruction: string | null   // the sentence the user sent, when it is not the fix the card offered
259  ignored: number
260  sent: number              // subagents its standing instruction reached (the row's "sent ×N")
261}
262export type PaneModel = {
263  header: Header
264  wasters: Card[]                   // undecided patterns, newest first
265  expanded: string | null
266  steering: string | null
267  steerDraft: string | null
268  decided: DecidedRow[]             // newest first
269  artifacts: Artifact[]
270}
271/** The band's one teaser line: which of the four states the session is in, and the figures that state names. */
272export type BandModel = {
273  state: 'died' | 'checking' | 'found' | 'saved' | 'watching'   // the first that applies: the last turn died, a judge run in flight, cards waiting, a saving credited, else watching
274  died: TurnEnd | null     // the last turn's `ended` when it was an error or a refusal and no turn has started since
275  running: { name: string; loops: number; calls: number; label: string | null } | null   // the newest active workflow run: its loops, their rows, the newest unended loop's stage
276  fresh: number            // cards awaiting a decision
277  costPct: number          // what those cards have already cost, as a share of the window
278  costMs: number           // and in wall time
279  savedPct: number
280  savedMs: number
281  calls: number            // ledger rows watched this session
282  paneOpen: boolean
283}
284export type BandProps = { ui: Ui; model: BandModel; site: Site; actions: Actions }
285export type PaneProps = { ui: Ui; model: PaneModel; site: Site; placement: 'dock' | 'inline'; actions: Actions }
286
hooks/host.ts 52 lines
1import type {
2  CommandSpec,
3  ModelForkResult,
4  PaneCloseArgs,
5  PaneOpenArgs,
6  SessionMessage,
7  SessionUsage,
8  SessionUsageArgs,
9  UiFocusArgs,
10  UiFocusResult,
11} from 'claude-code'
12
13/** Host table: one lambda per `$.noun.verb` call, bound in session.start. */
14export type Host = {
15  /** $.clock.now() — current time in milliseconds. */
16  now(): Promise<number>
17  /** $.clock.sleep(ms) — resolve after ms milliseconds. */
18  sleep(ms: number): Promise<void>
19  /** $.ui.invalidate('ui.render') — request a redraw (fire-and-forget). */
20  invalidate(): void
21  /** $.ui.toast(text) — show a transient notification (fire-and-forget). */
22  toast(text: string): void
23  /** $.ui.log(text) — emit a log line (fire-and-forget). */
24  log(text: string): void
25  /** $.ui.open(args) — open the named pane. */
26  openPane(args: PaneOpenArgs): Promise<void>
27  /** $.ui.close(args) — close the named pane. */
28  closePane(args: PaneCloseArgs): Promise<void>
29  /** $.ui.focus(args) — move a site's focus ring onto one of this plugin's elements. */
30  focusElement(args: UiFocusArgs): Promise<UiFocusResult>
31  /** $.command.register(spec) — register a slash command. */
32  registerCommand(spec: CommandSpec): Promise<{ command: string }>
33  /** $.session.usage(args?) — read context window usage. */
34  usage(args?: SessionUsageArgs): Promise<SessionUsage>
35  /** $.session.messages() — read the transcript so far, newest 4096 messages. */
36  messages(): Promise<SessionMessage[]>
37  /** $.store.get(key) — read a value from the plugin store. */
38  storeGet(key: string): Promise<unknown>
39  /** $.store.set(key, v) — write a value to the plugin store. */
40  storeSet(key: string, v: unknown): Promise<void>
41  /** $.model.fork({ prompt }) — run a detached model completion over the session transcript. */
42  fork(prompt: string): Promise<ModelForkResult | null>
43  /** $.fs.read(p) — read a file as a string. */
44  readFile(p: string): Promise<string>
45  /** $.fs.write(p, t) — write a string to a file. */
46  writeFile(p: string, t: string): Promise<void>
47  /** $.fs.exists(p) — check if a path exists. */
48  exists(p: string): Promise<boolean>
49  /** $.env.get('CONTEXTSAVER_DEBUG') — read the debug flag. */
50  debugFlag(): Promise<string | undefined>
51}
52