SLOPSHOPPER

claim-ledger

Checks end-of-turn claims like 'tests pass', 'builds cleanly' and 'I've pushed' against the tool calls that actually ran, and notes each claim that has no…

newguardcommand
★ 2v0.1.0GPL-3.0updated 2026-10-09bloknayrb/claudestuff/plugins/claim-ledger
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · claim-ledger
› fix the failing auth test and add an audit log call ● claim-ledger: claim-ledger: loaded, options {"tellModel":true,"testCommands":["pytest","uv run pytest","npm test","npm run test","npm run test:e2e","pnpm test","yarn test","vitest","jest","npx playwright test","cargo test","go test","claude plugin test"],"buildCommands":["tsc","cargo build","npm run build","npm run typecheck","npm run typecheck:tests","idf.py build"]} ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /claim-ledger ⎿ claim-ledger: Claim Ledger, all sessions: ⎿ claim-ledger: tests 0 flagged, 0 later backed (-), 0 repeated unbacked ⎿ claim-ledger: build 0 flagged, 0 later backed (-), 0 repeated unbacked ⎿ claim-ledger: shipped 0 flagged, 0 later backed (-), 0 repeated unbacked ⎿ claim-ledger: A tests or build flag later backed by a run, with no edit between, may have been true when made: a high rate m ⎿ claim-ledger: This session: 9 tool calls; last code edit 08:53 (cache.ts). ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts
README

Claim Ledger

Checks end-of-turn claims ("tests pass", "builds cleanly", "I've pushed") against the tool calls that actually ran, and puts one line beneath the answer for each claim with no evidence or only weak evidence. It makes no model calls and never blocks, changes or re-runs a tool call.

What it does

Claude's answers are read for three families of claim, and each needs its own evidence:

FamilyExample claimsEvidence that counts
tests"tests pass", "all tests passing", "tests are green"the latest foreground run of a test command after the last code edit, showing it passed
build"builds cleanly", "compiles", "type-checks"the latest foreground run of a build or type-check command after the last code edit, showing it passed
shipped"I committed", "I've pushed", "I merged", "opened PR #N", "the PR is merged"the matching git commit / git push / gh pr merge (or git merge) / gh pr create this turn, showing it succeeded

A run "shows it passed" through its exit code, or, when the exit code is masked (piped into tail, followed by ; echo and so on), through an echo of its own status or a result summary in its output. Quoted text, code spans, negations, hedges ("should pass", "will pass") and questions are not claims. Subagent answers are not checked; tool calls from every loop, subagents included, count as evidence.

What you see

One line beneath the answer for each unbacked claim, at most five per turn and then a count:

Claim Ledger: "all tests pass" has no test run after the last edit (17:39).
Claim Ledger: "the tests pass": the last test run after the last edit failed (17:45).
Claim Ledger: "all tests passing" rests on `npm test ...`, whose exit code is masked and whose output shows no result.
Claim Ledger: "I've pushed" has no git push this turn.

Times are local. A claim sentence is flagged once per evidence state, so restating it over the same evidence adds no second line.

With tellModel on, the same lines are also added to the conversation as a hidden note Claude reads on its next request. The note never starts a turn: it waits until you next send something.

Config

Set through /plugin (the plugin's options):

OptionDefaultMeaning
tellModeltrueAlso add each flag as a hidden note for Claude
testCommandspytest, uv run pytest, npm test, npm run test, npm run test:e2e, pnpm test, yarn test, vitest, jest, npx playwright test, cargo test, go test, claude plugin testCommands whose run backs a tests claim, matched as whole words in Bash and PowerShell commands
buildCommandstsc, cargo build, npm run build, npm run typecheck, npm run typecheck:tests, idf.py buildCommands whose run backs a build claim

/claim-ledger

Prints, across all sessions, how many claims of each family were flagged, how many of those were later backed, and how many were repeated while still unbacked. A tests or build flag is later backed when a passing run follows it with no edit between, so the claim may well have been true when it was made: the later-backed rate is the false-positive measure, and a high rate means the check is noisy. Shipped flags are never later backed, since they are judged on their own turn only. The report also shows this session's tool-call count, the last code edit, the current tests and build verdicts, and the last five flags.

Files

Under ~/.claude/state/mods/claim-ledger/:

  • <session id>.json, the health file: when the mod loaded (loadedAt) and the last hook failure (lastError, or null). Written at session start, when /clear or /resume moves to a new session, and on any failure.
  • trail/<session id>.jsonl, the trail: one JSON object per line, the last 200 events of the session. It records each code edit (edit), each test, build or git run with how it was judged (run: kind, ok, masked, background, basis), a background run judged when its result was read back (read-back), each claim's outcome (claim: backed, fired with its status, or repeated), and each hidden note sent (append, or append-failed). For a background run, the run row's ok is the launch; its result lands in the later read-back row. The trail is written only when something was recorded, so a session with no claim, edit or run has none.

Limits

  • Edits. Any successful Edit, Write or NotebookEdit counts as a code edit, in any folder, except .md files, memory folders (~/.claude/memory/, ~/.claude/projects/*/memory/, ~/.claude/rules/), scratch and temp folders, and a commit-message or PR-body file that git or gh then reads. So an edit in a sibling worktree voids an earlier run, and any other page or data file written outside scratch counts too.
  • Code changed through Bash counts as an edit only when the result lists the file (bashEditDiff, an internal field that may be absent).
  • Masked runs. A run whose exit code is hidden counts only when one of these shows how it ended: an echo of its own status ($? directly after it, ${PIPESTATUS[n]}, $LASTEXITCODE); its output's summary; a later && member's pass summary, for a run that ends its pipeline; or, for a piped tsc that surely ran into filters that keep its lines, the absence of error TS.
  • A status held in a variable before the echo is missed. git push > log 2>&1; ec=$?; echo "push exit $ec" echoes $ec, not $?, so the echo is not read as the push's exit status. Echo $? (or ${PIPESTATUS[n]}) directly.
  • Quiet commits. A quiet commit followed by git log in the same command counts on a sha line that carries the commit's subject, or, when the command does not show the subject, on any sha line with no branch, folder or repository change between the two. So a commit a hook blocked can still pass when its subject repeats the previous commit's, or when its message came from a file.
  • git ls-remote matching HEAD is not counted as push evidence. A push is confirmed by its own output (a..b x -> y, [new branch], set up to track), an echoed exit status, or the result's gitOperation.
  • Shipped claims need a known subject and this turn's op. "The PR is merged" is a claim; "P4 is merged" is not. A push made in an earlier turn does not back "I've pushed" now.
  • Background runs. A background run counts as evidence only when a later tool call reads its result back (its task output, or a file it wrote) and that read shows a status or summary. A background run that is never read back counts for nothing: the claim gets a "background run" line.
  • Task notifications are not read. The completion notification that arrives between turns is not a tool call, so a background run whose result was only reported that way stays weak.
  • Read-back keys are paths as the command names them. A path is resolved from the session root through the command's own cds, with $TEMP, $HOME and assignments made on the line expanded. A log named through a loop variable or an environment variable set elsewhere is matched only when the reader names it the same way.
  • Read-back keys keep .. segments as written. A log written as repo/../push.log and read as push.log (or the other way round) is not matched, so the run stays weak.
  • Temp-file names are shared machine-wide. A read-back matches a log by path; $TEMP/push.log written by another session at the same time would be read as this session's run.
  • Log paths resolve from the session root, not the shell's persisted working directory: relative log names written from two different persisted directories can collide.
  • A push's pre-push suite counts as test evidence only when a read-back of that push shows the runner's summary.
  • Some PowerShell loops can be credited. Runs inside a loop are never strong evidence, but PowerShell's % alias, do { ... } while (...) / do { ... } until (...) and .ForEach({ ... }) are not recognised as loops.
  • Some bash loops can be credited: a loop inside backtick substitution, until bash -c "...", a function called inside a loop, and select.
  • PowerShell loops weaken runs to the end of the command. Braces are not tracked, so a run after a foreach, ForEach-Object, while or for block in the same command is weak too.
  • A git operation the tool result records (gitOperation) can be credited inside a loop when the command itself was not matched: eval "git push" in a loop, or a push segment over 1,000 characters in a loop.
  • cmd /c, cargo +nightly test, node --test and yarn workspace X test are not recognised as runs.
  • A segment over 1,000 characters is not matched for runners or git operations: it yields no run, so a true claim resting on a very long command gets a "no run" note.
  • Not checked: subagent answers, and "verified", "works" and "fixed", which are too vague to match against a run.
  • Stop-hook continuations. Seen live: when a Stop hook blocks the end of a turn and Claude continues, the turn keeps one id and the end-of-turn event fires once, after the continuation, carrying only the newer text. Claims made before the block are still caught, because every response's text is read as it finishes, and their lines appear beneath the final answer. A response with two text blocks was not seen, so how they are joined is untested live.
  • A tool call can start before the response that issued it finishes (38 ms in the one case seen); each response's claims are judged against the evidence as it stood when that response began.
  • Hot reload was not verified live. That the ledger survives a hot reload, and that a hook which throws after a reload fails open (the answer shows, the error goes to the health file), were not checked in a live session, because /reload-plugins did not restart a mod loaded with --plugin-dir. A unit test covers restoring a saved ledger; a hook that throws has no test beyond a failing state write being recorded while the line still shows.
  • The hooks API is early access and changes between Claude Code releases.

Precision

Measured by replaying the detector over 107 real session transcripts (722 turns), split by file into a tuning set and a held-out set; no transcript text is kept here.

  • Detection (claims found that a labeller agreed were claims), held-out: tests 0.93, shipped 0.97 (build had 4 labelled claims, too few to judge), against a 0.85 target.
  • Flag justification (a flag the labeller agreed was unbacked) first came in at 0.14 held-out against a 0.70 target. 8 of the 13 unjustified flags, across both sets (held-out alone had 6), rested on background runs whose result a later call read back, so read-back became evidence, along with idf.py build and a few confirmation forms.
  • Independent re-label of the held-out set after that change, by a separate labeller that had not tuned the detector: detection tests 0.98, shipped 0.99 on 112 claims (0.98, 65 of 66, on the final combined labels); 96% agreement with the earlier labels; 2 spurious notes in 112 claims; read-back credits 11 of 11 correct. Too few held-out flags remained (3 to 5 in total) to judge justification.

Install

/plugin marketplace add bloknayrb/claudestuff
/plugin install claim-ledger@claudestuff-marketplace

Develop

Run claude --plugin-dir plugins/claim-ledger once; it lays down .claude-plugin/types/ (git-ignored). Then:

claude plugin test plugins/claim-ledger
claude plugin validate plugins/claim-ledger
npx -y -p typescript@5 tsc -p plugins/claim-ledger

node plugins/claim-ledger/scripts/precision.mjs <out> <list> replays the detector over a list of transcripts. Its output holds transcript text and must stay out of the repository.

Source 5 files
hooks/register.ts 275 lines
1import type { EngineInterface, Register } from 'claude-code'
2
3import type { Claim, Ledger } from '../types'
4import { findClaims } from './claims'
5import {
6  asOfNow,
7  classify,
8  closeTurn,
9  configOf,
10  DEFAULT_BUILD_COMMANDS,
11  DEFAULT_TEST_COMMANDS,
12  emptyLedger,
13  factsOf,
14  hhmm,
15  listOption,
16  markBacked,
17  noteClaims,
18  noteText,
19  readBack,
20  record,
21  restoreLedger,
22} from './evidence'
23import type { Config, Fired, Ran } from './evidence'
24import { applyFlush, countersOf, reportText, ringOf } from './stats'
25
26const LEDGER = { plugin: 'claim-ledger', key: 'ledger' } as const
27const REPORT_FAILED = 'Claim Ledger: the report failed; see the health file.'
28const TRAIL_CAP = 200
29
30// The module holds the authoritative ledger; $.state mirrors it so a hot reload can restore it.
31let ledger: Ledger = emptyLedger()
32let home: string | null = null
33// %TEMP%: writes under it are not code edits.
34let temp: string | null = null
35let loadedAt = 0
36let lastError: { ts: number; message: string } | null = null
37// The session id the health file was last written under: /clear changes it with no session.start.
38let healthId: string | null = null
39// This session's runs, edits, verdicts and append attempts, one JSON object per line. Starts over at a reload.
40let trail: string[] = []
41let trailDirty = false
42
43function trace(event: Record<string, unknown>): void {
44  trail.push(JSON.stringify(event))
45  if (trail.length > TRAIL_CAP) trail.splice(0, trail.length - TRAIL_CAP)
46  trailDirty = true
47}
48
49const errText = (err: unknown): string => (err instanceof Error ? err.message : String(err))
50const stateDir = (): string | null => (home === null ? null : `${home}/.claude/state/mods/claim-ledger`)
51
52async function homeOf($: EngineInterface): Promise<string | null> {
53  const profile = await $.env.get('USERPROFILE')
54  if (profile !== undefined && profile !== '') return profile.replace(/\\/g, '/')
55  const posix = await $.env.get('HOME')
56  return posix !== undefined && posix !== '' ? posix.replace(/\\/g, '/') : null
57}
58
59async function writeHealth($: EngineInterface): Promise<void> {
60  try {
61    const dir = stateDir()
62    if (dir === null) return
63    const id = await $.session.id()
64    healthId = id
65    await $.fs.write(`${dir}/${id}.json`, `${JSON.stringify({ loadedAt, lastError })}\n`)
66  } catch {
67    // Best effort: a missing heartbeat reads as unknown, never as ok.
68  }
69}
70
71/** The heartbeat under a new session id, at the first main-loop event after /clear. */
72async function followSession($: EngineInterface): Promise<void> {
73  try {
74    if ((await $.session.id()) !== healthId) await writeHealth($)
75  } catch {
76    // Best effort, as writeHealth.
77  }
78}
79
80async function noteFailure($: EngineInterface, where: string, err: unknown): Promise<void> {
81  const message = `${where}: ${errText(err)}`
82  try {
83    $.ui.log(`claim-ledger: ${message}`, { to: 'debug' })
84    lastError = { ts: await $.clock.now(), message }
85  } catch {
86    lastError = { ts: 0, message }
87  }
88  await writeHealth($)
89}
90
91async function writeTrail($: EngineInterface): Promise<void> {
92  if (!trailDirty) return
93  trailDirty = false
94  try {
95    const dir = stateDir()
96    if (dir === null) return
97    await $.fs.write(`${dir}/trail/${await $.session.id()}.jsonl`, `${trail.join('\n')}\n`)
98  } catch (err) {
99    await noteFailure($, 'trail', err)
100  }
101}
102
103async function mirror($: EngineInterface): Promise<void> {
104  try {
105    await $.state.set(LEDGER, ledger)
106  } catch (err) {
107    await noteFailure($, 'state.set', err)
108  }
109}
110
111async function start($: EngineInterface, optionsText: string): Promise<void> {
112  home = await homeOf($)
113  const tempDir = await $.env.get('TEMP')
114  temp = tempDir !== undefined && tempDir !== '' ? tempDir.replace(/\\/g, '/') : null
115  loadedAt = await $.clock.now()
116  ledger = restoreLedger(ledger, (await $.state.get(LEDGER)).value)
117  await $.command.register({ name: 'claim-ledger', description: 'Claim Ledger: how often a flagged claim was later backed by a run' })
118  $.ui.log(`claim-ledger: loaded, options ${optionsText}`, { to: 'debug' })
119  await writeHealth($)
120}
121
122async function flush($: EngineInterface, fired: readonly Fired[], repeated: readonly Claim[], now: number): Promise<void> {
123  const backed = ledger.backed
124  if (fired.length === 0 && repeated.length === 0 && backed.length === 0) return
125  ledger.backed = []
126  try {
127    const updated = applyFlush(countersOf(await $.store.get('counters')), ringOf(await $.store.get('ring')), fired, repeated, backed, now)
128    await $.store.set('counters', updated.counters)
129    await $.store.set('ring', updated.ring)
130  } catch (err) {
131    await noteFailure($, 'store', err)
132  }
133}
134
135async function track($: EngineInterface, input: Readonly<Record<string, unknown>>, ran: Ran, seq: number, ts: number, config: Config): Promise<void> {
136  const facts = factsOf(input, ran)
137  const agentId = typeof input.agentId === 'string' ? input.agentId : null
138  const places = { home, root: await $.session.root(), temp }
139  const classified = classify(facts, config, places)
140  const added = record(ledger, facts, classified, seq, ts, agentId)
141  for (const path of classified.mutations) trace({ ts, ev: 'edit', seq, path })
142  // For a background run, ok is the launch status; a later read-back row carries the run's result.
143  for (const e of added) trace({ ts, ev: 'run', seq, kind: e.kind, ok: e.ok, masked: e.masked, background: e.background, basis: e.basis, agentId, short: e.short })
144  // A background run whose result this call read back is judged now.
145  for (const e of readBack(ledger, facts, config, places)) trace({ ts, ev: 'read-back', seq: e.seq, by: seq, kind: e.kind, ok: e.ok, basis: e.basis, short: e.short })
146  markBacked(ledger)
147  await mirror($)
148}
149
150/** The seam around $.session.append: the trail records each attempt, which a test can read whether or not the kit serves the append. */
151async function deliver($: EngineInterface, text: string, now: number): Promise<void> {
152  trace({ ts: now, ev: 'append', text })
153  try {
154    const appended = await $.session.append({ message: { type: 'user', content: [{ type: 'text', text }] } })
155    if (appended.deny !== undefined) throw new Error(appended.deny)
156  } catch (err) {
157    trace({ ts: now, ev: 'append-failed', message: errText(err) })
158    await noteFailure($, 'session.append', err)
159  }
160}
161
162async function check($: EngineInterface, turnId: string, answer: string, tellModel: boolean): Promise<string[]> {
163  // The final text again, judged now: a fallback for a step the turn.step hook missed. Deduplicated by hash.
164  noteClaims(ledger, turnId, findClaims(answer), asOfNow(ledger))
165  const now = await $.clock.now()
166  const out = closeTurn(ledger, turnId, now, hhmm)
167  for (const c of out.backed) trace({ ts: now, ev: 'claim', turnId, hash: c.hash, phrase: c.phrase, outcome: 'backed' })
168  for (const f of out.fired) trace({ ts: now, ev: 'claim', turnId, hash: f.claim.hash, phrase: f.claim.phrase, outcome: 'fired', status: f.status })
169  for (const c of out.repeated) trace({ ts: now, ev: 'claim', turnId, hash: c.hash, phrase: c.phrase, outcome: 'repeated' })
170  await mirror($)
171  await flush($, out.fired, out.repeated, now)
172  if (tellModel && out.lines.length > 0) await deliver($, noteText(out.lines), now)
173  await writeTrail($)
174  return out.lines
175}
176
177async function clear($: EngineInterface): Promise<void> {
178  await flush($, [], [], await $.clock.now())
179  ledger = emptyLedger()
180  trail = []
181  trailDirty = false
182  // The next session's health file starts clean: an error belongs to the session it happened in.
183  lastError = null
184  await mirror($)
185}
186
187async function reportFor($: EngineInterface): Promise<string> {
188  await flush($, [], [], await $.clock.now())
189  return reportText(countersOf(await $.store.get('counters')), ringOf(await $.store.get('ring')), ledger, hhmm)
190}
191
192export const register: Register = (on, options) => {
193  const tellModel = options.tellModel !== false
194  const config = configOf(listOption(options.testCommands, DEFAULT_TEST_COMMANDS), listOption(options.buildCommands, DEFAULT_BUILD_COMMANDS))
195  const optionsText = JSON.stringify({ tellModel: options.tellModel, testCommands: options.testCommands, buildCommands: options.buildCommands })
196
197  on('session.start', async ($, e, next) => {
198    try {
199      await start($, optionsText)
200    } catch (err) {
201      await noteFailure($, 'session.start', err)
202    }
203    return next(e)
204  })
205
206  on('session.end', async ($, e, next) => {
207    try {
208      // /clear and /resume both carry on under another conversation: every session value goes.
209      // The ending session's trail is written before clear() empties it.
210      await writeTrail($)
211      if (e.reason === 'clear' || e.reason === 'resume') await clear($)
212      else await flush($, [], [], await $.clock.now())
213    } catch (err) {
214      await noteFailure($, 'session.end', err)
215    }
216    return next(e)
217  })
218
219  // All loops. Order is assigned at entry, before any await; the call itself is never changed.
220  on('tool.call', async ($, e, next) => {
221    ledger.seq += 1
222    const seq = ledger.seq
223    const ts = await $.clock.now()
224    const ran = await next(e)
225    try {
226      const input = e as unknown as Readonly<Record<string, unknown>>
227      if (input.agentId === undefined) await followSession($)
228      await track($, input, ran, seq, ts, config)
229    } catch (err) {
230      await noteFailure($, 'tool.call', err)
231    }
232    return ran
233  }).catch(($, e, next) => next(e))
234
235  // Main loop: claims in each response's text, judged against what the ledger knew when the response began.
236  on('turn.step', async function* ($, e, next) {
237    const asOf = asOfNow(ledger)
238    const response = yield* next(e)
239    if (e.agentId === undefined) {
240      try {
241        await followSession($)
242        noteClaims(ledger, e.turnId, findClaims(response.answer), asOf)
243      } catch (err) {
244        await noteFailure($, 'turn.step', err)
245      }
246    }
247    return response
248  })
249
250  // Main loop, answered turns only: flag what is still unbacked, beneath the answer.
251  on('turn.complete', async ($, e, next) => {
252    const done = await next(e)
253    if (e.agentId !== undefined || e.reason !== 'answer') return done
254    try {
255      await followSession($)
256      const lines = await check($, e.turnId, e.answer, tellModel)
257      if (lines.length === 0) return done
258      const above = done.text === e.answer ? '' : `${done.text}\n`
259      return { ...done, text: `${above}${lines.join('\n')}` }
260    } catch (err) {
261      await noteFailure($, 'turn.complete', err)
262      return done
263    }
264  })
265
266  on('command.run', { command: 'claim-ledger' }, async $ => {
267    try {
268      return { text: await reportFor($) }
269    } catch (err) {
270      await noteFailure($, 'command.run', err)
271      return { text: REPORT_FAILED }
272    }
273  }).catch(() => ({ text: REPORT_FAILED }))
274}
275
hooks/claims.ts 129 lines
1import type { Claim, Family, ShipOp } from '../types'
2
3// Pure: no `$`, no runtime imports. scripts/precision.mjs imports this file under Node.
4
5/** `gap`: the match spans a subject and the words before its verb (the present-state forms). */
6type Pattern = { family: Family; op: ShipOp | null; re: RegExp; gap?: boolean }
7
8// First person: "I committed", "I've pushed", "I have just merged".
9const I = String.raw`\bI(?:'ve|\s+have)?\s+(?:just\s+|now\s+|also\s+|already\s+)?`
10// A ship verb already said, joined to the next: "Committed and pushed", "committed, pushed".
11const THEN = String.raw`(?:(?:committed|pushed|force-pushed|merged|squash-merged)(?:\s*,\s*|\s+and\s+))?`
12// Sentence-initial, as a list item or in bold: "Committed.", "- **Merged** into main.".
13const START = String.raw`^(?:[-*]\s+)?(?:\*\*)?`
14const BY_ME = `(?:${START}${THEN}|${I}${THEN})`
15// Present state needs a git noun as its subject, so "rows are merged into the tracker" is not a claim.
16// "Everything", "the tag", "Task 3" and "round 2" are the other subjects real end-of-task claims use, as is
17// "Unit 8c".
18const GIT_NOUN = String.raw`(?<![\w#])(?:PR\s*#?\s*\d+|#\d+|pull\s+requests?|PRs?|branch(?:es)?|commits?|changes|fix(?:es)?|patch(?:es)?|work|everything|tags?|task\s+\d+|round\s+\d+|units?\s+\d+[a-z]?)(?![\w-])`
19
20const shipped = (verb: string): RegExp => new RegExp(`${BY_ME}${verb}\\b`, 'i')
21const present = (verb: string): RegExp => new RegExp(`${GIT_NOUN}[^.;:!?]{0,40}?\\b(?:is|are)\\s+(?:now\\s+|all\\s+|both\\s+)?${verb}\\b`, 'i')
22
23const PATTERNS: readonly Pattern[] = [
24  { family: 'tests', op: null, re: /\b(?:all\s+)?(?:\d+(?:\s*\/\s*\d+)?\s+)?(?:the\s+)?(?:unit\s+|integration\s+|e2e\s+|new\s+)?(?:tests?|test\s+files?)\s+(?:now\s+|still\s+|all\s+)?(?:pass(?:es|ing)?|are\s+(?:all\s+|now\s+)?(?:passing|green)|is\s+(?:now\s+)?(?:passing|green))\b/i },
25  { family: 'tests', op: null, re: /\b(?:test\s+)?suite\s+(?:now\s+|still\s+)?(?:passes|is\s+(?:now\s+)?(?:passing|green))\b/i },
26  { family: 'tests', op: null, re: /\ball\s+green\b/i },
27  { family: 'build', op: null, re: /\b(?:builds|compiles|type-?checks)\s+(?:cleanly|clean|successfully|fine|without\s+(?:errors?|warnings?))\b/i },
28  // "It builds its own stub" describes; it does not claim: no object may follow the verb.
29  { family: 'build', op: null, re: /\b(?:it|everything|the\s+(?:project|code|plugin|mod|crate|package|app|module))\s+(?:now\s+|still\s+)?(?:builds|compiles|type-?checks)\b(?!\s+(?:its|their|his|her|our|my|your|the|a|an|this|that|these|those|on|for|against|into|to|from|with)\b)/i },
30  { family: 'build', op: null, re: /\b(?:the\s+)?(?:build|type-?check|tsc)\s+(?:now\s+|still\s+)?(?:passes|succeeds|is\s+(?:now\s+)?(?:clean|green))\b/i },
31  // "I committed to the plan" is an idiom, not a commit.
32  { family: 'shipped', op: 'commit', re: shipped('committed(?!\\s+to\\b)') },
33  { family: 'shipped', op: 'commit', re: present('committed'), gap: true },
34  // "Committed" is git's word whatever the subject ("The helper scripts are committed on the spike branch"): any subject,
35  // with the hedges and negations in its clause, but not "committed to", and not a person or group ("the team is
36  // committed"). In real answers this was the form the git-noun rule missed most often.
37  { family: 'shipped', op: 'commit', re: /(?<!\b(?:we|they|you|he|she|team|everyone|everybody|people)\s)\b(?:is|are)\s+(?:now\s+|all\s+|both\s+)?committed\b(?!\s+to\b)/i },
38  { family: 'shipped', op: 'push', re: shipped('(?:force-)?pushed') },
39  { family: 'shipped', op: 'push', re: present('(?:force-)?pushed'), gap: true },
40  { family: 'shipped', op: 'merge', re: shipped('(?:squash-)?merged') },
41  { family: 'shipped', op: 'merge', re: present('(?:squash-)?merged'), gap: true },
42  { family: 'shipped', op: 'pr-create', re: new RegExp(`(?:${START}|${I}|\\band\\s+)opened\\s+(?:PR\\s*#\\s*\\d+|(?:a|the)\\s+(?:PR|pull\\s+request))\\b`, 'i') },
43]
44
45// A word here, in the same clause before the match, makes it a negation, a condition, a future, an instruction or a question.
46const HEDGE = /\b(?:not|never|no|nothing|none|neither|nor|nobody|don't|doesn't|didn't|won't|isn't|aren't|wasn't|weren't|haven't|hasn't|if|once|until|unless|when|whether|should|would|could|might|may|must|will|make|makes|making|ensure|ensures|verify|check|confirm|expect|expects|expected|need|needs|want|wait|before|get|getting|keep|so\s+that|to\s+see|what|how)\b/i
47// A word here, in the same clause after the match, negates it: "I've pushed nothing yet".
48const NEG_AFTER = /\b(?:yet|nothing|not)\b/i
49// After "merged": "into" a thing that is no branch. A branch is main, master, develop, trunk, a base, release or upstream
50// branch, an inline-code name (`CODE`), a slashed or hyphenated name (feat/x, my-branch), or anything the clause calls
51// a branch ("the feature branch").
52const FIGURATIVE_INTO = /^\W*into\s+(?!(?:the\s+|its\s+|their\s+)?(?:main|master|develop|dev|trunk|base|branch|release|origin|upstream|CODE)\b)(?![\w.-]+\/)(?![^.;:!?]*\bbranch\b)(?![\w.]+-[\w.-]+)[a-z]/i
53// Anywhere in the sentence: a description of how a test behaves under a mutation, not a claim that the suite passes.
54// "against that mutation" too: a negative test that passes against a mutation describes the test.
55const UNCLAIM = /\b(?:even\s+without|identically|(?:against|under)\s+(?:(?:the|a|each|every|that|this|these|those)\s+)?mutations?|with\s+(?:\S+\s+){0,3}removed)\b/i
56
57/** Removes what is not my own claim: fenced code, blockquotes, inline code, and double- or single-quoted text. */
58export function stripNonClaims(text: string): string {
59  return text
60    .replace(/[’‘]/g, "'")
61    .replace(/[“”]/g, '"')
62    .replace(/```[\s\S]*?(?:```|$)/g, '\n')
63    .replace(/^[ \t]*>.*$/gm, '\n')
64    .replace(/`[^`\n]*`/g, ' CODE ')
65    .replace(/"[^"\n]*"/g, ' QUOTE ')
66    // A single-quoted span opens and closes beside a non-letter; an apostrophe between two letters ('I've pushed') stays inside it.
67    .replace(/(?<![\p{L}\p{N}])'(?:[^'\n]|(?<=\p{L})'(?=\p{L}))+?'(?![\p{L}\p{N}])/gu, ' QUOTE ')
68}
69
70/** Splits on sentence ends and on line breaks (so list items are sentences). */
71export function sentencesOf(text: string): string[] {
72  return text
73    .split(/(?<=[.!?])\s+|\n+/)
74    .map(s => s.trim())
75    .filter(s => s.length > 0)
76}
77
78/** FNV-1a, 32-bit, as 8 hex digits. Not security: a stable key for "flagged once per session". */
79export function fnv1a(text: string): string {
80  let hash = 0x811c9dc5
81  for (let i = 0; i < text.length; i += 1) {
82    hash ^= text.charCodeAt(i)
83    hash = Math.imul(hash, 0x01000193) >>> 0
84  }
85  return hash.toString(16).padStart(8, '0')
86}
87
88const CLAUSE_MARKS = [',', ';', ':', '—', ' - ']
89
90function clauseBefore(sentence: string, index: number): string {
91  const head = sentence.slice(0, index)
92  const cut = Math.max(...CLAUSE_MARKS.map(m => head.lastIndexOf(m)))
93  return head.slice(cut + 1)
94}
95
96function clauseAfter(sentence: string, end: number): string {
97  const tail = sentence.slice(end)
98  const cuts = CLAUSE_MARKS.map(m => tail.indexOf(m)).filter(i => i >= 0)
99  return cuts.length === 0 ? tail : tail.slice(0, Math.min(...cuts))
100}
101
102function normalize(sentence: string): string {
103  return sentence.toLowerCase().replace(/\s+/g, ' ').replace(/[.!?]+$/, '').trim()
104}
105
106/** Every claim in my text, one per (family, op, sentence), in pattern order within a sentence. */
107export function findClaims(text: string): Claim[] {
108  const found = new Map<string, Claim>()
109  for (const sentence of sentencesOf(stripNonClaims(text))) {
110    if (sentence.endsWith('?') || UNCLAIM.test(sentence)) continue
111    for (const pattern of PATTERNS) {
112      const match = pattern.re.exec(sentence)
113      if (match === null) continue
114      if (HEDGE.test(clauseBefore(sentence, match.index))) continue
115      // A present-state match spans its subject and up to 40 characters before the verb: a hedge in that gap counts
116      // too, so "branch review runs before anything is pushed" is not a claim.
117      if (pattern.gap === true && HEDGE.test(match[0])) continue
118      if (NEG_AFTER.test(clauseAfter(sentence, match.index + match[0].length))) continue
119      // "Merged into the pipeline", "merged into the tracker": merged into something that is not a branch is not git.
120      if (pattern.op === 'merge' && FIGURATIVE_INTO.test(clauseAfter(sentence, match.index + match[0].length))) continue
121      const hash = fnv1a(`${pattern.family}|${pattern.op ?? ''}|${normalize(sentence)}`)
122      if (found.has(hash)) continue
123      const phrase = match[0].replace(/\s+/g, ' ').trim().slice(0, 60)
124      found.set(hash, { family: pattern.family, op: pattern.op, phrase, hash })
125    }
126  }
127  return [...found.values()]
128}
129
hooks/evidence.ts 1516 lines
1import type { AsOf, Basis, Claim, Entry, Family, Kind, Ledger, Pending, ShipOp, Status } from '../types'
2
3// Pure: no `$`, no runtime imports. scripts/precision.mjs imports this file under Node.
4
5// Equal to the manifest's userConfig defaults: change both together.
6export const DEFAULT_TEST_COMMANDS: readonly string[] = ['pytest', 'uv run pytest', 'npm test', 'npm run test', 'npm run test:e2e', 'pnpm test', 'yarn test', 'vitest', 'jest', 'npx playwright test', 'cargo test', 'go test', 'claude plugin test']
7export const DEFAULT_BUILD_COMMANDS: readonly string[] = ['tsc', 'cargo build', 'npm run build', 'npm run typecheck', 'npm run typecheck:tests', 'idf.py build']
8
9/** `firsts`: the words a run can start with (each runner's first word, and every launcher's), lowercased. */
10export type Config = { tests: readonly RegExp[]; build: readonly RegExp[]; firsts?: ReadonlySet<string> }
11/** What a tool call answered (the ToolCallResult arms, read loosely). `text` is the result as the model read it. */
12export type Ran = { readonly deny?: string; readonly isError?: boolean; readonly result?: unknown; readonly text?: string }
13export type Facts = {
14  tool: string
15  command: string | null
16  path: string | null
17  toolUseId: string | null
18  denied: boolean
19  isError: boolean
20  interrupted: boolean
21  background: boolean
22  /** The result as the model read it: where a runner's summary shows. */
23  output: string
24  git: Partial<Record<ShipOp, true>>
25  /** Files a Bash command changed, when its result lists them (bashEditDiff, @internal: may be absent). */
26  changed: string[]
27  /** A background command's task id (`backgroundTaskId`, or the id its notice names), when it has one. */
28  taskId?: string | null
29  /**
30   * False when `isError` is not this command's exit status: a background run re-judged from a read-back that shows no
31   * `[exited with code N]` line. Absent means known.
32   */
33  exitKnown?: boolean
34  /** The result says `gh pr merge` only enabled auto-merge (`gitOperation.pr.action`): nothing was merged yet. */
35  autoMerge?: boolean
36}
37export type Run = { kind: Kind; ok: boolean; masked: boolean; basis: Basis }
38/**
39 * `mutations`: code edits, as pathKey()s. `messageFiles`: files a git or gh command read a message or body from
40 * (`-F`, `--file`, `--body-file`), normalized; relative ones as written.
41 */
42export type Classified = { mutations: string[]; runs: Run[]; messageFiles: string[]; keys?: string[] }
43/** Where things live: the home folder (memory roots), the session's project root (relative paths) and %TEMP%. Each may be unknown. */
44export type Places = { home: string | null; root: string | null; temp: string | null }
45export type Word = { text: string; quoted: boolean }
46/** The operator after a segment; '' ends the line (a trailing `&` is kept: the line ran in the background). */
47export type Op = '' | ';' | '\n' | '&' | '&&' | '||' | '|'
48export type Segment = { words: Word[]; op: Op }
49
50const SHELLS = new Set(['Bash', 'PowerShell'])
51const EDITORS = new Set(['Edit', 'Write', 'NotebookEdit'])
52const SHIP_OPS: readonly ShipOp[] = ['commit', 'push', 'merge', 'pr-create']
53
54// ---- Parsing a command line ----
55
56// A heredoc's body is data: keep the line that opens it, drop the body and its closing delimiter. With no closing
57// delimiter it is not a heredoc (`python -c "print(1<<n)"` is a shift), and nothing is dropped.
58// Matched at one position (sticky): heredocsOf only starts it at an opener whose word closes a later line.
59const HEREDOC = /<<-?[ \t]*(['"]?)([A-Za-z_]\w*)\1([^\n]*)\n(?:[\s\S]*?\n)?[ \t]*\2[ \t\r]*(?=\n|$)/y
60const OPENER = /<<-?[ \t]*(['"]?)([A-Za-z_]\w*)\1/g
61
62/**
63 * The heredocs in a command, in order. An opener whose word is no later line of the command never enters the body
64 * scan, so many unterminated `<<` (shifts, comparisons) cost nothing; the lazy scan from each one to the end was
65 * O(openers x length).
66 */
67function heredocsOf(command: string): RegExpExecArray[] {
68  if (!command.includes('<<')) return []
69  // Where each trimmed line last occurs.
70  const lastAt = new Map<string, number>()
71  let at = 0
72  for (const line of command.split('\n')) {
73    // Strip exactly what HEREDOC's closing line allows ([ \t] before, [ \t\r] after), never trim()'s wider set.
74    let a = 0
75    let b = line.length
76    while (a < b && (line[a] === ' ' || line[a] === '\t')) a += 1
77    while (b > a && (line[b - 1] === ' ' || line[b - 1] === '\t' || line[b - 1] === '\r')) b -= 1
78    lastAt.set(line.slice(a, b), at)
79    at += line.length + 1
80  }
81  const found: RegExpExecArray[] = []
82  OPENER.lastIndex = 0
83  for (let m = OPENER.exec(command); m !== null; m = OPENER.exec(command)) {
84    if ((lastAt.get(m[2] ?? '') ?? -1) <= m.index) continue
85    HEREDOC.lastIndex = m.index
86    const h = HEREDOC.exec(command)
87    if (h === null) continue
88    found.push(h)
89    OPENER.lastIndex = h.index + h[0].length
90  }
91  return found
92}
93
94/** The command with each heredoc's body and closing line dropped; its opener becomes `<<HEREDOC` and keeps its line. */
95function withoutHeredocs(command: string): string {
96  let out = ''
97  let last = 0
98  for (const h of heredocsOf(command)) {
99    out += `${command.slice(last, h.index)}<<HEREDOC${h[3] ?? ''}`
100    last = h.index + h[0].length
101  }
102  return out + command.slice(last)
103}
104// PowerShell here-strings are data too.
105const HERESTRING = /@(['"])\r?\n[\s\S]*?\r?\n\1@/g
106// A line continuation (bash backslash, PowerShell backtick) joins two lines.
107const CONTINUATION = /\\\r?\n|`\r?\n/g
108const BREAKS = new Set([' ', '\t', '\r', '(', ')', '{', '}'])
109
110/** Splits a bash or PowerShell line into simple commands at top-level operators, outside quotes. */
111export function segmentsOf(command: string): Segment[] {
112  const src = withoutHeredocs(command).replace(HERESTRING, ' @HERE@ ').replace(CONTINUATION, ' ')
113  const segments: Segment[] = []
114  let words: Word[] = []
115  let text = ''
116  let quoted = false
117  let open = false
118  let quote: string | null = null
119  const endWord = (): void => {
120    if (open) words.push({ text, quoted })
121    text = ''
122    quoted = false
123    open = false
124  }
125  const endSegment = (op: Op): void => {
126    endWord()
127    if (words.length > 0) segments.push({ words, op })
128    words = []
129  }
130  for (let i = 0; i < src.length; i += 1) {
131    const ch = src.charAt(i)
132    const next = src.charAt(i + 1)
133    if (quote !== null) {
134      if (ch === quote) {
135        quote = null
136      } else if (ch === '\\' && quote === '"' && next !== '' && '"\\$`'.includes(next)) {
137        text += next
138        i += 1
139      } else {
140        text += ch
141      }
142      continue
143    }
144    if (ch === "'" || ch === '"') {
145      quote = ch
146      quoted = true
147      open = true
148    } else if (ch === '\\' && next !== '') {
149      text += next
150      open = true
151      i += 1
152    } else if (ch === '#' && !open) {
153      while (i + 1 < src.length && src.charAt(i + 1) !== '\n') i += 1
154    } else if (ch === ';' || ch === '\n') {
155      endSegment(ch === ';' ? ';' : '\n')
156    } else if (ch === '&') {
157      const prev = src.charAt(i - 1)
158      if (next === '&') {
159        endSegment('&&')
160        i += 1
161      } else if (prev === '>' || prev === '<' || next === '>') {
162        // A redirection: 2>&1, >&2, &>.
163        text += ch
164        open = true
165      } else if (open || words.length > 0) {
166        endSegment('&')
167      }
168      // Otherwise PowerShell's call operator, `& "C:/x.exe"`: nothing to record.
169    } else if (ch === '|') {
170      if (next === '|') {
171        endSegment('||')
172        i += 1
173      } else {
174        if (next === '&') i += 1 // `|&` pipes stderr too
175        endSegment('|')
176      }
177    } else if (BREAKS.has(ch)) {
178      endWord()
179    } else {
180      text += ch
181      open = true
182    }
183  }
184  endSegment('')
185  const last = segments[segments.length - 1]
186  if (last !== undefined && last.op !== '&') last.op = ''
187  return segments
188}
189
190const SHELL_PAYLOAD: Readonly<Record<string, readonly string[]>> = {
191  bash: ['-c', '-lc'],
192  sh: ['-c'],
193  zsh: ['-c'],
194  pwsh: ['-command', '-c'],
195  powershell: ['-command', '-c'],
196  cmd: ['/c'],
197}
198const WRAPPERS = new Set(['sudo', 'time', 'env', 'command', 'exec', 'nice', 'nohup', 'xvfb-run', 'cross-env', 'dotenv', 'watchexec'])
199// Wrappers whose own options end at `--` (`dotenv -e .env -- npm test`, `watchexec -e ts -- pytest`).
200const DASHDASH_WRAPPERS = new Set(['dotenv', 'watchexec'])
201// Commands whose arguments are data: a runner word inside them is not a run.
202const DATA_HEADS = new Set(['echo', 'printf', 'grep', 'egrep', 'fgrep', 'rg', 'sed', 'awk', 'cat', 'head', 'tail', 'less', 'jq', 'which', 'where', 'pip', 'pip3', 'write-output', 'write-host', 'select-string', 'sls', 'findstr', 'get-content', 'set-content', 'out-file'])
203const ECHO_HEADS = new Set(['echo', 'printf', 'write-output', 'write-host'])
204// Flags whose value is a message or a body, not a command.
205const PAYLOAD_FLAGS = new Set(['-m', '--message', '--body', '-b', '--title', '-t', '-F', '--file', '--body-file'])
206
207// Shell keywords that open a loop or conditional body: `do gh pr merge $n`, `then git push`.
208const KEYWORDS = new Set(['do', 'then', 'else', 'elif', 'if', 'while', 'until', '!'])
209
210/**
211 * The index past a wrapper's own options: `nice -n 10`, `xvfb-run -a`, `xvfb-run -s "-screen 0 1x1x24"`. An option's
212 * value is skipped when it is a number or quoted; a `--`-terminated wrapper skips everything up to its `--`.
213 */
214function pastWrapperOptions(segment: Segment, from: number, toDashDash: boolean): number {
215  const words = segment.words
216  if (toDashDash) {
217    const dd = words.findIndex((w, k) => k >= from && !w.quoted && w.text === '--')
218    if (dd >= 0) return dd + 1
219  }
220  let i = from
221  while (i < words.length) {
222    const w = words[i]
223    if (w === undefined || w.quoted || !w.text.startsWith('-')) break
224    i += 1
225    if (w.text === '--') break
226    const value = words[i]
227    if (value !== undefined && (value.quoted || /^\d+$/.test(value.text))) i += 1
228  }
229  return i
230}
231
232/** The index of a segment's command word, past keywords, `X=1` assignments (quoted or not) and wrappers such as `sudo` or `timeout 60`. */
233function headOf(segment: Segment): number {
234  let i = 0
235  while (i < segment.words.length) {
236    const word = segment.words[i]
237    if (word === undefined) break
238    const t = word.text.toLowerCase()
239    if (/^[A-Za-z_]\w*=/.test(word.text)) i += 1
240    // PowerShell's `$out = idf.py build`: the command is what is assigned.
241    else if (/^\$[\w:]+$/.test(word.text) && segment.words[i + 1]?.text === '=') i += 2
242    else if (KEYWORDS.has(t)) i += 1
243    else if (WRAPPERS.has(t)) i = pastWrapperOptions(segment, i + 1, DASHDASH_WRAPPERS.has(t))
244    else if (t === 'timeout') i += 2
245    else break
246  }
247  return i
248}
249
250// A loop's status is its last pass's (or 0 when it ran none), and its status lines print once per pass, so no run
251// inside a loop is strong evidence. A loop keyword counts as a segment's command word or as any word before it that
252// headOf skips (`then for`, `time for`, `do while`), never inside quoted text. A bash loop runs from its keyword to its
253// matching `done`; PowerShell's braces are not segment boundaries, so a `foreach` or `ForEach-Object` runs to the end.
254const BASH_LOOPS = new Set(['for', 'while', 'until'])
255const PS_LOOPS = new Set(['foreach', 'foreach-object'])
256
257/** Which segments sit inside a loop. */
258function loopSegments(segments: readonly Segment[]): boolean[] {
259  let depth = 0
260  let toEnd = false
261  return segments.map(s => {
262    const words = s.words.slice(0, headOf(s) + 1)
263    const lead = words.filter(w => !w.quoted).map(w => w.text.toLowerCase())
264    if (lead.some(w => PS_LOOPS.has(w))) toEnd = true
265    depth += lead.filter(w => BASH_LOOPS.has(w)).length
266    const inside = toEnd || depth > 0
267    // `done` closes a loop, glued to a redirection too (`done>log`, `done<list`, `done>"$TEMP/x"`, which reads as quoted).
268    depth = Math.max(0, depth - words.filter(w => /^done(?![\w-])/.test(w.text)).length)
269    return inside
270  })
271}
272
273const commandWord = (segment: Segment): string => (segment.words[headOf(segment)]?.text ?? '').toLowerCase().replace(/\.exe$/, '')
274
275/** segmentsOf, with `bash -c "..."` and `pwsh -Command "..."` payloads parsed as the commands they are. */
276export function commandsOf(command: string, depth = 0): Segment[] {
277  const out: Segment[] = []
278  for (const segment of segmentsOf(command)) {
279    const flags = SHELL_PAYLOAD[commandWord(segment)]
280    const head = headOf(segment)
281    const at = flags === undefined ? -1 : segment.words.findIndex((w, i) => i > head && flags.includes(w.text.toLowerCase()))
282    const payload = at < 0 ? undefined : segment.words[at + 1]
283    const inner = payload === undefined || depth >= 2 ? [] : commandsOf(payload.text, depth + 1)
284    const tail = inner[inner.length - 1]
285    if (tail === undefined) {
286      out.push(segment)
287      continue
288    }
289    tail.op = segment.op
290    out.push(...inner)
291  }
292  return out
293}
294
295const PYTHONS = /^(?:python[\d.]*|py)$/
296
297/** A segment as matching sees it: from the command word on, quoted and payload arguments as `Q`; '' for a data command. */
298function plainOf(segment: Segment): string {
299  const head0 = commandWord(segment)
300  if (DATA_HEADS.has(head0)) return ''
301  const head = headOf(segment)
302  const out: string[] = []
303  for (let i = head; i < segment.words.length; i += 1) {
304    const word = segment.words[i]
305    if (word === undefined) continue
306    const before = segment.words[i - 1]
307    // `python -m pytest`: there `-m` names a module to run, not a message.
308    const isPayload = word.quoted || (i > head && before !== undefined && !before.quoted && PAYLOAD_FLAGS.has(before.text) && !(before.text === '-m' && segment.words.slice(head, i - 1).some(w => PYTHONS.test(w.text.toLowerCase()))))
309    out.push(isPayload ? 'Q' : word.text)
310  }
311  return out.join(' ')
312}
313
314// ---- Runners, git ops and their outcome ----
315
316// After the runner, in its segment: flags that make it run nothing.
317const NON_RUN = /\s(?:--(?:version|help|init|co|collect-only|no-run|listTests)|-h)(?=\s|$)/i
318
319const GIT = String.raw`^git(?:\s+-[cC]\s+\S+)*\s+`
320const SHIP_RES: Readonly<Record<ShipOp, RegExp>> = {
321  commit: new RegExp(`${GIT}commit(?![\\w-])`, 'i'),
322  push: new RegExp(`${GIT}push(?![\\w-])`, 'i'),
323  merge: new RegExp(`^gh\\s+pr\\s+merge(?![\\w-])|${GIT}merge(?![\\w-])`, 'i'),
324  'pr-create': /^gh\s+pr\s+create(?![\w-])/i,
325}
326// An op that does nothing: `git merge --abort`, `git push --dry-run`, `git push -n`, `git commit --help`, and
327// `gh pr merge --auto`, which only enables auto-merge. (`git commit -n` is --no-verify and still commits.)
328const SHIP_SKIP: Readonly<Record<ShipOp, RegExp>> = {
329  commit: /\s(?:--abort|--dry-run|--help|-h)(?=\s|$)/i,
330  push: /\s(?:--dry-run|--help|-h|-n)(?=\s|$)/i,
331  merge: /\s(?:--abort|--dry-run|--help|-h|--auto)(?=\s|$)/i,
332  'pr-create': /\s(?:--dry-run|--help|-h)(?=\s|$)/i,
333}
334
335// What a runner's visible output says. `claude plugin test` prints "66 pass" / "4 fail"; svelte-check prints
336// "COMPLETED ... 0 ERRORS"; tsup prints "Build success"; tsc prints nothing on success.
337const SUMMARY: Readonly<Record<'tests' | 'build', { fail: RegExp; pass: RegExp }>> = {
338  tests: {
339    // `FAILED` counts at a line's start (pytest's `FAILED tests/x.py::t`), not inside a test title. TAP and node:test
340    // print `not ok N` and `# fail N`; unittest prints `OK` or `OK (skipped=1)`, not any line that starts with OK.
341    fail: /\b[1-9]\d*\s+(?:failed|failing|fail|failures?|errors?)\b|^FAILED\b|^[ \t]*FAIL\b|^\(fail\)|test result: FAILED|^not ok\b|^#\s*fail\s+[1-9]/m,
342    pass: /\b[1-9]\d*\s+(?:passed|passing|pass)\b|\btest result: ok\b|^OK(?: \(|$)|^ok\s+\S/m,
343  },
344  build: {
345    // The last three: the compiler never started (npx found no tsc, or it is not on PATH), so its silence proves nothing.
346    // ESP-IDF's `idf.py build`: "Project build complete" on success; ninja's "build stopped" or a FAILED step on failure.
347    fail: /\berror TS\d+|\bFound [1-9]\d* errors?\b|^error(?:\[E\d+\])?:|\bBuild failed\b|\bFailed to compile\b|\bCOMPLETED\b.*\b[1-9]\d* ERRORS\b|This is not the tsc command|command not found|is not recognized as|\bninja: build stopped\b|^FAILED: /im,
348    pass: /^[ \t]*Finished\b|\bCompiled successfully\b|\bbuilt in \d|\bFound 0 errors\b|\bCOMPLETED\b.*\b0 ERRORS\b|\bBuild success\b|\bProject build complete\b/im,
349  },
350}
351const SHIP_FAIL = /^(?:error|fatal):|\[rejected\]|\bnothing to commit\b|\bfailed to push\b|\bAutomatic merge failed\b|^CONFLICT \(/im
352// Each op's own confirmation, so one op's output cannot confirm another (`MERGED` from `gh pr view` is not a commit).
353const SHIP_PASS: Readonly<Record<ShipOp, RegExp>> = {
354  commit: /^\[[\w./-]+(?: \(root-commit\))? [0-9a-f]{7,}\]/m,
355  push: /^[ \t]*(?:\+[ \t]*)?[0-9a-f]{7,}\.\.\.?[0-9a-f]{7,}\s+\S+\s+->\s+\S+|^[ \t]*\*\s+\[new (?:branch|tag)\]|\bset up to track\b/im,
356  // `gh pr view --json state` prints MERGED bare, first in a --jq line, or as JSON. `git merge` prints "Merge made by",
357  // and a `git log` after it shows the merge commit's own subject (`git merge x | tail -5 && git log --oneline -3`).
358  merge: /\bMerged pull request\b|^[ \t]*MERGED\b|"state"\s*:\s*"MERGED"|\bstate\s*[=:]\s*"?MERGED\b|^Merge made by\b|^[0-9a-f]{7,40} +(?:\([^)]*\) +)?Merge (?:remote-tracking branch|branch|pull request #\d+)\b/im,
359  'pr-create': /github\.com\/[\w.-]+\/[\w.-]+\/pull\/\d+/i,
360}
361// `git commit -q` prints nothing; a later `git log --oneline` in the same command prints a sha line: the commit's own
362// when it carries the commit's subject (after any `(HEAD -> x)` decoration), or, with no subject known, when nothing
363// between the commit and the log moved to another branch, folder or repository.
364const SHA_LINE = /^[0-9a-f]{7,40} +(?:\([^)]*\) +)?(\S.*)$/
365// `git show --oneline HEAD` prints the same sha line as `git log --oneline -1`. Without `--oneline` it prints
366// `commit <sha>`, which SHA_LINE does not take.
367const GIT_LOG = /^git(?:\s+-[cC]\s+\S+)*\s+(?:log|show)(?![\w-])/i
368// `gh pr view` in a call of its own: its MERGED is merge evidence, read only from gh's own forms, never from a
369// `git log` line another segment printed.
370const PR_VIEW = /^gh\s+pr\s+view(?![\w-])/i
371// `--jq .state` prints MERGED bare; `--json state` prints `"state": "MERGED"`; a jq template can print `state=MERGED`.
372const VIEW_MERGED = /^[ \t]*MERGED\b|"state"\s*:\s*"MERGED"|\bstate\s*[=:]\s*"?MERGED\b/im
373const GIT_ELSEWHERE = /^git(?:\s+-[cC]\s+\S+)*\s+(?:checkout|switch|pull|merge|reset)(?![\w-])/
374const CD = new Set(['cd', 'pushd', 'popd', 'chdir', 'set-location', 'sl'])
375const TSC = /(?:^|\s)tsc(?=$|\s)/i
376// Pipe members that keep every line tsc prints, so a failure would still show its first `error TS` line. `tail` keeps
377// them only as `tail -n +K` (a last-N `tail` often holds only a diagnostic's indented elaboration lines),
378// and `Select-Object` only without `-Last`; `less` and `more` are left out.
379const FILTERS = new Set(['head', 'tail', 'grep', 'egrep', 'rg', 'cat', 'tee', 'sort', 'uniq', 'select-object', 'select', 'out-host', 'out-string', 'select-string', 'sls', 'findstr'])
380const GREPS = new Set(['grep', 'egrep', 'rg', 'select-string', 'sls', 'findstr'])
381// Grep flags that print a count, a file name or nothing instead of the lines.
382const COUNTING = /^-(?:-(?:count|quiet|silent|files-with(?:out)?-matches)$|[A-Za-z]*[cqlL][A-Za-z]*$)|^-Quiet$/
383
384/** A parsed command line, with what classify has learned so far about each segment. */
385type Line = {
386  segments: readonly Segment[]
387  plains: readonly string[]
388  kindsAt: readonly Kind[][]
389  /** Segment k holds a run already judged strong and ok (classify fills this in order, so earlier segments are known). */
390  strong: readonly boolean[]
391  /** The first line of the heredoc segment k reads, when it reads one. */
392  heredocAt: readonly (string | undefined)[]
393  /** Each segment's echo shape (`filled(echoText)`), or null for a non-echo: filled on first use. */
394  shapes: (string | null)[]
395  /** The echoes matching each status pattern seen so far, by pattern source. */
396  like: Map<string, number[]>
397}
398
399/** An `echo` of a run's own status: the line pattern (status as group 1), and which of the echoes printing lines of that shape it is. */
400type Echoed = { pattern: StatusShape; n: number; of: number }
401
402type Place = {
403  /** Something after the run can replace its exit status. */
404  masked: boolean
405  /** The run's pipeline ends the line, so the call's isError is its status. */
406  last: boolean
407  /** What an `&& echo` after the run printed (only for a run that ends its pipeline). */
408  echoes: string[]
409  /** An `echo` of this run's own exit status straight after its pipeline. */
410  status: Echoed | null
411  /**
412   * Kinds run by later members of this run's `&&` chain: one that showed its own pass summary ran, so this run exited 0.
413   * Only for a run that ends its pipeline: otherwise the chain went on because the last filter exited 0.
414   */
415  chained: Kind[]
416  /** A `git log` after this run can show a quiet commit's sha line (see SHA_LINE). */
417  logged: boolean
418  /** The commit's subject, when the command shows it (`-m`, or the heredoc's first line). */
419  subject: string | null
420  /** Piped only into filters that keep every line tsc prints. */
421  filtered: boolean
422  /** Every earlier member of this run's `&&` chain is a `cd` or a proven run, so this run surely started. */
423  ran: boolean
424}
425
426/** What an echo prints: its arguments, without `-n`/`-e` and without a redirection (`>> log`), which prints nothing. */
427function echoText(segment: Segment): string {
428  const out: string[] = []
429  const words = segment.words.slice(headOf(segment) + 1)
430  for (let i = 0; i < words.length; i += 1) {
431    const w = words[i]
432    if (w === undefined) continue
433    if (!w.quoted && /^[0-9&]?>/.test(w.text)) {
434      if (/^[0-9&]?>>?$/.test(w.text)) i += 1 // the target is the next word
435      continue
436    }
437    if (w.quoted || !/^-[neE]+$/.test(w.text)) out.push(w.text)
438  }
439  return out.join(' ').trim()
440}
441
442/**
443 * The line an echo of a run's status prints, as a pattern. `$?` is the status of the command just before the echo, so
444 * it counts only for the last member of that pipeline; `${PIPESTATUS[n]}` names the n-th member; `$LASTEXITCODE` is
445 * PowerShell's. Any other variable in the echo matches any text.
446 */
447function statusPattern(template: string, last: boolean, member: number, piped: boolean): StatusShape | null {
448  const own = [...(last ? [String.raw`\$\?`, String.raw`\$\{\?\}`] : []), ...(piped ? [String.raw`\$\{PIPESTATUS\[${member}\]\}`] : []), String.raw`\$LASTEXITCODE\b`]
449  const found = new RegExp(own.join('|'), 'i').exec(template)
450  if (found === null) return null
451  return statusShape(template.slice(0, found.index), template.slice(found.index + found[0].length))
452}
453
454/** The shape of a status line, matched like a regex (`source`, `test`, `exec` with the status as group 1). */
455type StatusShape = { source: string; test: (line: string) => boolean; exec: (line: string) => [string, string] | null }
456
457const VARIABLE = /\$\{[^}$]*\}|\$[A-Za-z_?][\w]*/
458
459/**
460 * A status line is the echo's literal pieces in order with anything at each variable (a glob with `*` per variable),
461 * then the status, then the tail's pieces. It is matched with indexOf, never a regex of lazy wildcards (which
462 * backtracks as length^k). The head's pieces before its last are matched leftmost from the line's start once, and the
463 * tail's pieces after its first rightmost from the line's end once; each candidate status position is then checked in
464 * constant time, so a line costs time linear in its length however many variables the echo has.
465 */
466function statusShape(before: string, after: string): StatusShape {
467  const head = before.split(VARIABLE)
468  const tail = after.split(VARIABLE)
469  // Lines are trimmed, so a leading variable that printed nothing leaves its line starting at the next piece with the
470  // space trimmed off (and a trailing one likewise at the end). That reading is tried only anchored at the line's
471  // start (or end), never in place of the separator when the variable printed something.
472  const heads = [head]
473  if (head.length > 1 && head[0] === '' && (head[1] ?? '') !== (head[1] ?? '').trimStart()) heads.push([(head[1] ?? '').trimStart(), ...head.slice(2)])
474  const tails = [tail]
475  if (tail.length > 1 && tail[tail.length - 1] === '') {
476    const piece = tail[tail.length - 2] ?? ''
477    if (piece !== piece.trimEnd()) tails.push([...tail.slice(0, -2), piece.trimEnd()])
478  }
479  const exec = (line: string): [string, string] | null => {
480    for (const h of heads) {
481      for (const t of tails) {
482        const hit = shapeMatch(h, t, line)
483        if (hit !== null) return hit
484      }
485    }
486    return null
487  }
488  return { source: `${head.join('\u0000')}\u0001${tail.join('\u0000')}`, test: line => exec(line) !== null, exec }
489}
490
491/** One reading of a status line against the head and tail globs (see statusShape). */
492function shapeMatch(head: readonly string[], tail: readonly string[], line: string): [string, string] | null {
493  const lastHead = head[head.length - 1] ?? ''
494  const firstTail = tail[0] ?? ''
495  {
496    // Where the head's last piece may start (minStart), and where the tail's first piece may end (maxEnd).
497    let minStart = 0
498    if (head.length > 1) {
499      const first = head[0] ?? ''
500      if (!line.startsWith(first)) return null
501      minStart = first.length
502      for (let h = 1; h < head.length - 1; h += 1) {
503        const piece = head[h] ?? ''
504        if (piece === '') continue
505        const k = line.indexOf(piece, minStart)
506        if (k < 0) return null
507        minStart = k + piece.length
508      }
509    }
510    let maxEnd = line.length
511    if (tail.length > 1) {
512      const last = tail[tail.length - 1] ?? ''
513      if (!line.endsWith(last)) return null
514      maxEnd = line.length - last.length
515      for (let t = tail.length - 2; t >= 1; t -= 1) {
516        const piece = tail[t] ?? ''
517        if (piece === '') continue
518        const k = line.lastIndexOf(piece, maxEnd - piece.length)
519        if (k < 0) return null
520        maxEnd = k
521      }
522    }
523    // The tail's first piece must occur somewhere after the head: a quick reject before any candidate scan.
524    if (firstTail !== '' && line.indexOf(firstTail, minStart) < 0) return null
525    // The status at `at`: the head ends there and the tail starts right after the digits.
526    const check = (at: number): [string, string] | null => {
527      DIGITS.lastIndex = at
528      const m = DIGITS.exec(line)
529      if (m === null) return null
530      const b = at + m[0].length
531      const tailOk = tail.length === 1 ? line.length - b === firstTail.length && line.startsWith(firstTail, b) : line.startsWith(firstTail, b) && b + firstTail.length <= maxEnd
532      return tailOk ? [line, m[0]] : null
533    }
534    if (head.length === 1) return line.startsWith(lastHead) ? check(lastHead.length) : null
535    // A tail with no variable fixes where the status ends: only the digits just before it are candidates.
536    if (tail.length === 1) {
537      const b = line.length - firstTail.length
538      if (b < 0 || !line.startsWith(firstTail, b)) return null
539      let a = b
540      while (a > 0 && line.charCodeAt(a - 1) >= 48 && line.charCodeAt(a - 1) <= 57) a -= 1
541      if (a > 0 && line.charCodeAt(a - 1) === 45) a -= 1
542      for (let at = a; at < b; at += 1) {
543        const start = at - lastHead.length
544        if (start < minStart || !line.startsWith(lastHead, start)) continue
545        // Everything from `a` to `b` is digits, but for a minus sign at `a`: a lone `-` is no status.
546        if (b - at === 1 && line.charCodeAt(at) === 45) continue
547        return [line, line.slice(at, b)]
548      }
549      return null
550    }
551    if (lastHead !== '') {
552      // Each place the head's last piece occurs (no earlier than its predecessors allow) is a candidate.
553      for (let p = line.indexOf(lastHead, minStart); p >= 0; p = line.indexOf(lastHead, p + 1)) {
554        const hit = check(p + lastHead.length)
555        if (hit !== null) return hit
556      }
557      return null
558    }
559    // The head ends in a variable: each maximal digit run (with its minus sign) from minStart on is a candidate. Every
560    // start inside one run ends at the same place, so the tail is checked once per run; the run's own start wins.
561    const isDigit = (c: number): boolean => c >= 48 && c <= 57
562    for (let p = minStart; p < line.length; ) {
563      const c = line.charCodeAt(p)
564      // A `-` is a minus sign only where it cannot be a hyphen inside a word (`run-0` holds no status -0).
565      const minus = c === 45 && isDigit(line.charCodeAt(p + 1)) && (p === minStart || !/[A-Za-z0-9_]/.test(line.charAt(p - 1)))
566      if (!isDigit(c) && !minus) {
567        p += 1
568        continue
569      }
570      let e = c === 45 ? p + 1 : p
571      while (e < line.length && isDigit(line.charCodeAt(e))) e += 1
572      const tailOk = tail.length === 1 ? line.length - e === firstTail.length && line.startsWith(firstTail, e) : line.startsWith(firstTail, e) && e + firstTail.length <= maxEnd
573      if (tailOk) return [line, line.slice(p, e)]
574      p = e
575    }
576    return null
577  }
578}
579
580const DIGITS = /-?\d+/y
581
582const STATUS_LINE_CAP = 300
583let statusCache: { output: string; lines: string[]; by: Map<string, { found: string[]; once: string[] }> } | null = null
584
585/** An echo's text with every variable as `0`: the shape of the line it prints. */
586const filled = (text: string): string => text.replace(/\$\{[^}$]*\}|\$[A-Za-z_?][\w]*/g, '0')
587
588/**
589 * The status an echo printed. Several echoes can print lines of one shape (`echo "exit=$?"` after each of two runs),
590 * so the n-th matching line is the n-th such echo's; when the count of matching lines differs from the count of those
591 * echoes (a loop, or a program printing the same shape), the status is unsure.
592 */
593function statusIn(output: string, status: Echoed): number | 'unsure' | null {
594  // One command's runs share its output and usually one status shape: split the output and scan it once per shape.
595  if (statusCache?.output !== output) {
596    // A status line is short; a long line is never one, and skipping it bounds the pattern's work.
597    const lines = output
598      .split(/\r?\n/)
599      .map(line => line.trim())
600      .filter(line => line.length <= STATUS_LINE_CAP)
601    statusCache = { output, lines, by: new Map() }
602  }
603  const cache = statusCache
604  let seen = cache.by.get(status.pattern.source)
605  if (seen === undefined) {
606    const matching = cache.lines.filter(line => status.pattern.test(line))
607    const values = (from: readonly string[]) => from.map(line => status.pattern.exec(line)?.[1]).filter((v): v is string => v !== undefined)
608    // A reader that prints the same log twice (grep, then tail) repeats each status line: identical lines are one.
609    seen = { found: values(matching), once: values([...new Set(matching)]) }
610    cache.by.set(status.pattern.source, seen)
611  }
612  if (seen.found.length === 0) return null
613  if (seen.found.length === status.of) return Number(seen.found[status.n])
614  return seen.once.length === status.of ? Number(seen.once[status.n]) : 'unsure'
615}
616
617/** A pipe member that keeps every line tsc prints (FILTERS). */
618function keepsLines(segment: Segment): boolean {
619  const word = commandWord(segment)
620  if (!FILTERS.has(word)) return false
621  const args = segment.words.slice(headOf(segment) + 1).filter(w => !w.quoted).map(w => w.text)
622  if (word === 'tail') return args.some(a => /^(?:-n|--lines=)?\+\d+$/.test(a))
623  if (word === 'select-object' || word === 'select') return !args.some(a => /^-l/i.test(a))
624  return !GREPS.has(word) || !args.some(a => COUNTING.test(a))
625}
626
627/** The repository a git segment names with `-C` before its verb, or null. */
628function gitDirOf(segment: Segment): string | null {
629  if (commandWord(segment) !== 'git') return null
630  for (let k = headOf(segment) + 1; k < segment.words.length; ) {
631    const word = segment.words[k]
632    if (word === undefined || !word.text.startsWith('-')) break
633    if (word.text === '-C') return segment.words[k + 1]?.text ?? null
634    k += word.text === '-c' ? 2 : 1
635  }
636  return null
637}
638
639/** A commit segment's subject: the first line of its `-m` value, or of the heredoc it reads; null when the command does not show it. */
640function subjectOf(segment: Segment, heredoc: string | undefined): string | null {
641  const words = segment.words
642  const at = words.findIndex(w => !w.quoted && (w.text === '-m' || w.text === '--message'))
643  let value = at >= 0 ? words[at + 1]?.text : words.map(w => /^--message=([\s\S]*)$/.exec(w.text)?.[1]).find(v => v !== undefined)
644  if (value === undefined || value.includes('<<HEREDOC')) value = words.some(w => w.text.includes('<<HEREDOC')) ? heredoc : undefined
645  else if (/[$`]/.test(value)) value = undefined // `-m "$(cat msg.txt)"`: the text is not in the command
646  const subject = (value ?? '').split(/\r?\n/)[0]?.trim() ?? ''
647  return subject === '' ? null : subject
648}
649
650/**
651 * The echoes on the line that print a line of this pattern's shape, in order. Each echo's shape is computed once per
652 * line and each distinct pattern is scanned once, so a line of many runs and echoes stays linear.
653 */
654function echoesLike(line: Line, pattern: StatusShape): number[] {
655  const cached = line.like.get(pattern.source)
656  if (cached !== undefined) return cached
657  if (line.shapes.length === 0) {
658    for (const s of line.segments) line.shapes.push(ECHO_HEADS.has(commandWord(s)) ? filled(echoText(s)) : null)
659  }
660  const found: number[] = []
661  line.shapes.forEach((shape, k) => {
662    if (shape !== null && shape.length <= STATUS_LINE_CAP && pattern.test(shape)) found.push(k)
663  })
664  line.like.set(pattern.source, found)
665  return found
666}
667
668/** What, after segment i, can replace its exit status, and what else on the line can show how it ended. */
669function placeOf(line: Line, i: number): Place {
670  const { segments, plains, kindsAt, strong } = line
671  let start = i
672  while (start > 0 && segments[start - 1]?.op === '|') start -= 1
673  let j = i
674  while (j < segments.length - 1 && segments[j]?.op === '|') j += 1
675  // A later `&&` member proves this run only when this run's status is its pipeline's (it ends the pipeline: a piped
676  // run's chain goes on because the last filter exited 0), and when no `||` before it could have skipped it
677  // (`x || run && echo ok` prints ok when x passed and run never ran).
678  const credits = j === i && segments[start - 1]?.op !== '||'
679  let masked = j > i // a later pipe member's status is the pipeline's
680  const echoes: string[] = []
681  // Still in this run's `&&` chain: an `&& echo` after a `;`, newline, `&` or `||` prints whatever this run did.
682  let chain = true
683  for (let k = j; k < segments.length; k += 1) {
684    const op = segments[k]?.op ?? ''
685    if (op === '') break
686    if (op === '|') continue // a pipe inside a later && command: this run's failure still stops the chain
687    if (op === '&&') {
688      const after = segments[k + 1]
689      if (after !== undefined && ECHO_HEADS.has(commandWord(after))) {
690        masked = true
691        if (credits && chain) echoes.push(echoText(after))
692      }
693      continue
694    }
695    masked = true // `;`, a newline, `&` or `||`: another command's status ends the line
696    chain = false
697    break // nothing later can unmask the run or join its chain
698  }
699  // An echo of the status straight after the pipeline (`; echo "rc=$?"`), never after `&&` or `||`.
700  const end = segments[j]?.op
701  const next = segments[j + 1]
702  const echoed = (end === ';' || end === '\n') && next !== undefined && ECHO_HEADS.has(commandWord(next))
703  const pattern = echoed ? statusPattern(echoText(next), j === i, i - start, j > start) : null
704  let status: Echoed | null = null
705  if (pattern !== null) {
706    // Every echo on the line that prints a line of this shape, in order; this one is the n-th.
707    const same = echoesLike(line, pattern)
708    const n = same.indexOf(j + 1)
709    if (n >= 0) status = { pattern, n, of: same.length }
710  }
711  // The members of this run's `&&` chain after it, each a pipeline.
712  const chained: Kind[] = []
713  for (let k = j; credits && segments[k]?.op === '&&'; ) {
714    let last = k + 1
715    while (last < segments.length - 1 && segments[last]?.op === '|') last += 1
716    for (let m = k + 1; m <= last; m += 1) chained.push(...(kindsAt[m] ?? []))
717    k = last
718  }
719  // A later `git log`. With the commit's subject known, its sha line must carry it (shaShows). Without it, nothing
720  // between the commit and the log may change branch, folder or repository.
721  const segment = segments[i]
722  const subject = segment !== undefined && (kindsAt[i] ?? []).includes('commit') ? subjectOf(segment, line.heredocAt[i]) : null
723  const here = segment === undefined ? null : gitDirOf(segment)
724  // Only a commit reads a later log, so only a commit pays for the scan.
725  const log = (kindsAt[i] ?? []).includes('commit') ? plains.findIndex((plain, m) => m > j && GIT_LOG.test(plain)) : -1
726  const stays =
727    log >= 0 &&
728    segments.slice(j + 1, log + 1).every((s, k) => !CD.has(commandWord(s)) && !GIT_ELSEWHERE.test(plains[j + 1 + k] ?? '') && (commandWord(s) !== 'git' || gitDirOf(s) === here))
729  // tsc's silence proves something only if tsc started: every earlier member of its `&&` chain is a `cd` or a run
730  // already proven, and no `||` could have skipped it.
731  let ran = true
732  for (let k = start - 1; k >= 0; ) {
733    const op = segments[k]?.op
734    if (op === '||') ran = false
735    if (op !== '&&') break
736    let first = k
737    while (first > 0 && segments[first - 1]?.op === '|') first -= 1
738    const before = segments[k]
739    const cd = first === k && before !== undefined && CD.has(commandWord(before))
740    if (!cd && !strong.slice(first, k + 1).some(Boolean)) {
741      ran = false
742      break
743    }
744    k = first - 1
745  }
746  return {
747    masked,
748    last: j === segments.length - 1,
749    echoes,
750    status,
751    chained,
752    logged: log >= 0 && (subject !== null || stays),
753    subject,
754    filtered: j > i && segments.slice(i + 1, j + 1).every(keepsLines),
755    ran,
756  }
757}
758
759/** A sha line in the output: one carrying the subject when it is known, else any. */
760function shaShows(output: string, subject: string | null): boolean {
761  // `git show --stat HEAD` and `git log -1` without `--oneline` print the subject indented by four spaces.
762  const head = subject === null ? '' : subject.slice(0, 40)
763  if (head !== '' && output.split(/\r?\n/).some(text => /^ {4}\S/.test(text) && text.trim().startsWith(head))) return true
764  return output.split(/\r?\n/).some(text => {
765    const match = SHA_LINE.exec(text.trim())
766    return match !== null && (subject === null || (match[1] ?? '').startsWith(subject.slice(0, 40)))
767  })
768}
769
770// One command's runs share its output: each summary pattern is tested against it once, and its lines are split once.
771let shownCache: { output: string; tested: Map<RegExp, boolean>; lines: string[] | null } | null = null
772function outputHas(re: RegExp, output: string): boolean {
773  if (shownCache?.output !== output) shownCache = { output, tested: new Map(), lines: null }
774  let hit = shownCache.tested.get(re)
775  if (hit === undefined) {
776    hit = re.test(output)
777    shownCache.tested.set(re, hit)
778  }
779  return hit
780}
781function outputLines(output: string): string[] {
782  if (shownCache?.output !== output) shownCache = { output, tested: new Map(), lines: null }
783  shownCache.lines ??= output.split(/\r?\n/).map(l => l.trim())
784  return shownCache.lines
785}
786
787function shownBy(kind: Kind, output: string, echoes: readonly string[], logged: boolean, subject: string | null): 'pass' | 'fail' | null {
788  const group = kind === 'tests' || kind === 'build' ? SUMMARY[kind] : { fail: SHIP_FAIL, pass: SHIP_PASS[kind] }
789  if (outputHas(group.fail, output)) return 'fail'
790  if (outputHas(group.pass, output)) return 'pass'
791  if (kind === 'commit' && logged && shaShows(output, subject)) return 'pass'
792  // The text an `&& echo` printed shows the run before it succeeded.
793  if (echoes.some(t => t.length >= 2) && echoes.some(t => t.length >= 2 && outputLines(output).includes(t))) return 'pass'
794  return null
795}
796
797function strengthOf(kind: Kind, facts: Facts, place: Place, plain: string): Omit<Run, 'kind'> {
798  if (facts.denied || facts.interrupted) return { ok: false, masked: false, basis: 'exit' }
799  // The exit status speaks for this run only when nothing after it can replace it; isError speaks for the last segment.
800  if (facts.exitKnown !== false && !place.masked && (place.last || !facts.isError)) return { ok: !facts.isError, masked: false, basis: 'exit' }
801  // An echo of this run's own status is its exit code; echoed lines that do not pair off with their echoes are unsure.
802  const status = place.status === null ? null : statusIn(facts.output, place.status)
803  // Unsure status lines still let a visible failure speak; a visible pass does not outvote them.
804  if (status === 'unsure') {
805    const seen = shownBy(kind, facts.output, [], false, null)
806    return seen === 'fail' ? { ok: false, masked: false, basis: 'output' } : { ok: true, masked: true, basis: 'none' }
807  }
808  if (status !== null) return { ok: status === 0, masked: false, basis: 'echo' }
809  const shown = shownBy(kind, facts.output, place.echoes, place.logged, place.subject)
810  if (shown === 'fail') return { ok: false, masked: false, basis: 'output' }
811  if (shown === 'pass') return { ok: true, masked: false, basis: 'output' }
812  // A later `&&` member that showed its own pass summary ran, so this run exited 0. (A failure summary is not used:
813  // a generic `error:` line may be this run's own.) `chained` is empty unless this run ends its pipeline.
814  if (place.chained.some(k => shownBy(k, facts.output, [], false, null) === 'pass')) return { ok: true, masked: false, basis: 'output' }
815  // tsc prints nothing on success, and its first line on failure is an `error TS` line. So no such line is a pass when
816  // tsc surely started (the call did not error, every earlier `&&` member is a `cd` or proven) and every later pipe
817  // member keeps all its lines (no last-N `tail`, no count).
818  if (kind === 'build' && facts.exitKnown !== false && place.filtered && place.ran && !facts.isError && TSC.test(plain)) return { ok: true, masked: false, basis: 'output' }
819  return { ok: true, masked: true, basis: 'none' }
820}
821
822/** A userConfig list: an array (the manifest's `multiple` form) or a comma string; empty means the defaults. */
823export function listOption(value: unknown, fallback: readonly string[]): string[] {
824  const raw: unknown[] = Array.isArray(value) ? value : typeof value === 'string' ? value.split(',') : []
825  const list = raw.filter((v): v is string => typeof v === 'string').map(v => v.trim()).filter(v => v.length > 0)
826  return list.length > 0 ? list : [...fallback]
827}
828
829function escapeRe(text: string): string {
830  return text.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')
831}
832
833// One option with at most one value. The option name starts with a word character, so `--opt` splits only one way
834// (`--?[\w-]+` could read it as `-` plus `-opt`, which doubles the work per option when a match fails).
835// A word that starts a launcher is never an option's value, so `pnpm -a pnpm -a ...` reads one way only. Only words
836// that really start one are kept out: a bare launcher word, `X run`, `npm exec`, `python -m`. A value such as `node`,
837// `py` or `python3.12` alone stays a value (`pnpm --filter node test`, `uv run --python python3.12 pytest`).
838const OPTION = String.raw`\s+--?\w[\w-]*(?:[= ](?!(?:npx|bunx|pnpx|pnpm|yarn|bun|uvx)(?:\s|$)|(?:uv|poetry|pdm|hatch|pipenv|pipx)\s+run(?:\s|$)|(?:pnpm|yarn)\s+dlx(?:\s|$)|npm\s+exec(?:\s|$)|(?:python[\d.]*|py)\s+-m(?:\s|$))[^\s-]\S*)?`
839// Options between a runner's words: `idf.py -C firmware build`.
840const OPTIONS_BETWEEN = String.raw`(?:${OPTION})*\s+`
841// What may come before a runner that is not the command word itself: a launcher that runs it.
842const LAUNCHER = String.raw`(?:(?:npx|bunx|pnpx)(?:${OPTION})*\s+|uvx(?:${OPTION})*\s+|(?:pnpm|yarn)\s+dlx(?:${OPTION})*\s+|(?:uv|poetry|pdm|hatch|pipenv|pipx)\s+run(?:${OPTION})*\s+|(?:pnpm|npm)\s+exec(?:${OPTION})*\s+(?:--\s+)?|(?:pnpm|yarn)(?:${OPTION})*\s+|bun(?:\s+run|\s+x)?(?:${OPTION})*\s+|python[\d.]*\s+-m\s+|py\s+-m\s+)`
843
844/**
845 * A runner as whole tokens at the start of a segment's command, or after a launcher: `pytest -q`, `npx vitest run`,
846 * `uv run --with pytest pytest`, `python -m pytest`. Never an argument of another command (`npm install -D vitest`,
847 * `find . -name jest`, `mkdir -p tsc`). Options may sit between its words (`idf.py -C firmware build`).
848 */
849export function runnerRe(command: string): RegExp {
850  const body = command.trim().split(/\s+/).map(escapeRe).join(OPTIONS_BETWEEN)
851  // At most three launchers deep (`uv run python -m pytest` is two): an unbounded repeat costs 2^n on a failing match.
852  return new RegExp(`^(?:${LAUNCHER}){0,3}${body}(?=$|\\s)`, 'i')
853}
854
855// Words that start a launcher (see LAUNCHER); `python3.12` and the like are matched by PYTHONS.
856const LAUNCH_WORDS = ['npx', 'bunx', 'pnpx', 'uvx', 'pnpm', 'yarn', 'bun', 'uv', 'poetry', 'pdm', 'hatch', 'pipenv', 'pipx', 'npm', 'py']
857
858export function configOf(tests: readonly string[], build: readonly string[]): Config {
859  const firsts = new Set([...tests, ...build].map(c => (c.trim().split(/\s+/)[0] ?? '').toLowerCase()).concat(LAUNCH_WORDS))
860  return { tests: tests.map(runnerRe), build: build.map(runnerRe), firsts }
861}
862
863function runnerIn(plain: string, res: readonly RegExp[]): boolean {
864  for (const re of res) {
865    const match = re.exec(plain)
866    if (match === null) continue
867    if (NON_RUN.test(plain.slice(match.index + match[0].length))) continue
868    return true
869  }
870  return false
871}
872
873const objectOf = (value: unknown): Record<string, unknown> =>
874  typeof value === 'object' && value !== null ? (value as Record<string, unknown>) : {}
875
876/** What a finished tool call tells the ledger. `input` is the tool.call event (tool name beside its arguments). */
877export function factsOf(input: Readonly<Record<string, unknown>>, ran: Ran): Facts {
878  const result = objectOf(ran.result)
879  const op = objectOf(result.gitOperation)
880  const pr = objectOf(op.pr)
881  const branch = objectOf(op.branch)
882  const git: Partial<Record<ShipOp, true>> = {}
883  if (op.commit !== undefined) git.commit = true
884  if (op.push !== undefined) git.push = true
885  if (pr.action === 'merged' || branch.action === 'merged') git.merge = true
886  if (pr.action === 'created') git['pr-create'] = true
887  const diff = objectOf(result.bashEditDiff)
888  const listed: unknown[] = Array.isArray(diff.changedFiles)
889    ? diff.changedFiles
890    : Array.isArray(diff.files)
891      ? diff.files.map(f => objectOf(f).filePath)
892      : []
893  const stdout = typeof result.stdout === 'string' ? result.stdout : null
894  const stderr = typeof result.stderr === 'string' ? result.stderr : ''
895  const output = typeof ran.text === 'string' ? ran.text : stdout !== null ? `${stdout}\n${stderr}` : typeof ran.result === 'string' ? ran.result : ''
896  return {
897    tool: String(input.tool),
898    command: typeof input.command === 'string' ? input.command : null,
899    path: typeof input.file_path === 'string' ? input.file_path : typeof input.notebook_path === 'string' ? input.notebook_path : null,
900    toolUseId: typeof input.tool_use_id === 'string' ? input.tool_use_id : null,
901    denied: ran.deny !== undefined,
902    isError: ran.isError === true,
903    interrupted: result.interrupted === true,
904    background: input.run_in_background === true || typeof result.backgroundTaskId === 'string',
905    output,
906    git,
907    changed: listed.filter((p): p is string => typeof p === 'string'),
908    taskId: typeof result.backgroundTaskId === 'string' ? result.backgroundTaskId : (BG_ID.exec(output)?.[1] ?? null),
909    autoMerge: pr.action === 'auto-merge-enabled',
910  }
911}
912
913// ---- Mutation scope ----
914
915// The user's temp folder comes from the environment (`places.temp`); these are the fixed ones.
916const SCRATCH_DIRS = /^\/tmp\/|\/scratchpad\/|^[a-z]:\/windows\/temp\//
917const trimEnd = (p: string): string => p.replace(/\/+$/, '')
918
919/** Forward slashes, `~/` and git-bash `/c/` expanded, lowercased (Windows paths compare case-insensitively). */
920export function normPath(path: string, home: string | null): string {
921  let p = path.replace(/\\/g, '/').replace(/\/{2,}/g, '/')
922  if (p.startsWith('~/') && home !== null) p = trimEnd(home.replace(/\\/g, '/')) + p.slice(1)
923  const drive = /^\/([a-zA-Z])\//.exec(p)
924  if (drive !== null) p = `${drive[1]}:/${p.slice(3)}`
925  return p.toLowerCase()
926}
927
928/** A path as the ledger keys it: normalized, and made absolute against the session root when it is relative and the root is known. */
929export function pathKey(path: string, places: Places): string {
930  const p = normPath(path, places.home)
931  if (/^(?:[a-z]:\/|\/)/.test(p) || places.root === null) return p
932  return `${trimEnd(normPath(places.root, places.home))}/${p.replace(/^\.\//, '')}`
933}
934
935/**
936 * A write that counts as a code edit: anywhere, sibling worktrees included, except `.md` files,
937 * memory roots, and scratch or temp folders (the session scratchpad and %TEMP%). A commit-message or PR-body file is
938 * taken back out when a git or gh command reads it (`record`).
939 */
940export function isCodeMutation(path: string, places: Places): boolean {
941  const p = pathKey(path, places)
942  if (p.endsWith('.md') || SCRATCH_DIRS.test(p)) return false
943  if (places.temp !== null) {
944    const temp = trimEnd(normPath(places.temp, places.home))
945    if (temp !== '' && p.startsWith(`${temp}/`)) return false
946  }
947  if (places.home !== null) {
948    const claude = `${trimEnd(normPath(places.home, null))}/.claude/`
949    if (p.startsWith(`${claude}memory/`) || p.startsWith(`${claude}rules/`)) return false
950    const projects = `${claude}projects/`
951    if (p.startsWith(projects) && p.slice(projects.length).split('/')[1] === 'memory') return false
952  }
953  return true
954}
955
956/** The kinds of run a segment is: test or build runners, and git or gh ops. */
957// A segment longer than this (as matching sees it, quoted arguments collapsed) is not matched at all: no real runner
958// or git command is that long, and it bounds the pattern work. It yields no run, so a claim on it reads as unknown.
959const SEGMENT_CAP = 1000
960
961function kindsIn(plain: string, config: Config): Kind[] {
962  if (plain === '' || plain.length > SEGMENT_CAP) return []
963  const found: Kind[] = []
964  // Runners are anchored at the segment's start: a first word that no runner or launcher begins with runs nothing,
965  // and skipping their patterns keeps a command of thousands of segments cheap.
966  const first = plain.slice(0, plain.indexOf(' ') < 0 ? plain.length : plain.indexOf(' ')).toLowerCase()
967  const mayRun = config.firsts === undefined || config.firsts.has(first) || PYTHONS.test(first)
968  if (mayRun && runnerIn(plain, config.tests)) found.push('tests')
969  if (mayRun && runnerIn(plain, config.build)) found.push('build')
970  for (const op of SHIP_OPS) if (SHIP_RES[op].test(plain) && !SHIP_SKIP[op].test(plain)) found.push(op)
971  return found
972}
973
974// A git or gh command that reads its message or body from a file: that file was a message, not code.
975const MESSAGE_CMD = /^(?:git(?:\s+-[cC]\s+\S+)*\s+(?:commit|tag|merge)|gh\s+(?:pr|issue|release)\s+(?:create|edit|merge|comment))(?![\w-])/i
976const MESSAGE_FLAGS = new Set(['-F', '--file', '--body-file'])
977
978/** The files git or gh read a message or body from, normalized; a relative one as written. `-F -` (stdin) is none. */
979function messageFilesOf(segments: readonly Segment[], plains: readonly string[], home: string | null): string[] {
980  const out: string[] = []
981  segments.forEach((segment, i) => {
982    if (!MESSAGE_CMD.test(plains[i] ?? '')) return
983    segment.words.forEach((word, k) => {
984      const inline = /^--(?:file|body-file)=(.+)$/.exec(word.text)
985      const value = inline !== null ? inline[1] : MESSAGE_FLAGS.has(word.text) ? segment.words[k + 1]?.text : undefined
986      if (value !== undefined && value !== '' && value !== '-') out.push(normPath(value, home).replace(/^\.\//, ''))
987    })
988  })
989  return out
990}
991
992export function classify(facts: Facts, config: Config, places: Places): Classified {
993  if (EDITORS.has(facts.tool)) {
994    const ok = !facts.denied && !facts.isError && !facts.interrupted
995    const edited = ok && facts.path !== null && isCodeMutation(facts.path, places)
996    return { mutations: edited && facts.path !== null ? [pathKey(facts.path, places)] : [], runs: [], messageFiles: [] }
997  }
998  const command = facts.command
999  if (!SHELLS.has(facts.tool) || command === null) return { mutations: [], runs: [], messageFiles: [] }
1000  const segments = commandsOf(command)
1001  const plains = segments.map(plainOf)
1002  const kindsAt = plains.map(plain => kindsIn(plain, config))
1003  // Each heredoc's first body line, in order; segmentsOf leaves `<<HEREDOC` where each one was.
1004  const heredocs = heredocsOf(command).map(m => (m[0].split('\n')[1] ?? '').trim())
1005  let seen = 0
1006  const heredocAt = segments.map(segment => {
1007    const count = segment.words.reduce((n, w) => n + w.text.split('<<HEREDOC').length - 1, 0)
1008    const first = count > 0 ? heredocs[seen] : undefined
1009    seen += count
1010    return first
1011  })
1012  const strong: boolean[] = segments.map(() => false)
1013  const line: Line = { segments, plains, kindsAt, strong, heredocAt, shapes: [], like: new Map() }
1014  const runs: Run[] = []
1015  // Every run inside a loop stays weak, whatever its own status or output showed; so does a `gh pr view` inside one.
1016  const inLoop = loopSegments(segments)
1017  const looped = new Set<Run>()
1018  kindsAt.forEach((kinds, i) => {
1019    if (kinds.length === 0) return
1020    const place = placeOf(line, i)
1021    for (const kind of kinds) {
1022      const run: Run = { kind, ...strengthOf(kind, facts, place, plains[i] ?? '') }
1023      runs.push(run)
1024      if (inLoop[i] === true) looped.add(run)
1025      else if (run.ok && !run.masked) strong[i] = true
1026    }
1027  })
1028  // A `gh pr view` that printed MERGED shows a merge, though the merge ran in an earlier call whose own output hid it
1029  // (`gh pr merge N | tail -5`, then `gh pr view N --json state`). Not a run otherwise: an OPEN PR is no merge.
1030  const view = plains.findIndex(p => PR_VIEW.test(p))
1031  if (!facts.denied && !facts.interrupted && view >= 0 && !runs.some(r => r.kind === 'merge') && VIEW_MERGED.test(facts.output)) {
1032    const run: Run = { kind: 'merge', ok: true, masked: false, basis: 'output' }
1033    runs.push(run)
1034    if (plains.every((p, k) => !PR_VIEW.test(p) || inLoop[k] === true)) looped.add(run)
1035  }
1036  // The result's gitOperation confirms an op whatever the command looked like.
1037  for (const op of SHIP_OPS) {
1038    if (facts.git[op] !== true) continue
1039    const seen = runs.filter(r => r.kind === op)
1040    if (seen.length === 0) runs.push({ kind: op, ok: true, masked: false, basis: 'gitOperation' })
1041    for (const run of seen) Object.assign(run, { ok: true, masked: false, basis: 'gitOperation' })
1042  }
1043  for (const run of looped) Object.assign(run, { ok: true, masked: true, basis: 'none' })
1044  return {
1045    mutations: facts.changed.filter(p => isCodeMutation(p, places)).map(p => pathKey(p, places)),
1046    // `gh pr merge` that only enabled auto-merge merged nothing yet.
1047    runs: facts.autoMerge === true ? runs.filter(r => r.kind !== 'merge') : runs,
1048    messageFiles: messageFilesOf(segments, plains, places.home),
1049    // A background command's names for a later read-back (full paths, so two worktrees' logs of one name do not collide).
1050    keys: facts.background ? outputKeys(command, facts.taskId ?? null, places) : [],
1051  }
1052}
1053
1054// ---- Recording and judging ----
1055
1056export const ENTRY_CAP = 100
1057export const EDIT_CAP = 20
1058export const FLAGGED_CAP = 500
1059export const PENDING_CAP = 50
1060export const STEP_CAP = 50
1061export const LINE_CAP = 5
1062export const NOTE_HEAD = "[claim-ledger plugin note, not the user's words. No reply needed.]"
1063
1064export type Verdict = { status: Status; entry: Entry | null }
1065export type Fired = { claim: Claim; status: Status }
1066export type Assessment = { lines: string[]; fired: Fired[]; repeated: Claim[]; backed: Claim[] }
1067
1068const NOUN: Record<Kind, string> = {
1069  tests: 'test run',
1070  build: 'build or type-check',
1071  commit: 'git commit',
1072  push: 'git push',
1073  merge: 'merge',
1074  'pr-create': 'gh pr create',
1075}
1076
1077const isShipKind = (kind: Kind): kind is ShipOp => kind !== 'tests' && kind !== 'build'
1078
1079export function emptyLedger(): Ledger {
1080  return { seq: 0, done: 0, calls: 0, turnFrom: 0, lastMutation: null, edits: [], entries: [], ships: {}, flagged: [], flaggedEvidence: [], pending: [], backed: [], steps: null }
1081}
1082
1083/** A hot reload: the saved ledger, unless this module already holds one. Fields an older build lacked are filled. */
1084export function restoreLedger(current: Ledger, saved: unknown): Ledger {
1085  if (current.calls > 0 || typeof saved !== 'object' || saved === null || Array.isArray(saved)) return current
1086  // A copy, because values read from the host may be frozen.
1087  const restored: Ledger = { ...emptyLedger(), ...(JSON.parse(JSON.stringify(saved)) as Partial<Ledger>) }
1088  // A ledger saved before `edits` existed keeps its last edit.
1089  if (restored.edits.length === 0 && restored.lastMutation !== null) restored.edits = [restored.lastMutation]
1090  return restored
1091}
1092
1093/** Whether a message file named to git or gh (normalized; relative as written) is this edit's file. */
1094function sameFile(edited: string, named: string): boolean {
1095  return edited === named || (!/^(?:[a-z]:\/|\/)/.test(named) && edited.endsWith(`/${named}`))
1096}
1097
1098export function shortOf(command: string): string {
1099  const one = command.replace(/\s+/g, ' ').trim()
1100  return one.length > 80 ? `${one.slice(0, 77)}...` : one
1101}
1102
1103/** Records one finished call; returns the entries it added. `seq` and `ts` are from hook entry, so order is call order. */
1104export function record(ledger: Ledger, facts: Facts, classified: Classified, seq: number, ts: number, agentId: string | null): Entry[] {
1105  ledger.calls += 1
1106  ledger.done += 1
1107  // A command's own edits are taken to come before its own runs (`sed -i ... && pytest`).
1108  const editSeq = classified.runs.length > 0 ? seq - 0.5 : seq
1109  for (const path of classified.mutations) {
1110    ledger.edits = ledger.edits.filter(m => m.path !== path)
1111    ledger.edits.push({ seq: editSeq, ts, path })
1112  }
1113  // Calls complete out of order: keep the edits in call order, one per path, newest last.
1114  ledger.edits.sort((a, b) => a.seq - b.seq)
1115  if (ledger.edits.length > EDIT_CAP) ledger.edits.splice(0, ledger.edits.length - EDIT_CAP)
1116  // A file git or gh read a commit message or PR body from was a message, not code: its write no longer counts.
1117  if (classified.messageFiles.length > 0) ledger.edits = ledger.edits.filter(m => !classified.messageFiles.some(f => sameFile(m.path, f)))
1118  const newest = ledger.edits[ledger.edits.length - 1]
1119  ledger.lastMutation = newest === undefined ? null : { ...newest }
1120  const added: Entry[] = []
1121  const keys = facts.background && facts.command !== null ? (classified.keys ?? []) : []
1122  const nth: Partial<Record<Kind, number>> = {}
1123  for (const run of classified.runs) {
1124    const n = nth[run.kind] ?? 0
1125    nth[run.kind] = n + 1
1126    const entry: Entry = {
1127      seq,
1128      done: ledger.done,
1129      ts,
1130      agentId,
1131      toolUseId: facts.toolUseId,
1132      tool: facts.tool,
1133      short: shortOf(facts.command ?? ''),
1134      kind: run.kind,
1135      ok: run.ok,
1136      background: facts.background,
1137      masked: run.masked,
1138      basis: run.basis,
1139    }
1140    if (keys.length > 0 && facts.command !== null) entry.watch = { tool: facts.tool, command: facts.command.slice(0, WATCH_CAP), keys, n }
1141    ledger.entries.push(entry)
1142    added.push(entry)
1143    if (isShipKind(run.kind) && entry.ok && !entry.masked && !entry.background) ledger.ships[run.kind] = entry
1144  }
1145  if (ledger.entries.length > ENTRY_CAP) ledger.entries.splice(0, ledger.entries.length - ENTRY_CAP)
1146  return added
1147}
1148
1149// ---- Background runs read back ----
1150
1151const WATCH_CAP = 4000
1152const BG_ID = /running in background with ID: ([\w-]+)/
1153const EXITED = /^\[exited with code (\d+)\]\s*$/m
1154const REDIRECT = /^(?:[0-9&]?>>?)(?!&)(.*)$/
1155// The Read tool numbers each line ("    12→text"); runner summaries are matched at line starts.
1156const READ_PREFIX = /^ *\d+(?:→|\t)/gm
1157
1158/**
1159 * The files a command writes (`> f`, `>> f`, `2> f`, `tee f`) and the other file-like words it names, each resolved to a
1160 * full, normalized path (a basename alone would let one worktree's `push.log` judge another's run). Resolution follows
1161 * the command's own `cd`s from the session root, expands `NAME=value` assignments made on the line, `$TEMP`, `$HOME` and
1162 * `~`; a path still holding a variable is kept as written (normalized), so a reader that names it the same way matches.
1163 */
1164function filesOf(command: string, places: Places): { written: string[]; named: string[] } {
1165  const vars = new Map<string, string>()
1166  if (places.temp !== null) {
1167    vars.set('TEMP', places.temp)
1168    vars.set('TMP', places.temp)
1169  }
1170  if (places.home !== null) {
1171    vars.set('HOME', places.home)
1172    vars.set('USERPROFILE', places.home)
1173  }
1174  const expand = (text: string): string => text.replace(/\$\{?([A-Za-z_]\w*)\}?/g, (all, name: string) => vars.get(name) ?? all)
1175  let cwd = places.root
1176  const resolve = (word: string): string => pathKey(expand(word), { ...places, root: cwd })
1177  const written = new Set<string>()
1178  const named = new Set<string>()
1179  for (const s of segmentsOf(command)) {
1180    for (const w of s.words) {
1181      const m = /^([A-Za-z_]\w*)=(\S+)$/.exec(w.text)
1182      if (m !== null && m[1] !== undefined && m[2] !== undefined) vars.set(m[1], expand(m[2]))
1183    }
1184    const word = commandWord(s)
1185    const head = headOf(s)
1186    if (CD.has(word)) {
1187      const to = s.words[head + 1]?.text
1188      if (to !== undefined && !to.startsWith('-')) cwd = resolve(to)
1189      continue
1190    }
1191    const tee = word === 'tee'
1192    for (let i = head + 1; i < s.words.length; i += 1) {
1193      const w = s.words[i]
1194      if (w === undefined) continue
1195      const m = w.quoted ? null : REDIRECT.exec(w.text)
1196      if (m !== null) {
1197        const inline = m[1] ?? ''
1198        const target = inline !== '' ? inline : s.words[i + 1]?.text
1199        if (inline === '') i += 1
1200        if (target !== undefined && FILE_LIKE.test(expand(target))) written.add(resolve(target))
hooks/stats.ts 72 lines
1import type { Claim, Counter, Counters, Family, Fire, Ledger, Pending } from '../types'
2import { judge } from './evidence'
3import type { Fired } from './evidence'
4
5export const RING_CAP = 200
6
7const FAMILIES: readonly Family[] = ['tests', 'build', 'shipped']
8const num = (n: unknown): number => (typeof n === 'number' && Number.isFinite(n) ? n : 0)
9const pct = (part: number, whole: number): string => (whole === 0 ? '-' : `${Math.round((100 * part) / whole)}%`)
10const tail = (path: string): string => path.replace(/\\/g, '/').split('/').pop() ?? path
11
12export function countersOf(value: unknown): Counters {
13  const v = (typeof value === 'object' && value !== null ? value : {}) as Record<string, Record<string, unknown> | undefined>
14  const one = (c: Record<string, unknown> | undefined): Counter => ({ fires: num(c?.fires), laterBacked: num(c?.laterBacked), repeated: num(c?.repeated) })
15  return { tests: one(v.tests), build: one(v.build), shipped: one(v.shipped) }
16}
17
18export function ringOf(value: unknown): Fire[] {
19  if (!Array.isArray(value)) return []
20  return value.filter(
21    (f): f is Fire => typeof f === 'object' && f !== null && typeof (f as Fire).input_hash === 'string' && FAMILIES.includes((f as Fire).guard),
22  )
23}
24
25export function applyFlush(
26  counters: Counters,
27  ring: readonly Fire[],
28  fired: readonly Fired[],
29  repeated: readonly Claim[],
30  backed: readonly Pending[],
31  now: number,
32): { counters: Counters; ring: Fire[] } {
33  const c: Counters = { tests: { ...counters.tests }, build: { ...counters.build }, shipped: { ...counters.shipped } }
34  const r: Fire[] = ring.map(f => ({ ...f }))
35  for (const f of fired) {
36    c[f.claim.family].fires += 1
37    r.push({ ts: now, guard: f.claim.family, input_hash: f.claim.hash, outcome: f.status })
38  }
39  for (const claim of repeated) c[claim.family].repeated += 1
40  for (const p of backed) {
41    c[p.family].laterBacked += 1
42    for (let i = r.length - 1; i >= 0; i -= 1) {
43      const f = r[i]
44      if (f !== undefined && f.input_hash === p.hash && f.laterBacked !== true) {
45        f.laterBacked = true
46        break
47      }
48    }
49  }
50  return { counters: c, ring: r.slice(-RING_CAP) }
51}
52
53export function reportText(counters: Counters, ring: readonly Fire[], ledger: Ledger, clock: (ms: number) => string): string {
54  const lines = ['Claim Ledger, all sessions:']
55  for (const family of FAMILIES) {
56    const c = counters[family]
57    lines.push(`  ${family.padEnd(8)}${c.fires} flagged, ${c.laterBacked} later backed (${pct(c.laterBacked, c.fires)}), ${c.repeated} repeated unbacked`)
58  }
59  lines.push('A tests or build flag later backed by a run, with no edit between, may have been true when made: a high rate means the check is noisy. Shipped flags are never later backed.')
60  const edit = ledger.lastMutation
61  lines.push(`This session: ${ledger.calls} tool calls; last code edit ${edit === null ? 'none' : `${clock(edit.ts)} (${tail(edit.path)})`}.`)
62  for (const kind of ['tests', 'build'] as const) {
63    const v = judge(kind, ledger)
64    lines.push(`  ${kind}: ${v.status}${v.entry === null ? '' : ` (${clock(v.entry.ts)}, ${v.entry.short})`}`)
65  }
66  const recent = ring.slice(-5)
67  if (recent.length > 0) {
68    lines.push(`Last flags: ${recent.map(f => `${clock(f.ts)} ${f.guard} ${f.outcome}${f.laterBacked === true ? ', later backed' : ''}`).join('; ')}`)
69  }
70  return lines.join('\n')
71}
72
types/index.d.ts 88 lines
1/** A claim family the detector knows. */
2export type Family = 'tests' | 'build' | 'shipped'
3
4/** The git operation a shipped claim names. */
5export type ShipOp = 'commit' | 'push' | 'merge' | 'pr-create'
6
7/** What a run is evidence for: tests, build, or one shipped op. */
8export type Kind = 'tests' | 'build' | ShipOp
9
10/** How a claim stands against the ledger. */
11export type Status = 'backed' | 'none' | 'failed' | 'masked' | 'background'
12
13/** One claim found in my text: its family, op (shipped only), the words as written, and a stable hash. */
14export type Claim = {
15  family: Family
16  op: ShipOp | null
17  phrase: string
18  hash: string
19}
20
21/** How a run's outcome was read: its exit status, an echo of that status, its visible output, the result's gitOperation, or not at all (weak). */
22export type Basis = 'exit' | 'echo' | 'output' | 'gitOperation' | 'none'
23
24/** One evidence-bearing run, recorded when it completed. `seq`: order at hook entry. `done`: order of completion. */
25export type Entry = {
26  seq: number
27  done: number
28  ts: number
29  agentId: string | null
30  toolUseId: string | null
31  tool: string
32  short: string
33  kind: Kind
34  ok: boolean
35  background: boolean
36  masked: boolean
37  basis: Basis
38  /**
39   * A background run's recipe for re-judging it when a later call reads its result back: the command, the
40   * names a reader would use (task id, basenames of files it wrote), and which run of its kind in the command it is.
41   */
42  watch?: Watch
43}
44
45export type Watch = { tool: string; command: string; keys: string[]; n: number }
46
47/** A successful code edit. `path` is normalized (pathKey). A Bash command's own edits sit half a step before its runs. */
48export type Mutation = { seq: number; ts: number; path: string }
49
50/** What the ledger knew when a step began. */
51export type AsOf = { done: number; lastMutation: Mutation | null }
52
53/** A flagged tests or build claim not yet backed. `edit`: the last edit's seq when it was flagged. */
54export type Pending = { hash: string; family: Family; op: ShipOp | null; ts: number; edit: number }
55
56/** The session's ledger: module-authoritative, mirrored to $.state. */
57export type Ledger = {
58  seq: number
59  done: number
60  calls: number
61  /** The shipped window opens after this seq: the seq when the previous answered main turn completed. */
62  turnFrom: number
63  /** The newest of `edits`; null when there are none. */
64  lastMutation: Mutation | null
65  /** Recent code edits, one per path, oldest first (capped): a commit-message file handed to git later is taken back out. */
66  edits: Mutation[]
67  entries: Entry[]
68  /** The latest strong entry of each shipped op, kept past the entry cap. */
69  ships: Partial<Record<ShipOp, Entry>>
70  flagged: string[]
71  flaggedEvidence: string[]
72  pending: Pending[]
73  backed: Pending[]
74  steps: { turnId: string; candidates: Claim[]; backed: Claim[] } | null
75}
76
77export type Counter = { fires: number; laterBacked: number; repeated: number }
78export type Counters = Record<Family, Counter>
79
80/** One entry of the 200-entry fire ring in $.store. */
81export type Fire = { ts: number; guard: Family; input_hash: string; outcome: Status; laterBacked?: true }
82
83declare module 'claude-code' {
84  interface PluginState {
85    'claim-ledger': { ledger: Ledger }
86  }
87}
88