SLOPSHOPPER

fine-print

Opens a card when Claude's final answer omits a recorded file change or a recognized command with side effects, or makes a test or Git claim the recorded…

newpaneguardcommandtoaststatus
v0.1.0MITupdated 2026-10-02bIackr0se/fine-print/plugins/fine-print
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · fine-print
│ ┃ Fine print ✕ › fix the failing auth test and add an audit log call │ ┃ FINE PRINT · turn 1 │ ┃ 2 things the summary left out · 1 claim the ⏺ Read(src/auth.ts) │ ┃ record does not back ⎿ Read 6 lines │ ┃ ⏺ Update(src/auth.ts) │ ┃ ┌─────────────────────────────────────────── ⎿ Added 2 lines, removed 1 line │ ┃ │ ✗ “I made `refresh` reject expired claims ⏺ Bash(bun test) │ ┃ │ added an audit call, and created ⎿ 3 pass, 1 fail │ ┃ │ `src/audit.ts`.” │ ┃ │ the edit did not land: result shape ● Done. refresh now rejects expired claims and logs an audit event. │ ┃ │ not recognised │ ┃ │ Summary: the record contradicts this ✻ Worked for 42s · done 4:20 PM │ ┃ │ claim │ ┃ └─────────────────────────────────────────── › /fineprint │ ┃ ⎿ fine-print: Fine print closed │ ┃ ┌─────────────────────────────────────────── │ ┃ │ ⚠ deleted files │ ┃ │ $ rm -rf build && git push --force │ ┃ │ origin main │ ┃ │ no readable exit status │ ┃ │ Summary: the summary never mentions it │ ┃ └─────────────────────────────────────────── │ ┃ │ ┃ [ Close ] Ask Claude puts one question in y │ ┃ box. You edit it and press Enter. │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⚠ fine-print: fine print: 2 things the summary left out · 1 claim the record does not back

Draws

Pane · Fine print · while holding a tool call
FINE PRINT · turn 1 2 things the summary left out · 1 claim the record does not back ┌─────────────────────────────────────────────────────────── │ ✗ “I made `refresh` reject expired claims, added an audit │ call, and created `src/audit.ts`.” │ the edit did not land: result shape not recognised │ Summary: the record contradicts this claim [ Ask Claud └─────────────────────────────────────────────────────────── ┌──────────────────────────────────────────────────────────┐ │ ⚠ deleted files │ │ $ rm -rf build && git push --force origin main │ │ no readable exit status │ │ Summary: the summary never mentions it [ Ask Claude ]│ └──────────────────────────────────────────────────────────┘ [ Close ] Ask Claude puts one question in your prompt box. and press Enter.
README

Fine Print

See the changes Claude left out of its answer.

Fine Print synthetic demo

Reconstructed UI on a synthetic repository. The callout explains the example; the mod shows the recorded evidence.

Fine Print opens a card when Claude's final answer omits a recorded file change or a recognized command with side effects, or makes a test or Git claim the recorded results do not fully back. Inspect the recorded evidence, step through file diffs, and prepare a question about the selected item. You edit and send the question yourself.

It runs inside Claude Code and makes no model, network, filesystem or process calls of its own. It observes the supported tool events from the main conversation. It cannot establish that every action or natural-language claim has been covered.

Install

Download Fine Print 0.1.0. Extract it, then follow the steps below.

Requires Claude Code 2.1.287 or later. Tested with 2.1.287 on macOS in Claude Desktop's local Code tab. Keep the extracted folder somewhere permanent: the installer registers it as a local plugin source. From that folder, run:

sh install.sh

This adds the local jr-mods marketplace and installs Fine Print for your user account. Start a new local session in Claude Desktop's Code tab. The mod opens automatically when it finds an item. Type /fineprint to open or close the card yourself.

To enable Fine Print in one repository only, run sh /absolute/path/to/install.sh local from that repository. The jr-mods marketplace is still registered for your user account. The terminal and local Desktop Code sessions share Claude Code's plugin settings. This package does not add the card to ordinary Claude chat, Cowork, the VS Code chat panel, or cloud Code sessions.

Controls

  • Full diff: step through the recorded hunks for omitted file changes.
  • Ask Claude: prepare an editable question about one item. Existing drafts are preserved. Nothing is submitted automatically.
  • /fineprint off: open the card only when you request it.
  • /fineprint on: restore automatic cards.

To disable a user installation:

claude plugin disable fine-print@jr-mods --scope user

For a local installation, use --scope local from the repository where you enabled it. To uninstall a user installation:

claude plugin uninstall fine-print@jr-mods --scope user
claude plugin marketplace remove jr-mods --scope user

For a local installation, replace the uninstall command's --scope user with --scope local, while keeping the marketplace removal at user scope. Remove the marketplace only when you no longer use it in any repository.

The card reads recorded events, not the live working tree. A path mention is evidence that a file was named, not proof that the answer explained its consequences. Unsupported command syntax appears as a coverage gap. Subagent events and earlier turns do not back the current main-turn audit.

Missing evidence shows as a grey question mark, partial support as an amber marker, and an observed contradiction as a red cross. Clearly framed double-quoted examples are excluded from claim matching; unframed quoted claims remain checkable. These are bounded text rules, not a complete interpretation of every possible answer.

Installation behavior follows the official plugin documentation and mods documentation.

Development

No build step or dependencies are required for the mod. With Claude Code 2.1.287 or later:

claude plugin validate plugins/fine-print
claude plugin test plugins/fine-print

Claude Code generates its version-specific type declarations locally. They are not redistributed here. The test suite covers audit matching, turn isolation, draft preservation and UI lifecycle behavior. Native consumer checks and their platform/version are recorded in VERIFICATION.md.

License

MIT.

Source 6 files
hooks/register.tsx 233 lines
1import type { On } from 'claude-code'
2
3import { type FinePrint, finePrintOf, headlineOf, questionFor } from './audit'
4import { eventOf, type ToolEvent } from './ledger'
5import { cardView, changesOf, type View } from './view'
6
7export const COMMAND = 'fineprint'
8export const PANE_ID = 'fine-print'
9
10/** Housekeeping bound on memory, not on what is shown. */
11const MAX_EVENTS = 5_000
12
13type Host = {
14  surfaces: () => Promise<readonly string[]>
15  status: (text: string | undefined) => void
16  log: (text: string) => void
17  debug: (text: string) => void
18  toast: (text: string) => void
19  redraw: () => void
20  openPane: () => Promise<void>
21  closePane: () => Promise<void>
22  /** The prompt box's current draft (empty where the session draws none). */
23  draft: () => Promise<string>
24  fill: (text: string, mode: 'replace' | 'append') => Promise<boolean>
25  /** Runs work outside the dispatch that asked for it (a press), so it is not cut short with that dispatch. */
26  later: (fn: () => void) => void
27}
28
29/**
30 * Fine Print: at the end of each turn, compares Claude's final answer with
31 * what the turn's tool calls actually did, and when the answer left something
32 * out (a change it never named, a command with side effects it never
33 * mentioned, a claim the record does not back) opens a card with it, a
34 * step-through of the unmentioned diffs, and a question per item for the prompt box.
35 * Observes only: no model, network or file calls; nothing is sent for you.
36 */
37export function register(on: On) {
38  let host: Host | null = null
39  let isAuto = true
40  let isOpen = false
41  /** Turns are counted by their ends: desktop sessions raise turn.complete but no turn.start. */
42  let turn = 1
43  let seq = 0
44  let events: ToolEvent[] = []
45  let last: FinePrint | null = null
46  let view: View = { mode: 'list' }
47  let cwd = ''
48  /** Bumped by every new card, /clear and /resume: a deferred Ask from an older card does nothing. */
49  let generation = 0
50  /** Only the latest open or close may settle the pane's state. */
51  let paneOp = 0
52
53  async function open() {
54    if (!host) return
55    const op = ++paneOp
56    await host.openPane().catch(() => undefined)
57    if (op === paneOp) isOpen = true
58    host.redraw()
59  }
60  async function close(): Promise<boolean> {
61    const h = host
62    if (!h || !isOpen) return true
63    const op = ++paneOp
64    try {
65      await h.closePane()
66      if (op === paneOp) isOpen = false
67      return true
68    } catch (error) {
69      h.toast('Could not close Fine Print. Try /fineprint again.')
70      h.debug(`fine-print: close failed: ${error instanceof Error ? error.message : String(error)}`)
71      return false
72    }
73  }
74  const actions = {
75    review: (index: number) => { view = { mode: 'review', index }; host?.redraw() },
76    back: () => { view = { mode: 'list' }; host?.redraw() },
77    step: (delta: number) => {
78      if (view.mode !== 'review' || !last) return
79      const n = changesOf(last).length
80      view = { mode: 'review', index: (view.index + delta + n) % n }
81      host?.redraw()
82    },
83    close: () => {
84      const closing = generation
85      host?.later(() => { if (closing === generation) void close() })
86    },
87    ask: (index: number) => {
88      const h = host
89      const card = last
90      const item = card?.items[index]
91      if (!h || !card || !item) return
92      const asked = generation
93      const question = questionFor(item)
94      h.later(() => {
95        void (async () => {
96          if (asked !== generation || last !== card) {
97            h.toast('That card is out of date. Ask again from the current one.')
98            return
99          }
100          const showQuestion = (why: string) => {
101            view = { mode: 'question', index }
102            h.redraw()
103            h.toast(why)
104          }
105          let draft: string
106          try {
107            draft = (await h.draft()).trim()
108          } catch {
109            showQuestion('Could not read your prompt box; the question is shown in the card.')
110            return
111          }
112          // the card may have been replaced while the box was being read
113          if (asked !== generation || last !== card) {
114            h.toast('That card is out of date. Ask again from the current one.')
115            return
116          }
117          const filled = await h.fill(draft ? `\n\n${question}` : question, draft ? 'append' : 'replace').catch(() => false)
118          if (asked !== generation) return
119          if (filled) {
120            h.toast(draft ? 'The question was added after your draft: edit it, then press Enter.' : 'The question is in your prompt box: edit it, then press Enter.')
121            await close()
122          } else {
123            showQuestion('Could not reach the prompt box; the question is shown in the card.')
124          }
125        })()
126      })
127    },
128  }
129
130  on('session.start', async ($, e, next) => {
131    try {
132      await $.command.register({ name: COMMAND, description: "Show what Claude's last answer left out: unmentioned changes, side-effect commands, unbacked claims", argumentHint: '[off | on]' })
133      host = {
134        surfaces: () => $.session.surfaces(),
135        status: text => $.ui.status(text),
136        log: text => $.ui.log(text),
137        debug: text => $.ui.log(text, { to: 'debug' }),
138        toast: text => $.ui.toast(text),
139        redraw: () => $.ui.invalidate('ui.render'),
140        openPane: () => $.ui.open({ id: PANE_ID, title: 'Fine print' }),
141        closePane: () => $.ui.close({ id: PANE_ID }),
142        draft: async () => (await $.prompt.read()).text,
143        fill: async (text, mode) => (await $.prompt.fill({ text, mode })).isFilled,
144        later: fn => { $.clock.after(0, fn) },
145      }
146      cwd = e.cwd
147    } catch (error) {
148      $.ui.log(`fine-print: /${COMMAND} not registered: ${error instanceof Error ? error.message : String(error)}`, { to: 'debug' })
149    }
150    return next(e)
151  })
152
153  on('tool.call', { tool: ['Edit', 'Write', 'NotebookEdit', 'MultiEdit', 'Bash', 'Read'] }, async ($, e, next) => {
154    const atMs = await $.clock.now()
155    let result: unknown
156    try {
157      result = await next(e)
158      return result as Awaited<ReturnType<typeof next>>
159    } finally {
160      const event = eventOf({ tool: e.tool, input: e as unknown as Record<string, unknown>, result }, { seq: ++seq, turn, atMs })
161      if (event) {
162        events.push(event)
163        if (events.length > MAX_EVENTS) events = events.slice(-MAX_EVENTS)
164      }
165    }
166  })
167
168  on('turn.complete', async ($, e, next) => {
169    if (e.agentId !== undefined) return next(e)
170    const h = host
171    const ended = turn
172    turn += 1
173    if (!h) return next(e)
174    generation += 1
175    last = finePrintOf(e.answer, events, ended)
176    view = { mode: 'list' }
177    const shown = last.items.length > 0
178
179    h.status(shown ? `fine print: ${headlineOf(last.items)}` : undefined)
180    const surfaces = await h.surfaces().catch(() => [] as readonly string[])
181    if (!surfaces.length) {
182      if (shown) h.log(`fine print · ${headlineOf(last.items)}\n${last.items.map((it, i) => `${i + 1}. ${questionFor(it)}`).join('\n')}`)
183    } else if (shown && isAuto) {
184      await open()
185    } else if (isOpen) {
186      h.redraw()
187    }
188    return next(e)
189  })
190
191  on('ui.render', { component: 'Pane' }, async ($, e, next) => {
192    if (e.requestId !== PANE_ID) return next(e)
193    const { Box, Text, Button } = $.ui.resolve(e)
194    return cardView({ Box, Text, Button }, last, view, actions, cwd)
195  })
196
197  on('command.run', { command: COMMAND }, async ($, e, next) => {
198    const h = host
199    if (!h) return next(e)
200    const arg = e.args.trim().toLowerCase()
201    if (arg === 'off' || arg === 'on') {
202      isAuto = arg === 'on'
203      return { text: `Fine print ${isAuto ? 'opens by itself when an answer leaves something out' : 'opens only on /fineprint'}` }
204    }
205    const surfaces = await h.surfaces().catch(() => [] as readonly string[])
206    if (!surfaces.length) return { text: last?.items.length ? `${headlineOf(last.items)}\n${last.items.map((it, i) => `${i + 1}. ${questionFor(it)}`).join('\n')}` : 'No fine print in the last answer' }
207    if (isOpen) {
208      return { text: (await close()) ? 'Fine print closed' : 'Fine print could not be closed' }
209    }
210    view = { mode: 'list' }
211    await open()
212    return { text: 'Fine print shown' }
213  })
214
215  on('ui.close', { id: PANE_ID }, async ($, e, next) => {
216    const result = await next(e)
217    if ((result as { deny?: string } | undefined)?.deny === undefined) isOpen = false
218    return result
219  })
220
221  on('command.run', { command: ['clear', 'resume'] }, async ($, e, next) => {
222    await close()
223    generation += 1
224    turn = 1
225    seq = 0
226    events = []
227    last = null
228    view = { mode: 'list' }
229    host?.status(undefined)
230    return next(e)
231  })
232}
233
hooks/audit.ts 213 lines
1/**
2 * The fine print of one turn: what actually happened that Claude's final
3 * answer did not say. Pure. Built only from the turn's recorded tool calls
4 * (ledger.ts) and the answer's own words; a thing the answer names is not
5 * fine print, and a turn whose answer covers everything has none.
6 */
7import { claimsOf, filesIn } from './claims'
8import { asksForNoRun, programOf, stepsOf } from './command'
9import { type Receipt, receiptsOf, type ToolEvent } from './ledger'
10
11type Edit = Extract<ToolEvent, { kind: 'edit' }>
12type Shell = Extract<ToolEvent, { kind: 'shell' }>
13
14export type Item =
15  | { kind: 'claim'; receipt: Receipt }
16  | { kind: 'change'; path: string; why: string | null; edits: Edit[]; shared: string | null }
17  | { kind: 'command'; command: string; effect: string; ok: boolean | null; observed: boolean | null; result: string; at: number }
18
19export type FinePrint = { turn: number; items: Item[] }
20
21/** Paths whose silent change matters more than most, with the reason shown beside them. */
22const SENSITIVE: [RegExp, string][] = [
23  [/(?:^|\/)\.env(?:\.[\w.-]+)?$|(?:^|\/)(?:secrets?|credentials?)(?:\.[\w]+)?$|\.(?:pem|key|p12)$/i, 'secrets'],
24  [/(?:^|\/)(?:package-lock\.json|yarn\.lock|pnpm-lock\.yaml|bun\.lockb?|poetry\.lock|uv\.lock|Cargo\.lock|go\.sum|Gemfile\.lock|composer\.lock)$/, 'lockfile'],
25  [/(?:^|\/)(?:package\.json|pyproject\.toml|requirements[\w.-]*\.txt|go\.mod|Cargo\.toml|Gemfile|composer\.json)$/, 'dependencies'],
26  [/(?:^|\/)\.github\/workflows\/|(?:^|\/)\.gitlab-ci\.yml$|(?:^|\/)\.circleci\/|(?:^|\/)Jenkinsfile$/, 'CI'],
27  [/(?:^|\/)migrations?\//i, 'migration'],
28  [/(?:^|\/)(?:Dockerfile|docker-compose[\w.-]*\.ya?ml|[\w.-]+\.tf)$/, 'infrastructure'],
29  [/(?:^|\/)(?:auth|security|permissions?|policy)[\w.-]*\.\w+$/i, 'auth'],
30]
31
32export const sensitivityOf = (path: string): string | null => SENSITIVE.find(([re]) => re.test(path))?.[1] ?? null
33
34/** Side effects a step can have, and the words an answer would use if it told you. */
35const EFFECTS: { label: string; is: (p: string[]) => boolean; told: RegExp }[] = [
36  { label: 'pushed to a remote', is: p => p[0] === 'git' && p.includes('push'), told: /\bpush(?:ed|ing)?\b/i },
37  { label: 'rewrote git history or discarded work', is: p => p[0] === 'git' && (p.includes('reset') && p.includes('--hard') || p.includes('clean') || (p.includes('checkout') && p.includes('--')) || p.includes('restore') || p.includes('rebase') || p.includes('--amend') || (p.includes('stash') && p.includes('drop')) || (p.includes('branch') && p.includes('-D'))), told: /\b(?:reset|clean(?:ed)?|discard(?:ed)?|restor(?:e|ed)|rebas(?:e|ed)|amend(?:ed)?|drop(?:ped)?)\b/i },
38  { label: 'deleted files', is: p => p[0] === 'rm' || p[0] === 'rmdir' || (p[0] === 'find' && p.includes('-delete')), told: /\b(?:delet(?:e|ed|ing)|remov(?:e|ed|ing)|rm)\b/i },
39  { label: 'installed packages', is: p => (['npm', 'pnpm', 'yarn', 'bun'].includes(p[0] ?? '') && ['install', 'i', 'add'].includes(p[1] ?? '')) || ((p[0] === 'pip' || p[0] === 'pip3') && p[1] === 'install') || (p[0] === 'uv' && (p[1] === 'add' || (p[1] === 'pip' && p[2] === 'install'))) || (p[0] === 'poetry' && p[1] === 'add') || (p[0] === 'brew' && p[1] === 'install') || (p[0] === 'cargo' && (p[1] === 'add' || p[1] === 'install')) || (p[0] === 'go' && p[1] === 'get'), told: /\binstall(?:ed|ing)?\b|\badd(?:ed)? (?:the )?(?:package|dependency|dep)/i },
40  { label: 'ran a script from the network', is: p => p[0] === 'sh' || p[0] === 'bash', told: /\binstall(?:er|ed)?\b|\bscript\b/i },
41  { label: 'changed containers or clusters', is: p => (p[0] === 'docker' && ['run', 'rm', 'rmi', 'push', 'compose', 'system'].includes(p[1] ?? '')) || (p[0] === 'kubectl' && ['apply', 'delete', 'rollout', 'scale'].includes(p[1] ?? '')) || (p[0] === 'terraform' && ['apply', 'destroy'].includes(p[1] ?? '')), told: /\b(?:docker|container|kubectl|cluster|deploy(?:ed)?|terraform|appl(?:y|ied)|destroy(?:ed)?)\b/i },
42  { label: 'acted on GitHub', is: p => p[0] === 'gh' && ((p[1] === 'pr' && ['create', 'merge', 'close', 'comment'].includes(p[2] ?? '')) || (p[1] === 'release' && p[2] === 'create') || (p[1] === 'issue' && ['create', 'close', 'comment'].includes(p[2] ?? '')) || (p[1] === 'repo' && p[2] === 'delete')), told: /\b(?:PR|pull request|merged|release|issue|comment(?:ed)?)\b/i },
43  { label: 'ran with sudo', is: p => p[0] === 'sudo', told: /\bsudo\b/i },
44]
45
46/** File tokens an answer names: dotted names with paths, dotfiles, and anything in backticks that looks like a path. */
47export function namesIn(answer: string): string[] {
48  const names = new Set(filesIn(answer))
49  for (const m of answer.matchAll(/`([^`\s]+)`/g)) if (/[./]/.test(m[1] ?? '')) names.add((m[1] ?? '').replace(/^\.\//, ''))
50  for (const m of answer.matchAll(/(?:^|[\s(`'"])((?:[\w.@-]+\/)*\.[\w][\w.-]*)(?=[\s`'"),.:;!?]|$)/g)) names.add(m[1] ?? '')
51  return [...names].filter(Boolean)
52}
53
54/** Does this answer token name this path: the whole path, or a suffix ending on a segment boundary? */
55const tokenMatches = (token: string, path: string) => path === token || path.endsWith('/' + token)
56
57/**
58 * Which of the turn's changed paths the answer names, and which it names only
59 * through a token that matches several of them (a bare `config.ts` for
60 * src/config.ts and tests/config.ts names neither; it is reported).
61 */
62export function namingOf(answer: string, paths: string[]): { named: Set<string>; shared: Map<string, string> } {
63  const named = new Set<string>()
64  const shared = new Map<string, string>()
65  for (const token of namesIn(answer)) {
66    const hits = paths.filter(p => tokenMatches(token, p))
67    if (hits.length === 1) named.add(hits[0] as string)
68    else if (hits.length > 1) for (const h of hits) if (!shared.has(h)) shared.set(h, token)
69  }
70  for (const p of named) shared.delete(p)
71  return { named, shared }
72}
73
74/** Does the answer name this path unambiguously among the given changed paths (default: just this one)? */
75export function isNamed(answer: string, path: string, among: string[] = [path]): boolean {
76  return namingOf(answer, among.includes(path) ? among : [...among, path]).named.has(path)
77}
78
79/** The text inside each top-level `$( … )` and backtick pair; null when the nesting is deeper or unbalanced. */
80export function substitutionsOf(command: string): string[] | null {
81  const inner: string[] = []
82  let quote: "'" | null = null
83  for (let i = 0; i < command.length; i++) {
84    const c = command[i]
85    if (quote) { if (c === quote) quote = null; continue }
86    if (c === "'") { quote = c; continue }
87    if (c === '\\') { i++; continue }
88    if (c === '$' && command[i + 1] === '(') {
89      let depth = 1
90      let j = i + 2
91      for (; j < command.length && depth > 0; j++) {
92        if (command[j] === '(') depth++
93        else if (command[j] === ')') depth--
94      }
95      if (depth !== 0) return null
96      const body = command.slice(i + 2, j - 1)
97      if (body.includes('$(') || body.includes('`')) return null
98      inner.push(body)
99      i = j - 1
100    } else if (c === '`') {
101      const end = command.indexOf('`', i + 1)
102      if (end < 0) return null
103      inner.push(command.slice(i + 1, end))
104      i = end
105    }
106  }
107  return inner
108}
109
110function topEffectOf(command: string): string | null {
111  const steps = stepsOf(command)
112  const piped = steps.some((s, i) => i > 0 && s.joiner === '|' && ['sh', 'bash', 'zsh'].includes(programOf(s.argv)[0] ?? '') && ['curl', 'wget'].includes(programOf(steps[i - 1]?.argv ?? [])[0] ?? ''))
113  if (piped) return 'ran a script from the network'
114  for (const s of steps) {
115    const p = programOf(s.argv)
116    if (asksForNoRun(p.slice(1), p[0] === 'git' && p.includes('push') ? /^-n$/ : null)) continue
117    const hit = EFFECTS.find(e => e.label !== 'ran a script from the network' && e.is(p))
118    if (hit) return hit.label
119  }
120  return null
121}
122
123export const UNREADABLE = 'a command substitution Fine Print cannot read'
124
125/** The side effect a recorded command had, reading one level of `$( … )` or backticks; UNREADABLE for deeper nesting. */
126function effectOf(command: string): string | null {
127  const inner = substitutionsOf(command)
128  if (inner === null) return UNREADABLE
129  for (const body of inner) {
130    const hit = topEffectOf(body)
131    if (hit) return hit
132  }
133  return topEffectOf(command)
134}
135
136/** For effects whose output shows them happening, whether it did; null where output cannot tell. */
137const SHOWN: Record<string, RegExp> = {
138  'pushed to a remote': /^To \S+[\s\S]*->/m,
139  'acted on GitHub': /https?:\/\/\S+\/(?:pull|issues|releases)\/|✓ Merged|Merged pull request/,
140}
141
142function resultOf(e: Shell, effect: string): { ok: boolean | null; observed: boolean | null; result: string } {
143  const ok = e.exit.kind === 'code' ? e.exit.code === 0 : null
144  const exit = e.exit.kind === 'code' ? `exit ${e.exit.code}` : e.exit.kind === 'unknown' ? `no readable exit status` : e.exit.kind
145  const shown = SHOWN[effect]
146  const observed = ok && shown ? shown.test(e.output) : null
147  const tail = e.output.split('\n').map(l => l.trim()).filter(Boolean).slice(-2).join(' ⏎ ')
148  const result = observed === false ? `${exit} · the output shows no sign it happened` : `${exit}${tail ? ` · ${tail.length > 90 ? tail.slice(0, 89) + '…' : tail}` : ''}`
149  return { ok, observed, result }
150}
151
152const short = (cmd: string) => (cmd.length > 60 ? cmd.slice(0, 59) + '…' : cmd)
153
154export function finePrintOf(answer: string, events: ToolEvent[], turn: number): FinePrint {
155  const items: Item[] = []
156  // only what the main loop did in this turn: a subagent's work and older turns are not this answer's to name
157  const now = events.filter(e => e.turn === turn && e.agentId === null)
158
159  for (const receipt of receiptsOf(claimsOf(answer), now, turn, 'this turn')) {
160    // a test or git claim with nothing behind it is fine print too (shown as unbacked, not as failed)
161    const matters = receipt.verdict === 'contradicted' || (receipt.claim.kind !== 'file' && (receipt.verdict === 'partial' || receipt.verdict === 'unknown'))
162    if (matters) items.push({ kind: 'claim', receipt })
163  }
164
165  const byPath = new Map<string, Edit[]>()
166  for (const e of now) {
167    if (e.kind !== 'edit' || !e.landed || e.path === null) continue
168    byPath.set(e.path, [...(byPath.get(e.path) ?? []), e])
169  }
170  const changes: Item[] = []
171  const { named, shared } = namingOf(answer, [...byPath.keys()])
172  for (const [path, edits] of byPath) {
173    if (!named.has(path)) changes.push({ kind: 'change', path, why: sensitivityOf(path), edits, shared: shared.get(path) ?? null })
174  }
175  changes.sort((a, b) => Number(b.kind === 'change' && b.why !== null) - Number(a.kind === 'change' && a.why !== null))
176  items.push(...changes)
177
178  for (const e of now) {
179    if (e.kind !== 'shell') continue
180    const effect = effectOf(e.command)
181    if (!effect) continue
182    const told = EFFECTS.find(x => x.label === effect)?.told
183    if (effect !== UNREADABLE && told?.test(answer)) continue
184    items.push({ kind: 'command', command: e.command, effect, ...resultOf(e, effect), at: e.seq })
185  }
186
187  return { turn, items }
188}
189
190const WAIT = 'Do not change or run anything yet; wait for my OK.'
191
192/** The question for one item: what happened, why, and the smallest correction that keeps the rest of my work. */
193export function questionFor(item: Item): string {
194  if (item.kind === 'claim') {
195    return `You said "${item.receipt.claim.quote}", but ${item.receipt.evidence}. What is actually true right now? If something needs fixing, propose the smallest fix. ${WAIT}`
196  }
197  if (item.kind === 'change') {
198    const hunk = item.edits.flatMap(e => e.patch).slice(0, 40)
199    return `This turn you ${item.edits.some(e => e.created) ? 'created' : 'changed'} \`${item.path}\`${item.why ? ` (${item.why})` : ''}, and your summary ${item.shared ? `named only \`${item.shared}\`, which matches more than one file you changed` : 'did not mention it'}.\n\`\`\`diff\n${hunk.join('\n')}\n\`\`\`\nWhy was this change needed? If it was not part of what I asked, propose the smallest correction that keeps the rest of my work. ${WAIT}`
200  }
201  return `This turn you ran \`${item.command}\` (${item.effect}; ${item.result}), and your summary did not mention it. Why did you run it, and what did it change? If anything needs putting right, propose the smallest correction that keeps the rest of my work. ${WAIT}`
202}
203
204export function headlineOf(items: Item[]): string {
205  const n = items.length
206  const claims = items.filter(i => i.kind === 'claim').length
207  const sensitive = items.filter(i => i.kind === 'change' && i.why).length
208  const parts = [`${n} thing${n === 1 ? '' : 's'} the summary left out`]
209  if (claims) parts.push(`${claims} claim${claims === 1 ? '' : 's'} the record does not back`)
210  if (sensitive) parts.push(`${sensitive} sensitive file${sensitive === 1 ? '' : 's'}`)
211  return parts.join(' · ')
212}
213
hooks/ledger.ts 351 lines
1/**
2 * The session's record of what tools actually did, and the receipts it gives
3 * each claim. Pure. Reads only the result shapes recorded on Claude Code
4 * 2.1.286 (demo/probe-shapes.md); any other shape is kept as `unknown` and
5 * drawn as such, never guessed into a pass or a fail.
6 */
7import type { Claim } from './claims'
8import { asksForNoRun, decidingStep, programOf, stepsOf } from './command'
9
10export const EDIT_TOOLS = ['Edit', 'Write', 'NotebookEdit', 'MultiEdit'] as const
11export const SHELL_TOOLS = ['Bash'] as const
12export const READ_TOOLS = ['Read'] as const
13
14export type Exit =
15  | { kind: 'code'; code: number }
16  | { kind: 'interrupted' }
17  | { kind: 'denied'; reason: string }
18  | { kind: 'unknown'; reason: string }
19
20export type ToolEvent =
21  | { seq: number; turn: number; atMs: number; agentId: string | null; kind: 'edit'; path: string | null; landed: boolean; reason: string | null; created: boolean; patch: string[] }
22  | { seq: number; turn: number; atMs: number; agentId: string | null; kind: 'read'; path: string | null }
23  | { seq: number; turn: number; atMs: number; agentId: string | null; kind: 'shell'; command: string; exit: Exit; output: string }
24
25type Raw = { tool: string; input: Record<string, unknown>; result: unknown }
26
27const isRecord = (v: unknown): v is Record<string, unknown> => typeof v === 'object' && v !== null && !Array.isArray(v)
28const str = (v: unknown): string | null => (typeof v === 'string' ? v : null)
29const firstLine = (s: string) => s.split('\n')[0] ?? ''
30
31/** The diff lines an Edit or Write result recorded: its structuredPatch hunks, or a new file's content as additions. */
32export function patchOf(inner: Record<string, unknown>): string[] {
33  const hunks = Array.isArray(inner.structuredPatch) ? inner.structuredPatch : []
34  const lines: string[] = []
35  for (const hunk of hunks) {
36    if (!isRecord(hunk) || !Array.isArray(hunk.lines)) continue
37    if (lines.length) lines.push('…')
38    for (const l of hunk.lines) if (typeof l === 'string') lines.push(l)
39  }
40  if (!lines.length && inner.type === 'create' && typeof inner.content === 'string') {
41    for (const l of inner.content.replace(/\n$/, '').split('\n')) lines.push('+' + l)
42  }
43  return lines
44}
45
46/** Turns one finished tool call into an event, by the recorded shapes only. */
47export function eventOf(raw: Raw, stamp: { seq: number; turn: number; atMs: number }): ToolEvent | null {
48  const agentId = typeof raw.input.agentId === 'string' ? raw.input.agentId : null
49  const at = { ...stamp, agentId }
50  const r = isRecord(raw.result) ? raw.result : {}
51  const isError = r.isError === true
52  const deny = str(r.deny)
53
54  if ((EDIT_TOOLS as readonly string[]).includes(raw.tool)) {
55    const inner = isRecord(r.result) ? r.result : null
56    const path = str(inner?.filePath) ?? str(raw.input.file_path) ?? str(raw.input.notebook_path)
57    if (deny !== null) return { ...at, kind: 'edit', path, landed: false, reason: `refused: ${deny}`, created: false, patch: [] }
58    if (isError) return { ...at, kind: 'edit', path, landed: false, reason: firstLine(str(r.result) ?? str(r.text) ?? 'error').replace(/^Error:\s*/, ''), created: false, patch: [] }
59    const landed = inner !== null && str(inner.filePath) !== null
60    return { ...at, kind: 'edit', path, landed, reason: landed ? null : 'result shape not recognised', created: inner?.type === 'create', patch: inner ? patchOf(inner) : [] }
61  }
62
63  if ((READ_TOOLS as readonly string[]).includes(raw.tool)) {
64    const inner = isRecord(r.result) ? r.result : null
65    const file = inner && isRecord(inner.file) ? inner.file : null
66    return { ...at, kind: 'read', path: str(file?.filePath) ?? str(raw.input.file_path) }
67  }
68
69  if ((SHELL_TOOLS as readonly string[]).includes(raw.tool)) {
70    const command = str(raw.input.command) ?? ''
71    if (deny !== null) return { ...at, kind: 'shell', command, exit: { kind: 'denied', reason: deny }, output: '' }
72    if (isError) {
73      const text = str(r.text) ?? ''
74      const m = /^Exit code (\d+)\n?([\s\S]*)$/.exec(text)
75      return m
76        ? { ...at, kind: 'shell', command, exit: { kind: 'code', code: Number(m[1]) }, output: m[2] ?? '' }
77        : { ...at, kind: 'shell', command, exit: { kind: 'unknown', reason: firstLine(text) || 'error without an exit code' }, output: text }
78    }
79    const inner = isRecord(r.result) ? r.result : null
80    if (inner && typeof inner.stdout === 'string' && typeof inner.interrupted === 'boolean') {
81      const output = [inner.stdout, str(inner.stderr) ?? ''].filter(Boolean).join('\n')
82      return { ...at, kind: 'shell', command, exit: inner.interrupted ? { kind: 'interrupted' } : { kind: 'code', code: 0 }, output }
83    }
84    return { ...at, kind: 'shell', command, exit: { kind: 'unknown', reason: 'result shape not recognised' }, output: '' }
85  }
86
87  return null
88}
89
90// --- receipts ---------------------------------------------------------------
91
92export type Verdict = 'backed' | 'contradicted' | 'partial' | 'unknown'
93
94export type Receipt = {
95  claim: Claim
96  verdict: Verdict
97  /** One line of evidence, as the strip draws it. */
98  evidence: string
99  /** The event it rests on, when there is one. */
100  at: number | null
101}
102
103/** Is `program` (from programOf) a test runner, and with which arguments after the runner's own words? */
104function testArgsOf(program: string[]): string[] | null {
105  const [name, ...args] = program
106  const sub = args[0]
107  switch (name) {
108    case 'pytest': case 'py.test': case 'jest': case 'vitest': case 'mocha': case 'ava': case 'tap':
109    case 'tox': case 'nox': case 'rspec': case 'phpunit': case 'ctest': case 'unittest':
110      return args
111    case 'go': case 'swift': case 'dotnet': case 'deno': case 'bun':
112      return sub === 'test' ? args.slice(1) : null
113    case 'cargo':
114      return sub === 'test' || sub === 'nextest' ? args.slice(sub === 'nextest' && args[1] === 'run' ? 2 : 1) : null
115    case 'npm': case 'pnpm': case 'yarn':
116      return sub === 'test' || sub === 't' ? args.slice(1) : sub === 'run' && args[1] === 'test' ? args.slice(2) : null
117    case 'make':
118      return sub === 'test' || sub === 'check' ? args.slice(1) : null
119    case 'claude':
120      return sub === 'plugin' && args[1] === 'test' ? args.slice(2) : null
121    case 'gradle': case 'gradlew': case './gradlew': case 'mvn': case 'mvnw':
122      return args.includes('test') ? args.filter(a => a !== 'test') : null
123    default:
124      return null
125  }
126}
127
128/**
129 * How a recorded command relates to running tests: `run` when the step that
130 * decides the exit status is a real test run, `ambiguous` when a runner is
131 * a step but another program may decide the status, `no-run` when the only
132 * runner asks for help, a version, a dry run or a listing, `none` otherwise.
133 */
134export function testRunOf(command: string): { kind: 'run'; args: string[] } | { kind: 'ambiguous' | 'no-run' | 'none' } {
135  const deciding = decidingStep(command)
136  if (deciding) {
137    const args = testArgsOf(deciding.program)
138    if (args) return asksForNoRun(args) ? { kind: 'no-run' } : { kind: 'run', args }
139  }
140  const anyRunner = stepsOf(command).some(s => testArgsOf(programOf(s.argv)) !== null)
141  return { kind: anyRunner ? 'ambiguous' : 'none' }
142}
143
144export const isTestCommand = (command: string) => testRunOf(command).kind === 'run'
145
146/** Arguments that narrow a run to part of the suite: a filter flag or a path. */
147const narrowing = (args: string[]) =>
148  args.filter((a, i) => /^(?:-k|-m|--filter|-t|--testNamePattern|--grep|-run|--test)$/.test(a) || (/[/:]|\.(?:py|ts|tsx|js|jsx|go|rs|rb)$/.test(a) && !/^-/.test(a) && !/^(?:-k|-m|--filter|-t|--grep|-run)$/.test(args[i - 1] ?? '')))
149
150export type Counts = { passed: number | null; failed: number | null }
151
152/** What the runner itself printed in its summary; nulls when it printed no summary Receipts can read. Test names are never counted. */
153export function countsOf(output: string): Counts {
154  const unittest = /^Ran (\d+) tests? in [\d.]+s\s*$/m.exec(output)
155  if (unittest) {
156    if (/^OK\b/m.test(output)) return { passed: Number(unittest[1]), failed: 0 }
157    const f = /^FAILED \((?:failures=(\d+))?(?:, )?(?:errors=(\d+))?/m.exec(output)
158    if (f) {
159      const failed = Number(f[1] ?? 0) + Number(f[2] ?? 0)
160      return { passed: Number(unittest[1]) - failed, failed }
161    }
162  }
163  const lines = output.split('\n')
164  const tally = (text: string): Counts => {
165    let passed: number | null = null
166    let failed: number | null = null
167    for (const m of text.matchAll(/(\d+)\s+(passed|passing|pass|failed|failing|fail|failures?)\b/gi)) {
168      const n = Number(m[1])
169      if (/^pass/i.test(m[2] ?? '')) passed = (passed ?? 0) + n
170      else failed = (failed ?? 0) + n
171    }
172    return { passed, failed }
173  }
174  const lastOf = (re: RegExp) => [...lines].reverse().find(l => re.test(l))
175  // bun / claude plugin test: " 42 pass" and " 0 fail" each on a line of its own
176  const pass = lastOf(/^\s*\d+\s+pass(?:ed)?\s*$/i)
177  const fail = lastOf(/^\s*\d+\s+fail(?:ed)?\s*$/i)
178  if (pass || fail) return { passed: pass ? tally(pass).passed : null, failed: fail ? tally(fail).failed : null }
179  // jest / vitest: the last "Tests:" line
180  const jest = lastOf(/^\s*Tests:?\s/)
181  if (jest) return tally(jest)
182  // cargo: one "test result:" line per test binary
183  const cargo = lines.filter(l => /^test result:/.test(l))
184  if (cargo.length) return tally(cargo.join('\n'))
185  // otherwise the last line that carries a count (pytest's closing "=== 3 failed, 9 passed in 0.4s ===")
186  const last = lastOf(/\d+\s+(?:passed|passing|failed|failing)\b/i)
187  return last ? tally(last) : { passed: null, failed: null }
188}
189
190const pad = (n: number) => String(n).padStart(2, '0')
191/** Wall-clock time in the session's own time zone. */
192const clock = (ms: number) => { const d = new Date(ms); return `${pad(d.getHours())}:${pad(d.getMinutes())}:${pad(d.getSeconds())}` }
193const short = (cmd: string) => (cmd.length > 48 ? cmd.slice(0, 47) + '…' : cmd)
194const base = (p: string | null) => (p ?? '?').split('/').pop() ?? '?'
195type Shell = Extract<ToolEvent, { kind: 'shell' }>
196type Edit = Extract<ToolEvent, { kind: 'edit' }>
197
198const pathMatches = (path: string | null, name: string) =>
199  path !== null && (path === name || path.endsWith('/' + name))
200
201function exitProblem(run: Shell): Receipt['verdict'] | null {
202  return run.exit.kind === 'code' ? null : run.exit.kind === 'unknown' ? 'unknown' : 'contradicted'
203}
204function exitWords(run: Shell): string {
205  switch (run.exit.kind) {
206    case 'interrupted': return 'was interrupted'
207    case 'denied': return `was refused (${run.exit.reason})`
208    case 'unknown': return `ended with no readable exit status (${run.exit.reason})`
209    case 'code': return `exited ${run.exit.code}`
210  }
211}
212
213function testReceipt(claim: Extract<Claim, { kind: 'tests' }>, events: ToolEvent[], turn: number): Receipt {
214  const shells = events.filter((e): e is Shell => e.kind === 'shell')
215  const runs = shells.filter(e => { const k = testRunOf(e.command).kind; return k === 'run' || k === 'ambiguous' })
216  const run = runs[runs.length - 1]
217  if (!run) {
218    const onlyNoRun = shells.some(e => testRunOf(e.command).kind === 'no-run')
219    // nothing recorded is missing evidence, not an observed failure
220    return { claim, verdict: 'unknown', evidence: `no test run was recorded in this session${onlyNoRun ? ' (only help, a dry run or a listing)' : ''}`, at: null }
221  }
222  const what = `\`${short(run.command)}\` at ${clock(run.atMs)}`
223  const shape = testRunOf(run.command)
224  if (shape.kind !== 'run') {
225    return { claim, verdict: 'unknown', evidence: `${what} is a compound command; its exit status may be another program's`, at: run.seq }
226  }
227  const problem = exitProblem(run)
228  if (problem) return { claim, verdict: problem, evidence: `${what} ${exitWords(run)}`, at: run.seq }
229
230  const counts = countsOf(run.output)
231  const tally = [counts.passed !== null && `${counts.passed} passed`, counts.failed !== null && `${counts.failed} failed`].filter(Boolean).join(', ')
232  if (run.exit.kind === 'code' && run.exit.code !== 0) {
233    return { claim, verdict: 'contradicted', evidence: `${what} exited ${run.exit.code}${tally ? ` · ${tally}` : ''}`, at: run.seq }
234  }
235  const partial = (why: string): Receipt => ({ claim, verdict: 'partial', evidence: `${what} exited 0, ${why}`, at: run.seq })
236  if ((counts.failed ?? 0) > 0) return partial(`but printed ${tally}`)
237  if (counts.passed === null) return partial('but printed no result line Receipts can read')
238  if (claim.count !== null && claim.count !== counts.passed) return partial(`printed ${counts.passed} passed; the claim says ${claim.count}`)
239  if (run.turn !== turn) return partial(`${tally}, in an earlier turn`)
240  const later = events.filter((e): e is Edit => e.kind === 'edit' && e.landed && e.seq > run.seq)
241  if (later.length) {
242    const names = [...new Set(later.map(e => base(e.path)))]
243    return partial(`${tally}, then ${later.length} edit${later.length === 1 ? '' : 's'} landed after it (${names.join(', ')})`)
244  }
245  const narrowed = narrowing(shape.args)
246  if (narrowed.length && /\ball\b/i.test(claim.quote)) return partial(`${tally}, but the run was narrowed (${narrowed.join(' ')}) and the claim says all`)
247  return { claim, verdict: 'backed', evidence: `${what} exited 0 · ${tally}`, at: run.seq }
248}
249
250function fileReceipt(claim: Extract<Claim, { kind: 'file' }>, events: ToolEvent[], turn: number): Receipt {
251  const edits = events.filter((e): e is Edit => e.kind === 'edit' && pathMatches(e.path, claim.name))
252  const landed = edits.filter(e => e.landed)
253  const last = edits[edits.length - 1]
254  const paths = [...new Set(landed.map(e => e.path))]
255  if (paths.length > 1) {
256    return { claim, verdict: 'partial', evidence: `edits landed on ${paths.length} different files with this name: ${paths.join(', ')}`, at: landed[landed.length - 1]?.seq ?? null }
257  }
258  if (last && !last.landed && landed.length) {
259    return { claim, verdict: 'partial', evidence: `an earlier edit landed, but the last one at ${clock(last.atMs)} did not: ${last.reason ?? 'unknown reason'}`, at: last.seq }
260  }
261  const lastLanded = landed[landed.length - 1]
262  if (lastLanded) {
263    const where = lastLanded.turn === turn ? '' : ', in an earlier turn'
264    return { claim, verdict: 'backed', evidence: `edit landed at ${clock(lastLanded.atMs)}${where}${landed.length > 1 ? ` (${landed.length} edits)` : ''}`, at: lastLanded.seq }
265  }
266  if (last) return { claim, verdict: 'contradicted', evidence: `the edit did not land: ${last.reason ?? 'unknown reason'}`, at: last.seq }
267  const name = claim.name.split('/').pop() ?? claim.name
268  const shell = [...events].reverse().find((e): e is Shell => e.kind === 'shell' && e.command.includes(name))
269  if (shell) return { claim, verdict: 'unknown', evidence: `no edit tool touched it; only a shell command named it (\`${short(shell.command)}\`)`, at: shell.seq }
270  const read = events.some(e => e.kind === 'read' && pathMatches(e.path, claim.name))
271  return { claim, verdict: 'contradicted', evidence: read ? 'read, but never edited in this session' : 'no edit to this file in this session', at: null }
272}
273
274type GitAction = Extract<Claim, { kind: 'git' }>['action']
275
276const GIT_LABEL: Record<GitAction, string> = { commit: 'git commit', push: 'git push', pr: 'gh pr create', merge: 'merge' }
277
278/** What a successful run prints when the action really happened. */
279const GIT_EFFECT: Record<GitAction, RegExp> = {
280  commit: /^\[[^\]\n]+ [0-9a-f]{7,40}\]/m,
281  push: /^To \S+[\s\S]*->/m,
282  pr: /https?:\/\/\S+\/pull\/\d+/,
283  merge: /Merged pull request|✓ Merged|Fast-forward|Merge made by/,
284}
285
286const GIT_NO_RUN: Record<GitAction, RegExp | null> = { commit: null, push: /^-n$/, pr: null, merge: null }
287
288function gitStepOf(program: string[], action: GitAction): string[] | null {
289  const [name, ...args] = program
290  if (name === 'gh') {
291    if (args[0] !== 'pr') return null
292    return (action === 'pr' && args[1] === 'create') || (action === 'merge' && args[1] === 'merge') ? args.slice(2) : null
293  }
294  if (name !== 'git') return null
295  let i = 0
296  while (i < args.length && /^-/.test(args[i] as string)) i += /^-(?:C|c)$/.test(args[i] as string) ? 2 : 1
297  const sub = args[i]
298  return (action === 'commit' && sub === 'commit') || (action === 'push' && sub === 'push') || (action === 'merge' && sub === 'merge') ? args.slice(i + 1) : null
299}
300
301export function gitRunOf(command: string, action: GitAction): { kind: 'run' } | { kind: 'ambiguous' | 'no-run' | 'none' } {
302  const deciding = decidingStep(command)
303  if (deciding) {
304    const args = gitStepOf(deciding.program, action)
305    if (args) return asksForNoRun(args, GIT_NO_RUN[action]) ? { kind: 'no-run' } : { kind: 'run' }
306  }
307  return { kind: stepsOf(command).some(s => gitStepOf(programOf(s.argv), action) !== null) ? 'ambiguous' : 'none' }
308}
309
310function gitReceipt(claim: Extract<Claim, { kind: 'git' }>, events: ToolEvent[]): Receipt {
311  const shells = events.filter((e): e is Shell => e.kind === 'shell')
312  const runs = shells.filter(e => { const k = gitRunOf(e.command, claim.action).kind; return k === 'run' || k === 'ambiguous' })
313  const run = runs[runs.length - 1]
314  const label = GIT_LABEL[claim.action]
315  if (!run) {
316    const onlyNoRun = shells.some(e => gitRunOf(e.command, claim.action).kind === 'no-run')
317    // nothing recorded is missing evidence, not an observed failure
318    return { claim, verdict: 'unknown', evidence: `no ${label} was recorded in this session${onlyNoRun ? ' (only help or a dry run)' : ''}`, at: null }
319  }
320  const what = `\`${short(run.command)}\` at ${clock(run.atMs)}`
321  if (gitRunOf(run.command, claim.action).kind !== 'run') {
322    return { claim, verdict: 'unknown', evidence: `${what} is a compound command; its exit status may be another program's`, at: run.seq }
323  }
324  const problem = exitProblem(run)
325  if (problem) return { claim, verdict: problem, evidence: `${what} ${exitWords(run)}`, at: run.seq }
326  if (run.exit.kind === 'code' && run.exit.code !== 0) return { claim, verdict: 'contradicted', evidence: `${what} exited ${run.exit.code}`, at: run.seq }
327  return GIT_EFFECT[claim.action].test(run.output)
328    ? { claim, verdict: 'backed', evidence: `${what} exited 0 and printed the ${claim.action === 'pr' ? 'PR link' : claim.action === 'commit' ? 'new commit' : claim.action === 'push' ? 'pushed ref' : 'merge'}`, at: run.seq }
329    : { claim, verdict: 'partial', evidence: `${what} exited 0, but printed no sign of the ${claim.action === 'pr' ? 'PR' : claim.action}`, at: run.seq }
330}
331
332/**
333 * Each claim's receipt over the given events. `scope` names what those events
334 * cover in the evidence text ('this session' by default; Fine Print passes
335 * only the turn's own main-loop events and 'this turn').
336 */
337export function receiptsOf(claims: Claim[], events: ToolEvent[], turn: number, scope = 'this session'): Receipt[] {
338  return claims.map(claim => {
339    const r = claim.kind === 'tests' ? testReceipt(claim, events, turn)
340      : claim.kind === 'file' ? fileReceipt(claim, events, turn)
341        : gitReceipt(claim, events)
342    return scope === 'this session' ? r : { ...r, evidence: r.evidence.replace(/in this session/g, `in ${scope}`) }
343  })
344}
345
346export function tallyOf(receipts: Receipt[]): Record<Verdict, number> {
347  const t: Record<Verdict, number> = { backed: 0, contradicted: 0, partial: 0, unknown: 0 }
348  for (const r of receipts) t[r.verdict] += 1
349  return t
350}
351
hooks/view.tsx 166 lines
1import type { Elements, RenderElement } from 'claude-code'
2
3import { type FinePrint, headlineOf, type Item, questionFor } from './audit'
4
5export type Kit = Pick<Elements['terminal'], 'Box' | 'Text' | 'Button'>
6
7export type View = { mode: 'list' } | { mode: 'review'; index: number } | { mode: 'question'; index: number }
8
9export type Actions = {
10  ask: (item: number) => void
11  review: (change: number) => void
12  step: (delta: number) => void
13  back: () => void
14  close: () => void
15}
16
17const INK = { bad: 'red', warn: 'yellow', add: 'green', del: 'red' } as const
18const HUNK_PREVIEW = 6
19
20const pad = (n: number) => String(n).padStart(2, '0')
21/** A path as the workspace sees it: relative to the session's folder when inside it. */
22export const relativeTo = (cwd: string, path: string) => (cwd && path.startsWith(cwd + '/') ? path.slice(cwd.length + 1) : path)
23
24const clock = (ms: number) => { const d = new Date(ms); return `${pad(d.getHours())}:${pad(d.getMinutes())}:${pad(d.getSeconds())}` }
25
26/** The changes a review steps through: unmentioned edits, in card order. */
27export const changesOf = (fp: FinePrint) => fp.items.filter((i): i is Extract<Item, { kind: 'change' }> => i.kind === 'change')
28
29function diffLines(kit: Kit, lines: string[], keyPrefix: string): RenderElement[] {
30  const { Text } = kit
31  return lines.map((l, i) => (
32    <Text key={`${keyPrefix}${i}`} color={l.startsWith('+') ? INK.add : l.startsWith('-') ? INK.del : undefined} dimColor={!l.startsWith('+') && !l.startsWith('-')}>{l === '' ? ' ' : l}</Text>
33  ))
34}
35
36function itemCard(kit: Kit, fp: FinePrint, item: Item, i: number, act: Actions, cwd: string): RenderElement {
37  const { Box, Text, Button } = kit
38  const ask = <Button key={`ask-${i}`} label="Ask Claude" hotkey={i < 9 ? String(i + 1) : undefined} variant={i === 0 ? 'primary' : undefined} onPress={() => act.ask(i)} />
39
40  let mark: RenderElement
41  let title: RenderElement
42  let body: RenderElement[]
43  let coverage: string
44
45  if (item.kind === 'claim') {
46    const v = item.receipt.verdict
47    mark = v === 'contradicted' ? <Text color={INK.bad} bold>✗</Text> : v === 'partial' ? <Text color={INK.warn} bold>◐</Text> : <Text dimColor bold>?</Text>
48    title = <Text bold italic>{`“${item.receipt.claim.quote}”`}</Text>
49    body = [<Text key="result">{item.receipt.evidence}</Text>]
50    coverage = v === 'contradicted' ? 'the record contradicts this claim' : v === 'partial' ? 'the record only partly backs this claim' : 'nothing recorded this turn backs this claim'
51  } else if (item.kind === 'change') {
52    const lines = item.edits.flatMap((e, k) => (k ? ['…', ...e.patch] : e.patch))
53    const first = item.edits[0]
54    const created = item.edits.some(e => e.created)
55    mark = <Text color={item.why ? INK.warn : undefined} bold>{item.why ? '⚠' : '•'}</Text>
56    title = (
57      <Box flexDirection="row" gap={1} flexWrap="wrap">
58        <Text bold>{relativeTo(cwd, item.path)}</Text>
59        {item.why ? <Text color={INK.warn}>{item.why}</Text> : null}
60      </Box>
61    )
62    body = [
63      ...(relativeTo(cwd, item.path) !== item.path ? [<Text key="path" dimColor>{item.path}</Text>] : []),
64      <Text key="action" dimColor>{`${created ? 'Write created it' : `${item.edits.length} edit${item.edits.length === 1 ? '' : 's'} landed`}${first ? ` at ${clock(first.atMs)}` : ''}`}</Text>,
65      <Box key="hunk" flexDirection="column" paddingLeft={1}>{diffLines(kit, lines.slice(0, HUNK_PREVIEW), `h${i}-`)}</Box>,
66      ...(lines.length > HUNK_PREVIEW ? [<Text key="more" dimColor>{`… ${lines.length - HUNK_PREVIEW} more line${lines.length - HUNK_PREVIEW === 1 ? '' : 's'}`}</Text>] : []),
67    ]
68    coverage = item.shared ? `the summary's \`${item.shared}\` matches more than one changed file` : 'the summary never names this file'
69  } else {
70    mark = <Text color={INK.warn} bold>⚠</Text>
71    title = <Text bold>{item.effect}</Text>
72    body = [
73      <Text key="action">{`$ ${item.command}`}</Text>,
74      <Text key="result" color={item.ok === false ? INK.bad : undefined} dimColor={item.ok !== false}>{item.result}</Text>,
75    ]
76    coverage = 'the summary never mentions it'
77  }
78
79  const changeIndex = item.kind === 'change' ? changesOf(fp).indexOf(item) : -1
80  return (
81    <Box key={`item-${i}`} flexDirection="column" paddingLeft={1} borderStyle="single" borderColor={item.kind === 'claim' ? (item.receipt.verdict === 'contradicted' ? INK.bad : item.receipt.verdict === 'partial' ? INK.warn : 'gray') : item.kind === 'command' || (item.kind === 'change' && item.why) ? INK.warn : 'gray'}>
82      <Box key="title" flexDirection="row" gap={1}>
83        <Box width={2} flexShrink={0}>{mark}</Box>
84        <Box flexShrink={1}>{title}</Box>
85      </Box>
86      <Box key="body" flexDirection="column" paddingLeft={3}>{body}</Box>
87      <Box key="foot" flexDirection="row" gap={2} paddingLeft={3} flexWrap="wrap">
88        <Text italic dimColor>{`Summary: ${coverage}`}</Text>
89        {ask}
90        {changeIndex >= 0 ? <Button key={`diff-${i}`} label="Full diff" onPress={() => act.review(changeIndex)} /> : null}
91      </Box>
92    </Box>
93  )
94}
95
96function listView(kit: Kit, fp: FinePrint, act: Actions, cwd: string): RenderElement {
97  const { Box, Text, Button } = kit
98  return (
99    <Box flexDirection="column">
100      <Box key="head" flexDirection="column" marginBottom={1}>
101        <Text bold color={fp.items.some(i => i.kind === 'claim' && i.receipt.verdict === 'contradicted') ? INK.bad : INK.warn}>{`FINE PRINT · turn ${fp.turn}`}</Text>
102        <Text>{headlineOf(fp.items)}</Text>
103      </Box>
104      <Box key="items" flexDirection="column" gap={1}>{fp.items.map((item, i) => itemCard(kit, fp, item, i, act, cwd))}</Box>
105      <Box key="actions" flexDirection="row" gap={2} marginTop={1}>
106        <Button key="close" label="Close" role="dismiss" onPress={act.close} />
107        <Text dimColor>Ask Claude puts one question in your prompt box. You edit it and press Enter.</Text>
108      </Box>
109    </Box>
110  )
111}
112
113function reviewView(kit: Kit, fp: FinePrint, index: number, act: Actions, cwd: string): RenderElement {
114  const { Box, Text, Button } = kit
115  const changes = changesOf(fp)
116  const at = Math.max(0, Math.min(index, changes.length - 1))
117  const item = changes[at]
118  if (!item) return listView(kit, fp, act, cwd)
119  const lines = item.edits.flatMap((e, k) => (k ? ['…', ...e.patch] : e.patch))
120  return (
121    <Box flexDirection="column">
122      <Box key="head" flexDirection="row" justifyContent="space-between" marginBottom={1}>
123        <Text bold color={item.why ? INK.warn : undefined}>{`${relativeTo(cwd, item.path)}${item.why ? `  ·  ${item.why}` : ''}`}</Text>
124        <Text dimColor>{`change ${at + 1} of ${changes.length}`}</Text>
125      </Box>
126      <Box key="diff" flexDirection="column">
127        {lines.length ? diffLines(kit, lines, 'd') : <Text dimColor>The tool result recorded no diff for this change.</Text>}
128      </Box>
129      <Box key="actions" flexDirection="row" gap={2} marginTop={1} flexWrap="wrap">
130        <Button key="prev" label="Prev" hotkey="p" onPress={() => act.step(-1)} />
131        <Button key="next" label="Next" hotkey="n" onPress={() => act.step(1)} />
132        <Button key="ask" label="Ask Claude" hotkey="a" variant="primary" onPress={() => act.ask(fp.items.indexOf(item))} />
133        <Button key="back" label="Back" hotkey="b" onPress={act.back} />
134        <Button key="close" label="Close" role="dismiss" onPress={act.close} />
135      </Box>
136    </Box>
137  )
138}
139
140function questionView(kit: Kit, fp: FinePrint, index: number, act: Actions, cwd: string): RenderElement {
141  const { Box, Text, Button } = kit
142  const item = fp.items[index]
143  if (!item) return listView(kit, fp, act, cwd)
144  return (
145    <Box flexDirection="column">
146      <Text key="head" bold>The question for this item</Text>
147      <Text key="why" dimColor>The prompt box could not be filled from here. Copy this into it, edit it, and press Enter.</Text>
148      <Box key="question" flexDirection="column" marginTop={1} paddingLeft={1} borderStyle="single" borderColor="gray"><Text>{questionFor(item)}</Text></Box>
149      <Box key="actions" flexDirection="row" gap={2} marginTop={1}>
150        <Button key="back" label="Back" hotkey="b" onPress={act.back} />
151        <Button key="close" label="Close" role="dismiss" onPress={act.close} />
152      </Box>
153    </Box>
154  )
155}
156
157export function cardView(kit: Kit, fp: FinePrint | null, view: View, act: Actions, cwd = ''): RenderElement {
158  const { Box, Text } = kit
159  if (!fp || !fp.items.length) {
160    return <Box flexDirection="column"><Text dimColor>No fine print in the last answer: it names every edit Fine Print recorded, mentions every recorded side-effect command, and no checked claim is contradicted. Fine Print reads Edit, Write, Read and Bash results from the main conversation only.</Text></Box>
161  }
162  if (view.mode === 'review') return reviewView(kit, fp, view.index, act, cwd)
163  if (view.mode === 'question') return questionView(kit, fp, view.index, act, cwd)
164  return listView(kit, fp, act, cwd)
165}
166
hooks/claims.ts 119 lines
1/**
2 * Finds the few claims in Claude's final answer that the session's own tool
3 * calls can settle: tests passing, a file changed, a commit / push / PR.
4 * Pure. Deliberately narrow: a sentence that is negated, conditional or
5 * about the future makes no claim, and anything else is left alone.
6 */
7
8export type GitAction = 'commit' | 'push' | 'pr' | 'merge'
9
10export type Claim =
11  | { kind: 'tests'; quote: string; count: number | null }
12  | { kind: 'file'; quote: string; name: string }
13  | { kind: 'git'; quote: string; action: GitAction }
14
15const NEGATED = /\b(?:not|never|unable|skipped|cannot)\b|n't\b/i
16const HEDGED = /\b(?:should|will|would|might|may|could|if|once|expect(?:ed|s)?|try|trying|need(?:s|ed)? to|want(?:s|ed)? to)\b/i
17
18const TESTS = [
19  /\b(?:all\s+)?(?:(\d+)\s+)?(?:(?:unit|integration|e2e|end-to-end|regression)\s+)?tests?\s+(?:now\s+|still\s+)?(?:pass(?:es|ed)?|are\s+(?:passing|green)|succeed(?:s|ed)?|went\s+green|are\s+all\s+passing)\b/i,
20  /\btest\s+suite\s+(?:now\s+)?(?:passes|passed|is\s+(?:green|passing))\b/i,
21  /\b(\d+)\s*(?:\/\s*\d+\s*)?(?:tests?\s+)?passed\b/i,
22]
23
24const CHANGE = /\b(?:updated|changed|edited|modified|added|created|fixed|wrote|rewrote|renamed|removed|deleted|refactored|moved|patched|implemented|replaced)\b/i
25
26const GIT: [RegExp, GitAction][] = [
27  [/\bcommitted\b/i, 'commit'],
28  [/\bpushed\b/i, 'push'],
29  [/\b(?:opened|created|raised)\s+(?:a\s+|the\s+)?(?:pr|pull\s+request)\b/i, 'pr'],
30  [/\bmerged\s+(?:a\s+|the\s+|it\b|this\b)?(?:pr|pull\s+request|branch)?/i, 'merge'],
31]
32
33/** File-like tokens: a dotted name with a lettered extension, optionally with a path. */
34const FILE = /(?:^|[\s(`'"])((?:[\w.@-]+\/)*[\w@-][\w.@-]*\.[a-z][a-z0-9]{0,6})(?=[\s`'"),.:;!?]|$)/gi
35const NOT_FILES = /^(?:e\.g|i\.e|etc|vs|v\d|node\.js|next\.js|vue\.js|react\.js)$|\.(?:com|org|net|io|dev|ai|app|co)$/i
36
37/** Words that mark a double-quoted span as someone else's words (reported speech), not the answer's own claim. */
38const FRAMING = /\b(?:said|says|say|saying|wrote|writes|reads|read|claim(?:s|ed)?|answer(?:ed)?|summary|message|reply|replied|line|text|caption|callout|title|tagline|headline|dialogue|quot(?:e|es|ed|ing)|example|e\.g\.|such as|like|called|labell?ed|shows?|displays?|prints?|outputs?)\b[^"“.!?\n]{0,40}\s*$/i
39
40/**
41 * Blanks the inside of double-quoted spans ("…" or “…”) that the text frames as
42 * reported speech: a framing word in the same sentence before the quote (`said`,
43 * `reads`, `the answer`, `for example`, …). A colon alone is not framing:
44 * `Tests: "All 12 tests pass."` is the answer's own claim. The answer quoting a
45 * line is describing it, not claiming it. An unframed quote is left as written,
46 * so a claim quoted for emphasis is still checked. Lengths are kept, so nothing
47 * else shifts.
48 */
49export function maskReportedSpeech(text: string): string {
50  return text.replace(/(["“])([^"“”\n]{1,600}?)(["”])/g, (whole, open: string, inner: string, close: string, at: number) => {
51    const before = text.slice(0, at)
52    const sentenceStart = Math.max(before.lastIndexOf('\n'), before.search(/[.!?]\s+[^.!?]*$/))
53    const lead = before.slice(sentenceStart + 1)
54    return FRAMING.test(lead) ? open + ' '.repeat(inner.length) + close : whole
55  })
56}
57
58export function sentencesOf(text: string): string[] {
59  return text
60    .replace(/```[\s\S]*?```/g, ' ')
61    .split(/\n+|(?<=[.!?])\s+(?=[A-Z`*(\-])/)
62    .map(s => s.replace(/^\s*(?:[-*•]|\d+\.)\s+/, '').trim())
63    .filter(s => s !== '')
64}
65
66function makesClaim(sentence: string): boolean {
67  return !NEGATED.test(sentence) && !HEDGED.test(sentence)
68}
69
70export function filesIn(sentence: string): string[] {
71  const names = new Set<string>()
72  for (const m of sentence.matchAll(FILE)) {
73    const name = (m[1] ?? '').replace(/^\.\//, '')
74    if (name && !NOT_FILES.test(name) && !/^\d+(?:\.\d+)+$/.test(name)) names.add(name)
75  }
76  return [...names]
77}
78
79export function claimsOf(answer: string): Claim[] {
80  const claims: Claim[] = []
81  const seenFiles = new Set<string>()
82  const seenGit = new Set<GitAction>()
83  let sawTests = false
84
85  for (const sentence of sentencesOf(maskReportedSpeech(answer))) {
86    if (!makesClaim(sentence)) continue
87    const plain = sentence.replace(/(["“])\s+(["”])/g, '$1…$2').replace(/\s{2,}/g, ' ')
88    const quote = plain.length > 160 ? plain.slice(0, 157) + '…' : plain
89
90    if (!sawTests) {
91      for (const re of TESTS) {
92        const m = re.exec(sentence)
93        if (m) {
94          const n = m[1] !== undefined ? Number(m[1]) : null
95          claims.push({ kind: 'tests', quote, count: n })
96          sawTests = true
97          break
98        }
99      }
100    }
101
102    if (CHANGE.test(sentence)) {
103      for (const name of filesIn(sentence)) {
104        if (seenFiles.has(name)) continue
105        seenFiles.add(name)
106        claims.push({ kind: 'file', quote, name })
107      }
108    }
109
110    for (const [re, action] of GIT) {
111      if (!seenGit.has(action) && re.test(sentence)) {
112        seenGit.add(action)
113        claims.push({ kind: 'git', quote, action })
114      }
115    }
116  }
117  return claims
118}
119
hooks/command.ts 98 lines
1/**
2 * Reads a Bash command string just far enough to say which program a step
3 * runs. Not a shell: quotes, `&&`, `||`, `;`, `|` and newlines are split; a
4 * leading `VAR=x` and a few runners (`uv run`, `npx`, `python -m`) are
5 * stepped through; anything else is taken literally. Pure.
6 */
7
8export type Joiner = '&&' | '||' | ';' | '|' | 'start'
9
10export type Step = {
11  /** How this step is joined to the one before it. */
12  joiner: Joiner
13  /** The program and its arguments, quotes removed. */
14  argv: string[]
15}
16
17/** Splits a command into steps, keeping quoted text together. */
18export function stepsOf(command: string): Step[] {
19  const steps: Step[] = []
20  let argv: string[] = []
21  let word = ''
22  let hasWord = false
23  let joiner: Joiner = 'start'
24  let quote: '"' | "'" | null = null
25
26  const endWord = () => {
27    if (hasWord) argv.push(word)
28    word = ''
29    hasWord = false
30  }
31  const endStep = (next: Joiner) => {
32    endWord()
33    if (argv.length) steps.push({ joiner, argv })
34    argv = []
35    joiner = next
36  }
37
38  for (let i = 0; i < command.length; i++) {
39    const c = command[i] as string
40    if (quote) {
41      if (c === quote) quote = null
42      else if (c === '\\' && quote === '"' && i + 1 < command.length) word += command[++i]
43      else word += c
44      continue
45    }
46    if (c === '"' || c === "'") { quote = c; hasWord = true; continue }
47    if (c === '\\' && i + 1 < command.length) { word += command[++i]; hasWord = true; continue }
48    const two = command.slice(i, i + 2)
49    if (two === '&&' || two === '||') { endStep(two); i++; continue }
50    if (c === ';' || c === '\n') { endStep(';'); continue }
51    if (c === '|') { endStep('|'); continue }
52    if (c === ' ' || c === '\t') { endWord(); continue }
53    word += c
54    hasWord = true
55  }
56  endStep(';')
57  return steps
58}
59
60const RUNNERS: [string[], number][] = [
61  [['uv', 'run'], 2], [['poetry', 'run'], 2], [['pipenv', 'run'], 2], [['pdm', 'run'], 2],
62  [['npx'], 1], [['pnpx'], 1], [['bunx'], 1], [['pnpm', 'exec'], 2], [['yarn', 'exec'], 2],
63]
64
65/** The program a step really runs, with its arguments, after env vars and runners. */
66export function programOf(argv: string[]): string[] {
67  let rest = argv
68  while (rest.length && /^[A-Za-z_][A-Za-z0-9_]*=/.test(rest[0] as string)) rest = rest.slice(1)
69  for (;;) {
70    const runner = RUNNERS.find(([words]) => words.every((w, i) => rest[i] === w))
71    if (!runner) break
72    rest = rest.slice(runner[1])
73  }
74  const head = (rest[0] ?? '').split('/').pop() ?? ''
75  if (/^python(?:3(?:\.\d+)?)?$/.test(head) && rest[1] === '-m' && rest[2]) return [rest[2], ...rest.slice(3)]
76  return [head, ...rest.slice(1)]
77}
78
79/**
80 * The step whose exit status is the command's: the last one, and only when
81 * every join before it is `&&` (each earlier step had to succeed for it to run).
82 * After a `|`, `||` or `;` the status may be another program's, so null.
83 */
84export function decidingStep(command: string): { program: string[]; isSole: boolean } | null {
85  const steps = stepsOf(command)
86  const last = steps[steps.length - 1]
87  if (!last) return null
88  if (steps.slice(1).some(s => s.joiner !== '&&')) return null
89  return { program: programOf(last.argv), isSole: steps.length === 1 }
90}
91
92/** True when any argument asks for help, a version, a dry run or a listing instead of a real run. */
93export const NO_RUN_FLAGS = /^(?:--help|-h|--version|-V|--dry-run|--dryrun|--collect-only|--co|--list|--list-tests|--listTests|--showConfig|--no-run)$/
94
95export function asksForNoRun(args: string[], extra: RegExp | null = null): boolean {
96  return args.some(a => NO_RUN_FLAGS.test(a) || (extra !== null && extra.test(a)))
97}
98