SLOPSHOPPER

tooling-retro

Observes sessions for moments where skills, agents or repo tooling let the agent down, pushes structured observations to chore/tooling-retro and runs…

newspinnerguardcommandtoaststatus
v0.3.0no licenseupdated 2026-10-08ItsReddi/pp-claude-tooling/plugins/tooling-retro
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · tooling-retro
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /tooling-retro ⎿ tooling-retro: Tooling retro over chore/tooling-retro starts once this session's observations are pushed. ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts
README

tooling-retro

Watches Claude Code sessions for moments where a repository's agent tooling let the agent down, and turns them into retrospectives on that tooling: CLAUDE.md files, skills, agents, rules, hooks and workflows.

Install: /plugin install tooling-retro@pp-claude-tooling (see the marketplace README for adding the marketplace).

How it works

Collect, locally. While a session runs, the mod writes one JSONL line per event to .git/tooling-retro/<session-id>.jsonl: skill loads, subagent starts, every tool call (what it touched, whether it failed, whether it read tooling), the person's prompts (clipped, redacted) and how each turn ended. Each observer run adds a line with its triggers, its outcome and the observer's reply (clipped to 4000 characters, redacted), so a NONE can be traced back too. This raw log never leaves the machine.

To see whether and why the observer ran in a session:

grep -E '"type":"(turn_end|observer)"' .git/tooling-retro/<session-id>.jsonl

Observe at triggers. At the end of a turn the mod checks four triggers:

TriggerDetected by
abortturn.complete reports the turn as interrupted
correctionthe person's next prompt after a turn that edited files or ran commands, classified by the engine's small model as a correction
tool-errorstwo or more failed tool calls in the turn
subagentthe turn started a subagent

When one fires, the mod asks a fork of the session ($.model.fork) whether the moment points at the repository's tooling. The fork sees the whole conversation from the prompt cache, runs on the session's own model, has no tools and leaves the conversation untouched. It answers NONE or up to three observations in a fixed template (prompts/observer.md):

{"category":"guidance-not-loaded","tooling":".claude/skills/pest-testing/SKILL.md","summary":"…","evidence":"…","suggestion":"…","confidence":"high"}

Categories: guidance-not-loaded, guidance-ignored, guidance-wrong, guidance-conflict, guidance-missing, subagent-context, tool-friction, other. Lines off the template are dropped, fields are redacted and clipped. At most 20 observer runs per session. A toast says when observations were noted, and the spinner shows the session's count.

Push. Observations go to observations/<day>/<session-id>.jsonl on the branch chore/tooling-retro, together with stats/<day>/<session-id>.json: counts of turns, aborts, corrections, tool calls and errors, tooling reads, and how often each skill and subagent type was used. No prompt or file content. A push happens whenever new observations exist, stats alone at most every 15 minutes and at session end, because cloud sessions lose anything that stays local. The branch name is fixed, has a history of its own and never merges into the code branches. Commits go through a temporary index and leave the working tree alone. A repository that opens pull requests for new branches automatically will open one for this branch too; keeping it out is that repository's call. A 403 or 407 on push stops pushing for the rest of the session.

Retro. /tooling-retro [focus] reads the observations and stats no earlier retro covered, takes an inventory of the repository's tooling, groups observations by tooling and category, and turns patterns into proposals. A proposal needs observations from at least two sessions, or one high-confidence observation backed by the stats. Each proposal comes with one or two variants that differ in substance. The report goes to retros/<day>.md on the same branch.

Decisions. The retro asks about each proposal on its own: a variant, Reject or Later. Every answer goes to decisions.jsonl on the branch before anything is implemented. Later retros read that file first: a rejected proposal does not come back, a later one does, and an accepted one returns only as a follow-up when the problem persists after the change. Accepted proposals are implemented on a normal working branch.

Nothing in the mod is specific to one repository: it reads whatever tooling the repository has.

Options

Set through /config or pluginConfigs in settings:

OptionDefaultEffect
pushtruePush observations and stats. Off: they stay in .git/tooling-retro/.

Reading the observations by hand

git fetch origin +refs/heads/chore/tooling-retro:refs/tooling-retro/remote
git ls-tree -r --name-only refs/tooling-retro/remote
git show refs/tooling-retro/remote:observations/<day>/<session-id>.jsonl

Develop

claude plugin validate plugins/tooling-retro
claude plugin test plugins/tooling-retro
claude --plugin-dir plugins/tooling-retro
Source 3 files
hooks/register.ts 486 lines
1import type { EngineInterface, Register } from 'claude-code'
2
3import { clip, describeCall, isGuidancePath, type LogEntry, toLine, withoutCredentials } from './log'
4import { parseObservations, type Trigger, triggersFor } from './observe'
5
6type Repo = { root: string; gitDir: string }
7type Entry = Omit<LogEntry, 'at'>
8
9type Stats = {
10  session: string
11  day: string
12  model: string | null
13  remote: string | null
14  turns: number
15  aborts: number
16  corrections: number
17  toolCalls: number
18  toolErrors: number
19  guidanceReads: number
20  skills: Record<string, number>
21  subagents: Record<string, number>
22  observerRuns: number
23  observations: number
24}
25
26type SessionState = {
27  id: string
28  raw: string
29  observations: string
30  stats: Stats
31  isStatsDirty: boolean
32  isObservationsDirty: boolean
33  lastPushAt: number
34}
35
36type TurnState = { hasActed: boolean; toolErrors: number; spawnedAgents: number }
37
38const COMMAND = 'tooling-retro'
39const LOG_BRANCH = 'chore/tooling-retro'
40const HUMAN_ORIGINS = new Set(['composer', 'bridge', 'sdk'])
41const ACTING_TOOLS = new Set(['Edit', 'Write', 'MultiEdit', 'NotebookEdit', 'Bash'])
42const CORRECTION_LABELS = ['correction', 'follow-up', 'new-task'] as const
43const MAX_OBSERVER_RUNS = 20
44const MAX_LOGGED_REPLY_CHARS = 4000
45const STATS_PUSH_INTERVAL_MS = 15 * 60_000
46const MAX_RAW_CHARS = 3_500_000
47const PUSH_ATTEMPTS = 3
48
49let isPushEnabled = true
50let repo: Repo | undefined | null = null
51let state: SessionState | undefined
52let turn: TurnState = { hasActed: false, toolErrors: 0, spawnedAgents: 0 }
53let hasPreviousTurnActed = false
54let pendingCorrection: Promise<boolean> = Promise.resolve(false)
55let queue: Promise<unknown> = Promise.resolve()
56
57function serial<T>(work: () => Promise<T>): Promise<T> {
58  const next = queue.then(work, work)
59  queue = next.catch(() => undefined)
60
61  return next
62}
63
64async function git($: EngineInterface, root: string, args: readonly string[], env?: Record<string, string>) {
65  return $.process.run(['git', ...args], { cwd: root, env, timeoutMs: 60_000 })
66}
67
68async function gitOut($: EngineInterface, root: string, args: readonly string[], env?: Record<string, string>) {
69  const ran = await git($, root, args, env)
70
71  if (ran.exitCode !== 0) {
72    throw new Error(`git ${args[0]} failed: ${ran.stderr.trim().slice(0, 300)}`)
73  }
74
75  return ran.stdout.trim()
76}
77
78async function resolveRepo($: EngineInterface): Promise<Repo | undefined> {
79  if (repo !== null) {
80    return repo
81  }
82
83  const found = await $.session.repo()
84
85  if (found === null || found.remote === null) {
86    repo = undefined
87
88    return repo
89  }
90
91  const ran = await git($, found.root, ['rev-parse', '--path-format=absolute', '--git-common-dir'])
92  repo = ran.exitCode === 0 ? { root: found.root, gitDir: ran.stdout.trim() } : undefined
93
94  return repo
95}
96
97function localPath(found: Repo, sessionId: string, suffix: string): string {
98  return `${found.gitDir}/tooling-retro/${sessionId}${suffix}`
99}
100
101async function readOrEmpty($: EngineInterface, path: string): Promise<string> {
102  return (await $.fs.exists(path)) ? $.fs.read(path) : ''
103}
104
105async function nowIso($: EngineInterface): Promise<string> {
106  return new Date(await $.clock.now()).toISOString()
107}
108
109/**
110 * The state of the session the engine runs now. A /clear starts a new session id, so the state
111 * is rebuilt from the local files whenever the id changes, and after a hot reload.
112 */
113async function currentSession($: EngineInterface, found: Repo, sessionId?: string): Promise<SessionState> {
114  const id = sessionId ?? (await $.session.id())
115
116  if (state !== undefined && state.id === id) {
117    return state
118  }
119
120  const statsText = await readOrEmpty($, localPath(found, id, '.stats.json'))
121  const remoteUrl = (await $.session.repo())?.remote
122  const remote = remoteUrl === undefined || remoteUrl === null ? null : withoutCredentials(remoteUrl)
123  const stats: Stats =
124    statsText === ''
125      ? {
126          session: id,
127          day: (await nowIso($)).slice(0, 10),
128          model: await $.session.model(),
129          remote,
130          turns: 0,
131          aborts: 0,
132          corrections: 0,
133          toolCalls: 0,
134          toolErrors: 0,
135          guidanceReads: 0,
136          skills: {},
137          subagents: {},
138          observerRuns: 0,
139          observations: 0,
140        }
141      : (JSON.parse(statsText) as Stats)
142
143  state = {
144    id,
145    raw: await readOrEmpty($, localPath(found, id, '.jsonl')),
146    observations: await readOrEmpty($, localPath(found, id, '.observations.jsonl')),
147    stats,
148    isStatsDirty: false,
149    isObservationsDirty: false,
150    lastPushAt: 0,
151  }
152
153  return state
154}
155
156function record($: EngineInterface, entry: Entry, update?: (stats: Stats) => void, sessionId?: string): Promise<void> {
157  return serial(() => writeRaw($, entry, update, sessionId))
158}
159
160async function writeRaw($: EngineInterface, entry: Entry, update?: (stats: Stats) => void, sessionId?: string) {
161  const found = await resolveRepo($)
162
163  if (found === undefined) {
164    return
165  }
166
167  const session = await currentSession($, found, sessionId)
168
169  if (update !== undefined) {
170    update(session.stats)
171    session.isStatsDirty = true
172  }
173
174  if (session.raw.length > MAX_RAW_CHARS) {
175    return
176  }
177
178  session.raw += toLine({ ...entry, at: await nowIso($) } as LogEntry)
179  await $.fs.write(localPath(found, session.id, '.jsonl'), session.raw)
180}
181
182async function persistStats($: EngineInterface, found: Repo, session: SessionState): Promise<void> {
183  await $.fs.write(localPath(found, session.id, '.stats.json'), `${JSON.stringify(session.stats, null, 2)}\n`)
184}
185
186function flush($: EngineInterface, isForced: boolean): Promise<void> {
187  return serial(() => pushIfDue($, isForced))
188}
189
190async function pushIfDue($: EngineInterface, isForced: boolean): Promise<void> {
191  const found = await resolveRepo($)
192
193  if (found === undefined || state === undefined) {
194    return
195  }
196
197  const session = state
198  await persistStats($, found, session)
199  const now = await $.clock.now()
200  const isStatsDue = session.isStatsDirty && (isForced || now - session.lastPushAt >= STATS_PUSH_INTERVAL_MS)
201
202  if (!isPushEnabled || session.stats.turns === 0 || (!session.isObservationsDirty && !isStatsDue)) {
203    return
204  }
205
206  try {
207    await pushSession($, found, session)
208    session.isStatsDirty = false
209    session.isObservationsDirty = false
210    session.lastPushAt = now
211    $.ui.status(undefined)
212  } catch (error) {
213    const message = error instanceof Error ? error.message : 'push failed'
214
215    if (/\b40[37]\b/.test(message)) {
216      isPushEnabled = false
217      $.ui.status(`tooling-retro: push to ${LOG_BRANCH} refused by policy, observations stay local`)
218
219      return
220    }
221
222    $.ui.status(`tooling-retro: ${message.slice(0, 120)}`)
223  }
224}
225
226/**
227 * Commits the session's observations and stats onto the log branch without touching the working
228 * tree or the real index: a temporary index, `commit-tree`, and a push of that commit. The raw
229 * event log never leaves the machine. Each session owns its files, so a rejected push only needs
230 * a refetch, never a merge. The commit uses the person's own git identity.
231 */
232async function pushSession($: EngineInterface, found: Repo, session: SessionState): Promise<void> {
233  const files = [{ source: localPath(found, session.id, '.stats.json'), target: `stats/${session.stats.day}/${session.id}.json` }]
234
235  if (session.observations !== '') {
236    files.push({
237      source: localPath(found, session.id, '.observations.jsonl'),
238      target: `observations/${session.stats.day}/${session.id}.jsonl`,
239    })
240  }
241
242  const index = { GIT_INDEX_FILE: `${found.gitDir}/tooling-retro/index.tmp` }
243  const root = found.root
244  let lastError = ''
245
246  for (let attempt = 0; attempt < PUSH_ATTEMPTS; attempt++) {
247    const fetched = await git($, root, ['fetch', '--quiet', 'origin', `+refs/heads/${LOG_BRANCH}:refs/tooling-retro/remote`])
248    const parent = fetched.exitCode === 0 ? await gitOut($, root, ['rev-parse', 'refs/tooling-retro/remote']) : undefined
249
250    await gitOut($, root, parent === undefined ? ['read-tree', '--empty'] : ['read-tree', parent], index)
251
252    for (const file of files) {
253      const blob = await gitOut($, root, ['hash-object', '-w', file.source])
254      await gitOut($, root, ['update-index', '--add', '--cacheinfo', `100644,${blob},${file.target}`], index)
255    }
256
257    const tree = await gitOut($, root, ['write-tree'], index)
258
259    if (parent !== undefined && tree === (await gitOut($, root, ['rev-parse', `${parent}^{tree}`]))) {
260      return
261    }
262
263    const commit = await gitOut($, root, [
264      'commit-tree',
265      tree,
266      ...(parent === undefined ? [] : ['-p', parent]),
267      '-m',
268      `tooling-retro: ${session.stats.observations} observations, session ${session.id}`,
269    ])
270    const pushed = await git($, root, ['push', '--quiet', 'origin', `${commit}:refs/heads/${LOG_BRANCH}`])
271
272    if (pushed.exitCode === 0) {
273      return
274    }
275
276    lastError = pushed.stderr.trim().slice(0, 300)
277  }
278
279  throw new Error(`push to ${LOG_BRANCH} failed: ${lastError}`)
280}
281
282async function classifyCorrection($: EngineInterface, text: string): Promise<boolean> {
283  try {
284    return (await $.model.classify(clip(text, 2000), CORRECTION_LABELS)) === 'correction'
285  } catch {
286    return false
287  }
288}
289
290/**
291 * Asks a fork of this very session, which sees the whole conversation from the prompt cache,
292 * whether the triggers point at the repository's tooling. Only the parsed observations are kept.
293 */
294async function observe($: EngineInterface, triggers: readonly Trigger[]): Promise<void> {
295  const found = await resolveRepo($)
296
297  if (found === undefined) {
298    return
299  }
300
301  try {
302    await runObserver($, found, triggers)
303  } finally {
304    await flush($, false)
305  }
306}
307
308async function runObserver($: EngineInterface, found: Repo, triggers: readonly Trigger[]): Promise<void> {
309  const template = await $.fs.read(`${$.plugin.root}/prompts/observer.md`)
310  const reply = await $.model.fork({ prompt: template.replaceAll('{{triggers}}', triggers.join(', ')) })
311
312  if (!reply.isAnswered) {
313    await record($, { type: 'observer', triggers, outcome: reply.reason })
314
315    return
316  }
317
318  const observations = parseObservations(reply.text)
319  await record($, {
320    type: 'observer',
321    triggers,
322    outcome: `${observations.length} observations`,
323    reply: clip(reply.text, MAX_LOGGED_REPLY_CHARS),
324  })
325
326  if (observations.length === 0) {
327    return
328  }
329
330  const at = await nowIso($)
331  const model = await $.session.model()
332
333  await serial(async () => {
334    const session = await currentSession($, found)
335    session.observations += observations.map(observation => toLine({ ...observation, triggers, model, at } as unknown as LogEntry)).join('')
336    session.stats.observations += observations.length
337    session.isObservationsDirty = true
338    session.isStatsDirty = true
339    await $.fs.write(localPath(found, session.id, '.observations.jsonl'), session.observations)
340  })
341
342  $.ui.toast(`tooling-retro: ${observations.length} ${observations.length === 1 ? 'observation' : 'observations'} noted`)
343  $.ui.invalidate('ui.render')
344}
345
346async function startRetro($: EngineInterface, focus: string): Promise<void> {
347  const template = await $.fs.read(`${$.plugin.root}/prompts/tooling-retro.md`)
348  const text = template
349    .replaceAll('{{branch}}', LOG_BRANCH)
350    .replaceAll('{{focus}}', focus === '' ? 'none, cover everything' : focus)
351
352  await flush($, true)
353  await $.prompt.submit({ text })
354}
355
356function increment(counts: Record<string, number>, key: string): void {
357  counts[key] = (counts[key] ?? 0) + 1
358}
359
360export const register: Register = (on, options) => {
361  isPushEnabled = options.push !== false
362
363  on('session.start', async ($, e, next) => {
364    await $.command.register({
365      name: COMMAND,
366      description: `Retro over the observations on ${LOG_BRANCH}: how skills, agents and repo tooling should change`,
367      argumentHint: '[focus]',
368    })
369    void record($, { type: 'session_start', model: await $.session.model(), interactive: e.isInteractive })
370
371    return next(e)
372  })
373
374  on('prompt.submit', async ($, e, next) => {
375    if (HUMAN_ORIGINS.has(e.origin.kind)) {
376      pendingCorrection = hasPreviousTurnActed ? classifyCorrection($, e.text) : Promise.resolve(false)
377      void record($, { type: 'prompt', origin: e.origin.kind, text: clip(e.text, 400) })
378    }
379
380    return next(e)
381  })
382
383  on('skill.prompt', async ($, e, next) => {
384    void record($, { type: 'skill', skill: e.skill }, stats => increment(stats.skills, e.skill))
385
386    return next(e)
387  })
388
389  on('agent.spawn', async ($, e, next) => {
390    turn.spawnedAgents++
391    void record(
392      $,
393      { type: 'agent', agentId: e.parentAgentId, subagentType: e.subagentType, description: clip(e.description, 120) },
394      stats => increment(stats.subagents, e.subagentType),
395    )
396
397    return next(e)
398  })
399
400  on('tool.call', async ($, e, next) => {
401    const ran = await next(e)
402    const tool = String(e.tool)
403    const target = describeCall(tool, e as unknown as Record<string, unknown>, repo?.root)
404    const hasFailed = ran.deny !== undefined || ran.isError === true
405    const isGuidance = isGuidancePath(target)
406
407    if (hasFailed) {
408      turn.toolErrors++
409    }
410
411    if (e.agentId === undefined && ACTING_TOOLS.has(tool)) {
412      turn.hasActed = true
413    }
414
415    void record(
416      $,
417      {
418        type: 'tool',
419        tool,
420        agentId: e.agentId,
421        target,
422        ...(isGuidance ? { guidance: true } : {}),
423        ...(hasFailed ? { error: clip(ran.deny ?? ran.text ?? '', 400) } : {}),
424      },
425      stats => {
426        stats.toolCalls++
427        stats.toolErrors += hasFailed ? 1 : 0
428        stats.guidanceReads += isGuidance ? 1 : 0
429      },
430    )
431
432    return ran
433  })
434
435  on('turn.complete', async ($, e, next) => {
436    const done = await next(e)
437
438    if (e.agentId !== undefined) {
439      return done
440    }
441
442    const finished = turn
443    turn = { hasActed: false, toolErrors: 0, spawnedAgents: 0 }
444    hasPreviousTurnActed = finished.hasActed
445
446    const isCorrection = await pendingCorrection
447    pendingCorrection = Promise.resolve(false)
448    const triggers = triggersFor({ reason: e.reason, toolErrors: finished.toolErrors, spawnedAgents: finished.spawnedAgents, isCorrection })
449    let isObserving = false
450
451    await record($, { type: 'turn_end', reason: e.reason, triggers }, stats => {
452      stats.turns++
453      stats.aborts += e.reason === 'aborted' ? 1 : 0
454      stats.corrections += isCorrection ? 1 : 0
455
456      if (triggers.length > 0 && stats.observerRuns < MAX_OBSERVER_RUNS) {
457        stats.observerRuns++
458        isObserving = true
459      }
460    })
461
462    void (isObserving ? observe($, triggers) : flush($, false))
463
464    return done
465  })
466
467  on('session.end', async ($, e, next) => {
468    await record($, { type: 'session_end', reason: e.reason }, undefined, e.sessionId)
469    await flush($, true)
470
471    return next(e)
472  })
473
474  on('ui.render', { component: 'Spinner' }, async ($, e, next) => {
475    const count = state?.stats.observations ?? 0
476
477    return count === 0 ? next(e) : next({ ...e, props: { ...e.props, suffix: `${e.props.suffix} · tooling-retro: ${count}` } })
478  })
479
480  on('command.run', { command: COMMAND }, async ($, e) => {
481    void startRetro($, e.args.trim())
482
483    return { text: `Tooling retro over ${LOG_BRANCH} starts once this session's observations are pushed.` }
484  })
485}
486
hooks/log.ts 72 lines
1export type LogEntry = { type: string; at: string; agentId?: string } & Record<string, unknown>
2
3const SECRET_PATTERNS: readonly RegExp[] = [
4  /\bBearer\s+[A-Za-z0-9._~+/=-]{8,}/g,
5  /\b(api[_-]?key|token|secret|password|passwd|pwd|authorization|client[_-]?secret)\b(\s*["']?\s*[:=]\s*["']?)[^\s"',;]+/gi,
6  /\b(sk|pk|rk)[-_](live|test|ant)?[-_]?[A-Za-z0-9_-]{16,}/g,
7  /\bgh[pousr]_[A-Za-z0-9]{20,}/g,
8  /\bAKIA[0-9A-Z]{16}\b/g,
9  /\beyJ[A-Za-z0-9_-]{10,}\.[A-Za-z0-9_-]{10,}\.[A-Za-z0-9_-]{10,}/g,
10  /-----BEGIN [A-Z ]*PRIVATE KEY-----[\s\S]*?-----END [A-Z ]*PRIVATE KEY-----/g,
11]
12
13export function redact(text: string): string {
14  return SECRET_PATTERNS.reduce(
15    (current, pattern) =>
16      current.replace(pattern, (match, key: unknown, separator: unknown) =>
17        typeof key === 'string' && typeof separator === 'string' && /[:=]/.test(separator)
18          ? `${key}${separator}[redacted]`
19          : '[redacted]',
20      ),
21    text,
22  )
23}
24
25export function clip(text: string, max: number): string {
26  const clean = redact(text).replace(/\s+/g, ' ').trim()
27
28  return clean.length > max ? `${clean.slice(0, max)}…` : clean
29}
30
31const GUIDANCE_PATH = /(^|\/)(CLAUDE\.md|AGENTS\.md|SKILL\.md)$|(^|\/)\.claude\/(skills|agents|commands|rules)\/|(^|\/)\.ai\/|(^|\/)\.github\/(workflows|copilot-instructions)/
32
33export function isGuidancePath(path: string): boolean {
34  return GUIDANCE_PATH.test(path)
35}
36
37export function relativeTo(root: string | undefined, path: string): string {
38  if (root === undefined || !path.startsWith(`${root}/`)) {
39    return path
40  }
41
42  return path.slice(root.length + 1)
43}
44
45/**
46 * The argument of a tool call that tells a reader what it touched, without the payload.
47 */
48export function describeCall(tool: string, input: Record<string, unknown>, root?: string): string {
49  const pick = (key: string): string | undefined =>
50    typeof input[key] === 'string' ? (input[key] as string) : undefined
51  const path = pick('file_path') ?? pick('notebook_path') ?? pick('path')
52
53  if (tool === 'Bash') {
54    return clip(pick('command') ?? '', 200)
55  }
56
57  if (path !== undefined) {
58    return relativeTo(root, path)
59  }
60
61  return clip(pick('pattern') ?? pick('url') ?? pick('query') ?? pick('skill') ?? pick('description') ?? '', 160)
62}
63
64/** A remote URL without the credentials a token-based clone embeds in it. */
65export function withoutCredentials(remote: string): string {
66  return remote.replace(/\/\/[^@/]+@/, '//')
67}
68
69export function toLine(entry: LogEntry): string {
70  return `${JSON.stringify(entry)}\n`
71}
72
hooks/observe.ts 116 lines
1import { clip } from './log'
2
3export const TRIGGERS = ['abort', 'correction', 'tool-errors', 'subagent'] as const
4export type Trigger = (typeof TRIGGERS)[number]
5
6export const CATEGORIES = [
7  'guidance-not-loaded',
8  'guidance-ignored',
9  'guidance-wrong',
10  'guidance-conflict',
11  'guidance-missing',
12  'subagent-context',
13  'tool-friction',
14  'other',
15] as const
16export type Category = (typeof CATEGORIES)[number]
17
18export const CONFIDENCES = ['high', 'medium', 'low'] as const
19export type Confidence = (typeof CONFIDENCES)[number]
20
21/**
22 * The template every observation follows, whichever session or model wrote it, so the retro
23 * can group them by `category` and `tooling` without reading prose.
24 */
25export type Observation = {
26  category: Category
27  tooling: string
28  summary: string
29  evidence: string
30  suggestion: string
31  confidence: Confidence
32}
33
34export type TurnFacts = {
35  reason: 'answer' | 'aborted' | 'refusal' | 'error'
36  toolErrors: number
37  spawnedAgents: number
38  isCorrection: boolean
39}
40
41/** A single failed call is routine (a grep without match); two in one turn are a pattern. */
42export const TOOL_ERROR_THRESHOLD = 2
43
44export function triggersFor(turn: TurnFacts): Trigger[] {
45  const triggers: Trigger[] = []
46
47  if (turn.reason === 'aborted') {
48    triggers.push('abort')
49  }
50
51  if (turn.isCorrection) {
52    triggers.push('correction')
53  }
54
55  if (turn.toolErrors >= TOOL_ERROR_THRESHOLD) {
56    triggers.push('tool-errors')
57  }
58
59  if (turn.spawnedAgents > 0) {
60    triggers.push('subagent')
61  }
62
63  return triggers
64}
65
66const LIMITS = { tooling: 160, summary: 240, evidence: 400, suggestion: 400 } as const
67
68function isOneOf<T extends string>(values: readonly T[], value: unknown): value is T {
69  return typeof value === 'string' && (values as readonly string[]).includes(value)
70}
71
72function toObservation(value: unknown): Observation | undefined {
73  if (typeof value !== 'object' || value === null) {
74    return undefined
75  }
76
77  const row = value as Record<string, unknown>
78  const text = (key: keyof typeof LIMITS): string =>
79    typeof row[key] === 'string' ? clip(row[key] as string, LIMITS[key]) : ''
80
81  if (!isOneOf(CATEGORIES, row.category) || !isOneOf(CONFIDENCES, row.confidence)) {
82    return undefined
83  }
84
85  const observation = {
86    category: row.category,
87    tooling: text('tooling'),
88    summary: text('summary'),
89    evidence: text('evidence'),
90    suggestion: text('suggestion'),
91    confidence: row.confidence,
92  }
93
94  return observation.summary === '' || observation.evidence === '' ? undefined : observation
95}
96
97/**
98 * Reads the observer's reply: one JSON object per line, or NONE. Lines that are not a complete
99 * observation are dropped rather than repaired, so a sloppy reply costs data, never shape.
100 */
101export function parseObservations(reply: string): Observation[] {
102  return reply
103    .split('\n')
104    .map(line => line.trim().replace(/^```(json)?$/, ''))
105    .filter(line => line.startsWith('{'))
106    .flatMap(line => {
107      try {
108        const observation = toObservation(JSON.parse(line))
109
110        return observation === undefined ? [] : [observation]
111      } catch {
112        return []
113      }
114    })
115}
116