Observes sessions for moments where skills, agents or repo tooling let the agent down, pushes structured observations to chore/tooling-retro and runs…

Watches Claude Code sessions for moments where a repository's agent tooling let the agent down, and turns them into retrospectives on that tooling: CLAUDE.md files, skills, agents, rules, hooks and workflows.
Install: /plugin install tooling-retro@pp-claude-tooling (see the marketplace README for adding the marketplace).
Collect, locally. While a session runs, the mod writes one JSONL line per event to .git/tooling-retro/<session-id>.jsonl: skill loads, subagent starts, every tool call (what it touched, whether it failed, whether it read tooling), the person's prompts (clipped, redacted) and how each turn ended. Each observer run adds a line with its triggers, its outcome and the observer's reply (clipped to 4000 characters, redacted), so a NONE can be traced back too. This raw log never leaves the machine.
To see whether and why the observer ran in a session:
grep -E '"type":"(turn_end|observer)"' .git/tooling-retro/<session-id>.jsonl
Observe at triggers. At the end of a turn the mod checks four triggers:
| Trigger | Detected by |
|---|---|
abort | turn.complete reports the turn as interrupted |
correction | the person's next prompt after a turn that edited files or ran commands, classified by the engine's small model as a correction |
tool-errors | two or more failed tool calls in the turn |
subagent | the turn started a subagent |
When one fires, the mod asks a fork of the session ($.model.fork) whether the moment points at the repository's tooling. The fork sees the whole conversation from the prompt cache, runs on the session's own model, has no tools and leaves the conversation untouched. It answers NONE or up to three observations in a fixed template (prompts/observer.md):
{"category":"guidance-not-loaded","tooling":".claude/skills/pest-testing/SKILL.md","summary":"…","evidence":"…","suggestion":"…","confidence":"high"}
Categories: guidance-not-loaded, guidance-ignored, guidance-wrong, guidance-conflict, guidance-missing, subagent-context, tool-friction, other. Lines off the template are dropped, fields are redacted and clipped. At most 20 observer runs per session. A toast says when observations were noted, and the spinner shows the session's count.
Push. Observations go to observations/<day>/<session-id>.jsonl on the branch chore/tooling-retro, together with stats/<day>/<session-id>.json: counts of turns, aborts, corrections, tool calls and errors, tooling reads, and how often each skill and subagent type was used. No prompt or file content. A push happens whenever new observations exist, stats alone at most every 15 minutes and at session end, because cloud sessions lose anything that stays local. The branch name is fixed, has a history of its own and never merges into the code branches. Commits go through a temporary index and leave the working tree alone. A repository that opens pull requests for new branches automatically will open one for this branch too; keeping it out is that repository's call. A 403 or 407 on push stops pushing for the rest of the session.
Retro. /tooling-retro [focus] reads the observations and stats no earlier retro covered, takes an inventory of the repository's tooling, groups observations by tooling and category, and turns patterns into proposals. A proposal needs observations from at least two sessions, or one high-confidence observation backed by the stats. Each proposal comes with one or two variants that differ in substance. The report goes to retros/<day>.md on the same branch.
Decisions. The retro asks about each proposal on its own: a variant, Reject or Later. Every answer goes to decisions.jsonl on the branch before anything is implemented. Later retros read that file first: a rejected proposal does not come back, a later one does, and an accepted one returns only as a follow-up when the problem persists after the change. Accepted proposals are implemented on a normal working branch.
Nothing in the mod is specific to one repository: it reads whatever tooling the repository has.
Set through /config or pluginConfigs in settings:
| Option | Default | Effect |
|---|---|---|
push | true | Push observations and stats. Off: they stay in .git/tooling-retro/. |
git fetch origin +refs/heads/chore/tooling-retro:refs/tooling-retro/remote
git ls-tree -r --name-only refs/tooling-retro/remote
git show refs/tooling-retro/remote:observations/<day>/<session-id>.jsonl
claude plugin validate plugins/tooling-retro
claude plugin test plugins/tooling-retro
claude --plugin-dir plugins/tooling-retrohooks/register.ts 486 lines1import type { EngineInterface, Register } from 'claude-code'
2
3import { clip, describeCall, isGuidancePath, type LogEntry, toLine, withoutCredentials } from './log'
4import { parseObservations, type Trigger, triggersFor } from './observe'
5
6type Repo = { root: string; gitDir: string }
7type Entry = Omit<LogEntry, 'at'>
8
9type Stats = {
10 session: string
11 day: string
12 model: string | null
13 remote: string | null
14 turns: number
15 aborts: number
16 corrections: number
17 toolCalls: number
18 toolErrors: number
19 guidanceReads: number
20 skills: Record<string, number>
21 subagents: Record<string, number>
22 observerRuns: number
23 observations: number
24}
25
26type SessionState = {
27 id: string
28 raw: string
29 observations: string
30 stats: Stats
31 isStatsDirty: boolean
32 isObservationsDirty: boolean
33 lastPushAt: number
34}
35
36type TurnState = { hasActed: boolean; toolErrors: number; spawnedAgents: number }
37
38const COMMAND = 'tooling-retro'
39const LOG_BRANCH = 'chore/tooling-retro'
40const HUMAN_ORIGINS = new Set(['composer', 'bridge', 'sdk'])
41const ACTING_TOOLS = new Set(['Edit', 'Write', 'MultiEdit', 'NotebookEdit', 'Bash'])
42const CORRECTION_LABELS = ['correction', 'follow-up', 'new-task'] as const
43const MAX_OBSERVER_RUNS = 20
44const MAX_LOGGED_REPLY_CHARS = 4000
45const STATS_PUSH_INTERVAL_MS = 15 * 60_000
46const MAX_RAW_CHARS = 3_500_000
47const PUSH_ATTEMPTS = 3
48
49let isPushEnabled = true
50let repo: Repo | undefined | null = null
51let state: SessionState | undefined
52let turn: TurnState = { hasActed: false, toolErrors: 0, spawnedAgents: 0 }
53let hasPreviousTurnActed = false
54let pendingCorrection: Promise<boolean> = Promise.resolve(false)
55let queue: Promise<unknown> = Promise.resolve()
56
57function serial<T>(work: () => Promise<T>): Promise<T> {
58 const next = queue.then(work, work)
59 queue = next.catch(() => undefined)
60
61 return next
62}
63
64async function git($: EngineInterface, root: string, args: readonly string[], env?: Record<string, string>) {
65 return $.process.run(['git', ...args], { cwd: root, env, timeoutMs: 60_000 })
66}
67
68async function gitOut($: EngineInterface, root: string, args: readonly string[], env?: Record<string, string>) {
69 const ran = await git($, root, args, env)
70
71 if (ran.exitCode !== 0) {
72 throw new Error(`git ${args[0]} failed: ${ran.stderr.trim().slice(0, 300)}`)
73 }
74
75 return ran.stdout.trim()
76}
77
78async function resolveRepo($: EngineInterface): Promise<Repo | undefined> {
79 if (repo !== null) {
80 return repo
81 }
82
83 const found = await $.session.repo()
84
85 if (found === null || found.remote === null) {
86 repo = undefined
87
88 return repo
89 }
90
91 const ran = await git($, found.root, ['rev-parse', '--path-format=absolute', '--git-common-dir'])
92 repo = ran.exitCode === 0 ? { root: found.root, gitDir: ran.stdout.trim() } : undefined
93
94 return repo
95}
96
97function localPath(found: Repo, sessionId: string, suffix: string): string {
98 return `${found.gitDir}/tooling-retro/${sessionId}${suffix}`
99}
100
101async function readOrEmpty($: EngineInterface, path: string): Promise<string> {
102 return (await $.fs.exists(path)) ? $.fs.read(path) : ''
103}
104
105async function nowIso($: EngineInterface): Promise<string> {
106 return new Date(await $.clock.now()).toISOString()
107}
108
109/**
110 * The state of the session the engine runs now. A /clear starts a new session id, so the state
111 * is rebuilt from the local files whenever the id changes, and after a hot reload.
112 */
113async function currentSession($: EngineInterface, found: Repo, sessionId?: string): Promise<SessionState> {
114 const id = sessionId ?? (await $.session.id())
115
116 if (state !== undefined && state.id === id) {
117 return state
118 }
119
120 const statsText = await readOrEmpty($, localPath(found, id, '.stats.json'))
121 const remoteUrl = (await $.session.repo())?.remote
122 const remote = remoteUrl === undefined || remoteUrl === null ? null : withoutCredentials(remoteUrl)
123 const stats: Stats =
124 statsText === ''
125 ? {
126 session: id,
127 day: (await nowIso($)).slice(0, 10),
128 model: await $.session.model(),
129 remote,
130 turns: 0,
131 aborts: 0,
132 corrections: 0,
133 toolCalls: 0,
134 toolErrors: 0,
135 guidanceReads: 0,
136 skills: {},
137 subagents: {},
138 observerRuns: 0,
139 observations: 0,
140 }
141 : (JSON.parse(statsText) as Stats)
142
143 state = {
144 id,
145 raw: await readOrEmpty($, localPath(found, id, '.jsonl')),
146 observations: await readOrEmpty($, localPath(found, id, '.observations.jsonl')),
147 stats,
148 isStatsDirty: false,
149 isObservationsDirty: false,
150 lastPushAt: 0,
151 }
152
153 return state
154}
155
156function record($: EngineInterface, entry: Entry, update?: (stats: Stats) => void, sessionId?: string): Promise<void> {
157 return serial(() => writeRaw($, entry, update, sessionId))
158}
159
160async function writeRaw($: EngineInterface, entry: Entry, update?: (stats: Stats) => void, sessionId?: string) {
161 const found = await resolveRepo($)
162
163 if (found === undefined) {
164 return
165 }
166
167 const session = await currentSession($, found, sessionId)
168
169 if (update !== undefined) {
170 update(session.stats)
171 session.isStatsDirty = true
172 }
173
174 if (session.raw.length > MAX_RAW_CHARS) {
175 return
176 }
177
178 session.raw += toLine({ ...entry, at: await nowIso($) } as LogEntry)
179 await $.fs.write(localPath(found, session.id, '.jsonl'), session.raw)
180}
181
182async function persistStats($: EngineInterface, found: Repo, session: SessionState): Promise<void> {
183 await $.fs.write(localPath(found, session.id, '.stats.json'), `${JSON.stringify(session.stats, null, 2)}\n`)
184}
185
186function flush($: EngineInterface, isForced: boolean): Promise<void> {
187 return serial(() => pushIfDue($, isForced))
188}
189
190async function pushIfDue($: EngineInterface, isForced: boolean): Promise<void> {
191 const found = await resolveRepo($)
192
193 if (found === undefined || state === undefined) {
194 return
195 }
196
197 const session = state
198 await persistStats($, found, session)
199 const now = await $.clock.now()
200 const isStatsDue = session.isStatsDirty && (isForced || now - session.lastPushAt >= STATS_PUSH_INTERVAL_MS)
201
202 if (!isPushEnabled || session.stats.turns === 0 || (!session.isObservationsDirty && !isStatsDue)) {
203 return
204 }
205
206 try {
207 await pushSession($, found, session)
208 session.isStatsDirty = false
209 session.isObservationsDirty = false
210 session.lastPushAt = now
211 $.ui.status(undefined)
212 } catch (error) {
213 const message = error instanceof Error ? error.message : 'push failed'
214
215 if (/\b40[37]\b/.test(message)) {
216 isPushEnabled = false
217 $.ui.status(`tooling-retro: push to ${LOG_BRANCH} refused by policy, observations stay local`)
218
219 return
220 }
221
222 $.ui.status(`tooling-retro: ${message.slice(0, 120)}`)
223 }
224}
225
226/**
227 * Commits the session's observations and stats onto the log branch without touching the working
228 * tree or the real index: a temporary index, `commit-tree`, and a push of that commit. The raw
229 * event log never leaves the machine. Each session owns its files, so a rejected push only needs
230 * a refetch, never a merge. The commit uses the person's own git identity.
231 */
232async function pushSession($: EngineInterface, found: Repo, session: SessionState): Promise<void> {
233 const files = [{ source: localPath(found, session.id, '.stats.json'), target: `stats/${session.stats.day}/${session.id}.json` }]
234
235 if (session.observations !== '') {
236 files.push({
237 source: localPath(found, session.id, '.observations.jsonl'),
238 target: `observations/${session.stats.day}/${session.id}.jsonl`,
239 })
240 }
241
242 const index = { GIT_INDEX_FILE: `${found.gitDir}/tooling-retro/index.tmp` }
243 const root = found.root
244 let lastError = ''
245
246 for (let attempt = 0; attempt < PUSH_ATTEMPTS; attempt++) {
247 const fetched = await git($, root, ['fetch', '--quiet', 'origin', `+refs/heads/${LOG_BRANCH}:refs/tooling-retro/remote`])
248 const parent = fetched.exitCode === 0 ? await gitOut($, root, ['rev-parse', 'refs/tooling-retro/remote']) : undefined
249
250 await gitOut($, root, parent === undefined ? ['read-tree', '--empty'] : ['read-tree', parent], index)
251
252 for (const file of files) {
253 const blob = await gitOut($, root, ['hash-object', '-w', file.source])
254 await gitOut($, root, ['update-index', '--add', '--cacheinfo', `100644,${blob},${file.target}`], index)
255 }
256
257 const tree = await gitOut($, root, ['write-tree'], index)
258
259 if (parent !== undefined && tree === (await gitOut($, root, ['rev-parse', `${parent}^{tree}`]))) {
260 return
261 }
262
263 const commit = await gitOut($, root, [
264 'commit-tree',
265 tree,
266 ...(parent === undefined ? [] : ['-p', parent]),
267 '-m',
268 `tooling-retro: ${session.stats.observations} observations, session ${session.id}`,
269 ])
270 const pushed = await git($, root, ['push', '--quiet', 'origin', `${commit}:refs/heads/${LOG_BRANCH}`])
271
272 if (pushed.exitCode === 0) {
273 return
274 }
275
276 lastError = pushed.stderr.trim().slice(0, 300)
277 }
278
279 throw new Error(`push to ${LOG_BRANCH} failed: ${lastError}`)
280}
281
282async function classifyCorrection($: EngineInterface, text: string): Promise<boolean> {
283 try {
284 return (await $.model.classify(clip(text, 2000), CORRECTION_LABELS)) === 'correction'
285 } catch {
286 return false
287 }
288}
289
290/**
291 * Asks a fork of this very session, which sees the whole conversation from the prompt cache,
292 * whether the triggers point at the repository's tooling. Only the parsed observations are kept.
293 */
294async function observe($: EngineInterface, triggers: readonly Trigger[]): Promise<void> {
295 const found = await resolveRepo($)
296
297 if (found === undefined) {
298 return
299 }
300
301 try {
302 await runObserver($, found, triggers)
303 } finally {
304 await flush($, false)
305 }
306}
307
308async function runObserver($: EngineInterface, found: Repo, triggers: readonly Trigger[]): Promise<void> {
309 const template = await $.fs.read(`${$.plugin.root}/prompts/observer.md`)
310 const reply = await $.model.fork({ prompt: template.replaceAll('{{triggers}}', triggers.join(', ')) })
311
312 if (!reply.isAnswered) {
313 await record($, { type: 'observer', triggers, outcome: reply.reason })
314
315 return
316 }
317
318 const observations = parseObservations(reply.text)
319 await record($, {
320 type: 'observer',
321 triggers,
322 outcome: `${observations.length} observations`,
323 reply: clip(reply.text, MAX_LOGGED_REPLY_CHARS),
324 })
325
326 if (observations.length === 0) {
327 return
328 }
329
330 const at = await nowIso($)
331 const model = await $.session.model()
332
333 await serial(async () => {
334 const session = await currentSession($, found)
335 session.observations += observations.map(observation => toLine({ ...observation, triggers, model, at } as unknown as LogEntry)).join('')
336 session.stats.observations += observations.length
337 session.isObservationsDirty = true
338 session.isStatsDirty = true
339 await $.fs.write(localPath(found, session.id, '.observations.jsonl'), session.observations)
340 })
341
342 $.ui.toast(`tooling-retro: ${observations.length} ${observations.length === 1 ? 'observation' : 'observations'} noted`)
343 $.ui.invalidate('ui.render')
344}
345
346async function startRetro($: EngineInterface, focus: string): Promise<void> {
347 const template = await $.fs.read(`${$.plugin.root}/prompts/tooling-retro.md`)
348 const text = template
349 .replaceAll('{{branch}}', LOG_BRANCH)
350 .replaceAll('{{focus}}', focus === '' ? 'none, cover everything' : focus)
351
352 await flush($, true)
353 await $.prompt.submit({ text })
354}
355
356function increment(counts: Record<string, number>, key: string): void {
357 counts[key] = (counts[key] ?? 0) + 1
358}
359
360export const register: Register = (on, options) => {
361 isPushEnabled = options.push !== false
362
363 on('session.start', async ($, e, next) => {
364 await $.command.register({
365 name: COMMAND,
366 description: `Retro over the observations on ${LOG_BRANCH}: how skills, agents and repo tooling should change`,
367 argumentHint: '[focus]',
368 })
369 void record($, { type: 'session_start', model: await $.session.model(), interactive: e.isInteractive })
370
371 return next(e)
372 })
373
374 on('prompt.submit', async ($, e, next) => {
375 if (HUMAN_ORIGINS.has(e.origin.kind)) {
376 pendingCorrection = hasPreviousTurnActed ? classifyCorrection($, e.text) : Promise.resolve(false)
377 void record($, { type: 'prompt', origin: e.origin.kind, text: clip(e.text, 400) })
378 }
379
380 return next(e)
381 })
382
383 on('skill.prompt', async ($, e, next) => {
384 void record($, { type: 'skill', skill: e.skill }, stats => increment(stats.skills, e.skill))
385
386 return next(e)
387 })
388
389 on('agent.spawn', async ($, e, next) => {
390 turn.spawnedAgents++
391 void record(
392 $,
393 { type: 'agent', agentId: e.parentAgentId, subagentType: e.subagentType, description: clip(e.description, 120) },
394 stats => increment(stats.subagents, e.subagentType),
395 )
396
397 return next(e)
398 })
399
400 on('tool.call', async ($, e, next) => {
401 const ran = await next(e)
402 const tool = String(e.tool)
403 const target = describeCall(tool, e as unknown as Record<string, unknown>, repo?.root)
404 const hasFailed = ran.deny !== undefined || ran.isError === true
405 const isGuidance = isGuidancePath(target)
406
407 if (hasFailed) {
408 turn.toolErrors++
409 }
410
411 if (e.agentId === undefined && ACTING_TOOLS.has(tool)) {
412 turn.hasActed = true
413 }
414
415 void record(
416 $,
417 {
418 type: 'tool',
419 tool,
420 agentId: e.agentId,
421 target,
422 ...(isGuidance ? { guidance: true } : {}),
423 ...(hasFailed ? { error: clip(ran.deny ?? ran.text ?? '', 400) } : {}),
424 },
425 stats => {
426 stats.toolCalls++
427 stats.toolErrors += hasFailed ? 1 : 0
428 stats.guidanceReads += isGuidance ? 1 : 0
429 },
430 )
431
432 return ran
433 })
434
435 on('turn.complete', async ($, e, next) => {
436 const done = await next(e)
437
438 if (e.agentId !== undefined) {
439 return done
440 }
441
442 const finished = turn
443 turn = { hasActed: false, toolErrors: 0, spawnedAgents: 0 }
444 hasPreviousTurnActed = finished.hasActed
445
446 const isCorrection = await pendingCorrection
447 pendingCorrection = Promise.resolve(false)
448 const triggers = triggersFor({ reason: e.reason, toolErrors: finished.toolErrors, spawnedAgents: finished.spawnedAgents, isCorrection })
449 let isObserving = false
450
451 await record($, { type: 'turn_end', reason: e.reason, triggers }, stats => {
452 stats.turns++
453 stats.aborts += e.reason === 'aborted' ? 1 : 0
454 stats.corrections += isCorrection ? 1 : 0
455
456 if (triggers.length > 0 && stats.observerRuns < MAX_OBSERVER_RUNS) {
457 stats.observerRuns++
458 isObserving = true
459 }
460 })
461
462 void (isObserving ? observe($, triggers) : flush($, false))
463
464 return done
465 })
466
467 on('session.end', async ($, e, next) => {
468 await record($, { type: 'session_end', reason: e.reason }, undefined, e.sessionId)
469 await flush($, true)
470
471 return next(e)
472 })
473
474 on('ui.render', { component: 'Spinner' }, async ($, e, next) => {
475 const count = state?.stats.observations ?? 0
476
477 return count === 0 ? next(e) : next({ ...e, props: { ...e.props, suffix: `${e.props.suffix} · tooling-retro: ${count}` } })
478 })
479
480 on('command.run', { command: COMMAND }, async ($, e) => {
481 void startRetro($, e.args.trim())
482
483 return { text: `Tooling retro over ${LOG_BRANCH} starts once this session's observations are pushed.` }
484 })
485}
486hooks/log.ts 72 lines1export type LogEntry = { type: string; at: string; agentId?: string } & Record<string, unknown>
2
3const SECRET_PATTERNS: readonly RegExp[] = [
4 /\bBearer\s+[A-Za-z0-9._~+/=-]{8,}/g,
5 /\b(api[_-]?key|token|secret|password|passwd|pwd|authorization|client[_-]?secret)\b(\s*["']?\s*[:=]\s*["']?)[^\s"',;]+/gi,
6 /\b(sk|pk|rk)[-_](live|test|ant)?[-_]?[A-Za-z0-9_-]{16,}/g,
7 /\bgh[pousr]_[A-Za-z0-9]{20,}/g,
8 /\bAKIA[0-9A-Z]{16}\b/g,
9 /\beyJ[A-Za-z0-9_-]{10,}\.[A-Za-z0-9_-]{10,}\.[A-Za-z0-9_-]{10,}/g,
10 /-----BEGIN [A-Z ]*PRIVATE KEY-----[\s\S]*?-----END [A-Z ]*PRIVATE KEY-----/g,
11]
12
13export function redact(text: string): string {
14 return SECRET_PATTERNS.reduce(
15 (current, pattern) =>
16 current.replace(pattern, (match, key: unknown, separator: unknown) =>
17 typeof key === 'string' && typeof separator === 'string' && /[:=]/.test(separator)
18 ? `${key}${separator}[redacted]`
19 : '[redacted]',
20 ),
21 text,
22 )
23}
24
25export function clip(text: string, max: number): string {
26 const clean = redact(text).replace(/\s+/g, ' ').trim()
27
28 return clean.length > max ? `${clean.slice(0, max)}…` : clean
29}
30
31const GUIDANCE_PATH = /(^|\/)(CLAUDE\.md|AGENTS\.md|SKILL\.md)$|(^|\/)\.claude\/(skills|agents|commands|rules)\/|(^|\/)\.ai\/|(^|\/)\.github\/(workflows|copilot-instructions)/
32
33export function isGuidancePath(path: string): boolean {
34 return GUIDANCE_PATH.test(path)
35}
36
37export function relativeTo(root: string | undefined, path: string): string {
38 if (root === undefined || !path.startsWith(`${root}/`)) {
39 return path
40 }
41
42 return path.slice(root.length + 1)
43}
44
45/**
46 * The argument of a tool call that tells a reader what it touched, without the payload.
47 */
48export function describeCall(tool: string, input: Record<string, unknown>, root?: string): string {
49 const pick = (key: string): string | undefined =>
50 typeof input[key] === 'string' ? (input[key] as string) : undefined
51 const path = pick('file_path') ?? pick('notebook_path') ?? pick('path')
52
53 if (tool === 'Bash') {
54 return clip(pick('command') ?? '', 200)
55 }
56
57 if (path !== undefined) {
58 return relativeTo(root, path)
59 }
60
61 return clip(pick('pattern') ?? pick('url') ?? pick('query') ?? pick('skill') ?? pick('description') ?? '', 160)
62}
63
64/** A remote URL without the credentials a token-based clone embeds in it. */
65export function withoutCredentials(remote: string): string {
66 return remote.replace(/\/\/[^@/]+@/, '//')
67}
68
69export function toLine(entry: LogEntry): string {
70 return `${JSON.stringify(entry)}\n`
71}
72hooks/observe.ts 116 lines1import { clip } from './log'
2
3export const TRIGGERS = ['abort', 'correction', 'tool-errors', 'subagent'] as const
4export type Trigger = (typeof TRIGGERS)[number]
5
6export const CATEGORIES = [
7 'guidance-not-loaded',
8 'guidance-ignored',
9 'guidance-wrong',
10 'guidance-conflict',
11 'guidance-missing',
12 'subagent-context',
13 'tool-friction',
14 'other',
15] as const
16export type Category = (typeof CATEGORIES)[number]
17
18export const CONFIDENCES = ['high', 'medium', 'low'] as const
19export type Confidence = (typeof CONFIDENCES)[number]
20
21/**
22 * The template every observation follows, whichever session or model wrote it, so the retro
23 * can group them by `category` and `tooling` without reading prose.
24 */
25export type Observation = {
26 category: Category
27 tooling: string
28 summary: string
29 evidence: string
30 suggestion: string
31 confidence: Confidence
32}
33
34export type TurnFacts = {
35 reason: 'answer' | 'aborted' | 'refusal' | 'error'
36 toolErrors: number
37 spawnedAgents: number
38 isCorrection: boolean
39}
40
41/** A single failed call is routine (a grep without match); two in one turn are a pattern. */
42export const TOOL_ERROR_THRESHOLD = 2
43
44export function triggersFor(turn: TurnFacts): Trigger[] {
45 const triggers: Trigger[] = []
46
47 if (turn.reason === 'aborted') {
48 triggers.push('abort')
49 }
50
51 if (turn.isCorrection) {
52 triggers.push('correction')
53 }
54
55 if (turn.toolErrors >= TOOL_ERROR_THRESHOLD) {
56 triggers.push('tool-errors')
57 }
58
59 if (turn.spawnedAgents > 0) {
60 triggers.push('subagent')
61 }
62
63 return triggers
64}
65
66const LIMITS = { tooling: 160, summary: 240, evidence: 400, suggestion: 400 } as const
67
68function isOneOf<T extends string>(values: readonly T[], value: unknown): value is T {
69 return typeof value === 'string' && (values as readonly string[]).includes(value)
70}
71
72function toObservation(value: unknown): Observation | undefined {
73 if (typeof value !== 'object' || value === null) {
74 return undefined
75 }
76
77 const row = value as Record<string, unknown>
78 const text = (key: keyof typeof LIMITS): string =>
79 typeof row[key] === 'string' ? clip(row[key] as string, LIMITS[key]) : ''
80
81 if (!isOneOf(CATEGORIES, row.category) || !isOneOf(CONFIDENCES, row.confidence)) {
82 return undefined
83 }
84
85 const observation = {
86 category: row.category,
87 tooling: text('tooling'),
88 summary: text('summary'),
89 evidence: text('evidence'),
90 suggestion: text('suggestion'),
91 confidence: row.confidence,
92 }
93
94 return observation.summary === '' || observation.evidence === '' ? undefined : observation
95}
96
97/**
98 * Reads the observer's reply: one JSON object per line, or NONE. Lines that are not a complete
99 * observation are dropped rather than repaired, so a sloppy reply costs data, never shape.
100 */
101export function parseObservations(reply: string): Observation[] {
102 return reply
103 .split('\n')
104 .map(line => line.trim().replace(/^```(json)?$/, ''))
105 .filter(line => line.startsWith('{'))
106 .flatMap(line => {
107 try {
108 const observation = toObservation(JSON.parse(line))
109
110 return observation === undefined ? [] : [observation]
111 } catch {
112 return []
113 }
114 })
115}
116