Collects what is waiting on you in a session: questions, Claude's notes and PR reviews

A Claude Code mod that gathers and intelligently manages outstanding questions, tasks, issues, and opportunities into one nicely organized place alongside your session. Just use /inbox.
https://github.com/user-attachments/assets/560f80be-e6fb-484e-8a3f-7b4828b98fb2
In a Claude Code session, type:
/plugin install inbox --marketplace petekp/inbox
Answer y to add the marketplace, then choose a scope. It starts working in that session.
gh, signed in.To update, run claude plugin update inbox, then /reload-plugins in any session that's already open.
A line above your prompt shows what the session is working on and how many things are waiting on you.
Type /inbox to see them. It opens a panel with three tabs:
Each item shows what you can do with it. To answer a question, press the letter next to the answer you want, and it goes to Claude as your message. You can also have Claude fix a finding, mark a task done, or dismiss anything you don't need.
In the panel, 1, 2 and 3 switch tabs, j and k move up and down, and an action's letter runs it. You can also click. ctrl+x tab moves between the panel and the session. When the tab you're on is empty and something new arrives on another tab, the panel switches to it.
After 15 minutes with no activity, or when you resume a session, the line above the prompt expands into a short summary of where things stand.
Claude runs each test or build as a command of its own, so the mod knows exactly how it ended. If Claude says something passes when the last run failed, or hasn't run since an edit, the mod asks Claude to run it again or say the change is untested.
To see every part with sample items, run /inbox demo. Run it again to go back to your own.
You might be wondering how many extra tokens the inbox uses. It adds two things:
| What the inbox adds | Tokens | Cost at API prices |
|---|---|---|
| A Sonnet call after each of Claude's replies, to update the inbox | 3–4k in, about 2k of it cached. 100–250 out. | About 0.5¢ per reply, or 1¢ after 5 idle minutes |
| Its instructions and tools, sent with every request Claude makes | About 1,350, cached | Under 0.1¢ per request |
In one long Opus session, this came to about 1% of the total cost. A shorter or cheaper session pays a larger share. On a Claude plan, these count toward your usage limits.
Run the mod from a clone:
claude --plugin-dir /path/to/inbox
The session reloads the mod when you change one of its files. ./scripts/check.sh runs Prettier, claude plugin validate, a type check and the tests. The type check needs the types Claude Code writes to .claude-plugin/types/ the first time it loads the mod.
hooks/register.tsx 3736 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register, SessionAppendMessage, UiPressArgument, UiScrollResult } from 'claude-code'
3
4import type {
5 Arrival,
6 Check,
7 Checks,
8 Cursor,
9 Closed,
10 Dialog,
11 Help,
12 Item,
13 Ledger,
14 Finding,
15 PrCheck,
16 PrFixSent,
17 PrThread,
18 PrView,
19 PrViews,
20 Presence,
21 Previous,
22 LastAction,
23 Settled,
24 Snapshot,
25 Stop,
26 Tab,
27} from '../types'
28import {
29 checkName,
30 checkRun,
31 checkLine,
32 checksIn,
33 claimMessage,
34 distinctSummary,
35 failureLines,
36 failCount,
37 failureList,
38 fixMessage,
39 readResult,
40 resolvePath,
41} from './checks'
42import type { Contradiction } from './checks'
43import {
44 NO_CHECKS,
45 addRepo,
46 contains,
47 bandLines,
48 changed,
49 checkKey,
50 claimAgainst,
51 dismissed,
52 fixSent,
53 needsYou,
54 pruned,
55 recorded,
56 turnEnded,
57 upgradeChecks,
58} from './check-tracking'
59import { demoView } from './demo'
60import type { View } from './demo'
61import type { Exchange, Press, Update } from './ledger'
62import type { Handoffs, PrStatus } from './prs'
63import {
64 THREADS_QUERY,
65 VIEW_FIELDS,
66 checkCounts,
67 commentLine,
68 failingChecks,
69 namedPrs,
70 parseRef,
71 prAttention,
72 prRefs,
73 prRowsOnYou,
74 prompts,
75 threadsOnYou,
76 readThreads,
77 readView,
78 readableComment,
79 readiness,
80 threadWhere,
81 waitingThreads,
82} from './prs'
83import {
84 EMPTY,
85 ago,
86 applyUpdate,
87 buildPrompt,
88 carryText,
89 closeItem,
90 type Closing,
91 commandRowLine,
92 dialogLine,
93 parseReply,
94 catchUpPrompt,
95 CLAUDE_CODE,
96 isLapsed,
97 promptNotes,
98 resetTime,
99 readCommandRow,
100 readKind,
101 screenText,
102 statusLine,
103 stopFix,
104 stopKindOf,
105 stopText,
106 systemText,
107 tasksRunBy,
108 toolActivity,
109 transcriptCatchUpPrompt,
110 transcriptText,
111 TOLD_NOTHING,
112 type Told,
113 upgradeLedger,
114} from './ledger'
115import { baseName, clipLabel, helpLabel, isTaskHandedOff, localPath, messages, openCommands, steps } from './presses'
116import type { HelpStep } from './presses'
117import {
118 CLOSE_DESCRIPTION,
119 CLOSE_SCHEMA,
120 FINDING_SCHEMA,
121 findingDescription,
122 recordClose,
123 recordFinding,
124} from './tools'
125import { readRepo, type TreeIO } from './tree'
126
127const LEDGER = atom({ plugin: 'inbox', key: 'ledger' } as const, EMPTY)
128const PRESENCE = atom(
129 { plugin: 'inbox', key: 'presence' } as const,
130 {
131 lastActiveAt: 0,
132 isAway: false,
133 isUpdating: false,
134 ledgerState: 'current',
135 turnsStarted: 0,
136 turnsApplied: 0,
137 minute: 0,
138 } as Presence,
139)
140const PREVIOUS = atom({ plugin: 'inbox', key: 'previous' } as const, null as Previous | null)
141const TAB = atom({ plugin: 'inbox', key: 'tab' } as const, 'needsYou' as Tab)
142const NO_CURSOR: Cursor = { id: null, index: 0 }
143const NO_SELECTION: Record<Tab, Cursor> = { needsYou: NO_CURSOR, findings: NO_CURSOR, prs: NO_CURSOR }
144const SELECTION = atom({ plugin: 'inbox', key: 'selection' } as const, NO_SELECTION)
145// Whether Claude Code uses its `dark` theme, which picks the selected row's tint.
146const THEME = atom({ plugin: 'inbox', key: 'theme' } as const, '')
147const PR_VIEWS = atom(
148 { plugin: 'inbox', key: 'prViews' } as const,
149 { views: {}, branchRef: null, isFetching: false } as PrViews,
150)
151const PR_FIXES_SENT = atom({ plugin: 'inbox', key: 'prFixesSent' } as const, {} as Record<string, PrFixSent>)
152const STOP = atom({ plugin: 'inbox', key: 'stop' } as const, null as Stop | null)
153const DIALOGS = atom({ plugin: 'inbox', key: 'dialogs' } as const, [] as Dialog[])
154const EDIT_TOOLS = new Set(['Edit', 'Write', 'NotebookEdit'])
155const CHECKS = atom({ plugin: 'inbox', key: 'checks' } as const, NO_CHECKS)
156const SNAPSHOTS = atom({ plugin: 'inbox', key: 'snapshots' } as const, {} as Record<string, Snapshot>)
157const TYPING = atom({ plugin: 'inbox', key: 'typing' } as const, null as string | null)
158const SETTLED = atom({ plugin: 'inbox', key: 'settled' } as const, [] as Settled[])
159const ARRIVAL = atom({ plugin: 'inbox', key: 'arrival' } as const, null as Arrival | null)
160const LAST_ACTIONS = atom({ plugin: 'inbox', key: 'lastActions' } as const, {} as Record<string, LastAction>)
161const UNFOLDED = atom({ plugin: 'inbox', key: 'unfolded' } as const, [] as Item['kind'][])
162const SHOWN_DETAILS = atom({ plugin: 'inbox', key: 'shownDetails' } as const, [] as string[])
163const IS_KEY_LIST_SHOWN = atom({ plugin: 'inbox', key: 'isKeyListShown' } as const, false)
164// How long a closed item's row stays in place, with its outcome, before it moves to Closed.
165const SETTLED_MS = 5120
166// The bar under a just-closed row, in cells. It loses half a cell per step of SETTLED_MS, so the pane redraws that often.
167const LEAVE_BAR_CELLS = 12
168const LEAVE_BAR_STEPS = LEAVE_BAR_CELLS * 2
169// After a jump, how long the new tab's top edge takes to draw in, in steps, and how long the new row's bar keeps the tab's color.
170const DRAW_IN_MS = 200
171const DRAW_IN_STEPS = 4
172const ARRIVAL_MS = 1500
173const IS_DEMO = atom({ plugin: 'inbox', key: 'isDemo' } as const, false)
174const SAMPLE_PRESS = 'Sample entry: nothing was sent. Run /inbox demo to go back.'
175// The pane's buttons that only move around it, which work in the demo. Every other press there sends nothing.
176const DEMO_PRESSES = /^(tab-|title-|select-|typekey-|fold-|key-list$|next$|previous$)/
177const PR_POLL_MS = 2 * 60_000
178const MAX_PRS = 6
179
180const FINDING_TOOL = 'mcp__inbox__record_finding'
181const GUIDANCE = `# Inbox
182The inbox plugin shows the user what waits on them: your open questions and the tasks only they can do, in a band above their prompt and in the /inbox pane, and your findings, in the pane's Findings tab.
183
184A finding is something you noticed that deserves the user's attention but is outside the current task: a bug, a risk, missing tests, tech debt, or a chance to improve something. Record it with mcp__inbox__record_finding the moment you notice it, then go on with the task; fixing it waits until the user asks. Record one too at two moments that are easy to pass over while focused on the task:
185- You work around a problem instead of fixing it, such as copying files by hand because a tool does not reach them.
186- Part of your change could not be tested or verified.
187State those two in your reply as well. Mention any other finding only when it bears on what the user asked.
188
189The latest "inbox:" text beside the user's prompt is the current state. It lists every open item and finding with its id, such as [i35] or [f12], and anything it does not list is closed. Check it before telling the user that something is open or waits on them.
190
191Close with mcp__inbox__close:
192- When the user's message answers an open item, close it first, before other work, with their answer.
193- When an item or finding is done or no longer applies, close it with a short reason, without waiting to be asked.
194The answer to a question is the user's to give. Close a question with their answer, or once it no longer applies, and never with an answer of your own.`
195const CLOSE_TOOL = 'mcp__inbox__close'
196
197const RUN_CHECK_TOOL = 'mcp__inbox__run_check'
198const RUN_CHECK_DESCRIPTION = `Run tests, type checks, lints, builds and validation scripts with this tool, not with Bash. Each command is one check, such as "npm test", "npx vitest run src/a.test.ts" or "tsc --noEmit", with no pipe, redirect or other command. They run in order in the shell's current folder; to check another folder, cd there with Bash first. Each check's full output goes to a log file, and you get back whether it passed, its summary, the lines that name what failed, and the log's path.`
199const RUN_CHECK_SCHEMA = {
200 type: 'object',
201 properties: {
202 checks: {
203 type: 'array',
204 items: { type: 'string' },
205 description: 'The check commands to run, in order, one check each.',
206 },
207 },
208 required: ['checks'],
209}
210const BASH_CHECK_REFUSAL = `Not run. Run each check on its own, with no pipe, redirect or other command, so its exit status is the check's: with ${RUN_CHECK_TOOL}, which saves the output to a log file, or alone with Bash.`
211const SUBAGENT_CHECK_REFUSAL = `Not run: ${RUN_CHECK_TOOL} runs checks in the main conversation's folder, not a subagent's. Run each check with Bash instead.`
212
213const PANE = 'inbox'
214// Theme keys, so the colors follow the person's Claude Code theme.
215const ACCENT = 'claude'
216const NEEDS_YOU = 'warning'
217// A child's place under its section: a middle child, the last, or a block the tree passes.
218type TreePos = 'mid' | 'last' | 'pass'
219// More tree lines than a row wraps to; the tree's Box clips the rest.
220const TREE_DEPTH = 200
221// The bar that marks a selected row where the palette has no selection color.
222const SELECTION_BAR = Array<string>(TREE_DEPTH).fill('▌').join('\n')
223// The tree column's text for each place and lead, built once each.
224const treeTexts = new Map<string, string>()
225
226/** The tree column for a row: `lead` lines of │, its branch, then │ or blank below. */
227function treeText(pos: TreePos, lead: number): string {
228 const key = `${pos}:${lead}`
229 const cached = treeTexts.get(key)
230 if (cached !== undefined) return cached
231 const text = [
232 ...Array<string>(lead).fill('│'),
233 pos === 'last' ? '└─' : pos === 'mid' ? '├─' : '│',
234 ...Array<string>(TREE_DEPTH).fill(pos === 'last' ? ' ' : '│'),
235 ].join('\n')
236 treeTexts.set(key, text)
237
238 return text
239}
240const DONE = 'success'
241const FINDINGS = 'autoAccept'
242const PRS = 'planMode'
243
244// A status color in the pane. Each tab has the tone of its own name.
245type Tone = Tab | 'done' | 'error'
246/**
247 * The pane's colors in one Claude Code theme. Each text color meets WCAG AA,
248 * 4.5:1, on every background the pane draws it on, and each marker meets 3:1.
249 * The terminal draws Button labels in its own text color, and the theme's dim
250 * gray for `dimColor`, so the pane gives Buttons no dimColor. The footer's Keys
251 * toggle and the Closed and Details folds are the exceptions: secondary
252 * controls, they brighten under the pointer.
253 */
254type Palette = {
255 /** Painted under the whole pane, over the theme's own background. */
256 body?: string
257 /** The sections' cards. */
258 card?: string
259 /** An unselected tab's panel. */
260 tab?: string
261 /** The selected tab's panel, and an unselected tab's under the pointer. */
262 raised: string
263 /** The text on `raised`; without one, the default text color. */
264 raisedText?: string
265 /** The selected row's background. Without one, a bar in `key` marks the row. */
266 selection?: string
267 /** Secondary text: ages, counts, labels and closed items. */
268 muted: string
269 /** The letters of a selected row's keys. */
270 key: string
271 /** The tree's lines. */
272 line: string
273 /** The rules between rows. */
274 divider: string
275 /** Text in a status color. A tone left out draws its text in the default color. */
276 tone: Partial<Record<Tone, string>>
277 /** A marker in a status color, such as ✓, and the selected tab's top line. */
278 mark: Record<Tone, string>
279}
280
281// Hex tuned for one theme, each tone checked against the pane, the cards, the
282// selected tab and the selected row. They hold only for the theme they name.
283const DARK_TONES = { needsYou: '#ffc107', findings: '#cab0ff', prs: '#8cc8c0', done: '#7ecd8f', error: '#ffa3b0' }
284const DARK_DALTONIZED_TONES = {
285 needsYou: '#ffcc00',
286 findings: '#cab0ff',
287 prs: '#a3c2c2',
288 done: '#82c1ff',
289 error: '#ffa3a3',
290}
291const LIGHT_TONES = { needsYou: '#745417', findings: '#7b00e8', prs: '#006363', done: '#25652f', error: '#a5293d' }
292const LIGHT_DALTONIZED_TONES = {
293 needsYou: '#7a4900',
294 findings: '#7400db',
295 prs: '#2d5a5a',
296 done: '#005885',
297 error: '#ad0000',
298}
299const DARK_PALETTE: Palette = {
300 card: '#373737',
301 tab: '#373737',
302 raised: '#4c4c4c',
303 selection: '#3b4654',
304 muted: '#bdbdbd',
305 key: '#b1b9f9',
306 line: '#505050',
307 divider: '#444444',
308 tone: DARK_TONES,
309 // A marker needs only 3:1, so a failing check's ✗ takes Claude Code's own red, which text here could not.
310 mark: { ...DARK_TONES, error: '#ff6b80' },
311}
312const LIGHT_PALETTE: Palette = {
313 card: '#e6e6e6',
314 tab: '#e6e6e6',
315 raised: '#d0d0d0',
316 selection: '#b4d5ff',
317 muted: '#595959',
318 key: '#243bf5',
319 line: '#afafaf',
320 divider: '#c8c8c8',
321 tone: LIGHT_TONES,
322 mark: LIGHT_TONES,
323}
324// The ANSI themes take the terminal's 16 colors, and Claude Code draws the
325// pane on a palette gray that default text fails against. The pane paints its
326// own body in the palette color nearest the terminal's background, draws text
327// only in the default and muted colors, and puts status colors on markers.
328// A color name such as "black" is drawn as fixed RGB, so each color is a theme
329// key whose value in that ANSI theme is the palette color named beside it.
330// Checked with Ghostty's default palette and its "Apple System Colors Light".
331const DARK_ANSI_PALETTE: Palette = {
332 body: 'inverseText', // black
333 raised: 'inactive', // white
334 raisedText: 'inverseText', // black
335 muted: 'inactive', // white
336 key: 'suggestion', // bright blue
337 line: 'userMessageBackground', // bright black
338 divider: 'userMessageBackground',
339 tone: {},
340 mark: { needsYou: NEEDS_YOU, findings: FINDINGS, prs: PRS, done: DONE, error: 'error' }, // the bright colors
341}
342const LIGHT_ANSI_PALETTE: Palette = {
343 body: 'userMessageBackgroundHover', // bright white
344 raised: 'text', // black
345 raisedText: 'userMessageBackgroundHover', // bright white
346 muted: 'inactive', // bright black
347 key: 'suggestion', // blue
348 line: 'userMessageBackground', // white
349 divider: 'userMessageBackground',
350 tone: {},
351 // Red, magenta, blue, green and red: the theme's yellow and cyan fall under 3:1 on bright white.
352 mark: { needsYou: 'error', findings: FINDINGS, prs: 'suggestion', done: DONE, error: 'error' },
353}
354const PALETTES: Record<string, Palette> = {
355 dark: DARK_PALETTE,
356 'dark-daltonized': {
357 ...DARK_PALETTE,
358 key: '#99ccff',
359 tone: DARK_DALTONIZED_TONES,
360 mark: { ...DARK_DALTONIZED_TONES, error: '#ff6666' },
361 },
362 light: LIGHT_PALETTE,
363 'light-daltonized': {
364 ...LIGHT_PALETTE,
365 card: '#dcdcdc',
366 tab: '#dcdcdc',
367 raised: '#c8c8c8',
368 muted: '#545454',
369 key: '#003ae8',
370 divider: '#c0c0c0',
371 tone: LIGHT_DALTONIZED_TONES,
372 mark: LIGHT_DALTONIZED_TONES,
373 },
374 'dark-ansi': DARK_ANSI_PALETTE,
375 'light-ansi': LIGHT_ANSI_PALETTE,
376}
377// Auto and custom themes: theme keys whose values pass in both Claude Code's
378// dark and light themes, since the mod cannot tell which one Auto chose. Text
379// sits on the pane's own background, as in the ANSI palettes.
380const THEME_KEY_PALETTE: Palette = {
381 tab: 'userMessageBackground',
382 raised: 'subtle',
383 raisedText: 'text',
384 muted: 'inactive',
385 key: 'remember',
386 line: 'subtle',
387 divider: 'subtle',
388 tone: {},
389 mark: { needsYou: NEEDS_YOU, findings: FINDINGS, prs: PRS, done: DONE, error: 'error' },
390}
391const TABS: { id: Tab; label: string; hotkey: string }[] = [
392 { id: 'needsYou', label: 'Needs you', hotkey: '1' },
393 { id: 'findings', label: 'Findings', hotkey: '2' },
394 { id: 'prs', label: 'PRs', hotkey: '3' },
395]
396// The blank columns on each side of a docked tab's name, which a click there also selects.
397const TAB_PAD = ' '
398// The Needs you tab lists questions first, because each takes one key.
399const NEEDS_YOU_GROUPS: { kind: Item['kind']; title: string; empty: string }[] = [
400 { kind: 'question', title: 'Questions', empty: 'No questions are waiting on you.' },
401 { kind: 'task', title: 'Your tasks', empty: 'No tasks are waiting on you.' },
402]
403// How many recently closed items each group lists under its open ones.
404const CLOSED_SHOWN = 3
405const MODEL = 'sonnet'
406const AWAY_MS = 15 * 60_000
407const PREVIOUS_MAX_AGE_MS = 7 * 24 * 60 * 60_000
408const KEPT_SESSIONS = 40
409const LOCAL_URL = /https?:\/\/(?:localhost|127\.0\.0\.1|0\.0\.0\.0|[\w.-]+\.localhost)(?::\d+)?[^\s"'`)\]]*/g
410
411// One turn's exchange, gathered across hooks. Module variables reset on a hot
412// reload, which only loses the turn in progress.
413let activity: string[] = []
414let person: string | null = null
415let trigger: string | null = null
416let isTurnRunning = false
417let isOn = false
418// The press behind this turn's prompt.
419let press: Press | null = null
420// Replies a Stop hook sent Claude back from this turn, before the final one.
421let sentBack: string[] = []
422// Set in session.start, which a hot reload runs again.
423let sessionId = ''
424let root = ''
425let home = ''
426// The repo's top folder, where git reads the working tree; null outside git.
427let top: string | null = null
428// Each folder outside the session's where a check ran, with its repo's top folder.
429const repoTops = new Map<string, Promise<string | null>>()
430let refreshing: Promise<void> = Promise.resolve()
431// The PR fetches, one at a time, and the next one when it is waiting to start: whether it looks up the branch's PR.
432let fetchingPrs: Promise<void> = Promise.resolve()
433let nextFetchFindsBranchPr: boolean | null = null
434// Each tab's row ids as of the last write that can add a row, so a row that appears later counts as new.
435let recordedRows: { isDemo: boolean; ids: Partial<Record<Tab, Set<string>>> } | null = null
436let isSaved = false
437let queue: Promise<void> = Promise.resolve()
438// Counts the conversations this process has run, so an update queued in one never writes into the next.
439let conversation = 0
440// What Claude last read of the inbox beside a prompt, so it is sent again only when it changed.
441let told: Told = TOLD_NOTHING
442// The line last published to this pane's Herdr sidebar row, and the chain that publishes in order.
443let published: string | null = null
444let publishing: Promise<void> = Promise.resolve()
445// Context for prompts this mod sent, by text, appended just before each prompt's row.
446const contextFor = new Map<string, string[]>()
447// The `!` command whose output row comes next.
448let shellCommand: string | null = null
449
450/** Clears everything gathered for the turn's exchange. */
451function resetTurn() {
452 activity = []
453 person = null
454 trigger = null
455 press = null
456 sentBack = []
457}
458
459/** Sends Claude back when the reply claims a check passes against its latest run. */
460async function blockOnClaim<R extends { block?: string }>($: EngineInterface, r: R, reply: string): Promise<R> {
461 if ((await read($, CHECKS)).results.length === 0) return r
462 await refreshTree($)
463 const sent: { claim: Contradiction | null } = { claim: null }
464 await update($, CHECKS, c => {
465 const found = claimAgainst(c, reply, root)
466 sent.claim = found.claim
467
468 return found.checks
469 })
470
471 return sent.claim ? { ...r, block: claimMessage(sent.claim) } : r
472}
473
474/** Adds a message or a command of the person's to what they sent this turn. */
475function notePerson(line: string) {
476 person = person === null ? line : `${person}\n\n${line}`
477}
478
479function noteActivity(line: string) {
480 if (activity.length < 40 && !activity.includes(line)) activity.push(line)
481}
482
483/** A transcript row's text, its text blocks joined. */
484function messageText(message: SessionAppendMessage): string {
485 return message.content.map(b => (b.type === 'text' ? b.text : '')).join('')
486}
487
488/** A tool result's fields, or none when the result is not an object. */
489function resultFields(ran: { result?: unknown }): Record<string, unknown> {
490 return ran.result && typeof ran.result === 'object' ? (ran.result as Record<string, unknown>) : {}
491}
492
493/** Reads which Claude Code theme draws the pane. */
494async function syncTheme($: EngineInterface) {
495 const theme = (await $.config.list().catch(() => [])).find(row => row.key === 'theme')?.value
496 await update($, THEME, () => (typeof theme === 'string' ? theme : ''))
497}
498
499async function save($: EngineInterface, ledger: Ledger) {
500 const savedAt = await $.clock.now()
501 const key = `s:${sessionId}`
502 // The first save in a process deletes the key first, moving it to the end of
503 // the store's insertion order, which the pruning treats as most recent. Later
504 // saves only overwrite, so pruning runs once per process.
505 if (!isSaved) await $.store.delete(key)
506 await Promise.all([
507 $.store.set(key, { savedAt, ledger }),
508 ledger.card ? $.store.set(`p:${root}`, { savedAt, ledger }) : undefined,
509 ])
510 if (isSaved) return
511 isSaved = true
512 const sessions = (await $.store.keys()).filter(k => k.startsWith('s:'))
513 await Promise.all(sessions.slice(0, Math.max(0, sessions.length - KEPT_SESSIONS)).map(old => $.store.delete(old)))
514}
515
516/**
517 * Publishes the session's status line to its Herdr pane, where a sidebar row
518 * showing the `inbox` token reads it. Outside Herdr it does nothing. Calls run
519 * in order and read the state when they run, so an older line never lands
520 * after a newer one, as when a dialog opens and closes in quick succession.
521 */
522function publishStatus($: EngineInterface, isEnding = false): Promise<void> {
523 publishing = publishing
524 .then(async () => {
525 const pane = await $.env.get('HERDR_PANE_ID')
526 if (!pane) return
527 const [ledger, stop, dialogs, lastActions, presence] = await Promise.all([
528 read($, LEDGER),
529 read($, STOP),
530 read($, DIALOGS),
531 read($, LAST_ACTIONS),
532 read($, PRESENCE),
533 ])
534 // A task handed to Claude waits on no one, as in the band.
535 const items = ledger.items.filter(i => i.kind !== 'task' || !isTaskHandedOff(lastActions[i.id], presence))
536 const line = isEnding ? '' : statusLine({ ...ledger, items }, stop, dialogs)
537 if (line === published) return
538 published = line
539 const token = line ? ['--token', `inbox=${line}`] : ['--clear-token', 'inbox']
540 await $.process.run(['herdr', 'pane', 'report-metadata', pane, '--source', 'inbox', ...token], {
541 timeoutMs: 5000,
542 })
543 })
544 .catch(() => undefined)
545
546 return publishing
547}
548
549async function setStop($: EngineInterface, stop: Stop | null) {
550 if (stop === null && (await read($, STOP)) === null) return
551 await update($, STOP, () => stop)
552 void publishStatus($)
553}
554
555/**
556 * A tool call's identity: the tool and its main argument. A dialog keeps it,
557 * so it closes when its own call ends, not when a like call does.
558 */
559function callKey(tool: string, input: Record<string, unknown>): string {
560 const main = ['command', 'file_path', 'notebook_path', 'url', 'pattern', 'query', 'skill', 'description'].find(
561 k => typeof input[k] === 'string',
562 )
563
564 return main ? `${tool} ${main}=${String(input[main])}` : `${tool} ${JSON.stringify(input)}`
565}
566
567async function openDialog($: EngineInterface, dialog: Dialog) {
568 await update($, DIALOGS, d => [...d.filter(x => x.key !== dialog.key), dialog])
569 void publishStatus($)
570}
571
572/** Closes the dialogs of one call, or every dialog. */
573async function closeDialogs($: EngineInterface, key: string | null) {
574 if (!(await read($, DIALOGS)).some(d => key === null || d.key === key)) return
575 await update($, DIALOGS, d => (key === null ? [] : d.filter(x => x.key !== key)))
576 void publishStatus($)
577}
578
579/** Runs a tool call and closes its dialog when the call resolves, since `next` returns only after the person answers any prompt for it. */
580async function clearingDialog<T>($: EngineInterface, key: string, run: () => Promise<T>): Promise<T> {
581 try {
582 return await run()
583 } finally {
584 await closeDialogs($, key)
585 }
586}
587
588/** What a permission prompt asks, in a few words: "push main to origin", "edit README.md". */
589function permissionText(tool: string, input: Record<string, unknown>): string {
590 const text = (key: string) => (typeof input[key] === 'string' ? (input[key] as string) : '')
591 const file = (key: string) => baseName(text(key))
592 switch (tool) {
593 case 'Bash': {
594 const said = text('description') || text('command')
595 // "Push main to origin" reads as "Allow push main to origin?"; an acronym keeps its case.
596 return clipLabel(/^[A-Z][a-z]/.test(said) ? said.charAt(0).toLowerCase() + said.slice(1) : said, 60)
597 }
598 case 'Edit':
599 case 'MultiEdit':
600 case 'Write':
601 return `edit ${file('file_path')}`
602 case 'NotebookEdit':
603 return `edit ${file('notebook_path')}`
604 case 'Read':
605 return `read ${file('file_path')}`
606 case 'Glob':
607 case 'Grep':
608 return `search ${text('path') ? `${file('path')} ` : ''}for ${clipLabel(text('pattern'), 30)}`
609 case 'WebFetch':
610 return `fetch ${
611 text('url')
612 .replace(/^https?:\/\//, '')
613 .split('/')[0]
614 }`
615 case 'WebSearch':
616 return `search the web for ${clipLabel(text('query'), 30)}`
617 case 'Skill':
618 return `use the ${text('skill')} skill`
619 case 'Agent':
620 case 'Task':
621 return text('description') ? `run an agent: ${text('description')}` : 'run an agent'
622 default:
623 return tool.startsWith('mcp__') ? tool.slice(5).replace(/__/g, ' ') : tool
624 }
625}
626
627/** Changes the ledger and saves it, so a resumed session finds it. */
628async function commitLedger($: EngineInterface, change: (l: Ledger) => Ledger): Promise<void> {
629 let before: Ledger = EMPTY
630 const after = await update($, LEDGER, l => {
631 before = l
632 return change(l)
633 })
634 await save($, after)
635 void publishStatus($)
636 await showSettled($, before, after)
637 // After showSettled, so a row this change closed holds its tab.
638 void followNewRows($).catch(() => undefined)
639}
640
641/**
642 * The open items of one kind in the order the Needs you tab lists them.
643 * Questions go newest first, so a new batch's numbers match Claude's in its reply.
644 */
645function listedItems(items: Item[], kind: Item['kind']): Item[] {
646 const listed = items.filter(i => i.kind === kind)
647 return kind === 'question' ? listed.sort((a, b) => b.turn - a.turn) : listed
648}
649
650/**
651 * Keeps each item that just closed in its row for a few seconds, with its
652 * outcome, so the person sees it was registered, however it closed.
653 */
654async function showSettled($: EngineInterface, before: Ledger, after: Ledger) {
655 const open = new Set(after.items.map(i => i.id))
656 await settle(
657 $,
658 before.items.flatMap(item => {
659 const d = open.has(item.id) ? undefined : after.closed.find(x => x.id === item.id)
660 return d ? [{ ...d, index: listedItems(before.items, item.kind).indexOf(item) }] : []
661 }),
662 )
663}
664
665/** Keeps rows that just closed in place, each with a leave bar, until SETTLED_MS passes. */
666async function settle($: EngineInterface, settled: Settled[]) {
667 if (settled.length === 0) return
668 await update($, SETTLED, s => [...s.filter(x => !settled.some(y => y.id === x.id)), ...settled])
669 expireSettled($, settled, SETTLED_MS)
670 redrawWhileLeaving($, () => update($, SETTLED, s => [...s]), SETTLED_MS)
671}
672
673/** What a just-closed row says: what it was, and how it closed. */
674function settledText(s: Settled): { what: string; outcome: string } {
675 return s.kind === 'check' ? { what: s.title, outcome: s.outcome } : { what: s.ask, outcome: outcomeText(s) }
676}
677
678/** Redraws the pane at each step of a just-closed row's leave bar, for `waitMs`. A fresh value is what redraws it. */
679function redrawWhileLeaving($: EngineInterface, refresh: () => Promise<unknown>, waitMs: number) {
680 // Whole milliseconds, so the last wait ends at `waitMs` and its redraw finds the row gone.
681 const step = Math.ceil(SETTLED_MS / LEAVE_BAR_STEPS)
682 void (async () => {
683 for (let left = waitMs; left > 0; left -= step) {
684 await $.clock.sleep(Math.min(step, left))
685 await refresh()
686 }
687 })().catch(() => undefined)
688}
689
690/** Records what a row's action did. A row the press removed shows it in its place until SETTLED_MS passes. */
691async function recordLastAction($: EngineInterface, id: string, last: Omit<LastAction, 'at' | 'turnsStarted'>) {
692 const [at, { turnsStarted }] = await Promise.all([$.clock.now(), read($, PRESENCE)])
693 await update($, LAST_ACTIONS, a => ({ ...a, [id]: { ...last, at, turnsStarted } }))
694 if (last.isHandoff) void publishStatus($)
695 // A finding the press removed shows in its place, with a leave bar, until the settle time is over.
696 // A fresh object redraws the pane, so the bar shrinks and the place then clears.
697 redrawWhileLeaving($, () => update($, LAST_ACTIONS, a => ({ ...a })), SETTLED_MS)
698}
699
700/** Removes just-closed rows after a wait. A reload cancels the wait, so session.start sets it again. */
701function expireSettled($: EngineInterface, settled: Settled[], waitMs: number) {
702 void $.clock
703 .sleep(waitMs)
704 .then(() => update($, SETTLED, s => s.filter(x => !settled.some(y => y.id === x.id && y.at === x.at))))
705 .catch(() => undefined)
706}
707
708type ModelResult = Awaited<ReturnType<EngineInterface['model']['fork']>>
709
710type LedgerState = Presence['ledgerState']
711
712/** Applies the ledger model's reply. */
713async function applyLedgerReply(
714 $: EngineInterface,
715 at: number,
716 r: ModelResult,
717 source: string | null,
718 change: (l: Ledger, u: Update, now: number) => Ledger,
719): Promise<LedgerState> {
720 if (!r.isAnswered) {
721 // A fork has nothing to read until this process sends its first request,
722 // as after a resume. The catch-up runs again after the next reply.
723 return r.reason === 'nothing-to-fork' ? 'behind' : 'failed'
724 }
725 const parsed = parseReply(r.text, source)
726 if (!parsed || at !== conversation) return 'failed'
727 const now = await $.clock.now()
728 await commitLedger($, l => change(l, parsed, now))
729
730 return 'current'
731}
732
733/** Asks the ledger model to update the ledger, with systemText as its instructions, cached for the next call within five minutes. */
734function askLedgerModel($: EngineInterface, prompt: string, timeoutMs: number): Promise<ModelResult> {
735 const system = [{ text: systemText(CLAUDE_CODE), cache: true as const }]
736
737 return $.model.complete({ model: MODEL, system, prompt, maxTokens: 1600, effort: 'low', timeoutMs })
738}
739
740/** Updates the ledger from one exchange. */
741async function runUpdate($: EngineInterface, at: number, ex: Exchange): Promise<LedgerState> {
742 const promptAt = await $.clock.now()
743 const r = await askLedgerModel($, buildPrompt(await read($, LEDGER), ex), 45_000)
744 // An Explain turn talks about its item without deciding it.
745 const explained = ex.press?.action === 'explain' ? ex.press.id : null
746
747 return applyLedgerReply($, at, r, [ex.reply, ...ex.activity].join('\n'), (l, u, now) =>
748 applyUpdate(l, { ...u, closed: u.closed.filter(c => c.id !== explained) }, now, ex.turn, promptAt),
749 )
750}
751
752/**
753 * Brings the ledger up to date over the whole conversation, for turns the
754 * per-turn update missed: it closes what was handled and adds what still waits.
755 */
756async function catchUp($: EngineInterface, at: number): Promise<LedgerState> {
757 const [ledger, shown] = await Promise.all([read($, LEDGER), screen($)])
758 const change = (l: Ledger, u: Update, now: number) => applyUpdate(l, u, now, l.turn)
759 const forked = await $.model.fork({ prompt: catchUpPrompt(CLAUDE_CODE, ledger, shown) })
760 if (forked.isAnswered || forked.reason !== 'nothing-to-fork') return applyLedgerReply($, at, forked, null, change)
761 // A resumed conversation has no request of this process's to fork until its
762 // first reply, so the model reads the transcript instead, uncached.
763 const transcript = transcriptText(await $.session.messages())
764 if (!transcript) return 'behind'
765 const r = await askLedgerModel($, transcriptCatchUpPrompt(ledger, shown, transcript), 90_000)
766
767 return applyLedgerReply($, at, r, transcript, change)
768}
769
770/**
771 * Queues the next update: a catch-up while the ledger may have missed a turn,
772 * else this exchange alone.
773 */
774function queueUpdate($: EngineInterface, next: { ex: Exchange; turnsStarted: number } | null) {
775 const at = conversation
776 queue = queue.then(async () => {
777 if (at !== conversation) return
778 const { ledgerState, turnsStarted } = await update($, PRESENCE, p => ({ ...p, isUpdating: true }))
779 const state = await (ledgerState !== 'current' || next === null ? catchUp($, at) : runUpdate($, at, next.ex)).catch(
780 (): LedgerState => 'failed',
781 )
782 if (at !== conversation) return
783 // An exchange covers the turns started by its end; a catch-up, those started before it ran.
784 const applied = next?.turnsStarted ?? turnsStarted
785 await update($, PRESENCE, p => ({
786 ...p,
787 ledgerState: state,
788 isUpdating: false,
789 turnsApplied: state === 'current' ? Math.max(p.turnsApplied, applied) : p.turnsApplied,
790 }))
791 // A task handed to Claude unfolds once the update for its turn applies.
792 void publishStatus($)
793 })
794 // A rejected link would skip every later update, so the chain swallows it.
795 queue = queue.catch(() => undefined)
796}
797
798async function tick($: EngineInterface) {
799 const now = await $.clock.now()
800 const p = await read($, PRESENCE)
801 if (p.isAway) {
802 await update($, PRESENCE, q => ({ ...q, minute: Math.floor(now / 60_000) }))
803 } else if (!isTurnRunning && p.lastActiveAt > 0 && now - p.lastActiveAt > AWAY_MS) {
804 await update($, PRESENCE, q => ({ ...q, isAway: true, minute: Math.floor(now / 60_000) }))
805 }
806}
807
808/**
809 * Sends an answer to Claude as the person's own message. The item closes at
810 * once, so a second press cannot send it twice. A prompt sent mid-turn waits
811 * for the turn to end.
812 */
813async function sendAnswer($: EngineInterface, item: Item, answer: string) {
814 await close($, item.id, { how: 'answered', outcome: answer })
815 await send($, messages.answer(item, answer), { id: item.id, action: 'answer' })
816}
817
818/**
819 * Sends a prompt as the person's own message. A plugin's own prompt skips that
820 * plugin's prompt.submit hook, so the bookkeeping happens here, and what
821 * Claude reads beside the prompt goes just before it, in a row only the model sees.
822 */
823async function send($: EngineInterface, text: string, sentBy: Press | null = null) {
824 const context = await notePrompt($, text, sentBy)
825 // The engine may run the prompt now, after the running turn, or inside it, so
826 // the context goes in when the prompt's own row is stored (session.append).
827 if (context.length > 0) contextFor.set(text, context)
828 await $.prompt.submit({ text, asUser: true })
829}
830
831/** Appends context as a row only the model reads. Without it Claude still gets the prompt, so a refused append is ignored. */
832async function appendContext($: EngineInterface, context: string[]) {
833 await $.session
834 .append({ message: { type: 'user', content: [{ type: 'text', text: context.join('\n\n') }] } })
835 .catch(() => undefined)
836}
837
838/**
839 * Records a prompt in the person's words and returns what Claude reads beside
840 * it: the previous session's card when they continue from it, the questions a
841 * numbered answer refers to, and the inbox when it changed.
842 */
843async function notePrompt($: EngineInterface, text: string, sentBy: Press | null): Promise<string[]> {
844 const [now, ledger, prev, isShown] = await Promise.all([
845 $.clock.now(),
846 update($, LEDGER, l => ({ ...l, turn: l.turn + 1 })),
847 read($, PREVIOUS),
848 isPaneShown($),
849 ])
850 await Promise.all([
851 update($, PRESENCE, p => ({ ...p, lastActiveAt: now, isAway: false })),
852 prev ? update($, PREVIOUS, () => null) : undefined,
853 ])
854 const notes: string[] = []
855
856 if (prev?.isBroughtIn) {
857 const carried = carryText(
858 prev.ledger,
859 `inbox: the user chose to continue from the previous session in this folder (${ago(now - prev.savedAt)}). Where it stood:`,
860 )
861 if (carried) notes.push(carried)
862 }
863 const r = promptNotes(CLAUDE_CODE, ledger, { text, isPress: sentBy !== null, isOpen: isShown }, told)
864 notes.push(...r.notes)
865 told = r.told
866
867 if (sentBy) press = sentBy
868 notePerson(text)
869 trigger = null
870
871 return notes
872}
873
874/** Asks Claude what an item is about. The item stays open, since nothing was closed. */
875async function explain($: EngineInterface, item: Item) {
876 await send($, messages.explain(item), { id: item.id, action: 'explain' })
877}
878
879async function close($: EngineInterface, id: string, closing: Closing) {
880 const now = await $.clock.now()
881 await commitLedger($, l => closeItem(l, id, closing, now))
882}
883
884/** Opens a path an item names, as `openCommands` decides. */
885async function openPath($: EngineInterface, raw: string) {
886 const path = localPath(raw, root, (await $.env.get('HOME')) ?? '')
887 const name = baseName(path)
888 if (!(await $.fs.exists(path))) {
889 $.ui.toast(`${name} is not there anymore`)
890 return
891 }
892 const stat = await $.fs.stat(path)
893 const isFile = stat.kind === 'file'
894 const isExecutable = isFile && (await $.process.run(['test', '-x', path])).exitCode === 0
895 const { argv, fallback } = openCommands(path, isFile, isExecutable)
896 const r = await $.process.run(argv)
897 const retry = r.exitCode !== 0 && fallback ? await $.process.run(fallback) : r
898 if (retry.exitCode !== 0) $.ui.toast(`Could not open ${name}: ${retry.stderr.trim()}`)
899}
900
901async function useHelp($: EngineInterface, item: Item, help: Help, press: UiPressArgument) {
902 if (help.kind === 'open') {
903 await openPath($, help.path)
904 } else if (help.kind === 'copy') {
905 const r = await $.ui.copy({ text: help.text, surface: press.surface })
906 $.ui.toast(r.isCopied ? `Copied ${help.name ?? 'snippet'}` : 'Could not copy to the clipboard')
907 } else if (help.kind === 'run') {
908 await send($, messages.run(item, help.command), { id: item.id, action: 'run' })
909 } else if (help.kind === 'terminal') {
910 // A filled "! command" reaches the model as text; only a typed "!" switches the prompt to shell mode.
911 const r = await $.ui.copy({ text: help.command, surface: press.surface })
912 const what = help.name ?? 'the command'
913 $.ui.toast(
914 r.isCopied
915 ? `Copied ${what}. Run it in a terminal, or type ! here and paste.`
916 : 'Could not copy to the clipboard',
917 )
918 } else {
919 await openUrl($, help.url)
920 }
921}
922
923/** Each step's helps in order: one press copies then opens, for example. */
924async function useStep($: EngineInterface, item: Item, step: Help[], press: UiPressArgument) {
925 for (const help of step) await useHelp($, item, help, press)
926}
927
928/** A path inside the repo, relative to its top folder; any other path as given. */
929function repoPath(path: string): string {
930 const base = top ?? root
931
932 return base && path.startsWith(`${base}/`) ? path.slice(base.length + 1) : path
933}
934
935/** Text with the home folder written as ~, as in a shell prompt. */
936function tilde(text: string): string {
937 return home ? text.replaceAll(`${home}/`, '~/') : text
938}
939
940/**
941 * Brings state an earlier version of the mod wrote up to date, and returns the
942 * ledger. A reload keeps $.state, so a running session can still hold findings
943 * under `notes`, and `notes` as its tab, and the Needs you tab as `waiting`,
944 * with item kinds `decide` and `do`, PR threads without a time or
945 * `isLinesChanged`, and last actions that don't say whether they handed work off.
946 */
947async function upgradeState($: EngineInterface): Promise<Ledger> {
948 const [ledger] = await Promise.all([
949 update($, LEDGER, upgradeLedger),
950 update($, PREVIOUS, p => p && { ...p, ledger: upgradeLedger(p.ledger) }),
951 update($, TAB, t => ((t as string) === 'notes' ? 'findings' : (t as string) === 'waiting' ? 'needsYou' : t)),
952 update($, SELECTION, ({ waiting, ...s }: Record<Tab, Cursor> & { waiting?: Cursor }) => ({
953 ...NO_SELECTION,
954 ...(waiting ? { needsYou: waiting } : {}),
955 ...s,
956 })),
957 update($, UNFOLDED, u => u.map(readKind)),
958 // Before last actions said whether they handed work off, Address, a task's Run step and a typed reply did.
959 // An earlier build saved one for Open and Open log, which record nothing now, as the page they open shows the press.
960 update($, LAST_ACTIONS, a =>
961 Object.fromEntries(
962 Object.entries(a)
963 .filter(([, last]) => !/^(open|log)-/.test(last.action))
964 .map(([id, last]) => [
965 id,
966 {
967 ...last,
968 // Earlier versions saved "Discuss sent" or "Sent to Claude to fix"; a last action is now the label pressed.
969 text: /^Sent to Claude to /.test(last.text) ? 'Address' : last.text.replace(/ sent$/, ''),
970 isHandoff: last.isHandoff ?? /^(address-|help-|typed$)/.test(last.action),
971 },
972 ]),
973 ),
974 ),
975 update($, SETTLED, s => s.map(x => (x.kind === 'check' ? x : { ...x, kind: readKind(x.kind) }))),
976 // Presence from before the turn counts existed cannot say whether the last
977 // turn's update landed, so that load catches up once.
978 update($, PRESENCE, p => ({ ...p, turnsStarted: p.turnsStarted ?? 1, turnsApplied: p.turnsApplied ?? 0 })),
979 update($, CHECKS, c => upgradeChecks(c, root, home)),
980 update($, PR_VIEWS, v => ({
981 ...v,
982 views: Object.fromEntries(
983 Object.entries(v.views).map(([ref, pr]) => [
984 ref,
985 {
986 ...pr,
987 threads: pr.threads.map(t => ({
988 ...t,
989 isLinesChanged: t.isLinesChanged ?? false,
990 at: t.at ?? null,
991 reply: t.reply && { ...t.reply, at: t.reply.at ?? null },
992 })),
993 },
994 ]),
995 ),
996 })),
997 ])
998
999 return ledger
1000}
1001
1002/** Runs a record_finding or close call on the ledger. Returns the text the agent reads back. */
1003async function runTool(
1004 $: EngineInterface,
1005 tool: typeof recordFinding,
1006 input: Record<string, unknown>,
1007): Promise<string> {
1008 const now = await $.clock.now()
1009 let result = ''
1010 await commitLedger($, l => {
1011 const r = tool(CLAUDE_CODE, l, input, now)
1012 result = r.result
1013 return r.ledger
1014 })
1015
1016 return result
1017}
1018
1019/**
1020 * Opens the pane, or raises it, asking for the keyboard. The surface grants
1021 * that only while the prompt holds the keys over an empty composer.
1022 */
1023function openPane($: EngineInterface) {
1024 return $.ui.open({ id: PANE, title: 'Inbox', focus: true, closeOnEscape: true })
1025}
1026
1027/** Opens the free-text field under a row and gives it the keyboard. */
1028async function startTyping($: EngineInterface, id: string) {
1029 await update($, TYPING, () => id)
1030 // A click leaves the keys with the prompt, and `ui.focus` is refused in a pane that lacks them.
1031 await openPane($)
1032 await $.ui.focus({ requestId: PANE, key: `type-${id}` }).catch(() => undefined)
1033}
1034
1035/**
1036 * Sends the person's own words about an item as their message. A question
1037 * closes with those words as its answer, as an option press does. A task
1038 * stays open, since Claude closes it once the message settles it.
1039 */
1040async function sendTypedForItem($: EngineInterface, item: Item, text: string) {
1041 await update($, TYPING, () => null)
1042 const words = text.trim()
1043 if (!words) return
1044 if (item.kind === 'task') await send($, messages.taskReply(item, words))
1045 else await sendAnswer($, item, words)
1046}
1047
1048/** Sends the person's own words about a finding, which leaves the Findings tab as Address does. */
1049async function sendTypedForFinding($: EngineInterface, finding: Finding, text: string) {
1050 await update($, TYPING, () => null)
1051 const words = text.trim()
1052 if (!words) return
1053 await removeFinding($, finding.id)
1054 await send($, messages.finding(finding, 'typed', words))
1055}
1056
1057async function removeFinding($: EngineInterface, id: string) {
1058 await commitLedger($, l => ({ ...l, findings: l.findings.filter(f => f.id !== id) }))
1059}
1060
1061/**
1062 * How many things wait on the person: open questions, tasks not handed to
1063 * Claude, and the session's checks left failing with no fix sent. The band and
1064 * the Needs you tab both read this.
1065 */
1066function needsYouCount(
1067 ledger: Ledger,
1068 checks: Checks,
1069 lastActions: Record<string, LastAction>,
1070 turns: View['turns'],
1071): number {
1072 const items = ledger.items.filter(i => i.kind !== 'task' || !isTaskHandedOff(lastActions[i.id], turns))
1073
1074 return items.length + needsYou(checks, root).count
1075}
1076
1077/** Hides a failing check's row in the Needs you tab until the check runs again. */
1078async function dismissCheck($: EngineInterface, check: Check) {
1079 await update($, CHECKS, c => dismissed(c, check))
1080}
1081
1082/**
1083 * Asks Claude to fix a failing check. Its row stays in place with a ✓ and
1084 * "Fix" for a few seconds, then leaves until the check runs again.
1085 */
1086async function sendFix($: EngineInterface, check: Check) {
1087 await send($, fixMessage(check, root))
1088 const at = await $.clock.now()
1089 const id = checkRowId(check)
1090 let index = -1
1091 await update($, CHECKS, c => {
1092 index = waitingChecks(c).findIndex(x => checkRowId(x) === id)
1093 return fixSent(c, check, at)
1094 })
1095 if (index >= 0) await settle($, [{ id, kind: 'check', title: checkName(check, root), outcome: 'Fix', at, index }])
1096}
1097
1098/** The session's failing checks the Needs you tab lists: left failing, not dismissed, and not handed to Claude with Fix. */
1099function waitingChecks(checks: Checks): Check[] {
1100 return needsYou(checks, root).rows.filter(c => c.fixSentAt === null)
1101}
1102
1103/** Sends the finding back to Claude, to fix it or to talk it through first. */
1104async function actOnFinding($: EngineInterface, finding: Finding, how: 'address' | 'discuss') {
1105 await removeFinding($, finding.id)
1106 await send($, messages.finding(finding, how))
1107}
1108
1109async function showTab($: EngineInterface, tab: Tab) {
1110 if (tab === 'prs') await findPrs($)
1111 await update($, TAB, () => tab)
1112 await $.ui.scroll({ in: PANE, to: 'start' }).catch(() => undefined)
1113}
1114
1115/**
1116 * Refreshes the PRs tab as it comes into view, asking gh for the branch's PR.
1117 * Call it before the tab draws: it marks the fetch first, so the tab opens on
1118 * "Checking for PRs…" and not on "No PRs". The demo shows sample PRs, so it asks nothing.
1119 */
1120async function findPrs($: EngineInterface) {
1121 if (await read($, IS_DEMO)) return
1122 await update($, PR_VIEWS, v => ({ ...v, isFetching: true }))
1123 void fetchPrs($, true)
1124}
1125
1126/**
1127 * Selects a row and scrolls the pane the least that shows it whole. The row's
1128 * scroll target is drawn only once the row redraws expanded, and the engine
1129 * refuses a key it has not drawn, so the scroll retries for a few frames.
1130 * Measured before the redraw, the row's end would land out of view.
1131 */
1132async function select($: EngineInterface, tab: Tab, id: string, index: number) {
1133 await update($, SELECTION, s => ({ ...s, [tab]: { id, index } }))
1134 await update($, TYPING, t => (t === id ? t : null))
1135 for (let tries = 0; tries < 10; tries++) {
1136 const result = await $.ui
1137 .scroll({ in: PANE, to: { key: `view-${id}` }, block: 'nearest' })
1138 .catch((): UiScrollResult => ({}))
1139 if (result.deny === undefined) return
1140 await $.clock.sleep(30)
1141 }
1142}
1143
1144/**
1145 * The selected row's position: the cursor's row while it exists, else the row
1146 * now at its old position, so closing an item selects the one after it.
1147 */
1148function selectedIndex(ids: string[], cursor: Cursor): number {
1149 if (ids.length === 0) return -1
1150 const at = cursor.id === null ? -1 : ids.indexOf(cursor.id)
1151
1152 return at >= 0 ? at : Math.min(cursor.index, ids.length - 1)
1153}
1154
1155/**
1156 * Records each tab's row ids and returns, by tab, the ids it did not have at
1157 * the last record. A tab passed as null is loading: it keeps its record and
1158 * returns nothing. The first record, and the first after the demo turns on or
1159 * off, return nothing.
1160 */
1161function newRows(isDemo: boolean, current: Record<Tab, string[] | null>): Record<Tab, string[]> {
1162 const before = recordedRows?.isDemo === isDemo ? recordedRows.ids : {}
1163 const ids = { ...before }
1164 const added: Record<Tab, string[]> = { needsYou: [], findings: [], prs: [] }
1165 for (const { id: tab } of TABS) {
1166 const listed = current[tab]
1167 if (listed === null) continue
1168 const seen = before[tab]
1169 if (seen) added[tab] = listed.filter(id => !seen.has(id))
1170 ids[tab] = new Set(listed)
1171 }
1172 recordedRows = { isDemo, ids }
1173
1174 return added
1175}
1176
1177/**
1178 * Moves the pane to a tab and selects a row there, or shows the tab's top when
1179 * there is none. The pane then redraws while the tab's top edge draws in, and
1180 * once more when the row's arrival bar ends.
1181 */
1182async function jumpTo($: EngineInterface, tab: Tab, row: { id: string; index: number } | null) {
1183 const at = await $.clock.now()
1184 await update($, ARRIVAL, () => ({ tab, id: row?.id ?? null, at }))
1185 await update($, TAB, () => tab)
1186 if (row) await select($, tab, row.id, row.index)
1187 else await $.ui.scroll({ in: PANE, to: 'start' }).catch(() => undefined)
1188 const redrawAt = [
1189 ...Array.from({ length: DRAW_IN_STEPS }, (_, n) => ((n + 1) * DRAW_IN_MS) / DRAW_IN_STEPS),
1190 ARRIVAL_MS,
1191 ]
1192 void (async () => {
1193 for (const due of redrawAt) {
1194 await $.clock.sleep(Math.max(0, at + due - (await $.clock.now())))
1195 await update($, ARRIVAL, a => a && { ...a })
1196 }
1197 })().catch(() => undefined)
1198}
1199
1200/** Each tab's row ids, in the order the pane lists them. */hooks/checks.ts 551 lines1// Checks measured from Claude's commands: which commands were tests, type
2// checks, lints, builds or validations, how each ended, whether the files
3// changed after it, and which successes Claude's reply claims. Nothing here
4// asks a model.
5
6import type { Check, CheckKind, Target } from '../types'
7
8/** One recognized check inside a command, such as `bun test` or `tsc`. */
9export type CheckCall = { name: string; kind: CheckKind }
10
11/** A package manager's options before its script, such as `--prefix app` or `-w pkg`. */
12const PM_OPTIONS = String.raw`(?:-{1,2}[\w-]+(?:=\S+|\s+(?!-|(?:run|test|build|lint)\b)\S+)?\s+)*`
13/** A script such as `test` or `lint:css`; the match's group 2 is its `:css` part, so `lint` and `lint:css` are different checks. */
14const pm = (managers: string, script: string) =>
15 new RegExp(String.raw`\b(${managers})\s+${PM_OPTIONS}(?:run\s+)?(?:${script})(:[\w.:-]*\w)?\b`)
16
17const CHECKS: { kind: CheckKind; pattern: RegExp; name: (m: RegExpMatchArray) => string }[] = [
18 { kind: 'tests', pattern: /\bclaude\s+plugin\s+test\b/, name: () => 'plugin tests' },
19 { kind: 'validate', pattern: /\bclaude\s+plugin\s+validate\b/, name: () => 'plugin validate' },
20 { kind: 'tests', pattern: pm('bun|npm|pnpm|yarn|deno', 'test'), name: m => `${m[1]} test${m[2] ?? ''}` },
21 { kind: 'tests', pattern: /\bnode\s+(?:-{1,2}[\w-]+(?:=\S+)?\s+)*--test\b/, name: () => 'node --test' },
22 { kind: 'tests', pattern: /\b(vitest|jest|pytest|rspec|mocha|phpunit|ava|tap)\b/, name: m => m[1] ?? 'tests' },
23 { kind: 'tests', pattern: /\b(go|cargo|swift|mix|dotnet)\s+test\b/, name: m => `${m[1]} test` },
24 { kind: 'tests', pattern: /\bplaywright\s+test\b/, name: () => 'playwright' },
25 { kind: 'tests', pattern: /\bxcodebuild\b[^|;&]*\btest\b/, name: () => 'xcodebuild test' },
26 { kind: 'tests', pattern: /\bmake\s+(?:check|test)\b/, name: m => m[0] },
27 { kind: 'types', pattern: /\btsc\b/, name: () => 'tsc' },
28 { kind: 'types', pattern: /\b(mypy|pyright)\b/, name: m => m[1] ?? 'types' },
29 {
30 kind: 'types',
31 pattern: pm('bun|npm|pnpm|yarn', 'typecheck|type-check|check-types'),
32 name: m => `typecheck${m[2] ?? ''}`,
33 },
34 { kind: 'types', pattern: /\bcargo\s+check\b/, name: () => 'cargo check' },
35 {
36 kind: 'lint',
37 pattern: /\b(eslint|biome|ruff|shellcheck|swiftlint|clippy|golangci-lint|stylelint)\b/,
38 name: m => m[1] ?? 'lint',
39 },
40 { kind: 'lint', pattern: pm('bun|npm|pnpm|yarn', 'lint'), name: m => `lint${m[2] ?? ''}` },
41 { kind: 'lint', pattern: /\bprettier\b[^|;&]*--check\b/, name: () => 'prettier' },
42 { kind: 'build', pattern: pm('bun|npm|pnpm|yarn', 'build'), name: m => `${m[1]} build${m[2] ?? ''}` },
43 { kind: 'build', pattern: /\b(cargo|go|swift)\s+build\b/, name: m => `${m[1]} build` },
44 { kind: 'build', pattern: /\bxcodebuild\b(?![^|;&]*\btest\b)/, name: () => 'xcodebuild' },
45 { kind: 'build', pattern: /\b(?:next|vite)\s+build\b/, name: m => m[0] },
46 {
47 kind: 'validate',
48 pattern: /(?:^|\s)(?:\S*\/)?[\w.-]*(?:validate|doctor)[\w.-]*\.sh\b/,
49 name: m => m[0].trim().replace(/^.*\//, ''),
50 },
51 { kind: 'all', pattern: /(?:^|\s)(?:\S*\/)?checks?\.sh\b/, name: m => m[0].trim().replace(/^.*\//, '') },
52]
53
54/** A piece of a shell command and the operator after it: `&&`, `||`, `;`, a newline, `|`, or '' at the end. */
55type Segment = { text: string; next: string }
56
57/**
58 * The pieces of a shell command that run one after another or in a pipe, with
59 * heredoc bodies removed and continued lines joined. A command substitution
60 * stays inside its piece, so `-xctestrun $(ls *.xctestrun | head -1)` is not a
61 * pipe. Quoted strings and substitutions are blanked, so `git commit -m "fix tsc"`
62 * runs no tsc, unless `keepQuotes`, for reading the paths a command names.
63 */
64function segments(command: string, keepQuotes = false): Segment[] {
65 const quoted: string[] = []
66 const substituted: string[] = []
67 let masked = command
68 .replace(/<<-?\s*(["']?)(\w+)\1[^\n]*\n[\s\S]*?\n\2\b/g, ' ')
69 .replace(/(["'])(?:\\.|(?!\1)[\s\S])*\1/g, q => `\0${quoted.push(q) - 1}\0`)
70 .replace(/\\\n/g, ' ')
71 // Innermost first, so a substitution inside another one is masked too.
72 for (let before = ''; before !== masked;) {
73 before = masked
74 masked = masked.replace(/\$\([^()]*\)|`[^`]*`/g, s => `\x01${substituted.push(s) - 1}\x01`)
75 }
76 const unmask = (text: string): string => {
77 let out = text
78 while (/\x01\d+\x01/.test(out))
79 out = out.replace(/\x01(\d+)\x01/g, (_, n: string) => (keepQuotes ? (substituted[Number(n)] ?? '') : '""'))
80
81 return out.replace(/\0(\d+)\0/g, (_, n: string) => (keepQuotes ? (quoted[Number(n)] ?? '') : '""'))
82 }
83 const parts = masked.split(/(&&|\|\||;|\n|\|)/)
84 const found: Segment[] = []
85 for (let i = 0; i < parts.length; i += 2) {
86 const text = unmask(parts[i] ?? '').trim()
87 if (text) found.push({ text, next: parts[i + 1] ?? '' })
88 }
89
90 return found
91}
92
93/** A segment without the shell syntax before its command, such as `if !` in `if ! pgrep -x xcodebuild` or `do`. */
94function commandOf(segment: string): string {
95 return segment.replace(/^(?:[({]\s*|!\s+|(?:if|then|elif|else|do|while|until|time)\s+)+/, '')
96}
97
98/**
99 * Commands that only show text, move around or look for a process; a check
100 * named in one of them is not run, as in `pgrep -x xcodebuild`.
101 */
102const NOT_A_CHECK =
103 /^(?:echo|printf|cat|grep|rg|tail|head|less|ls|cd|git|gh|brew|man|which|type|open|pgrep|pkill|killall|ps|lsof)\b/
104
105/** The check one segment runs, the word that starts it, such as `npm` or `tsc`, and how many words after it the match used, such as `run lint` after `npm`. */
106function checkIn(segment: string): { call: CheckCall; runner: string; consumed: number } | null {
107 const command = commandOf(segment)
108 if (NOT_A_CHECK.test(command)) return null
109 for (const c of CHECKS) {
110 const m = command.match(c.pattern)
111 if (m) {
112 const [runner = '', ...used] = m[0].trim().split(/\s+/)
113
114 return { call: { name: c.name(m), kind: c.kind }, runner, consumed: used.length }
115 }
116 }
117
118 return null
119}
120
121/** The checks a command runs, in order, at most one per segment. */
122export function checksIn(command: string): CheckCall[] {
123 const found: CheckCall[] = []
124 for (const segment of segments(command)) {
125 const call = checkIn(segment.text)?.call
126 if (call && !found.some(f => f.name === call.name)) found.push(call)
127 }
128
129 return found
130}
131
132/** Options that name the folder a check runs in, as in `npm --prefix app test`. */
133const FOLDER_OPTIONS = new Set(['--prefix', '-C', '--cwd', '--dir', '--directory', '--package-path', '--manifest-path'])
134/** tsc's `-p` names its project; other runners use `-p` for a package or plugin. */
135const TSC_FOLDER_OPTIONS = new Set([...FOLDER_OPTIONS, '-p', '--project'])
136
137/** A segment's words, with their quotes removed. */
138function wordsOf(segment: string): string[] {
139 return (segment.match(/(?:"(?:\\.|[^"\\])*"|'[^']*'|[^\s"']+)+/g) ?? []).map(w =>
140 w.replace(/"((?:\\.|[^"\\])*)"|'([^']*)'/g, (_, double?: string, single?: string) => double ?? single ?? ''),
141 )
142}
143
144/** `path` resolved against the folder `from`, with `.` and `..` worked out. */
145export function resolvePath(from: string, path: string): string {
146 const parts: string[] = []
147 for (const part of `${path.startsWith('/') ? '' : from}/${path}`.split('/')) {
148 if (part === '..') parts.pop()
149 else if (part && part !== '.') parts.push(part)
150 }
151
152 return `/${parts.join('/')}`
153}
154
155/** The target of a run of the whole suite. */
156export const NO_TARGET: Target = { paths: [], filters: [] }
157
158/** Flags that change how a run reports or how fast it goes, not what it runs. */
159const OUTPUT_FLAGS = new Set([
160 '--reporter',
161 '--verbose',
162 '-v',
163 '--silent',
164 '--quiet',
165 '-q',
166 '--color',
167 '--no-color',
168 '--ci',
169 '--bail',
170 '-x',
171 '--coverage',
172 '--pretty',
173 '--noEmit',
174 '--passWithNoTests',
175 '--tb',
176 '--maxWorkers',
177 '-j',
178 '--run',
179 '--no-coverage',
180 '--runInBand',
181 '--watch',
182 '--watchAll',
183 '--no-watch',
184 '--forceExit',
185 '--detectOpenHandles',
186 '--no-cache',
187 '--format',
188 '--timeout',
189 '--testTimeout',
190 '-s',
191 '-vv',
192])
193/** The output flags that take their value as the next word. */
194const VALUE_FLAGS = new Set(['--reporter', '--tb', '--maxWorkers', '-j', '--format', '--timeout', '--testTimeout'])
195/** Flags that pick tests by name or by path pattern. */
196const FILTER_FLAGS = new Set(['-t', '--testNamePattern', '--grep', '-g', '-k', '--filter', '-run', '--testPathPattern'])
197/** A word a runner takes as its command, not as a target, as in `vitest run`. */
198const COMMAND_WORDS: Record<string, string[]> = {
199 vitest: ['run'],
200 'golangci-lint': ['run'],
201 ruff: ['check'],
202 biome: ['check', 'lint', 'ci'],
203}
204const REDIRECT = /^(?:\d*|&)[<>]|^&$/
205const isPath = (word: string) => word === '.' || word.includes('/') || /\.[A-Za-z0-9]+$/.test(word)
206
207/**
208 * The part of a suite a run covered, read from the words after its runner
209 * and script. A file or folder narrows it, as does a test-name filter. A flag
210 * that only changes output or speed does not. Any other argument counts as a
211 * filter, so a run read wrongly never clears a failure. `resolve` gives a
212 * path's absolute form, null when it cannot; `folder` is the run's folder.
213 */
214function readTarget(
215 words: string[],
216 runner: string,
217 folder: string | null,
218 resolve: (word: string) => string | null,
219): Target {
220 const paths: string[] = []
221 const filters: string[] = []
222 const add = (list: string[], item: string) => void (list.includes(item) || list.push(item))
223 const rest = words[0] === '--' ? words.slice(1) : words
224 for (let i = 0; i < rest.length; i++) {
225 const word = rest[i] ?? ''
226 if (REDIRECT.test(word)) break
227 const isFlag = word.startsWith('-')
228 const [flag = '', inline] = isFlag && word.includes('=') ? word.split(/=(.*)/s) : [word, undefined]
229 if (OUTPUT_FLAGS.has(flag)) {
230 if (inline === undefined && VALUE_FLAGS.has(flag)) i++
231 } else if (FILTER_FLAGS.has(flag)) {
232 const value = inline ?? rest[++i]
233 add(filters, value === undefined ? flag : `${flag}=${value}`)
234 } else if (COMMAND_WORDS[runner]?.includes(word)) {
235 continue
236 } else if (!isFlag && isPath(word)) {
237 // Go's `./pkg/...` means the folder and everything below it.
238 const path = folder === null ? null : resolve(word.replace(/\/\.\.\.$/, '') || '.')
239 if (path === null) add(filters, word)
240 else if (path !== folder) add(paths, path)
241 } else {
242 add(filters, word)
243 }
244 }
245
246 return { paths, filters }
247}
248
249/** The target of the check run a saved command, run in `folder`, describes; the whole suite when it reads as none. */
250export function targetOfCommand(command: string, folder: string, home: string): Target {
251 return checkRun(command, folder, home)?.target ?? NO_TARGET
252}
253
254/** A check a command runs: the check, the part of its suite the run covered, the command as written, and the folder it ran in. */
255export type CheckRun = {
256 call: CheckCall
257 target: Target
258 command: string
259 folder: string | null
260}
261
262/**
263 * The first check a command runs, with its folder, absolute: where the shell
264 * started, or the folder the check names, as in `claude plugin test <folder>`
265 * or `npm --prefix <folder> test`. The folder is null when the command hides
266 * it, as with a variable other than HOME or `$(...)`.
267 */
268export function checkRun(command: string, cwd: string, home: string): CheckRun | null {
269 const dir = resolvePath('/', cwd)
270 const expand = (word: string): string | null => {
271 if (/\$\(|`/.test(word)) return null
272 const out = word.replace(/^~(?=\/|$)/, home).replace(/\$\{?HOME\}?/g, home)
273
274 return out.includes('$') ? null : out
275 }
276 const blank = segments(command)
277 const kept = segments(command, true)
278 for (const [i, { text: segment }] of kept.entries()) {
279 const found = checkIn(blank[i]?.text ?? '')
280 if (!found) continue
281 const { call, runner, consumed } = found
282 const words = wordsOf(commandOf(segment))
283 const after = words.slice(words.findIndex(w => w === runner || w.endsWith(`/${runner}`)) + 1)
284 const options = runner === 'tsc' ? TSC_FOLDER_OPTIONS : FOLDER_OPTIONS
285 // The words that name the folder are not target words.
286 const folderWords = new Set<number>()
287 let named: string | undefined
288 if (runner === 'claude') {
289 // `claude plugin test <folder>`: the first word after `plugin test` that is not an option or a redirect.
290 const j = after.findIndex((w, k) => k >= 2 && !/^-|^\d*[<>&]/.test(w))
291 if (j >= 0) {
292 named = after[j]
293 folderWords.add(j)
294 }
295 }
296 after.forEach((w, j) => {
297 const isInline = w.startsWith('--') && w.includes('=')
298 const [option = '', value] = isInline ? w.split(/=(.*)/s) : [w, after[j + 1]]
299 if (options.has(option) && value !== undefined) {
300 named = value
301 folderWords.add(j)
302 if (!isInline) folderWords.add(j + 1)
303 }
304 })
305 // A project file, such as tsconfig.build.json or Cargo.toml, names its folder.
306 if (named && /\.(?:json|toml)$/.test(named)) named = named.replace(/\/?[^/]*$/, '') || '.'
307 let folder: string | null = dir
308 if (named !== undefined) {
309 const path = expand(named)
310 folder = path === null ? null : resolvePath(dir, path)
311 }
312 const resolve = (word: string): string | null => {
313 const path = expand(word)
314
315 return path === null || folder === null ? null : resolvePath(folder, path)
316 }
317 const target = readTarget(
318 after.filter((_, j) => j >= consumed && !folderWords.has(j)),
319 runner,
320 folder,
321 resolve,
322 )
323
324 return { call, target, command: segment, folder }
325 }
326
327 return null
328}
329
330/** Whether a folder is a temporary one, where a throwaway copy of a project lives. */
331export function isTemporary(folder: string): boolean {
332 return /^(?:\/private)?\/(?:tmp|var\/tmp|var\/folders)(?:\/|$)/.test(folder)
333}
334
335const SUMMARY = [
336 /^.*\b\d+\s+(?:pass(?:ed|ing)?|fail(?:ed|ing)?)\b.*$/im,
337 /^.*\bRan \d+ tests?\b.*$/im,
338 /^.*\bTests?:\s+.*$/im,
339 /^.*\bFound \d+ (?:errors?|problems?)\b.*$/im,
340 /^.*\b\d+ (?:errors?|warnings?|problems?)\b.*$/im,
341]
342
343/** The marks a test runner puts before a failing test. node's "✖ failing tests:" heading names none. */
344const FAIL_MARK = String.raw`\(fail\)|FAIL(?:ED)?\b|✗|✖(?! failing tests:)|×|✘`
345
346/** A line that names what failed: a failing test, a compiler error, or an error message. */
347const FAILURE_LINE = new RegExp(
348 String.raw`^\s*(?:${FAIL_MARK})|\berror\s+TS\d+\b|^\s*(?:[A-Z]\w*Error|error)(?:\[\w+\])?:`,
349)
350
351/** Up to three lines of a check's output that name what failed, each cut to 160 characters. */
352export function failureLines(output: string): string[] {
353 return output
354 .split('\n')
355 .filter(line => FAILURE_LINE.test(line))
356 .map(line => line.trim().slice(0, 160))
357 .filter((line, i, all) => all.indexOf(line) === i)
358 .slice(0, 3)
359}
360
361/** Pass and fail counts on lines of their own: bun's " 24 pass", node's "ℹ pass 24", TAP's "# pass 24". */
362function counts(output: string): { pass: number; fail: number } | null {
363 const count = (word: string) => {
364 const m = output.match(new RegExp(`^\\s*(\\d+)\\s+${word}\\s*$|^[ℹ#]\\s*${word}\\s+(\\d+)\\s*$`, 'm'))
365
366 return m ? Number(m[1] ?? m[2]) : null
367 }
368 const pass = count('pass')
369 const fail = count('fail')
370
371 return pass === null || fail === null ? null : { pass, fail }
372}
373
374/** How a check ended, from the engine's exit status, with a summary read from its output. */
375export function readResult(output: string, isError: boolean): { result: Check['result']; summary: string } {
376 const tally = counts(output)
377 // failCount reads this back for the pane's "2 of 176 failed".
378 const summary = tally ? `${tally.pass} pass, ${tally.fail} fail` : summaryLine(output)
379
380 return { result: isError ? 'fail' : 'pass', summary: summary || (isError ? 'exited with an error' : '') }
381}
382
383function summaryLine(output: string): string {
384 for (const p of SUMMARY) {
385 const m = output.match(p)
386 if (m) return m[0].trim().slice(0, 120)
387 }
388 const firstError = output.split('\n').find(line => /\berror\b/i.test(line))
389
390 return (firstError ?? '').trim().slice(0, 120)
391}
392
393/** A check's result as one mark: ✓ passed, ✗ failed. */
394function checkMark(check: Check): string {
395 return check.result === 'pass' ? '✓' : '✗'
396}
397
398/** What follows a check's name: ", 24 pass", then ", before the last edit" when the files changed after it ran. */
399function checkDetail(check: Check): string {
400 return `${check.summary ? `, ${check.summary}` : ''}${check.isStale ? ', before the last edit' : ''}`
401}
402
403/** What a check's target adds to its name: its paths relative to the check's folder (`root` for the session's own), then its filters. */
404function targetText(check: Check, root: string): string {
405 const base = check.folder ?? root
406 const paths = check.target.paths.map(p => (p.startsWith(`${base}/`) ? p.slice(base.length + 1) : p))
407
408 return [...paths, ...check.target.filters].join(' ')
409}
410
411/** "npm test", "vitest a.test.ts" for a narrower run, and "npm test in app" for one that ran outside the session's folder `root`. */
412export function checkName(check: Check, root: string): string {
413 const target = targetText(check, root)
414 const name = target ? `${check.name} ${target}` : check.name
415
416 return check.folder ? `${name} in ${check.folder.replace(/^.*\//, '')}` : name
417}
418
419/** "✓ npm test, 24 pass, before the last edit". */
420export function checkLine(check: Check, root: string): string {
421 return `${checkMark(check)} ${checkName(check, root)}${checkDetail(check)}`
422}
423
424/** The reply's prose, as its sentences, without code blocks and inline code. */
425function sentences(reply: string): string[] {
426 return reply
427 .replace(/```[\s\S]*?```/g, ' ')
428 .replace(/`[^`\n]*`/g, ' ')
429 .split(/\n+|(?<=[.!?])\s+/)
430 .map(s =>
431 s
432 .replace(/^[\s>#*-]*(?:\d+[.)]\s*)?/, '')
433 .replace(/[*_]+/g, '')
434 .trim(),
435 )
436 .filter(Boolean)
437}
438
439const CLAIMS: { kind: CheckKind; pattern: RegExp }[] = [
440 {
441 kind: 'tests',
442 pattern:
443 /\b(?:all\s+)?(?:\d+\s+)?(?:the\s+|unit\s+)?tests?\s+(?:now\s+|all\s+|still\s+)*(?:pass(?:es|ed|ing)?|are passing|are green|succeed(?:s|ed)?)\b/i,
444 },
445 {
446 kind: 'types',
447 pattern:
448 /\b(?:tsc|type[- ]?check(?:s|ing)?|types?)\s+(?:is\s+|are\s+|now\s+)*(?:pass(?:es|ed)?|clean|green|succeed(?:s|ed)?)\b|\bno type errors\b/i,
449 },
450 { kind: 'lint', pattern: /\blint(?:ing|er)?\s+(?:is\s+|now\s+)*(?:pass(?:es|ed)?|clean|green)\b/i },
451 {
452 kind: 'build',
453 pattern: /\bbuild\s+(?:is\s+|now\s+)*(?:pass(?:es|ed)?|succeed(?:s|ed)?|green|works)\b|\bbuilds cleanly\b/i,
454 },
455]
456
457/** A sentence of Claude's reply that claims a success, and what the check results say against it. */
458export type Contradiction = { claim: string; problem: string }
459
460/** Words that make a claim conditional, negative or about the future. */
461const HEDGE = /\b(?:not|fail\w*|if|should|would|will|until|unless|once|might|may|expect\w*|untested)\b|n't\b/i
462
463/** Text in double quotes, curly double quotes or backticks: a phrase the reply mentions, not one it says. */
464const QUOTED = /"[^"]*"|“[^”]*”|`[^`]*`/g
465
466/**
467 * The successes the reply claims, in sentence order and then `CLAIMS` order.
468 * A hedged sentence claims nothing, and neither does a check phrase in quotes.
469 * Single quotes stay, since they are also apostrophes.
470 */
471export function claimsIn(reply: string): { sentence: string; kind: CheckKind }[] {
472 return sentences(reply).flatMap(sentence => {
473 const said = sentence.replace(QUOTED, ' ')
474
475 return HEDGE.test(said) ? [] : CLAIMS.filter(c => c.pattern.test(said)).map(c => ({ sentence, kind: c.kind }))
476 })
477}
478
479/** The output's summary line, unless it repeats a failure line, as tsc's first error does. */
480export function distinctSummary(check: Check): string | null {
481 const isRepeated = check.failures.some(f => f.includes(check.summary) || check.summary.includes(f))
482
483 return check.summary && !isRepeated ? check.summary : null
484}
485
486/** What a failed check's output says: the lines that name the failure, then the summary unless it repeats one. */
487export function outputLines(check: Check): string[] {
488 const summary = distinctSummary(check)
489
490 return [...check.failures, ...(summary ? [summary] : [])]
491}
492
493/** A path with an extension, then the line it names, if any: `src/a.ts(12,5)`, `/src/a.swift:30:27`, `src/a.test.ts`. */
494const LOCATION = /((?:[\w@.~-]*\/)*[\w@~-][\w@.~-]*\.[A-Za-z]\w{0,5})(?:\((\d+),\d+\)|:(\d+)(?::\d+)?)?(?![\w/])/g
495
496/** A failure mark, or vitest's ❯, at the start of a failure line. */
497const LEADING_MARK = new RegExp(String.raw`^\s*(?:${FAIL_MARK}|❯)\s*`)
498
499/**
500 * What a failed check's failure lines say, one entry per failure: the file a
501 * line names, as its base name and line, and its text without marker,
502 * location and timing. A line whose text ends with an earlier one's, as
503 * vitest's FAIL line repeats its × line, only adds its file to that entry.
504 */
505export function failureList(check: Check): { file: string | null; text: string }[] {
506 const list: { file: string | null; text: string }[] = []
507 for (const failure of check.failures) {
508 // A token counts as a file when it names a line or a folder, so a word like "foo.bar" does not.
509 const found = [...failure.matchAll(LOCATION)].find(m => m[2] ?? m[3] ?? m[1]?.includes('/'))
510 const line = found?.[2] ?? found?.[3]
511 const file = found?.[1] ? `${found[1].replace(/^.*\//, '')}${line ? `:${line}` : ''}` : null
512 const text = failure
513 .replace(found?.[0] ?? '', '')
514 .replace(LEADING_MARK, '')
515 .replace(/^\|[\w-]+\|\s*/, '')
516 .replace(/^[\s:>-]*(?:error(?:\s+TS\d+|\[\w+\])?:\s*)?/i, '')
517 .replace(/\s*[[(]?\d+(?:\.\d+)?\s?m?s[\])]?$/, '')
518 .trim()
519 const earlier = list.find(f => text.endsWith(f.text))
520 if (earlier) earlier.file ??= file
521 else if (text) list.push({ file, text })
522 }
523
524 return list
525}
526
527/** How many of a check's tests failed, read from its "N pass, M fail" summary; null when the output gave no counts. */
528export function failCount(check: Check): { fail: number; total: number } | null {
529 const m = check.summary.match(/^(\d+) pass, (\d+) fail$/)
530
531 return m ? { fail: Number(m[2]), total: Number(m[1]) + Number(m[2]) } : null
532}
533
534/** What `a: Fix` on a failing check sends Claude. */
535export function fixMessage(check: Check, root: string): string {
536 const output = outputLines(check)
537 const target = targetText(check, root)
538
539 return [
540 `${check.name}${target ? ` ${target}` : ''} failed when you last ran it${check.folder ? ` in ${check.folder}` : ''}.`,
541 ...(check.command ? [`Command: ${check.command}`] : []),
542 ...(output.length > 0 ? ['Output:', ...output] : []),
543 'Find the cause, fix it, and run it again to verify.',
544 ].join('\n')
545}
546
547/** What sends Claude back to correct a claim before its turn ends. */
548export function claimMessage(c: Contradiction): string {
549 return `inbox: your reply says "${c.claim}", but ${c.problem}. Before you end your turn, run the check and report what it shows, or say plainly that the latest change is untested.`
550}
551hooks/check-tracking.ts 266 lines1// Check tracking: each check's latest result, how a result changes (recorded,
2// stale, left failing, dismissed, fix sent), and what the band, the pane and
3// the Stop hook read from the results. Pure functions; the I/O stays in register.tsx.
4
5import type { Check, CheckKind, Checks, Target } from '../types'
6import { checkName, claimsIn, isTemporary, targetOfCommand } from './checks'
7import type { Contradiction } from './checks'
8
9export const NO_CHECKS: Checks = { results: [], repos: [] }
10
11/** One check run that ended, with its folder (absolute) and repo already resolved. */
12type Run = {
13 name: string
14 kind: CheckKind
15 folder: string
16 result: Check['result']
17 summary: string
18 command: string
19 failures: string[]
20 repo: string | null
21 target: Target
22}
23
24/** The one identity rule: a check is known by its name, the folder it ran in and, when it narrows anything, its target. */
25export function checkKey(check: Check): string {
26 const { paths, filters } = check.target
27 const key = `${check.folder ?? '.'}:${check.name}`
28
29 return paths.length === 0 && filters.length === 0
30 ? key
31 : `${key}:${JSON.stringify([[...paths].sort(), [...filters].sort()])}`
32}
33
34/** Whether the folder `path` is, or contains, `inner`. */
35export const contains = (path: string, inner: string) => inner === path || inner.startsWith(`${path}/`)
36
37/**
38 * Whether `run` covers `result`: it is the same check in the same folder, its
39 * filters are none or the same ones, and its paths are none or contain every
40 * path the result has (which must have some).
41 */
42function covers(run: Check, result: Check): boolean {
43 if (run.name !== result.name || run.folder !== result.folder) return false
44 const a = run.target
45 const b = result.target
46 const isSameFilters = a.filters.length === b.filters.length && a.filters.every(f => b.filters.includes(f))
47 const hasPaths = b.paths.length > 0 && b.paths.every(q => a.paths.some(p => contains(p, q)))
48
49 return (a.filters.length === 0 || isSameFilters) && (a.paths.length === 0 || hasPaths)
50}
51
52/**
53 * Keeps each check's latest results. A pass or fail replaces every result it
54 * covers; a check script that passes, such as check.sh, replaces every earlier
55 * result in its folder.
56 */
57function recordCheck(results: Check[], check: Check): Check[] {
58 const isAll = check.kind === 'all' && check.result === 'pass'
59
60 return [...results.filter(c => !(covers(check, c) || (isAll && c.folder === check.folder))), check]
61}
62
63/**
64 * Records the runs that ended at `at`. A run in a temporary folder says
65 * nothing about the session's work, unless the session itself is in one.
66 * The session's own folder, `root`, is saved as null.
67 */
68export function recorded(checks: Checks, runs: Run[], at: number, root: string): Checks {
69 return {
70 ...checks,
71 results: runs
72 .filter(r => !(isTemporary(r.folder) && !isTemporary(root)))
73 .reduce(
74 (all, r) =>
75 recordCheck(all, {
76 name: r.name,
77 kind: r.kind,
78 target: r.target,
79 folder: r.folder === root ? null : r.folder,
80 result: r.result,
81 summary: r.summary,
82 ranAt: at,
83 command: r.command,
84 failures: r.failures,
85 repo: r.repo,
86 isStale: false,
87 isLeftFailing: false,
88 isDismissed: false,
89 isSentBack: false,
90 fixSentAt: null,
91 }),
92 checks.results,
93 ),
94 }
95}
96
97/** Whether a result of this kind reads the file at `path`: Markdown matters only to a lint, a validation or a check script. */
98function readsFile(kind: CheckKind, path: string): boolean {
99 return ['lint', 'validate', 'all'].includes(kind) || !/\.md$/i.test(path)
100}
101
102/**
103 * The files in `repo` changed at `paths`, relative to `repo`; null means git
104 * could not tell, which makes every result there stale. Otherwise a result is
105 * stale when a file it reads changed in its folder, `root` for the session's own.
106 */
107export function changed(checks: Checks, repo: string, paths: string[] | null, root: string): Checks {
108 if (paths !== null && paths.length === 0) return checks
109 const isStaleBy = (c: Check) => {
110 if (paths === null) return true
111 const folder = c.folder ?? root
112
113 return paths.some(p => {
114 const path = `${repo}/${p}`
115
116 return readsFile(c.kind, p) && (path === folder || path.startsWith(`${folder}/`))
117 })
118 }
119
120 return {
121 ...checks,
122 results: checks.results.map(c => (c.repo === repo && isStaleBy(c) ? { ...c, isStale: true } : c)),
123 }
124}
125
126/** The session started in this repo, or Claude edited a file in it. */
127export function addRepo(checks: Checks, repo: string): Checks {
128 return checks.repos.includes(repo) ? checks : { ...checks, repos: [...checks.repos, repo] }
129}
130
131/** Drops the results from folders that are gone. */
132export function pruned(checks: Checks, gone: ReadonlySet<string>): Checks {
133 return { ...checks, results: checks.results.filter(c => !c.folder || !gone.has(c.folder)) }
134}
135
136/** Claude's turn ended: every failing result is now left failing. */
137export function turnEnded(checks: Checks): Checks {
138 return { ...checks, results: checks.results.map(c => (c.result === 'fail' ? { ...c, isLeftFailing: true } : c)) }
139}
140
141/** The person dismissed the check, whichever run it came from. */
142export function dismissed(checks: Checks, check: Check): Checks {
143 const key = checkKey(check)
144
145 return { ...checks, results: checks.results.map(c => (checkKey(c) === key ? { ...c, isDismissed: true } : c)) }
146}
147
148/** The person asked Claude to fix the check's run at `at`. Matched by run, so a rerun recorded meanwhile keeps its own state. */
149export function fixSent(checks: Checks, check: Check, at: number): Checks {
150 const key = checkKey(check)
151
152 return {
153 ...checks,
154 results: checks.results.map(c => (checkKey(c) === key && c.ranAt === check.ranAt ? { ...c, fixSentAt: at } : c)),
155 }
156}
157
158/**
159 * Results saved before a field existed get its default. A result from before
160 * each one kept its repo never goes stale, and one from before targets takes
161 * its target from its saved command. `root` is the session's folder.
162 */
163export function upgradeChecks(saved: Checks, root: string, home: string): Checks {
164 return {
165 // State saved before the session's repos were kept counted every result's repo.
166 repos: saved.repos ?? [...new Set(saved.results.flatMap(r => (r.repo ? [r.repo] : [])))],
167 // Results read as unknown, before run_check, say nothing about the work.
168 results: saved.results
169 .filter(r => (r.result as string) !== 'unknown')
170 .map(r => ({
171 ...r,
172 repo: r.repo ?? null,
173 isStale: r.isStale ?? false,
174 command: r.command ?? '',
175 target: r.target ?? targetOfCommand(r.command ?? '', r.folder ?? root, home),
176 failures: r.failures ?? [],
177 isLeftFailing: r.isLeftFailing ?? false,
178 isDismissed: r.isDismissed ?? false,
179 isSentBack: r.isSentBack ?? false,
180 fixSentAt: r.fixSentAt ?? null,
181 })),
182 }
183}
184
185/**
186 * Whether a result is the session's: it ran in one of the session's repos, or,
187 * outside git, in the session's folder `root` or below it.
188 */
189function counts(checks: Checks, c: Check, root: string): boolean {
190 if (c.repo !== null) return checks.repos.includes(c.repo)
191
192 return c.folder === null || c.folder === root || c.folder.startsWith(`${root}/`)
193}
194
195/** What the Needs you tab lists: the session's results left failing and not dismissed, and how many of them wait on the person (no fix sent). */
196export function needsYou(checks: Checks, root: string): { rows: Check[]; count: number } {
197 const rows = checks.results.filter(c => counts(checks, c, root) && c.isLeftFailing && !c.isDismissed)
198
199 return { rows, count: rows.filter(c => c.fixSentAt === null).length }
200}
201
202/** What the band shows from the session's results: the failing ones not dismissed, one line each, and the rest on the dim summary line. No result shows twice. */
203export function bandLines(checks: Checks, root: string): { failing: Check[]; summary: Check[] } {
204 const own = checks.results.filter(c => counts(checks, c, root))
205
206 return {
207 failing: own.filter(c => c.result === 'fail' && !c.isDismissed),
208 summary: own.filter(c => c.result !== 'fail'),
209 }
210}
211
212/** A contradicted claim, with the result it is about. */
213type Contradicted = Contradiction & { check: Check }
214
215/**
216 * Each result that contradicts a success the reply claims, claim by claim:
217 * every failure of that kind, or of a script that runs them all, latest
218 * first, then the latest result of that kind if the files changed after it
219 * ran. A later pass removes the failures it covers, so a failure still
220 * recorded is one no later run cleared. A claim with no check of its kind is
221 * left alone, since the check may have run in a way the mod cannot see.
222 */
223function* contradictions(reply: string, checks: Checks, root: string): Generator<Contradicted> {
224 for (const { sentence, kind } of claimsIn(reply)) {
225 const ofKind = checks.results.filter(c => c.kind === kind || c.kind === 'all').sort((a, b) => b.ranAt - a.ranAt)
226 for (const failed of ofKind.filter(c => c.result === 'fail'))
227 yield {
228 claim: sentence,
229 problem: `${checkName(failed, root)} failed when it last ran${failed.summary ? ` (${failed.summary})` : ''}`,
230 check: failed,
231 }
232 const [latest] = ofKind
233 if (latest && latest.result !== 'fail' && latest.isStale)
234 yield { claim: sentence, problem: `the files changed after ${checkName(latest, root)} last ran`, check: latest }
235 }
236}
237
238/**
239 * The Stop hook's check: the first claim in the reply that the session's
240 * results contradict and that has not yet sent Claude back, with that result
241 * marked so it sends Claude back once.
242 */
243export function claimAgainst(
244 checks: Checks,
245 reply: string,
246 root: string,
247): { checks: Checks; claim: Contradiction | null } {
248 const own = { ...checks, results: checks.results.filter(c => counts(checks, c, root)) }
249 for (const { claim, problem, check } of contradictions(reply, own, root)) {
250 if (check.isSentBack) continue
251 const key = checkKey(check)
252
253 return {
254 checks: {
255 ...checks,
256 results: checks.results.map(c =>
257 checkKey(c) === key && c.ranAt === check.ranAt ? { ...c, isSentBack: true } : c,
258 ),
259 },
260 claim: { claim, problem },
261 }
262 }
263
264 return { checks, claim: null }
265}
266hooks/demo.ts 459 lines1// Sample entries for every section of the band and the /inbox pane, shown by
2// `/inbox demo` for work on the layout. Nothing here reaches the session's
3// real inbox, its store, or Claude.
4
5import type { Checks, LastAction, Ledger, Presence, PrFixSent, PrViews, Settled, Stop } from '../types'
6
7export type View = {
8 ledger: Ledger
9 stop: Stop | null
10 checks: Checks
11 settled: Settled[]
12 prViews: PrViews
13 prFixesSent: Record<string, PrFixSent>
14 lastActions: Record<string, LastAction>
15 /** The turn counts a task handed to Claude folds by. */
16 turns: Pick<Presence, 'turnsStarted' | 'turnsApplied'>
17}
18
19const MIN = 60_000
20const REPO = 'https://github.com/petekp/inbox'
21
22export function demoView(now: number): View {
23 return {
24 ledger: {
25 card: {
26 goal: 'Give the inbox pane keyboard shortcuts and a cleaner tab bar',
27 done: ['Tabs switch with 1, 2 and 3', 'Answer keys moved to letters', 'PR #31 opened'],
28 now: 'Waiting on where the Keys list goes and a gh sign-in',
29 running: ['live test session: tmux attach -t inbox-live'],
30 updatedAt: now - 2 * MIN,
31 },
32 items: [
33 {
34 id: 'd11',
35 kind: 'question',
36 label: '1',
37 ask: 'Show the Keys list in a footer, or under the tab bar?',
38 options: ['Footer', 'Under the tab bar'],
39 rec: 'Footer',
40 helps: [],
41 turn: 14,
42 at: now - 12 * MIN,
43 },
44 {
45 id: 'd12',
46 kind: 'question',
47 label: '2',
48 ask: 'Keep findings open until you confirm a fix?',
49 options: ['Keep it open', 'Close it'],
50 rec: 'Keep it open',
51 helps: [],
52 turn: 14,
53 at: now - 12 * MIN,
54 },
55 {
56 id: 'd13',
57 kind: 'question',
58 label: '3',
59 ask: 'Expire questions after 20 prompts, not 12?',
60 options: ['Yes', 'No'],
61 rec: 'No',
62 helps: [],
63 turn: 14,
64 at: now - 12 * MIN,
65 },
66 {
67 id: 'd14',
68 kind: 'task',
69 label: null,
70 ask: 'Sign in to gh for the PRs tab',
71 options: [],
72 rec: null,
73 helps: [{ kind: 'terminal', command: 'gh auth login', name: 'gh auth login' }],
74 turn: 14,
75 at: now - 5 * MIN,
76 },
77 {
78 id: 'd15',
79 kind: 'question',
80 label: null,
81 ask: 'Which theme should the README screenshot use?',
82 options: [],
83 rec: null,
84 helps: [],
85 turn: 11,
86 at: now - 40 * MIN,
87 },
88 {
89 id: 'd16',
90 kind: 'task',
91 label: null,
92 ask: 'Check the pane in the light theme',
93 options: [],
94 rec: null,
95 helps: [{ kind: 'copy', text: '/theme', name: 'theme command' }],
96 turn: 14,
97 at: now - 5 * MIN,
98 },
99 {
100 id: 'd18',
101 kind: 'task',
102 label: null,
103 ask: 'Run the load script and send back its output lines',
104 options: [],
105 rec: null,
106 helps: [{ kind: 'run', command: './scripts/load.sh', name: 'load script' }],
107 turn: 14,
108 at: now - 4 * MIN,
109 },
110 {
111 id: 'd17',
112 kind: 'task',
113 label: null,
114 ask: 'Run /reload-plugins in your other sessions',
115 options: [],
116 rec: null,
117 helps: [],
118 turn: 12,
119 at: now - 25 * MIN,
120 },
121 ],
122 closed: [
123 {
124 id: 'd5',
125 kind: 'question',
126 ask: 'Rename the Waiting tab?',
127 outcome: 'Needs you',
128 how: 'answered',
129 at: now - 90 * MIN,
130 },
131 {
132 id: 'd6',
133 kind: 'question',
134 ask: 'Commit the reload fix and the rename as two commits?',
135 outcome: 'yes',
136 how: 'answered',
137 at: now - 60 * MIN,
138 },
139 {
140 id: 'd9',
141 kind: 'task',
142 ask: 'Update Claude Code to 2.1.292',
143 outcome: 'done',
144 how: 'done',
145 at: now - 45 * MIN,
146 },
147 {
148 id: 'd7',
149 kind: 'question',
150 ask: 'Keep the darker body behind the section cards?',
151 outcome: 'closed by Claude: no longer applies: the body matches the tab bar',
152 how: 'claude',
153 at: now - 30 * MIN,
154 },
155 {
156 id: 'd8',
157 kind: 'question',
158 ask: 'Add a fourth Session tab?',
159 outcome: 'dismissed',
160 how: 'dismissed',
161 at: now - 20 * MIN,
162 },
163 {
164 id: 'd10',
165 kind: 'question',
166 ask: 'Draw the tabs on the pane’s own background?',
167 outcome: 'Yes, as part of the title bar',
168 how: 'answered',
169 at: now,
170 },
171 ],
172 findings: [
173 {
174 id: 'd20',
175 kind: 'opportunity',
176 title: 'One contrast check could cover all six themes',
177 detail:
178 'A capture of the pane in each theme holds every cell’s colors. A script could flag any text under 4.5:1 and run in check.sh.',
179 path: 'scripts/check.sh',
180 at: now - 3 * 60 * MIN,
181 },
182 {
183 id: 'd21',
184 kind: 'opportunity',
185 title: 'Catch-up could read only the turns it missed',
186 detail:
187 'After a reload, the catch-up call reads the whole conversation. Starting from the last turn the ledger applied would cut its tokens on long sessions.',
188 path: 'hooks/ledger.ts',
189 at: now - 90 * MIN,
190 },
191 {
192 id: 'd22',
193 kind: 'issue',
194 title: 'Hotkeys vanish on selected rows in ANSI',
195 detail:
196 'The engine draws hotkeys in the same blue the selected row uses as its background, so a: and b: disappear there.',
197 path: 'hooks/register.tsx',
198 at: now - 45 * MIN,
199 },
200 {
201 id: 'd23',
202 kind: 'issue',
203 title: 'Selected tab is unreadable in dark ANSI',
204 detail:
205 'The selected tab is drawn on ANSI white with the default light text. Its name needs a dark color there, or an inverse style.',
206 path: 'hooks/register.tsx',
207 at: now - 30 * MIN,
208 },
209 ],
210 prs: ['petekp/inbox#31', 'petekp/inbox#29', 'petekp/inbox#33'],
211 nextId: 24,
212 turn: 14,
213 batchTurn: 14,
214 },
215 stop: null,
216 checks: {
217 repos: [],
218 results: [
219 {
220 name: 'prettier',
221 kind: 'lint',
222 folder: null,
223 target: { paths: [], filters: [] },
224 result: 'pass',
225 summary: '',
226 ranAt: now - 20 * MIN,
227 command: 'npx -y prettier@3.9.9 --check .',
228 failures: [],
229 repo: null,
230 isStale: true,
231 isLeftFailing: false,
232 isDismissed: false,
233 isSentBack: false,
234 fixSentAt: null,
235 },
236 {
237 name: 'claude plugin test',
238 kind: 'tests',
239 folder: null,
240 target: { paths: [], filters: [] },
241 result: 'pass',
242 summary: '64 pass, 0 fail',
243 ranAt: now - 6 * MIN,
244 command: 'claude plugin test .',
245 failures: [],
246 repo: null,
247 isStale: false,
248 isLeftFailing: false,
249 isDismissed: false,
250 isSentBack: false,
251 fixSentAt: null,
252 },
253 {
254 name: 'tsc',
255 kind: 'types',
256 folder: null,
257 target: { paths: [], filters: [] },
258 result: 'fail',
259 summary: 'hooks/register.tsx(2310,7): error TS2322',
260 ranAt: now - 5 * MIN,
261 command: 'npx -y -p typescript tsc --noEmit -p .',
262 failures: [
263 "hooks/register.tsx(2310,7): error TS2322: Type 'string | undefined' is not assignable to type 'string'.",
264 ],
265 repo: null,
266 isStale: false,
267 isLeftFailing: true,
268 isDismissed: false,
269 isSentBack: false,
270 fixSentAt: null,
271 },
272 {
273 name: 'vitest',
274 kind: 'tests',
275 folder: null,
276 target: { paths: [], filters: ['tests/parse.test.ts'] },
277 result: 'fail',
278 summary: '1 failed | 11 passed',
279 ranAt: now - 3 * MIN,
280 command: 'npx vitest run tests/parse.test.ts',
281 failures: ['× parses a dated heading 2ms', 'FAIL tests/parse.test.ts > headings > parses a dated heading'],
282 repo: null,
283 isStale: false,
284 isLeftFailing: true,
285 isDismissed: false,
286 isSentBack: false,
287 fixSentAt: now - 1 * MIN,
288 },
289 ],
290 },
291 settled: [
292 {
293 id: 'd10',
294 ask: 'Draw the tabs on the pane’s own background?',
295 outcome: 'Yes, as part of the title bar',
296 how: 'answered',
297 at: now,
298 kind: 'question',
299 index: 1,
300 },
301 ],
302 prViews: {
303 branchRef: 'petekp/inbox#31',
304 isFetching: false,
305 views: {
306 'petekp/inbox#31': {
307 ref: 'petekp/inbox#31',
308 number: 31,
309 title: 'Switch tabs with 1, 2 and 3, and letter the answers',
310 url: `${REPO}/pull/31`,
311 isDraft: false,
312 state: 'OPEN',
313 base: 'main',
314 mergeable: 'CONFLICTING',
315 reviewDecision: 'CHANGES_REQUESTED',
316 checks: [
317 { name: 'prettier', bucket: 'pass', url: `${REPO}/actions/runs/1` },
318 { name: 'plugin tests', bucket: 'fail', url: `${REPO}/actions/runs/2` },
319 { name: 'lint', bucket: 'fail', url: `${REPO}/actions/runs/3` },
320 { name: 'plugin validate', bucket: 'pending', url: null },
321 ],
322 threads: [
323 {
324 id: 'DT1',
325 author: 'sam',
326 reply: null,
327 isWaiting: true,
328 isOutdated: false,
329 isLinesChanged: false,
330 path: 'hooks/register.tsx',
331 line: 1147,
332 body: 'Why do the answer keys skip d? A question has no Done action.',
333 replies: 0,
334 url: `${REPO}/pull/31#discussion_r1`,
335 at: now - 3 * 60 * MIN,
336 },
337 {
338 id: 'DT2',
339 author: 'sam',
340 reply: {
341 author: 'robin',
342 body: 'Agreed, the footer reads better than a list under the tabs.',
343 url: `${REPO}/pull/31#discussion_r3`,
344 at: now - 5 * 60 * MIN,
345 },
346 isWaiting: true,
347 isOutdated: true,
348 isLinesChanged: false,
349 path: 'README.md',
350 line: 87,
351 body: 'Should the Keys list live in the footer?',
352 replies: 1,
353 url: `${REPO}/pull/31#discussion_r2`,
354 at: now - 26 * 60 * MIN,
355 },
356 {
357 id: 'DT3',
358 author: 'robin',
359 reply: {
360 author: 'you',
361 body: 'Done in the latest push.',
362 url: `${REPO}/pull/31#discussion_r5`,
363 at: now - 2 * 60 * MIN,
364 },
365 isWaiting: false,
366 isOutdated: false,
367 isLinesChanged: false,
368 path: 'hooks/register.tsx',
369 line: 2296,
370 body: 'Can the tab width come from the label alone now?',
371 replies: 1,
372 url: `${REPO}/pull/31#discussion_r4`,
373 at: now - 26 * 60 * MIN,
374 },
375 {
376 id: 'DT4',
377 author: 'review-bot',
378 reply: null,
379 isWaiting: true,
380 isOutdated: true,
381 isLinesChanged: true,
382 path: 'hooks/prs.ts',
383 line: 114,
384 body: 'Outdated threads still count as waiting on the person.',
385 replies: 0,
386 url: `${REPO}/pull/31#discussion_r6`,
387 at: now - 50 * MIN,
388 },
389 ],
390 fetchedAt: now - MIN,
391 error: null,
392 },
393 'petekp/inbox#29': {
394 ref: 'petekp/inbox#29',
395 number: 29,
396 title: 'Count turns to tell when a reload cut off an update',
397 url: `${REPO}/pull/29`,
398 isDraft: false,
399 state: 'OPEN',
400 base: 'main',
401 mergeable: 'MERGEABLE',
402 reviewDecision: 'APPROVED',
403 checks: [
404 { name: 'prettier', bucket: 'pass', url: null },
405 { name: 'plugin tests', bucket: 'pass', url: null },
406 ],
407 threads: [],
408 fetchedAt: now - MIN,
409 error: null,
410 },
411 'petekp/inbox#33': {
412 ref: 'petekp/inbox#33',
413 number: 33,
414 title: 'Draw the tabs on the pane’s own background',
415 url: `${REPO}/pull/33`,
416 isDraft: true,
417 state: 'OPEN',
418 base: 'main',
419 mergeable: 'MERGEABLE',
420 reviewDecision: '',
421 checks: [
422 { name: 'prettier', bucket: 'pass', url: null },
423 { name: 'plugin tests', bucket: 'pending', url: null },
424 ],
425 threads: [],
426 fetchedAt: now - MIN,
427 error: null,
428 },
429 },
430 },
431 prFixesSent: {
432 'petekp/inbox#31 check lint': { at: now - 2 * MIN, url: `${REPO}/actions/runs/3` },
433 },
434 lastActions: {
435 d18: {
436 action: 'help-d18-0',
437 text: 'Run load script',
438 at: now - 1 * MIN,
439 isHandoff: true,
440 turnsStarted: 14,
441 tab: 'needsYou',
442 title: 'Run the load script and send back its output lines',
443 index: 0,
444 },
445 'petekp/inbox#31 thread DT1': {
446 action: 'address-DT1',
447 text: 'Address',
448 at: now - 2 * MIN,
449 isHandoff: true,
450 turnsStarted: 14,
451 tab: 'prs',
452 title: 'hooks/register.tsx:1147',
453 index: 0,
454 },
455 },
456 turns: { turnsStarted: 14, turnsApplied: 14 },
457 }
458}
459hooks/ledger.ts 1038 lines1import type { SessionMessage } from 'claude-code'
2
3import type { Card, Closed, Dialog, Finding, Help, Item, Ledger, Stop } from '../types'
4
5export const EMPTY: Ledger = {
6 card: null,
7 items: [],
8 closed: [],
9 findings: [],
10 prs: [],
11 nextId: 1,
12 turn: 0,
13 batchTurn: 0,
14}
15
16const NL = '\n'
17const MAX_OPEN = 20
18/** Prompts after which an unanswered item is dropped as moot. */
19const STALE_AFTER = 12
20const MAX_CLOSED = 12
21/** The outcome of an item left unanswered until it went stale. */
22const EXPIRED = 'expired, unanswered'
23const MAX_HELPS = 3
24const MAX_FINDINGS = 30
25
26/** The inbox model's protocol names the two kinds decide and do. */
27const MODEL_KIND: Record<Item['kind'], 'decide' | 'do'> = { question: 'decide', task: 'do' }
28
29/** Reads a kind the model wrote or an older version saved; anything but a task is a question. */
30export function readKind(kind: string | null | undefined): Item['kind'] {
31 return kind === 'task' || kind === 'do' ? 'task' : 'question'
32}
33
34/**
35 * A ledger saved by an earlier version of the mod, in the current shape.
36 * Findings were once saved as `notes`; items once had no time, and closed
37 * items no kind or `how`, so an old closed item counts as a question. Closed
38 * items were saved as `decided`, and kinds as `decide` and `do`.
39 */
40export function upgradeLedger(ledger: Ledger): Ledger {
41 const { notes, decided, ...rest } = ledger as Ledger & { notes?: Finding[]; decided?: Ledger['closed'] }
42
43 return {
44 ...rest,
45 findings: [...(rest.findings ?? []), ...(notes ?? [])],
46 items: rest.items.map(i => ({
47 ...i,
48 kind: readKind(i.kind),
49 at: i.at ?? null,
50 rec: recommendedOption(i.options, i.rec),
51 })),
52 closed: (rest.closed ?? decided ?? []).map(d => ({
53 ...d,
54 kind: readKind(d.kind),
55 how: d.how ?? howFromOutcome(d.outcome),
56 })),
57 }
58}
59
60function words(text: string): string[] {
61 return text.toLowerCase().match(/[\p{L}\p{N}]+/gu) ?? []
62}
63
64/**
65 * The option a recommendation names: the one whose words all appear in it,
66 * the longest when several do. "Symlink into a PATH folder" names "Symlink
67 * into PATH". Null when it names none, since a recommendation that is not one
68 * of the answers came from the wrong field or from no recommendation at all.
69 */
70export function recommendedOption(options: string[], rec: string | null): string | null {
71 if (!rec) return null
72 const named = new Set(words(rec))
73 let best: string | null = null
74 for (const option of options) {
75 const w = words(option)
76 if (w.length > words(best ?? '').length && w.every(x => named.has(x))) best = option
77 }
78
79 return best
80}
81
82/** How a closed item saved before `how` existed closed, read from its outcome's wording. */
83function howFromOutcome(outcome: string): Closed['how'] {
84 if (outcome === 'dismissed') return 'dismissed'
85 if (outcome === EXPIRED) return 'expired'
86 if (outcome === 'done' || outcome === 'you ran it') return 'done'
87 if (outcome.startsWith(CLOSED_BY_CLAUDE)) return 'claude'
88
89 return 'update'
90}
91/**
92 * Commands that sign in or ask for a password. They need the person's own
93 * terminal, so their button copies them instead of asking Claude to run them.
94 */
95const NEEDS_PERSON =
96 /\b(login|logout|auth|signin|sign-in|sudo|passwd|ssh-add|ssh-keygen|configure|init --interactive)\b/i
97
98/** The inbox model's instructions in a host. */
99export function systemText(host: Host): string {
100 const band = host.band ? ` They always see the items in <open> in ${host.band}.` : ''
101
102 return `You keep a short ledger for a person who works with a coding agent across many parallel sessions. They glance at your ledger between tasks, or after time away, to see where this session stands. You read one exchange and update the ledger.
103
104Input:
105- <card>: the ledger before this exchange (may be empty)
106- <open>: items still waiting on the person, each with an id
107- <findings>: findings the agent recorded for the person to review later, outside the current task, each with an id
108- <decided>: items the person already settled, and how. Never add one of these again as NEW, even when the reply asks it again.
109- <person>: what the person just sent, and the commands they ran themselves: "$ cmd" for a shell command, with its output, and "/name" for a slash command
110- <activity>: what the agent did this turn (files edited, commands, URLs)
111- <reply>: the agent's final reply
112- <screen>: what the person has on screen besides the conversation.${band}
113- <checks>: the latest result of each test, type check, lint or build the agent ran, read from the commands themselves. "before the last edit" means files changed after it ran. These results override the reply: never write in DONE or NOW that a check passes unless <checks> shows it passing and not before the last edit.
114
115Answer with lines only, each starting with one of these keys. No other text.
116
117GOAL: what this session is for, at most 12 words. Keep the previous goal unless the person clearly changed direction. "-" until the person has asked for something.
118DONE: one finished outcome, at most 8 words. Up to 4 DONE lines, oldest first, keeping the most recent. Outcomes, not activity: "PR #12 opened", not "ran gh".
119NOW: where the work stands at the end of this reply, at most 12 words. Name what it waits on, if anything. "-" when no work has started.
120RUNNING: something still running that the person may open, as "name: URL or port". Dev servers, simulators, background jobs. Omit anything the agent stopped. Zero or more lines.
121CLOSED: <id> | what was decided, at most 8 words. For each item in <open> the person answered in <person> (including "all recommended", "go", "yes to all", numbered answers), or that the reply or <activity> shows is done or no longer applies. A person asking what an item means has not answered it, and a reply explaining it does not close it. When <person> asks to run an item's command and the reply says it ran, that item is done. So is an item whose command the person ran themselves, per <person>, when its output shows it worked. Also one line for each finding in <findings> that the reply or <activity> shows was fixed, or that the person dealt with or set aside.
122NEW: <kind> | <label> | <ask> | <options> | <rec>
123 One line per thing in <reply> that waits on the person and is not already in <open> or <findings>. A finding the agent recorded is not NEW unless the reply asks the person to decide on it now. When the reply restates, rewords or narrows an item in <open>, it is not new: add HELP lines to that item's id instead.
124 kind: "decide" (a choice, approval, or information only the person has, explicitly put to them, without which the agent cannot go on with its task) or "do" (an action only the person can take, without which the agent cannot continue or finish: sign in, run a command needing their password, test on their device, reply to a teammate).
125 label: the reply's own number or id for it ("1", "D3"), or "-".
126 ask: plain words, readable without the reply, at most 12 words, or up to 16 when 12 would lose meaning. Keep the question's meaning and every alternative it names. Replace any term the reply coined with what it means.
127 options: the answers the person can pick, separated by " / ", at most 5 words each; "-" for kind "do". For kind "decide", always at least one, so one press can answer. Use the choices the reply offers, plus the answer it recommends when that is not one of them. When the reply offers none, predict the answers the person would most likely give: the yes and the no for an approval or a yes-or-no question ("Approve / Not yet"); the 1 to 3 likeliest answers to an open question, from the conversation; or, when the conversation suggests no answer, what the person would most likely ask the agent to do instead ("List the choices", "Pick one for me").
128 rec: the option the reply states it recommends for this question ("I'd go with X", "I recommend X"), copied from options, or "-" when it states none for this question.
129 Skip: rhetorical questions; offers to continue ("Want me to start?") when continuing is the obvious default; FYIs; generic "let me know"; invitations to look at, try or check finished work ("Open X to see it", "reload to check") unless the agent waits on the person's verdict before going on; optional suggestions; anything the person already has on screen, per <screen>; questions asking the person to describe what they saw, did or meant, even when the answer would help diagnose a problem ("What happens when you click it?", "Which file did you mean?"). The person answers those by replying.
130HELP: <item> | <kind> | <value> | <name>
131 A step that does part of an item's work in one press, when the reply or activity already spells it out: the file to edit, the text to paste, the command to run, the page to visit. Not background reading. Up to 3 per item, most useful first.
132 item: "new N" for the Nth NEW line in your answer, or an id from <open>.
133 kind and value:
134 open | a file or folder the person needs to open or edit, the path as written
135 copy | text for the person to paste into a file or form, never a command: "block N" for the code block marked [block N] in <reply>, or one line of text
136 run | a shell command the reply asks the person to run or approve: "block N" or the command
137 link | an https URL the person needs to visit
138 name: what a copy, command or link is, at most 3 words ("settings snippet", "removal command", "token page"), or "-".
139 Use only paths, commands, text and URLs that appear in <reply> or <activity>. Never invent one.
140
141Write plainly. No jargon, no filler, no markdown.`
142}
143
144const CATCH_UP =
145 'The ledger below may have missed turns. Close every item in <open> and every finding in <findings> that the conversation shows answered, done, dealt with, or no longer relevant. Add as NEW only what still waits on the user and is not already in <open> or <findings>.'
146
147/**
148 * Asks a fork of the main conversation to bring the ledger up to date at once,
149 * for turns the per-turn update missed.
150 */
151export function catchUpPrompt(host: Host, ledger: Ledger, screen: string): string {
152 return [
153 'Pause the task. Do not use tools. Instead, act as the ledger keeper described below, over this whole conversation.',
154 `Treat the whole conversation as the exchange. ${CATCH_UP}`,
155 ...ledgerBlocks(ledger),
156 `<screen>${NL}${screen}${NL}</screen>`,
157 'Code blocks here carry no [block N] marker, so a copy HELP must be one line of text.',
158 '',
159 systemText(host),
160 ].join(NL)
161}
162
163/**
164 * The catch-up as an ordinary call over the transcript, for a conversation
165 * this process cannot fork yet. systemText goes in the call's system prompt.
166 */
167export function transcriptCatchUpPrompt(ledger: Ledger, screen: string, transcript: string): string {
168 return [
169 `Treat the conversation in <conversation> as the exchange. Its start may be cut. ${CATCH_UP}`,
170 ...ledgerBlocks(ledger),
171 `<conversation>${NL}${numberBlocks(transcript)}${NL}</conversation>`,
172 `<screen>${NL}${screen}${NL}</screen>`,
173 ].join(NL)
174}
175
176/** What the agent did with one tool call, as an activity line; null for a call the ledger has no use for. */
177export function toolActivity(tool: string, input: Record<string, unknown>): string | null {
178 const text = (key: string) => (typeof input[key] === 'string' ? (input[key] as string) : '')
179 switch (tool) {
180 case 'Bash':
181 return `${input.run_in_background ? 'started in background' : 'ran'}: ${text('command').slice(0, 140)}`
182 case 'Edit':
183 case 'Write':
184 return `edited ${text('file_path')}`
185 case 'Skill':
186 return `used skill ${text('skill')}`
187 case 'Agent':
188 return `started agent: ${text('description')}`
189 }
190
191 return tool.startsWith('mcp__') ? `called ${tool.slice(5)}` : null
192}
193
194/** What the agent asked in a question dialog and how it was answered, as an activity line. */
195export function dialogLine(answer: string, isTimedOut: boolean): string {
196 // A dialog that resolved on its own while the person was away holds no answer of theirs.
197 const how = isTimedOut
198 ? 'it timed out while the user was away, so the user did not answer; it went on with'
199 : 'answer'
200
201 return `asked the user in a dialog; ${how}: ${answer.slice(0, 400)}`
202}
203
204const TRANSCRIPT_MAX = 80_000
205const MESSAGE_MAX = 6000
206
207/**
208 * The conversation as the transcript catch-up reads it: the person's messages
209 * and commands, the agent's replies and what it did. Only the most recent
210 * part that fits is kept, since what still waits is near the end. Rows this
211 * mod added beside a prompt, which start with "inbox:", are left out.
212 */
213export function transcriptText(messages: SessionMessage[]): string {
214 const lines: string[] = []
215 for (const m of messages) {
216 const text = m.text.trim()
217 if (m.role === 'user') {
218 if (!text || text.startsWith('inbox:')) continue
219 const row = readCommandRow(text)
220 const line = row ? commandRowLine(row) : clip(text, MESSAGE_MAX)
221 if (line) lines.push(`Person: ${line}`)
222 continue
223 }
224 if (text) lines.push(`Agent: ${clip(text, MESSAGE_MAX)}`)
225 for (const use of m.toolUses) {
226 const timedOut = typeof (use.result as { afkTimeoutMs?: unknown } | undefined)?.afkTimeoutMs === 'number'
227 const line =
228 use.tool === 'AskUserQuestion' ? dialogLine(use.text ?? '', timedOut) : toolActivity(use.tool, use.input)
229 if (line) lines.push(`Agent ${line}`)
230 }
231 }
232 const kept: string[] = []
233 let size = 0
234 for (const line of lines.reverse()) {
235 size += line.length + 1
236 if (size > TRANSCRIPT_MAX) {
237 kept.push('(earlier conversation left out)')
238 break
239 }
240 kept.push(line)
241 }
242
243 return kept.reverse().join(NL)
244}
245
246export type Exchange = {
247 /** What the person typed; null when something else started the turn. */
248 person: string | null
249 /** What started the turn when the person didn't (a task notification, a peer). */
250 trigger: string | null
251 activity: string[]
252 reply: string
253 /** The person-prompt count when the reply was written. */
254 turn: number
255 /** The item button the person pressed to send this turn's prompt, if any. */
256 press: Press | null
257 /** What the person has on screen besides the conversation, from screenText. */
258 screen: string
259 /** The latest result of each check the agent ran, from checkLine. */
260 checks: string[]
261}
262
263/** A prompt the mod sent for an item: an answer, an Explain, or a Run. */
264export type Press = { id: string; action: 'answer' | 'explain' | 'run' }
265
266export type Update = {
267 card: Omit<Card, 'updatedAt'>
268 closed: { id: string; outcome: string }[]
269 added: (Omit<Item, 'id' | 'turn' | 'at' | 'helps'> & { helps: Help[] })[]
270 /** Helps for items already open. */
271 helped: { id: string; help: Help }[]
272}
273
274const FENCE = /```[^\n]*\n([\s\S]*?)```/g
275
276/** The reply's fenced code blocks, in the order buildPrompt numbers them. */
277function codeBlocks(reply: string): string[] {
278 return [...reply.matchAll(FENCE)].map(m => (m[1] ?? '').replace(/\n$/, ''))
279}
280
281function numberBlocks(reply: string): string {
282 let n = 0
283
284 return reply.replace(FENCE, block => `[block ${(n += 1)}]${NL}${block}`)
285}
286
287/**
288 * Checks one HELP line; null when it is unusable. A path, command, URL or line
289 * of text must appear in `source`, the reply and activity, so the model cannot
290 * invent one. A null source (a catch-up over the whole conversation) skips that.
291 */
292function readHelp(
293 kind: string,
294 value: string,
295 name: string | null,
296 blocks: string[],
297 source: string | null,
298): Help | null {
299 const isQuoted = (text: string) => source === null || source.includes(text)
300 // "block N" names a code block of the reply, already quoted; anything else is the value as written.
301 const block = value.match(/^block\s+(\d+)$/i)
302 const quoted = block ? (blocks[Number(block[1]) - 1] ?? '') : value.replace(/^`|`$/g, '')
303 const isFromReply = (text: string) => block !== null || isQuoted(text)
304 if (kind === 'open') {
305 const path = value.replace(/^`|`$/g, '')
306 return path !== '' && path.length <= 300 && !/^[a-z]+:/i.test(path) && isQuoted(path) ? { kind, path } : null
307 }
308 if (kind === 'copy') {
309 return quoted.trim() !== '' && quoted.length <= 8000 && isFromReply(quoted) ? { kind, text: quoted, name } : null
310 }
311 if (kind === 'run') {
312 const command = commandText(quoted)
313 const isUsable = command !== '' && command.length <= 4000 && isFromReply(command)
314 return isUsable ? { kind: NEEDS_PERSON.test(command) ? 'terminal' : 'run', command, name } : null
315 }
316 if (kind === 'link') {
317 try {
318 const url = new URL(value)
319 return url.protocol === 'https:' && isQuoted(value) ? { kind, url: url.href, name } : null
320 } catch {
321 return null
322 }
323 }
324
325 return null
326}
327
328/** A command as written, without surrounding space or a leading `!`. */
329function commandText(text: string): string {
330 return text.trim().replace(/^!\s*/, '')
331}
332
333/** The command a help runs or copies, so a copy of a command yields to its command button. */
334function commandOf(help: Help): string | null {
335 if (isCommand(help)) return help.command
336 if (help.kind === 'copy') return commandText(help.text)
337
338 return null
339}
340
341function isCommand(help: Help): help is Extract<Help, { kind: 'run' | 'terminal' }> {
342 return help.kind === 'run' || help.kind === 'terminal'
343}
344
345/** What a help acts on, without the name the model gave it, so one command under two names counts once. */
346function helpTarget(help: Help): string {
347 if (help.kind === 'open') return `open ${help.path}`
348 if (help.kind === 'link') return `link ${help.url}`
349
350 return `command ${commandOf(help)}`
351}
352
353function withHelp(helps: Help[], help: Help): Help[] {
354 const command = commandOf(help)
355 const kept = isCommand(help) ? helps.filter(h => h.kind !== 'copy' || commandOf(h) !== command) : helps
356 // A copy and a run of one command share a target, so either covers the other.
357 const isCovered = kept.some(h => helpTarget(h) === helpTarget(help))
358
359 return isCovered || kept.length >= MAX_HELPS ? kept : [...kept, help]
360}
361
362function clip(text: string, max: number): string {
363 return text.length <= max ? text : text.slice(0, max) + ' …[cut]'
364}
365
366/** The ledger as the model reads it: the card, the open items and findings with their ids, and recently settled items. */
367function ledgerBlocks(ledger: Ledger): string[] {
368 const card = ledger.card
369 ? [
370 `GOAL: ${ledger.card.goal}`,
371 ...ledger.card.done.map(d => `DONE: ${d}`),
372 `NOW: ${ledger.card.now}`,
373 ...ledger.card.running.map(r => `RUNNING: ${r}`),
374 ].join(NL)
375 : ''
376 const open = ledger.items.map(i => `${i.id} | ${MODEL_KIND[i.kind]} | ${i.label ?? '-'} | ${i.ask}`).join(NL)
377 const findings = ledger.findings.map(f => `${f.id} | ${f.kind}: ${f.title}`).join(NL)
378 // An expired item was never answered, so the reply may ask it again.
379 const closed = ledger.closed
380 .filter(d => d.how !== 'expired')
381 .slice(-8)
382 .map(d => `${d.ask} → ${outcomeText(d)}`)
383 .join(NL)
384
385 return [
386 `<card>${NL}${card}${NL}</card>`,
387 `<open>${NL}${open}${NL}</open>`,
388 `<findings>${NL}${findings}${NL}</findings>`,
389 `<decided>${NL}${closed}${NL}</decided>`,
390 ]
391}
392
393export function buildPrompt(ledger: Ledger, ex: Exchange): string {
394 const person = ex.person ?? `(The person sent nothing. The turn was started by: ${ex.trigger ?? 'unknown'}.)`
395
396 return [
397 ...ledgerBlocks(ledger),
398 `<person>${NL}${clip(person, 4000)}${NL}</person>`,
399 `<activity>${NL}${clip(ex.activity.join(NL), 2500)}${NL}</activity>`,
400 `<reply>${NL}${clip(numberBlocks(ex.reply), 12000)}${NL}</reply>`,
401 `<screen>${NL}${ex.screen}${NL}</screen>`,
402 `<checks>${NL}${ex.checks.length > 0 ? ex.checks.join(NL) : '(none ran)'}${NL}</checks>`,
403 ].join(NL)
404}
405
406/** The app the inbox runs in, as its texts name it: the agent, and where the person sees the inbox. */
407export type Host = {
408 agent: string
409 surface: string
410 /** Where the person sees the open items at all times, besides the surface; null when only the surface shows them. */
411 band: string | null
412 /** Where the person sees the findings, within the surface. */
413 findingsIn: string
414}
415
416export const CLAUDE_CODE: Host = {
417 agent: 'Claude',
418 surface: 'the /inbox pane',
419 band: 'a band above their prompt',
420 findingsIn: 'the Findings tab of the /inbox pane',
421}
422
423/** The host's surface at the start of a sentence. */
424function surfaceAtStart(host: Host): string {
425 return host.surface.charAt(0).toUpperCase() + host.surface.slice(1)
426}
427
428/** What the person sees besides the conversation, as the ledger model reads it. `tab` is the tab shown, when the surface has tabs the model should hear about. */
429export function screenText(host: Host, isOpen: boolean, tab: string | null): string {
430 return isOpen
431 ? `${surfaceAtStart(host)} is open beside the conversation, ${tab ? `on its ${tab} tab, ` : ''}listing every open item.`
432 : `${surfaceAtStart(host)} is closed.`
433}
434
435/**
436 * The open tasks a shell command the person ran completes: those whose run or
437 * sign-in step is that exact command, so nothing closes on a guess.
438 */
439export function tasksRunBy(ledger: Ledger, command: string): Item[] {
440 const typed = squash(command)
441
442 return ledger.items.filter(i => i.kind === 'task' && i.helps.some(h => isCommand(h) && squash(h.command) === typed))
443}
444
445/** A transcript row of the person's own command, as session.append carries it. */
446export type CommandRow =
447 | { kind: 'shell'; command: string }
448 | { kind: 'output'; stdout: string; stderr: string }
449 | { kind: 'slash'; name: string; args: string }
450
451/** Reads a `!` command, its output, or a slash command from a row's text; null for any other row. */
452export function readCommandRow(text: string): CommandRow | null {
453 const input = text.match(/^<bash-input>([\s\S]*)<\/bash-input>$/)
454 if (input) return { kind: 'shell', command: (input[1] ?? '').trim() }
455 const output = text.match(/<bash-stdout>([\s\S]*?)<\/bash-stdout><bash-stderr>([\s\S]*?)<\/bash-stderr>/)
456 if (output) return { kind: 'output', stdout: (output[1] ?? '').trim(), stderr: (output[2] ?? '').trim() }
457 const name = text.match(/<command-name>\/?([^<\s]+)<\/command-name>/)
458 if (name)
459 return {
460 kind: 'slash',
461 name: name[1] ?? '',
462 args: (text.match(/<command-args>([\s\S]*?)<\/command-args>/)?.[1] ?? '').trim(),
463 }
464
465 return null
466}
467
468/** Slash commands that say nothing about the work. */
469const QUIET_COMMANDS = new Set(['inbox', 'clear'])
470
471/** A command row as the ledger model reads it in <person>; null for one that says nothing. */
472export function commandRowLine(row: CommandRow): string | null {
473 if (row.kind === 'shell') return `$ ${row.command}`
474 if (row.kind === 'slash') return QUIET_COMMANDS.has(row.name) ? null : `/${row.name} ${row.args}`.trim()
475 const output = [row.stdout, row.stderr].filter(Boolean).join(NL)
476
477 return output ? `output: ${clip(output, 600)}` : null
478}
479
480function squash(command: string): string {
481 return command.trim().replace(/\s+/g, ' ')
482}
483
484function dash(value: string | undefined): string | null {
485 const v = (value ?? '').trim()
486
487 return v === '' || v === '-' ? null : v
488}
489
490/**
491 * Reads the model's line format; null when no known key appears. `source` is
492 * the reply followed by the activity: HELP values must appear in it, and
493 * "copy | block N" names its code blocks. Null skips both, so block copies drop.
494 */
495export function parseReply(text: string, source: string | null = null): Update | null {
496 const blocks = source === null ? [] : codeBlocks(source)
497 const card = { goal: '', done: [] as string[], now: '', running: [] as string[] }
498 const closed: Update['closed'] = []
499 const added: Update['added'] = []
500 const helped: Update['helped'] = []
501 const helps: { target: string; help: Help }[] = []
502 let seen = 0
503
504 for (const raw of text.split(NL)) {
505 const m = raw.match(/^\s*[-*]?\s*(GOAL|DONE|NOW|RUNNING|CLOSED|NEW|HELP)\s*:\s*(.*)$/)
506 if (!m) continue
507 seen += 1
508 const key = m[1]
509 // The model writes "-" for "nothing", for a whole value ("NOW: -") or one field ("NEW: do | - | - | - | -").
510 const value = dash(m[2]) ?? ''
511 if (key === 'GOAL') card.goal = value
512 else if (key === 'NOW') card.now = value
513 else if (key === 'DONE' && value) card.done.push(value)
514 else if (key === 'RUNNING' && value) card.running.push(value)
515 else if (key === 'CLOSED') {
516 const [id, outcome] = value.split('|').map(s => s.trim())
517 if (id) closed.push({ id, outcome: outcome ?? '' })
518 } else if (key === 'NEW') {
519 // The recommendation is the last field. The model sometimes writes one answer after "|" instead of " / ".
520 const [kind, label, ask, ...rest] = value.split('|').map(dash)
521 const rec = rest.length > 1 ? rest.pop() : null
522 if (!ask) continue
523 const options = rest.flatMap(o => o?.split(/\s+\/\s+/) ?? []).filter(Boolean)
524 added.push({
525 kind: readKind(kind),
526 label: label ?? null,
527 ask,
528 options,
529 rec: recommendedOption(options, rec ?? null),
530 helps: [],
531 })
532 } else if (key === 'HELP') {
533 const [target, kind, help, name] = value.split('|').map(s => s.trim())
534 const parsed = readHelp(kind ?? '', help ?? '', dash(name), blocks, source)
535 if (parsed && target) helps.push({ target, help: parsed })
536 }
537 }
538 // HELP lines name their item, and may come before or after its NEW line.
539 for (const { target, help } of helps) {
540 // Lenient about extra words: the model sometimes writes "open i9" for "i9".
541 const n = target.match(/\bnew\s*(\d+)\b/i)
542 const id = target.match(/\b(i\d+)\b/)?.[1]
543 const item = n ? added[Number(n[1]) - 1] : undefined
544 if (item) item.helps = withHelp(item.helps, help)
545 else if (!n && id) helped.push({ id, help })
546 }
547
548 return seen === 0 ? null : { card: { ...card, done: card.done.slice(-4) }, closed, added, helped }
549}
550
551const FILLER = new Set([
552 'the',
553 'and',
554 'for',
555 'with',
556 'use',
557 'into',
558 'from',
559 'that',
560 'this',
561 'your',
562 'you',
563 'should',
564 'make',
565 'add',
566 'all',
567])
568
569function keyWords(text: string): Set<string> {
570 return new Set((text.toLowerCase().match(/[a-z0-9]+/g) ?? []).filter(w => w.length > 2 && !FILLER.has(w)))
571}
572
573/**
574 * Whether one ask rewords the other: most of the shorter one's key words
575 * appear in the longer. Asks with one key word only match exactly, since
576 * "Deploy?" and "Deploy now?" can be different questions.
577 */
578function restates(a: string, b: string): boolean {
579 const x = keyWords(a)
580 const y = keyWords(b)
581 const fewer = Math.min(x.size, y.size)
582 const shared = [...x].filter(w => y.has(w)).length
583
584 return fewer >= 2 ? shared / fewer >= 0.7 : sameAsk(a, b)
585}
586
587function sameAsk(a: string, b: string): boolean {
588 const norm = (s: string) =>
589 s
590 .toLowerCase()
591 .replace(/[^a-z0-9]+/g, ' ')
592 .trim()
593
594 return norm(a) === norm(b)
595}
596
597/**
598 * The open item a NEW line repeats, if any. The summary model often restates
599 * an open item, labelled with its id or reworded, instead of leaving it alone.
600 */
601function matchOpen(items: Pick<Item, 'id' | 'kind' | 'ask'>[], a: Update['added'][number]): number {
602 return items.findIndex(
603 i => i.id === a.label || sameAsk(i.ask, a.ask) || (i.kind === a.kind && restates(i.ask, a.ask)),
604 )
605}
606
607/** Whether a NEW line repeats an item that closed at or after `since`, by matchOpen's rule or its label. */
608function repeatsRecentlyClosed(closed: Closed[], a: Update['added'][number], since: number): boolean {
609 const recent = closed.filter(c => c.at >= since)
610
611 return matchOpen(recent, a) >= 0 || (a.label !== null && recent.some(c => c.label === a.label))
612}
613
614function closedRecord(item: Item, closing: Closing, now: number): Closed {
615 const { id, kind, ask, label } = item
616
617 return { id, kind, ask, ...(label ? { label } : {}), ...closing, at: now }
618}
619
620/** How an outcome the mod saved before `how` existed marks an item Claude closed. */
621const CLOSED_BY_CLAUDE = 'closed by Claude'
622
623/**
624 * Closes an open item or finding for the agent: one the user answered in their
625 * own message, with that answer as its outcome, or one that is done or no
626 * longer applies, with the agent's reason. A finding, which the agent recorded
627 * itself, is removed. `closed` is null when no open one has the id. Its `how`
628 * is the stored value 'claude' for either agent.
629 */
630export function closeByAgent(
631 host: Host,
632 ledger: Ledger,
633 id: string,
634 how: { answer: string } | { reason: string },
635 now: number,
636): { ledger: Ledger; closed: 'item' | 'finding' | null } {
637 const closing: Closing =
638 'answer' in how
639 ? { how: 'answered', outcome: how.answer }
640 : { how: 'claude', outcome: `closed by ${host.agent}: ${how.reason}` }
641 if (ledger.items.some(i => i.id === id)) return { ledger: closeItem(ledger, id, closing, now), closed: 'item' }
642 if (ledger.findings.some(f => f.id === id))
643 return { ledger: { ...ledger, findings: ledger.findings.filter(f => f.id !== id) }, closed: 'finding' }
644
645 return { ledger, closed: null }
646}
647
648/** How an item is closing: what the pane shows, and how it closed. */
649export type Closing = Pick<Closed, 'how' | 'outcome'>
650
651/** Closes one item, recording how it closed. */
652export function closeItem(ledger: Ledger, id: string, closing: Closing, now: number): Ledger {
653 return {
654 ...ledger,
655 items: ledger.items.filter(i => i.id !== id),
656 closed: [...ledger.closed, ...ledger.items.filter(i => i.id === id).map(i => closedRecord(i, closing, now))].slice(
657 -MAX_CLOSED,
658 ),
659 }
660}
661
662/**
663 * `promptAt` is when the update's prompt was built. A NEW line that repeats an
664 * item closed since then is dropped: the person answered it while the model ran.
665 */
666export function applyUpdate(ledger: Ledger, u: Update, now: number, turn: number, promptAt = now): Ledger {
667 const prev = ledger.card
668 const card: Card = {
669 goal: u.card.goal || prev?.goal || '',
670 done: u.card.done.length > 0 ? u.card.done : (prev?.done ?? []),
671 now: u.card.now || prev?.now || '',
672 running: u.card.running,
673 updatedAt: now,
674 }
675 const closing = new Map(u.closed.map(c => [c.id, c.outcome]))
676 const closed = [...ledger.closed]
677 const items: Item[] = []
678 for (const item of ledger.items) {
679 const outcome = closing.get(item.id)
680 const helps = u.helped.filter(h => h.id === item.id).reduce((all, h) => withHelp(all, h.help), item.helps)
681 if (outcome === undefined) items.push({ ...item, helps })
682 else closed.push(closedRecord(item, { outcome, how: 'update' }, now))
683 }
684
685 let nextId = ledger.nextId
686 let added = 0
687 for (const a of u.added) {
688 if (repeatsRecentlyClosed(ledger.closed, a, promptAt)) continue
689 const at = matchOpen(items, a)
690 const restated = items[at]
691 if (restated) {
692 items[at] = { ...restated, helps: a.helps.reduce((all, h) => withHelp(all, h), restated.helps) }
693 continue
694 }
695 items.push({ ...a, id: `i${nextId}`, turn, at: now })
696 nextId += 1
697 added += 1
698 }
699
700 // An item left unanswered too long, or pushed out by newer ones, closes as expired.
701 const kept = items.filter(i => turn - i.turn <= STALE_AFTER).slice(-MAX_OPEN)
702 for (const i of items) if (!kept.includes(i)) closed.push(closedRecord(i, { outcome: EXPIRED, how: 'expired' }, now))
703
704 return {
705 ...ledger,
706 card,
707 items: kept,
708 closed: closed.slice(-MAX_CLOSED),
709 findings: ledger.findings.filter(f => !closing.has(f.id)),
710 nextId,
711 batchTurn: added > 0 ? turn : ledger.batchTurn,
712 }
713}
714
715/** The items the agent's latest reply added, which "1. yes" answers refer to. */
716export function latestBatch(ledger: Ledger): Item[] {
717 return ledger.batchTurn === 0 ? [] : ledger.items.filter(i => i.turn === ledger.batchTurn)
718}
719
720function describe(item: Item): string {
721 const parts = [`"${item.ask}"`]
722 if (item.options.length > 0) parts.push(`options: ${item.options.join(' / ')}`)
723 if (item.rec) parts.push(`recommended: ${item.rec}`)
724 if (item.kind === 'task') parts.push('an action for the user')
725
726 return parts.join('; ')
727}
728
729function numberOf(label: string | null): number | null {
730 const m = label?.match(/(\d+)/)
731
732 return m ? Number(m[1]) : null
733}
734
735const LINE_ANSWER = /(?:^|\n)\s*(?:[QqDd#]\s?)?(\d{1,2})\s*[.):\-–]\s*\S/g
736const INLINE_ANSWER = /\s(?:[QqDd#]\s?)?(\d{1,2})\s*[.)]\s+\S/g
737const ACCEPT_ALL =
738 /^\s*(go|go ahead|yes|yep|yeah|sure|ok|okay|sgtm|lgtm|sounds good|do it|proceed|all good|ship it)\s*[.!]*\s*$/i
739const ACCEPT_RECS = /\b(all|both|everything|your)\b[^.\n]{0,40}\b(recommend\w*|recs?|suggest\w*|picks?|calls?)\b/i
740
741/**
742 * The note attached to the person's prompt when it answers open items:
743 * numbered answers mapped to the latest batch, an item quoted from the band,
744 * or a blanket "go" over the latest recommendations. Null when it answers none.
745 *
746 * `ledger.turn` already counts this prompt, so the latest batch is only
747 * offered when it came from the reply just before it.
748 */
749export function answerNote(ledger: Ledger, text: string): string | null {
750 const lines: string[] = []
751 const batch = ledger.batchTurn === ledger.turn - 1 ? latestBatch(ledger) : []
752
753 if (batch.length > 0) {
754 const numbers = new Set<number>()
755 for (const m of text.matchAll(LINE_ANSWER)) numbers.add(Number(m[1]))
756 // "1. node 2. yes" on one line counts only when the message opens with an
757 // answer, so "bullet 4." mid-sentence is not read as answering item 4.
758 if (/^\s*(?:[QqDd#]\s?)?\d{1,2}\s*[.):\-–]\s/.test(text)) {
759 for (const m of text.matchAll(INLINE_ANSWER)) numbers.add(Number(m[1]))
760 }
761 const hasLabels = batch.some(i => numberOf(i.label) !== null)
762 for (const n of [...numbers].sort((a, b) => a - b)) {
763 const item = hasLabels ? batch.find(i => numberOf(i.label) === n) : batch[n - 1]
764 if (item) lines.push(`- ${n} → ${describe(item)}`)
765 }
766 if (lines.length === 0 && (ACCEPT_ALL.test(text) || ACCEPT_RECS.test(text))) {
767 const withRecs = batch.filter(i => i.rec)
768 if (withRecs.length > 0) {
769 return [
770 'inbox: if the user is accepting your recommendations, these are the open ones:',
771 ...withRecs.map(i => `- ${i.label ?? '•'} → ${describe(i)}`),
772 ].join(NL)
773 }
774 }
775 }
776
777 return lines.length === 0
778 ? null
779 : ["inbox: the user's message answers these open items from your earlier replies:", ...lines].join(NL)
780}
781
782/**
783 * Adds a finding Claude recorded. One whose title matches an open one is
784 * skipped, so the same finding is not listed twice.
785 */
786export function addFinding(
787 ledger: Ledger,
788 finding: Omit<Finding, 'id'>,
789): { ledger: Ledger; id: string; isAdded: boolean } {
790 const findings = ledger.findings
791 const same = findings.find(f => sameAsk(f.title, finding.title))
792 if (same) return { ledger, id: same.id, isAdded: false }
793 const id = `f${ledger.nextId}`
794
795 return {
796 ledger: {
797 ...ledger,
798 findings: [...findings, { ...finding, id }].slice(-MAX_FINDINGS),
799 nextId: ledger.nextId + 1,
800 },
801 id,
802 isAdded: true,
803 }
804}
805
806/** An open item as Claude reads it, with its id when Claude may close it. */
807function itemLine(item: Item, withId: boolean): string {
808 return `- ${withId ? `[${item.id}] ` : ''}${item.label ? `(${item.label}) ` : ''}${describe(item)}`
809}
810
811function findingLine(finding: Finding, withId: boolean): string {
812 return `- ${withId ? `[${finding.id}] ` : ''}${finding.kind}: ${finding.title}`
813}
814
815/**
816 * What the model reads after a compaction, so open items and the goal
817 * survive it. Ids go in only for this session's own ledger, since Claude
818 * closes items by id; the previous session's items are not this one's.
819 */
820export function carryText(ledger: Ledger, title: string, isOwn = false): string | null {
821 const findings = ledger.findings
822 if (!ledger.card && ledger.items.length === 0 && findings.length === 0) return null
823 const out = [title]
824 if (ledger.card) {
825 if (ledger.card.goal) out.push(`Goal: ${ledger.card.goal}`)
826 if (ledger.card.done.length > 0) out.push(`Done: ${ledger.card.done.join('; ')}`)
827 if (ledger.card.now) out.push(`Now: ${ledger.card.now}`)
828 if (ledger.card.running.length > 0) out.push(`Running: ${ledger.card.running.join('; ')}`)
829 }
830 if (ledger.items.length > 0) {
831 out.push('Waiting on the user:')
832 for (const item of ledger.items) out.push(itemLine(item, isOwn))
833 }
834 if (findings.length > 0) {
835 out.push(isOwn ? 'Findings you recorded, still open:' : 'Findings that session recorded, still open:')
836 for (const f of findings) out.push(findingLine(f, isOwn))
837 }
838 if (ledger.closed.length > 0) {
839 out.push('Recently closed:')
840 for (const d of ledger.closed.slice(-6)) out.push(`- "${d.ask}" → ${outcomeText(d)}`)
841 }
842
843 return out.join(NL)
844}
845
846/**
847 * The inbox as the agent reads it beside a prompt: what is open now and whether
848 * the host's surface shows it. prompt.context reaches only the first message,
849 * so this is how the agent learns what changed since.
850 */
851export function inboxText(host: Host, ledger: Ledger, isOpen: boolean): string {
852 const out = ['inbox: what waits on the user, as of your last reply. Anything not listed here is closed.']
853 if (ledger.items.length > 0) {
854 out.push('Waiting on the user:')
855 for (const item of ledger.items) out.push(itemLine(item, true))
856 } else {
857 out.push('Nothing is waiting on the user.')
858 }
859 if (ledger.findings.length > 0) {
860 out.push('Findings you recorded, still open:')
861 for (const f of ledger.findings) out.push(findingLine(f, true))
862 }
863 out.push(isOpen ? `The user has ${host.surface} open beside the conversation.` : `${surfaceAtStart(host)} is closed.`)
864
865 return out.join(NL)
866}
867
868/** An item that closed without the person deciding it: dismissed, expired, or overtaken by the work. */
869export function isLapsed(d: Closed): boolean {
870 if (d.how === 'dismissed' || d.how === 'expired' || d.how === 'claude') return true
871 // The per-reply update writes its own outcome, so only its wording says the work overtook the item.
872 return d.how === 'update' && /^(no longer applies|replaced|superseded|moot)/i.test(d.outcome)
873}
874
875/** How an item closed, in words a model reads without the mod's vocabulary. */
876function outcomeText(d: Closed): string {
877 if (d.how === 'dismissed') return 'dismissed by the user'
878 if (d.how === 'expired') return 'expired before the user answered'
879
880 return d.outcome
881}
882
883/**
884 * The items closed since Claude last read the inbox, so Claude can tell a
885 * dismissed question from one still waiting, and an expired one from one answered.
886 */
887export function closedText(closed: Closed[]): string | null {
888 if (closed.length === 0) return null
889 const advice = (d: Closed) =>
890 d.how === 'dismissed'
891 ? '. Ask it again only if the user brings it up.'
892 : d.how === 'expired'
893 ? '. Ask it again if it still matters.'
894 : ''
895
896 return [
897 'Closed since you last read the inbox:',
898 ...closed.map(d => `- "${d.ask}" → ${outcomeText(d)}${advice(d)}`),
899 ].join(NL)
900}
901
902/** The inbox text the agent last read beside a prompt, and the closed items it was told about, by id. */
903export type Told = { inbox: string | null; closed: string[] }
904
905export const TOLD_NOTHING: Told = { inbox: null, closed: [] }
906
907/**
908 * What the agent reads beside a prompt: whether it answers an open item,
909 * unless a press sent it, and the inbox when it changed since the agent last
910 * read it or when an item closed since. An empty inbox with nothing closed
911 * says nothing new. Also returns what the agent has now been told.
912 */
913export function promptNotes(
914 host: Host,
915 ledger: Ledger,
916 prompt: { text: string; isPress: boolean; isOpen: boolean },
917 told: Told,
918): { notes: string[]; told: Told } {
919 const notes: string[] = []
920 // A press's message already says what it does; an Explain, for one, quotes its item without answering it.
921 const answer = prompt.isPress ? null : answerNote(ledger, prompt.text)
922 if (answer) notes.push(answer)
923 const inbox = inboxText(host, ledger, prompt.isOpen)
924 const toldClosed = new Set(told.closed)
925 const closed = closedText(ledger.closed.filter(d => !toldClosed.has(d.id)))
926 const isEmpty = ledger.items.length === 0 && ledger.findings.length === 0
927 if (!closed && (inbox === told.inbox || (isEmpty && told.inbox === null))) return { notes, told }
928 notes.push(closed ? `${inbox}${NL}${closed}` : inbox)
929
930 return { notes, told: { inbox, closed: ledger.closed.map(d => d.id) } }
931}
932
933/**
934 * The stop's kind, from the error word and the message Claude Code showed.
935 * A reached usage limit and a busy server are both `rate_limit`; only the
936 * message, "You've hit your weekly limit · resets 7:33pm", tells them apart.
937 */
938export function stopKindOf(error: string, message: string): Stop['kind'] {
939 if (error === 'rate_limit') return /\bhit your\b.*\blimit\b/i.test(message) ? 'usage-limit' : 'api-error'
940 switch (error) {
941 case 'authentication_failed':
942 case 'oauth_org_not_allowed':
943 case 'verification_required':
944 case 'cloud_credential_error':
945 return 'sign-in'
946 case 'billing_error':
947 case 'account_on_hold':
948 return 'billing'
949 default:
950 return 'api-error'
951 }
952}
953
954/** When a usage limit resets, as Claude Code's message says it: "7:33pm". */
955export function resetTime(message: string): string | null {
956 return message.match(/\bresets\s+(?:at\s+)?([^·(\n]+?)\s*(?:\(|·|$)/i)?.[1]?.trim() ?? null
957}
958
959/** "sign-in expired" */
960export function stopText(stop: Stop): string {
961 switch (stop.kind) {
962 case 'sign-in':
963 return 'sign-in expired'
964 case 'billing':
965 return 'billing problem'
966 case 'usage-limit':
967 return 'usage limit reached'
968 default:
969 return stop.detail === 'rate_limit'
970 ? 'the API is limiting requests'
971 : stop.detail === 'overloaded'
972 ? 'the API is overloaded'
973 : `API error (${stop.detail})`
974 }
975}
976
977/** What the person does to clear a stop. */
978export function stopFix(stop: Stop): string {
979 switch (stop.kind) {
980 case 'sign-in':
981 return 'Run /login, then send a message to resume.'
982 case 'billing':
983 return 'Check billing in the Claude console, then resume.'
984 case 'usage-limit':
985 return stop.resets ? `Resume after ${stop.resets}.` : 'Resume when the limit resets.'
986 default:
987 return 'Send a message to resume.'
988 }
989}
990
991/** The stop and its fix in a few words, for the sidebar. */
992function stopShort(stop: Stop): string {
993 switch (stop.kind) {
994 case 'sign-in':
995 return 'Signed out: /login'
996 case 'billing':
997 return 'Billing problem'
998 case 'usage-limit':
999 return stop.resets ? `Limit: resets ${stop.resets}` : 'Usage limit reached'
1000 default:
1001 return 'API error: resume'
1002 }
1003}
1004
1005/** "Allow push main to origin?", or a question dialog's question. */
1006function dialogText(dialog: Dialog): string {
1007 return dialog.kind === 'permission' ? `Allow ${dialog.text}?` : dialog.text
1008}
1009
1010/**
1011 * The session's line in the Herdr sidebar: a stop and its fix, else an open
1012 * dialog, else how many items wait and the first of the latest reply's, else
1013 * where the work stands. The count leads, because the sidebar cuts long lines
1014 * at its edge. Empty when there is nothing to say.
1015 */
1016export function statusLine(ledger: Ledger, stop: Stop | null, dialogs: Dialog[]): string {
1017 if (stop) return `! ${stopShort(stop)}`
1018 const dialog = dialogs[0]
1019 if (dialog) return dialogText(dialog)
1020 const first = latestBatch(ledger)[0] ?? ledger.items[0]
1021 if (first) return ledger.items.length > 1 ? `${ledger.items.length} · ${first.ask}` : first.ask
1022 const count = ledger.findings.length
1023 const findings = count === 0 ? null : `${count} finding${count === 1 ? '' : 's'}`
1024
1025 return [ledger.card?.now, findings].filter(Boolean).join(' · ')
1026}
1027
1028/** "just now", "12m ago", "3h ago", "2d ago". */
1029export function ago(ms: number): string {
1030 const min = Math.round(ms / 60000)
1031 if (min < 1) return 'just now'
1032 if (min < 60) return `${min}m ago`
1033 const h = Math.round(min / 60)
1034 if (h < 48) return `${h}h ago`
1035
1036 return `${Math.round(h / 24)}d ago`
1037}
1038hooks/prs.ts 355 lines1import type { PrCheck, PrThread, PrView } from '../types'
2
3const NL = '\n'
4const PR_URL = /https:\/\/github\.com\/([\w.-]+\/[\w.-]+)\/pull\/(\d+)/g
5
6/** "owner/repo#123" for each GitHub PR link in the text, in order, without repeats. */
7export function prRefs(text: string): string[] {
8 return [...new Set([...text.matchAll(PR_URL)].map(m => `${m[1]}#${m[2]}`))]
9}
10
11/**
12 * The PRs of `refs` that the text names as "#123", in its order. A number two
13 * of `refs` share names neither, since the text does not say which repo.
14 */
15export function namedPrs(text: string, refs: string[]): string[] {
16 const numbers = new Set([...text.matchAll(/#(\d+)\b/g)].map(m => m[1]))
17 const known = [...new Set(refs)]
18
19 return [...numbers].flatMap(n => {
20 const named = known.filter(ref => parseRef(ref).number === n)
21 return named.length === 1 ? named : []
22 })
23}
24
25/** "owner/repo#123" as its parts. */
26export function parseRef(ref: string): { repo: string; owner: string; name: string; number: string } {
27 const [repo = '', number = ''] = ref.split('#')
28 const [owner = '', name = ''] = repo.split('/')
29
30 return { repo, owner, name, number }
31}
32
33/** The fields `gh pr view --json` returns that the tab uses. */
34export const VIEW_FIELDS = 'number,title,url,isDraft,state,baseRefName,mergeable,reviewDecision,statusCheckRollup'
35
36/**
37 * Unresolved review threads: the first comment, which states the finding, and
38 * who wrote the last one, which says whose turn it is. The latest commit's
39 * time tells a reply that came after the lines changed from one before.
40 */
41export const THREADS_QUERY = `query($owner: String!, $repo: String!, $number: Int!) {
42 viewer { login }
43 repository(owner: $owner, name: $repo) {
44 pullRequest(number: $number) {
45 commits(last: 1) { nodes { commit { committedDate } } }
46 reviewThreads(first: 100) {
47 nodes {
48 id isResolved isOutdated path line originalLine
49 comments(first: 1) { totalCount nodes { author { login } body url createdAt } }
50 last: comments(last: 1) { nodes { author { login } body url createdAt } }
51 }
52 }
53 }
54 }
55}`
56
57type RollupEntry = {
58 __typename?: string
59 name?: string
60 context?: string
61 status?: string
62 conclusion?: string
63 state?: string
64 detailsUrl?: string
65 targetUrl?: string
66}
67
68/** gh's check rollup as pass, fail, pending or skip; a check run and a commit status spell it differently. */
69function bucket(e: RollupEntry): PrCheck['bucket'] {
70 const result = (e.conclusion || e.state || '').toUpperCase()
71 if (e.status && e.status.toUpperCase() !== 'COMPLETED') return 'pending'
72 if (['SUCCESS', 'NEUTRAL'].includes(result)) return 'pass'
73 if (['SKIPPED', 'STALE'].includes(result)) return 'skip'
74 if (['PENDING', 'EXPECTED', 'QUEUED', 'IN_PROGRESS', ''].includes(result)) return 'pending'
75
76 return 'fail'
77}
78
79/** Reads `gh pr view --json VIEW_FIELDS` output; null when it is not that. */
80export function readView(ref: string, json: string): Omit<PrView, 'threads' | 'fetchedAt' | 'error'> | null {
81 try {
82 const v = JSON.parse(json) as Record<string, unknown>
83 if (typeof v.number !== 'number') return null
84 const rollup = Array.isArray(v.statusCheckRollup) ? (v.statusCheckRollup as RollupEntry[]) : []
85
86 return {
87 ref,
88 number: v.number,
89 title: String(v.title ?? ''),
90 url: String(v.url),
91 isDraft: v.isDraft === true,
92 state: String(v.state ?? 'OPEN'),
93 base: String(v.baseRefName ?? ''),
94 mergeable: String(v.mergeable ?? 'UNKNOWN'),
95 reviewDecision: String(v.reviewDecision ?? ''),
96 checks: rollup.map(e => ({
97 name: e.name ?? e.context ?? 'check',
98 bucket: bucket(e),
99 url: e.detailsUrl ?? e.targetUrl ?? null,
100 })),
101 }
102 } catch {
103 return null
104 }
105}
106
107type ThreadNode = {
108 id: string
109 isResolved: boolean
110 isOutdated: boolean
111 path: string
112 line: number | null
113 originalLine: number | null
114 comments: {
115 totalCount: number
116 nodes: { author: { login: string } | null; body: string; url: string; createdAt?: string }[]
117 }
118 last: { nodes: { author: { login: string } | null; body: string; url: string; createdAt?: string }[] }
119}
120
121type ThreadsAnswer = {
122 data?: {
123 viewer?: { login: string }
124 repository?: {
125 pullRequest?: {
126 commits?: { nodes?: { commit?: { committedDate?: string } }[] }
127 reviewThreads?: { nodes?: ThreadNode[] }
128 }
129 }
130 }
131}
132
133/** Reads the THREADS_QUERY answer: open threads only, oldest first. */
134export function readThreads(json: string): PrThread[] {
135 try {
136 const data = (JSON.parse(json) as ThreadsAnswer).data
137 const viewer = data?.viewer?.login
138 const pr = data?.repository?.pullRequest
139 const headAt = Date.parse(pr?.commits?.nodes?.[0]?.commit?.committedDate ?? '') || null
140
141 return (pr?.reviewThreads?.nodes ?? [])
142 .filter(t => !t.isResolved)
143 .map(t => {
144 const first = t.comments.nodes[0]
145 const last = t.comments.totalCount > 1 ? t.last.nodes[0] : undefined
146 const reply = last
147 ? {
148 author: last.author?.login ?? 'ghost',
149 body: last.body.trim(),
150 url: last.url,
151 at: Date.parse(last.createdAt ?? '') || null,
152 }
153 : null
154
155 // With either time unknown, someone else's reply keeps the thread waiting.
156 const isAnsweredSince =
157 reply !== null && reply.author !== viewer && (headAt === null || reply.at === null || reply.at > headAt)
158
159 return {
160 id: t.id,
161 author: first?.author?.login ?? 'ghost',
162 reply,
163 isWaiting: (reply?.author ?? first?.author?.login) !== viewer,
164 isOutdated: t.isOutdated,
165 isLinesChanged: t.isOutdated && !isAnsweredSince,
166 path: t.path,
167 line: t.line ?? t.originalLine,
168 body: (first?.body ?? '').trim(),
169 replies: Math.max(0, t.comments.totalCount - 1),
170 url: first?.url ?? '',
171 at: Date.parse(first?.createdAt ?? '') || null,
172 }
173 })
174 } catch {
175 return []
176 }
177}
178
179/** Open threads whose last comment someone else wrote: the PRs tab lists these. */
180export function waitingThreads(pr: PrView): PrThread[] {
181 return pr.threads.filter(t => t.isWaiting)
182}
183
184/** Which of a PR's rows the person handed to Claude: a thread sent to it, or a failing check whose fix was sent. */
185export type Handoffs = {
186 isThreadSent: (pr: PrView, t: PrThread) => boolean
187 isFixSent: (pr: PrView, c: PrCheck) => boolean
188}
189
190export const NO_HANDOFFS: Handoffs = { isThreadSent: () => false, isFixSent: () => false }
191
192/** Waiting threads still on the person: not sent to Claude, and not on lines a later commit changed. */
193export function threadsOnYou(pr: PrView, h: Handoffs): PrThread[] {
194 return waitingThreads(pr).filter(t => !t.isLinesChanged && !h.isThreadSent(pr, t))
195}
196
197/** The PR rows that wait on the person, for every count of them: failing checks with no fix sent, and threads on them. */
198export function prRowsOnYou(pr: PrView, h: Handoffs): number {
199 return failingChecks(pr).filter(c => !h.isFixSent(pr, c)).length + threadsOnYou(pr, h).length
200}
201
202export function checkCounts(pr: PrView): Record<PrCheck['bucket'], number> {
203 const counts = { pass: 0, fail: 0, pending: 0, skip: 0 }
204 for (const c of pr.checks) counts[c.bucket] += 1
205
206 return counts
207}
208
209/** "1 thread waiting on you", "3 threads waiting on you". */
210function threadsWaiting(count: number): string {
211 return `${count} ${count === 1 ? 'thread' : 'threads'} waiting on you`
212}
213
214/** Where a PR stands. A draft is open and not ready, whatever else blocks it. */
215export type PrStatus = 'ready' | 'blocked' | 'draft' | 'merged' | 'closed'
216
217function plural(count: number, one: string, many: string): string {
218 return `${count} ${count === 1 ? one : many}`
219}
220
221/** "2 failing checks", or with how many have a fix sent. A check stays a blocker until it passes. */
222function failingText(failing: number, sent: number): string {
223 const checks = plural(failing, 'failing check', 'failing checks')
224 if (sent === 0) return checks
225
226 return sent === failing ? `${checks}, fix sent` : `${checks}, ${sent} with a fix sent`
227}
228
229/**
230 * What stands between the PR and merging, or that it is ready. A thread sent
231 * to Claude or on changed lines still blocks, as it is still open on GitHub.
232 */
233export function readiness(pr: PrView, h: Handoffs = NO_HANDOFFS): { status: PrStatus; text: string } {
234 if (pr.state === 'MERGED') return { status: 'merged', text: 'Merged' }
235 if (pr.state !== 'OPEN') return { status: 'closed', text: 'Closed' }
236 const { fail: failing, pending } = checkCounts(pr)
237 const fixesSent = failingChecks(pr).filter(c => h.isFixSent(pr, c)).length
238 const waiting = waitingThreads(pr)
239 const open = threadsOnYou(pr, h).length
240 const sent = waiting.filter(t => h.isThreadSent(pr, t)).length
241 const changed = waiting.filter(t => t.isLinesChanged && !h.isThreadSent(pr, t)).length
242 const blockers = [
243 pr.isDraft ? 'draft' : null,
244 pr.mergeable === 'CONFLICTING' ? `conflicts with ${pr.base}` : null,
245 failing > 0 ? failingText(failing, fixesSent) : null,
246 pr.reviewDecision === 'CHANGES_REQUESTED' ? 'changes requested' : null,
247 open > 0 ? threadsWaiting(open) : null,
248 sent > 0 ? `${plural(sent, 'thread', 'threads')} sent to Claude` : null,
249 changed > 0 ? `${plural(changed, 'thread', 'threads')} on changed lines` : null,
250 pr.reviewDecision === 'REVIEW_REQUIRED' ? 'needs approval' : null,
251 pending > 0 ? `${pending} ${pending === 1 ? 'check' : 'checks'} running` : null,
252 ].filter((b): b is string => b !== null)
253
254 return blockers.length === 0
255 ? {
256 status: 'ready',
257 text: `Ready to merge${pr.reviewDecision === 'APPROVED' ? ': approved' : ''}, checks pass, no threads waiting on you`,
258 }
259 : { status: pr.isDraft ? 'draft' : 'blocked', text: `Blocked: ${blockers.join(', ')}` }
260}
261
262/** The band's one-line PR alert: the first open PR that needs the person, or null. */
263export function prAttention(views: PrView[], h: Handoffs = NO_HANDOFFS): string | null {
264 for (const pr of views) {
265 if (pr.state !== 'OPEN') continue
266 const open = threadsOnYou(pr, h).length
267 if (failingChecks(pr).some(c => !h.isFixSent(pr, c))) return `PR #${pr.number} CI failing`
268 if (pr.reviewDecision === 'CHANGES_REQUESTED') return `PR #${pr.number} changes requested`
269 if (open > 0) return `PR #${pr.number} ${threadsWaiting(open)}`
270 }
271
272 return null
273}
274
275/**
276 * The HTML tags review bots write in comments, such as `<sub>` and `<details>`.
277 * Lowercase only, so a type parameter such as `Props<P>` stays.
278 */
279const HTML_TAG =
280 /<\/?(?:a|b|br|code|details|div|em|h[1-6]|hr|i|img|kbd|li|ol|p|pre|span|strong|sub|summary|sup|table|tbody|td|th|thead|tr|ul)\b[^>]*>/g
281
282/**
283 * A review comment's Markdown without the HTML that review bots add, and with
284 * each image as its alt text, so a bot's `` reads "P1".
285 */
286export function readableComment(body: string): string {
287 return body
288 .replace(/<!--[\s\S]*?-->/g, '')
289 .replace(HTML_TAG, '')
290 .replace(/!\[([^\]]*)\]\([^)]*\)/g, (_, alt: string) => alt.replace(/\s*badge$/i, ''))
291 .trim()
292}
293
294/** A review comment as one line of plain text, for a row that shows only its start. */
295export function commentLine(body: string): string {
296 return readableComment(body)
297 .replace(/\[([^\]]*)\]\([^)]*\)/g, '$1')
298 .replace(/^\s*(?:#{1,6}|>)\s*/gm, '')
299 .replace(/\*\*|__|`/g, '')
300 .replace(/\s+/g, ' ')
301}
302
303/** Where a thread sits in the diff: "path:line", or the path alone. */
304export function threadWhere(t: PrThread, path = t.path): string {
305 return `${path}${t.line ? `:${t.line}` : ''}`
306}
307
308export function failingChecks(pr: PrView): PrCheck[] {
309 return pr.checks.filter(c => c.bucket === 'fail')
310}
311
312function threadText(t: PrThread): string {
313 const where = threadWhere(t)
314
315 const opening = [`${where}, from @${t.author}${t.isOutdated ? ' (on an outdated diff)' : ''}:`, t.body, t.url]
316 const reply = t.reply ? ['', `Latest reply, from @${t.reply.author}:`, t.reply.body, t.reply.url] : []
317
318 return [...opening, ...reply].join(NL)
319}
320
321/** The message each PR button sends to Claude. */
322export const prompts = {
323 fix: (pr: PrView, check: PrCheck) =>
324 [
325 `The CI check "${check.name}" is failing on PR #${pr.number} (${pr.url}).`,
326 'Find out why from its logs, fix it, and verify the fix. If the fix needs a change to CI configuration, ask me before making it.',
327 ...(check.url ? [`Check details: ${check.url}`] : []),
328 ].join(NL),
329 resolve: (pr: PrView) =>
330 [
331 `PR #${pr.number} (${pr.url}) conflicts with ${pr.base}. Update its branch from ${pr.base}, resolve the conflicts, and verify.`,
332 'Ask me before you push, and tell me how you resolved each conflict.',
333 ].join(NL),
334 address: (pr: PrView, threads: PrThread[]) =>
335 [
336 threads.length === 1
337 ? `Address this review comment on PR #${pr.number} (${pr.url}). Fix it in code and verify.`
338 : `Address these ${threads.length} review comments on PR #${pr.number} (${pr.url}). Fix each in code and verify.`,
339 "Don't reply on GitHub or resolve the threads; tell me what you changed for each.",
340 ...threads.map(t => NL + threadText(t)),
341 ].join(NL),
342 draft: (pr: PrView, t: PrThread) =>
343 [
344 `Draft a reply to this review comment on PR #${pr.number} (${pr.url}) for me to review. Don't post it.`,
345 '',
346 threadText(t),
347 ].join(NL),
348 discuss: (pr: PrView, t: PrThread) =>
349 [
350 `Let's talk through this review comment on PR #${pr.number} (${pr.url}) before changing anything.`,
351 '',
352 threadText(t),
353 ].join(NL),
354}
355hooks/presses.ts 127 lines1// What the person's presses send and how their buttons are labeled, the same
2// in every host: each message goes to the agent as the person's own words.
3
4import type { Finding, Help, Item } from '../types'
5
6export function clipLabel(text: string, max: number): string {
7 return text.length > max ? `${text.slice(0, max - 1)}…` : text
8}
9
10export function baseName(path: string): string {
11 return path.replace(/\/+$/, '').split('/').pop() ?? path
12}
13
14export function helpLabel(help: Help): string {
15 const label =
16 help.kind === 'open'
17 ? `Open ${baseName(help.path)}`
18 : help.kind === 'copy'
19 ? `Copy ${help.name ?? 'snippet'}`
20 : help.kind === 'run'
21 ? `Run ${help.name ?? help.command}`
22 : help.kind === 'terminal'
23 ? `Copy ${help.name ?? help.command}`
24 : `Open ${help.name ?? new URL(help.url).host}`
25
26 return clipLabel(label, 32)
27}
28
29/** A path an item names, made absolute: `~/` from the home folder, a relative one from the session's folder. */
30export function localPath(raw: string, root: string, home: string): string {
31 if (raw.startsWith('~/')) return home + raw.slice(1)
32 if (raw.startsWith('/')) return raw
33
34 return `${root.replace(/\/$/, '')}/${raw.replace(/^\.\//, '')}`
35}
36
37// Files macOS `open` would run or install instead of showing.
38const LAUNCHES =
39 /\.(app|command|tool|terminal|workflow|scpt|scptd|applescript|pkg|mpkg|dmg|webloc|inetloc|fileloc|prefpane|kext)$/i
40
41/**
42 * The commands that open a local path. A folder, an executable, or anything
43 * `open` would launch is shown in Finder instead. A file opens in the app
44 * macOS assigns to its type, and `fallback` opens one with no assigned app
45 * in the default text editor.
46 */
47export function openCommands(
48 path: string,
49 isFile: boolean,
50 isExecutable: boolean,
51): { argv: string[]; fallback: string[] | null } {
52 if (!isFile || isExecutable || LAUNCHES.test(path)) return { argv: ['open', '-R', path], fallback: null }
53
54 return { argv: ['open', path], fallback: ['open', '-t', path] }
55}
56
57/** One button's helps, used in order by one press. */
58export type HelpStep = { label: string; step: Help[] }
59
60/**
61 * The item's helps as buttons. A snippet to copy and a file to open become one
62 * step, "Copy env line and open .env.local", since the snippet goes in that file.
63 */
64export function steps(helps: Help[]): HelpStep[] {
65 const copy = helps.find(h => h.kind === 'copy')
66 const open = helps.find(h => h.kind === 'open')
67 if (!copy || !open || copy.kind !== 'copy' || open.kind !== 'open')
68 return helps.map(h => ({ label: helpLabel(h), step: [h] }))
69
70 return helps
71 .filter(h => h !== copy)
72 .map(h =>
73 h === open
74 ? { label: clipLabel(`Copy ${copy.name ?? 'snippet'} and open ${baseName(open.path)}`, 48), step: [copy, open] }
75 : { label: helpLabel(h), step: [h] },
76 )
77}
78
79/**
80 * A task handed to the agent waits on the agent until the update for a turn
81 * started after the press has applied. Still open then, the agent's reply did
82 * not finish it, so it waits on the person again. Counts lower than at the
83 * press were reset, so the task no longer folds.
84 */
85export function isTaskHandedOff(
86 last: { isHandoff?: boolean; turnsStarted?: number } | undefined,
87 turns: { turnsStarted: number; turnsApplied: number },
88): boolean {
89 const pressed = last?.isHandoff === true ? last.turnsStarted : undefined
90
91 return pressed !== undefined && pressed <= turns.turnsStarted && turns.turnsApplied <= pressed
92}
93
94/** A finding as the agent reads it back: its kind and title, its detail, and its file. */
95function findingBody(finding: Finding): string[] {
96 return [
97 `${finding.kind === 'issue' ? 'Issue' : 'Opportunity'}: ${finding.title}`,
98 finding.detail,
99 ...(finding.path ? [`File: ${finding.path}`] : []),
100 ]
101}
102
103/** The message each press sends. */
104export const messages = {
105 answer: (item: Item, answer: string) => `Re "${item.ask}": ${answer}`,
106 /** Asks what an item is about. The item stays open, since nothing was decided. */
107 explain: (item: Item) => {
108 const what = item.kind === 'task' ? 'this task you left for me' : 'this question you asked me'
109 const options = item.options.length > 0 ? `\nOptions: ${item.options.join(' / ')}` : ''
110
111 return `Remind me what ${what} is about: why it came up, and what each choice would mean. Don't act on it yet.\n"${item.ask}"${options}`
112 },
113 run: (item: Item, command: string) => `For "${item.ask}", run this:\n\`\`\`\n${command}\n\`\`\``,
114 taskReply: (item: Item, words: string) => `Re the task you left for me, "${item.ask}": ${words}`,
115 /** Sends a finding back to the agent: to fix it, to talk it through first, or with the person's own words. */
116 finding: (finding: Finding, how: 'address' | 'discuss' | 'typed', words = '') => {
117 const opening =
118 how === 'address'
119 ? 'Please address this finding you recorded:'
120 : how === 'discuss'
121 ? "Let's talk through this finding you recorded before changing anything:"
122 : 'About this finding you recorded:'
123
124 return [opening, ...findingBody(finding), ...(how === 'typed' ? ['', words] : [])].join('\n')
125 },
126}
127hooks/tools.ts 96 lines1// The record_finding and close tools the agent calls, the same in every host
2// except for where the texts say the person sees the result.
3
4import { addFinding, closeByAgent } from './ledger'
5import type { Host } from './ledger'
6import type { Ledger } from '../types'
7
8export const FINDING_SCHEMA = {
9 type: 'object',
10 properties: {
11 kind: {
12 type: 'string',
13 enum: ['issue', 'opportunity'],
14 description: 'issue: something wrong or risky. opportunity: something that could be better.',
15 },
16 title: { type: 'string', description: 'What it is, in at most 12 plain words.' },
17 detail: { type: 'string', description: 'Why it matters and what you would do, in one or two sentences.' },
18 path: { type: 'string', description: 'The file it is about, if one.' },
19 },
20 required: ['kind', 'title', 'detail'],
21}
22
23export const CLOSE_SCHEMA = {
24 type: 'object',
25 properties: {
26 id: { type: 'string', description: 'The id of the open item or finding, such as i35 or f32.' },
27 answer: {
28 type: 'string',
29 description: "The user's answer, in their words and at most 8, when their own message answered it.",
30 },
31 reason: {
32 type: 'string',
33 description: 'Otherwise, why it is closed, in at most 8 words, such as "no longer applies: Inbox kept".',
34 },
35 },
36 required: ['id'],
37}
38
39export function findingDescription(host: Host): string {
40 return `Record a finding for the user. It waits in ${host.findingsIn} until it is closed, and from there the user can ask you to address it or discuss it. Record what a careful senior engineer would flag to a teammate, and leave out style nits and anything the user already decided.`
41}
42
43export const CLOSE_DESCRIPTION = `Close an open item or finding by its id, such as i35 or f12, as listed in the latest "inbox:" text beside the user's prompt. Pass the user's answer when their message answered it, and a reason otherwise.`
44
45function cut(value: unknown, max: number): string {
46 return typeof value === 'string' ? value.trim().slice(0, max) : ''
47}
48
49/** Adds the finding a record_finding call describes. `result` is the text the agent reads back. */
50export function recordFinding(
51 host: Host,
52 ledger: Ledger,
53 input: Record<string, unknown>,
54 now: number,
55): { ledger: Ledger; result: string } {
56 const title = cut(input.title, 120)
57 const detail = cut(input.detail, 600)
58 if (title === '' || detail === '') return { ledger, result: 'Not recorded: a finding needs a title and a detail.' }
59 const path = cut(input.path, 300)
60 const r = addFinding(ledger, {
61 kind: input.kind === 'opportunity' ? 'opportunity' : 'issue',
62 title,
63 detail,
64 path: path || null,
65 at: now,
66 })
67
68 return {
69 ledger: r.ledger,
70 result: r.isAdded ? `Recorded as ${r.id}. The user sees it in ${host.findingsIn}.` : `Already recorded as ${r.id}.`,
71 }
72}
73
74/** Closes the item or finding a close call names. `result` is the text the agent reads back. */
75export function recordClose(
76 host: Host,
77 ledger: Ledger,
78 input: Record<string, unknown>,
79 now: number,
80): { ledger: Ledger; result: string } {
81 const id = cut(input.id, 80).replace(/^\[|\]$/g, '')
82 const answer = cut(input.answer, 80)
83 const reason = cut(input.reason, 80)
84 if (id === '' || (answer === '' && reason === ''))
85 return { ledger, result: "Not closed: give the id, and the user's answer or a reason." }
86 const r = closeByAgent(host, ledger, id, answer ? { answer } : { reason }, now)
87 if (r.closed === 'item')
88 return { ledger: r.ledger, result: `Closed ${id}. The user sees it in ${host.surface} with its outcome.` }
89 if (r.closed === 'finding') return { ledger: r.ledger, result: `Closed finding ${id}.` }
90
91 return {
92 ledger,
93 result: `Not closed: no open item or finding has the id ${id}. The open ones are listed beside the user's latest message.`,
94 }
95}
96hooks/tree.ts 113 lines1// Reads a repo's working tree with git, to tell when a check's result went
2// stale. Each app passes in how it runs a command and reads a path's kind.
3
4import { candidates, changedPaths, readChanged, readLsTree, sameSnapshot } from './git'
5import type { Snapshot } from '../types'
6
7// A dirtier tree is read only this far, so changes past it go unseen.
8const SNAPSHOT_MAX = 2000
9
10export type TreeIO = {
11 /** Runs a command in `cwd`: its stdout when it exits 0, otherwise null. Never throws. */
12 run: (args: string[], opts: { cwd: string; stdin?: string; timeoutMs: number }) => Promise<string | null>
13 /** What a path leads to; 'other' when it is neither a file nor a folder, or cannot be read. */
14 kindOf: (path: string) => Promise<'file' | 'dir' | 'other'>
15}
16
17/** Runs a git command that reads a repo's working tree; null when it fails. Optional locks are off, so it never takes the index lock from a commit. */
18function git(io: TreeIO, repo: string, args: string[], stdin?: string): Promise<string | null> {
19 return io.run(['git', '--no-optional-locks', ...args], {
20 cwd: repo,
21 timeoutMs: 15_000,
22 ...(stdin === undefined ? {} : { stdin }),
23 })
24}
25
26/** Each file's blob id, as a commit would store it. A folder, such as a submodule's, gets a mark of its own. */
27async function hashFiles(io: TreeIO, repo: string, paths: string[]): Promise<string[] | null> {
28 const out = await git(io, repo, ['hash-object', '--stdin-paths'], `${paths.join('\n')}\n`)
29 if (out !== null) return out.split('\n')
30 // One path git cannot hash fails the whole call, so hash the files alone.
31 const kinds = await Promise.all(paths.map(p => io.kindOf(`${repo}/${p}`)))
32 const files = paths.filter((_p, i) => kinds[i] === 'file')
33 const hashed = files.length > 0 ? await git(io, repo, ['hash-object', '--stdin-paths'], `${files.join('\n')}\n`) : ''
34 if (hashed === null) return null
35 const ids = hashed.split('\n')
36
37 return paths.map((p, i) => (kinds[i] === 'file' ? (ids[files.indexOf(p)] ?? '') : `${kinds[i]}`))
38}
39
40/** A repo's working tree content, read without writing to the repo. */
41async function readSnapshot(io: TreeIO, repo: string): Promise<Snapshot | null> {
42 const [out, headOut] = await Promise.all([
43 git(io, repo, ['status', '--porcelain=v1', '-z', '-uall']),
44 git(io, repo, ['rev-parse', '--verify', '-q', 'HEAD']),
45 ])
46 if (out === null) return null
47 const head = headOut?.trim() || null
48 const changes = readChanged(out).slice(0, SNAPSHOT_MAX)
49 const dirty: Snapshot['dirty'] = {}
50 for (const c of changes) if (c.isDeleted) dirty[c.path] = ''
51 const present = changes.filter(c => !c.isDeleted && !c.path.includes('\n')).map(c => c.path)
52 if (present.length > 0) {
53 const ids = await hashFiles(io, repo, present)
54 if (!ids) return null
55 present.forEach((path, i) => {
56 dirty[path] = ids[i] ?? ''
57 })
58 }
59
60 return { head, dirty }
61}
62
63/** Each path's object id in a commit, for the paths it has. */
64async function lsTree(io: TreeIO, repo: string, head: string, paths: string[]): Promise<Record<string, string> | null> {
65 const ids: Record<string, string> = {}
66 for (let i = 0; i < paths.length; i += 200) {
67 const out = await git(io, repo, ['ls-tree', '-z', '--full-tree', head, '--', ...paths.slice(i, i + 200)])
68 if (out === null) return null
69 Object.assign(ids, readLsTree(out))
70 }
71
72 return ids
73}
74
75/** The paths whose content differs between two snapshots of a repo, or null when git could not tell. A commit only moves content into HEAD, so it changes nothing. */
76async function contentChanges(io: TreeIO, repo: string, a: Snapshot, b: Snapshot): Promise<string[] | null> {
77 if (a.head !== b.head && !b.head) return null
78 let committed: string[] = []
79 if (a.head && b.head && a.head !== b.head) {
80 const out = await git(io, repo, ['diff', '--name-only', '-z', '--no-renames', a.head, b.head])
81 if (out === null) return null
82 committed = out.split('\0').filter(Boolean)
83 }
84 const paths = candidates(a, b, committed)
85 if (paths.length === 0) return []
86 const none = Promise.resolve<Record<string, string>>({})
87 const reading = a.head ? lsTree(io, repo, a.head, paths) : none
88 const [before, after] = await Promise.all([
89 reading,
90 b.head === a.head ? reading : b.head ? lsTree(io, repo, b.head, paths) : none,
91 ])
92 if (!before || !after) return null
93
94 return changedPaths(a, b, paths, before, after)
95}
96
97/**
98 * Reads a repo's working tree and compares it with its last reading. Returns
99 * null when it could not be read or did not change. Otherwise returns the new
100 * reading and the paths whose content changed: none on a first reading, and
101 * null when git could not tell, which counts as a change to code.
102 */
103export async function readRepo(
104 io: TreeIO,
105 repo: string,
106 last: Snapshot | undefined,
107): Promise<{ snapshot: Snapshot; changes: string[] | null } | null> {
108 const snapshot = await readSnapshot(io, repo)
109 if (!snapshot || (last && sameSnapshot(last, snapshot))) return null
110
111 return { snapshot, changes: last ? await contentChanges(io, repo, last, snapshot) : [] }
112}
113hooks/git.ts 67 lines1// Reading the working tree's content from git output, so a check can tell
2// whether the files changed after it ran. Parsing is pure; register.tsx runs
3// the commands, none of which write to the repo.
4
5import type { Snapshot } from '../types'
6
7/** Reads `git status --porcelain=v1 -z -uall`: each path whose content differs from HEAD. A rename's old path is a deleted one. */
8export function readChanged(stdout: string): { path: string; isDeleted: boolean }[] {
9 const fields = stdout.split('\0')
10 const changed: { path: string; isDeleted: boolean }[] = []
11 for (let i = 0; i < fields.length; i++) {
12 const field = fields[i] ?? ''
13 if (field === '') continue
14 const x = field[0]
15 const y = field[1]
16 changed.push({ path: field.slice(3), isDeleted: y === 'D' || (x === 'D' && y === ' ') })
17 // A rename or copy is followed by its source path, which a rename leaves absent.
18 if (x === 'R' || x === 'C') {
19 i++
20 const source = fields[i]
21 if (x === 'R' && source) changed.push({ path: source, isDeleted: true })
22 }
23 }
24
25 return changed
26}
27
28/** Reads `git ls-tree -z <tree> -- <paths>`: each path's object id. */
29export function readLsTree(stdout: string): Record<string, string> {
30 const ids: Record<string, string> = {}
31 for (const entry of stdout.split('\0')) {
32 const tab = entry.indexOf('\t')
33 if (tab < 0) continue
34 const id = entry.slice(0, tab).split(' ')[2]
35 if (id) ids[entry.slice(tab + 1)] = id
36 }
37
38 return ids
39}
40
41export function sameSnapshot(a: Snapshot, b: Snapshot): boolean {
42 const keys = Object.keys(a.dirty)
43
44 return a.head === b.head && keys.length === Object.keys(b.dirty).length && keys.every(k => b.dirty[k] === a.dirty[k])
45}
46
47/** The paths whose content can differ between two snapshots, given the paths the commits between their heads changed. */
48export function candidates(a: Snapshot, b: Snapshot, committed: string[]): string[] {
49 return [...new Set([...Object.keys(a.dirty), ...Object.keys(b.dirty), ...committed])]
50}
51
52/**
53 * The paths whose content differs between two snapshots, given each candidate
54 * path's object id in each one's HEAD. A path's content is its uncommitted
55 * hash, else its id in HEAD; a missing path counts as empty either way. So a
56 * commit, which moves content from uncommitted to HEAD, changes nothing.
57 */
58export function changedPaths(
59 a: Snapshot,
60 b: Snapshot,
61 paths: string[],
62 before: Record<string, string>,
63 after: Record<string, string>,
64): string[] {
65 return paths.filter(p => (a.dirty[p] ?? before[p] ?? '') !== (b.dirty[p] ?? after[p] ?? ''))
66}
67types/index.d.ts 305 lines1/** Where the session stands, as the ledger model last summarized it. */
2export type Card = {
3 goal: string
4 done: string[]
5 now: string
6 running: string[]
7 updatedAt: number
8}
9
10/** Something the agent put to the person that is still unanswered. */
11export type Item = {
12 id: string
13 kind: 'question' | 'task'
14 /** The agent's own number or id for it ("1", "D3"), so "1. yes" maps back. */
15 label: string | null
16 ask: string
17 options: string[]
18 /** The option the agent recommended, exactly as `options` has it. */
19 rec: string | null
20 /** Steps the agent's reply spelled out, one press each. */
21 helps: Help[]
22 /** The person-prompt count when the agent asked it. */
23 turn: number
24 /** When the agent asked it; null for an item saved before the mod kept the time. */
25 at: number | null
26}
27
28/**
29 * A one-press step toward finishing an item. The mod never runs a command:
30 * `run` asks Claude to run it, under the usual permission checks, and
31 * `terminal` copies one that only the person can run.
32 */
33export type Help =
34 | { kind: 'open'; path: string }
35 | { kind: 'copy'; text: string; name: string | null }
36 | { kind: 'run'; command: string; name: string | null }
37 | { kind: 'terminal'; command: string; name: string | null }
38 | { kind: 'link'; url: string; name: string | null }
39
40/** An item that closed, and how. */
41export type Closed = {
42 id: string
43 kind: Item['kind']
44 ask: string
45 /** The reply's own number for it, when it had one. */
46 label?: string
47 /** What the pane shows for how it closed, such as the person's answer. */
48 outcome: string
49 /**
50 * How it closed: the person answered it, marked a task done or ran its
51 * command, or dismissed it; it expired unanswered; Claude closed it with a
52 * reason; or the per-reply update closed it, with an outcome the model wrote.
53 */
54 how: 'answered' | 'done' | 'dismissed' | 'expired' | 'claude' | 'update'
55 at: number
56}
57
58/** A failing check's row that left the list: a passing run cleared it, or Fix handed it to Claude. */
59export type LeftCheck = { id: string; kind: 'check'; title: string; outcome: 'Passed' | 'Fix'; at: number }
60
61/** A row that just closed or left, kept where it stood in its group for a few seconds. */
62export type Settled = (Closed | LeftCheck) & { index: number }
63
64/** The row a jump moved the pane to, so the tab and row can show where it went. */
65export type Arrival = { tab: Tab; id: string | null; at: number }
66
67/** What a row's last action did, shown on the row so the person sees the press went through. */
68export type LastAction = {
69 /** The action's key, so that action reads "… again". */
70 action: string
71 /** The label of the action pressed, as in "Discuss". The row shows it after a ✓. */
72 text: string
73 at: number
74 /** The press handed the row's work to Claude, so the row folds and stops waiting on the person. */
75 isHandoff?: boolean
76 /** The turns started when it was pressed, so a task folds until the update for a later turn applies. */
77 turnsStarted?: number
78 /** Where the row was, and its title, for a row the press removed: it shows its last action in its place for a few seconds. */
79 tab: Tab
80 title: string
81 index: number
82}
83
84/** Something Claude noticed outside the current task and recorded for the person. */
85export type Finding = {
86 id: string
87 kind: 'issue' | 'opportunity'
88 title: string
89 detail: string
90 /** The file it is about, as Claude gave it. */
91 path: string | null
92 at: number
93}
94
95export type PrCheck = { name: string; bucket: 'pass' | 'fail' | 'pending' | 'skip'; url: string | null }
96
97/** An unresolved review thread, told by its first comment. */
98export type PrThread = {
99 id: string
100 author: string
101 /** The thread's latest comment, when anyone answered the first. Its `at` is as the thread's. */
102 reply: { author: string; body: string; url: string; at: number | null } | null
103 /** Someone other than the viewer wrote its last comment, so it waits on them. */
104 isWaiting: boolean
105 isOutdated: boolean
106 /**
107 * Outdated, and no one else replied after the PR's latest commit: a later
108 * commit changed the lines it was on, and nothing since says that left it unfixed.
109 */
110 isLinesChanged: boolean
111 path: string
112 line: number | null
113 body: string
114 replies: number
115 url: string
116 /** When its first comment was written; null when GitHub gave no time, or for a thread saved before the mod kept it. */
117 at: number | null
118}
119
120/** A PR as gh last reported it. */
121export type PrView = {
122 /** "owner/repo#123" */
123 ref: string
124 number: number
125 title: string
126 url: string
127 isDraft: boolean
128 state: string
129 base: string
130 mergeable: string
131 reviewDecision: string
132 checks: PrCheck[]
133 threads: PrThread[]
134 fetchedAt: number
135 /** Why the last fetch failed; the rest is from the fetch before. */
136 error: string | null
137}
138
139export type Ledger = {
140 card: Card | null
141 items: Item[]
142 closed: Closed[]
143 findings: Finding[]
144 /** PRs this session created or linked, as "owner/repo#123". */
145 prs: string[]
146 nextId: number
147 /** How many prompts the person has sent this session. */
148 turn: number
149 /** The turn whose reply added the newest items; 0 when none. */
150 batchTurn: number
151}
152
153/** Why the session halted outside the conversation, from the error Claude Code classified it with. */
154export type Stop = {
155 kind: 'sign-in' | 'billing' | 'usage-limit' | 'api-error'
156 /** Claude Code's word for the error, such as `authentication_failed`. */
157 detail: string
158 /** When a usage limit resets, as Claude Code said it: "7:33pm". */
159 resets: string | null
160 at: number
161}
162
163/** A permission prompt or question dialog in the session that waits on the person. */
164export type Dialog = {
165 kind: 'permission' | 'question'
166 text: string
167 /** The tool call it is for, so it closes when that call ends. */
168 key: string
169}
170
171/** `all`: a script that runs the project's checks together, such as `check.sh`. */
172export type CheckKind = 'tests' | 'types' | 'lint' | 'build' | 'validate' | 'all'
173
174/**
175 * The part of a check's suite one run covered. Both lists empty means the
176 * whole suite.
177 */
178export type Target = {
179 /** Files and folders the run named, absolute. */
180 paths: string[]
181 /** Test-name filters and any argument the inbox could not read, such as `-t=parses dates`. */
182 filters: string[]
183}
184
185/** The latest result of one check command Claude ran, such as `npm test`, in one folder, for one target. */
186export type Check = {
187 name: string
188 kind: CheckKind
189 /** The part of the suite the run covered; empty for the whole suite. */
190 target: Target
191 /** The folder it ran in, absolute; null for the session's own folder. */
192 folder: string | null
193 result: 'pass' | 'fail'
194 /** The output's summary line, such as "24 pass, 1 fail". */
195 summary: string
196 ranAt: number
197 /** The top folder of the git repo it ran in, whose edits make it stale; null outside git. */
198 repo: string | null
199 /**
200 * Files it reads changed after it ran: any file for a lint, a validation or
201 * a check script, and a file other than Markdown for the rest.
202 */
203 isStale: boolean
204 /** The piece of Claude's command that ran it, as written, such as `npm test > t.log`. */
205 command: string
206 /** Up to three output lines that name what failed, when the command ran no other check. */
207 failures: string[]
208 /** A turn of Claude's ended with it failing, so the Needs you tab lists it. */
209 isLeftFailing: boolean
210 /** The person dismissed its row in the Needs you tab. */
211 isDismissed: boolean
212 /** The Stop hook already sent Claude back over this result, so it does not again. */
213 isSentBack: boolean
214 /** When the person pressed Fix on its row, so it no longer waits on them; null before. The next run replaces the result. */
215 fixSentAt: number | null
216}
217
218/**
219 * The working tree's content, read without writing to the repo: HEAD, and the
220 * blob id of each path that differs from it ('' for a deleted path).
221 */
222export type Snapshot = { head: string | null; dirty: Record<string, string> }
223
224/** The checks Claude ran. */
225export type Checks = {
226 results: Check[]
227 /** The session's repos: the one it started in, and each repo where Claude edited a file. Only their results count. */
228 repos: string[]
229}
230
231export type Presence = {
232 lastActiveAt: number
233 isAway: boolean
234 isUpdating: boolean
235 /**
236 * `behind` and `failed`: the ledger may have missed a turn, so the next
237 * update re-reads the whole conversation. Only `failed` shows in the pane.
238 */
239 ledgerState: 'current' | 'behind' | 'failed'
240 /**
241 * The turns started in this process, and how many of them reached the
242 * ledger. A load that finds more started than applied catches up.
243 */
244 turnsStarted: number
245 turnsApplied: number
246 /** The minute of the last clock tick, so the "last active" text redraws. */
247 minute: number
248}
249
250/** The most recent other session in this project, offered on a fresh start. */
251export type Previous = {
252 savedAt: number
253 ledger: Ledger
254 isBroughtIn: boolean
255}
256
257export type Tab = 'needsYou' | 'findings' | 'prs'
258
259/** A tab's selected row: its id, and its position for when that row goes away. */
260export type Cursor = { id: string | null; index: number }
261
262/** The PR tab's data: each PR's latest view, the current branch's PR, and whether a fetch runs. */
263export type PrViews = { views: Record<string, PrView>; branchRef: string | null; isFetching: boolean }
264
265/** A failing PR check the person pressed Fix on: when, and the check's URL then. A rerun has a new URL, so the mark lapses. */
266export type PrFixSent = { at: number; url: string | null }
267
268declare module 'claude-code' {
269 interface PluginState {
270 inbox: {
271 /** The Claude Code theme in use, such as `dark` or `light-ansi`; '' until read. */
272 theme: string
273 ledger: Ledger
274 presence: Presence
275 previous: Previous | null
276 tab: Tab
277 prViews: PrViews
278 /** The PR checks the person pressed Fix on, by their row's id. */
279 prFixesSent: Record<string, PrFixSent>
280 selection: Record<Tab, Cursor>
281 stop: Stop | null
282 dialogs: Dialog[]
283 checks: Checks
284 /** Each checked repo's working tree as last read, by its top folder, kept apart from `checks` because no drawing reads it. */
285 snapshots: Record<string, Snapshot>
286 /** The row whose free-text field is open, if any. */
287 typing: string | null
288 /** Items that just closed, and checks that just passed, shown in place for a few seconds. */
289 settled: Settled[]
290 /** The pane's latest jump to a new row. */
291 arrival: Arrival | null
292 /** Each row's last action, by the row's id. */
293 lastActions: Record<string, LastAction>
294 /** The Needs you groups whose closed items show. Each starts folded. */
295 unfolded: Item['kind'][]
296 /** The rows whose folded details show, by the row's id. */
297 shownDetails: string[]
298 /** The pane's list of the keys no row shows is unfolded. */
299 isKeyListShown: boolean
300 /** The band and pane show sample entries instead of the session's own, for `/inbox demo`. */
301 isDemo: boolean
302 }
303 }
304}
305