The agents kit inside Claude Code: the /flow pane

One setup for every coding agent you use (Claude Code, Codex): the same instructions, skills, guardrails and checks, whichever harness runs the work. Agents are told to work like a staff engineer, and hooks check that they did: irreversible and outward-facing commands stop and wait for you, every turn that edits code is verified, and a pull request can't be opened until the exact change has been self-reviewed. You can switch any of those phases off, and back on, for a branch, a repo or everywhere: phase off verify as the first line of a message, or /flow off verify in Claude Code. review-rounds 2 the same way sets how many re-checks the self-review runs after its first pass (1 by default).
docs/framework.md is the full reference: every hook, guard rule, check, skill, tool and scheduled job, with diagrams. It is generated from the code on every commit, so it describes what the kit does now.
macOS only for now (the scheduled jobs use launchd).
git clone https://github.com/alopezari/agents-kit.git ~/.agents && ~/.agents/install.sh
install.sh wires every installed harness and is safe to re-run. To check the setup afterwards:
~/.agents/install.sh --doctor # report what's missing, change nothing
~/.agents/tests/run.sh # regression suite (passes once --doctor is clean)
The kit must live at ~/.agents. The installer:
deps.txt through Homebrew, asking once (--yes skips the question; without a terminal it only reports them);AGENTS.md as each harness's global instructions and the skills into each harness;skills.external, a pinned one at its commit into vendor/, and installs the Node dependencies of tools/ and site/;docs/framework.md current);/flow mod (see Mods contract);bin/gh into /usr/local/bin (asks for your password).It also adds the kit's baseline harness settings (no fast mode, effort defaults) wherever a key is missing, without overwriting one you set. Files it replaces are backed up under backups/.
To remove it, ~/.agents/uninstall.sh lists every link, hook, status line, scheduled job and the kit's Claude Code plugin that still point at the kit, asks, and removes them (--yes skips the question). Your own hooks and settings stay, and so do the baseline settings and the programs installed through Homebrew. It backs up the settings files it edits to ~/.agents-uninstall-backups/, and leaves ~/.agents itself for you to delete.
What it doesn't do, on a new machine:
claude, codex) and gh. Install them before running install.sh, which only wires the harnesses it finds.install.sh --doctor warns until you do.Requirements: deps.txt lists every program the kit runs. install.sh offers to install the missing required ones (python3, git, jq, node, gh) and recommended ones (semgrep, gitleaks, php) through Homebrew, and says how to install the optional ones, each needed by one feature. install.sh --doctor reports what is missing, and tests/run.sh expects it to report nothing.
AGENTS.md. It is plain Markdown and every harness reads the same file.repos/<checkout-directory-name>/notes.md with the context agents should read there. Add an executable verify when the automatic checks aren't enough. Nothing is written into the repository itself. See repos/README.md.install.sh --profile <dir>. It can hold repo overlays, work-only skills, research, MCP write rules and the repositories to learn from; examples/sample-profile/ has one of each to start from. The core stays generic and shareable.Every hook is a small program with one contract: a JSON payload on stdin, a JSON decision on stdout. Claude Code and Codex share this contract natively. Other harnesses need an adapter that translates their events into it.
Payload fields used: tool_name, tool_input.command, tool_input.file_path (or notebook_path, or the file list in Codex's apply_patch input), cwd, session_id, stop_hook_active, and turn_id to tell Codex apart in the log.
Decisions:
{"hookSpecificOutput": {"permissionDecision": "deny", "permissionDecisionReason": ...}}.{"decision": "block", "reason": ...}, optionally with systemMessage for the user.Claude Code also loads the kit as a plugin, mods/kit/, whose mod runs inside Claude Code: it can draw panes, run commands without a turn and define agent types. Mods exist only in Claude Code, so the kit keeps working the same with or without them. Every mod follows these rules:
.git/agents/, or a temporary directory when .git can't be written) and go through bin/ and hooks/review_stamp.py, so a change started in Claude Code can be finished in Codex, and the other way round. What those tools cache or migrate follows their own rules.mods/kit/hooks/features.js, which docs/framework.md prints, and the mod's tests fail when the commands or agents it registers, or the buttons it draws, differ from the list.claude --safe-mode turns off the kit's settings hooks too, so it isn't a supported way to run the kit.Its /flow pane shows where the branch is in the flow, with the phases switched off, and runs verify without a turn; /flow off <phase> [--repo|--global] and /flow on … switch a phase as the message line does. While the branch waits on its staging guide, the pane lists each step before the merge with a Pass and a Fail button that record your verdict; from a terminal, ~/.agents/bin/staging mark <step> PASS --by user does the same. When a guard blocks an MCP write or a DROP/TRUNCATE command, the mod asks you in a dialog that shows the call; allowing it, for that call or until your next message, writes the approval your message would have (bin/approve). docs/framework.md lists every capability with its Codex equivalent.
install.sh registers ~/.agents as the agents-kit plugin marketplace and installs kit@agents-kit. The plugin loads in place, so a pull reaches it at the next session start or /reload-plugins. To work on the mod, run claude --plugin-dir ~/.agents/mods/kit, which reloads it on save; tests/run.sh mods validates and tests it.
AGENTS.md, and its skills directory at skills/ if it has one.adapters/<harness>/ to translate its events into the payloads above:guard_bash.py;post_edit.py;stop_checks.py, with a guard so the stop check continues the agent only once per user turn.install.sh and run install.sh --doctor. Then test with a harmless blocked command in a scratch repo.bin/docs where the new harness registers its hooks (HARNESS_OF_SETTINGS and harness_hooks), then commit: the pre-commit hook regenerates docs/framework.md.tests/ feeds the hooks hand-made payloads. To see a branch's hooks run inside Codex itself, with its real payloads, point HOME at a directory whose .agents is the branch's checkout. Codex's hooks.json runs python3 $HOME/.agents/hooks/<hook>.py, and it trusts a hook by that command text, so the branch's code runs through the entries you already approved. CODEX_HOME keeps Codex's own settings, login and trust, and GH_CONFIG_DIR keeps gh logged in; they come before HOME= because bash and zsh expand ~ with the HOME assigned before it:
H=$(mktemp -d) && ln -s ~/.agents-worktree-<name> "$H/.agents" # the branch's checkout
git init -q /tmp/hook-probe && git -C /tmp/hook-probe commit -q --allow-empty -m init
echo 'Run this shell command once and report what happened: <a command the change should block or allow>' \
| CODEX_HOME=~/.codex GH_CONFIG_DIR=~/.config/gh HOME="$H" \
codex exec -C /tmp/hook-probe --skip-git-repo-check -s workspace-write --ephemeral -
tail -3 ~/.agents-worktree-<name>/logs/hooks.jsonl
The hooks log to the branch checkout's logs/ (not in git), so the last lines show each decision with "harness": "codex". The checkout has no repos/<repo>/verify overlay, so the stop hook falls back to the automatic verify. Use a scratch repo: a hook that fails to block lets the command run.
logs/, backups/, research/, approvals/, monitors/state/, review-mining/runs/, review-mining/baseline.json and usage/*.json hold local, possibly private data; site/dist/ is the built website. Profiles are linked in and never committed here.
MIT, see LICENSE.
hooks/register.js 608 lines1// The kit inside Claude Code. Everything here shows or runs the kit's own tools: the state stays with them
2// and the decisions in the Python hooks, so Codex and a session without this mod get the same flow.
3import { FEATURES } from './features.js'
4
5const PANE = 'flow'
6const PROCESS_TIMEOUT_MS = 30_000
7// phase_switches.py's exit for a switch it refused; 1 is Python's own for a crash.
8const SWITCH_REFUSED = 3
9// ci-wait --once makes one GitHub round trip per check source, each bounded at 30 s.
10const CI_TIMEOUT_MS = 90_000
11const STAMPS = [['verify', 'verify'], ['self-review', 'review'], ['validate', 'validate'], ['staging', 'staging']]
12const CI_OUTCOMES = { 0: 'passed', 1: 'failed', 2: 'running', 3: 'no checks', 4: 'unreadable' }
13const REPORT_LINE = /^(ran|skipped|warning|error):/
14const VERIFY_TAIL_LINES = 8
15// Claude Code refuses a Text string over 10,000 characters; a verify line can be any length.
16const LINE_CHARS = 500
17// stop_checks.py bounds the repo's verify at 600 s, but not its own git reads and stamp writes around it.
18const VERIFY_DEADLINE_MS = 660_000
19// What `stop_checks.py verify` prints when the run passed without checking anything.
20const CHECKED_NOTHING = 'verify passed, but it checked nothing'
21// A gathering whose branch or HEAD moved while it ran is gathered again, this many times in all.
22const GATHER_ATTEMPTS = 2
23const NEEDS_ATTENTION = /couldn't read|^Couldn't|^CI: failed|^Run verify.* failed/
24
25// Each self-review lens becomes the agent type `kit:review-<key>`, briefed from what `bin/triage --lens-briefs` cuts
26// out of lenses.md. No edit tools is a convenience for the reviewer, not a guard (Bash can still write): the hooks
27// run for subagents too.
28const REVIEWER_TOOLS = ['Read', 'Grep', 'Glob', 'Bash']
29// Registering takes tens of milliseconds and must land before the first turn (`claude -p` starts one at once), so the
30// session start waits for it, but never longer than this.
31const REVIEWERS_WAIT_MS = 2_000
32
33// What the pane draws. One gathering runs at a time: a request during one gathers again once it ends.
34let shown = null
35let gathering = false
36let gatherAgain = false
37let verifyRun = { state: 'idle' }
38// A verify stopped at its deadline holds the button until its child is gone, so two never write one report.
39let verifyStopping = false
40// Marks run one at a time, in the order pressed, each shown as pending until it lands.
41let marking = Promise.resolve()
42let pendingMarks = []
43// { root, text }: shown while the pane shows that checkout, whatever its branch now: a refusal says the branch moved.
44let markFailure = null
45// A step's title is cut to this, so its result and evidence stay on the line.
46const TITLE_CHARS = 60
47
48async function run($, argv, cwd, timeoutMs = PROCESS_TIMEOUT_MS) {
49 try {
50 return await $.process.run(argv, { cwd, timeoutMs })
51 } catch (error) {
52 return { failure: messageOf(error) }
53 }
54}
55
56function messageOf(error) {
57 return String(error?.message ?? error)
58}
59
60function failureOf(result) {
61 if (result.failure) return result.failure
62 return 'exit ' + result.exitCode + (result.stderr?.trim() ? ': ' + result.stderr.trim().split('\n')[0] : '')
63}
64
65async function kitDir($) {
66 return (await $.env.get('HOME')) + '/.agents'
67}
68
69async function stampLine($, kit, root, label, kind) {
70 const check = (stampKind) => run($, ['python3', kit + '/hooks/review_stamp.py', 'check', '--kind', stampKind], root)
71 const result = await check(kind)
72 if (result.exitCode === 0) return label + ': current'
73 if (result.exitCode !== 1) return label + ": couldn't read: " + failureOf(result)
74 if (kind === 'verify') {
75 const empty = await check('verify-empty')
76 if (empty.exitCode === 0) return label + ': current, but it checked nothing'
77 if (empty.exitCode !== 1) return label + ": couldn't read: " + failureOf(empty)
78 }
79 return label + ': not current'
80}
81
82async function verifyReportLines($, kit, root) {
83 const path = await run($, [kit + '/bin/reports', 'path', 'verify'], root)
84 if (path.failure || path.exitCode !== 0) return ["couldn't read: " + failureOf(path)]
85 const file = path.stdout.trim()
86 try {
87 if (!(await $.fs.exists(file))) return ['no verify run yet']
88 const { mtimeMs } = await $.fs.stat(file)
89 const lines = (await $.fs.read(file)).split('\n')
90 return [lines[0].replace(/^# /, '') + ' · ' + dateTimeOf(mtimeMs), ...lines.filter((line) => REPORT_LINE.test(line))]
91 } catch (error) {
92 return ["couldn't read: " + messageOf(error)]
93 }
94}
95
96// { steps } for the guide's steps before the merge, or { failure } when bin/staging can't list them.
97async function stagingSteps($, kit, root) {
98 const listed = await run($, [kit + '/bin/staging', 'steps', '--json'], root)
99 if (listed.failure || listed.exitCode !== 0) return { failure: "couldn't read: " + failureOf(listed) }
100 try {
101 return { steps: JSON.parse(listed.stdout) }
102 } catch (error) {
103 return { failure: "couldn't read: bin/staging printed something not JSON: " + messageOf(error) }
104 }
105}
106
107function stepLine(step) {
108 const title = step.title.length > TITLE_CHARS ? step.title.slice(0, TITLE_CHARS - 1) + '…' : step.title
109 const result = step.result ? step.result + (step.by ? ' (' + step.by + ')' : '') : 'not marked'
110 const evidence = step.evidence.length ? step.evidence.map((path) => path.split('/').pop()).join(', ') : 'no evidence yet'
111 return [step.id + ' ' + title, result, evidence].join(' · ')
112}
113
114function markStep($, id, result) {
115 if (!shown?.root) return marking
116 // What the pressed button showed: a gathering can move the pane before this mark's turn comes.
117 const { root, branch } = shown
118 const label = id + ' ' + (result === 'PASS' ? 'Pass' : 'Fail')
119 pendingMarks.push(label)
120 $.ui.invalidate('ui.render')
121 marking = marking.then(async () => {
122 try {
123 const kit = await kitDir($)
124 // --branch: the checkout itself may have moved since the pane showed this step.
125 const marked = await run($, [kit + '/bin/staging', 'mark', id, result, '--by', 'user', '--branch', branch], root)
126 if (marked.failure || marked.exitCode !== 0) markFailure = { root, text: `Couldn't mark ${label}: ` + failureOf(marked) }
127 else if (markFailure?.root === root) markFailure = null
128 } finally {
129 pendingMarks.splice(pendingMarks.indexOf(label), 1)
130 }
131 await gather($)
132 }).catch((error) => {
133 // A broken chain would leave every later press doing nothing.
134 markFailure = { root, text: `Marked ${label}, or not: the pane failed while it ran: ` + messageOf(error) }
135 $.ui.invalidate('ui.render')
136 })
137 return marking
138}
139
140async function ciLine($, kit, root, head) {
141 const upstream = await run($, ['git', 'rev-parse', '--verify', '--quiet', '@{u}'], root)
142 if (upstream.exitCode === 1) return 'CI: HEAD not pushed (no upstream)'
143 if (upstream.failure || upstream.exitCode !== 0) return "CI: couldn't read: " + failureOf(upstream)
144 const pushed = await run($, ['git', 'merge-base', '--is-ancestor', head, '@{u}'], root)
145 if (pushed.failure || pushed.exitCode > 1) return "CI: couldn't read: " + failureOf(pushed)
146 if (pushed.exitCode === 1) return 'CI: HEAD not pushed'
147 const ci = await run($, [kit + '/bin/ci-wait', '--sha', head, '--once', '--no-log'], root, CI_TIMEOUT_MS)
148 const outcome = ci.failure ? undefined : CI_OUTCOMES[ci.exitCode]
149 if (!outcome) return "CI: couldn't read: " + failureOf(ci)
150 const reason = ci.exitCode === 4 ? (ci.stderr.trim() || ci.stdout.trim()).split('\n')[0] : ''
151 const line = 'CI: ' + outcome + (reason ? ': ' + reason : '')
152 const status = await run($, ['git', 'status', '--porcelain'], root)
153 if (status.failure || status.exitCode !== 0) {
154 return line + ", for " + head.slice(0, 7) + "; couldn't tell whether there are uncommitted changes: " + failureOf(status)
155 }
156 return line + (status.stdout.trim() !== '' ? ', for ' + head.slice(0, 7) + ' without the uncommitted changes' : '')
157}
158
159async function contextLine($) {
160 try {
161 const { context } = await $.session.usage()
162 return 'context: ' + (context?.percent === undefined ? '–' : Math.round(context.percent) + '%')
163 } catch {
164 return 'context: –'
165 }
166}
167
168function timeOf(ms) {
169 return new Date(ms).toTimeString().slice(0, 8)
170}
171
172function dateTimeOf(ms) {
173 const date = new Date(ms)
174 const pad = (n) => String(n).padStart(2, '0')
175 return date.getFullYear() + '-' + pad(date.getMonth() + 1) + '-' + pad(date.getDate()) + ' ' + timeOf(ms)
176}
177
178// The checkout at `cwd`: { root, branch, head, inMain? }, or { none } saying why there is no change to follow.
179async function checkout($, cwd) {
180 const top = await run($, ['git', 'rev-parse', '--show-toplevel'], cwd)
181 if (top.failure) return { none: "Couldn't read the repository: " + top.failure }
182 // git exits 128 for "not a git repository", and for a refused or broken one too.
183 if (top.exitCode !== 0 && /not a git repository/.test(top.stderr)) return { none: 'Not in a git repository: no change to follow.' }
184 if (top.exitCode !== 0) return { none: "Couldn't read the repository: " + failureOf(top) }
185 const at = await branchAt($, top.stdout.trim())
186 if (at.none || at.branch) return at
187 return (await mainCheckout($, at.root)) ?? { none: 'Detached HEAD: no change to follow.' }
188}
189
190async function branchAt($, root) {
191 const [branch, head] = await Promise.all([
192 run($, ['git', 'branch', '--show-current'], root),
193 run($, ['git', 'rev-parse', 'HEAD'], root),
194 ])
195 if (head.failure || head.exitCode !== 0) return { none: "Couldn't read HEAD: " + failureOf(head) }
196 if (branch.failure || branch.exitCode !== 0) return { none: "Couldn't read the branch: " + failureOf(branch) }
197 return { root, branch: branch.stdout.trim(), head: head.stdout.trim() }
198}
199
200// validate step 7 frees the branch by detaching the session's worktree, and the user checks it out in the main one.
201async function mainCheckout($, root) {
202 const listed = await run($, ['git', 'worktree', 'list', '--porcelain'], root)
203 if (listed.failure || listed.exitCode !== 0) return { none: "Detached HEAD, and couldn't find the main checkout: " + failureOf(listed) }
204 const first = listed.stdout.split('\n\n')[0]
205 const main = /^worktree (.+)$/m.exec(first)?.[1]
206 if (!main || main === root || /^bare$/m.test(first)) return null
207 const at = await branchAt($, main)
208 return at.none || at.branch ? { ...at, inMain: !at.none } : null
209}
210
211function clip(text) {
212 return text.length > LINE_CHARS ? text.slice(0, LINE_CHARS - 1) + '…' : text
213}
214
215function nameOf(at) {
216 return at.branch + ' @ ' + at.head.slice(0, 7) + (at.inMain ? ' in the main checkout' : '')
217}
218
219function sameCheckout(a, b) {
220 return !a.none && !b.none && a.root === b.root && a.branch === b.branch && a.head === b.head
221}
222
223// The facts for the change the session is on: { none } when there is no change to follow, otherwise a
224// heading naming the branch and HEAD they describe, and one list of lines per section.
225async function collect($) {
226 const kit = await kitDir($)
227 for (let attempt = 1; ; attempt++) {
228 const at = await checkout($, await $.session.cwd())
229 if (at.none) return at
230 // The brief comes first: on the default branch it is the only answer, and CI can take ninety seconds.
231 const brief = await run($, [kit + '/bin/reports', 'brief'], at.root)
232 let flow
233 if (brief.failure || brief.exitCode !== 0) flow = ["couldn't read: " + failureOf(brief)]
234 else {
235 flow = brief.stdout.split('\n').filter((line) => line && !/^\s/.test(line) && !line.startsWith('Read one in full'))
236 if (flow[0] === 'phase: ') {
237 return { none: at.inMain ? 'Detached HEAD, and the main checkout is on the default branch: no change to follow.' : 'On the default branch: no change to follow.' }
238 }
239 }
240 const [stamps, report, staging, ci, context, gatheredAt] = await Promise.all([
241 Promise.all(STAMPS.map(([label, kind]) => stampLine($, kit, at.root, label, kind))),
242 verifyReportLines($, kit, at.root),
243 stagingSteps($, kit, at.root),
244 ciLine($, kit, at.root, at.head),
245 contextLine($),
246 $.clock.now(),
247 ])
248 // Read again from the session's directory: Claude Code's /cd can move it while this gathers.
249 const after = await checkout($, await $.session.cwd())
250 if (after.none) return after
251 const moved = !sameCheckout(at, after)
252 if (moved && attempt < GATHER_ATTEMPTS) continue
253 return {
254 root: at.root,
255 branch: at.branch,
256 heading: nameOf(at) + ' · gathered ' + timeOf(gatheredAt) + (moved ? ' · the checkout moved while gathering: Refresh' : ''),
257 sections: [['Flow', flow], ['Stamps', stamps], ['Last verify report', report], ['CI and context', [ci, context]]],
258 staging,
259 }
260 }
261}
262
263async function collectOrSayWhy($) {
264 try {
265 return await withSwitches($, await collect($))
266 } catch (error) {
267 return { none: "Couldn't gather the flow: " + messageOf(error) }
268 }
269}
270
271async function gather($) {
272 if (gathering) {
273 gatherAgain = true
274 return
275 }
276 gathering = true
277 try {
278 $.ui.invalidate('ui.render')
279 do {
280 gatherAgain = false
281 shown = await collectOrSayWhy($)
282 $.ui.invalidate('ui.render')
283 } while (gatherAgain)
284 } finally {
285 gathering = false
286 }
287 $.ui.invalidate('ui.render')
288}
289
290// With no change to follow, the repo's and every repo's switches still apply.
291async function withSwitches($, facts) {
292 if (!facts.none) return facts
293 const listed = await run($, [(await kitDir($)) + '/bin/phase', 'switches'], await $.session.cwd())
294 if (listed.failure || listed.exitCode !== 0) return { none: facts.none + "\nPhases off: couldn't read: " + failureOf(listed) }
295 const off = listed.stdout.trim().split('\n').filter(Boolean).map((line) => line.replace(': off (', ' ('))
296 return off.length ? { none: facts.none + '\nPhases off: ' + off.join(', ') } : facts
297}
298
299function asText(facts) {
300 if (facts.none) return clip(facts.none)
301 return [clip(facts.heading), ...facts.sections.flatMap(([title, lines]) => ['', title + ':', ...lines.map((l) => ' ' + clip(l))])].join('\n')
302}
303
304function verifyLines() {
305 if (verifyRun.state === 'idle') return []
306 if (verifyRun.state === 'running') return ['Run verify: running…']
307 return ['Run verify ' + verifyRun.verdict, ...(verifyStopping ? ['stopping it: the button works again once it has ended'] : []), ...verifyRun.tail]
308}
309
310// One verify on the checkout the pane follows now: its verdict and the last lines of its own output.
311async function verifyHere($) {
312 const at = await checkout($, await $.session.cwd())
313 if (at.none) return { verdict: 'did not run: ' + at.none, tail: [] }
314 const kit = await kitDir($)
315 const startedAt = await $.clock.now()
316 const named = (verdict) => ({ verdict: 'on ' + nameOf(at) + ' at ' + timeOf(startedAt) + ': ' + verdict, tail })
317 const tail = []
318 const partial = { stdout: '', stderr: '' }
319 let checkedNothing = false
320 // The marker is stop_checks.py's own, on stderr; the repo's verify output reaches stdout.
321 const keep = (line, stream) => {
322 if (stream === 'stderr' && line.startsWith(CHECKED_NOTHING)) checkedNothing = true
323 if (!line.trim()) return
324 tail.push(line)
325 if (tail.length > VERIFY_TAIL_LINES) tail.shift()
326 }
327 const child = $.process.spawn({ argv: ['python3', kit + '/hooks/stop_checks.py', 'verify'], cwd: at.root })
328 const timer = new AbortController()
329 const deadline = $.clock.sleep(VERIFY_DEADLINE_MS, { signal: timer.signal }).then(() => 'deadline', () => 'cancelled')
330 try {
331 for (;;) {
332 const step = await Promise.race([child.next(), deadline])
333 if (step === 'deadline') {
334 // Not awaited: a child stuck mid-step finishes its return only after that step.
335 verifyStopping = true
336 child.return()
337 .catch((error) => $.ui.log('flow: stopping verify: ' + messageOf(error), { to: 'debug' }))
338 .finally(() => {
339 verifyStopping = false
340 $.ui.invalidate('ui.render')
341 })
342 return named('failed (no result after ' + VERIFY_DEADLINE_MS / 1000 + ' s)')
343 }
344 if (step.done) {
345 keep(partial.stdout, 'stdout')
346 keep(partial.stderr, 'stderr')
347 const { code, signal } = step.value
348 if (code !== 0) return named('failed (' + (signal ? 'killed by ' + signal : 'exit ' + code) + ')')
349 return named(checkedNothing ? 'passed, but it checked nothing' : 'passed')
350 }
351 const stream = step.value.stream === 'stderr' ? 'stderr' : 'stdout'
352 const lines = (partial[stream] + step.value.text).split('\n')
353 partial[stream] = lines.pop().slice(0, LINE_CHARS + 1)
354 lines.forEach((line) => keep(line, stream))
355 }
356 } finally {
357 timer.abort()
358 }
359}
360
361async function runVerify($) {
362 if (verifyRun.state === 'running' || verifyStopping) return
363 verifyRun = { state: 'running' }
364 $.ui.invalidate('ui.render')
365 try {
366 verifyRun = { state: 'done', ...(await verifyHere($)) }
367 } catch (error) {
368 verifyRun = { state: 'done', verdict: "failed: couldn't run: " + messageOf(error), tail: [] }
369 }
370 await gather($)
371}
372
373function reviewerPrompt(lens) {
374 return [
375 `You are one reviewer in the kit's self-review, with one lens: ${lens.title}. You read files and run read-only`,
376 'commands; you never edit files. The spawn prompt gives the base ref, the goal, the spec when there is one, and how',
377 "to report. The repository's own AGENTS.md or CLAUDE.md, when it has one, holds its conventions: read it when a",
378 'finding depends on them.',
379 '',
380 lens.brief,
381 ].join('\n')
382}
383
384async function registerReviewers($) {
385 const kit = await kitDir($)
386 const briefs = await run($, [kit + '/bin/triage', '--lens-briefs'], kit)
387 if (briefs.failure || briefs.exitCode !== 0) {
388 return $.ui.log('kit: no review agents: bin/triage --lens-briefs: ' + failureOf(briefs), { to: 'debug' })
389 }
390 let lenses
391 try {
392 lenses = JSON.parse(briefs.stdout)
393 } catch (error) {
394 return $.ui.log('kit: no review agents: bin/triage --lens-briefs printed something not JSON: ' + messageOf(error), { to: 'debug' })
395 }
396 if (!Array.isArray(lenses)) {
397 return $.ui.log('kit: no review agents: bin/triage --lens-briefs printed JSON that is not a list of lenses', { to: 'debug' })
398 }
399 for (const lens of lenses) {
400 await $.agent
401 .register({
402 name: 'review-' + lens.key,
403 description: `The self-review's ${lens.title} lens: a read-only reviewer of a change. Give it the base ref, ` +
404 'the goal, the spec path and the evidence instruction.',
405 prompt: reviewerPrompt(lens),
406 tools: REVIEWER_TOOLS,
407 omitClaudeMd: true,
408 })
409 .catch((error) => $.ui.log(`kit: review-${lens.key} not registered: ` + messageOf(error), { to: 'debug' }))
410 }
411}
412
413// A kit hook's PreToolUse deny as Claude Code words it when next(e) hands it back: the call never ran. Only the
414// start counts, so a tool's own error that quotes a deny is never run again.
415const KIT_BLOCK = /^PreToolUse:\S+ hook error: Blocked by ~\/\.agents\/hooks\//
416const ALLOW_ONCE = 'Allow once'
417const KEEP_BLOCKED = 'Keep it blocked'
418// The user approves what they read: a cut preview says how much it leaves out.
419const PREVIEW_CHARS = 2_000
420
421async function approveHelper($, kit, args, stdin = '') {
422 try {
423 const done = await $.process.run([kit + '/bin/approve', ...args], { stdin, timeoutMs: PROCESS_TIMEOUT_MS })
424 return done.exitCode === 0 ? done : { failure: failureOf(done) }
425 } catch (error) {
426 return { failure: messageOf(error) }
427 }
428}
429
430// `/flow off|on <phase> [--repo|--global]` and `/flow review-rounds <0-9> [--repo|--global]`: phase_switches.py decides
431// and records, in the session's own directory, as for the message lines `phase off …` and `review-rounds …`, which the
432// prompt hook reads from the same directory.
433async function switchPhase($, e) {
434 const args = e.args.trim().split(/\s+/)
435 const action = args[0].toLowerCase()
436 const rounds = action === 'review-rounds'
437 if (e.origin?.kind !== 'composer') {
438 const line = [...(rounds ? [] : ['phase']), action, ...args.slice(1).map((word) => word.replace(/^--/, ''))].join(' ')
439 return { text: `/flow ${action} ${rounds ? 'sets the re-check rounds' : 'switches a phase'} only when typed at this terminal's prompt: nothing was switched. A message opening with \`${line}\` does it from anywhere.` }
440 }
441 const kit = await kitDir($)
442 const done = await run($, ['python3', kit + '/hooks/phase_switches.py', 'set', action, ...args.slice(1)], await $.session.cwd())
443 // After a failure too: a run that died or timed out may have written the switch first.
444 refreshIfOpen($, 'a switch')
445 const said = done.stdout?.trim()
446 if (done.exitCode === 0 && rounds) return { text: said, context: [`The user set the review re-check rounds with /flow: ${said} The self-review runs that many re-checks.`] }
447 if (done.exitCode === 0) return { text: said, context: [`The user switched a phase with /flow: ${said} Skip the steps of the phases off.`] }
448 if (done.exitCode === SWITCH_REFUSED) return { text: said }
449 if (rounds) {
450 return {
451 text: `Couldn't tell whether review-rounds was set to ${args[1] ?? 'a number'}: ${failureOf(done)}. Run \`~/.agents/bin/phase review-rounds\` to see the value.`,
452 context: ['A /flow review-rounds setting may or may not have been recorded: read `~/.agents/bin/phase review-rounds` before the self-review.'],
453 }
454 }
455 return {
456 text: `Couldn't tell whether ${args[1] ?? 'the phase'} was switched ${action}: ${failureOf(done)}. Run /flow to see what is off.`,
457 context: ['A /flow switch may or may not have been recorded: check `~/.agents/bin/phase switches` before a step of the flow.'],
458 }
459}
460
461function refreshIfOpen($, after) {
462 // Asked each time: a pane whose drawing threw is dropped without a ui.close this mod hears.
463 $.ui.panes()
464 .then((panes) => panes.some((pane) => pane.id === PANE) && gather($))
465 .catch((error) => $.ui.log(`flow: refresh after ${after}: ` + messageOf(error), { to: 'debug' }))
466}
467
468function previewOf(tool, input) {
469 const call = tool === 'Bash' ? String(input.command) : tool + ' ' + JSON.stringify(input, null, 2)
470 if (call.length <= PREVIEW_CHARS) return call
471 return call.slice(0, PREVIEW_CHARS) + `… (${call.length - PREVIEW_CHARS} more characters not shown)`
472}
473
474// The guards decide; this only carries the user's answer to them, as their next message would.
475async function askToLiftBlock($, e, next) {
476 const blocked = await next(e)
477 if (!blocked.isError || !KIT_BLOCK.test(String(blocked.text ?? ''))) return blocked
478 const { tool, tool_use_id, agentId, ...input } = e
479 const kit = await kitDir($)
480 const session = await $.session.id()
481 const asked = await approveHelper($, kit, ['needed'],
482 JSON.stringify({ tool_name: tool, tool_input: input, session_id: session, cwd: await $.session.cwd(), deny: blocked.text }))
483 let needed = null
484 try {
485 needed = asked.failure ? null : JSON.parse(asked.stdout)
486 } catch (error) {
487 asked.failure = 'printed something not JSON: ' + messageOf(error)
488 }
489 if (asked.failure) $.ui.log('kit: approve needed: ' + asked.failure, { to: 'debug' })
490 if (!needed?.names?.length) return blocked
491 const allowTurn = `Allow ${needed.scope} until my next message`
492 const answer = await $.ui
493 .ask(`A kit guard blocked this call: ${needed.what}.\n\n${previewOf(tool, input)}\n\nAllow it?`,
494 { header: 'Approval', options: [ALLOW_ONCE, allowTurn, KEEP_BLOCKED] })
495 .then((given) => given || null)
496 .catch(() => null) // dismissed, interrupted, or a -p run with no one to ask: the block stands, as without the mod
497 const outcome = answer === null ? 'dismissed, or no one to ask' : [ALLOW_ONCE, allowTurn, KEEP_BLOCKED].includes(answer) ? answer : 'answered in their own words'
498 $.ui.log(`kit: approval dialog for ${needed.names.join(', ')}: ${outcome}`, { to: 'debug' })
499 if (answer === null) return blocked
500 if (answer !== ALLOW_ONCE && answer !== allowTurn) {
501 const said = answer === KEEP_BLOCKED ? '' : ` They answered: ${JSON.stringify(answer)}.`
502 return { deny: `${blocked.text}\nThe user was asked in a dialog and kept it blocked.${said} Don't ask them to approve it again this turn.` }
503 }
504 const granted = await approveHelper($, kit, [answer === ALLOW_ONCE ? 'once' : 'grant', session, ...needed.names])
505 if (granted.failure) {
506 $.ui.log("kit: couldn't record your approval, so the call stays blocked: " + granted.failure)
507 return blocked
508 }
509 try {
510 return await next(e)
511 } finally {
512 if (answer === ALLOW_ONCE) {
513 const revoked = await approveHelper($, kit, ['revoke', session, ...needed.names])
514 if (revoked.failure) $.ui.log(`kit: couldn't take back the one-time approval of ${needed.scope}, so it lasts until your next message: ` + revoked.failure)
515 }
516 }
517}
518
519export function register(on) {
520 on('tool.call', { tool: ['Bash', /^mcp__/] }, askToLiftBlock)
521
522 on('session.start', async ($, e, next) => {
523 const flow = FEATURES.find((feature) => feature.command === 'flow')
524 await $.command.register({ name: 'flow', description: flow.description, argumentHint: '[off|on <phase> | review-rounds <0-9> [--repo|--global]]', immediate: true })
525 const waited = new AbortController()
526 const registering = registerReviewers($)
527 .catch((error) => $.ui.log('kit: review agents: ' + messageOf(error), { to: 'debug' }))
528 .finally(() => waited.abort())
529 const timer = $.clock.sleep(REVIEWERS_WAIT_MS, { signal: waited.signal }).then(() => 'timed out', () => 'registered')
530 if ((await Promise.race([registering.then(() => 'registered'), timer])) === 'timed out') {
531 await $.ui.log('kit: review agents: still registering after 2 s, so the session started without them', { to: 'debug' })
532 }
533 return next(e)
534 })
535
536 on('command.run', { command: 'flow' }, async ($, e) => {
537 if (/^(off|on|review-rounds)(\s|$)/i.test(e.args.trim())) return switchPhase($, e)
538 if ((await $.session.surfaces()).length === 0) return { text: asText(await collectOrSayWhy($)) }
539 // Without closeOnEscape: Escape hands the keys back to the prompt and the pane stays, refreshing after each turn.
540 await $.ui.open({ id: PANE, title: 'flow', focus: true })
541 void gather($)
542 return {}
543 })
544
545 on('turn.complete', async ($, e, next) => {
546 if (!e.agentId) refreshIfOpen($, 'the turn')
547 return next(e)
548 })
549
550 on('ui.render', { component: 'Pane' }, async ($, e, next) => {
551 if (e.requestId !== PANE) return next(e)
552 const { Box, Text, Button } = $.ui.resolve(e)
553 const line = (text) => Text({ bold: NEEDS_ATTENTION.test(text), wrap: 'wrap', children: [clip(text)] })
554 const refreshing = shown && gathering ? ' · refreshing…' : ''
555 const top = !shown
556 ? [Text({ children: ['Gathering…'] })]
557 : shown.none
558 ? [line(shown.none + refreshing)]
559 : [Text({ bold: true, children: [clip(shown.heading + refreshing)] })]
560 const sections = (shown?.sections ?? []).flatMap(([title, lines]) => [
561 Text({ children: [' '] }),
562 Text({ dimColor: true, children: [title] }),
563 ...lines.map((text, i) => (title === 'Flow' && i === 0 ? Text({ bold: true, children: [clip(text)] }) : line(text))),
564 ])
565 // A press starts the work and returns: Claude Code skips a hook still running after 10 s.
566 // Run verify only where there is a change: elsewhere a press would do nothing.
567 const buttons = [
568 ...(shown?.root ? [Button({ key: 'run-verify', label: 'Run verify', hotkey: 'v', plain: true, onPress: () => void runVerify($) })] : []),
569 Button({ key: 'refresh', label: 'Refresh', hotkey: 'r', plain: true, onPress: () => void gather($) }),
570 ]
571 // One row per step before the merge, each with its own Pass and Fail: what a press records is the user's.
572 // The buttons lead the row, so they line up whatever the step's line holds.
573 const staging = shown?.staging
574 const failure = markFailure && markFailure.root === shown?.root ? markFailure.text : null
575 const steps = staging?.steps ?? []
576 const stagingRows = !staging || (!staging.failure && !steps.length && !failure && !pendingMarks.length) ? [] : [
577 Text({ children: [' '] }),
578 Text({ dimColor: true, children: ['Staging before the merge'] }),
579 ...(pendingMarks.length ? [Text({ dimColor: true, children: ['Marking ' + pendingMarks.join(', ') + '…'] })] : []),
580 ...(failure ? [line(failure)] : []),
581 ...(staging.failure ? [line(staging.failure)] : steps.map((step) => Box({
582 flexDirection: 'row',
583 columnGap: 2,
584 children: [
585 ...['PASS', 'FAIL'].map((result) => Button({
586 key: `staging-${step.id}-${result}`,
587 label: result === 'PASS' ? 'Pass' : 'Fail',
588 plain: true,
589 onPress: () => void markStep($, step.id, result),
590 })),
591 Text({ bold: step.result === 'FAIL', wrap: 'wrap', children: [clip(stepLine(step))] }),
592 ],
593 }))),
594 ]
595 // The buttons and the run they started come first: a long report scrolls the bottom of the pane away.
596 return Box({
597 flexDirection: 'column',
598 children: [
599 ...top,
600 Box({ flexDirection: 'row', columnGap: 2, children: buttons }),
601 ...verifyLines().map(line),
602 ...stagingRows,
603 ...sections,
604 ],
605 })
606 })
607}
608hooks/features.js 41 lines1// What each command, hook and agent type of the mod offers, and what Codex has instead: bin/docs prints it in
2// docs/framework.md, and tests/flow.test.ts fails when the mod registers a command or agent, or draws a button,
3// this list doesn't name.
4// Kept as one JSON value after the `=`, which bin/docs reads without running JavaScript.
5export const FEATURES = [
6 {
7 "command": "flow",
8 "description": "Show where this branch is in the kit's flow, run verify without a turn, and switch a phase off or on or set the self-review's re-check rounds",
9 "capabilities": [
10 {"name": "Phase and reports", "codex": "`bin/reports brief`"},
11 {"name": "The phases switched off, with their scope, also on the default branch", "codex": "`bin/phase switches`"},
12 {"name": "`/flow off|on <phase> [--repo|--global]`, typed at the prompt, switches a phase for the branch, the repo or every repo", "codex": "a message whose first line is `phase off|on <phase> [branch|repo|global]`"},
13 {"name": "`/flow review-rounds <0-9> [--repo|--global]`, typed at the prompt, sets how many re-checks the self-review runs after its first pass (1 when unset)", "codex": "a message whose first line is `review-rounds <0-9> [branch|repo|global]`; `bin/phase review-rounds` reads it"},
14 {"name": "Follows the branch into the main checkout once validate detaches the session's own", "codex": "the commands of this list, run from the main checkout"},
15 {"name": "Stamps (verify, self-review, validate, staging), current or not, and a verify that checked nothing", "codex": "`python3 ~/.agents/hooks/review_stamp.py check --kind <kind>`"},
16 {"name": "The last verify report's ran/skipped/warning/error lines", "codex": "`cat \"$(~/.agents/bin/reports path verify)\"`"},
17 {"name": "CI of the pushed HEAD", "codex": "`bin/ci-wait --once --no-log`"},
18 {"name": "Context use", "codex": "nothing (Codex shows its own)"},
19 {"name": "Run verify button, without a turn, on the checkout the pane shows", "codex": "`python3 ~/.agents/hooks/stop_checks.py verify` in a terminal"},
20 {"name": "Refresh button, and a refresh after each turn while the pane is open", "codex": "running the commands above again"},
21 {"name": "The staging guide's steps before the merge, each with its latest result, who gave it, and its evidence files", "codex": "`bin/staging steps --json`"},
22 {"name": "Pass button on each step before the merge, recorded as the user's verdict", "codex": "`bin/staging mark <step> PASS --by user` in a terminal (the shell guard refuses it from the agent)"},
23 {"name": "Fail button on each step before the merge, recorded as the user's verdict", "codex": "`bin/staging mark <step> FAIL --by user` in a terminal"}
24 ]
25 },
26 {
27 "hook": "tool.call",
28 "description": "Ask the user in a dialog when a kit guard blocks a call they could approve",
29 "capabilities": [
30 {"name": "A blocked MCP write or `DROP`/`TRUNCATE` command shown in a dialog: allowed once, until the user's next message, or kept blocked", "codex": "the user's next message naming the service or opening with `allow <statement>`, or `touch ~/.agents/approvals/<service>`"}
31 ]
32 },
33 {
34 "agentPrefix": "review",
35 "description": "One reviewer agent type per self-review lens",
36 "capabilities": [
37 {"name": "One for each lens `bin/triage --lens-briefs` prints: its brief from lenses.md, no edit tools, no CLAUDE.md block", "codex": "a subagent given the lens text from `skills/self-review/lenses.md`"}
38 ]
39 }
40]
41