SLOPSHOPPER

kit

The agents kit inside Claude Code: the /flow pane

newpaneguardcommandprocess
v0.5.0MITupdated 2026-10-06alopezari/agents-kit/mods/kit
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · kit
│ ┃ flow ✕ › fix the failing auth test and add an audit log call │ ┃ Gathering… │ ┃ r: Refresh ● kit: kit: review agents: still registering after 2 s, so the sessio │ ● kit: kit: no review agents: bin/triage --lens-briefs printed someth │ ⏺ Read(src/auth.ts) │ ⎿ Read 6 lines │ ⏺ Update(src/auth.ts) │ ⎿ Added 2 lines, removed 1 line │ ⏺ Bash(bun test) │ ⎿ 3 pass, 1 fail │ │ ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ │ › /flow │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · flow
Gathering… r: Refresh
README

agents-kit

One setup for every coding agent you use (Claude Code, Codex): the same instructions, skills, guardrails and checks, whichever harness runs the work. Agents are told to work like a staff engineer, and hooks check that they did: irreversible and outward-facing commands stop and wait for you, every turn that edits code is verified, and a pull request can't be opened until the exact change has been self-reviewed. You can switch any of those phases off, and back on, for a branch, a repo or everywhere: phase off verify as the first line of a message, or /flow off verify in Claude Code. review-rounds 2 the same way sets how many re-checks the self-review runs after its first pass (1 by default).

docs/framework.md is the full reference: every hook, guard rule, check, skill, tool and scheduled job, with diagrams. It is generated from the code on every commit, so it describes what the kit does now.

macOS only for now (the scheduled jobs use launchd).

Install

git clone https://github.com/alopezari/agents-kit.git ~/.agents && ~/.agents/install.sh

install.sh wires every installed harness and is safe to re-run. To check the setup afterwards:

~/.agents/install.sh --doctor   # report what's missing, change nothing
~/.agents/tests/run.sh          # regression suite (passes once --doctor is clean)

The kit must live at ~/.agents. The installer:

  • offers to install missing required and recommended programs from deps.txt through Homebrew, asking once (--yes skips the question; without a terminal it only reports them);
  • links AGENTS.md as each harness's global instructions and the skills into each harness;
  • registers the hooks next to any hooks already there;
  • fetches the third-party skills in skills.external, a pinned one at its commit into vendor/, and installs the Node dependencies of tools/ and site/;
  • sets the Claude Code status line and enables the kit's git hooks (they keep docs/framework.md current);
  • registers the kit as a Claude Code plugin marketplace and installs its plugin, the /flow mod (see Mods contract);
  • installs the scheduled jobs;
  • when git uses a per-host proxy, offers to link bin/gh into /usr/local/bin (asks for your password).

It also adds the kit's baseline harness settings (no fast mode, effort defaults) wherever a key is missing, without overwriting one you set. Files it replaces are backed up under backups/.

To remove it, ~/.agents/uninstall.sh lists every link, hook, status line, scheduled job and the kit's Claude Code plugin that still point at the kit, asks, and removes them (--yes skips the question). Your own hooks and settings stay, and so do the baseline settings and the programs installed through Homebrew. It backs up the settings files it edits to ~/.agents-uninstall-backups/, and leaves ~/.agents itself for you to delete.

What it doesn't do, on a new machine:

  1. Install or log in to the harnesses (claude, codex) and gh. Install them before running install.sh, which only wires the harnesses it finds.
  2. Trust the Codex hooks. Open Codex once and approve them; install.sh --doctor warns until you do.
  3. Install harness plugins, other than the kit's own, or MCP servers. Add the ones you use yourself, or keep their setup in a profile.

Requirements: deps.txt lists every program the kit runs. install.sh offers to install the missing required ones (python3, git, jq, node, gh) and recommended ones (semgrep, gitleaks, php) through Homebrew, and says how to install the optional ones, each needed by one feature. install.sh --doctor reports what is missing, and tests/run.sh expects it to report nothing.

Make it yours

  • Instructions: edit AGENTS.md. It is plain Markdown and every harness reads the same file.
  • A repository: add repos/<checkout-directory-name>/notes.md with the context agents should read there. Add an executable verify when the automatic checks aren't enough. Nothing is written into the repository itself. See repos/README.md.
  • A profile: keep anything tied to one employer or client in a separate private repository and layer it in with install.sh --profile <dir>. It can hold repo overlays, work-only skills, research, MCP write rules and the repositories to learn from; examples/sample-profile/ has one of each to start from. The core stays generic and shareable.

Hook contract

Every hook is a small program with one contract: a JSON payload on stdin, a JSON decision on stdout. Claude Code and Codex share this contract natively. Other harnesses need an adapter that translates their events into it.

Payload fields used: tool_name, tool_input.command, tool_input.file_path (or notebook_path, or the file list in Codex's apply_patch input), cwd, session_id, stop_hook_active, and turn_id to tell Codex apart in the log.

Decisions:

  • Deny a command: {"hookSpecificOutput": {"permissionDecision": "deny", "permissionDecisionReason": ...}}.
  • Block after an edit or at stop: {"decision": "block", "reason": ...}, optionally with systemMessage for the user.

Mods contract

Claude Code also loads the kit as a plugin, mods/kit/, whose mod runs inside Claude Code: it can draw panes, run commands without a turn and define agent types. Mods exist only in Claude Code, so the kit keeps working the same with or without them. Every mod follows these rules:

  1. It owns no workflow state. Specs, stamps, reports and evidence stay where the kit's tools keep them (.git/agents/, or a temporary directory when .git can't be written) and go through bin/ and hooks/review_stamp.py, so a change started in Claude Code can be finished in Codex, and the other way round. What those tools cache or migrate follows their own rules.
  2. It owns no safety. The Python hooks decide what is blocked, in every harness. A mod can add to that in Claude Code, or carry the user's own answer to a hook through the kit's tools, as their next message would; it never decides one. A disabled or failed mod leaves the behavior as it was without it, which is what Codex has.
  3. Every capability names what Codex has instead, or says it has nothing. They are listed in mods/kit/hooks/features.js, which docs/framework.md prints, and the mod's tests fail when the commands or agents it registers, or the buttons it draws, differ from the list.
  4. The flow works fully with the mod disabled or failed. claude --safe-mode turns off the kit's settings hooks too, so it isn't a supported way to run the kit.
  5. Guarantees go by tier. A capability that gives Claude Code a stronger guarantee than Codex has records where each fact came from, such as a step the user marked with a button versus one the agent wrote down. A record without that provenance counts as the weaker tier.
  6. It calls the kit's Python instead of reimplementing it, so one rule never has two implementations that drift apart.

Its /flow pane shows where the branch is in the flow, with the phases switched off, and runs verify without a turn; /flow off <phase> [--repo|--global] and /flow on … switch a phase as the message line does. While the branch waits on its staging guide, the pane lists each step before the merge with a Pass and a Fail button that record your verdict; from a terminal, ~/.agents/bin/staging mark <step> PASS --by user does the same. When a guard blocks an MCP write or a DROP/TRUNCATE command, the mod asks you in a dialog that shows the call; allowing it, for that call or until your next message, writes the approval your message would have (bin/approve). docs/framework.md lists every capability with its Codex equivalent.

install.sh registers ~/.agents as the agents-kit plugin marketplace and installs kit@agents-kit. The plugin loads in place, so a pull reaches it at the next session start or /reload-plugins. To work on the mod, run claude --plugin-dir ~/.agents/mods/kit, which reloads it on save; tests/run.sh mods validates and tests it.

Adding a harness

  1. Point its global instructions file at AGENTS.md, and its skills directory at skills/ if it has one.
  2. If it has no hook settings of its own, write adapters/<harness>/ to translate its events into the payloads above:
  3. before a command → guard_bash.py;
  4. after an edit → post_edit.py;
  5. at idle or stop → stop_checks.py, with a guard so the stop check continues the agent only once per user turn.
  6. Add a section to install.sh and run install.sh --doctor. Then test with a harmless blocked command in a scratch repo.
  7. Teach bin/docs where the new harness registers its hooks (HARNESS_OF_SETTINGS and harness_hooks), then commit: the pre-commit hook regenerates docs/framework.md.

Testing a hook change through a real harness

tests/ feeds the hooks hand-made payloads. To see a branch's hooks run inside Codex itself, with its real payloads, point HOME at a directory whose .agents is the branch's checkout. Codex's hooks.json runs python3 $HOME/.agents/hooks/<hook>.py, and it trusts a hook by that command text, so the branch's code runs through the entries you already approved. CODEX_HOME keeps Codex's own settings, login and trust, and GH_CONFIG_DIR keeps gh logged in; they come before HOME= because bash and zsh expand ~ with the HOME assigned before it:

H=$(mktemp -d) && ln -s ~/.agents-worktree-<name> "$H/.agents"   # the branch's checkout
git init -q /tmp/hook-probe && git -C /tmp/hook-probe commit -q --allow-empty -m init
echo 'Run this shell command once and report what happened: <a command the change should block or allow>' \
  | CODEX_HOME=~/.codex GH_CONFIG_DIR=~/.config/gh HOME="$H" \
    codex exec -C /tmp/hook-probe --skip-git-repo-check -s workspace-write --ephemeral -
tail -3 ~/.agents-worktree-<name>/logs/hooks.jsonl

The hooks log to the branch checkout's logs/ (not in git), so the last lines show each decision with "harness": "codex". The checkout has no repos/<repo>/verify overlay, so the stop hook falls back to the automatic verify. Use a scratch repo: a hook that fails to block lets the command run.

Not in git

logs/, backups/, research/, approvals/, monitors/state/, review-mining/runs/, review-mining/baseline.json and usage/*.json hold local, possibly private data; site/dist/ is the built website. Profiles are linked in and never committed here.

License

MIT, see LICENSE.

Source 2 files
hooks/register.js 608 lines
1// The kit inside Claude Code. Everything here shows or runs the kit's own tools: the state stays with them
2// and the decisions in the Python hooks, so Codex and a session without this mod get the same flow.
3import { FEATURES } from './features.js'
4
5const PANE = 'flow'
6const PROCESS_TIMEOUT_MS = 30_000
7// phase_switches.py's exit for a switch it refused; 1 is Python's own for a crash.
8const SWITCH_REFUSED = 3
9// ci-wait --once makes one GitHub round trip per check source, each bounded at 30 s.
10const CI_TIMEOUT_MS = 90_000
11const STAMPS = [['verify', 'verify'], ['self-review', 'review'], ['validate', 'validate'], ['staging', 'staging']]
12const CI_OUTCOMES = { 0: 'passed', 1: 'failed', 2: 'running', 3: 'no checks', 4: 'unreadable' }
13const REPORT_LINE = /^(ran|skipped|warning|error):/
14const VERIFY_TAIL_LINES = 8
15// Claude Code refuses a Text string over 10,000 characters; a verify line can be any length.
16const LINE_CHARS = 500
17// stop_checks.py bounds the repo's verify at 600 s, but not its own git reads and stamp writes around it.
18const VERIFY_DEADLINE_MS = 660_000
19// What `stop_checks.py verify` prints when the run passed without checking anything.
20const CHECKED_NOTHING = 'verify passed, but it checked nothing'
21// A gathering whose branch or HEAD moved while it ran is gathered again, this many times in all.
22const GATHER_ATTEMPTS = 2
23const NEEDS_ATTENTION = /couldn't read|^Couldn't|^CI: failed|^Run verify.* failed/
24
25// Each self-review lens becomes the agent type `kit:review-<key>`, briefed from what `bin/triage --lens-briefs` cuts
26// out of lenses.md. No edit tools is a convenience for the reviewer, not a guard (Bash can still write): the hooks
27// run for subagents too.
28const REVIEWER_TOOLS = ['Read', 'Grep', 'Glob', 'Bash']
29// Registering takes tens of milliseconds and must land before the first turn (`claude -p` starts one at once), so the
30// session start waits for it, but never longer than this.
31const REVIEWERS_WAIT_MS = 2_000
32
33// What the pane draws. One gathering runs at a time: a request during one gathers again once it ends.
34let shown = null
35let gathering = false
36let gatherAgain = false
37let verifyRun = { state: 'idle' }
38// A verify stopped at its deadline holds the button until its child is gone, so two never write one report.
39let verifyStopping = false
40// Marks run one at a time, in the order pressed, each shown as pending until it lands.
41let marking = Promise.resolve()
42let pendingMarks = []
43// { root, text }: shown while the pane shows that checkout, whatever its branch now: a refusal says the branch moved.
44let markFailure = null
45// A step's title is cut to this, so its result and evidence stay on the line.
46const TITLE_CHARS = 60
47
48async function run($, argv, cwd, timeoutMs = PROCESS_TIMEOUT_MS) {
49  try {
50    return await $.process.run(argv, { cwd, timeoutMs })
51  } catch (error) {
52    return { failure: messageOf(error) }
53  }
54}
55
56function messageOf(error) {
57  return String(error?.message ?? error)
58}
59
60function failureOf(result) {
61  if (result.failure) return result.failure
62  return 'exit ' + result.exitCode + (result.stderr?.trim() ? ': ' + result.stderr.trim().split('\n')[0] : '')
63}
64
65async function kitDir($) {
66  return (await $.env.get('HOME')) + '/.agents'
67}
68
69async function stampLine($, kit, root, label, kind) {
70  const check = (stampKind) => run($, ['python3', kit + '/hooks/review_stamp.py', 'check', '--kind', stampKind], root)
71  const result = await check(kind)
72  if (result.exitCode === 0) return label + ': current'
73  if (result.exitCode !== 1) return label + ": couldn't read: " + failureOf(result)
74  if (kind === 'verify') {
75    const empty = await check('verify-empty')
76    if (empty.exitCode === 0) return label + ': current, but it checked nothing'
77    if (empty.exitCode !== 1) return label + ": couldn't read: " + failureOf(empty)
78  }
79  return label + ': not current'
80}
81
82async function verifyReportLines($, kit, root) {
83  const path = await run($, [kit + '/bin/reports', 'path', 'verify'], root)
84  if (path.failure || path.exitCode !== 0) return ["couldn't read: " + failureOf(path)]
85  const file = path.stdout.trim()
86  try {
87    if (!(await $.fs.exists(file))) return ['no verify run yet']
88    const { mtimeMs } = await $.fs.stat(file)
89    const lines = (await $.fs.read(file)).split('\n')
90    return [lines[0].replace(/^# /, '') + ' · ' + dateTimeOf(mtimeMs), ...lines.filter((line) => REPORT_LINE.test(line))]
91  } catch (error) {
92    return ["couldn't read: " + messageOf(error)]
93  }
94}
95
96// { steps } for the guide's steps before the merge, or { failure } when bin/staging can't list them.
97async function stagingSteps($, kit, root) {
98  const listed = await run($, [kit + '/bin/staging', 'steps', '--json'], root)
99  if (listed.failure || listed.exitCode !== 0) return { failure: "couldn't read: " + failureOf(listed) }
100  try {
101    return { steps: JSON.parse(listed.stdout) }
102  } catch (error) {
103    return { failure: "couldn't read: bin/staging printed something not JSON: " + messageOf(error) }
104  }
105}
106
107function stepLine(step) {
108  const title = step.title.length > TITLE_CHARS ? step.title.slice(0, TITLE_CHARS - 1) + '…' : step.title
109  const result = step.result ? step.result + (step.by ? ' (' + step.by + ')' : '') : 'not marked'
110  const evidence = step.evidence.length ? step.evidence.map((path) => path.split('/').pop()).join(', ') : 'no evidence yet'
111  return [step.id + ' ' + title, result, evidence].join(' · ')
112}
113
114function markStep($, id, result) {
115  if (!shown?.root) return marking
116  // What the pressed button showed: a gathering can move the pane before this mark's turn comes.
117  const { root, branch } = shown
118  const label = id + ' ' + (result === 'PASS' ? 'Pass' : 'Fail')
119  pendingMarks.push(label)
120  $.ui.invalidate('ui.render')
121  marking = marking.then(async () => {
122    try {
123      const kit = await kitDir($)
124      // --branch: the checkout itself may have moved since the pane showed this step.
125      const marked = await run($, [kit + '/bin/staging', 'mark', id, result, '--by', 'user', '--branch', branch], root)
126      if (marked.failure || marked.exitCode !== 0) markFailure = { root, text: `Couldn't mark ${label}: ` + failureOf(marked) }
127      else if (markFailure?.root === root) markFailure = null
128    } finally {
129      pendingMarks.splice(pendingMarks.indexOf(label), 1)
130    }
131    await gather($)
132  }).catch((error) => {
133    // A broken chain would leave every later press doing nothing.
134    markFailure = { root, text: `Marked ${label}, or not: the pane failed while it ran: ` + messageOf(error) }
135    $.ui.invalidate('ui.render')
136  })
137  return marking
138}
139
140async function ciLine($, kit, root, head) {
141  const upstream = await run($, ['git', 'rev-parse', '--verify', '--quiet', '@{u}'], root)
142  if (upstream.exitCode === 1) return 'CI: HEAD not pushed (no upstream)'
143  if (upstream.failure || upstream.exitCode !== 0) return "CI: couldn't read: " + failureOf(upstream)
144  const pushed = await run($, ['git', 'merge-base', '--is-ancestor', head, '@{u}'], root)
145  if (pushed.failure || pushed.exitCode > 1) return "CI: couldn't read: " + failureOf(pushed)
146  if (pushed.exitCode === 1) return 'CI: HEAD not pushed'
147  const ci = await run($, [kit + '/bin/ci-wait', '--sha', head, '--once', '--no-log'], root, CI_TIMEOUT_MS)
148  const outcome = ci.failure ? undefined : CI_OUTCOMES[ci.exitCode]
149  if (!outcome) return "CI: couldn't read: " + failureOf(ci)
150  const reason = ci.exitCode === 4 ? (ci.stderr.trim() || ci.stdout.trim()).split('\n')[0] : ''
151  const line = 'CI: ' + outcome + (reason ? ': ' + reason : '')
152  const status = await run($, ['git', 'status', '--porcelain'], root)
153  if (status.failure || status.exitCode !== 0) {
154    return line + ", for " + head.slice(0, 7) + "; couldn't tell whether there are uncommitted changes: " + failureOf(status)
155  }
156  return line + (status.stdout.trim() !== '' ? ', for ' + head.slice(0, 7) + ' without the uncommitted changes' : '')
157}
158
159async function contextLine($) {
160  try {
161    const { context } = await $.session.usage()
162    return 'context: ' + (context?.percent === undefined ? '–' : Math.round(context.percent) + '%')
163  } catch {
164    return 'context: –'
165  }
166}
167
168function timeOf(ms) {
169  return new Date(ms).toTimeString().slice(0, 8)
170}
171
172function dateTimeOf(ms) {
173  const date = new Date(ms)
174  const pad = (n) => String(n).padStart(2, '0')
175  return date.getFullYear() + '-' + pad(date.getMonth() + 1) + '-' + pad(date.getDate()) + ' ' + timeOf(ms)
176}
177
178// The checkout at `cwd`: { root, branch, head, inMain? }, or { none } saying why there is no change to follow.
179async function checkout($, cwd) {
180  const top = await run($, ['git', 'rev-parse', '--show-toplevel'], cwd)
181  if (top.failure) return { none: "Couldn't read the repository: " + top.failure }
182  // git exits 128 for "not a git repository", and for a refused or broken one too.
183  if (top.exitCode !== 0 && /not a git repository/.test(top.stderr)) return { none: 'Not in a git repository: no change to follow.' }
184  if (top.exitCode !== 0) return { none: "Couldn't read the repository: " + failureOf(top) }
185  const at = await branchAt($, top.stdout.trim())
186  if (at.none || at.branch) return at
187  return (await mainCheckout($, at.root)) ?? { none: 'Detached HEAD: no change to follow.' }
188}
189
190async function branchAt($, root) {
191  const [branch, head] = await Promise.all([
192    run($, ['git', 'branch', '--show-current'], root),
193    run($, ['git', 'rev-parse', 'HEAD'], root),
194  ])
195  if (head.failure || head.exitCode !== 0) return { none: "Couldn't read HEAD: " + failureOf(head) }
196  if (branch.failure || branch.exitCode !== 0) return { none: "Couldn't read the branch: " + failureOf(branch) }
197  return { root, branch: branch.stdout.trim(), head: head.stdout.trim() }
198}
199
200// validate step 7 frees the branch by detaching the session's worktree, and the user checks it out in the main one.
201async function mainCheckout($, root) {
202  const listed = await run($, ['git', 'worktree', 'list', '--porcelain'], root)
203  if (listed.failure || listed.exitCode !== 0) return { none: "Detached HEAD, and couldn't find the main checkout: " + failureOf(listed) }
204  const first = listed.stdout.split('\n\n')[0]
205  const main = /^worktree (.+)$/m.exec(first)?.[1]
206  if (!main || main === root || /^bare$/m.test(first)) return null
207  const at = await branchAt($, main)
208  return at.none || at.branch ? { ...at, inMain: !at.none } : null
209}
210
211function clip(text) {
212  return text.length > LINE_CHARS ? text.slice(0, LINE_CHARS - 1) + '…' : text
213}
214
215function nameOf(at) {
216  return at.branch + ' @ ' + at.head.slice(0, 7) + (at.inMain ? ' in the main checkout' : '')
217}
218
219function sameCheckout(a, b) {
220  return !a.none && !b.none && a.root === b.root && a.branch === b.branch && a.head === b.head
221}
222
223// The facts for the change the session is on: { none } when there is no change to follow, otherwise a
224// heading naming the branch and HEAD they describe, and one list of lines per section.
225async function collect($) {
226  const kit = await kitDir($)
227  for (let attempt = 1; ; attempt++) {
228    const at = await checkout($, await $.session.cwd())
229    if (at.none) return at
230    // The brief comes first: on the default branch it is the only answer, and CI can take ninety seconds.
231    const brief = await run($, [kit + '/bin/reports', 'brief'], at.root)
232    let flow
233    if (brief.failure || brief.exitCode !== 0) flow = ["couldn't read: " + failureOf(brief)]
234    else {
235      flow = brief.stdout.split('\n').filter((line) => line && !/^\s/.test(line) && !line.startsWith('Read one in full'))
236      if (flow[0] === 'phase: ') {
237        return { none: at.inMain ? 'Detached HEAD, and the main checkout is on the default branch: no change to follow.' : 'On the default branch: no change to follow.' }
238      }
239    }
240    const [stamps, report, staging, ci, context, gatheredAt] = await Promise.all([
241      Promise.all(STAMPS.map(([label, kind]) => stampLine($, kit, at.root, label, kind))),
242      verifyReportLines($, kit, at.root),
243      stagingSteps($, kit, at.root),
244      ciLine($, kit, at.root, at.head),
245      contextLine($),
246      $.clock.now(),
247    ])
248    // Read again from the session's directory: Claude Code's /cd can move it while this gathers.
249    const after = await checkout($, await $.session.cwd())
250    if (after.none) return after
251    const moved = !sameCheckout(at, after)
252    if (moved && attempt < GATHER_ATTEMPTS) continue
253    return {
254      root: at.root,
255      branch: at.branch,
256      heading: nameOf(at) + ' · gathered ' + timeOf(gatheredAt) + (moved ? ' · the checkout moved while gathering: Refresh' : ''),
257      sections: [['Flow', flow], ['Stamps', stamps], ['Last verify report', report], ['CI and context', [ci, context]]],
258      staging,
259    }
260  }
261}
262
263async function collectOrSayWhy($) {
264  try {
265    return await withSwitches($, await collect($))
266  } catch (error) {
267    return { none: "Couldn't gather the flow: " + messageOf(error) }
268  }
269}
270
271async function gather($) {
272  if (gathering) {
273    gatherAgain = true
274    return
275  }
276  gathering = true
277  try {
278    $.ui.invalidate('ui.render')
279    do {
280      gatherAgain = false
281      shown = await collectOrSayWhy($)
282      $.ui.invalidate('ui.render')
283    } while (gatherAgain)
284  } finally {
285    gathering = false
286  }
287  $.ui.invalidate('ui.render')
288}
289
290// With no change to follow, the repo's and every repo's switches still apply.
291async function withSwitches($, facts) {
292  if (!facts.none) return facts
293  const listed = await run($, [(await kitDir($)) + '/bin/phase', 'switches'], await $.session.cwd())
294  if (listed.failure || listed.exitCode !== 0) return { none: facts.none + "\nPhases off: couldn't read: " + failureOf(listed) }
295  const off = listed.stdout.trim().split('\n').filter(Boolean).map((line) => line.replace(': off (', ' ('))
296  return off.length ? { none: facts.none + '\nPhases off: ' + off.join(', ') } : facts
297}
298
299function asText(facts) {
300  if (facts.none) return clip(facts.none)
301  return [clip(facts.heading), ...facts.sections.flatMap(([title, lines]) => ['', title + ':', ...lines.map((l) => '  ' + clip(l))])].join('\n')
302}
303
304function verifyLines() {
305  if (verifyRun.state === 'idle') return []
306  if (verifyRun.state === 'running') return ['Run verify: running…']
307  return ['Run verify ' + verifyRun.verdict, ...(verifyStopping ? ['stopping it: the button works again once it has ended'] : []), ...verifyRun.tail]
308}
309
310// One verify on the checkout the pane follows now: its verdict and the last lines of its own output.
311async function verifyHere($) {
312  const at = await checkout($, await $.session.cwd())
313  if (at.none) return { verdict: 'did not run: ' + at.none, tail: [] }
314  const kit = await kitDir($)
315  const startedAt = await $.clock.now()
316  const named = (verdict) => ({ verdict: 'on ' + nameOf(at) + ' at ' + timeOf(startedAt) + ': ' + verdict, tail })
317  const tail = []
318  const partial = { stdout: '', stderr: '' }
319  let checkedNothing = false
320  // The marker is stop_checks.py's own, on stderr; the repo's verify output reaches stdout.
321  const keep = (line, stream) => {
322    if (stream === 'stderr' && line.startsWith(CHECKED_NOTHING)) checkedNothing = true
323    if (!line.trim()) return
324    tail.push(line)
325    if (tail.length > VERIFY_TAIL_LINES) tail.shift()
326  }
327  const child = $.process.spawn({ argv: ['python3', kit + '/hooks/stop_checks.py', 'verify'], cwd: at.root })
328  const timer = new AbortController()
329  const deadline = $.clock.sleep(VERIFY_DEADLINE_MS, { signal: timer.signal }).then(() => 'deadline', () => 'cancelled')
330  try {
331    for (;;) {
332      const step = await Promise.race([child.next(), deadline])
333      if (step === 'deadline') {
334        // Not awaited: a child stuck mid-step finishes its return only after that step.
335        verifyStopping = true
336        child.return()
337          .catch((error) => $.ui.log('flow: stopping verify: ' + messageOf(error), { to: 'debug' }))
338          .finally(() => {
339            verifyStopping = false
340            $.ui.invalidate('ui.render')
341          })
342        return named('failed (no result after ' + VERIFY_DEADLINE_MS / 1000 + ' s)')
343      }
344      if (step.done) {
345        keep(partial.stdout, 'stdout')
346        keep(partial.stderr, 'stderr')
347        const { code, signal } = step.value
348        if (code !== 0) return named('failed (' + (signal ? 'killed by ' + signal : 'exit ' + code) + ')')
349        return named(checkedNothing ? 'passed, but it checked nothing' : 'passed')
350      }
351      const stream = step.value.stream === 'stderr' ? 'stderr' : 'stdout'
352      const lines = (partial[stream] + step.value.text).split('\n')
353      partial[stream] = lines.pop().slice(0, LINE_CHARS + 1)
354      lines.forEach((line) => keep(line, stream))
355    }
356  } finally {
357    timer.abort()
358  }
359}
360
361async function runVerify($) {
362  if (verifyRun.state === 'running' || verifyStopping) return
363  verifyRun = { state: 'running' }
364  $.ui.invalidate('ui.render')
365  try {
366    verifyRun = { state: 'done', ...(await verifyHere($)) }
367  } catch (error) {
368    verifyRun = { state: 'done', verdict: "failed: couldn't run: " + messageOf(error), tail: [] }
369  }
370  await gather($)
371}
372
373function reviewerPrompt(lens) {
374  return [
375    `You are one reviewer in the kit's self-review, with one lens: ${lens.title}. You read files and run read-only`,
376    'commands; you never edit files. The spawn prompt gives the base ref, the goal, the spec when there is one, and how',
377    "to report. The repository's own AGENTS.md or CLAUDE.md, when it has one, holds its conventions: read it when a",
378    'finding depends on them.',
379    '',
380    lens.brief,
381  ].join('\n')
382}
383
384async function registerReviewers($) {
385  const kit = await kitDir($)
386  const briefs = await run($, [kit + '/bin/triage', '--lens-briefs'], kit)
387  if (briefs.failure || briefs.exitCode !== 0) {
388    return $.ui.log('kit: no review agents: bin/triage --lens-briefs: ' + failureOf(briefs), { to: 'debug' })
389  }
390  let lenses
391  try {
392    lenses = JSON.parse(briefs.stdout)
393  } catch (error) {
394    return $.ui.log('kit: no review agents: bin/triage --lens-briefs printed something not JSON: ' + messageOf(error), { to: 'debug' })
395  }
396  if (!Array.isArray(lenses)) {
397    return $.ui.log('kit: no review agents: bin/triage --lens-briefs printed JSON that is not a list of lenses', { to: 'debug' })
398  }
399  for (const lens of lenses) {
400    await $.agent
401      .register({
402        name: 'review-' + lens.key,
403        description: `The self-review's ${lens.title} lens: a read-only reviewer of a change. Give it the base ref, ` +
404          'the goal, the spec path and the evidence instruction.',
405        prompt: reviewerPrompt(lens),
406        tools: REVIEWER_TOOLS,
407        omitClaudeMd: true,
408      })
409      .catch((error) => $.ui.log(`kit: review-${lens.key} not registered: ` + messageOf(error), { to: 'debug' }))
410  }
411}
412
413// A kit hook's PreToolUse deny as Claude Code words it when next(e) hands it back: the call never ran. Only the
414// start counts, so a tool's own error that quotes a deny is never run again.
415const KIT_BLOCK = /^PreToolUse:\S+ hook error: Blocked by ~\/\.agents\/hooks\//
416const ALLOW_ONCE = 'Allow once'
417const KEEP_BLOCKED = 'Keep it blocked'
418// The user approves what they read: a cut preview says how much it leaves out.
419const PREVIEW_CHARS = 2_000
420
421async function approveHelper($, kit, args, stdin = '') {
422  try {
423    const done = await $.process.run([kit + '/bin/approve', ...args], { stdin, timeoutMs: PROCESS_TIMEOUT_MS })
424    return done.exitCode === 0 ? done : { failure: failureOf(done) }
425  } catch (error) {
426    return { failure: messageOf(error) }
427  }
428}
429
430// `/flow off|on <phase> [--repo|--global]` and `/flow review-rounds <0-9> [--repo|--global]`: phase_switches.py decides
431// and records, in the session's own directory, as for the message lines `phase off …` and `review-rounds …`, which the
432// prompt hook reads from the same directory.
433async function switchPhase($, e) {
434  const args = e.args.trim().split(/\s+/)
435  const action = args[0].toLowerCase()
436  const rounds = action === 'review-rounds'
437  if (e.origin?.kind !== 'composer') {
438    const line = [...(rounds ? [] : ['phase']), action, ...args.slice(1).map((word) => word.replace(/^--/, ''))].join(' ')
439    return { text: `/flow ${action} ${rounds ? 'sets the re-check rounds' : 'switches a phase'} only when typed at this terminal's prompt: nothing was switched. A message opening with \`${line}\` does it from anywhere.` }
440  }
441  const kit = await kitDir($)
442  const done = await run($, ['python3', kit + '/hooks/phase_switches.py', 'set', action, ...args.slice(1)], await $.session.cwd())
443  // After a failure too: a run that died or timed out may have written the switch first.
444  refreshIfOpen($, 'a switch')
445  const said = done.stdout?.trim()
446  if (done.exitCode === 0 && rounds) return { text: said, context: [`The user set the review re-check rounds with /flow: ${said} The self-review runs that many re-checks.`] }
447  if (done.exitCode === 0) return { text: said, context: [`The user switched a phase with /flow: ${said} Skip the steps of the phases off.`] }
448  if (done.exitCode === SWITCH_REFUSED) return { text: said }
449  if (rounds) {
450    return {
451      text: `Couldn't tell whether review-rounds was set to ${args[1] ?? 'a number'}: ${failureOf(done)}. Run \`~/.agents/bin/phase review-rounds\` to see the value.`,
452      context: ['A /flow review-rounds setting may or may not have been recorded: read `~/.agents/bin/phase review-rounds` before the self-review.'],
453    }
454  }
455  return {
456    text: `Couldn't tell whether ${args[1] ?? 'the phase'} was switched ${action}: ${failureOf(done)}. Run /flow to see what is off.`,
457    context: ['A /flow switch may or may not have been recorded: check `~/.agents/bin/phase switches` before a step of the flow.'],
458  }
459}
460
461function refreshIfOpen($, after) {
462  // Asked each time: a pane whose drawing threw is dropped without a ui.close this mod hears.
463  $.ui.panes()
464    .then((panes) => panes.some((pane) => pane.id === PANE) && gather($))
465    .catch((error) => $.ui.log(`flow: refresh after ${after}: ` + messageOf(error), { to: 'debug' }))
466}
467
468function previewOf(tool, input) {
469  const call = tool === 'Bash' ? String(input.command) : tool + ' ' + JSON.stringify(input, null, 2)
470  if (call.length <= PREVIEW_CHARS) return call
471  return call.slice(0, PREVIEW_CHARS) + `… (${call.length - PREVIEW_CHARS} more characters not shown)`
472}
473
474// The guards decide; this only carries the user's answer to them, as their next message would.
475async function askToLiftBlock($, e, next) {
476  const blocked = await next(e)
477  if (!blocked.isError || !KIT_BLOCK.test(String(blocked.text ?? ''))) return blocked
478  const { tool, tool_use_id, agentId, ...input } = e
479  const kit = await kitDir($)
480  const session = await $.session.id()
481  const asked = await approveHelper($, kit, ['needed'],
482    JSON.stringify({ tool_name: tool, tool_input: input, session_id: session, cwd: await $.session.cwd(), deny: blocked.text }))
483  let needed = null
484  try {
485    needed = asked.failure ? null : JSON.parse(asked.stdout)
486  } catch (error) {
487    asked.failure = 'printed something not JSON: ' + messageOf(error)
488  }
489  if (asked.failure) $.ui.log('kit: approve needed: ' + asked.failure, { to: 'debug' })
490  if (!needed?.names?.length) return blocked
491  const allowTurn = `Allow ${needed.scope} until my next message`
492  const answer = await $.ui
493    .ask(`A kit guard blocked this call: ${needed.what}.\n\n${previewOf(tool, input)}\n\nAllow it?`,
494      { header: 'Approval', options: [ALLOW_ONCE, allowTurn, KEEP_BLOCKED] })
495    .then((given) => given || null)
496    .catch(() => null)  // dismissed, interrupted, or a -p run with no one to ask: the block stands, as without the mod
497  const outcome = answer === null ? 'dismissed, or no one to ask' : [ALLOW_ONCE, allowTurn, KEEP_BLOCKED].includes(answer) ? answer : 'answered in their own words'
498  $.ui.log(`kit: approval dialog for ${needed.names.join(', ')}: ${outcome}`, { to: 'debug' })
499  if (answer === null) return blocked
500  if (answer !== ALLOW_ONCE && answer !== allowTurn) {
501    const said = answer === KEEP_BLOCKED ? '' : ` They answered: ${JSON.stringify(answer)}.`
502    return { deny: `${blocked.text}\nThe user was asked in a dialog and kept it blocked.${said} Don't ask them to approve it again this turn.` }
503  }
504  const granted = await approveHelper($, kit, [answer === ALLOW_ONCE ? 'once' : 'grant', session, ...needed.names])
505  if (granted.failure) {
506    $.ui.log("kit: couldn't record your approval, so the call stays blocked: " + granted.failure)
507    return blocked
508  }
509  try {
510    return await next(e)
511  } finally {
512    if (answer === ALLOW_ONCE) {
513      const revoked = await approveHelper($, kit, ['revoke', session, ...needed.names])
514      if (revoked.failure) $.ui.log(`kit: couldn't take back the one-time approval of ${needed.scope}, so it lasts until your next message: ` + revoked.failure)
515    }
516  }
517}
518
519export function register(on) {
520  on('tool.call', { tool: ['Bash', /^mcp__/] }, askToLiftBlock)
521
522  on('session.start', async ($, e, next) => {
523    const flow = FEATURES.find((feature) => feature.command === 'flow')
524    await $.command.register({ name: 'flow', description: flow.description, argumentHint: '[off|on <phase> | review-rounds <0-9> [--repo|--global]]', immediate: true })
525    const waited = new AbortController()
526    const registering = registerReviewers($)
527      .catch((error) => $.ui.log('kit: review agents: ' + messageOf(error), { to: 'debug' }))
528      .finally(() => waited.abort())
529    const timer = $.clock.sleep(REVIEWERS_WAIT_MS, { signal: waited.signal }).then(() => 'timed out', () => 'registered')
530    if ((await Promise.race([registering.then(() => 'registered'), timer])) === 'timed out') {
531      await $.ui.log('kit: review agents: still registering after 2 s, so the session started without them', { to: 'debug' })
532    }
533    return next(e)
534  })
535
536  on('command.run', { command: 'flow' }, async ($, e) => {
537    if (/^(off|on|review-rounds)(\s|$)/i.test(e.args.trim())) return switchPhase($, e)
538    if ((await $.session.surfaces()).length === 0) return { text: asText(await collectOrSayWhy($)) }
539    // Without closeOnEscape: Escape hands the keys back to the prompt and the pane stays, refreshing after each turn.
540    await $.ui.open({ id: PANE, title: 'flow', focus: true })
541    void gather($)
542    return {}
543  })
544
545  on('turn.complete', async ($, e, next) => {
546    if (!e.agentId) refreshIfOpen($, 'the turn')
547    return next(e)
548  })
549
550  on('ui.render', { component: 'Pane' }, async ($, e, next) => {
551    if (e.requestId !== PANE) return next(e)
552    const { Box, Text, Button } = $.ui.resolve(e)
553    const line = (text) => Text({ bold: NEEDS_ATTENTION.test(text), wrap: 'wrap', children: [clip(text)] })
554    const refreshing = shown && gathering ? ' · refreshing…' : ''
555    const top = !shown
556      ? [Text({ children: ['Gathering…'] })]
557      : shown.none
558        ? [line(shown.none + refreshing)]
559        : [Text({ bold: true, children: [clip(shown.heading + refreshing)] })]
560    const sections = (shown?.sections ?? []).flatMap(([title, lines]) => [
561      Text({ children: [' '] }),
562      Text({ dimColor: true, children: [title] }),
563      ...lines.map((text, i) => (title === 'Flow' && i === 0 ? Text({ bold: true, children: [clip(text)] }) : line(text))),
564    ])
565    // A press starts the work and returns: Claude Code skips a hook still running after 10 s.
566    // Run verify only where there is a change: elsewhere a press would do nothing.
567    const buttons = [
568      ...(shown?.root ? [Button({ key: 'run-verify', label: 'Run verify', hotkey: 'v', plain: true, onPress: () => void runVerify($) })] : []),
569      Button({ key: 'refresh', label: 'Refresh', hotkey: 'r', plain: true, onPress: () => void gather($) }),
570    ]
571    // One row per step before the merge, each with its own Pass and Fail: what a press records is the user's.
572    // The buttons lead the row, so they line up whatever the step's line holds.
573    const staging = shown?.staging
574    const failure = markFailure && markFailure.root === shown?.root ? markFailure.text : null
575    const steps = staging?.steps ?? []
576    const stagingRows = !staging || (!staging.failure && !steps.length && !failure && !pendingMarks.length) ? [] : [
577      Text({ children: [' '] }),
578      Text({ dimColor: true, children: ['Staging before the merge'] }),
579      ...(pendingMarks.length ? [Text({ dimColor: true, children: ['Marking ' + pendingMarks.join(', ') + '…'] })] : []),
580      ...(failure ? [line(failure)] : []),
581      ...(staging.failure ? [line(staging.failure)] : steps.map((step) => Box({
582        flexDirection: 'row',
583        columnGap: 2,
584        children: [
585          ...['PASS', 'FAIL'].map((result) => Button({
586            key: `staging-${step.id}-${result}`,
587            label: result === 'PASS' ? 'Pass' : 'Fail',
588            plain: true,
589            onPress: () => void markStep($, step.id, result),
590          })),
591          Text({ bold: step.result === 'FAIL', wrap: 'wrap', children: [clip(stepLine(step))] }),
592        ],
593      }))),
594    ]
595    // The buttons and the run they started come first: a long report scrolls the bottom of the pane away.
596    return Box({
597      flexDirection: 'column',
598      children: [
599        ...top,
600        Box({ flexDirection: 'row', columnGap: 2, children: buttons }),
601        ...verifyLines().map(line),
602        ...stagingRows,
603        ...sections,
604      ],
605    })
606  })
607}
608
hooks/features.js 41 lines
1// What each command, hook and agent type of the mod offers, and what Codex has instead: bin/docs prints it in
2// docs/framework.md, and tests/flow.test.ts fails when the mod registers a command or agent, or draws a button,
3// this list doesn't name.
4// Kept as one JSON value after the `=`, which bin/docs reads without running JavaScript.
5export const FEATURES = [
6  {
7    "command": "flow",
8    "description": "Show where this branch is in the kit's flow, run verify without a turn, and switch a phase off or on or set the self-review's re-check rounds",
9    "capabilities": [
10      {"name": "Phase and reports", "codex": "`bin/reports brief`"},
11      {"name": "The phases switched off, with their scope, also on the default branch", "codex": "`bin/phase switches`"},
12      {"name": "`/flow off|on <phase> [--repo|--global]`, typed at the prompt, switches a phase for the branch, the repo or every repo", "codex": "a message whose first line is `phase off|on <phase> [branch|repo|global]`"},
13      {"name": "`/flow review-rounds <0-9> [--repo|--global]`, typed at the prompt, sets how many re-checks the self-review runs after its first pass (1 when unset)", "codex": "a message whose first line is `review-rounds <0-9> [branch|repo|global]`; `bin/phase review-rounds` reads it"},
14      {"name": "Follows the branch into the main checkout once validate detaches the session's own", "codex": "the commands of this list, run from the main checkout"},
15      {"name": "Stamps (verify, self-review, validate, staging), current or not, and a verify that checked nothing", "codex": "`python3 ~/.agents/hooks/review_stamp.py check --kind <kind>`"},
16      {"name": "The last verify report's ran/skipped/warning/error lines", "codex": "`cat \"$(~/.agents/bin/reports path verify)\"`"},
17      {"name": "CI of the pushed HEAD", "codex": "`bin/ci-wait --once --no-log`"},
18      {"name": "Context use", "codex": "nothing (Codex shows its own)"},
19      {"name": "Run verify button, without a turn, on the checkout the pane shows", "codex": "`python3 ~/.agents/hooks/stop_checks.py verify` in a terminal"},
20      {"name": "Refresh button, and a refresh after each turn while the pane is open", "codex": "running the commands above again"},
21      {"name": "The staging guide's steps before the merge, each with its latest result, who gave it, and its evidence files", "codex": "`bin/staging steps --json`"},
22      {"name": "Pass button on each step before the merge, recorded as the user's verdict", "codex": "`bin/staging mark <step> PASS --by user` in a terminal (the shell guard refuses it from the agent)"},
23      {"name": "Fail button on each step before the merge, recorded as the user's verdict", "codex": "`bin/staging mark <step> FAIL --by user` in a terminal"}
24    ]
25  },
26  {
27    "hook": "tool.call",
28    "description": "Ask the user in a dialog when a kit guard blocks a call they could approve",
29    "capabilities": [
30      {"name": "A blocked MCP write or `DROP`/`TRUNCATE` command shown in a dialog: allowed once, until the user's next message, or kept blocked", "codex": "the user's next message naming the service or opening with `allow <statement>`, or `touch ~/.agents/approvals/<service>`"}
31    ]
32  },
33  {
34    "agentPrefix": "review",
35    "description": "One reviewer agent type per self-review lens",
36    "capabilities": [
37      {"name": "One for each lens `bin/triage --lens-briefs` prints: its brief from lenses.md, no edit tools, no CLAUDE.md block", "codex": "a subagent given the lens text from `skills/self-review/lenses.md`"}
38    ]
39  }
40]
41