SLOPSHOPPER

bughunt

Runs proof-driven bug hunt rounds with /bughunt: each round proves one bug with a failing command the mod runs itself, fixes it, and proves the fix, and the…

newguardcommandprompttoolprocess
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · bughunt
› fix the failing auth test and add an audit log call ● bughunt: round 1/1 started ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /bughunt ⎿ bughunt: hunt started: 1 round(s) over the whole project ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts
README

bughunt

A bug hunt prompt tells the model to prove a bug before it fixes it, and the model often writes the fix first and a test that passes after. This mod runs the hunt in rounds and holds each round to its proof. The model writes a proof command, the mod runs it itself, and production code stays unchanged until the proof exits non-zero with a FAIL line. The round counts as fixed only when the same command then exits 0 with a PASS line, and fails again once the mod reverts the fix for a moment. The loop reads each round's outcome and stops at a blocked or unverified round instead of starting the next one.

/bughunt collab <paths> is a second, read-only mode: a scanner, a planner and a critic subagent review the paths in turn, and the report carries the findings, the plan and the critic's verdict.

What it does

Rounds

  1. /bughunt [--rounds N] [target] starts a hunt of 1 to 25 rounds over the target, or over the whole project. --rounds may stand anywhere among the arguments. The hunt starts at any point of the session.
  2. The mod sends each round as a prompt: the round number, the scope, the round's proof directory (.temp_files/bughunt/<round>/), the fingerprints of the earlier rounds and the protocol. The model first opens the bughunt:hunt skill, which holds the full rules; while a round runs, the mod appends the round's block to the skill's text.
  3. While a round runs:
  4. Edit, Write and NotebookEdit stop until the skill is open in the round.
  5. An edit outside the proof directory stops until the mod recorded a FAIL.
  6. With a target, an edit outside it stops after the FAIL too. A test file (tests/, __tests__/, *.test.*, *.spec.*, *_test.*, test_*.py) passes, so the regression test can go into the suite.
  7. Every subagent spawn stops: a round runs in one conversation.
  8. The model calls mcp__bughunt__proof with phase: "before" and the proof command's argv. The mod runs the command (5 minutes at most) and records FAIL only when it exits non-zero and prints a line that starts with FAIL. A setup or import error that prints no such line is rejected. After the fix, phase: "after" with the same argv runs the command again and needs exit 0 and a line that starts with PASS. The model reads the exit code, the last 20 lines of output and the reason.
  9. Before it records the PASS, the mod checks that the proof reaches the fix. When the FAIL was recorded, it took a snapshot of the working tree (git stash create, which adds no stash entry, or HEAD on a clean tree). After an accepted PASS run it lists the files modified since that snapshot (git diff --diff-filter=M), leaves out the proof directory and test files, and restores them to the snapshot in the working tree (git restore --source, the index stays as it was). It runs the proof again and then puts the fix back. The PASS counts only when that reverted run exits non-zero with a FAIL line. Otherwise the call is rejected and the model can fix the proof and call again. The model reads the reason in each case:
  10. the proof passes with the fix reverted, so it does not reach the fixed code;
  11. no production file was modified since the FAIL;
  12. the directory is not a git repository;
  13. a git command failed.

While the fix is reverted, the mod keeps a record in $.store. When a crash cuts the check short, the next session in the same directory puts the fix back, if the reverted files are still unchanged. If they changed since, it does not overwrite them and tells you the git restore command that brings the fix back.

  1. When the round's turn ends, the mod reads the answer's outcome line: the first line that begins with an outcome label, so a sentence before it does not hide it:
  2. fixed-and-verified goes on only when the mod recorded FAIL then PASS in the round; otherwise the hunt stops.
  3. no-proven-bug goes on.
  4. blocked, fixed-verification-incomplete, no outcome line, an interrupt or an API error stop the hunt.
  5. After the last round the hunt ends.

The fingerprint: line of each answer goes into the next rounds' prompts, so the same root cause is not counted twice.

  1. A prompt you write yourself ends the hunt; /bughunt commands do not.

Collab

  1. /bughunt collab <paths>, or the model's mcp__bughunt__collab tool, starts three subagents in turn. They are the mod's own agent types (bughunt:scanner, bughunt:planner, bughunt:critic), hidden from the model's agent list, and each can only read: Read, Grep, Glob.
  2. The scanner reports each finding at once with mcp__bughunt__found (file, line, severity, description, optional fix). The mod checks the fields and keeps the finding. Only the running scanner may call the tool.
  3. The planner receives the findings and the scanner's report, and writes a fix plan. The critic receives the findings and the plan, and begins its answer with verdict: approve, revise or reject.
  4. Each step has a time limit (scanner 10, planner 8, critic 6 minutes). A step that runs out is named in the report as timed-out, and the findings the scanner sent before that stay in the report.
  5. When the critic gives no verdict line, the verdict is no-verdict, never approve.
  6. The command and the tool return once the scanner started; the report arrives later as one message, and the model reads it as a read-only review. A step's hand-back message is taken by the mod and dropped, so it does not start a turn of its own.
  7. A collab does not start while a round runs.

What you see

The sidebar shows a standing bughunt section: the round, whether the skill is open, the proof state and each finished round's outcome and fingerprint. Stopped edits, proof results and the hunt's end go to the sidebar stream. Without the sidebar, each of them is one transcript line such as bughunt: edit stopped (proof): src/a.ts.

Command

/bughunt [--rounds N] [target] start a hunt of N rounds (1 by default, 25 at most) /bughunt collab <paths> a read-only scanner, planner and critic review /bughunt stop end the running hunt /bughunt status on or off, and the running round /bughunt on | off on by default; off starts nothing and holds no edit

Install

claude plugin marketplace add KilimcininKorOglu/claude-code-mods claude plugin install bughunt@kilimcininkoroglu-mods

Function hooks are early access. Claude Code 2.1.288 and later load them by default, so there is nothing to switch on.

After installing

  1. Restart Claude Code.
  2. Run /bughunt in a repository whose tests you can run from the command line. The proof command runs with your permissions, so read what the model proposes.

What it can reach

Validated with claude plugin validate on Claude Code 2.1.284:

❯ ./register.ts hooks: session.start, command.run{command=bughunt}, agent.offer{agent=/"^bughunt:(scanner|planner|critic)$"/}, tool.describe{tool=/"^mcp__bughunt__(proof|found|collab)$"/}, tool.call{tool=/"^mcp__bughunt__proof$"/}, tool.call{tool=/"^mcp__bughunt__found$"/}, tool.call{tool=/"^mcp__bughunt__collab$"/}, prompt.submit, skill.prompt{skill=bughunt:hunt}, tool.call{tool=Skill}, tool.call{tool=Edit}, tool.call{tool=Write}, tool.call{tool=NotebookEdit}, agent.spawn, turn.complete ❯ ./register.ts calls: $.agent.register (via declare), $.agent.spawn (via portsOf), $.clock.after (via portsOf, send), $.command.register (via declare), $.command.run (via send), $.process.run (via git, recoverFix, runProof), $.prompt.submit (via send), $.sidebar.clear (via show), $.sidebar.set (via show, toPerson), $.store.delete (via putBack, recoverFix), $.store.get (via readSettings, recoverFix), $.store.set (via revertCheck, setEnabled), $.tool.register (via declare), $.ui.log (via launchCollab, send, toPerson)

Reach L2, it runs the proof command the model names.

  1. Reads: the path of each Edit, Write and NotebookEdit call; your prompts, only to see whether you wrote one; each round's final answer; the collab subagents' answers
  2. Runs: the proof command the model passes to mcp__bughunt__proof, as argv without a shell, in the working directory or the cwd it names, for 5 minutes at most, a second time with the fix reverted; git stash create, rev-parse, diff and restore --worktree in the working directory for that check; three read-only subagents for a collab
  3. Sends: each round's prompt and each collab report to the model as a message, a deny text for a stopped edit or spawn, the round's block after the skill's text, sidebar sections and lines or transcript lines to you; nothing leaves the machine
  4. Persists: in $.store, the on/off setting, and while a revert check runs the snapshot that holds the fix and the reverted files; the hunt itself lives in memory and ends with the session
  5. Hostile input: the proof command is the model's and runs with your permissions, as a Bash call would, but without a shell; a finding's fields are checked before they are kept

Limits

  • The gate reads Edit, Write and NotebookEdit. A file changed through Bash (sed -i, a redirect, a script) is not held.
  • The revert check shows that the proof depends on the files the fix modified. It cannot tell whether the proof asserts the right behaviour.
  • The revert check reverts only tracked files the fix modified. A fix that only adds files has nothing to revert, so its PASS is rejected. Outside a git repository, no PASS is recorded.
  • While the check runs (at most one more proof run), the fixed files hold the old code. Another process that reads them in that window sees the bug.
  • The mod measures that the skill was delivered, not that the model read it.
  • A collab step that runs out of time keeps running in the background until it ends; the mod no longer waits for it.
  • The test engine cannot start a subagent, so the collab waits (hand-back, early answer, time limit, failed end) are tested in pipeline.ts with fake engine calls, and the whole run is checked live. Live on 2.1.284: a two-round hunt recorded FAIL then PASS, carried the fingerprint into round 2 and ended there; a collab kept the scanner's finding, read the critic's verdict, and its three hand-backs started no turn. A round in a git repository recorded PASS after the reverted run failed, the fix was back in place afterwards, and git stash list stayed empty.
  • There is no bypass of a round's gates. /bughunt stop ends the hunt, and /bughunt off turns the mod off.

Development

make install # eslint, typescript-eslint, typescript make lint # complexity limit 10, the build fails above it make typecheck # needs .claude/types/ from /plugin-types make validate make test # claude plugin test

Source 8 files
hooks/register.ts 418 lines
1import type { EngineInterface, Register } from 'claude-code'
2import { parseArgs } from './args.ts'
3import { COLLAB_SCHEMA, CRITIC_PROMPT, findingOf, FOUND_SCHEMA, PLANNER_PROMPT, SCANNER_PROMPT, type Step } from './collab.ts'
4import { runCollab, settleAgent, takeHandBack, type Ports, type Waits } from './pipeline.ts'
5import { editRule, proofDir, relativeTo } from './paths.ts'
6import { judgeProof, judgeReverted, proofInput, revertTargets, savedFix, tailOf, type Phase, type ProofJudgement, type ProofRun, type SavedFix } from './proof.ts'
7import { advance, decide, newHunt, outcomeOf, roundId, type Hunt, type RoundEnd } from './round.ts'
8import { editDenyText, editLog, roundLine, roundText, SKILL, skillBlock, SPAWN_DENY } from './texts.ts'
9
10const ENABLED_KEY = 'enabled'
11const SEND_COMMAND = 'bughunt:send'
12const FOUND_TOOL = 'mcp__bughunt__found'
13const PROOF_MS = 300_000
14const GIT_MS = 30_000
15const RESTORE_KEY = 'restore'
16const MAX_TURNS: Record<Step, number> = { scanner: 200, planner: 120, critic: 80 }
17const SECTION = { consumer: 'bughunt', key: 'hunt', order: 22 } as const
18
19/** Prompt origins that are the person's own words. */
20const PERSON = new Set(['composer', 'bridge', 'sdk'])
21
22type Line = { text: string; kind: 'ok' | 'warn' | 'error' | 'info' | 'dim' }
23
24/** The on/off setting as the store held it at the last read; the running hunt; the working directory; the collab waits. */
25type State = Waits & {
26  enabled: boolean
27  hunt: Hunt | undefined
28  cwd: string
29  lastHunt: string[]
30}
31
32async function readSettings($: EngineInterface, state: State): Promise<void> {
33  state.enabled = (await $.store.get(ENABLED_KEY)) !== false
34}
35
36/** One line for the person: a stream entry in the sidebar while it is open, else the transcript line. */
37async function toPerson($: EngineInterface, text: string, kind: 'ok' | 'warn' | 'error' | 'info'): Promise<void> {
38  try {
39    if (await $.sidebar.set({ consumer: SECTION.consumer, key: 'log', title: 'bughunt', lines: [{ text, kind }], until: 'stream' })) return
40  } catch {
41    // The sidebar mod is not installed.
42  }
43  $.ui.log(text)
44}
45
46function huntLines(hunt: Hunt): Line[] {
47  const proof = hunt.proof === undefined ? 'no proof yet' : hunt.proof.passed ? 'FAIL then PASS recorded' : 'FAIL recorded'
48  return [
49    { text: `round ${hunt.round}/${hunt.rounds}${hunt.target === '' ? '' : ` · ${hunt.target}`}`, kind: 'info' },
50    { text: `skill ${hunt.skill ? 'open' : 'not open'} · ${proof}`, kind: hunt.proof?.failed === true ? 'ok' : 'warn' },
51    ...hunt.history.map(r => ({ text: `r${r.round}: ${r.outcome}${r.fingerprint === undefined ? '' : ` · ${r.fingerprint}`}`, kind: r.outcome === 'fixed-and-verified' ? ('ok' as const) : ('dim' as const) })),
52  ]
53}
54
55/** The standing section: the running hunt or collab, else the last hunt's summary; cleared when there is none. */
56async function show($: EngineInterface, state: State): Promise<void> {
57  const lines = state.hunt !== undefined ? huntLines(state.hunt) : state.collab !== undefined ? [{ text: `collab: ${state.collab.step} · ${state.collab.findings.length} finding(s)`, kind: 'info' as const }] : state.lastHunt.map(text => ({ text, kind: 'dim' as const }))
58  try {
59    if (lines.length === 0) await $.sidebar.clear({ consumer: SECTION.consumer, key: SECTION.key })
60    else await $.sidebar.set({ ...SECTION, title: state.hunt === undefined ? 'bughunt' : 'bughunt running', lines, until: 'session' })
61  } catch {
62    // The sidebar mod is not installed; the stream lines reach the transcript instead.
63  }
64}
65
66/** Sends a round or a report as the person's own words, through the send command; a refusal falls back to a plain prompt. */
67function send($: EngineInterface, text: string): void {
68  $.clock.after(0, () => {
69    $.command.run({ command: SEND_COMMAND, args: text }).catch(async (err: unknown) => {
70      $.ui.log(`the send command failed (${err instanceof Error ? err.message : String(err)}), submitting the prompt instead`)
71      await $.prompt.submit({ text }).catch((e: unknown) => $.ui.log(`the prompt was not submitted: ${e instanceof Error ? e.message : String(e)}`))
72    })
73  })
74}
75
76async function startRound($: EngineInterface, state: State, hunt: Hunt): Promise<void> {
77  state.hunt = hunt
78  send($, roundText(hunt, proofDir(roundId(hunt))))
79  await toPerson($, roundLine(hunt), 'info')
80  await show($, state)
81}
82
83async function endHunt($: EngineInterface, state: State, text: string, kind: 'ok' | 'error' | 'warn'): Promise<void> {
84  const hunt = state.hunt
85  state.hunt = undefined
86  state.lastHunt = hunt === undefined ? [] : [text, ...hunt.history.map(r => `r${r.round}: ${r.outcome}${r.fingerprint === undefined ? '' : ` · ${r.fingerprint}`}`)]
87  await toPerson($, text, kind)
88  await show($, state)
89}
90
91/** Reads a round's end and starts the next round, or ends the hunt saying why. */
92async function onRoundEnd($: EngineInterface, state: State, hunt: Hunt, end: RoundEnd): Promise<void> {
93  const verdict = decide(hunt, end)
94  const next = advance(hunt, end.answer)
95  state.hunt = { ...hunt, history: next.history }
96  if (verdict.next === 'stop') return endHunt($, state, `hunt stopped after round ${hunt.round}/${hunt.rounds}: ${verdict.reason}`, 'error')
97  if (verdict.next === 'done') return endHunt($, state, `hunt finished: ${hunt.rounds} round(s), last ${outcomeOf(end.answer) ?? 'none'}`, 'ok')
98  await startRound($, state, next)
99}
100
101async function startHunt($: EngineInterface, state: State, rounds: number, target: string): Promise<string> {
102  if (!state.enabled) return 'off: /bughunt on turns it on'
103  if (state.hunt !== undefined) return `a hunt is running (round ${state.hunt.round}/${state.hunt.rounds}); /bughunt stop ends it`
104  if (state.collab !== undefined) return 'a collab run is in progress; wait for its report'
105  await startRound($, state, newHunt(Date.now().toString(36), rounds, target))
106  return `hunt started: ${rounds} round(s) over ${target === '' ? 'the whole project' : target}`
107}
108
109async function setEnabled($: EngineInterface, state: State, on: boolean): Promise<string> {
110  await $.store.set(ENABLED_KEY, on)
111  state.enabled = on
112  if (!on && state.hunt !== undefined) await endHunt($, state, 'hunt stopped: the mod was turned off', 'warn')
113  return on ? 'on: /bughunt starts a hunt' : 'off: /bughunt starts nothing and no gate holds edits'
114}
115
116function statusText(state: State): string {
117  const onOff = state.enabled ? 'on' : 'off'
118  if (state.hunt !== undefined) return `${onOff} · round ${state.hunt.round}/${state.hunt.rounds}${state.hunt.target === '' ? '' : ` in ${state.hunt.target}`}`
119  if (state.collab !== undefined) return `${onOff} · collab ${state.collab.step}`
120  return `${onOff} · no hunt running`
121}
122
123async function runCommand($: EngineInterface, state: State, args: string): Promise<string> {
124  await readSettings($, state)
125  const p = parseArgs(args)
126  if (p.kind === 'error') return p.text
127  if (p.kind === 'start') return startHunt($, state, p.rounds, p.target)
128  if (p.kind === 'collab') return startCollabCommand($, state, p.paths)
129  if (p.kind === 'on' || p.kind === 'off') return setEnabled($, state, p.kind === 'on')
130  if (p.kind === 'status') return statusText(state)
131  if (state.hunt === undefined) return 'no hunt is running'
132  await endHunt($, state, `hunt stopped by /bughunt stop in round ${state.hunt.round}/${state.hunt.rounds}`, 'warn')
133  return 'hunt stopped'
134}
135
136/** The edit gate of a running round: the rule an edit of `path` breaks, told to the model and the person. */
137async function gateEdit($: EngineInterface, state: State, path: string | undefined, agentId: string | undefined): Promise<string | undefined> {
138  const hunt = state.hunt
139  if (hunt === undefined || agentId !== undefined || path === undefined) return undefined
140  const dir = proofDir(roundId(hunt))
141  const rel = relativeTo(state.cwd, path)
142  const rule = editRule(rel, { skill: hunt.skill, failed: hunt.proof?.failed === true, target: hunt.target, dir })
143  if (rule === undefined) return undefined
144  await toPerson($, editLog(rule, rel), 'error')
145  return editDenyText(rule, dir, hunt.target)
146}
147
148function joinCwd(cwd: string, dir: string | undefined): string {
149  if (dir === undefined || dir === '') return cwd
150  return dir.startsWith('/') ? dir : `${cwd.replace(/\/+$/, '')}/${dir}`
151}
152
153async function runProof($: EngineInterface, argv: string[], cwd: string): Promise<ProofRun | string> {
154  try {
155    const r = await $.process.run(argv, { cwd, timeoutMs: PROOF_MS })
156    return { exitCode: r.exitCode, stdout: r.stdout, stderr: r.stderr }
157  } catch (err) {
158    return `the proof did not run: ${err instanceof Error ? err.message : String(err)}`
159  }
160}
161
162/** Runs one git command in `cwd` and answers its stdout; a non-zero exit throws with git's own message. */
163async function git($: EngineInterface, cwd: string, args: string[]): Promise<string> {
164  const r = await $.process.run(['git', ...args], { cwd, timeoutMs: GIT_MS })
165  if (r.exitCode !== 0) throw new Error(`git ${args[0]} exited ${r.exitCode}: ${r.stderr.trim()}`)
166  return r.stdout
167}
168
169/** A commit that holds the working tree's tracked files as they are now: a stash commit, or HEAD on a clean tree. */
170async function snapshot($: EngineInterface, cwd: string): Promise<string> {
171  const stash = (await git($, cwd, ['stash', 'create'])).trim()
172  return stash !== '' ? stash : (await git($, cwd, ['rev-parse', 'HEAD'])).trim()
173}
174
175/** The snapshot a recorded FAIL keeps, or undefined outside a git repository; the round's first FAIL wins. */
176async function baseOf($: EngineInterface, state: State, hunt: Hunt): Promise<string | undefined> {
177  if (hunt.proof?.base !== undefined) return hunt.proof.base
178  try {
179    return await snapshot($, state.cwd)
180  } catch {
181    // No git repository: the revert check refuses the PASS later and says why.
182    return undefined
183  }
184}
185
186const restoreKey = (cwd: string): string => `${RESTORE_KEY}:${cwd}`
187
188/** Puts the fix back and forgets the record; a failure throws, so the person hears where the fix is. */
189async function putBack($: EngineInterface, cwd: string, saved: SavedFix): Promise<void> {
190  await git($, cwd, ['restore', `--source=${saved.fixed}`, '--worktree', '--', ...saved.files])
191  await $.store.delete(restoreKey(cwd))
192}
193
194/** Reverts the fix, runs the proof again and puts the fix back; undefined when the reverted run failed as it must. */
195async function revertCheck($: EngineInterface, state: State, hunt: Hunt, argv: string[], proofCwd: string): Promise<string | undefined> {
196  const base = hunt.proof?.base
197  if (base === undefined) return 'the revert check needs a git repository, and none answered when the FAIL was recorded'
198  const changed = (await git($, state.cwd, ['diff', '--name-only', '-z', '--diff-filter=M', '--relative', base])).split('\0')
199  const files = revertTargets(changed, proofDir(roundId(hunt)))
200  if (files.length === 0) return 'no production file was modified since the FAIL, so there is no fix to revert'
201  const saved: SavedFix = { base, fixed: await snapshot($, state.cwd), files }
202  await $.store.set(restoreKey(state.cwd), saved)
203  try {
204    await git($, state.cwd, ['restore', `--source=${base}`, '--worktree', '--', ...files])
205    const reverted = await runProof($, argv, proofCwd)
206    return typeof reverted === 'string' ? `the reverted run failed to start: ${reverted}` : judgeReverted(reverted, files)
207  } finally {
208    await putBack($, state.cwd, saved)
209  }
210}
211
212/** The PASS a judge accepted stands only when the revert check holds; its refusal or failure rejects it. */
213async function checkedPass($: EngineInterface, state: State, hunt: Hunt, argv: string[], proofCwd: string): Promise<string | undefined> {
214  try {
215    return await revertCheck($, state, hunt, argv, proofCwd)
216  } catch (err) {
217    const why = err instanceof Error ? err.message : String(err)
218    await toPerson($, `revert check failed: ${why}`, 'error')
219    return `the revert check failed: ${why}`
220  }
221}
222
223/**
224 * A revert check a crash cut short left the fix reverted: at the session's start the fix is put back when the
225 * files still hold the reverted text, else the person is told where the fix is, because a later edit would be lost.
226 */
227async function recoverFix($: EngineInterface, cwd: string): Promise<void> {
228  const saved = savedFix(await $.store.get(restoreKey(cwd)))
229  if (saved === undefined) return
230  try {
231    const untouched = await $.process.run(['git', 'diff', '--quiet', saved.base, '--', ...saved.files], { cwd, timeoutMs: GIT_MS })
232    if (untouched.exitCode === 0) {
233      await putBack($, cwd, saved)
234      await toPerson($, `put back the fix a cut-short revert check left reverted: ${saved.files.join(', ')}`, 'warn')
235      return
236    }
237    await $.store.delete(restoreKey(cwd))
238    await toPerson($, `a cut-short revert check left ${saved.files.join(', ')} reverted, and they changed since; the fix is in ${saved.fixed}: git restore --source=${saved.fixed} --worktree -- ${saved.files.join(' ')}`, 'error')
239  } catch (err) {
240    await toPerson($, `the fix a revert check left reverted was not put back (${err instanceof Error ? err.message : String(err)}); it is in ${saved.fixed}`, 'error')
241  }
242}
243
244/** The judgement of one proof run, with the base a FAIL keeps and the revert check an accepted PASS needs. */
245async function judged($: EngineInterface, state: State, hunt: Hunt, input: { phase: Phase; argv: string[]; cwd: string }, run: ProofRun): Promise<ProofJudgement> {
246  const j = judgeProof(input.phase, input.argv, run, hunt.proof)
247  if (!j.accepted || j.state === undefined) return j
248  if (input.phase === 'before') return { ...j, state: { ...j.state, base: await baseOf($, state, hunt) } }
249  const refused = await checkedPass($, state, hunt, input.argv, input.cwd)
250  return refused === undefined ? j : { accepted: false, reason: refused, state: hunt.proof }
251}
252
253/** The proof tool: runs the command itself and records FAIL or PASS for the round. */
254async function onProof($: EngineInterface, state: State, e: Record<string, unknown>): Promise<{ result: string } | { deny: string }> {
255  const hunt = state.hunt
256  if (hunt === undefined || e.agentId !== undefined) return { deny: 'The proof tool works only inside a /bughunt round, in the main conversation.' }
257  const input = proofInput(e)
258  if (typeof input === 'string') return { deny: input }
259  const proofCwd = joinCwd(state.cwd, input.cwd)
260  const run = await runProof($, input.argv, proofCwd)
261  if (typeof run === 'string') return { result: `rejected: ${run}` }
262  const j = await judged($, state, hunt, { phase: input.phase, argv: input.argv, cwd: proofCwd }, run)
263  state.hunt = { ...hunt, proof: j.state }
264  await toPerson($, `proof ${input.phase}: ${j.accepted ? j.reason : `rejected, ${j.reason}`}`, j.accepted ? 'ok' : 'warn')
265  await show($, state)
266  return { result: `${j.accepted ? 'accepted' : 'rejected'}: ${j.reason}\nexit code ${run.exitCode}\n--- last lines ---\n${tailOf(run)}` }
267}
268
269function collabRefusal(state: State): string | undefined {
270  if (!state.enabled) return 'bughunt is off; /bughunt on turns it on'
271  if (state.hunt !== undefined) return 'a hunt round is running; collab waits until it ends'
272  if (state.collab !== undefined) return 'a collab run is in progress'
273  return undefined
274}
275
276/**
277 * Starts collab and returns once the scanner started. A hook has a 10 s budget, so the run goes on after it
278 * and its report arrives through the send command. The scanner is spawned inside the calling hook's dispatch,
279 * because a subagent spawned outside one skips this mod's own tool.call hooks and its found calls would go
280 * unanswered (measured on 2.1.284).
281 */
282async function launchCollab($: EngineInterface, state: State, paths: string[]): Promise<void> {
283  let launched = () => {}
284  const started = new Promise<void>(resolve => { launched = resolve })
285  runCollab(state, portsOf($, state), paths, launched)
286    .then(report => send($, report))
287    .catch((err: unknown) => $.ui.log(`collab failed: ${err instanceof Error ? err.message : String(err)}`))
288    .finally(launched)
289  await started
290}
291
292/** The engine calls the collab pipeline makes. */
293function portsOf($: EngineInterface, state: State): Ports {
294  return {
295    spawn: (step, prompt) => $.agent.spawn({ subagentType: `bughunt:${step}`, description: `bughunt ${step}`, prompt }),
296    after: (ms, fn) => $.clock.after(ms, fn),
297    show: () => show($, state),
298    tell: (text, kind) => toPerson($, text, kind),
299  }
300}
301
302async function startCollabCommand($: EngineInterface, state: State, paths: string[]): Promise<string> {
303  const refused = collabRefusal(state)
304  if (refused !== undefined) return refused
305  await launchCollab($, state, paths)
306  return `collab started over ${paths.join(', ')}; the report arrives as a message`
307}
308
309async function onCollabTool($: EngineInterface, state: State, e: Record<string, unknown>): Promise<{ result: string } | { deny: string }> {
310  await readSettings($, state)
311  const refused = collabRefusal(state)
312  if (refused !== undefined) return { deny: refused }
313  const paths = Array.isArray(e.paths) ? e.paths.filter((p): p is string => typeof p === 'string' && p !== '') : []
314  if (paths.length === 0) return { deny: 'paths must name at least one file or directory' }
315  await launchCollab($, state, paths)
316  return { result: `collab started over ${paths.join(', ')}; the report arrives as a message` }
317}
318
319async function onFound($: EngineInterface, state: State, e: Record<string, unknown>): Promise<{ result: string } | { deny: string }> {
320  const run = state.collab
321  if (run === undefined || run.scanner === undefined || e.agentId !== run.scanner) return { deny: 'Only the running bughunt scanner reports findings with this tool.' }
322  const f = findingOf(e)
323  if (typeof f === 'string') return { deny: f }
324  run.findings.push(f)
325  await show($, state)
326  return { result: `recorded (${run.findings.length})` }
327}
328
329async function declare($: EngineInterface): Promise<void> {
330  await $.command.register({ name: 'bughunt', description: 'Proof-driven bug hunt rounds, and a read-only collab review (bughunt)', argumentHint: '[--rounds N] [target] | collab <paths> | stop | status | on | off' })
331  await $.tool.register({ name: 'proof', description: 'Runs the round\'s proof command and records FAIL (phase before: non-zero exit and a line starting with FAIL) or PASS (phase after: same argv, exit 0 and a line starting with PASS). Only inside a /bughunt round.', inputSchema: { type: 'object', properties: { phase: { type: 'string', enum: ['before', 'after'] }, argv: { type: 'array', items: { type: 'string' }, minItems: 1 }, cwd: { type: 'string' } }, required: ['phase', 'argv'] } })
332  await $.tool.register({ name: 'found', description: 'Reports one bug finding. Only the bughunt scanner subagent calls it.', inputSchema: FOUND_SCHEMA })
333  await $.tool.register({ name: 'collab', description: 'Runs a read-only bug review of the given paths: a scanner, a planner and a critic subagent in turn, and returns their report with findings, plan and verdict.', inputSchema: COLLAB_SCHEMA })
334  await $.agent.register({ name: 'scanner', description: 'bughunt collab scanner (started by the bughunt mod only)', prompt: SCANNER_PROMPT, tools: ['Read', 'Grep', 'Glob', FOUND_TOOL], skills: [SKILL], maxTurns: MAX_TURNS.scanner })
335  await $.agent.register({ name: 'planner', description: 'bughunt collab planner (started by the bughunt mod only)', prompt: PLANNER_PROMPT, tools: ['Read', 'Grep', 'Glob'], maxTurns: MAX_TURNS.planner })
336  await $.agent.register({ name: 'critic', description: 'bughunt collab critic (started by the bughunt mod only)', prompt: CRITIC_PROMPT, tools: ['Read', 'Grep', 'Glob'], maxTurns: MAX_TURNS.critic })
337}
338
339export const register: Register = on => {
340  const state: State = { enabled: true, hunt: undefined, cwd: '', collab: undefined, waiters: new Map(), early: new Map(), lastHunt: [] }
341
342  on('session.start', async ($, e, next) => {
343    const r = await next(e)
344    state.cwd = e.cwd
345    await declare($)
346    await readSettings($, state)
347    await recoverFix($, e.cwd)
348    return r
349  })
350
351  // The engine prints the plugin name in front of command text and log lines, so the texts do not repeat it.
352  on('command.run', { command: 'bughunt' }, async ($, e) => ({ text: await runCommand($, state, String(e.args ?? '')) }))
353
354  // The collab agents are the mod's own; the model never dispatches them.
355  on('agent.offer', { agent: /^bughunt:(scanner|planner|critic)$/ }, () => ({ isOffered: false }))
356  on('tool.describe', { tool: /^mcp__bughunt__(proof|found|collab)$/ }, async (_, e, next) => ({ ...(await next(e)), isDeferred: false }))
357
358  on('tool.call', { tool: /^mcp__bughunt__proof$/ }, async ($, e) => onProof($, state, e as Record<string, unknown>))
359  on('tool.call', { tool: /^mcp__bughunt__found$/ }, async ($, e) => onFound($, state, e as Record<string, unknown>))
360  on('tool.call', { tool: /^mcp__bughunt__collab$/ }, async ($, e) => onCollabTool($, state, e as Record<string, unknown>))
361
362  // A person's own prompt ends the hunt: the loop runs only on the mod's own round prompts. A collab step's
363  // hand-back arrives here as a peer prompt: the mod takes its report, and the model reads it in the report.
364  on('prompt.submit', async ($, e, next) => {
365    if (e.origin?.kind === 'peer' && takeHandBack(state, e.text)) return { drop: 'bughunt collab step' }
366    if (state.hunt !== undefined && PERSON.has(e.origin?.kind ?? '') && !e.text.trimStart().startsWith('/bughunt')) {
367      await endHunt($, state, `hunt stopped: you wrote a prompt in round ${state.hunt.round}/${state.hunt.rounds}`, 'warn')
368    }
369    return next(e)
370  })
371
372  on('skill.prompt', { skill: 'bughunt:hunt' }, async (_, e, next) => {
373    const r = await next(e)
374    const hunt = state.hunt
375    return hunt === undefined ? r : { text: `${r.text}\n${skillBlock(hunt, proofDir(roundId(hunt)))}` }
376  })
377
378  on('tool.call', { tool: 'Skill' }, async ($, e, next) => {
379    const r = await next(e)
380    if (state.hunt !== undefined && e.skill === SKILL && e.agentId === undefined && r.deny === undefined && r.isError !== true) {
381      state.hunt = { ...state.hunt, skill: true }
382      await show($, state)
383    }
384    return r
385  })
386
387  on('tool.call', { tool: 'Edit' }, async ($, e, next) => {
388    const deny = await gateEdit($, state, e.file_path, e.agentId)
389    return deny === undefined ? next(e) : { deny }
390  })
391  on('tool.call', { tool: 'Write' }, async ($, e, next) => {
392    const deny = await gateEdit($, state, e.file_path, e.agentId)
393    return deny === undefined ? next(e) : { deny }
394  })
395  on('tool.call', { tool: 'NotebookEdit' }, async ($, e, next) => {
396    const deny = await gateEdit($, state, e.notebook_path, e.agentId)
397    return deny === undefined ? next(e) : { deny }
398  })
399
400  // A round runs in one conversation; the mod's own collab spawns skip this hook, as the calling one.
401  on('agent.spawn', async ($, e, next) => {
402    if (state.hunt === undefined) return next(e)
403    await toPerson($, `subagent stopped: ${e.subagentType}`, 'error')
404    return { deny: SPAWN_DENY }
405  })
406
407  // A subagent's answered turn does not carry its report (measured: the hand-back does), so only a failed
408  // end settles a collab step here.
409  on('turn.complete', async ($, e, next) => {
410    const r = await next(e)
411    const end: RoundEnd = { reason: e.reason, isAborted: e.isAborted, answer: e.answer }
412    if (e.agentId !== undefined) {
413      if (e.reason !== 'answer') settleAgent(state, e.agentId, end)
414    } else if (state.hunt !== undefined) await onRoundEnd($, state, state.hunt, end)
415    return r
416  })
417}
418
hooks/args.ts 49 lines
1/** The most rounds one hunt may run. */
2export const MAX_ROUNDS = 25
3
4export type Parsed =
5  | { kind: 'start'; rounds: number; target: string }
6  | { kind: 'collab'; paths: string[] }
7  | { kind: 'stop' | 'status' | 'on' | 'off' }
8  | { kind: 'error'; text: string }
9
10const WORDS = new Set(['stop', 'status', 'on', 'off'])
11
12export const USAGE = 'usage: /bughunt [--rounds 1..25] [target] | collab <paths> | stop | status | on | off'
13
14/** Reads `/bughunt` arguments; `--rounds` may stand anywhere among them. */
15export function parseArgs(raw: string): Parsed {
16  const tokens = raw.trim().split(/\s+/).filter(t => t !== '')
17  const first = tokens[0] ?? ''
18  if (tokens.length === 1 && WORDS.has(first)) return { kind: first as 'stop' | 'status' | 'on' | 'off' }
19  if (first === 'collab') return collabOf(tokens.slice(1))
20  return startOf(tokens)
21}
22
23function collabOf(paths: string[]): Parsed {
24  return paths.length === 0 ? { kind: 'error', text: 'usage: /bughunt collab <paths>' } : { kind: 'collab', paths }
25}
26
27function startOf(tokens: string[]): Parsed {
28  const rest: string[] = []
29  let value: string | undefined
30  for (let i = 0; i < tokens.length; i += 1) {
31    const t = tokens[i]
32    if (t === '--rounds') {
33      value = tokens[i + 1] ?? ''
34      i += 1
35    } else if (t.startsWith('--rounds=')) value = t.slice('--rounds='.length)
36    else rest.push(t)
37  }
38  if (value === undefined) return { kind: 'start', rounds: 1, target: rest.join(' ') }
39  const rounds = roundsOf(value)
40  if (rounds === undefined) return { kind: 'error', text: `--rounds takes a whole number from 1 to ${MAX_ROUNDS}, not "${value}". ${USAGE}` }
41  return { kind: 'start', rounds, target: rest.join(' ') }
42}
43
44function roundsOf(value: string): number | undefined {
45  if (!/^\d+$/.test(value)) return undefined
46  const n = Number(value)
47  return n >= 1 && n <= MAX_ROUNDS ? n : undefined
48}
49
hooks/collab.ts 134 lines
1export const SEVERITIES = ['critical', 'high', 'medium', 'low'] as const
2export type Severity = (typeof SEVERITIES)[number]
3export type Finding = { file: string; line: number; severity: Severity; description: string; suggestedFix?: string }
4
5export type Step = 'scanner' | 'planner' | 'critic'
6export type StepStatus = 'done' | 'timed-out' | 'failed' | 'skipped'
7export type StepResult = { step: Step; status: StepStatus; answer: string; reason?: string }
8export type CollabVerdict = 'approve' | 'revise' | 'reject' | 'no-verdict'
9
10export const FOUND_SCHEMA = {
11  type: 'object',
12  properties: {
13    file: { type: 'string', description: 'Path of the file, as read' },
14    line: { type: 'integer', minimum: 1, description: 'Line the finding is on' },
15    severity: { type: 'string', enum: [...SEVERITIES] },
16    description: { type: 'string', description: 'The trigger and what breaks' },
17    suggestedFix: { type: 'string' },
18  },
19  required: ['file', 'line', 'severity', 'description'],
20} as const
21
22export const COLLAB_SCHEMA = {
23  type: 'object',
24  properties: { paths: { type: 'array', items: { type: 'string' }, minItems: 1, description: 'Files or directories to examine' } },
25  required: ['paths'],
26} as const
27
28type Raw = { file?: unknown; line?: unknown; severity?: unknown; description?: unknown; suggestedFix?: unknown }
29
30/** A finding from the found tool's input, or why it is refused. */
31export function findingOf(e: Raw): Finding | string {
32  const broken = FINDING_CHECKS.find(([ok]) => !ok(e))
33  if (broken !== undefined) return broken[1]
34  const f: Finding = { file: e.file as string, line: e.line as number, severity: e.severity as Severity, description: e.description as string }
35  return typeof e.suggestedFix === 'string' ? { ...f, suggestedFix: e.suggestedFix } : f
36}
37
38const text = (v: unknown) => typeof v === 'string' && v.trim() !== ''
39
40const FINDING_CHECKS: [(e: Raw) => boolean, string][] = [
41  [e => text(e.file), 'file must be a non-empty string'],
42  [e => Number.isInteger(e.line) && (e.line as number) >= 1, 'line must be a whole number from 1'],
43  [e => SEVERITIES.includes(e.severity as Severity), `severity must be one of ${SEVERITIES.join(', ')}`],
44  [e => text(e.description), 'description must be a non-empty string'],
45  [e => e.suggestedFix === undefined || typeof e.suggestedFix === 'string', 'suggestedFix must be a string'],
46]
47
48/**
49 * A subagent's hand-back as the main loop receives it: a peer prompt `<agent-message from="<id>">`, the
50 * engine's frame, then the report indented by two spaces after `The report follows:` (measured on 2.1.284).
51 */
52export function handBackOf(text: string): { from: string; report: string } | undefined {
53  const m = /^\s*<agent-message from="([^"]+)">/.exec(text)
54  if (m === null) return undefined
55  const at = text.indexOf('The report follows:')
56  const body = at < 0 ? text.slice(m[0].length) : text.slice(at + 'The report follows:'.length)
57  const report = body.replace(/<\/agent-message>\s*$/, '').split('\n').map(l => l.replace(/^ {2}/, '')).join('\n').trim()
58  return { from: m[1], report }
59}
60
61/** The `verdict:` line of the critic's answer. */
62export function verdictOf(answer: string): CollabVerdict {
63  const m = /^[\s*`#>-]*verdict[*`]*\s*:[*`\s]*(approve|revise|reject)\b/im.exec(answer)
64  return m === null ? 'no-verdict' : (m[1].toLowerCase() as CollabVerdict)
65}
66
67export function findingLine(f: Finding): string {
68  const fix = f.suggestedFix === undefined ? '' : ` Fix: ${f.suggestedFix}`
69  return `- [${f.severity}] ${f.file}:${f.line}: ${f.description}${fix}`
70}
71
72function findingsBlock(findings: readonly Finding[]): string {
73  return findings.length === 0 ? '(none reported)' : findings.map(findingLine).join('\n')
74}
75
76export const SCANNER_PROMPT = [
77  'You are the bughunt scanner. You are read-only: you read files and report bugs, you never edit.',
78  'Apply the bughunt skill in its collab scanner mode: every finding carries a file:line you read, a named trigger and a consequence; style is not a bug.',
79  'Report each finding at once with the mcp__bughunt__found tool, then answer with a short markdown report.',
80].join('\n')
81
82export const PLANNER_PROMPT = [
83  'You are the bughunt planner. You are read-only.',
84  'You receive the scanner\'s findings. Read the code they point at and write a fix plan: one step per root cause, ordered by severity, each naming the files and the smallest change.',
85  'Mark a finding you could not confirm from the code as unconfirmed instead of planning a fix for it.',
86].join('\n')
87
88export const CRITIC_PROMPT = [
89  'You are the bughunt critic. You are read-only.',
90  'You receive the findings and the fix plan. Check each against the code: is the finding real, does the plan fix the root cause, does it break a caller?',
91  'Begin your answer with exactly one line: "verdict: approve", "verdict: revise" or "verdict: reject". Then give the reasons.',
92].join('\n')
93
94export function scannerTask(paths: readonly string[]): string {
95  return `Scan these paths for bugs:\n${paths.map(p => `- ${p}`).join('\n')}`
96}
97
98export function plannerTask(paths: readonly string[], findings: readonly Finding[], report: string): string {
99  return [`Paths: ${paths.join(', ')}`, '', 'Findings:', findingsBlock(findings), '', 'Scanner report:', report || '(none)'].join('\n')
100}
101
102export function criticTask(paths: readonly string[], findings: readonly Finding[], plan: string): string {
103  return [`Paths: ${paths.join(', ')}`, '', 'Findings:', findingsBlock(findings), '', 'Fix plan:', plan || '(none)'].join('\n')
104}
105
106function stepLine(r: StepResult): string {
107  return `- ${r.step}: ${r.status}${r.reason === undefined ? '' : ` (${r.reason})`}`
108}
109
110/** The collab report the model reads. */
111export function reportText(paths: readonly string[], findings: readonly Finding[], steps: readonly StepResult[], verdict: CollabVerdict): string {
112  const answer = (s: Step) => steps.find(r => r.step === s)?.answer.trim() || '(no answer)'
113  return [
114    `# bughunt collab report`,
115    '',
116    `Paths: ${paths.join(', ')}`,
117    `Verdict: ${verdict}${verdict === 'no-verdict' ? ' (the critic gave no verdict line; this is not an approval)' : ''}`,
118    '',
119    '## Steps',
120    ...steps.map(stepLine),
121    '',
122    `## Findings (${findings.length})`,
123    findingsBlock(findings),
124    '',
125    '## Plan',
126    answer('planner'),
127    '',
128    '## Critique',
129    answer('critic'),
130    '',
131    'This is a read-only review. Do not edit files unless the person asks for the fix.',
132  ].join('\n')
133}
134
hooks/pipeline.ts 102 lines
1/**
2 * The collab pipeline: scanner, planner and critic in turn, each waiting for its subagent's answer or its
3 * time limit. Pure code: the hooks module hands in the engine calls as `Ports`, so tests can drive every
4 * wait with fakes (the test engine starts no subagent).
5 */
6import { criticTask, handBackOf, plannerTask, reportText, scannerTask, verdictOf, type Finding, type Step, type StepResult } from './collab.ts'
7import type { RoundEnd } from './round.ts'
8
9export const STEP_MS: Record<Step, number> = { scanner: 600_000, planner: 480_000, critic: 360_000 }
10
11/** A running collab: its paths, the step running, the agents it started, the scanner and what it found. */
12export type CollabRun = { paths: string[]; step: Step; agents: Set<string>; scanner?: string; findings: Finding[]; launched: () => void }
13
14type Waiter = (end: RoundEnd) => void
15
16/** The collab run, the subagent answers a step waits for, and those that came before the wait. */
17export type Waits = { collab: CollabRun | undefined; waiters: Map<string, Waiter>; early: Map<string, RoundEnd> }
18
19/** What the pipeline needs from the engine. */
20export type Ports = {
21  /** Starts the step's subagent: its id, or why none started. */
22  spawn: (step: Step, prompt: string) => Promise<{ agentId?: string; deny?: string }>
23  /** Calls `fn` after `ms`, unless cancelled first. */
24  after: (ms: number, fn: () => void) => { cancel: () => void }
25  /** Redraws the standing section. */
26  show: () => Promise<void>
27  /** One line for the person. */
28  tell: (text: string, kind: 'ok' | 'warn') => Promise<void>
29}
30
31/** Waits for a subagent's answer, or its step's time limit. */
32export function answerOf(w: Waits, ports: Ports, agentId: string, ms: number): Promise<RoundEnd | undefined> {
33  const early = w.early.get(agentId)
34  if (early !== undefined) {
35    w.early.delete(agentId)
36    return Promise.resolve(early)
37  }
38  return new Promise(resolve => {
39    const timer = ports.after(ms, () => {
40      w.waiters.delete(agentId)
41      resolve(undefined)
42    })
43    w.waiters.set(agentId, end => {
44      timer.cancel()
45      resolve(end)
46    })
47  })
48}
49
50async function runStep(w: Waits, ports: Ports, run: CollabRun, step: Step, prompt: string): Promise<StepResult> {
51  run.step = step
52  await ports.show()
53  const spawned = await ports.spawn(step, prompt)
54  if (spawned.deny !== undefined || spawned.agentId === undefined) return { step, status: 'failed', answer: '', reason: spawned.deny ?? 'no agent started' }
55  run.agents.add(spawned.agentId)
56  if (step === 'scanner') run.scanner = spawned.agentId
57  run.launched()
58  const end = await answerOf(w, ports, spawned.agentId, STEP_MS[step])
59  if (end === undefined) return { step, status: 'timed-out', answer: '', reason: `no answer within ${STEP_MS[step] / 60_000} min` }
60  if (end.reason !== 'answer') return { step, status: 'failed', answer: end.answer, reason: `ended with ${end.reason}` }
61  return { step, status: 'done', answer: end.answer }
62}
63
64const skipped = (step: Step): StepResult => ({ step, status: 'skipped', answer: '', reason: 'the scanner produced nothing' })
65
66/** Scanner, planner and critic in turn; each reads what the ones before it produced. The answer is the report. */
67export async function runCollab(w: Waits, ports: Ports, paths: string[], launched: () => void): Promise<string> {
68  const run: CollabRun = { paths, step: 'scanner', agents: new Set(), findings: [], launched }
69  w.collab = run
70  try {
71    const scan = await runStep(w, ports, run, 'scanner', scannerTask(paths))
72    if (scan.status !== 'done' && run.findings.length === 0) {
73      return reportText(paths, run.findings, [scan, skipped('planner'), skipped('critic')], 'no-verdict')
74    }
75    const plan = await runStep(w, ports, run, 'planner', plannerTask(paths, run.findings, scan.answer))
76    const critique = await runStep(w, ports, run, 'critic', criticTask(paths, run.findings, plan.answer))
77    const verdict = critique.status === 'done' ? verdictOf(critique.answer) : 'no-verdict'
78    await ports.tell(`collab finished: ${run.findings.length} finding(s), verdict ${verdict}`, verdict === 'approve' ? 'ok' : 'warn')
79    return reportText(paths, run.findings, [scan, plan, critique], verdict)
80  } finally {
81    w.collab = undefined
82    await ports.show()
83  }
84}
85
86/** A subagent's turn end: a collab step waits for it, or it came before the wait began. */
87export function settleAgent(w: Waits, agentId: string, end: RoundEnd): void {
88  const waiter = w.waiters.get(agentId)
89  if (waiter !== undefined) {
90    w.waiters.delete(agentId)
91    waiter(end)
92  } else if (w.collab !== undefined) w.early.set(agentId, end)
93}
94
95/** Settles the collab step whose hand-back `text` is; false when it is no collab step's hand-back. */
96export function takeHandBack(w: Waits, text: string): boolean {
97  const back = handBackOf(text)
98  if (back === undefined || w.collab?.agents.has(back.from) !== true) return false
99  settleAgent(w, back.from, { reason: 'answer', isAborted: false, answer: back.report })
100  return true
101}
102
hooks/paths.ts 43 lines
1/** The directory a round keeps its proof in, relative to the working directory. */
2export function proofDir(id: string): string {
3  return `.temp_files/bughunt/${id}`
4}
5
6/** A path relative to `cwd` when it lies under it, `./` and trailing `/` removed. */
7export function relativeTo(cwd: string, path: string): string {
8  const base = cwd.replace(/\/+$/, '')
9  const inside = path.startsWith(`${base}/`) ? path.slice(base.length + 1) : path
10  return inside.replace(/^(\.\/)+/, '').replace(/\/+$/, '')
11}
12
13/** Whether `path` is `dir` or lies under it. */
14export function under(path: string, dir: string): boolean {
15  const d = dir.replace(/^(\.\/)+/, '').replace(/\/+$/, '')
16  return d === '' || path === d || path.startsWith(`${d}/`)
17}
18
19const TEST_PATH = [/(^|\/)(tests?|__tests__|specs?)\//, /\.(test|spec)\.[^/]+$/, /_test\.[^/]+$/, /(^|\/)test_[^/]+\.py$/]
20
21/** Whether the path looks like a test file. */
22export function isTestPath(path: string): boolean {
23  return TEST_PATH.some(r => r.test(path))
24}
25
26/** Whether the path lies in the hunt's target; an empty target is the whole project. */
27export function inScope(path: string, target: string): boolean {
28  const parts = target.split(/\s+/).filter(t => t !== '')
29  return parts.length === 0 || parts.some(t => under(path, t))
30}
31
32export type EditGate = { skill: boolean; failed: boolean; target: string; dir: string }
33export type EditRule = 'skill' | 'proof' | 'scope'
34
35/** The rule an edit of `path` breaks in a running round, or undefined. */
36export function editRule(path: string, gate: EditGate): EditRule | undefined {
37  if (!gate.skill) return 'skill'
38  if (under(path, gate.dir)) return undefined
39  if (!gate.failed) return 'proof'
40  if (isTestPath(path)) return undefined
41  return inScope(path, gate.target) ? undefined : 'scope'
42}
43
hooks/proof.ts 76 lines
1import { isTestPath, under } from './paths.ts'
2import type { ProofState } from './round.ts'
3
4export type Phase = 'before' | 'after'
5export type ProofRun = { exitCode: number; stdout: string; stderr: string }
6export type ProofJudgement = { accepted: boolean; reason: string; state: ProofState | undefined }
7
8const FAIL_LINE = /^\s*FAIL\b/m
9const PASS_LINE = /^\s*PASS\b/m
10const TAIL_LINES = 20
11
12/** The input of the proof tool, or the reason it is refused. */
13export function proofInput(e: { phase?: unknown; argv?: unknown; cwd?: unknown }): { phase: Phase; argv: string[]; cwd?: string } | string {
14  if (e.phase !== 'before' && e.phase !== 'after') return 'phase must be "before" or "after"'
15  if (!Array.isArray(e.argv) || e.argv.length === 0 || !e.argv.every(a => typeof a === 'string')) return 'argv must be a non-empty array of strings'
16  if (e.cwd !== undefined && typeof e.cwd !== 'string') return 'cwd must be a string'
17  return { phase: e.phase, argv: e.argv, cwd: e.cwd }
18}
19
20/** Judges one run of the proof command against the round's record. */
21export function judgeProof(phase: Phase, argv: string[], run: ProofRun, state: ProofState | undefined): ProofJudgement {
22  return phase === 'before' ? judgeBefore(argv, run, state) : judgeAfter(argv, run, state)
23}
24
25function judgeBefore(argv: string[], run: ProofRun, state: ProofState | undefined): ProofJudgement {
26  const out = `${run.stdout}\n${run.stderr}`
27  if (run.exitCode === 0) return { accepted: false, reason: 'the proof exited 0, so it does not show the bug', state }
28  if (!FAIL_LINE.test(out)) return { accepted: false, reason: 'the proof exited non-zero but printed no line starting with FAIL', state }
29  return { accepted: true, reason: 'FAIL recorded', state: { argv, failed: true, passed: false } }
30}
31
32function judgeAfter(argv: string[], run: ProofRun, state: ProofState | undefined): ProofJudgement {
33  if (state === undefined || !state.failed) return { accepted: false, reason: 'no FAIL was recorded in this round; run phase "before" first', state }
34  if (JSON.stringify(argv) !== JSON.stringify(state.argv)) return { accepted: false, reason: `argv differs from the recorded proof ${JSON.stringify(state.argv)}`, state }
35  const out = `${run.stdout}\n${run.stderr}`
36  if (run.exitCode !== 0) return { accepted: false, reason: `the proof still exits ${run.exitCode}`, state }
37  if (!PASS_LINE.test(out)) return { accepted: false, reason: 'the proof exited 0 but printed no line starting with PASS', state }
38  return { accepted: true, reason: 'PASS recorded', state: { ...state, passed: true } }
39}
40
41/**
42 * The production files a revert check puts back: those modified since the FAIL, outside the proof
43 * directory and not tests, because the check asks whether the fix, not the proof, turns FAIL into PASS.
44 */
45export function revertTargets(changed: readonly string[], dir: string): string[] {
46  return changed.filter(path => path !== '' && !under(path, dir) && !isTestPath(path))
47}
48
49/**
50 * The proof run with the fix reverted must fail as the FAIL did; a pass there means the proof does not
51 * reach the fixed code. Undefined when it failed, else why the PASS is refused.
52 */
53export function judgeReverted(run: ProofRun, files: readonly string[]): string | undefined {
54  const out = `${run.stdout}\n${run.stderr}`
55  if (run.exitCode !== 0 && FAIL_LINE.test(out)) return undefined
56  return `with the fix reverted (${files.join(', ')}) the proof still exits ${run.exitCode}${FAIL_LINE.test(out) ? '' : ' without a FAIL line'}, so it does not show that the fix makes it pass`
57}
58
59/** A revert check in progress: the files it reverted to `base` and the snapshot commit `fixed` that holds the fix. */
60export type SavedFix = { base: string; fixed: string; files: string[] }
61
62/** The stored record of a revert check, or undefined when the value is none. */
63export function savedFix(value: unknown): SavedFix | undefined {
64  if (typeof value !== 'object' || value === null) return undefined
65  const v = value as Record<string, unknown>
66  if (typeof v.base !== 'string' || typeof v.fixed !== 'string' || !Array.isArray(v.files)) return undefined
67  const files = v.files.filter((f): f is string => typeof f === 'string' && f !== '')
68  return files.length === 0 ? undefined : { base: v.base, fixed: v.fixed, files }
69}
70
71/** The last lines of a run's output, for the model. */
72export function tailOf(run: ProofRun): string {
73  const lines = `${run.stdout}${run.stderr === '' ? '' : `\n${run.stderr}`}`.trimEnd().split('\n')
74  return lines.slice(-TAIL_LINES).join('\n')
75}
76
hooks/round.ts 83 lines
1export const OUTCOMES = ['fixed-and-verified', 'fixed-verification-incomplete', 'no-proven-bug', 'blocked'] as const
2export type Outcome = (typeof OUTCOMES)[number]
3
4/**
5 * What the proof tool recorded in the running round; `base` is the working tree's snapshot commit taken
6 * when the FAIL was recorded, which the revert check puts the fixed files back to.
7 */
8export type ProofState = { argv: string[]; failed: boolean; passed: boolean; base?: string }
9
10export type RoundRecord = { round: number; outcome: Outcome | 'none'; fingerprint?: string; note?: string }
11
12/** One hunt: `rounds` rounds over `target`, the running round `round`. */
13export type Hunt = {
14  id: string
15  round: number
16  rounds: number
17  target: string
18  skill: boolean
19  proof: ProofState | undefined
20  history: RoundRecord[]
21}
22
23export type RoundEnd = { reason: string; isAborted: boolean; answer: string }
24
25export type Verdict = { next: 'continue' } | { next: 'done' } | { next: 'stop'; reason: string }
26
27/** A new hunt at its first round. */
28export function newHunt(id: string, rounds: number, target: string): Hunt {
29  return { id, round: 1, rounds, target, skill: false, proof: undefined, history: [] }
30}
31
32/** The id of the running round, used for its proof directory. */
33export function roundId(hunt: Hunt): string {
34  return `${hunt.id}-r${hunt.round}`
35}
36
37/**
38 * The outcome label of the first line that begins with one, markdown marks stripped. The model sometimes
39 * writes a sentence before the label (measured), so every line is read, not only the first.
40 */
41export function outcomeOf(answer: string): Outcome | undefined {
42  for (const line of answer.split('\n')) {
43    const bare = line.replace(/[*`#>_]/g, '').replace(/^\s*[-:]?\s*/, '').trim().toLowerCase()
44    const hit = OUTCOMES.find(o => bare === o || new RegExp(`^${o}(?![\\w-])`).test(bare))
45    if (hit !== undefined) return hit
46  }
47  return undefined
48}
49
50/** The `fingerprint:` line of a round's answer. */
51export function fingerprintOf(answer: string): string | undefined {
52  const m = /^[\s*`>-]*fingerprint[*`]*\s*:[*`\s]*(.+?)[`\s]*$/im.exec(answer)
53  return m === null ? undefined : m[1]
54}
55
56/** Whether the hunt goes on after a round ended as `end` said. */
57export function decide(hunt: Hunt, end: RoundEnd): Verdict {
58  if (end.isAborted) return { next: 'stop', reason: 'the round was interrupted' }
59  if (end.reason !== 'answer') return { next: 'stop', reason: `the round ended with ${end.reason}` }
60  const outcome = outcomeOf(end.answer)
61  if (outcome === undefined) return { next: 'stop', reason: 'the answer did not begin with an outcome line' }
62  if (outcome === 'blocked' || outcome === 'fixed-verification-incomplete') return { next: 'stop', reason: `the round reported ${outcome}` }
63  if (outcome === 'fixed-and-verified' && !proven(hunt.proof)) {
64    return { next: 'stop', reason: 'the round reported fixed-and-verified, but the mod recorded no FAIL followed by a PASS' }
65  }
66  return hunt.round >= hunt.rounds ? { next: 'done' } : { next: 'continue' }
67}
68
69function proven(proof: ProofState | undefined): boolean {
70  return proof !== undefined && proof.failed && proof.passed
71}
72
73/** Records the ended round and moves the hunt to the next one. */
74export function advance(hunt: Hunt, answer: string): Hunt {
75  const record: RoundRecord = { round: hunt.round, outcome: outcomeOf(answer) ?? 'none', fingerprint: fingerprintOf(answer) }
76  return { ...hunt, round: hunt.round + 1, skill: false, proof: undefined, history: [...hunt.history, record] }
77}
78
79/** The fingerprints of the rounds so far. */
80export function fingerprints(hunt: Hunt): string[] {
81  return hunt.history.flatMap(r => (r.fingerprint === undefined ? [] : [r.fingerprint]))
82}
83
hooks/texts.ts 64 lines
1import type { EditRule } from './paths.ts'
2import { fingerprints, type Hunt } from './round.ts'
3
4export const SKILL = 'bughunt:hunt'
5export const PROOF_TOOL = 'mcp__bughunt__proof'
6
7const PROTOCOL = [
8  'Protocol (the bughunt skill has the full text):',
9  '1. Survey: record the starting revision and dirty paths; keep other changes untouched.',
10  `2. Prove: write a proof in the proof directory that runs the real code path, prints a FAIL line and exits non-zero; call ${PROOF_TOOL} with phase "before". Without a recorded FAIL, change no production code.`,
11  '3. Fix the root cause with the smallest change.',
12  `4. Verify: the same proof prints PASS and exits 0; call ${PROOF_TOOL} with phase "after" and the same argv. The mod then reverts the production files modified since the FAIL, runs the proof again and puts them back; the PASS counts only when that run fails. Add a regression test to the suite and run the checks.`,
13  '5. Report: begin the answer with one line, fixed-and-verified, fixed-verification-incomplete, no-proven-bug or blocked, and give a "fingerprint: <file>:<symbol>: <cause>" line. Then stop.',
14  'No subagents in a round. One root cause per round. Do not count a fingerprint listed below again.',
15].join('\n')
16
17/** The prompt that starts or continues a round. */
18export function roundText(hunt: Hunt, dir: string): string {
19  const seen = fingerprints(hunt)
20  return [
21    `Proof-driven bug hunt, round ${hunt.round}/${hunt.rounds}.`,
22    `Scope: ${hunt.target === '' ? 'the whole project' : `${hunt.target} and everything under it`}.`,
23    `Proof directory: ${dir}`,
24    `First invoke the ${SKILL} skill with the Skill tool; edits are refused until it is open in this round.`,
25    '',
26    PROTOCOL,
27    '',
28    `Fingerprints of earlier rounds: ${seen.length === 0 ? 'none' : ''}`,
29    ...seen.map(f => `- ${f}`),
30  ].join('\n')
31}
32
33/** The block the skill text gets while a round runs. */
34export function skillBlock(hunt: Hunt, dir: string): string {
35  const seen = fingerprints(hunt)
36  return [
37    '',
38    '## Current round (bughunt)',
39    `Round ${hunt.round}/${hunt.rounds}. Scope: ${hunt.target === '' ? 'the whole project' : hunt.target}. Proof directory: ${dir}`,
40    `Earlier fingerprints: ${seen.length === 0 ? 'none' : seen.join('; ')}`,
41  ].join('\n')
42}
43
44const EDIT_DENY: Record<EditRule, (dir: string, target: string) => string> = {
45  skill: () => `A bughunt round is running and the ${SKILL} skill is not open in it. Invoke it with the Skill tool first.`,
46  proof: dir => `No FAIL is recorded in this round. Write the proof in ${dir} and call ${PROOF_TOOL} with phase "before"; production code stays unchanged until it fails.`,
47  scope: (_, target) => `This file is outside the round's scope (${target}). Keep the fix inside the scope, or put a regression test in a test file.`,
48}
49
50export function editDenyText(rule: EditRule, dir: string, target: string): string {
51  return `${EDIT_DENY[rule](dir, target)} There is no way around this gate; /bughunt stop ends the hunt.`
52}
53
54export const SPAWN_DENY = 'A bughunt round is running, and a round uses no subagents. Do the work in this conversation. There is no way around this gate; /bughunt stop ends the hunt.'
55
56/** The one-line log of a stopped edit. */
57export function editLog(rule: EditRule, path: string): string {
58  return `edit stopped (${rule}): ${path}`
59}
60
61export function roundLine(hunt: Hunt): string {
62  return `round ${hunt.round}/${hunt.rounds} started${hunt.target === '' ? '' : ` in ${hunt.target}`}`
63}
64