Sends risky-looking Bash commands to an LLM judge and denies the ones it reads as destructive; regex only decides what is worth judging. Fails closed. Needs…

危なそうな Bash コマンドを LLM に意図判定させて止める mod。
正規表現だけのガードは言い換えで抜けられる。この mod では正規表現は「判定にかける価値が あるか」だけを決め、実行するかどうかは Haiku が safe / destructive で判定する。
tool.call を { tool: 'Bash' } で受ける。rm / git push --force / DROP TABLE / mkfs / curl … | sh など) に当たらなければ、そのまま next(e)。モデル呼び出しは一切しない。$.model.classify(…, ['safe', 'destructive'], { model: 'haiku' }) にかける。destructive なら { deny: … } で拒否。safe / ラベル無しなら next(e)。判定したコマンドは所要時間つきで PROOF-guard.log に記録する。実測でおよそ 500〜750 ms。
コマンド文字列をそのまま classify に渡すと、Haiku は rm -rf <フォルダ> を safe と 答えた (実測)。destructive / safe の定義を短く添えた文に包んで初めて destructive に なる。hooks/register.ts の question を参照。
❯ ./register.ts hooks: tool.call{tool=Bash} ❯ ./register.ts calls: $.fs.read, $.fs.write, $.model.classify
Bash の呼び出ししか見ない。コマンド文字列は判定のためセッション自身の API クライアントへ 送られる (外部サービスへは出ない)。書き込みは自分のログ 1 ファイルだけ。
cd mods/intent-guard
claude plugin validate .
claude plugin test .
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 が要る。
find <dir> -type f -delete は素通りしてファイルを消した (\bdel\b は delete に 当たらないため)。言い換え耐性があるのは「judge に届いた後」だけ。hooks/register.ts 62 lines1import type { Register } from 'claude-code'
2
3
4// A cheap filter, not the decision: it only picks what is worth one model
5// call. Note the trailing \b -- it costs this pattern `dd if=/dev/zero`,
6// because `=` is not a word character (see NOTES.md).
7const TRIGGER =
8 /\b(rm|rmdir|del|rd|git\s+(push\s+.*--force|reset\s+--hard|clean\s+-f|branch\s+-D)|DROP\s+(TABLE|DATABASE)|truncate|mkfs|dd\s+if=|chmod\s+-R\s+777|curl[^|]*\|\s*(sh|bash)|Remove-Item|Format-Volume)\b/i
9
10export const register: Register = on => {
11 on('tool.call', { tool: 'Bash' }, async ($, e, next) => {
12 const command = typeof e.command === 'string' ? e.command : ''
13 if (!TRIGGER.test(command)) return next(e)
14
15 // The bare command alone reads as routine housekeeping to the judge: it
16 // labelled `rm -rf <folder>` "safe" until this framing was added.
17 const question =
18 'A coding agent wants to run this shell command on a developer\'s machine.\n' +
19 'destructive = it irreversibly deletes, overwrites or discards data or history.\n' +
20 'safe = it only reads, lists, prints or creates.\n\n' +
21 `Command: ${command}`
22
23 const started = Date.now()
24 let verdict: string | undefined
25 let failure = ''
26 try {
27 verdict = await $.model.classify(question, ['safe', 'destructive'], { model: 'haiku' })
28 } catch (error) {
29 failure = error instanceof Error ? error.message : String(error)
30 }
31 const ms = Date.now() - started
32
33 const decision =
34 failure !== '' ? 'DENY(judge-unavailable)'
35 : verdict === 'destructive' ? 'DENY(destructive)'
36 : verdict === undefined ? 'allow(no-label)'
37 : 'allow(safe)'
38
39 const before = await $.fs.read(`${$.plugin.root}/PROOF-guard.log`).catch(() => '')
40 const note = failure === '' ? '' : ` error=${failure}`
41 await $.fs.write(
42 `${$.plugin.root}/PROOF-guard.log`,
43 `${before}${new Date().toISOString()} ${ms}ms verdict=${verdict ?? 'undefined'} ` +
44 `decision=${decision} cmd=${JSON.stringify(command)}${note}\n`,
45 )
46
47 // Fail closed: an unreachable judge is not a pass.
48 if (failure !== '') {
49 return { deny: `intent-guard: the intent judge could not answer (${failure}), so the command was refused rather than guessed at` }
50 }
51 if (verdict === 'destructive') {
52 return { deny: `intent-guard: an LLM judge read this command as destructive, so it was not run: ${command}` }
53 }
54 return next(e)
55 }).catch(($, e, next) => {
56 // The hook threw, misreturned or outran its 10 s budget. Same rule: a
57 // guard that cannot decide refuses. Never replay next(e) here -- that is
58 // the unjudged command running after the guard has already failed.
59 return { deny: `intent-guard: the guard itself failed (${next.error.kind}: ${next.error.message ?? 'no message'}), so the Bash command was refused unjudged` }
60 })
61}
62