SLOPSHOPPER

intent-guard

Sends risky-looking Bash commands to an LLM judge and denies the ones it reads as destructive; regex only decides what is worth judging. Fails closed. Needs…

newguardmodel
A shopper browsing a rack in a slop shop
README

intent-guard

危なそうな Bash コマンドを LLM に意図判定させて止める mod。

正規表現だけのガードは言い換えで抜けられる。この mod では正規表現は「判定にかける価値が あるか」だけを決め、実行するかどうかは Haiku が safe / destructive で判定する。

仕組み

  1. tool.call を { tool: 'Bash' } で受ける。
  2. 安価なトリガ正規表現 (rm / git push --force / DROP TABLE / mkfs / curl … | sh など) に当たらなければ、そのまま next(e)。モデル呼び出しは一切しない。
  3. 当たったら $.model.classify(…, ['safe', 'destructive'], { model: 'haiku' }) にかける。
  4. destructive なら { deny: … } で拒否。safe / ラベル無しなら next(e)。
  5. judge が失敗 (API エラー等) したら deny。判定できないものは通さない (fail closed)。

判定したコマンドは所要時間つきで PROOF-guard.log に記録する。実測でおよそ 500〜750 ms。

判定文の作り方が要

コマンド文字列をそのまま classify に渡すと、Haiku は rm -rf <フォルダ> を safe と 答えた (実測)。destructive / safe の定義を短く添えた文に包んで初めて destructive に なる。hooks/register.ts の question を参照。

届く範囲 (Claude Code 2.1.278 で検証)

❯ ./register.ts hooks: tool.call{tool=Bash} ❯ ./register.ts calls: $.fs.read, $.fs.write, $.model.classify

Bash の呼び出ししか見ない。コマンド文字列は判定のためセッション自身の API クライアントへ 送られる (外部サービスへは出ない)。書き込みは自分のログ 1 ファイルだけ。

動かし方

cd mods/intent-guard
claude plugin validate .
claude plugin test .

CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 が要る。

限界

  • トリガ正規表現に当たらないものは judge にすら届かない。 実測で find <dir> -type f -delete は素通りしてファイルを消した (\bdel\b は delete に 当たらないため)。言い換え耐性があるのは「judge に届いた後」だけ。
  • judge は 1 回のモデル呼び出し。同じコマンドでも答えがぶれる可能性がある。
  • Bash 1 回につき最大 1 回のモデル呼び出しぶん遅くなる (実測 500〜750 ms)。
Source 1 files
hooks/register.ts 62 lines
1import type { Register } from 'claude-code'
2
3
4// A cheap filter, not the decision: it only picks what is worth one model
5// call. Note the trailing \b -- it costs this pattern `dd if=/dev/zero`,
6// because `=` is not a word character (see NOTES.md).
7const TRIGGER =
8  /\b(rm|rmdir|del|rd|git\s+(push\s+.*--force|reset\s+--hard|clean\s+-f|branch\s+-D)|DROP\s+(TABLE|DATABASE)|truncate|mkfs|dd\s+if=|chmod\s+-R\s+777|curl[^|]*\|\s*(sh|bash)|Remove-Item|Format-Volume)\b/i
9
10export const register: Register = on => {
11  on('tool.call', { tool: 'Bash' }, async ($, e, next) => {
12    const command = typeof e.command === 'string' ? e.command : ''
13    if (!TRIGGER.test(command)) return next(e)
14
15    // The bare command alone reads as routine housekeeping to the judge: it
16    // labelled `rm -rf <folder>` "safe" until this framing was added.
17    const question =
18      'A coding agent wants to run this shell command on a developer\'s machine.\n' +
19      'destructive = it irreversibly deletes, overwrites or discards data or history.\n' +
20      'safe = it only reads, lists, prints or creates.\n\n' +
21      `Command: ${command}`
22
23    const started = Date.now()
24    let verdict: string | undefined
25    let failure = ''
26    try {
27      verdict = await $.model.classify(question, ['safe', 'destructive'], { model: 'haiku' })
28    } catch (error) {
29      failure = error instanceof Error ? error.message : String(error)
30    }
31    const ms = Date.now() - started
32
33    const decision =
34      failure !== '' ? 'DENY(judge-unavailable)'
35      : verdict === 'destructive' ? 'DENY(destructive)'
36      : verdict === undefined ? 'allow(no-label)'
37      : 'allow(safe)'
38
39    const before = await $.fs.read(`${$.plugin.root}/PROOF-guard.log`).catch(() => '')
40    const note = failure === '' ? '' : ` error=${failure}`
41    await $.fs.write(
42      `${$.plugin.root}/PROOF-guard.log`,
43      `${before}${new Date().toISOString()} ${ms}ms verdict=${verdict ?? 'undefined'} ` +
44        `decision=${decision} cmd=${JSON.stringify(command)}${note}\n`,
45    )
46
47    // Fail closed: an unreachable judge is not a pass.
48    if (failure !== '') {
49      return { deny: `intent-guard: the intent judge could not answer (${failure}), so the command was refused rather than guessed at` }
50    }
51    if (verdict === 'destructive') {
52      return { deny: `intent-guard: an LLM judge read this command as destructive, so it was not run: ${command}` }
53    }
54    return next(e)
55  }).catch(($, e, next) => {
56    // The hook threw, misreturned or outran its 10 s budget. Same rule: a
57    // guard that cannot decide refuses. Never replay next(e) here -- that is
58    // the unjudged command running after the guard has already failed.
59    return { deny: `intent-guard: the guard itself failed (${next.error.kind}: ${next.error.message ?? 'no message'}), so the Bash command was refused unjudged` }
60  })
61}
62