SLOPSHOPPER

protocol-guard

Required workflow checks with Claude Code function hooks. Refuses a knowledge-file write, or a memory-service save in external memory mode, until…

newguardprompt
v0.3.0no licenseupdated 2026-09-23Mar5929/claude-toolkit/plugins/protocol-guard
A shopper browsing a rack in a slop shop
README

protocol-guard plugin

Required workflow checks for Claude Code. The plugin refuses a tool call, or holds a final reply once, when a required step did not happen. It decides from facts Claude Code reports: which skill was opened, which file a tool wrote, which program a shell command ran, and whether the call succeeded. No model is called, and the agent's words are never read.

A check proves that a step happened. It does not prove the step was done well, and it never proves that the owner approved anything.

Last tested on Claude Code 2.1.280.

What it checks

The shipped list is protocols.default.json. Each entry names the skill that owns the step.

IDRequired stepWhen it is checkedOn failure
K4Knowledge files change only with knowledge-save openA write to knowledge/memory-inbox.md, knowledge/memory/, knowledge/prds/ or knowledge/memory-self-improvement.md, by a file tool or a shell commandThe call is refused. The inbox needs the skill opened this turn; the other files need it opened since the last reset
CWWorking memory follows work-item changesThe turn ends after a work item was created, closed, or moved to another stage (gh issue, a label edit with a stage label such as 08-build, gh project item-edit, GitHub issue tools, or the work command)The reply is held once and the agent is sent to knowledge-save, until knowledge/memory/current.md is written after the change
K5Generated indexes are not edited by handA write to knowledge/memory/memory-index.md, knowledge/prds/prd-index.md or ai-external-knowledge/README.mdThe call is refused; the agent runs the index builder instead
K6Indexes rebuilt and checker run after a knowledge writeThe turn ends after a K4 writeThe reply is held once until build-knowledge-index.mjs and then check-knowledge.mjs both exited 0 after the last write
K7Save review before a pull request, a close, or a mergegh pr create, gh issue close, gh pr merge (also after -R/--repo; not with --help, or --dry-run for create), work finish, and the GitHub tools create_pull_request, issue_write with state closed, merge_pull_request and enable_pr_auto_mergeThe call is refused until knowledge-save was opened this turn
P2Closing a work item goes through the work skillgh issue close, work finish, issue_write with state closedThe call is refused until work was opened this turn. The skill asks the owner for approval; the check does not prove approval
K4Xexternal mode: memory records and PRD files change only with knowledge-save openA call to the memory service's write tools (memory-write, below), or a write to prds/ by a file tool or a shell commandThe call is refused. A record with toolkit_kind pending needs the skill opened this turn; other records and prds/ files need it opened since the last reset
CWXexternal mode: working memory follows work-item changesThe turn ends after a work-item change, read as for CWThe reply is held once until a successful memory-write call with toolkit_kind working, or a delete of a working record (see "Memory mode"), follows the change
K5Xexternal mode: generated indexes are not edited by handA write to prds/prd-index.md or ai-external-knowledge/README.mdThe call is refused; the agent runs the index builder instead
K6Xexternal mode: indexes rebuilt and checker run after a PRD writeThe turn ends after a write to prds/The reply is held once until build-knowledge-index.mjs and then check-knowledge.mjs both exited 0 after the last write
P3A merge goes through merge-and-clean-upgh pr merge, merge_pull_request, enable_pr_auto_mergeThe call is refused until merge-and-clean-up was opened this session. Compaction does not clear it

K4, CW, K5 and K6 apply only in a project with knowledge/knowledge-manual.md in memory mode files. K4X, CWX, K5X and K6X apply only in memory mode external. K7 applies in both: in files mode with knowledge/knowledge-manual.md, and in external mode. See "Memory mode" below. A check whose owner skill is not installed is switched off: the agent is told, and the owner sees one line in the first turn. Every check with an owner skill tells the agent "If the skill is not installed, tell the owner and stop." K5 needs no skill.

Opening the owner skill is the whole check. For K7, P2 and P3 the engine refuses the action instead of holding it once, as decision 10 approved; the older hold-once hooks are the backup.

How each fact is read:

  • A skill opened: a successful Skill tool call, a typed /name, or a Read of the skill's SKILL.md.
  • A file written: a successful Write, Edit, MultiEdit or NotebookEdit, or a shell command that names the file and is not a read-only program such as cat, grep, cut, sort without -o, awk without -i inplace, sed without -i, or git diff (git's -C and -c options are skipped before the subcommand). cp, install and ln write only their target.
  • A memory-service call (external mode): a call to a tool named mcp__<server>__<tool>, where <server> is the config's server. A valid server holds only A-Z, a-z, 0-9, _ and -, so it appears in the tool name unchanged. Its kind is the call's metadata.toolkit_kind, or a tag toolkit_kind:<kind> in its tags. This is an exact field comparison; the record's text is never read.
  • A shell command: its program, subcommand and file arguments (hooks/shell-reader.mjs). Quoted text is one argument, and heredoc bodies are dropped, so a commit message that mentions a path is not that path.
  • Success and order: only a call that succeeded counts. A shell command counts toward a requirement only when its exit code 0 proves it ran to the end, such as the parts of a final && chain. "After" means a later call in the same session.
  • Paths: a file counts only inside the session's project root. A write to knowledge/memory/current.md in another checkout of the same repository, such as a sibling worktree, is not seen. CW then holds that turn once and shows its notice; later turns are not held for it.

Memory mode

.toolkit-memory.json at the project root says where memory lives. The engine reads it once per session.

{ "format": 1, "memory": "external", "service": "mem0", "server": "mem0", "project": "demo" }
  • A missing file means memory mode files: the knowledge/ checks run as before.
  • A file that cannot be read, or is not valid, also means files. The agent is told once. A valid file has "format": 1 and memory set to files or external. external also needs service (mem0 or hindsight), server (letters, digits, _ and - only) and a non-empty project, each on one line. hooks/memory-config.mjs holds these rules. They are the same rules as second-brain's readMemoryConfig, and tests/knowledge-schema.test.mjs checks that every reader agrees.
  • In external mode the engine sorts the service's tools into two classes:
Classmem0Hindsight
memory-writeadd_memory, update_memory, delete_memory, delete_all_memories, delete_entitiesretain, sync_retain, delete_document, clear_memories, update_memory, invalidate_memory, delete_bank
memory-readget_memories, get_memory, search_memorieslist_documents, get_document, recall, list_memories, get_memory

Hindsight update_bank is in neither class. It changes the bank's settings once at setup and holds no memory text.

A delete call carries no toolkit_kind. Which kind requirement it meets:

  • mem0 delete_memory and delete_all_memories carry only an id, so they meet a requirement for any kind. Removing a finished working-memory record meets CWX.
  • Hindsight delete_document meets a requirement for kind K only when its document_id starts with K:. A delete of working:issue-42 meets CWX; a delete of pending:<reference> does not.
  • Hindsight clear_memories, update_memory, invalidate_memory and delete_bank, and mem0 delete_entities, meet no kind requirement. They name no record key.

Agents and resets

  • A skill counts for the agent that opened it. A helper agent also starts with what the main agent had opened when the helper first acted.
  • A helper's successful writes, commands and memory-service calls count for the main agent's turn-end checks (CW, K6, CWX, K6X). While a helper that saved memory this turn is still running (a write under knowledge/ in files mode; a write under prds/ or a memory-write call in external mode), the turn-end checks wait: the reply is not held now, and the checks are carried to the end of the next turn.
  • At each turn start, every agent's "opened this turn" record is cleared.
  • While a helper that wrote knowledge/memory/current.md or the inbox is still running, another agent's write of the same file is refused.
  • Session start, /clear, resume and a plugin reload start the record over. Compaction clears which skills were opened.
  • A turn-end check looks only at the facts of the current turn, except a check carried from a pending helper save, which also counts a helper that wrote in the turn it was carried from. A check still unmet after its one hold is shown as a notice and is not held again in later turns.

How a reply is held

The engine drops the draft reply before it is shown. When Claude Code then asks for a visible reply, that draft is dropped too. The classic Stop event continues the turn with a note, the way a Stop hook's block does; a block from another Stop hook is kept beside it. The note says which step is missing and which skill to open, and the agent does the step and writes its reply again. Only the final reply is shown. If the turn then ends with no reply (an error or an interrupt), the first held draft is shown beneath it.

The design planned to insert a Skill call for the owner instead. In print-mode runs on Claude Code 2.1.280 the model treated that call, which it had not made, as a prompt injection and refused the step (2 of 2 runs). The Stop continuation is Claude Code's own channel for this, and the model followed it.

When the engine fails

  • A tool check that cannot decide refuses the call.
  • The reply hold is the exception: on an error, the reply is shown.
  • After two engine errors in one turn, the engine stops for the rest of that turn and one line says: "Workflow checks are off for this turn after an error."
  • When a check is still unmet after its one hold, the reply is shown with the line "Required workflow check not met after one retry: <IDs>."
  • A held-reply note ends "Do not mention this check in your reply." A tool refusal does not: it names protocol-guard as its source and says what to do. In print-mode runs on Claude Code 2.1.280 the model treated a refusal that carried the line as a prompt injection and stopped (1 of 10 followed it); without the line it followed 4 of 4. This departs from design decision 4 for refusals only, pending Mike's confirmation.

Turning it on in a project

Function hooks load only when CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1. The project-init and project-sync skills write both keys into the project's committed .claude/settings.json:

"env": { "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1" },
"enabledPlugins": { "protocol-guard@claude-toolkit": true }

machine-sync does not set the variable. In user settings it would turn on function hooks for every installed plugin on the computer.

A project changes the list in .claude/protocols.json:

{ "off": ["K6"], "protocols": [] }

off names default checks to skip. protocols adds entries, and an entry with a default's name replaces it.

The words an entry can use:

  • appliesIf: { "exists": "<path>" } and/or { "memory": "files" | "external" }. With both keys, both must hold. A list of such objects holds when any one holds (K7 uses this). Without appliesIf the entry always applies.
  • on: action (pr-create, pr-merge, issue-close, work-finish), write (paths; a path ending in / covers its folder), shell, oneWriter, call (tool classes: memory-write, memory-read), and turnEnd (afterWorkItemChange, afterWrite).
  • require: { "opened": "owner", "within": "turn" | "reset" | "session" }, optionally with paths (only writes to those paths) or kind (only memory-service calls with that toolkit_kind); { "never": true }; { "wrote": [paths] }; { "ran": [scripts] }; and { "called": "<class>", "kind"?: "<kind>" }, met by a later successful call in that class that carries the kind, or is a delete of that kind as "Memory mode" describes. An entry that uses a word the engine does not

know is ignored, and the agent is told once. The list is read once per session.

The older command hooks

The engine adds a toolkit_protocol_engine field, { version, active }, to the input of the classic UserPromptSubmit and Stop hooks. The second-brain command hooks read it:

  • knowledge-completion.mjs skips its end-of-turn check while K4 and K6 are active.
  • memory-reminder.mjs leaves out the turn-review line while CW, K4 and K6 are active. It still starts the review checkpoint, so if the engine stops mid-turn the old end-of-turn check runs in full. When function hooks are on in a project that enables this plugin but the field is absent, it shows the owner one line per session (systemMessage) saying the checks are not running, and gives the agent the same line. Codex runs this hook without CLAUDE_PROJECT_DIR, so it never shows the line there.

The command hooks that run before a tool get no engine field, so the engine also sets TOOLKIT_PROTOCOL_ENGINE to <session id>:<check names> for every process Claude Code starts, and unsets it while it is off. A hook acts on it only when the id matches its own session id, so a child claude or Codex process that inherits the variable keeps its hooks in full:

  • work-item-close.mjs skips its hold-once for gh issue close, gh pr merge and work finish while K7 is active.
  • save-reminder.mjs skips its general hold-once for gh pr create while K7 is active, and keeps the message for a branch that changes only knowledge/.

With the variable off, the field and the environment variable are absent and every old hook runs as before. After two engine errors, the Stop input has no field and the environment variable is unset, so the old checks run for the rest of that turn.

Codex

Codex cannot run function hooks. This plugin has no .codex-plugin manifest and no entry in .agents/plugins/marketplace.json. Codex sessions follow the same steps from their instructions and skills.

Files

FileWhat it is
.claude-plugin/plugin.jsonThe manifest
hooks/hooks.jsonRegisters the function-hook module
hooks/engine.tsThe engine
hooks/shell-reader.mjsReads a shell command's program, arguments and redirections
hooks/memory-config.mjsThe rules for a valid .toolkit-memory.json
protocols.default.jsonThe shipped list
tests/engine.test.tsOffline tests for each check
tests/defaults.fixture.mjsA copy of the list for the offline tests, which cannot read files
tsconfig.jsonType-checks against the declarations /plugin-types writes to .claude/types/

Testing

node tests/protocol-guard-check.mjs
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude plugin test plugins/protocol-guard

The check regenerates the declarations, type-checks the engine, validates the plugin, runs the offline tests, checks that each owner skill exists, and compares the shell reader with the command hooks' reader. It skips the Claude Code steps when claude is not installed. Run it after every Claude Code update: the function-hook API is early access and may change in any release.

Not yet tested: Windows, the desktop app, a first session in an untrusted folder, and turns started by background tasks.

Source 3 files
hooks/engine.ts 899 lines
1// protocol-guard engine: required workflow checks (issue #396, roadmap step 5).
2//
3// One module applies every protocol in a list. The list is data: the shipped
4// defaults (protocols.default.json beside this plugin) plus the project's own
5// changes (.claude/protocols.json: `off` and `protocols`). Changing a protocol
6// that uses the words below is a list edit; a new word is a change here.
7//
8// Every check reads facts Claude Code reports: which skill was opened, which
9// file a tool wrote, which program a shell command ran, whether a call
10// succeeded, and whether a turn is ending. Nothing reads the agent's words,
11// and nothing calls a model. The checks prove a step happened, not that it was
12// done well, and never that the owner approved anything.
13//
14// On a failed check the main agent redoes the step:
15// - A tool call is refused with the protocol's `tell`. Tool checks fail
16//   closed: a check that cannot decide refuses the call.
17// - A final reply is held once per protocol per turn: dropped before display,
18//   then the classic Stop event continues the turn with a note that says what
19//   is missing and which skill to open. The reply hold fails open: on an
20//   error the reply is shown with a notice.
21// After two engine errors in one turn the engine stops for that turn, leaves
22// the `toolkit_protocol_engine` field off the classic Stop hook input so the
23// old command hooks run in full, and shows one notice.
24//
25// Memory mode (issue #404): `.toolkit-memory.json` at the project root says
26// whether working and lasting memory live in `knowledge/` files ("files") or
27// in a memory service reached through an MCP server ("external"). A missing
28// or invalid file means "files". In "external" mode the engine also reads the
29// service's tool calls by class (`memory-write`, `memory-read`) and the
30// `toolkit_kind` field those calls carry: an exact field comparison.
31//
32// State is kept in module memory, by session. A new session id (/clear, a
33// resume) starts clean; session start and a plugin reload clear it;
34// compaction clears which skills were opened.
35import type { Register } from 'claude-code'
36import { STAGE_LABEL, changesWorkItem, fileWords, nodeScript, onlyReads, readCommand, reviewActions } from './shell-reader.mjs'
37import { validMemory } from './memory-config.mjs'
38
39const VERSION = '0.3.0'
40const FILE_TOOLS = new Set(['Write', 'Edit', 'MultiEdit', 'NotebookEdit'])
41const ACTIONS = ['pr-create', 'pr-merge', 'issue-close', 'work-finish']
42const NO_MENTION = 'Do not mention this check in your reply.'
43const OFF_NOTICE = 'Workflow checks are off for this turn after an error.'
44// Who is speaking: a held reply comes back as a Stop hook's continuation, so
45// the note names its source and why it asks for silence.
46const HOLD_INTRO = 'This note is from protocol-guard, the required workflow check that the project owner turned on in .claude/settings.json.'
47const OWNER_ASKED = 'The owner set up these checks and wants replies about the work, not about the checks.'
48const REFUSAL_INTRO = 'This refusal is from protocol-guard, the required workflow check that the project owner turned on in .claude/settings.json.'
49
50// A refusal names its source and leaves out the "Do not mention" line. In
51// print-mode runs the model took a refusal with that line for a prompt
52// injection (1 of 10 followed it; 4 of 4 without it). This departs from
53// design decision 4 for refusals only, pending the owner's confirmation; the
54// held-reply note keeps the line.
55function refusalText(failed: string[]): string {
56  return `${REFUSAL_INTRO}\n${[...new Set(failed)].join('\n')}`
57}
58
59type Owner = { skill?: string; file?: string }
60type Require =
61  | { opened: 'owner'; within: 'turn' | 'reset' | 'session'; paths?: string[]; kind?: string }
62  | { never: true }
63  | { wrote: string[] }
64  | { ran: string[] }
65  | { called: string; kind?: string }
66type MemoryMode = 'files' | 'external'
67// Every key must hold. A list of conditions holds when any one of them holds.
68type Condition = { exists?: string; memory?: MemoryMode }
69type Protocol = {
70  name: string
71  why?: string
72  owner: Owner
73  appliesIf?: Condition | Condition[]
74  on: {
75    action?: string[]
76    write?: string[]
77    shell?: boolean
78    oneWriter?: string[]
79    call?: string[]
80    turnEnd?: { afterWorkItemChange?: boolean; afterWrite?: string[] }
81  }
82  require: Require[]
83  tell: string
84}
85type NewFact =
86  | { kind: 'write'; path: string; agent: string; certain: boolean }
87  | { kind: 'work-item'; agent: string }
88  | { kind: 'ran'; script: string; agent: string }
89  | { kind: 'call'; classes: string[]; memoryKind?: string; deletes?: string; agent: string }
90type Fact = NewFact & { seq: number }
91type Loop = { turn: Set<string>; reset: Set<string>; session: Set<string> }
92type Turn = { holds: Set<string>; errors: number; off: boolean; offNotice: boolean }
93type Memory = { mode: MemoryMode; service?: string; server?: string }
94type State = {
95  key: string
96  root: string
97  memory: Memory
98  protocols: Protocol[]
99  notes: string[]
100  loops: Map<string, Loop>
101  facts: Fact[]
102  seq: number
103  turn: Turn
104  writers: Map<string, string>
105  pendingStop?: string
106  pendingDrops: number
107  // The last fact before this turn: a turn-end check looks only at facts of
108  // this turn, except for a check carried over from a helper save that was
109  // still running when the last turn ended.
110  turnSeq: number
111  carried: Map<string, number>
112  carriedNow: Map<string, number>
113  // Checks switched off because their owner skill is not installed; the
114  // owner is told once.
115  missing: string[]
116  missingShown: boolean
117  // The first draft reply held this turn, shown if the turn ends without one.
118  heldDraft?: string
119}
120
121// Memory-service tools by class, per service (`mcp__<server>__<tool>`).
122const TOOL_CLASSES: Record<string, Record<string, string[]>> = {
123  mem0: {
124    'memory-write': ['add_memory', 'update_memory', 'delete_memory', 'delete_all_memories', 'delete_entities'],
125    'memory-read': ['get_memories', 'get_memory', 'search_memories'],
126  },
127  hindsight: {
128    // update_bank is left out: it changes the bank's settings at setup, not
129    // memory text.
130    'memory-write': ['retain', 'sync_retain', 'delete_document', 'clear_memories', 'update_memory', 'invalidate_memory', 'delete_bank'],
131    'memory-read': ['list_documents', 'get_document', 'recall', 'list_memories', 'get_memory'],
132  },
133}
134const CALL_CLASSES = ['memory-write', 'memory-read']
135// The kind a delete call removes: '*' for any kind (a mem0 delete_memory or
136// delete_all_memories carries only an id), the prefix of a Hindsight
137// document_id before ':' (the key names the kind, such as "working:issue-42"),
138// or undefined. Every other write without a kind meets no kind requirement:
139// Hindsight clear_memories, update_memory, invalidate_memory and delete_bank
140// (they name no document), and mem0 delete_entities.
141function deletedKind(service: string | undefined, name: string, e: any): string | undefined {
142  if (service === 'mem0' && (name === 'delete_memory' || name === 'delete_all_memories')) return '*'
143  if (service === 'hindsight' && name === 'delete_document' && typeof e?.document_id === 'string') {
144    const at = e.document_id.indexOf(':')
145    return at > 0 ? e.document_id.slice(0, at) : undefined
146  }
147  return undefined
148}
149const MEMORY_CONFIG = '.toolkit-memory.json'
150// Where a helper's save lands, by mode: a helper that wrote here (or, in
151// external mode, called a memory-write tool) and is still running is a
152// pending save.
153const MEMORY_PATHS: Record<MemoryMode, string[]> = { files: ['knowledge/'], external: ['prds/'] }
154
155const states = new Map<string, State>()
156let skillCallsInFlight = 0
157
158// ---------- paths and names (identifiers only; no language is read) ----------
159
160function slashes(p: string): string {
161  let s = String(p ?? '').split('\\').join('/')
162  if (/^[A-Za-z]:\//.test(s)) s = s[0].toLowerCase() + s.slice(1)
163  return s
164}
165
166function normalize(p: string): string {
167  const abs = p.startsWith('/') || /^[a-z]:\//.test(p)
168  const out: string[] = []
169  for (const part of p.split('/')) {
170    if (part === '' || part === '.') continue
171    if (part === '..') { out.pop(); continue }
172    out.push(part)
173  }
174  const joined = out.join('/')
175  return abs && !/^[a-z]:/.test(joined) ? `/${joined}` : joined
176}
177
178function absolute(base: string, p: string): string {
179  const s = slashes(p)
180  if (s.startsWith('/') || /^[a-z]:\//.test(s)) return normalize(s)
181  if (s.startsWith('~')) return s
182  return normalize(`${slashes(base)}/${s}`)
183}
184
185function relative(root: string, abs: string): string {
186  const r = normalize(slashes(root))
187  const a = normalize(slashes(abs))
188  if (a === r) return ''
189  return a.startsWith(r + '/') ? a.slice(r.length + 1) : a
190}
191
192function covers(rel: string, entries: readonly string[]): boolean {
193  // An entry ending in "/" covers that folder and everything under it.
194  return entries.some((f) => (f.endsWith('/') ? rel.startsWith(f) || rel === f.slice(0, -1) : rel === f))
195}
196
197function skillName(raw: unknown): string {
198  // A skill id is "plugin:skill" or "skill".
199  const s = String(raw ?? '')
200  const i = s.lastIndexOf(':')
201  return i >= 0 ? s.slice(i + 1) : s
202}
203
204function ownerKey(o: Owner): string {
205  return o.skill !== undefined ? `skill:${o.skill}` : `file:${o.file ?? ''}`
206}
207
208// ---------- the protocol list ----------
209
210function isStrings(v: unknown): v is string[] {
211  return Array.isArray(v) && v.every((x) => typeof x === 'string')
212}
213
214// A protocol that uses a word this engine does not know is dropped with a note.
215function validProtocol(p: any): p is Protocol {
216  if (p === null || typeof p !== 'object') return false
217  if (typeof p.name !== 'string' || typeof p.tell !== 'string') return false
218  if (p.owner === null || typeof p.owner !== 'object' || (typeof p.owner.skill !== 'string' && typeof p.owner.file !== 'string')) return false
219  if (p.appliesIf !== undefined) {
220    const conds = Array.isArray(p.appliesIf) ? p.appliesIf : [p.appliesIf]
221    if (conds.length === 0 || !conds.every(validCondition)) return false
222  }
223  const on = p.on
224  if (on === null || typeof on !== 'object') return false
225  for (const k of Object.keys(on)) if (!['action', 'write', 'shell', 'oneWriter', 'call', 'turnEnd'].includes(k)) return false
226  if (on.call !== undefined && !(isStrings(on.call) && on.call.length > 0 && on.call.every((x: string) => CALL_CLASSES.includes(x)))) return false
227  if (on.action !== undefined && !(isStrings(on.action) && on.action.every((x: string) => ACTIONS.includes(x)))) return false
228  if (on.write !== undefined && !isStrings(on.write)) return false
229  if (on.oneWriter !== undefined && !isStrings(on.oneWriter)) return false
230  if (on.shell !== undefined && typeof on.shell !== 'boolean') return false
231  if (on.turnEnd !== undefined) {
232    const t = on.turnEnd
233    for (const k of Object.keys(t)) if (!['afterWorkItemChange', 'afterWrite'].includes(k)) return false
234    if (t.afterWrite !== undefined && !isStrings(t.afterWrite)) return false
235  }
236  if (on.write === undefined && on.turnEnd === undefined && on.action === undefined && on.call === undefined) return false
237  if (!Array.isArray(p.require) || p.require.length === 0) return false
238  for (const r of p.require) {
239    if (r === null || typeof r !== 'object') return false
240    const keys = Object.keys(r).sort().join(',')
241    if (keys === 'never' && r.never === true) continue
242    if (keys === 'wrote' && isStrings(r.wrote)) continue
243    if (keys === 'ran' && isStrings(r.ran) && r.ran.length > 0) continue
244    if ((keys === 'called' || keys === 'called,kind') && CALL_CLASSES.includes(r.called) && (r.kind === undefined || (typeof r.kind === 'string' && r.kind !== ''))) continue
245    if (['opened,within', 'opened,paths,within', 'kind,opened,within', 'kind,opened,paths,within'].includes(keys) && r.opened === 'owner' && (r.within === 'turn' || r.within === 'reset' || r.within === 'session') && (r.paths === undefined || isStrings(r.paths)) && (r.kind === undefined || (typeof r.kind === 'string' && r.kind !== ''))) continue
246    return false
247  }
248  return true
249}
250
251function validCondition(c: any): boolean {
252  if (c === null || typeof c !== 'object' || Array.isArray(c)) return false
253  const keys = Object.keys(c)
254  if (keys.length === 0 || !keys.every((k) => k === 'exists' || k === 'memory')) return false
255  if (c.exists !== undefined && typeof c.exists !== 'string') return false
256  if (c.memory !== undefined && c.memory !== 'files' && c.memory !== 'external') return false
257  return true
258}
259
260// The memory mode from .toolkit-memory.json, by the rules in
261// memory-config.mjs. Missing means files. A file that cannot be read or is
262// not valid also means files (the checks
263// fail open to today's behaviour), and the agent is told once.
264async function readMemory($: any, s: State): Promise<Memory> {
265  const path = `${s.root}/${MEMORY_CONFIG}`
266  let exists = false
267  try { exists = await $.fs.exists(path) } catch { exists = false }
268  if (!exists) return { mode: 'files' }
269  let c: any
270  try { c = JSON.parse(await $.fs.read(path)) } catch (err) {
271    s.notes.push(`The memory config ${MEMORY_CONFIG} could not be read (${String(err)}). The workflow checks treat this project as memory mode "files". Tell the owner once.`)
272    return { mode: 'files' }
273  }
274  const mode = validMemory(c)
275  if (mode === 'files') return { mode: 'files' }
276  if (mode === 'external') return { mode: 'external', service: c.service, server: c.server.trim() }
277  s.notes.push(`The memory config ${MEMORY_CONFIG} is not valid: it needs "format": 1 and "memory" set to "files" or "external", and "external" needs "service" (mem0 or hindsight), "server" (letters, digits, _ and - only) and "project". The workflow checks treat this project as memory mode "files". Tell the owner once.`)
278  return { mode: 'files' }
279}
280
281async function conditionHolds($: any, s: State, c: Condition): Promise<boolean> {
282  if (c.memory !== undefined && c.memory !== s.memory.mode) return false
283  if (c.exists !== undefined) {
284    try { return await $.fs.exists(`${s.root}/${c.exists}`) } catch { return false }
285  }
286  return true
287}
288
289// The classes a tool belongs to for the configured memory service, if any.
290// A valid server name holds only A-Z, a-z, 0-9, _ and -, so the tool name
291// Claude Code builds from it is mcp__<server>__<tool> unchanged.
292function toolClasses(s: State, tool: unknown): { classes: string[]; name: string } {
293  if (s.memory.mode !== 'external' || typeof tool !== 'string') return { classes: [], name: '' }
294  const prefix = `mcp__${s.memory.server}__`
295  if (!tool.startsWith(prefix)) return { classes: [], name: '' }
296  const name = tool.slice(prefix.length)
297  const table = TOOL_CLASSES[s.memory.service ?? ''] ?? {}
298  return { classes: Object.keys(table).filter((c) => table[c].includes(name)), name }
299}
300
301// A write call's kind: args.metadata.toolkit_kind, or a tag
302// "toolkit_kind:<kind>". An exact field comparison; no text is interpreted.
303function memoryKind(e: any): string | undefined {
304  let meta = e?.metadata
305  if (typeof meta === 'string') { try { meta = JSON.parse(meta) } catch { meta = undefined } }
306  if (meta !== null && typeof meta === 'object' && typeof meta.toolkit_kind === 'string') return meta.toolkit_kind
307  if (Array.isArray(e?.tags)) {
308    const tag = e.tags.find((t: unknown) => typeof t === 'string' && t.startsWith('toolkit_kind:'))
309    if (tag !== undefined) return tag.slice('toolkit_kind:'.length)
310  }
311  return undefined
312}
313
314async function load($: any, s: State) {
315  let defaults: unknown[] = []
316  try {
317    const parsed = JSON.parse(await $.fs.read(`${$.plugin.root}/protocols.default.json`))
318    defaults = Array.isArray(parsed?.protocols) ? parsed.protocols : []
319  } catch (err) {
320    s.notes.push(`The shipped protocol list could not be read (${String(err)}). No required workflow checks run in this session. Tell the owner once.`)
321  }
322  let off: string[] = []
323  let extra: unknown[] = []
324  const projectFile = `${s.root}/.claude/protocols.json`
325  let exists = false
326  try { exists = await $.fs.exists(projectFile) } catch { exists = false }
327  if (exists) {
328    try {
329      const p = JSON.parse(await $.fs.read(projectFile))
330      off = isStrings(p?.off) ? p.off : []
331      extra = Array.isArray(p?.protocols) ? p.protocols : []
332    } catch (err) {
333      s.notes.push(`The project protocol list .claude/protocols.json could not be read (${String(err)}). The shipped defaults apply.`)
334    }
335  }
336  const byName = new Map<string, Protocol>()
337  for (const p of [...defaults, ...extra]) {
338    if (!validProtocol(p)) {
339      s.notes.push(`A protocol entry was ignored because it uses a word the engine does not know: ${JSON.stringify((p as any)?.name ?? p)}.`)
340      continue
341    }
342    byName.set(p.name, p) // a project entry with a default's name replaces it
343  }
344  s.memory = await readMemory($, s)
345  const active: Protocol[] = []
346  for (const p of byName.values()) {
347    if (off.includes(p.name)) continue
348    if (p.appliesIf !== undefined) {
349      let ok = false
350      for (const c of Array.isArray(p.appliesIf) ? p.appliesIf : [p.appliesIf]) if (await conditionHolds($, s, c)) { ok = true; break }
351      if (!ok) continue
352    }
353    active.push(p)
354  }
355  // A check whose owner skill is not installed would trap the agent: it is
356  // switched off, and the owner is told once. When the list cannot be read,
357  // every check stays on.
358  let installed: Set<string> | undefined
359  try {
360    const list = (await $.command.list()) as { name: string }[]
361    if (list.length > 0) installed = new Set(list.map((c) => skillName(c.name)))
362  } catch { installed = undefined }
363  s.protocols = active.filter((p) => {
364    if (installed === undefined || p.owner.skill === undefined || installed.has(p.owner.skill)) return true
365    s.missing.push(`${p.name} (${p.owner.skill})`)
366    return false
367  })
368  if (s.missing.length > 0) s.notes.push(`These required workflow checks are off because their skill is not installed: ${s.missing.join(', ')}.`)
369}
370
371function freshTurn(): Turn {
372  return { holds: new Set(), errors: 0, off: false, offNotice: false }
373}
374
375async function state($: any): Promise<State> {
376  const key = String(await $.session.id())
377  let s = states.get(key)
378  if (s === undefined) {
379    s = {
380      key,
381      root: slashes(String(await $.session.root())),
382      memory: { mode: 'files' },
383      protocols: [],
384      notes: [],
385      loops: new Map(),
386      facts: [],
387      seq: 0,
388      turn: freshTurn(),
389      writers: new Map(),
390      pendingDrops: 0,
391      turnSeq: 0,
392      carried: new Map(),
393      carriedNow: new Map(),
394      missing: [],
395      missingShown: false,
396    }
397    states.set(key, s)
398    await load($, s)
399  }
400  return s
401}
402
403// The command hooks that run before a tool (PreToolUse) get no engine field,
404// so the engine also names its active checks in an environment variable that
405// every process it starts inherits. It is unset while the engine is off.
406async function signal($: any, s: State) {
407  // `<session id>:<names>`: a child claude or Codex process that inherits the
408  // variable has another session id, so its hooks still run in full.
409  const names = s.turn.off ? '' : s.protocols.map((p) => p.name).join(',')
410  await $.env.set('TOOLKIT_PROTOCOL_ENGINE', names === '' ? undefined : `${s.key}:${names}`)
411}
412
413function loopOf(s: State, agentId: string | undefined): Loop {
414  const key = agentId ?? 'main'
415  let l = s.loops.get(key)
416  if (l === undefined) {
417    // A helper starts with what the main agent had opened when the helper
418    // first acted, so the main agent's opening counts for a helper it starts.
419    const main = agentId === undefined ? undefined : loopOf(s, undefined)
420    l = { turn: new Set(main?.turn ?? []), reset: new Set(main?.reset ?? []), session: new Set(main?.session ?? []) }
421    s.loops.set(key, l)
422  }
423  return l
424}
425
426function engineError(s: State | undefined, $: any, where: string, err: unknown) {
427  try { $.ui.log(`protocol-guard error in ${where}: ${String(err)}`, { to: 'debug' }) } catch { /* never throws */ }
428  if (s === undefined) return
429  s.turn.errors++
430  if (s.turn.errors >= 2) {
431    s.turn.off = true
432    s.turn.offNotice = true
433    signal($, s).catch(() => undefined)
434  }
435}
436
437function record(s: State, fact: NewFact) {
438  s.seq++
439  s.facts.push({ ...fact, seq: s.seq })
440}
441
442// ---------- reading one tool call ----------
443
444type Touch = { path: string; shell: boolean }
445
446// The project files a call would change. A shell command counts every file
447// word of every command except commands that only read.
448function touches(s: State, cwd: string, e: any): Touch[] | undefined {
449  if (FILE_TOOLS.has(e.tool)) {
450    const p = e.file_path ?? e.notebook_path
451    if (typeof p !== 'string' || p === '') return undefined // cannot decide
452    return [{ path: relative(s.root, absolute(cwd, p)), shell: false }]
453  }
454  if (e.tool === 'Bash') {
455    if (typeof e.command !== 'string') return undefined
456    const out: Touch[] = []
457    let dir = cwd
458    for (const cmd of readCommand(e.command).commands) {
459      if (cmd.program === 'cd' && typeof cmd.args[0] === 'string') { dir = absolute(dir, cmd.args[0]); continue }
460      if (onlyReads(cmd)) continue
461      for (const w of fileWords(cmd)) out.push({ path: relative(s.root, absolute(dir, w)), shell: true })
462    }
463    return out
464  }
465  return []
466}
467
468// The facts a successful call adds. Shell facts used to satisfy a
469// requirement come only from commands that must have exited 0.
470function factsOf(s: State, cwd: string, e: any, agent: string): NewFact[] {
471  const out: NewFact[] = []
472  if (FILE_TOOLS.has(e.tool)) {
473    const p = e.file_path ?? e.notebook_path
474    if (typeof p === 'string') out.push({ kind: 'write', path: relative(s.root, absolute(cwd, p)), agent, certain: true })
475    return out
476  }
477  if (e.tool === 'Bash' && typeof e.command === 'string') {
478    const { commands, certain } = readCommand(e.command)
479    let dir = cwd
480    for (const cmd of commands) {
481      if (cmd.program === 'cd' && typeof cmd.args[0] === 'string') { dir = absolute(dir, cmd.args[0]); continue }
482      const sure = certain.includes(cmd)
483      if (changesWorkItem(cmd)) out.push({ kind: 'work-item', agent })
484      const script = nodeScript(cmd)
485      if (script !== undefined && sure) out.push({ kind: 'ran', script, agent })
486      if (!onlyReads(cmd)) for (const w of fileWords(cmd)) out.push({ kind: 'write', path: relative(s.root, absolute(dir, w)), agent, certain: sure })
487    }
488    return out
489  }
490  if (typeof e.tool === 'string' && e.tool.startsWith('mcp__')) {
491    const mem = toolClasses(s, e.tool)
492    if (mem.classes.length > 0) {
493      out.push({ kind: 'call', classes: mem.classes, memoryKind: memoryKind(e), deletes: deletedKind(s.memory.service, mem.name, e), agent })
494      return out
495    }
496    const name = e.tool.slice(e.tool.lastIndexOf('__') + 2)
497    // A label change counts only when a label has the stage form, such as 08-build.
498    const stage = Array.isArray(e.labels) && e.labels.some((l: unknown) => STAGE_LABEL.test(String(l)))
499    const changesState = e.state !== undefined || stage
500    if ((name === 'issue_write' && (e.method === 'create' || (e.method === 'update' && changesState)))
501      || name === 'create_issue'
502      || (name === 'update_issue' && changesState)) out.push({ kind: 'work-item', agent })
503  }
504  return out
505}
506
507// The review actions a call takes: shell commands, or GitHub tools by name.
508function actionsOf(e: any): string[] {
509  if (e.tool === 'Bash') {
510    if (typeof e.command !== 'string') return []
511    return readCommand(e.command).commands.flatMap((c: any) => reviewActions(c))
512  }
513  if (typeof e.tool === 'string' && e.tool.startsWith('mcp__')) {
514    const name = e.tool.slice(e.tool.lastIndexOf('__') + 2)
515    if (name === 'create_pull_request') return ['pr-create']
516    if (name === 'merge_pull_request' || name === 'enable_pr_auto_merge') return ['pr-merge']
517    if ((name === 'issue_write' && e.method === 'update' && e.state === 'closed') || (name === 'update_issue' && e.state === 'closed')) return ['issue-close']
518  }
519  return []
520}
521
522function succeeded(e: any, r: any): boolean {
523  if (r === undefined || r === null || r.deny !== undefined || r.isError === true) return false
524  if (e.tool === 'Bash') {
525    const res = r.result ?? {}
526    if (res.interrupted === true || res.backgroundTaskId !== undefined) return false
527  }
528  return true
529}
530
531// ---------- the checks ----------
532
533function windowOf(loop: Loop, within: 'turn' | 'reset' | 'session'): Set<string> {
534  return within === 'turn' ? loop.turn : within === 'reset' ? loop.reset : loop.session
535}
536
537// The refusal for a call, or undefined to let it run.
538async function refusal($: any, s: State, e: any, loop: Loop, agent: string): Promise<string | undefined> {
539  const failed: string[] = []
540  const actions = actionsOf(e)
541  if (actions.length > 0) {
542    for (const p of s.protocols) {
543      if (p.on.action === undefined || !actions.some((a) => p.on.action!.includes(a))) continue
544      const opened = p.require.find((r): r is Extract<Require, { opened: 'owner' }> => 'opened' in r)
545      if (opened === undefined) continue
546      const window = opened.within === 'turn' ? loop.turn : opened.within === 'reset' ? loop.reset : loop.session
547      if (!window.has(ownerKey(p.owner))) failed.push(`Required workflow check ${p.name}: ${p.tell}`)
548    }
549  }
550  const mem = toolClasses(s, e.tool)
551  if (mem.classes.length > 0) {
552    const kind = memoryKind(e)
553    for (const p of s.protocols) {
554      if (p.on.call === undefined || !mem.classes.some((c) => p.on.call!.includes(c))) continue
555      // A requirement scoped to a kind applies only to a call that carries it.
556      const opened = p.require.find((r): r is Extract<Require, { opened: 'owner' }> => 'opened' in r && (r.kind === undefined || r.kind === kind))
557      if (opened === undefined) continue
558      if (!windowOf(loop, opened.within).has(ownerKey(p.owner))) failed.push(`Required workflow check ${p.name}: ${p.tell}`)
559    }
560  }
561  const writers = s.protocols.filter((p) => p.on.write !== undefined)
562  if (writers.length === 0 || !(FILE_TOOLS.has(e.tool) || e.tool === 'Bash')) return failed.length === 0 ? undefined : refusalText(failed)
563  const cwd = slashes(String(await $.session.cwd()))
564  const t = touches(s, cwd, e)
565  if (t === undefined) return refusalText(['Required workflow check: the file this call changes could not be read, so the call was refused. Name the file path in the call.'])
566  for (const p of writers) {
567    for (const touch of t) {
568      if (touch.shell && p.on.shell !== true) continue
569      if (!covers(touch.path, p.on.write ?? [])) continue
570      if (p.require.some((r) => 'never' in r)) { failed.push(`Required workflow check ${p.name}: ${p.tell}`); break }
571      const opened = p.require.find((r): r is Extract<Require, { opened: 'owner' }> => 'opened' in r && r.kind === undefined && (r.paths === undefined || covers(touch.path, r.paths)))
572      if (opened !== undefined) {
573        const has = (opened.within === 'turn' ? loop.turn : opened.within === 'reset' ? loop.reset : loop.session).has(ownerKey(p.owner))
574        if (!has) { failed.push(`Required workflow check ${p.name}: ${p.tell}`); break }
575      }
576      if (p.on.oneWriter !== undefined && covers(touch.path, p.on.oneWriter)) {
577        const other = s.writers.get(touch.path)
578        if (other !== undefined && other !== agent && other !== 'main') {
579          const running = ((await $.agent.list()) as any[]).some((a) => a.id === other && a.status === 'running')
580          if (running) { failed.push(`Required workflow check ${p.name}: another agent's save of ${touch.path} is still running. Wait for it to finish, then read the file again before changing it.`); break }
581        }
582      }
583    }
584  }
585  if (failed.length === 0) return undefined
586  return refusalText(failed)
587}
588
589// Why a turn-end protocol is not met, or undefined when it is met or not due.
590function unmetReason(s: State, p: Protocol): string | undefined {
591  const t = p.on.turnEnd
592  if (t === undefined) return undefined
593  let trigger: Fact | undefined
594  const from = s.carriedNow.get(p.name) ?? s.turnSeq
595  for (const f of s.facts) {
596    if (f.seq <= from) continue
597    if (t.afterWorkItemChange === true && f.kind === 'work-item') trigger = f
598    if (t.afterWrite !== undefined && f.kind === 'write' && covers(f.path, t.afterWrite)) trigger = f
599  }
600  if (trigger === undefined) return undefined
601  const after = s.facts.filter((f) => f.seq > trigger!.seq)
602  const what = trigger.kind === 'write' ? `${trigger.path} changed` : 'a work item changed'
603  for (const r of p.require) {
604    if ('wrote' in r) {
605      if (!after.some((f) => f.kind === 'write' && f.certain && covers(f.path, r.wrote))) return `${what}, and ${r.wrote.join(', ')} was not written after that`
606    }
607    if ('called' in r) {
608      const met = after.some((f) => f.kind === 'call' && f.classes.includes(r.called) && (r.kind === undefined || f.memoryKind === r.kind || f.deletes === '*' || f.deletes === r.kind))
609      if (!met) return `${what}, and no ${r.called} call${r.kind === undefined ? '' : ` with toolkit_kind ${r.kind}`} to the memory service followed`
610    }
611    if ('ran' in r) {
612      let from = trigger.seq
613      for (const script of r.ran) {
614        const hit = after.find((f) => f.seq > from && f.kind === 'ran' && f.script === script)
615        if (hit === undefined) return `${what}, and ${r.ran.join(' then ')} did not both finish with exit code 0 after that`
616        from = hit.seq
617      }
618    }
619  }
620  return undefined
621}
622
623// A save by a helper that is still running is pending: the obligation stays
624// open for the next turn, and the reply is not held for it now.
625// Only a helper that saved memory since `since` counts: this turn, or, for a
626// carried check, since the turn it was first carried from. A save is a write
627// under the mode's memory paths or, in external mode, a memory-write call.
628async function helperSavePending($: any, s: State, since: number): Promise<boolean> {
629  const paths = MEMORY_PATHS[s.memory.mode]
630  const saved = (f: Fact) => (f.kind === 'write' && covers(f.path, paths)) || (f.kind === 'call' && f.classes.includes('memory-write'))
631  const savers = new Set(s.facts.filter((f) => f.seq > since && f.agent !== 'main' && saved(f)).map((f) => f.agent))
632  if (savers.size === 0) return false
633  const running = ((await $.agent.list()) as any[]).filter((a) => a.status === 'running').map((a) => a.id)
634  return running.some((id) => savers.has(id))
635}
636
637async function openAtTurnEnd($: any, s: State): Promise<{ p: Protocol; reason: string }[]> {
638  const out: { p: Protocol; reason: string }[] = []
639  for (const p of s.protocols) {
640    const reason = unmetReason(s, p)
641    if (reason === undefined) { s.carried.delete(p.name); continue }
642    // Pending: not held now; checked again at the end of the next turn, from
643    // the turn it was first carried.
644    const since = s.carriedNow.get(p.name) ?? s.turnSeq
645    if (await helperSavePending($, s, since)) { s.carried.set(p.name, since); continue }
646    out.push({ p, reason })
647  }
648  return out
649}
650
651// ---------- the engine ----------
652
653export const register: Register = (on) => {
654  on('session.start', async ($, e, next) => {
655    try {
656      // A plugin reload fires this again: start this session's record over.
657      states.delete(String(await $.session.id()))
658      await signal($, await state($))
659    } catch (err) { engineError(undefined, $, 'session.start', err) }
660    return next(e)
661  })
662
663  on('session.end', async ($, e, next) => {
664    states.delete(String(e.sessionId))
665    return next(e)
666  })
667
668  on('turn.start', async ($, e, next) => {
669    try {
670      const s = await state($)
671      s.turn = freshTurn()
672      s.pendingStop = undefined
673      s.pendingDrops = 0
674      s.heldDraft = undefined
675      s.turnSeq = s.seq
676      s.carriedNow = s.carried
677      s.carried = new Map()
678      loopOf(s, undefined)
679      for (const l of s.loops.values()) l.turn.clear()
680      await signal($, s)
681    } catch (err) { engineError(undefined, $, 'turn.start', err) }
682    return next(e)
683  })
684
685  on('session.compact', async ($, e, next) => {
686    const r: any = await next(e)
687    try {
688      if (e.trigger !== 'precompute' && r?.messages !== undefined) {
689        const s = await state($)
690        // The owner's text may no longer be in context: skills must be opened again.
691        const loops = e.agentId === undefined ? [...s.loops.values()] : [loopOf(s, e.agentId)]
692        for (const l of loops) { l.turn.clear(); l.reset.clear() }
693      }
694    } catch (err) { engineError(undefined, $, 'session.compact', err) }
695    return r
696  })
697
698  // A load problem is told to the agent once, with the next prompt.
699  on('prompt.submit', async ($, e, next) => {
700    try {
701      const s = await state($)
702      if (s.notes.length > 0) {
703        const notes = s.notes.splice(0)
704        return next({ ...e, context: [...(e.context ?? []), ...notes] })
705      }
706    } catch (err) { engineError(undefined, $, 'prompt.submit', err) }
707    return next(e)
708  })
709
710  // Old command hooks read this field to know which checks run now.
711  on('classic.UserPromptSubmit', async ($, e, next) => {
712    try {
713      const s = await state($)
714      return next({ ...e, toolkit_protocol_engine: { version: VERSION, active: s.protocols.map((p) => p.name) } } as typeof e)
715    } catch { return next(e) }
716  })
717  on('classic.Stop', async ($, e, next) => {
718    let s: State | undefined
719    try { s = await state($) } catch { return next(e) }
720    const pending = s.pendingStop
721    s.pendingStop = undefined
722    const r = s.turn.off
723      ? await next(e)
724      : await next({ ...e, toolkit_protocol_engine: { version: VERSION, active: s.protocols.map((p) => p.name) } } as typeof e)
725    // A held reply: continue the turn with the note.
726    if (pending !== undefined && !s.turn.off) return { ...r, block: r.block ? `${r.block}\n\n${pending}` : pending }
727    return r
728  })
729
730  // A skill typed as /name, or preloaded, opens it for the main agent.
731  on('skill.prompt', async ($, e, next) => {
732    const r = await next(e)
733    try {
734      if (skillCallsInFlight === 0) {
735        const s = await state($)
736        const main = loopOf(s, undefined)
737        const key = `skill:${skillName(e.skill)}`
738        main.turn.add(key); main.reset.add(key); main.session.add(key)
739      }
740    } catch (err) { engineError(undefined, $, 'skill.prompt', err) }
741    return r
742  })
743
744  // The tool guard and the fact recorder.
745  on('tool.call', async ($, e, next) => {
746    const ev: any = e
747    let s: State | undefined
748    let loop: Loop | undefined
749    const agent = e.agentId ?? 'main'
750    try {
751      s = await state($)
752      loop = loopOf(s, e.agentId)
753    } catch (err) {
754      engineError(undefined, $, 'tool.call', err)
755      return { deny: refusalText(["Required workflow check could not run. Try the call again."]) }
756    }
757    if (s.turn.off) return next(e)
758
759    // Before the call: fails closed.
760    if (FILE_TOOLS.has(ev.tool) || ev.tool === 'Bash' || String(ev.tool).startsWith('mcp__')) {
761      try {
762        const why = await refusal($, s, ev, loop, agent)
763        if (why !== undefined) return { deny: why }
764      } catch (err) {
765        engineError(s, $, 'tool check', err)
766        return { deny: refusalText(["Required workflow check could not run, so the call was refused. Try it again."]) }
767      }
768    }
769
770    if (ev.tool === 'Skill') skillCallsInFlight++
771    let r: any
772    try { r = await next(e) } finally { if (ev.tool === 'Skill') skillCallsInFlight-- }
773
774    // After the call: only a call that succeeded counts, in order.
775    try {
776      if (!succeeded(ev, r)) return r
777      if (ev.tool === 'Skill') {
778        const key = `skill:${skillName(ev.skill)}`
779        loop.turn.add(key); loop.reset.add(key); loop.session.add(key)
780        return r
781      }
782      if (ev.tool === 'Read' && typeof ev.file_path === 'string') {
783        const abs = absolute(slashes(String(await $.session.cwd())), ev.file_path)
784        if (abs.endsWith('/SKILL.md')) {
785          const parts = abs.split('/')
786          const key = `skill:${parts[parts.length - 2]}`
787          loop.turn.add(key); loop.reset.add(key); loop.session.add(key)
788        }
789        const rel = relative(s.root, abs)
790        for (const p of s.protocols) {
791          if (p.owner.file !== undefined && p.owner.file === rel) { loop.turn.add(ownerKey(p.owner)); loop.reset.add(ownerKey(p.owner)); loop.session.add(ownerKey(p.owner)) }
792        }
793        return r
794      }
795      const cwd = slashes(String(await $.session.cwd()))
796      for (const f of factsOf(s, cwd, ev, agent)) {
797        record(s, f)
798        if (f.kind === 'write') s.writers.set(f.path, agent)
799      }
800    } catch (err) { engineError(s, $, 'tool record', err) }
801    return r
802  }).catch(async ($, e, next) => {
803    if (next.called) return undefined // the tool already ran; nothing left to refuse
804    const s = states.get(String(await $.session.id()))
805    engineError(s, $, 'tool.call timeout', next.error.message ?? next.error.kind)
806    if (s?.turn.off === true) return undefined
807    return { deny: refusalText(["Required workflow check could not run, so the call was refused. Try it again."]) }
808  })
809
810  // The reply hold. Fails open: on any problem the reply is shown.
811  on('turn.step', async function* ($, e, next) {
812    if (e.agentId !== undefined) return yield* next(e) // helpers' replies are not shown to the reader
813    let s: State | undefined
814    let open: { p: Protocol; reason: string }[] = []
815    try {
816      s = await state($)
817      if (!s.turn.off) open = (await openAtTurnEnd($, s)).filter((o) => !s!.turn.holds.has(o.p.name))
818    } catch (err) {
819      engineError(s, $, 'reply check', err)
820      if (s !== undefined) s.turn.offNotice = true
821      open = []
822    }
823    // A reply already held this turn waits for the classic Stop event, which
824    // continues the turn with the note. Claude Code may ask for a visible
825    // reply first; that one is held back too.
826    if (s !== undefined && s.pendingStop !== undefined && s.pendingDrops < 2) {
827      s.pendingDrops++
828      const waiting = next(e)
829      for await (const c of waiting) if (c.kind !== 'text') yield c
830      return { ...(await waiting.result), answer: '' }
831    }
832    if (s === undefined || open.length === 0) return yield* next(e) // streams as normal
833
834    const stream = next(e)
835    const held: any[] = []
836    for await (const c of stream) held.push(c)
837    const result: any = await stream.result
838
839    let note: string | undefined
840    try {
841      if (result.stopReason === 'end_turn') {
842        const hold = (await openAtTurnEnd($, s)).filter((o) => !s!.turn.holds.has(o.p.name))
843        if (hold.length > 0) {
844          const owner = hold[0].p.owner
845          const lines = hold.map((o) => `${o.p.name}: ${o.reason}. ${o.p.tell}`)
846          note = `${HOLD_INTRO} It held back your reply before the reader saw it.\nMissing step:\n${lines.join('\n')}\nOpen ${owner.skill !== undefined ? `the ${owner.skill} skill` : owner.file}, do the missing step now, then write your reply again. ${OWNER_ASKED} ${NO_MENTION}`
847          for (const o of hold) s.turn.holds.add(o.p.name)
848          s.pendingStop = note
849          s.pendingDrops = 0
850          if (s.heldDraft === undefined) s.heldDraft = held.filter((c) => c.kind === 'text').map((c) => c.text).join('')
851        }
852      }
853    } catch (err) {
854      engineError(s, $, 'reply check', err)
855      s.turn.offNotice = true
856      note = undefined
857    }
858
859    if (note === undefined) {
860      for (const c of held) yield c
861      return result
862    }
863
864    // Drop the draft reply before display. The classic Stop event that follows
865    // then continues the turn with the note, as a Stop hook's block does, and
866    // the agent does the step and writes the reply again itself.
867    for (const c of held) if (c.kind !== 'text') yield c
868    return { ...result, answer: '' }
869  })
870
871  // One line beneath the answer: the checks were off after an error, or a
872  // required step was still missing after its one hold.
873  on('turn.complete', async ($, e, next) => {
874    const r: any = await next(e)
875    if (e.agentId !== undefined) return r
876    try {
877      const s = states.get(String(await $.session.id()))
878      if (s === undefined) return r
879      const lines: string[] = []
880      // The continuation after a hold ended with no reply (an error or an
881      // interrupt): show the held draft rather than nothing.
882      const draft = s.heldDraft
883      s.heldDraft = undefined
884      if (draft !== undefined && draft !== '' && (e.answer === '' || e.reason !== 'answer')) lines.push(draft)
885      if (s.missing.length > 0 && !s.missingShown) {
886        s.missingShown = true
887        lines.push(`Workflow checks off because their skill is not installed: ${s.missing.join(', ')}.`)
888      }
889      if (s.turn.offNotice) lines.push(OFF_NOTICE)
890      else if (e.reason === 'answer') {
891        const still = (await openAtTurnEnd($, s)).filter((o) => s.turn.holds.has(o.p.name))
892        if (still.length > 0) lines.push(`Required workflow check not met after one retry: ${still.map((o) => o.p.name).join(', ')}.`)
893      }
894      if (lines.length === 0) return r
895      return { ...r, text: lines.join('\n') }
896    } catch (err) { engineError(undefined, $, 'turn.complete', err); return r }
897  })
898}
899
hooks/shell-reader.mjs 317 lines
1// Shell command reader for the protocol-guard engine (issue #396, decision 7).
2//
3// Reads what a shell command runs: each command's program, its arguments,
4// and the files it redirects output to. It never reads the agent's words:
5// quoted text is one argument, and a heredoc body is dropped, so a commit
6// message that mentions a path is never taken for that path.
7//
8// Plain JavaScript with no imports, because a function-hook module has no
9// Node. tests/protocol-guard-check.mjs compares it with the command hooks'
10// reader, plugins/second-brain/hooks/command-parsing.mjs, on one command list.
11
12const OPERATORS = ['&&', '||', ';;', ';', '|&', '|', '&', '\n', '(', ')']
13const REDIRECT = /^(\d*|&)(>>?|>\||<)$/
14const WRAPPERS = new Set(['sudo', 'env', 'command', 'time', 'nohup', 'exec'])
15
16// Splits a command into words and operators. A word keeps its quotes removed;
17// `quoted` marks a word that had any quoted part.
18export function tokenize(command) {
19  const tokens = []
20  let word = ''
21  let inWord = false
22  let quoted = false
23  const heredocs = []
24  const push = () => {
25    if (inWord) tokens.push({ kind: 'word', text: word, quoted })
26    word = ''
27    inWord = false
28    quoted = false
29  }
30  let i = 0
31  const s = String(command ?? '')
32  while (i < s.length) {
33    const c = s[i]
34    if (c === "'") {
35      const end = s.indexOf("'", i + 1)
36      word += s.slice(i + 1, end < 0 ? s.length : end)
37      inWord = true
38      quoted = true
39      i = end < 0 ? s.length : end + 1
40      continue
41    }
42    if (c === '"') {
43      let j = i + 1
44      while (j < s.length && s[j] !== '"') {
45        if (s[j] === '\\' && j + 1 < s.length) { word += s[j + 1]; j += 2; continue }
46        word += s[j]
47        j++
48      }
49      inWord = true
50      quoted = true
51      i = j + 1
52      continue
53    }
54    if (c === '\\' && i + 1 < s.length) {
55      if (s[i + 1] !== '\n') { word += s[i + 1]; inWord = true }
56      i += 2
57      continue
58    }
59    if (c === '#' && !inWord) {
60      while (i < s.length && s[i] !== '\n') i++
61      continue
62    }
63    if (c === ' ' || c === '\t') { push(); i++; continue }
64    // A heredoc: `<<WORD`, `<<-WORD`, `<<'WORD'`. Its body is data and is dropped.
65    const here = /^<<(-|~)?[ \t]*(['"]?)([A-Za-z_][A-Za-z0-9_]*)\2/.exec(s.slice(i))
66    if (here !== null) {
67      push()
68      heredocs.push(here[3])
69      i += here[0].length
70      continue
71    }
72    if (c === '<' && s[i + 1] === '<' && s[i + 2] === '<') { push(); tokens.push({ kind: 'word', text: '<<<', quoted: false }); i += 3; continue }
73    const op = s.startsWith('&>', i) ? undefined : OPERATORS.find((o) => s.startsWith(o, i))
74    if (op !== undefined) {
75      push()
76      tokens.push({ kind: 'op', text: op })
77      i += op.length
78      if (op === '\n' && heredocs.length > 0) {
79        // Skip each pending heredoc body up to its closing line.
80        for (const tag of heredocs.splice(0)) {
81          while (i < s.length) {
82            const lineEnd = s.indexOf('\n', i)
83            const line = s.slice(i, lineEnd < 0 ? s.length : lineEnd)
84            i = lineEnd < 0 ? s.length : lineEnd + 1
85            if (line.trim() === tag) break
86          }
87        }
88      }
89      continue
90    }
91    if (c === '>' || c === '<' || (c === '&' && s[i + 1] === '>')) {
92      // A redirection operator stands alone even when written against a word.
93      // `2>&1` and `>&2` point at another descriptor, not a file: dropped.
94      let j = c === '&' ? i + 1 : i
95      while (j < s.length && (s[j] === '>' || s[j] === '<' || s[j] === '|')) j++
96      const prefix = c === '&' ? '&' : inWord && /^\d+$/.test(word) ? word : ''
97      if (prefix === '' || c === '&') push()
98      word = ''
99      inWord = false
100      const dup = /^&(\d+|-)/.exec(s.slice(j))
101      if (dup !== null) { i = j + dup[0].length; continue }
102      tokens.push({ kind: 'word', text: prefix + s.slice(c === '&' ? i + 1 : i, j), quoted: false, redirect: true })
103      i = j
104      continue
105    }
106    word += c
107    inWord = true
108    i++
109  }
110  push()
111  return tokens
112}
113
114function basename(path) {
115  const p = String(path).split('\\').join('/')
116  const i = p.lastIndexOf('/')
117  return i >= 0 ? p.slice(i + 1) : p
118}
119
120// One simple command: `{ program, args, redirects, words }`.
121// `program` is the basename of the first word after `VAR=value` prefixes and
122// wrappers such as `sudo` or `env`. `args` are the words after it. `redirects`
123// are the files output is written to (`>`, `>>`); `inputs` the files read (`<`).
124function simple(words) {
125  const redirects = []
126  const inputs = []
127  const plain = []
128  for (let i = 0; i < words.length; i++) {
129    const w = words[i]
130    if (w.redirect === true && REDIRECT.test(w.text)) {
131      const target = words[i + 1]
132      if (target !== undefined) (w.text.endsWith('<') ? inputs : redirects).push(target.text)
133      i++
134      continue
135    }
136    if (w.redirect === true) continue
137    plain.push(w)
138  }
139  let k = 0
140  while (k < plain.length && !plain[k].quoted && /^[A-Za-z_][A-Za-z0-9_]*=/.test(plain[k].text)) k++
141  while (k < plain.length && WRAPPERS.has(basename(plain[k].text))) {
142    k++
143    while (k < plain.length && (plain[k].text.startsWith('-') || /^[A-Za-z_][A-Za-z0-9_]*=/.test(plain[k].text))) k++
144  }
145  const program = k < plain.length ? basename(plain[k].text) : ''
146  return { program, args: plain.slice(k + 1).map((w) => w.text), redirects, inputs }
147}
148
149// Reads a whole command line.
150//
151// Returns `{ commands, certain }`. `commands` lists every simple command in
152// order. `certain` lists the ones that must have exited 0 when the whole
153// command exited 0: the commands of the last `&&` chain, leaving out any
154// command whose output was piped to another. A command before `;`, `||`, `&`
155// or a newline may have failed without failing the whole line.
156export function readCommand(command) {
157  const tokens = tokenize(command)
158  const commands = []
159  let certain = []
160  let newChain = false
161  let words = []
162  const end = (op) => {
163    if (words.length > 0) {
164      const cmd = simple(words)
165      commands.push(cmd)
166      if (newChain) { certain = []; newChain = false }
167      // Piped output, or a command sent to the background: its exit code is
168      // not the line's exit code.
169      if (op !== '|' && op !== '|&' && op !== '&') certain.push(cmd)
170    }
171    words = []
172    if (op === ';' || op === '||' || op === '&' || op === '\n' || op === ';;') newChain = true
173  }
174  for (const t of tokens) {
175    if (t.kind === 'op') {
176      end(t.text === '(' || t.text === ')' ? '&&' : t.text)
177      continue
178    }
179    words.push(t)
180  }
181  end('end')
182  return { commands, certain }
183}
184
185// The words of a command that name files it may write: its arguments that are
186// not flags, and its redirection targets. A flag's own value (`-m text`) is not
187// known to be a file, so a word after a flag still counts: a word is only a
188// file when it matches a path the caller asks about. `cp`, `install` and `ln`
189// write only their target: `-t <dir>` or `--target-directory=<dir>`, or else
190// the last word.
191const TARGET_ONLY = new Set(['cp', 'install', 'ln'])
192export function fileWords(cmd) {
193  const words = cmd.args.filter((a) => a !== '' && !a.startsWith('-'))
194  if (TARGET_ONLY.has(cmd.program)) {
195    const a = cmd.args
196    const t = a.findIndex((x) => x === '-t' || x === '--target-directory')
197    const eq = a.find((x) => x.startsWith('--target-directory='))
198    const target = t >= 0 ? a[t + 1] : eq !== undefined ? eq.slice('--target-directory='.length) : words[words.length - 1]
199    return [...(target === undefined ? [] : [target]), ...cmd.redirects]
200  }
201  return [...words, ...cmd.redirects]
202}
203
204// Programs that only read the files they name. A command with one of these as
205// its program and no output redirection changes no named file. `sed` counts
206// only without `-i`; `git` only with the subcommands listed.
207const READERS = new Set([
208  'cat', 'head', 'tail', 'less', 'more', 'grep', 'egrep', 'fgrep', 'rg', 'wc', 'ls', 'stat', 'file',
209  'diff', 'cmp', 'md5sum', 'sha256sum', 'shasum', 'test', '[', 'realpath', 'readlink', 'basename',
210  'dirname', 'echo', 'printf', 'jq', 'nl', 'column', 'du', 'tree', 'find', 'cd', 'pwd',
211  'cut', 'tr', 'uniq', 'od', 'xxd', 'bat', 'sort', 'awk', 'gawk',
212])
213const GIT_READERS = new Set(['status', 'diff', 'log', 'show', 'blame', 'ls-files', 'grep', 'add', 'commit', 'rev-parse', 'cat-file', 'check-ignore'])
214
215// git's subcommand: the first word after its global options. `-C <dir>`,
216// `-c <key=value>`, `--git-dir <dir>` and `--work-tree <dir>` take a value.
217export function gitSubcommand(cmd) {
218  const a = cmd.args
219  for (let i = 0; i < a.length; i++) {
220    if (a[i] === '-C' || a[i] === '-c' || a[i] === '--git-dir' || a[i] === '--work-tree' || a[i] === '--namespace') { i++; continue }
221    if (a[i].startsWith('-')) continue
222    return a[i]
223  }
224  return undefined
225}
226
227export function onlyReads(cmd) {
228  if (cmd.redirects.some((r) => r !== '/dev/null')) return false
229  const a = cmd.args
230  if (cmd.program === 'sed') return !a.some((x) => /^-[a-zA-Z]*i/.test(x) || x.startsWith('--in-place'))
231  if (cmd.program === 'find') return !a.some((x) => x === '-delete' || x === '-exec' || x === '-execdir' || x === '-fprint')
232  if (cmd.program === 'sort') return !a.some((x) => /^-[a-zA-Z]*o/.test(x) || x.startsWith('--output'))
233  if (cmd.program === 'awk' || cmd.program === 'gawk') {
234    return !a.some((x, i) => x === '--inplace' || x === '-i' && a[i + 1] === 'inplace' || x === '-iinplace' || x === '--include=inplace')
235  }
236  if (cmd.program === 'git') {
237    const sub = gitSubcommand(cmd)
238    return sub !== undefined && GIT_READERS.has(sub)
239  }
240  return READERS.has(cmd.program)
241}
242
243// The script a `node` command runs: its first argument that is not a flag.
244export function nodeScript(cmd) {
245  if (cmd.program !== 'node' && cmd.program !== 'node.exe') return undefined
246  const script = cmd.args.find((a) => !a.startsWith('-'))
247  return script === undefined ? undefined : basename(script)
248}
249
250// True when an `--add-label` or `--remove-label` value names a lifecycle stage.
251export const STAGE_LABEL = /^\d{2}-/
252function stageLabels(a) {
253  for (let i = 0; i < a.length; i++) {
254    const m = /^--(add-label|remove-label)(?:=(.*))?$/.exec(a[i])
255    if (m === null) continue
256    const value = m[2] ?? a[i + 1] ?? ''
257    if (value.split(',').some((l) => STAGE_LABEL.test(l.trim()))) return true
258  }
259  return false
260}
261
262// gh's arguments after a leading `-R <repo>`, `--repo <repo>` or `--repo=<repo>`.
263export function ghArgs(a) {
264  let i = 0
265  while (i < a.length) {
266    if (a[i] === '-R' || a[i] === '--repo') { i += 2; continue }
267    if (a[i].startsWith('--repo=')) { i += 1; continue }
268    break
269  }
270  return a.slice(i)
271}
272
273// True when the command creates, closes, reopens, deletes or moves a work item,
274// or changes its stage, through `gh` or the work tracker's `work` command.
275export function changesWorkItem(cmd) {
276  let a = cmd.args
277  if (cmd.program === 'gh') {
278    a = ghArgs(a)
279    if (a[0] === 'issue' && ['create', 'new', 'close', 'reopen', 'delete', 'transfer'].includes(a[1] ?? '')) return true
280    // A label edit counts only when a label has the stage form, such as 08-build.
281    if (a[0] === 'issue' && a[1] === 'edit') return stageLabels(a)
282    if (a[0] === 'project' && a[1] === 'item-edit') return true
283    return false
284  }
285  let sub
286  if (cmd.program === 'work') sub = a.filter((x) => !x.startsWith('-'))[0]
287  else if (nodeScript(cmd) === 'work.mjs') {
288    const rest = a.slice(a.findIndex((x) => !x.startsWith('-')) + 1)
289    sub = rest.filter((x) => !x.startsWith('-'))[0]
290  } else return false
291  const rest = cmd.program === 'work' ? a : a.slice(a.findIndex((x) => !x.startsWith('-')) + 1)
292  if (['add', 'finish', 'archive', 'unarchive', 'start'].includes(sub ?? '')) return true
293  if (sub === 'update') return rest.some((x) => /^--(stage|status)(=|$)/.test(x))
294  if (sub === 'requirements') return rest.some((x) => x === '--finalize' || x === '--reopen')
295  return false
296}
297
298// The review actions a command takes: `pr-create` (gh pr create), `pr-merge`
299// (gh pr merge, with or without --auto), `issue-close` (gh issue close) and
300// `work-finish` (the work tracker's `work finish`).
301export function reviewActions(cmd) {
302  let a = cmd.args
303  if (cmd.program === 'gh') {
304    a = ghArgs(a)
305    // Help and a dry run change nothing.
306    if (a.includes('--help') || a.includes('-h')) return []
307    if (a[0] === 'pr' && a[1] === 'create') return a.includes('--dry-run') ? [] : ['pr-create']
308    if (a[0] === 'pr' && a[1] === 'merge') return ['pr-merge']
309    if (a[0] === 'issue' && a[1] === 'close') return ['issue-close']
310    return []
311  }
312  let sub
313  if (cmd.program === 'work') sub = a.filter((x) => !x.startsWith('-'))[0]
314  else if (nodeScript(cmd) === 'work.mjs') sub = a.slice(a.findIndex((x) => !x.startsWith('-')) + 1).filter((x) => !x.startsWith('-'))[0]
315  return sub === 'finish' ? ['work-finish'] : []
316}
317
hooks/memory-config.mjs 22 lines
1// The rules for a valid `.toolkit-memory.json` (issue #404). They are the same
2// rules as second-brain's readMemoryConfig
3// (plugins/second-brain/hooks/knowledge-manual.mjs). protocol-guard ships
4// without second-brain, so it keeps its own copy.
5// tests/knowledge-schema.test.mjs checks that every reader agrees.
6
7export const MEMORY_SERVICES = ['mem0', 'hindsight']
8
9const line = (v) => typeof v === 'string' && v.trim() !== '' && !/[\r\n]/.test(v)
10
11// "files" or "external" for a valid config; undefined for an invalid one.
12export function validMemory(c) {
13  if (c === null || typeof c !== 'object' || Array.isArray(c)) return undefined
14  if (c.format !== 1) return undefined
15  if (c.memory === 'files') return 'files'
16  if (c.memory !== 'external') return undefined
17  if (!MEMORY_SERVICES.includes(c.service)) return undefined
18  if (!line(c.server) || !/^[A-Za-z0-9_-]+$/.test(c.server)) return undefined
19  if (!line(c.project)) return undefined
20  return 'external'
21}
22