Implement an agreed plan with an enforced adversarial review loop: every phase commit is gated on an external review.

A Claude Code plugin that implements an agreed plan phase by phase, with an external adversarial review as an enforcement gate. Claude cannot commit a phase until a separate reviewer — a second claude process by default, or OpenCode — has reviewed the exact tree it is about to commit and passed it.
The reviewer is external in the sense that matters: a fresh process, isolated from your ambient plugins, skills and MCP servers, given the diff as evidence and a fixed prompt, with no way to write anything and no memory of the session that wrote the code.
The review is not advice Claude may weigh up. It is a PreToolUse gate on the commit itself: findings come back as a denial, and the commit only proceeds once they are resolved.
/adversarial-review-loop:implement plan.md
-> arms (freezes the baseline and the plan) before Claude has a turn
-> Claude proposes phases and freezes them
-> phase N implemented -> git commit -> INTERCEPTED
snapshot the whole working state into a tree
the reviewer reads the delta since the last approved tree
approved -> commit proceeds -> phase advances
findings -> commit denied, findings returned inline
-> all phases committed -> turn ends -> COMPLETE
(plus a final cumulative review first, if final_review is on)
PreToolUse, PostToolUse, PostToolUseFailure, Stop, SessionStart and UserPromptSubmit hooks in hooks/hooks.json)PATH and authenticated: claude by default (you already have it), or opencode with harness set to opencode. Arming refuses if the configured one is missing.python3 3.12 or newer — the gate itself. No install step: the standard library plus a vendored, lint-excluded copy of bashlex is everything it needs. On macOS the bundled /usr/bin/python3 is 3.9 and will not do; install one from Homebrew or python.org.git, bash 3.2 or newer, and an outer watchdog for the hooks: timeout/gtimeout (GNU or uutils coreutils) or perl. macOS ships perl and no timeout, so it is covered out of the box. Arming refuses if none of them is present.Add this repository as a plugin marketplace, then install the plugin:
$ claude
> /plugin marketplace add leinardi/adversarial-review-loop
> /plugin install adversarial-review-loop
Then raise the Stop-hook block cap so a long loop is not cut short. Claude Code caps consecutive Stop blocks (default 8) and overrides by ending the turn, which reads as success — so give the loop headroom in settings.json and restart Claude Code, since the variable is read at process start:
{
"env": {
"CLAUDE_CODE_STOP_HOOK_BLOCK_CAP": "40"
}
}
Sandboxing is off unless you turned it on, so most people can skip this. If you did enable it, the loop needs two write permissions that a default-deny setup will refuse, and the symptoms do not look like permissions:
| It needs to write | Because | What it looks like when denied |
|---|---|---|
the repository's .git | the snapshot layer runs git add -A against a throwaway index, which writes blobs | arming refuses a clean worktree, saying it cannot establish whether it is clean |
$XDG_STATE_HOME/adversarial-review-loop | every activation, approval and report lives there, never inside the repo under review | ARMING FAILED, naming the state directory |
A blanket "denyWrite": ["~/"] covers both, and allowWrite does not re-open it — measured: paths listed in allowWrite are still refused when the home directory is denied. Narrow the deny instead of adding exceptions:
{
"sandbox": {
"enabled": true,
"filesystem": {
"denyWrite": ["~/.ssh", "~/.aws", "~/.gnupg"],
"allowWrite": ["~/.local/state/adversarial-review-loop/**"]
}
}
}
excludedCommands does not help here: the loop's git runs inside the gate process, not as a git … command the sandbox can pattern-match. If you would rather not loosen anything, run the loop in a session with the sandbox off — it is a per-session setting under /sandbox.
On Claude Code 2.1.272 and later, a slash command's shell step runs only if Claude Code's permission check allows it. Otherwise Claude Code hands the command to Claude as [run this first, exactly as written, and use its output: …]. For /adversarial-review-loop:implement that ends in "arming never ran", because the gate will not let Claude arm the loop itself. Each skill ships its own allowed-tools rule for exactly the arl.sh subcommand it runs, so this only happens with an install older than that fix, or when your own permissions.deny or permissions.ask rule matches arl.sh. Update the plugin, or remove the rule. If a session is already stuck in NEEDS_HUMAN from a failed arm, leave the mode from a terminal outside Claude Code, in the repository: <plugin-root>/scripts/arl.sh deactivate --session <session-id>. The handed-off command shows both values.
/adversarial-review-loop:implement plan.md. Arming happens before Claude gets a turn: the baseline and the plan are frozen, and the reviewer is probed for reachability.git add -A && git commit -m "…". That commit is intercepted, reviewed, and either allowed through or denied with the findings inline. Repeat until the phase passes, then on to phase 2./adversarial-review-loop:status at any time; /adversarial-review-loop:report [n] for a review in full. While the loop is armed, a band above the prompt shows its status, phase, round and failures, and an alert (an unreviewed commit, NEEDS_HUMAN, a state that cannot be read) appears as a line in the transcript. The band is a display only: if it fails, the gate is unaffected.Pause after phase 5 with --until 5, or mid-run with /adversarial-review-loop:pause. Pick a plan back up in a new session with /adversarial-review-loop:resume, which runs to the end of the plan unless you pass --until N again — never a second implement, which re-baselines and throws away every approval.
| Command | Who | What it does | ||
|---|---|---|---|---|
/adversarial-review-loop:implement <plan.md> [--allow-dirty] [--until N] [--harness H] [--model X] [--variant V] [--guide <path>] | you | Arms the loop for this worktree and starts the phased implementation | ||
/adversarial-review-loop:resume [--until N] [--plan <path>] [--guide <path>] [--replan] [--allow-dirty] [--abandon-pending] [--harness H] [--model X] [--variant V] | you | Continues an armed activation — in a new session, or adjusts it in this one — without losing the baseline or any approval. Clears the pause target unless --until N names a new one | ||
/adversarial-review-loop:status | anyone | Current state: phase, baseline, approvals, counters, stored reports | ||
/adversarial-review-loop:report [n] | anyone | Prints a stored review in full, untruncated | ||
| `/adversarial-review-loop:pause [N \ | 0 \ | all]` | you | Moves the pause target without a re-arm: with no argument, the loop finishes and commits the phase it is on and then stops instead of continuing |
/adversarial-review-loop:finish | you | Runs the final cumulative review now, even with phases outstanding — and regardless of final_review, which makes it the way to get one on a default install | ||
/adversarial-review-loop:accept [reason] | you | Manually approves the current working tree for the current phase, without a review | ||
/adversarial-review-loop:stop | you | Leaves the mode. Nothing is reverted | ||
/adversarial-review-loop:config [<key> <value> [--repo]] [<key> --unset [--repo]] | you | Reads or writes the review-loop configuration. Unrelated to any armed activation — never registers the gate |
Every command except status and report is disable-model-invocation: true: Claude can never arm, resume, finish, stop, accept, pause, or run config itself — only the exact slash command does, and no natural-language phrasing invokes any of them. A mode whose whole point is enforcement must not be self-enabling; the cost is that "use the review loop for this" does nothing. That stops the command, not ordinary file edits to what it writes — see the honest-agent bar under Known limitations.
Resolution order: ARL_* environment → repo .adversarial-review-loop.json → $XDG_CONFIG_HOME/adversarial-review-loop/config.json → defaults. Environment variables are the upper-cased key with an ARL_ prefix (ARL_BLOCK_SEVERITY, ARL_MODEL, …).
The keys most people touch:
| Key | Default | Purpose |
|---|---|---|
harness | claude-code | which reviewer CLI runs the review — claude-code or opencode |
model | the harness's own (opus for claude-code, openai/gpt-5.6-sol for opencode) | probed for reachability at arm time |
variant | unset | reasoning effort — --variant on OpenCode, --effort (low…max) on Claude Code |
block_severity | medium | blocks when actionable=yes AND severity >= this |
verify_cmd | unset | run by the hook, output attached to the review as evidence |
review_guide | unset | a Markdown file spliced into the reviewer's prompt as repo-specific guidance — see configuration.md |
ignore_globs | [] | paths whose sole change skips a review. A full bypass, not a relaxation |
final_review | false | run the final cumulative review at Stop |
ttl_hours | 24 | after this, every mutation is denied and each turn ends with a message — resume is usually the fix, not implement |
timeout_sec | 900 | per review run |
{
"model": "openai/gpt-5.6-sol",
"variant": "high",
"verify_cmd": "make test",
"ignore_globs": ["CHANGELOG.md", "docs/**"]
}
Every key, the full precedence rules, and what a review actually costs are in docs/configuration.md. State lives under $XDG_STATE_HOME/adversarial-review-loop/: no hook, and nothing Claude can invoke itself, ever writes inside the repository under review — the one exception is config <key> <value> --repo, an explicit user-only write.
Each question links to its full answer.
.md, run /adversarial-review-loop:implement plan.md./adversarial-review-loop:pause, then continue./cleared mid-phase — how do I pick it back up? Same session id (claude --resume, or /resume back to it): just continue. New session: /adversarial-review-loop:resume --allow-dirty. /adversarial-review-loop:status tells you which you're in.--until N on implement or resume; carry on later with a bare resume, which clears the target.review_guide — a Markdown file added to the reviewer's prompt. It cannot change the contract or what blocks./adversarial-review-loop:accept [reason] approves the current tree without another review, and records that it did.More — RECONCILE, NEEDS_HUMAN, rate limits, cost, running your tests as evidence, revising a plan mid-run — in docs/faq.md.
defer, or edit .adversarial-review-loop.json directly — disable-model-invocation blocks the config command, not ordinary edits to the file it writes. Policy written there can weaken the gate silently: ignore_globs: ["**"] skips the reviewer call on every commit, a raised block_severity stops findings from blocking, and final_review false removes the cumulative backstop. None of these run unreviewed code; all of them change what the gate does. Nothing here defends against a deliberately hostile agent, and this design does not pretend to — see security.md.PostToolUse, progress-aware counting and CLAUDE_CODE_STOP_HOOK_BLOCK_CAP mitigate it, but a run that repeatedly ends its turn without progress can still exhaust the cap — and exhaustion ends the turn.$ARGUMENTS is substituted into the skill body textually, with no shell escaping, so a plan path like x"; id; echo " used to run id. Every skill that takes an argument now hands it over inside a here-document with a quoted delimiter — the one shell construct whose body is never parsed — so the argument arrives verbatim and nothing inside it runs. The substitution itself is still not fixable here: it happens before any shell sees the body, so the containment remains that these skills are disable-model-invocation: true and only the person typing the slash command can supply the argument. An install also serves skill bodies from a version-pinned cache, so one installed before this fix keeps the old, breakable body until it is updated./clear, a crash, or a fresh claude leaves you in a session that is not bound to it. Nothing is disarmed and nothing is lost — the activation stays armed and enforcing, and the unbound session is simply denied every mutation in that worktree until /adversarial-review-loop:resume binds it, carrying the baseline and every approval across. Returning to the same session (claude --resume) needs no command at all. The cost is that a second session resuming the worktree retires the first for good — see edge-cases.md.PreToolUse hook itself adds ~111 ms to every tool call — the measurements are in architecture.md.| Page | For | Covers |
|---|---|---|
| how-it-works.md | anyone, no engineering background needed | What problem this solves and why, in plain language |
| faq.md | anyone using it day to day | The questions that come up first, answered short |
| configuration.md | anyone running the loop day to day | Every setting, precedence, cost, examples |
| architecture.md | engineers working on or integrating with the plugin | Components, data flow, the state machine, what blocks, on-disk layout |
| edge-cases.md | anyone debugging unexpected behaviour | What happens when things go sideways, and why |
| security.md | anyone assessing whether this is safe to trust | The threat model, what is and is not enforced, and why |
make test # 3900+ tests against scratch repos, plus the shim selftest; no model is called
make test-unit # the pytest half only
make test-filter FILTER=watchdog # one shim selftest section
make dry-run # print the exact reviewer argv and prompt, without invoking it
make check # pre-commit (shellcheck, yamllint, markdownlint, …)
tests/unit/ drives the hook entrypoints with synthetic payloads against scratch git repos — through the same bootstrap production uses — and replaces the reviewer with tests/fixtures/fake-reviewer.sh (ARL_REVIEWER_CMD), so loop logic costs nothing to iterate on. It covers the snapshot layer, the command-shape table, every arm-failure mode, the fail-closed guards, commit divergence and reconcile, the findings cap, the Stop accounting and the TTL.
tests/selftest.sh is a separate, much smaller suite in bash, for the one layer pytest structurally cannot reach because it runs inside the Python being launched: the interpreter probe, the shim contract, the watchdog layers, socket stdin, and the hot path's process budget.
Running the tests needs jq, and comparing the chunker against the real GNU split needs GNU coreutils (gsplit); both are development-only — the gate itself uses neither, and those comparisons skip cleanly where coreutils is absent, as on a stock Mac.
With the plugin enabled, its hooks run in every session, armed or not, and they deny any Claude Bash command that names a user-only subcommand (arl.sh arm, finish, deactivate, resume, config, accept, pause). When developing this repository, set "enabledPlugins": {"adversarial-review-loop@adversarial-review-loop": false} in its .claude/settings.local.json, and load the working tree with claude --plugin-dir . only when you mean to use the loop.
AGENTS.md is the contract any change to this project has to honour — the five non-negotiable rules, the invariants, and the hazards that silently reopen a closed hole if reverted. Before the first real run, work through tests/STEP0.md: the harness assumptions that only a live Claude Code session can settle.
GPL-3.0-or-later. See LICENSE.
The gate parses commands with bashlex, which is GPLv3 and is vendored under scripts/arl/_vendor/ so the plugin works straight from a checkout with no install step. See that directory's README for the version, the upstream commit, and the one change made to it.
hooks/register.js 230 lines1// This file is part of adversarial-review-loop.
2//
3// Copyright (c) 2026 Roberto Leinardi
4//
5// adversarial-review-loop is free software: you can redistribute it and/or modify
6// it under the terms of the GNU General Public License as published by
7// the Free Software Foundation, either version 3 of the License, or
8// (at your option) any later version.
9//
10// adversarial-review-loop is distributed in the hope that it will be useful,
11// but WITHOUT ANY WARRANTY; without even the implied warranty of
12// MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
13// GNU General Public License for more details.
14//
15// You should have received a copy of the GNU General Public License
16// along with adversarial-review-loop. If not, see <http://www.gnu.org/licenses/>.
17
18// The review loop's display: a band above the prompt with this session's activation, and a
19// transcript line when an alert appears. It decides nothing. The gate is the command hooks
20// in hooks.json; a module that fails is skipped by the engine, which is why nothing here may
21// ever matter to a verdict. See docs/design/mod.md before changing it.
22//
23// Everything it shows comes from `arl.sh status --json --session <id>`, which reads only and
24// computes every alert. This module never re-derives one, and draws only enums and integers.
25//
26// Plain JavaScript that imports nothing; the JSDoc below is what `tsc --checkJs` reads, against
27// the API types the engine lays in .claude-plugin/types/.
28
29/** @typedef {import('claude-code').EngineInterface} Engine */
30
31/**
32 * One `status --json` answer, as this module has checked it.
33 * @typedef {object} Answer
34 * @property {'pending' | 'bound' | 'unbound' | 'unarmed' | 'unknown'} binding
35 * @property {string[]} alerts
36 * @property {string} [status]
37 * @property {string} [cause]
38 * @property {number} [phase]
39 * @property {number} [phase_count]
40 * @property {number} [rounds_this_phase]
41 * @property {number} [failures]
42 * @property {number} [max_failures]
43 */
44
45const BINDINGS = new Set(['bound', 'unbound', 'unarmed', 'unknown'])
46const STATUSES = new Set(['ARMED', 'ACTIVE', 'RECONCILE', 'NEEDS_HUMAN', 'STALE', 'ARM_FAILED', 'COMPLETE', 'DISARMED', 'RESUMED'])
47const CAUSES = new Set(['document_missing', 'document_unreadable', 'document_malformed', 'version_conflict', 'repo_unresolvable'])
48const BOUND_COUNTS = ['phase', 'phase_count', 'rounds_this_phase', 'failures', 'max_failures']
49
50// What each alert says in the transcript. The keys are the only alerts `status` emits.
51/** @type {Record<string, string>} */
52const ALERTS = {
53 needs_human: 'the loop is NEEDS_HUMAN and only you can clear it: /adversarial-review-loop:status says why.',
54 unreviewed_at_exit: 'the mode ended with work committed that no review approved: /adversarial-review-loop:status names the commit.',
55 ended_record_malformed: "this activation's end record was not written by the gate: state.json was edited. Look at the history yourself.",
56 ended_unverifiable: 'the repository could not be read when the mode ended, so nothing says whether that work was reviewed.',
57 ended_unborn: 'HEAD no longer existed when the mode ended: the history this activation gated is gone.',
58 state_unknown: "the loop's state could not be read. That does NOT mean it is off: /adversarial-review-loop:status says why.",
59}
60
61// The last answer, and the alerts already reported. Module state starts over on a reload,
62// which only means the next refresh reports the current alerts once more.
63/** @type {Answer} */
64let last = { binding: 'pending', alerts: [] }
65/** @type {Set<string>} */
66let reported = new Set()
67
68/**
69 * @param {string} cause
70 * @returns {Answer}
71 */
72function unknown(cause) {
73 return { binding: 'unknown', cause, alerts: ['state_unknown'] }
74}
75
76/** @param {unknown} value */
77function isCount(value) {
78 return typeof value === 'number' && Number.isInteger(value) && value >= 0
79}
80
81// `status --json`'s answer, checked field by field. Anything this module does not recognise
82// is an unknown state, never an unarmed one.
83/**
84 * @param {string} stdout
85 * @returns {Answer}
86 */
87function parse(stdout) {
88 /** @type {any} */
89 let answer
90 try {
91 answer = JSON.parse(stdout)
92 } catch {
93 return unknown('unparseable')
94 }
95 if (answer === null || typeof answer !== 'object' || !BINDINGS.has(answer.binding)) {
96 return unknown('unparseable')
97 }
98 // An alert this module has no line for is not dropped: a newer `status` raising one this
99 // module predates must not leave the band drawing a quiet ACTIVE.
100 if (!Array.isArray(answer.alerts) || !answer.alerts.every((/** @type {unknown} */ alert) => typeof alert === 'string' && Object.hasOwn(ALERTS, alert))) {
101 return unknown('unparseable')
102 }
103 /** @type {string[]} */
104 const alerts = answer.alerts
105 if (answer.binding === 'unknown') {
106 return unknown(CAUSES.has(answer.cause) ? answer.cause : 'unparseable')
107 }
108 if (answer.binding === 'unbound') {
109 return STATUSES.has(answer.status) ? { binding: 'unbound', status: answer.status, alerts } : unknown('unparseable')
110 }
111 if (answer.binding === 'unarmed') {
112 return { binding: 'unarmed', alerts: [] }
113 }
114 if (!STATUSES.has(answer.status) || !BOUND_COUNTS.every(key => isCount(answer[key]))) {
115 return unknown('unparseable')
116 }
117 return {
118 binding: 'bound',
119 status: answer.status,
120 alerts,
121 phase: answer.phase,
122 phase_count: answer.phase_count,
123 rounds_this_phase: answer.rounds_this_phase,
124 failures: answer.failures,
125 max_failures: answer.max_failures,
126 }
127}
128
129/**
130 * @param {Engine} $
131 * @returns {Promise<Answer>}
132 */
133async function ask($) {
134 try {
135 const session = await $.session.id()
136 const cwd = await $.session.cwd()
137 const ran = await $.process.run([$.plugin.root + '/scripts/arl.sh', 'status', '--json', '--session', session], { cwd, timeoutMs: 15000 })
138 return ran.exitCode === 0 ? parse(ran.stdout) : unknown('status_failed')
139 } catch {
140 return unknown('status_failed')
141 }
142}
143
144/** @param {Engine} $ */
145async function refresh($) {
146 const answer = await ask($)
147 const current = new Set(answer.alerts)
148 for (const alert of current) {
149 if (!reported.has(alert)) {
150 $.ui.log(ALERTS[alert] ?? alert)
151 }
152 }
153 reported = current
154 last = answer
155 $.ui.invalidate('ui.render')
156}
157
158/**
159 * The band's one line, or null to draw nothing.
160 * @param {Answer} answer
161 * @returns {{ text: string, color: string | undefined } | null}
162 */
163function bandLine(answer) {
164 if (answer.binding === 'bound') {
165 const alert = answer.alerts.length > 0 ? ' · alert: ' + answer.alerts.join(', ') : ''
166 // Past the last phase, `phase` is total+1: what is left is the turn end, not a phase.
167 const count = answer.phase_count ?? 0
168 const phase = count > 0 && (answer.phase ?? 0) > count ? `all ${count} phases committed` : `phase ${answer.phase}/${count}`
169 return {
170 text: `ARL ${answer.status} · ${phase} · round ${answer.rounds_this_phase} · failures ${answer.failures}/${answer.max_failures}${alert}`,
171 color: alert ? 'red' : undefined,
172 }
173 }
174 if (answer.binding === 'unbound') {
175 return { text: `ARL: this worktree is armed by another session (${answer.status}); /adversarial-review-loop:resume binds this one`, color: 'yellow' }
176 }
177 if (answer.binding === 'unknown') {
178 return { text: `ARL: state unknown (${answer.cause})`, color: 'yellow' }
179 }
180 return null
181}
182
183// `Register` itself is not used as the type: it returns `unknown`, which a block body has to
184// return explicitly. The parameter is what carries the types into every hook.
185/** @param {import('claude-code').On} on */
186export const register = on => {
187 on('session.start', async ($, e, next) => {
188 await refresh($)
189 return next(e)
190 })
191
192 // /clear starts a new session id and fires no session.start (tests/STEP0.md item 22), so
193 // every refresh reads the id afresh rather than keeping one from the start.
194 on('turn.start', async ($, e, next) => {
195 await refresh($)
196 return next(e)
197 })
198
199 on('turn.complete', async ($, e, next) => {
200 const result = await next(e)
201 await refresh($)
202 return result
203 })
204
205 // A commit lands inside a turn, so a Bash call is the one place the phase moves mid-turn.
206 // Nothing is added to an unarmed session's tool calls.
207 on('tool.call', { tool: 'Bash' }, async ($, e, next) => {
208 const result = await next(e)
209 if (last.binding !== 'unarmed') {
210 await refresh($)
211 }
212 return result
213 })
214
215 // The band is shared: whatever the engine and other plugins draw there stays, with this
216 // line above it. A survey holds the band alone, so the line yields to it.
217 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
218 const below = await next(e)
219 const line = bandLine(last)
220 if (line === null || e.props.hasSurvey) {
221 return below
222 }
223 const { Box, Text } = $.ui.resolve(e)
224 // JSX would type this as a RenderElement; a direct `h` call returns the looser RenderNode.
225 return /** @type {import('claude-code').RenderElement} */ (
226 h(Box, { flexDirection: 'column' }, h(Text, { color: line.color, dimColor: line.color === undefined, wrap: 'truncate-end' }, line.text), below)
227 )
228 })
229}
230